Most Internet business people are grass-roots entrepreneurship, this time there is no powerful servers, and no money to buy a very expensive massive database. Under such severe conditions, batch after batch of entrepreneurs to be successful from the start, this and the current open-source technology, will 2015 Nike Free 5.0 have a massive data architecture inseparable relationship. For example, we use mysql, nginx and other open source software, through architecture and low-cost servers can also build ten million users visited system. Sina microblogging, Taobao, Tencent and other large Internet companies use a lot of free open source system to build their platform. So, what it does not matter, as long as they adopt a reasonable solution under reasonable circumstances. Then how to build a good system structure? This topic is too large, this is mainly talk about data distribution approach. Such as our database server can store 200 data, suddenly engage in an activity estimated to reach 600 data. It can be used in two ways: horizontal expansion or vertical extensions. Longitudinal extension is to upgrade the server hardware resources. But with the performance configuration of the machine, the higher the price, the price for the average small company can not afford. Scale is the use of multiple low-cost machines to provide services. Such a machine can only handle Nike Air Jordan 5 Women 200 data, three machine can handle 600 data, and if the traffic increases in the future may also be configured to increase rapidly. In most cases, choose the lateral extensions. As shown below: Now there is a problem, how this 600 data is routed to the corresponding machine. If you need to consider a balanced distribution, we assume 600 data are uniform increment id data, from 1 to 600, divided into three reactors can be used id mod 3 mode. In fact, in Nike Basketball the real world may not be the id string. Need to be converted to a string hashcode then modulo. Now, it seems is not the solution to our problem, all the data is very good and did not meet the load distribution system. But if we need to store data, so there is no need to read easily. Increase business how to do, we know that in accordance with the above scale need to increase a server. However, this is due to increased server brought some problems. Look at the Mens Nike Free 3.0 V2 Shoes Black Blue following example, a total of nine numbers, you need to put two machines (1, 2). Each machine is stored as follows: the 1st machine storage 1,3,5,7,9 No. 2 machine storage 2,4,6,8. If the extension is a machine 3 how data should a major migration, the 1st machine store number 1,4,7, 2,5,8 No. 2 machine storage, 3 machine store 3,6,9. Figure: As can be seen from the figure 1, 5, 9 migrate out of the machine, 4,6 migration 2 good machine out, according to the new order and Mens Nike Free 3.0 V2 Shoes Grey Orange re-allocate it again. A small amount of data, then redistribute over the cost is not large, but if we have a million T-level data on the operating costs are quite high, ranging from a few hours to more than a few days. And migrated when the original database machine load is relatively high, that we have doubts, this level is not scalable architecture is not reasonable manner? ---------- ------------- Consistency hash gorgeous dividing line in this application is made to the background, it is now widely used in distributed cache, like memcached ʱ?? The following outlines the basic principles under the hash of consistency. The earliest version http://dl.acm.org/citation.cfm?id=258660. Domestic There are many online articles are written better. Such as: http://blog.csdn.net/x15594/article/details/6270242 following simple example to illustrate the consistency of hash. Preparation: 1,2,3 three machines have yet to be allocated nine numbers 1,2,3,4,5,6,7,8,9 consistency hash algorithm architecture Step one, constructed out of 2 ^ 32 months virtual node out, because there are 01 computers in the world, using the power of 2, when the data is divided easily balanced distribution. Another 32 th power of 2 is 4.2 billion, we even have a super large number of servers can not be more than 4.2 billion Taiwan Bar, Air Jordan Outlet expansion and balance are guaranteed. Second, the three machines were taken Mens Nike Free 3.0 V2 Shoes Black Gold hashcode IP for calculations (where you can also take the hostname, as long as the Nike Blazers only difference between Lebron Slide 2 Elite the individual machines on it), and then mapped to the 32 th power of 2 up. For example, the 1st machine Nike Air Max counted out hashcode and mod (2 ^ 32) to 123 (this is fictional), the 2nd machine counted out the number of machines is 2300420,3 calculated as 90,203,920. Thus the three machines are mapped to the virtual node on the ring structure of the 4.2 billion. Third, the data (1-9) also used the same method of calculating hashcode and 4.2 billion modulo its configuration to the ring node. Assuming that several nodes calculated value of 1: 10,2: 23564,3: 57,4: 6984,5: 5689632,6: 86546845,7: 122,8: 3300689,9: 135,468. It can be Nike Air Max 2011 Men seen less than 123 1,3,7, 4, 9 less than 2,300,420 is greater than 123, more than 2,300,420 5,6,8 less than 90,203,920. Mapping from the data to find the location clockwise, to save data to the first node Cache found. If you still can not find more than 2 ^ 32 Cache node, it will be 378037 010 Original?CBlack-True Red-White Air Jordan 11 Retro Nike Kids Sneakers Factory Outlet Nike Air Max 2011 Men saved to the first Cache node. 1,3,7 is assigned to the 1st machine, 4, 9 will be assigned to the 2nd machine, 5,6,8 will be assigned to the 3rd machine. This time it might ask, I still do not see any good hash consistency than conventional modulo also adds complexity. Now do some key process immediately, for example, we add a machine. According to the original all the data we need to be reassigned to four machines. Consistency hash how to do it? Now add to the mix the 4th machine, he came out of the hash value modulo operator is 12,302,012. More than 2,300,420 12,302,012 less than 5,8, 6 more than 12,302,012 less than 90,203,920. Such adjustments just 5,8 deleted from the 3rd machine, the 4th machine was added 5,6. Similarly, to remove the machine how to do it, assuming the 2nd machine hang affected only the 2nd machine data is migrated to the node from it, the picture shows the 4th machine. Everyone should understand the basic principles of consistency hash of it. However, this method still has defects, such as fewer nodes in machines, when large volumes of data, data distribution may not be very balanced, it will lead to one of the servers where the data is a lot more than other machines. To solve this problem, the need to introduce mechanisms for virtual server node. As we have a total of only three machines, 1,2,3. But actually you can not have so many machines how to solve it? Each of these virtual machines out of three machines, which is 1a 1b 1c 2a 2b 2c 3a 3b 3c, so it becomes a 9 machine. Actual 1a 1b 1c 1 or correspondence. But the actual distribution to the ring node becomes a 9 machine. Data also can be more dispersed distribution point. Figure: write so much consistency hash, what is this, and distributed search slightest relationship? We now use solr4 build a distributed search, test data submitted 20 based on distributed platforms solrcloud actually require tens of seconds, so we abandoned solrcloud. Hack solr using their own platform, do not zookeeper distributed consistency management platform to Mens Nike Free 3.0 V2 Shoes Grey Orange manage data distribution mechanism. Since the need to manage the distribution of data, we need to consider the creation of the index, the index update. So that we will spend the consistency hash. Overall architecture as shown below: Establish the need to maintain and update the machine's position, according to the data to find the corresponding key distribution and updating data. Here need to consider is how efficient and reliable data to establish, update to the index in. Backup server to prevent the establishment of a server hang up, you can quickly recover according to the backup server. Read mainly to write separate server use, prevent write index affect query data. Server Status cluster management server to manage the entire cluster, and alarms. With the increase in traffic throughout the cluster can also be classified according to the type of data, such as user, microblogging. Each type of architecture built in accordance with the figure, to meet the general performance of distributed search. Search for solr and distributed subject subsequent talk later. Further reading: java hashmap with the increase of the amount of data will appear map adjustment Mens Nike Free 3.0 V2 Shoes Grey Green problems, when necessary, to initialize a large enough size to prevent insufficient capacity of existing data re-hash calculations. Vaccines: Optimizing Java HashMap infinite loop http://coolshell.cn/articles/9606.html consistent hashing algorithm - on how to ensure the new node is added to the ring, the hit rate is not affected (formerly pat colleague scott) http://scottina.iteye.com/blog/650380 language: http: //weblogs.java.net/blog/2007/11/27/consistent-hashing java version of the example http: //blog.csdn .net / mayongzhan / archive / 2009/06/25 / 4298834.aspx PHP version examples Examples http://www.codeproject.com/KB/recipes/lib-conhash.aspx C language versionconsistency hash and solr ten million data distributed search engine application