A method and device for processing data sharding at the ten-million level
By decoupling the computing function of the shard machine from the business processing function and adopting the method of the shard machine actively preempting the shard data, the problem of strong coupling between shard processing and business in the existing technology is solved, and data processing continuity and load balancing are achieved when the business node is down, improving the overall processing efficiency.
Patent Information
- Application Number
- CN202111019977.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-01
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-09-01
AI Technical Summary
When processing massive data, the shard processing is strongly coupled with the service, resulting in the processing efficiency dependent on the expansion service node. Data is lost when a certain node goes down, resource allocation is unbalanced, and overall processing capacity cannot be effectively improved.
The sharding machine's computing function is decoupled from the business processing function, and the sharding machine actively performs the sharding data preemption task, applies sharding data in real time according to its own performance, and processes it in parallel through multi-threads. The last sharding machine performs post-processing, and learns from the idea of the "ticket grab" model to achieve decoupling of sharding method and business.
Without affecting the service nodes of the service, the processing efficiency is greatly improved by expanding the capacity shard machine cluster, and the sharding method is decoupled from the service to ensure that the data can still be processed normally when some business nodes are down, load balancing, and overall business processing efficiency is improved.
Smart Images

Figure CN113722099B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sharding algorithms in big data processing, and particularly to a method and device for processing sharding of tens of millions of data.
Background Art
[0002] Most early software systems served the internal business needs of the company, and most of the system users were internal personnel of the company operating. The traditional software architecture could basically meet the daily needs. However, with the continuous expansion of the company's scale, the development of software has gradually shifted from unified integration to the microservices direction, and the deployment method has also changed from a single node to a multi-node cluster mode. This not only brings system complexity but also generates an increasing amount of data over time.
[0003] With the rapid development of the Internet in recent years, after realizing information interconnection, the functional requirements and processing capabilities of software systems are getting higher and higher. Especially in many Internet companies, from traditional enterprise services to personal services, the number of customers has basically increased exponentially, and at the same time, the maintenance of customer data will also increase day by day. Since it is customer-oriented service, once the system problem causes data errors or poor service experience, the affected user group range is large, which may even lead to a large number of customer complaints and even a large loss of customer groups, bringing immeasurable reputation and economic losses to the enterprise.
[0004] Due to the continuous increase in customer traffic, the data volume of the corresponding service system table is easily reached tens of millions or even hundreds of millions. When processing massive data (such as pushing service information to 10 million users), the current industry processing method usually adopts a sharding scheduling mode. The scheduling job is triggered by the Job mechanism, and requests are sent to the servers of all business Service nodes respectively. Different sharding parameters are assigned to different servers when sending requests. When each business Service node server gets the sharding parameters, it queries the sharding domain data it needs to process and then processes it on the local node server. The sharding processing in this way is strongly coupled with the business, and the processing efficiency can only be improved by expanding the number of business Service nodes. Once a business Service node fails, all the data processed by this node will be lost.
[0005] In summary, the existing technology mainly has the following three disadvantages:
[0006] 1) For the processing of massive data, the industry currently schedules and processes through Job services. The business Service nodes are deployed with a distributed cluster. After all business Service nodes receive the requests sent by the Job (with sharding parameters), the data range processed by each business Service node has been determined. Then each business Service node processes its own data independently. Once a certain business Service node crashes during the processing, all the data processed by this business Service node will be lost.
[0007] 2) After each business Service node receives the request sent by the Job, the sharded domain data obtained is basically average. However, there may be performance differences between the servers of each business Service node. The services of the business Service nodes with high performance will finish processing the data of their own nodes in advance, while those with low performance will take a longer time to finish processing the data of their own nodes. This will also lead to the problem of uneven distribution of resource processing.
[0008] 3) Since the sharding logic is in the code of each business Service node, if we want to improve the processing efficiency, we can only rely on horizontally expanding the business Service nodes to improve the overall processing capacity. This may cause the problem that the resource loads used by some low-frequency but important services (high load) and daily services (general load) cannot be balanced.
[0009] In view of this, how to overcome the defects of the existing technology and solve the above technical problems is a difficult problem to be solved in this technical field.
Summary of the Invention
[0010] In response to the above defects or improvement requirements of the existing technology, the present invention provides a method and device for processing shards of tens of millions of data, which can greatly improve the processing efficiency by expanding the shard computer cluster without affecting the business Service nodes. The sharding method is decoupled from the business, and the data can still be processed normally when some business Service nodes crash.
[0011] The embodiments of the present invention adopt the following technical solutions:
[0012] In the first aspect, the present invention provides a method for processing shards of tens of millions of data, including:
[0013] Decouple the computing function of the shard computer and the business processing function, and let the shard computer actively perform the task of preempting sharded data;
[0014] The shard computer applies for sharded data in real time according to its own performance;
[0015] Divide the data obtained by each shard computer application into multiple threads for parallel processing;
[0016] The last processed slicing machine performs post-processing work.
[0017] Furthermore, decoupling the computing function and business processing function of the slicing machine, and the slicing machine actively performing the task of preempting shard data specifically includes:
[0018] A random slicing machine receives a business processing request;
[0019] After this slicing machine finishes the pre-work, it notifies Zookeeper to broadcast the request;
[0020] After all slicing machines monitor the broadcast, they simultaneously start to actively preempt shard data.
[0021] Furthermore, during the process of preempting shard data, if a slicing machine crashes and restarts, it will perform automatic detection and continue to join the task cluster.
[0022] Furthermore, the pre-work specifically includes whitelist testing, setting the start and end positions of the data to be processed this time, task locking, and recording relevant task information.
[0023] Furthermore, the slicing machine applying for shard data in real time according to its own performance specifically includes:
[0024] Number all slicing machines;
[0025] Determine the processing efficiency of each slicing machine itself;
[0026] The slicing machine applies for corresponding shard data from redis according to its own processing efficiency.
[0027] Furthermore, the processing efficiency of the slicing machine is expressed as:
[0028] Shard data (x) = n * cpu, where x represents the size of the shard domain data, n represents the single-cpu batch processing size, and cpu represents the number of cpu cores;
[0029] The slicing machine applying for corresponding shard data from redis according to its own processing efficiency specifically includes: the slicing machine actively applies for incr(x) from redis, where incr(x) represents the size of the shard data processed by each slicing machine in a single batch.
[0030] Furthermore, dividing the data obtained by each slicing machine application into multiple threads for parallel processing specifically includes:
[0031] After the slicing machine obtains a single batch of data domain, it marks this batch of data in redis, and then distributes it to the business Service nodes in multiple threads in sequence;
[0032] After this batch process is completed, the mark is cleared and the next batch of sharding domains starts to be preempted.
[0033] Furthermore, the number of threads divided by each sharding machine is calculated based on the CPU. The default calculation formula is: CPU-intensive threads = number of CPU cores + 1.
[0034] Furthermore, the post-processing work performed by the last sharding machine that has completed processing specifically includes:
[0035] By comparing the number of zk temporary nodes with the number of redis processed, it is determined whether it is the last sharding machine that has completed processing. If so, it is called the post-processor;
[0036] The post-processor detects the integrity of this processing, conducts statistics, and gives notifications;
[0037] When the entire task ends, the post-processor releases the task lock.
[0038] On the other hand, the present invention provides a ten-million-level data sharding processing device, specifically including: at least one processor and a memory. The at least one processor and the memory are connected through a data bus. The memory stores instructions that can be executed by the at least one processor. After the instructions are executed by the processor, they are used to complete the ten-million-level data sharding processing method in the first aspect.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows: Without affecting the business Service nodes, the processing efficiency can be significantly improved by expanding the sharding machine cluster. The sharding method is decoupled from the business, and data can still be processed normally when some business nodes are down. Drawing on the idea of the "ticket grabbing" mode, different processing machines can process different amounts of data according to their own performance. That is, the higher the performance, the faster the processing, and the more sharding domains will be "grabbed". To a certain extent, the load of the sharding machines is balanced, and the overall business processing efficiency is improved.
Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0041] Figure 1 It is a flowchart of a ten-million-level data sharding processing method provided in Embodiment 1 of the present invention;
[0042] Figure 2 It is a specific flowchart of step 100 provided in Embodiment 1 of the present invention;
[0043] Figure 3 It is the specific flowchart of step 200 provided in Embodiment 1 of the present invention;
[0044] Figure 4 It is the schematic diagram of the data sharding algorithm provided in Embodiment 1 of the present invention;
[0045] Figure 5 It is the specific flowchart of step 300 provided in Embodiment 1 of the present invention;
[0046] Figure 6 It is the schematic diagram of multi-threaded parallel processing provided in Embodiment 1 of the present invention;
[0047] Figure 7 It is the specific flowchart of step 400 provided in Embodiment 1 of the present invention;
[0048] Figure 8 It is the architecture diagram of a ten-million-level data sharding processing system provided in Embodiment 2 of the present invention;
[0049] Figure 9 It is the block diagram of a ten-million-level data sharding processing system module provided in Embodiment 2 of the present invention;
[0050] Figure 10 It is the schematic diagram of the structure of a ten-million-level data sharding processing device provided in Embodiment 4 of the present invention.
Specific Embodiments
[0051] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] The present invention is an architecture of a specific functional system. Therefore, in specific embodiments, the functional logic relationships of each structural module are mainly described, and the specific software and hardware implementation manners are not limited.
[0053] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. The present invention will be described in detail below with reference to the drawings and embodiments.
[0054] Embodiment 1:
[0055] As Figure 1 shown, Embodiment 1 of the present invention provides a ten-million-level data sharding processing method, and the specific steps are as follows.
[0056] Step 100: Decouple the calculation function and the service processing function of the sharding machine, and let the sharding machine actively perform the task of preempting sharded data.
[0057] Step 200: The slicing machine applies for sliced data in real time according to its own performance.
[0058] Step 300: The data obtained by each slicing machine's application is divided into multiple threads for parallel processing.
[0059] Step 400: The last slicing machine that finishes processing performs post-processing work.
[0060] Through the above steps, the embodiment of the present invention decouples the slicing method and the service, so that the processing efficiency can be greatly improved by expanding the slicing machine, and the service is unaware. In addition, the embodiment of the present invention is different from the traditional method of issuing tasks to the slicing machine. Instead, the slicing machine actively applies for sliced data for processing according to its own performance. This embodiment of the present invention draws on the idea of the "ticket grabbing" mode, enabling different processors to process different amounts of data according to their own performance. That is, the higher the performance, the faster the processing, and the more sliced domains will be "grabbed". To a certain extent, the load of the slicing machine is balanced, and the overall service processing efficiency is improved.
[0061] Specifically, as Figure 2 shown, in this preferred embodiment, Step 100 (decoupling the computing function and the service processing function of the slicing machine, and having the slicing machine actively perform the sliced data preemption task) specifically includes:
[0062] Step 101: A random slicing machine receives a service processing request. In this step, the processing request of the service data is sent to a random slicing machine through the Job cluster or the browser.
[0063] Step 102: After this slicing machine finishes the pre-work, it notifies Zookeeper to broadcast the request. In this step, the slicing machine that receives the processing request first finishes the pre-work, and then notifies Zookeeper (Zookeeper is a distributed, open-source distributed application coordination service) of the processing request. After receiving the notification, Zookeeper broadcasts the start execution command.
[0064] In this preferred embodiment, the pre-work processed by the slicing machine includes:
[0065] (1) Whitelist test (verify the business correctness with the whitelist before batch processing and then start sliced batch processing).
[0066] (2) Set the start and end positions of the data to be processed this time.
[0067] (3) Task locking (it is not allowed to repeat execution during the processing of the same batch of tasks).
[0068] (4) Record relevant task information (such as start execution time, data volume, etc.).
[0069] The slicing machine for this pre - processing work is also called the pre - processor.
[0070] Step 103: After all slicing machines monitor the broadcast, they simultaneously start to actively preempt the sliced data. In this step, all slicing machines have subscribed to the broadcast notifications of Zookeeper in advance. After the slicing machines monitor the broadcast sent by Zookeeper, they start to actively preempt the sliced data. This operation does not require Zookeeper to allocate slicing parameters, but the slicing machines independently apply for pre - emption of the sliced data.
[0071] Combining the above steps 101 - 103, the embodiment of the present invention realizes the functions of one - stop wake - up and service decoupling of the slicing machine, and has the following advantages: By decoupling the slicing calculation function and the service processing function of the slicing machine, when a certain service Service node fails, it does not affect data processing; The data range processed by each slicing machine each time is calculated in real - time, and even if a slicing machine fails, it does not affect the continuity of the entire data processing; Only by expanding the slicing machine cluster can the high - speed distribution ability be achieved.
[0072] The above steps 101 - 103 (step 100) are the core deployment architecture of the embodiment of the present invention, which can make the slicing machine independent of the service Service node, not invade the service, and is convenient for horizontal expansion. In addition, during the process of preempting the sliced data, if a slicing machine fails, after restarting, it will perform automatic detection and continue to join the task cluster. The automatic detection process is specifically embodied as: After restarting after a failure, it automatically detects whether there is a task being processed currently. If there is, it automatically joins the task and continues slicing processing.
[0073] As Figure 3 shown, in this preferred embodiment, step 200 (the slicing machine applies for sliced data in real - time according to its own performance) specifically includes:
[0074] Step 201: Number all slicing machines. In this step, slicing machine 1, slicing machine 2... slicing machine n can be called S1, S2... Sn.
[0075] Step 202: Determine the processing efficiency of each slicing machine itself. In this step, the processing of the slicing machine is mainly used for calculation and distribution, without disk I / O operations, and belongs to CPU - intensive. The processing efficiency depends on the number of CPU cores. Therefore, the processing efficiency of the slicing machine can be expressed as: sliced data (x)=n * cpu. Where x represents the size of the sliced domain data, n represents the size of single - CPU batch processing, and cpu represents the number of CPU cores.
[0076] Step 203: The sharding machine applies to Redis for corresponding sharded data according to its own processing efficiency. In this step, the sharding machine actively applies to Redis (Redis is a key-value storage system) for incr(x), where incr(x) represents the size of the sharded data processed by each sharding machine in a single batch. If the single-batch processing volume of S1 is 4000 and that of S2 is 8000, then the data sharding is as Figure 4 shown.
[0077] Figure 4 There are two segments of S2 because the machine performance of sharding machine S2 is high and the processing speed is fast. After processing the data from 4001 to 12000, it continues to apply to Redis for processing the next batch of data (i.e., 12001 to 20000). Figure 4 In it, Sn represents one of the sharding machines. Suppose the number of sharding machine clusters is 3, then Sn represents one of S1, S2, and S3, and n is the representative of a certain sharding machine number; the failure handling marked above Sn means that the sharding machine crashes during the processing, and the batch processed by this sharding machine will be recorded as unfinished (or failure marked). If the sharding machine restarts after crashing, it will give priority to processing the failed batches; if the sharding machine remains crashed all the time, after other sharding machines finish processing the subsequent data, they will re-compensate and process the failed data.
[0078] Specifically, Figure 4 it can be expressed as: Suppose there are two sharding machines (the same for multiple machines) S1 and S2 in the current sharding machine cluster. S1 has lower performance and can process 4000 data items in a single batch, while S2 has higher performance and can process 8000 data items in a single batch. If the tasks start simultaneously and S1 grabs the first batch of 4000 data with the range of 1 to 4000, then S2 grabs 8000 data with the range of 4001 to 12000 at this time. When S2 finishes the first batch, if S1 has not finished the first batch yet, then S2 can continue to apply for the next batch of 8000 data, that is, 12001 to 20000. At this time, when S1 finishes and applies for 4000 data again, the data range is 20001 to 24000, and so on. If there is no crash or exception in the middle, the data processing ends when reaching the maximum value (end shown in the figure); if S2 crashes during the processing of the data from 12001 to 20000, then the data in this range is marked as failed data. At this time, S1 will skip this section of data and continue to execute the subsequent data (until end); when finishing processing the end data, a detection will be made on the batches of data marked as failed ( Figure 4 the dotted line connecting the failed handling dashed box represents the detection process). If there is failed data, compensation processing will be performed, and at this time, Sn represents S1 (a surviving sharding machine).
[0079] Through the above steps 201-203 (step 200), the embodiment of the present invention provides a method for grab-ticket-based dynamic sharding processing. This method is the core sharding method of the embodiment of the present invention. Combining the performance of the sharding machine, it adopts the "grab-ticket" sharding method. Each sharding is applied for in real time. Even if the machine crashes midway, it does not affect subsequent processing. In addition, in this preferred embodiment, when the sharding machine rejoins the task after crashing and restarting, it preferentially processes the data that was not successfully processed during the crash.
[0080] As Figure 5 shown, in this preferred embodiment, step 300 (dividing the data obtained by each sharding machine application into multiple threads for parallel processing) specifically includes:
[0081] Step 301: After the sharding machine obtains a single batch of data fields, it marks the data of this batch in Redis, and then distributes it to the business Service nodes in multiple threads in sequence. In this step, the sharding machine that grabs the sharded data needs to mark this data in Redis, and then distributes it to the business Service nodes in multiple threads in sequence. Specifically, multiple threads will be used inside a sharding machine to process a single batch of data. For example, if S1 processes 4000 pieces of data at a time, then it will be divided into multiple threads (such as 20) inside S1. Then these 20 threads will process these 4000 pieces of data simultaneously, and each thread gets 200 pieces of data, and then distributes them to the business Service nodes in sequence. It should be noted that each sharding machine corresponds to the entire business Service cluster. As for which business Service node the final request reaches for processing, it is determined by the business cluster load balancing. In addition, in practical applications, it can also be extended that a sharding machine cluster can correspond to multiple business service clusters, and task marking is used to select the business service cluster to which the request needs to be sent, but the tasks of the same batch can only correspond to the same business service cluster.
[0082] In this preferred embodiment, the number of threads divided by each slicing machine can be calculated according to the CPU. The specific calculation method is: CPU-intensive threads = number of CPU cores + 1. It should be noted that the number of CPUs in each slicing machine is not limited (it can be single or multiple), and it depends on the configuration of the slicing machine. The present invention only cares about the number of CPU cores that can process tasks in parallel at the same time. Additionally, it should be introduced that slicing calculation belongs to CPU-intensive (also known as compute-intensive). In order to improve efficiency and minimize the frequent switching of the CPU, the ideal state is that the number of threads processed simultaneously = the number of CPU cores. The reason for suggesting the number of CPU cores + 1 in this preferred embodiment is referred to in "Java Concurrency in Practice": when a compute-intensive thread happens to pause at a certain time due to a page fault or other reasons, there is just an "extra" thread, which can ensure that the CPU cycles will not be interrupted in this case. Therefore, the number of CPU cores + 1 is an empirical value and can be used as the default formula for slicing calculation in the present invention. Additionally, considering that the number of CPU cores is generally even, and after adding 1, it becomes odd, and the number allocated to each thread calculated cannot be divided evenly. Therefore, the relationship formula between the CPU and the threads can be set by the user, and the user can customize the calculation formula of the number of threads in the form of configuration when generating tasks.
[0083] Specifically, the schematic diagram of multi-thread parallel processing is as Figure 6 shown. Figure 6 What is shown is the multi-thread processing mechanism inside a slicing machine. How many threads are needed at one time to process is determined by the number of CUP cores. Suppose the number of CUP cores is 19, then the number of threads is 20 ( Figure 6 The 20 threads here only represent an example. If the total number of processed data and the number of threads cannot be divided evenly, then the last thread processes all the remaining data. For example, when processing 4000 at a time, for an 8-core slicing machine, the thread calculation is 9. Then the first 8 threads process 4000 / 9 = 444 (rounded down), and the last thread processes 4000 - 444 * 8 = 448).
[0084] Step 302: After this batch of processing is completed, clear the mark and start preempting the next batch of slicing domains. In this step, after each slicing machine processes and completes the data obtained by "grabbing tickets" through multi-threads, it clears the mark made in Redis beforehand, and then starts to preempt a batch of slicing domain data. After preemption, it continues to process.
[0085] Through the above steps 301 - 302 (step 300), the embodiment of the present invention can improve the processing speed inside each slicing machine, and marks will also be added to the data during the processing. If a machine crash occurs, after the machine is restarted after the crash, it can directly continue processing from the marked place.
[0086] As Figure 7As shown in the figure, in this preferred embodiment, step 400 (the last processed slicing machine performs post-processing work) specifically includes:
[0087] Step 401: Compare the number of zk temporary nodes with the number of processed items in redis to determine whether it is the last processed slicing machine. If so, it is called the post-processor. In this step, the zk temporary node is the Zookeeper temporary node. When the number of zk temporary nodes is the same as the number of processed items in redis, it is determined that this slicing machine is the post-processor.
[0088] Regarding the above judgment process, it should be added that each slicing machine will register with zk (Zookeeper) when starting (that is, form a temporary node), so the number of zk temporary nodes is equal to the number of normally operating slicing machines (normally equal to the total number of slicing machines in the slicing machine cluster). For example, there are 3 slicing machines in the slicing machine cluster. Assume 1: The total number of slicing machines is 3 and all slicing machines are operating normally, then the number of zk temporary nodes is equal to 3. Assume 2: The total number of slicing machines is 3, but one is always down and cannot process tasks, then the number of zk temporary nodes is equal to 2. In this embodiment, after each slicing machine processes to the maximum value of this task, a +1 operation is performed in redis (the initial number of processed items in redis is 0), and it is judged whether the number after +1 in redis is equal to the total number of slicing machines. If they are equal, it means it is the last processing machine. If the number of zk nodes is greater than the number of processed items in redis, it means there are still slicing machines processing at this time. For example, if there are 5 slicing machines in total, the number of nodes is equal to 5. If the number of processed items in redis is 4, it means there is still one slicing machine processing tasks. Only after all slicing machines perform +1 can it be equal to the total number of slicing machines. When judging non-last machines, it is less than the total number of slicing machines.
[0089] Step 402: The post-processor detects the integrity of this processing, counts the report data of this processing, and sends an email notification.
[0090] In this preferred embodiment, each single batch processing of the slicing machine will make a mark in redis and clear the mark after processing. When the post-processor detects, it can judge whether there is still unprocessed or failed-to-process data with marks. If not, it means all data has been processed; if the processing is incomplete, the post-processor continues to perform compensation processing. The recipients of the emails sent by the post-processor are the relevant business responsible persons for this task, which can be configured when generating the task.
[0091] Step 403: The entire task ends, and the post-processor releases the task lock. For example, for a task of sending service information to 10 million users, after the task job or the operation background triggers it, the pre-processor locks this task and waits until the task is processed. Then the post-processor releases the lock. The purpose of locking is to prevent the same task from being executed multiple times. Among them, the pre-processor is Figure 8 a random machine among those sent from the job or the browser to the sharding machine cluster in
[0092] Through the above steps 401 - 403 (step 400), the embodiments of the present invention handle some post-processing requirements. The main functions of the post-processor are to release the task lock, generate statistical reports, send email notifications, etc.
[0093] In summary, as can be seen from this preferred embodiment, the present invention decouples the sharding method from the service, so that the processing efficiency can be greatly improved by expanding the sharding machines, and the service is unaware. Moreover, even when some service nodes are down, the data can still be processed normally. When the service is down, the present invention only loses the data of atomic operations each time (that is, the data of single-step operations, and the data operations of the sharding machines in the embodiments of the present invention have atomicity), and a record compensation will be made on the sharding machines to ensure the integrity of data processing. The sharding machines of the present invention obtain sharding domain data through real-time calculation. Even if a sharding machine crashes during the sharding calculation, a large amount of data will not be lost, ensuring the continuity of data processing. Eventually, the failed data will be reprocessed in the form of compensation. According to the performance of different sharding machines, the data domains calculated by each sharding are different, and then they are sent to the service Service nodes. Through service load distribution balance, finally all sharding machine nodes and service Service nodes can process different amounts of data according to their own performance, reasonably distributing the load as a whole.
[0094] Embodiment 2:
[0095] Based on the ten-million-level data sharding processing method provided in Embodiment 1, this Embodiment 2 provides a corresponding ten-million-level data sharding processing system, as Figure 8 shown. The system architecture includes a job cluster or a browser as the sending end of mass data processing requests, a sharding machine cluster for sharding data processing, and a service Service cluster as the data processing end.
[0096] Specifically, the job cluster includes several jobs, the slicing machine cluster includes several slicing machines (slicing machine 1, slicing machine 2... slicing machine n) and Zookeeper, and the business Service cluster includes several services (service 1, service2... service n). During specific operations, the processing request for business data is sent to a random slicing machine through the Job cluster or the browser. For example Figure 8 it is sent to slicing machine 1 in Figure 8 . After slicing machine 1 finishes the preliminary work, it notifies Zookeeper to broadcast the processing request. After receiving the notification from slicing machine 1, Zookeeper broadcasts. Under the broadcast of Zookeeper, all slicing machines receive this broadcast message and then start to actively preempt the sliced data. During the specific preemption process, first determine the processing efficiency of each slicing machine itself, and then the slicing machine applies for the corresponding single batch of sliced data according to its own processing efficiency. After obtaining the single batch of data domain, the slicing machine divides the batch of data into multiple threads and distributes it to the business Service nodes for processing. After the processing is completed, the slicing machine starts to preempt the next batch of sliced domains.
[0097] Based on the processing process of the above system architecture, as Figure 9 shown, the ten-million-level data slicing processing system in this embodiment further includes the following processing modules: request processing module, real-time application module, multi-thread processing module, and post-processing module. Among them, the request processing module includes a request receiving module, a request notification module, a broadcast module, and a broadcast monitoring module; the real-time application module includes a slicing machine number module, a slicing machine efficiency calculation module, and a sliced data application module; the multi-thread processing module includes a multi-thread distribution module, a thread number calculation module, and a marking module; the post-processing module includes a post-processor judgment module, a detection module, and a task lock release module.
[0098] Specifically, among the above modules, the request receiving module is based on the sharding machine and is used to receive service processing requests; the request notification module is based on the sharding machine and is used to notify Zookeeper to broadcast requests; the broadcast module is based on Zookeeper and is used to perform broadcasts; the broadcast listening module is based on the sharding machine and is used to listen for broadcasts and notify the sharding machine to actively preempt shard data after a broadcast is detected; the sharding machine numbering module is based on the sharding machine and is used to number each sharding machine; the sharding machine efficiency calculation module is based on the sharding machine and is used to determine the processing efficiency of each sharding machine itself; the shard data application module is based on the sharding machine and is used to apply for corresponding shard data from redis according to its own processing efficiency; the multi-threaded distribution module is based on the sharding machine and is used to distribute processing data to the business Service nodes according to the planned multi-threads; the thread number calculation module is based on the sharding machine and is used to calculate the number of threads according to its own cpu; the marking module is based on the sharding machine and is used to mark the single batch of data obtained in redis, clear the mark after this batch is processed, and then notify the shard data application module to start preempting the next shard domain; the post-machine judgment module is based on the sharding machine and is used to judge whether this sharding machine is the last one to finish processing by comparing the number of zk temporary nodes with the number of redis processed. If so, it is the post-machine; the detection module is based on the sharding machine and is used for the post-machine to detect the integrity of this processing, count the report data of this processing, and send an email notification; the task lock release module is based on the sharding machine and is used for the post-machine to release the task lock at the end of the entire task.
[0099] The specific functions and interaction processes of the above modules correspond to the specific processes of each step in Embodiment 1 and will not be elaborated here.
[0100] Embodiment 3:
[0101] Based on the ten-million-level data sharding processing method and system provided in the above Embodiment 1 to Embodiment 2, the present invention also provides a ten-million-level data sharding processing device that can be used to implement the above method and system. As Figure 10 shown, it is a schematic diagram of the device architecture of an embodiment of the present invention. The ten-million-level data sharding processing device of this embodiment includes one or more processors 21 and a memory 22. Among them, Figure 10 One processor 21 is taken as an example in
[0102] The processor 21 and the memory 22 can be connected through a bus or other means. Figure 10 Taking the connection through the bus as an example in
[0103] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the ten-million-level data sharding processing methods and systems in Embodiments 1 to 2. The processor 21 executes various functional applications and data processing of the ten-million-level data sharding processing device by running the non-volatile software programs, instructions, and modules stored in the memory 22, that is, implements the ten-million-level data sharding processing methods and systems in Embodiments 1 to 2.
[0104] The memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 optionally includes a memory remotely set relative to the processor 21, and these remote memories can be connected to the processor 21 through a network. Examples of the above networks include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0105] The program instructions / modules are stored in the memory 22 and, when executed by one or more processors 21, execute the ten-million-level data sharding processing methods and systems in the above Embodiments 1 to 2. For example, execute the Figure 1 and Figure 9 various steps and modules shown.
[0106] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: Read Only Memory (ROM), Random Access Memory (RAM), magnetic disk, optical disc, etc.
[0107] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for processing data sharding at the ten-million level, characterized in that, Including: Decouple the calculation function and business processing function of the sharding machine, and the sharding machine actively performs the task of preempting sharded data; Specifically including: Any one sharding machine receives a business processing request; after this sharding machine finishes the pre-work, it notifies Zookeeper to broadcast the request; after all sharding machines monitor the broadcast, they simultaneously start to actively preempt sharded data; the pre-work specifically includes white list testing, setting the start and end positions of the data to be processed this time, task locking, and recording relevant task information; The sharding machine applies for sharded data in real time according to its own performance; specifically including: The sharding machine actively applies to redis for incr(x), where incr(x) represents the size of the sharded data processed by each sharding machine in a single batch; sharded data (x) = n * cpu, where x represents the size of the sharded domain data, n represents the single-cpu batch processing size, and cpu represents the number of cpu cores; Divide the data obtained by each sharding machine application into multiple threads for parallel processing; The last sharding machine that finishes processing performs post-processing work, specifically including: comparing the number of zk temporary nodes with the number of redis processed to determine whether it is the last sharding machine that finishes processing. If so, it is called the post-processor; the post-processor detects the integrity of this processing, conducts statistics and notifications; when the entire task ends, the post-processor releases the task lock; among them, each sharding machine will register with Zookeeper when it starts, that is, form a temporary node; after each sharding machine processes to the maximum value of this task, it performs a +1 operation in redis.
2. The method for processing ten-million-level data sharding according to claim 1, wherein During the process of preempting sharded data, if the sharding machine crashes and restarts, it will be automatically detected and continue to join the task cluster.
3. The method for processing ten-million-level data sharding according to claim 1, wherein The specific method of dividing the data obtained by each sharding machine application into multiple threads for parallel processing includes: After the sharding machine obtains a single batch of data domains, it marks this batch of data in redis, and then distributes it to the business Service nodes in multiple threads in sequence; After this batch of processing is completed, the mark is cleared, and the next batch of sharded domains is preempted.
4. The method for processing ten-million-level data sharding according to claim 3, wherein The number of threads divided by each sharding machine is calculated according to the cpu, and the calculation formula is default: cpu-intensive threads = the number of cpu cores + 1.
5. A ten-million-level data sharding processing device, characterized in that: It includes at least one processor and a memory, the at least one processor and the memory are connected through a data bus, the memory stores instructions that can be executed by the at least one processor, and after the instructions are executed by the processor, they are used to complete the ten-million-level data sharding processing method described in any one of claims 1-4.
Citation Information
Patent Citations
Data processing method and device
CN111752961A