A data processing method, system and computer-readable storage medium
Through the combination mechanism of distributed lock and cursor ID, multi-threaded data retrieval is serially controlled, which solves the problems of search failure and resource waste in the existing technology, and achieves efficient and reliable data processing consistency.
Patent Information
- Application Number
- CN202210025697.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-11
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-01-11
AI Technical Summary
When the prior art faces massive data processing, multi-threaded retrieval can easily lead to search failure or increase search time, and relying on downstream idempotent control and database unique index leads to wasted bandwidth and computing resources, making data consistency difficult to guarantee.
The combination mechanism of distributed lock and cursor ID is adopted to obtain the cursor ID by processing thread locking distributed locks, and data retrieval is serially controlled, and the cursor ID is used to record the end position of each round of search to avoid repeated searches and ensure data consistency.
It improves the data retrieval efficiency and reliability of multi-processing threads, avoids performance waste caused by repeated retrieval, and ensures the consistency of data processing.
Smart Images

Figure CN114356999B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of databases, and particularly to a data processing method, system and computer-readable storage medium. Background Art
[0002] In enterprise-level applications, it is often necessary to retrieve and process a large amount of data. Most of the data will be processed on the spot, but a small number of data cannot be processed on the spot and need to be processed by cleaning the database subsequently. Since the unprocessed data is distributed at different positions in the data table, a solution is needed to quickly find and process them.
[0003] In the related art, multiple threads are usually used to retrieve the data table. However, each thread needs to process from the head of the data table during retrieval. In the face of a huge amount of data, this retrieval method is likely to lead to retrieval failure or increase the retrieval time. In addition, during retrieval, data concurrency is allowed, and only downstream idempotency control or database unique indexes are relied on to ensure data consistency, which is likely to waste bandwidth, downstream system computing power and database computing resources. Summary of the Invention
[0004] The purpose of the present invention is to provide a data processing method, system and computer-readable storage medium, which can avoid performance waste caused by repeated retrieval by using a distributed lock and a cursor ID, and effectively ensure data processing consistency.
[0005] To solve the above technical problems, the present invention provides a data processing method, including:
[0006] After the processing thread is started by the processing engine, it locks the distributed lock, and when it determines that the lock is successful, it obtains the cursor ID of the specified data table in the first database; the distributed lock and the cursor ID are stored in the second database;
[0007] Determine the first data corresponding to the cursor ID and the second data of a preset number of lines after the first data in the data table, and retrieve the data to be processed from the first data and the second data; the data in the data table is provided with a data ID;
[0008] After the retrieval is completed, update the cursor ID to the data ID of the last second data, and release the distributed lock so that other processing threads can lock the distributed lock;
[0009] If the data to be processed is retrieved, process the data to be processed and send the processing result to the processing engine for merging;
[0010] Relock the distributed lock.
[0011] Optionally, before processing the data to be processed, it further includes:
[0012] The processing thread sends the ID of the data to be processed of the data to be processed to the second database;
[0013] The second database determines whether the ID of the data to be processed already exists;
[0014] If so, it records the ID of the data to be processed and sends a write success message to the processing thread, so that the processing thread executes the step of processing the data to be processed;
[0015] If not, it sends a write failure message to the processing thread, so that the processing thread does not execute the step of processing the data to be processed.
[0016] Optionally, after sending the write success message to the processing thread, it further includes:
[0017] The second database records the write time of the ID of the data to be processed;
[0018] Determine whether the ID of the data to be processed has expired according to the write time;
[0019] If so, clear the ID of the data to be processed.
[0020] Optionally, before updating the cursor ID to the ID of the last piece of the second data, it further includes:
[0021] The processing thread determines whether the data to be processed is retrieved;
[0022] If so, it executes the step of updating the cursor ID to the ID of the last piece of the second data;
[0023] If not, it retrieves the data to be processed again from the first data and the second data, and determines whether the data to be processed is retrieved;
[0024] If so, it updates the cursor ID to the ID of the data before the first piece of the data to be processed;
[0025] If not, it executes the step of updating the cursor ID to the ID of the last piece of the second data.
[0026] Optionally, the second database is a Redis database.
[0027] Optionally, before the processing thread is started by the processing engine, it further includes:
[0028] The processing engine receives the configuration information sent by the scheduling center; the configuration information includes data table parameters, retrieval conditions and retrieval methods of the data to be processed;
[0029] Configure the processing thread according to the configuration information, so that the processing thread searches for data in the first data table according to the data table parameters, and retrieves the data to be processed according to the data to be processed retrieval condition and the retrieval method;
[0030] When receiving the task execution instruction issued by the scheduling center, start multiple processing threads.
[0031] Optionally, after starting multiple processing threads, it further includes:
[0032] The processing engine records the task working time and determines whether the task working time is less than a preset threshold;
[0033] If so, continue to record the task working time;
[0034] If not, stop recording the task working time and close the processing thread.
[0035] Optionally, it further includes:
[0036] When the processing engine receives the reset cursor instruction, reset the cursor ID to a preset default value.
[0037] The present invention further provides a data processing system, including: a processing engine, a processing thread, a first database and a second database, wherein,
[0038] The processing thread is used to lock the distributed lock after being started by the processing engine, and obtain the cursor ID of the specified data table in the first database when it is determined that the lock is successful; the distributed lock and the cursor ID are stored in the second database; determine the first data corresponding to the cursor ID and the second data of a preset number of items after the first data in the data table, and retrieve the data to be processed from the first data and the second data; the data in the data table is set with a data ID; after the retrieval is completed, update the cursor ID to the data ID of the last second data, and release the distributed lock so that other processing threads can lock the distributed lock; if the data to be processed is retrieved, process the data to be processed and send the processing result to the processing engine for merging; re-lock the distributed lock;
[0039] The processing engine is used to start the processing thread; receive and merge the processing results sent by the processing thread;
[0040] The first database is used to store the data table;
[0041] The second database is used to store the distributed lock and the cursor ID.
[0042] The present invention also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, the data processing method as described above is implemented.
[0043] The present invention provides a data processing method, including: after a processing thread is started by a processing engine, locking a distributed lock, and obtaining a cursor ID of a specified data table in a first database when it is determined that the lock is successful; the distributed lock and the cursor ID are stored in a second database; determining first data corresponding to the cursor ID and second data of a preset number of entries after the first data in the data table, and retrieving data to be processed from the first data and the second data; data IDs are set for the data in the data table; after the retrieval is completed, updating the cursor ID to the data ID of the last piece of the second data, and releasing the distributed lock so that other processing threads can lock the distributed lock; if the data to be processed is retrieved, processing the data to be processed, and sending the processing result to the processing engine for merging; re-locking the distributed lock.
[0044] It can be seen that the present invention uses a distributed lock and a cursor ID to improve the retrieval ability of multiple processing threads. First, after a processing thread is started by a processing engine, it will lock the distributed lock and start executing data retrieval when it is determined that the lock is successful. Since only a single processing thread can successfully lock the distributed lock, the present invention can serially control multiple processing threads to perform data retrieval in the data table, ensuring that only a single processing thread can access and retrieve the database at a time; in addition, when starting to execute data retrieval, the processing thread will obtain the cursor ID, which is the starting position of the current round of retrieval by the processing thread. The thread will determine the first data corresponding to the cursor ID and the second data of a preset number of entries after the first data in the data table, perform retrieval in the first data and the second data, and after the retrieval is completed, update the cursor ID with the data ID of the last piece of the second data. In other words, the present invention will use the cursor ID to record the end position of each round of data retrieval and use it as the data retrieval position for the next round, avoiding performance waste caused by repeated retrieval. Coupled with the distributed lock, it can effectively ensure the consistency of data processing and greatly improve the efficiency and reliability of multiple processing threads in processing data. The present invention also provides a data processing system and a computer-readable storage medium, which have the above beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.
[0046] Figure 1 Flowchart of a data processing method provided by an embodiment of the present invention;
[0047] Figure 2 Execution schematic diagram of the existing remainder sharding scheme provided by an embodiment of the present invention when facing a small amount of data;
[0048] Figure 3 Execution schematic diagram of the technical solution of the present invention provided by an embodiment of the present invention when facing a small amount of data;
[0049] Figure 4 Execution schematic diagram of the existing remainder sharding scheme provided by an embodiment of the present invention when running to a large segment;
[0050] Figure 5 Execution schematic diagram of the technical solution of the present invention provided by an embodiment of the present invention when running to a large segment;
[0051] Figure 6 Execution schematic diagram of the technical solution of the present invention provided by an embodiment of the present invention when running to a large segment after introducing a custom starting point;
[0052] Figure 7 Execution flowchart of a data processing engine provided by an embodiment of the present invention;
[0053] Figure 8a Structural block diagram of a data processing system provided by an embodiment of the present invention;
[0054] Figure 8b Structural block diagram of another data processing system provided by an embodiment of the present invention. Detailed implementation manners
[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0056] In the related art, multiple threads are usually used to retrieve data tables. However, each thread needs to process from the data table header during retrieval. When faced with a huge amount of data, this retrieval method is likely to lead to retrieval failure or an increase in retrieval time. In addition, during retrieval, data concurrency is allowed, and only downstream idempotency control or database unique indexes are relied on to ensure data consistency, which is likely to waste bandwidth, downstream system computing power, and database computing resources. In view of this, the present invention provides a data processing method that can avoid performance waste caused by repeated retrieval by using a distributed lock and a cursor ID, and effectively ensure the consistency of data processing. Please refer to Figure 1 , Figure 1 is a flowchart of a data processing method provided by an embodiment of the present invention, and the method may include:
[0057] S101. After the processing thread is started by the processing engine, lock the distributed lock, and obtain the cursor ID of the specified data table in the first database when it is determined that the lock is successful; the distributed lock and the cursor ID are stored in the second database.
[0058] In an embodiment of the present invention, the processing thread is set in the processing engine, and one or more processing threads can be set in each processing engine. When each processing thread is started, it will first try to lock the distributed lock, where the distributed lock is used to specify the processing thread that can perform data retrieval. Simply put, the distributed lock can be understood as a data field in the database, which corresponds to an execution action, and only one thread can be allowed to insert data each time; for a thread, only after inserting data into the distributed lock, that is, after completing the locking of the distributed lock, can the action corresponding to the distributed lock be executed. In an embodiment of the present invention, each processing thread can only perform data retrieval in the first database after successfully locking the distributed lock. Obviously, after setting the distributed lock, all processing threads can only queue up for data retrieval, so the embodiment of the present invention can achieve serial control of the processing threads. It should be noted that the embodiment of the present invention does not limit the specific implementation manner of the distributed lock, and related technologies of the distributed lock can be referred to. It can be understood that, for the convenience of each thread to lock, the distributed lock is set in the common second database, and this database can be shared by multiple processing engines. In other words, the processing threads in multiple processing engines can perform the same work.
[0059] Further, to avoid performance waste caused by repeated retrievals, the embodiments of the present invention also set a cursor ID for recording the end position of each round of data retrieval. This is because in existing data retrieval solutions, the thread only performs a complete retrieval of the data table. When the data table is large or the data to be processed is located at the end of the data table, the complete retrieval is likely to increase the retrieval time. At the same time, during multi-threaded concurrent retrievals, the complete retrieval is likely to result in duplicate queries, which will not only waste the computing performance of the retrieval engine and the database, but also, if different threads process the same data, even lead to data inconsistency issues. Therefore, based on the distributed lock, the embodiments of the present invention also set a cursor ID for recording the end position of each round of data retrieval. Specifically, the processing thread in this round will obtain the cursor ID of the previous round as the retrieval starting point to retrieve a preset number of data, and update the cursor ID to the data ID of the last piece of data in this round, so that the processing thread in the next round can continue to retrieve backward starting from the last piece of data in this round. Obviously, after setting the distributed lock and the cursor ID, the present invention can ensure that each processing thread can obtain non-duplicate data in sequence, avoid the problem of repeated retrievals in the prior art, not only avoid wasting the computing performance of the retrieval engine and the database, improve the retrieval efficiency of the retrieval engine, but also guarantee data consistency. It can be understood that for the convenience of each thread to perform locking, the cursor ID is also set in the common second database.
[0060] Furthermore, it can be understood that the cursor ID should have a default value when it is first used. The embodiments of the present invention do not limit the specific initial value, which can be set according to actual application requirements. Further, to improve the flexibility of the retrieval engine, the engine can have the ability to restore the cursor ID to the default value. Of course, the retrieval engine can also be further added with the ability to set the cursor ID to a custom value, which can be set according to actual application requirements.
[0061] S102. Determine the first data corresponding to the cursor ID and the second data of a preset number of entries after the first data in the data table, and retrieve the data to be processed from the first data and the second data; the data in the data table is set with data IDs.
[0062] It should be noted that the embodiments of the present invention do not limit the specific value of the preset number of entries, which can be set according to actual application requirements. The present invention also does not limit the specific method for finding the data to be processed. For example, it can be determined whether the data has a to-be-processed mark, and when it has the to-be-processed mark, it is determined as the data to be processed.
[0063] S103. After the retrieval is completed, update the cursor ID to the data ID of the last piece of the second data, and release the distributed lock so that other processing threads can lock the distributed lock.
[0064] It should be noted that the retrieval is determined to be completed as long as the retrieval of the first data and the second data is completed. In other words, even if the first data and the second data do not contain the data to be processed, the retrieval is determined to be completed after the retrieval is completed. Further, since the data in the database is prone to state changes, when the data to be retrieved is not found in a certain round of retrieval, the data to be processed may reappear in the first data and the second data. To avoid omitting processing, the first data and the second data can be retrieved again after the retrieval is completed to determine whether there is newly appeared data to be processed. To improve the utilization rate of computing resources, the re-retrieval can be executed after it is determined that no data to be processed is retrieved in this round, that is, after it is determined that the retrieval is completed, it is judged whether the data to be processed is retrieved in this round. If not, the data retrieval is re-executed. The embodiments of the present invention do not limit whether to perform a complete re-retrieval of the first data and the second data, that is, to re-extract all the data to be processed in the first data and the second data. When the quantity of the first data and the second data is small, or when the newly obtained data to be processed needs to be processed in this round, a complete retrieval of the first data and the second data can be performed; when the quantity of the first data and the second data is large and only simply determine whether the data to be processed reappears in the first data and the second data, a complete retrieval may not be performed, and as long as it is determined that there is data to be processed in the first data and the second data, the re-retrieval is completed. In the embodiments of the present invention, considering that in practical applications, the amount of data to be retrieved by each processing thread is large, if a complete re-retrieval is performed, it is easy to waste computing resources. Therefore, a complete re-retrieval of the first data and the second data will not be performed, but as long as it is determined that there is data to be processed in the first data and the second data, the re-retrieval is completed. Specifically, the first data and the second data will be re-retrieved in the arranged order, and it is judged whether each data is the data to be processed. If so, it is determined that the data to be processed reappears in the first data and the second data. At this time, the cursor ID can be updated to the data ID of the data to be processed so that the next processing thread can continue to process.
[0065] In a possible situation, before updating the cursor ID to the data ID of the last second data, it may further include:
[0066] Step 11: The processing thread determines whether the data to be processed is retrieved; if so, go to Step 12; if not, go to Step 13;
[0067] Step 12: Execute the step of updating the cursor ID to the data ID of the last second data;
[0068] Step 13: Re-retrieve the data to be processed in the first data and the second data, and determine whether the data to be processed is retrieved; if so, go to Step 14; if not, go to Step 15;
[0069] Step 14: Update the cursor ID to the data ID of the data immediately preceding the first piece of data to be processed;
[0070] Step 15: Execute the step of updating the cursor ID to the data ID of the last piece of secondary data.
[0071] Furthermore, after determining that the retrieval is complete, the processing thread should release the distributed lock so that other processing threads can lock the distributed lock. It should be noted that the embodiments of the present invention do not limit how to release the distributed lock, and relevant technologies of distributed locks can be referred to.
[0072] S104. Determine whether the data to be processed is retrieved; if so, proceed to step S105; if not, proceed to step S106.
[0073] S105. If the data to be processed is retrieved, process the data to be processed and send the processing result to the processing engine for merging.
[0074] If the processing thread retrieves the data to be processed in this round, it should process the data to be processed according to the preset business logic and send the processing result to the processing engine for merging. It should be noted that since the processing thread starts to process the data to be processed after releasing the distributed lock, the processing threads in the embodiments of the present invention can process data in parallel. Simply put, the embodiments of the present invention can use the distributed lock and the cursor ID to ensure that the processing threads fetch data serially and process it in parallel, which can not only effectively guarantee data consistency but also effectively improve the data retrieval and data processing efficiency.
[0075] Furthermore, since data changes are likely to occur in the database, it is possible that the data to be processed retrieved by the processing thread in this round has been processed by other processing threads before. To further avoid duplicate processing, in the embodiments of the present invention, the processing thread can temporarily record the data ID of the data to be processed retrieved in this round in the database, and the database determines whether the data ID of the data to be processed has been recorded. If it has been recorded, the database sends a write failure message to the processing thread, and the processing thread will not process the data to be processed retrieved in this round; if not, the database sends a write success message to the processing thread, and the processing thread will process the data to be processed retrieved in this round. It can be understood that, for the convenience of each thread to record, the data ID of the data to be processed can be temporarily stored in the secondary database.
[0076] In a possible situation, before processing the data to be processed, it may further include:
[0077] Step 21: The processing thread sends the data ID of the data to be processed to the secondary database;
[0078] Step 22: The second database determines whether the ID of the data to be processed already exists; if so, proceed to Step 23; if not, proceed to Step 24;
[0079] Step 23: Record the ID of the data to be processed, and send a write success message to the processing thread so that the processing thread executes the steps for processing the data to be processed;
[0080] Step 24: Send a write failure message to the processing thread so that the processing thread does not execute the steps for processing the data to be processed.
[0081] Furthermore, the second database can also record the write time of each ID of the data to be processed, and determine whether the ID of the data to be processed has expired (i.e., exceeds the preset expiration time) based on the write time. If so, the ID of the data to be processed is cleared. It should be noted that the embodiments of the present invention do not limit the specific preset expiration time, which can be set according to actual application requirements.
[0082] In a possible case, after sending the write success message to the processing thread, it may further include:
[0083] Step 31: The second database records the write time of the ID of the data to be processed;
[0084] Step 32: Determine whether the ID of the data to be processed has expired based on the write time; if so, proceed to Step 33; if not, proceed to Step 32;
[0085] Step 33: Clear the ID of the data to be processed.
[0086] S106. Relock the distributed lock.
[0087] After completing the data processing or determining that no data to be processed is obtained in this round, the retrieval thread can start relocking the distributed lock again to enter the loop.
[0088] Furthermore, in the present invention, considering that it is convenient to set the distributed lock and store the cursor ID and the ID of the data to be processed in the Redis database, the second database can be set as the Redis database.
[0089] To better demonstrate the advantages of this solution, the following will be introduced based on specific comparison diagrams. For ease of understanding, first, the Figures 2 to 6Explain the structure in it. The above-mentioned drawings can all be divided into upper and lower parts. The upper part is an execution schematic diagram of each thread node retrieving physical data in the database. The spatial axis is used to represent the physical data in the database, the black vertical line is used to represent the starting scan position of each thread node, and the dark gray box is used to represent the valid data retrieved by the thread node; the lower part is a schematic diagram of the time-consuming of each thread node executing the data retrieval task, where the time axis represents the time unit.
[0090] Based on the above description, please refer to Figure 2 and Figure 3 , Figure 2 is an execution schematic diagram of the existing remainder sharding scheme provided by the embodiments of the present invention when facing a small amount of data. Figure 3 is an execution schematic diagram of the technical solution of the present invention provided by the embodiments of the present invention when facing a small amount of data. Among them, each thread node in the remainder sharding performs remainder sharding to obtain data. It can be seen that in the remainder sharding scheme, each thread node retrieves data in a parallel manner of obtaining data, while this scheme retrieves data in a serial manner of obtaining data. It should be noted that in the remainder sharding scheme, the number of data rows scanned by a single thread node each time is positively correlated with the number of shards. When the number of shards is large, each thread node needs to scan a large amount of data to obtain the required data, which is likely to increase the single scan time of the thread node. For this scheme, since serial data acquisition is adopted, the number of table scans of a single thread node in the database can be controlled to the lowest, thereby effectively shortening the single scan time of a single thread node.
[0091] Please refer to Figure 4 and Figure 5 , Figure 4 is an execution schematic diagram of the existing remainder sharding scheme provided by the embodiments of the present invention when running to a large segment. Figure 5 is an execution schematic diagram of the technical solution of the present invention provided by the embodiments of the present invention when running to a large segment. Among them, the large segment represents that the data to be processed is located at the end of the database table. It can be seen that for the remainder sharding scheme, since each thread retrieves data in parallel, after the retrieval task (JOB) is started, each thread node will have a slow query once; for this scheme, since each thread retrieves data serially, after the retrieval task is started, only the first thread node will have a slow query once.
[0092] To solve the slow start problem faced by the distributed lock + Redis global ID cursor scheme, the embodiments of the present invention can also be optimized by adopting a custom starting point, that is, setting the ID of the data to be processed as a custom value. Please refer to Figure 6 , Figure 6This is an execution schematic diagram of the technical solution of the present invention when a custom starting point is introduced and the invention runs to a large segment. It can be seen that after introducing the custom starting point, the slow query problem that may occur in the first thread node can be further avoided.
[0093] Based on the above embodiments, the present invention uses a distributed lock and a cursor ID to improve the retrieval ability of multiple processing threads. First, after the processing thread is started by the processing engine, it locks the distributed lock and starts to execute data retrieval when it determines that the lock is successfully obtained. Since only a single processing thread can successfully lock the distributed lock, the present invention can serially control multiple processing threads to perform data retrieval in the data table, ensuring that only a single processing thread can access and retrieve the database each time; in addition, when starting to execute data retrieval, the processing thread will obtain the cursor ID, which is the starting position of the current round of retrieval of the processing thread. The thread will determine the first data corresponding to the cursor ID and the second data of a preset number of records after the first data in the data table, retrieve among the first data and the second data, and after the retrieval is completed, update the cursor ID with the data ID of the last second data. In other words, the present invention will use the cursor ID to record the end position of each round of data retrieval and use it as the data retrieval position for the next round, avoiding the performance waste caused by repeated retrieval. Coupled with the distributed lock, it can effectively ensure the consistency of data processing and greatly improve the efficiency and reliability of multiple processing threads in processing data.
[0094] Based on the above embodiments, to improve the scheduling efficiency, the embodiment of the present invention can also additionally set up a scheduling center to uniformly schedule the processing engine. The following introduces the specific process of the scheduling center scheduling the processing engine.
[0095] In a possible situation, before the processing thread is started by the processing engine, it may further include:
[0096] S201. The processing engine receives the configuration information sent by the scheduling center; the configuration information includes data table parameters, retrieval conditions for data to be processed, and retrieval methods.
[0097] Specifically, the scheduling center first sends configuration information to the processing engine for relevant configuration. The configuration information may include data table parameters (table name and primary key field), retrieval conditions for the data to be processed (such as the data status being "to be processed"), and retrieval methods (i.e., the custom query actions provided by the user). It should be noted that the embodiments of the present invention do not limit the specific content of the data table parameters, retrieval conditions for the data to be processed, and retrieval methods, and relevant database retrieval technologies can be referred to. The embodiments of the present invention also do not limit the number of processing engines that the scheduling center can manage, which can be one or multiple, and can be set according to actual application requirements. The embodiments of the present invention also do not limit the specific scheduling center, and relevant distributed scheduling center technologies can be referred to and set in combination with actual application requirements. Considering that the XXL-JOB platform is a lightweight distributed task scheduling platform that is easy to learn and expand, the scheduling center in the embodiments of the present invention can adopt XXL-JOB.
[0098] S202. Configure processing threads according to the configuration information, so that the processing threads search for data in the data table in the first data table according to the data table parameters, and retrieve the data to be processed according to the retrieval conditions and retrieval methods of the data to be processed.
[0099] S203. When receiving the task execution instruction sent by the scheduling center, start multiple processing threads.
[0100] Further, when the scheduling center issues a task execution instruction, the processing engine starts multiple processing threads. It should be noted that the embodiments of the present invention do not limit the manner in which the scheduling center sends task execution instructions to multiple processing engines. For example, they can be sent simultaneously, or periodically (i.e., after sending an instruction to the first processing engine, waiting for a preset time before sending an instruction to the next processing engine), and can be set according to actual application requirements.
[0101] Further, to avoid long-term work, the processing engine also records the task working time after starting the processing threads, determines whether the task working time exceeds a preset threshold, and automatically stops the processing threads from working when it is determined that the preset threshold is exceeded.
[0102] In a possible situation, after starting multiple processing threads, it may further include:
[0103] Step 41: The processing engine records the task working time and determines whether the task working time is less than the preset threshold; if so, go to Step 42; if not, go to Step 43;
[0104] Step 42: Continue to record the task working time;
[0105] Step 43: Stop recording the task working time and close the processing threads.
[0106] Further, to restore the cursor ID for retrieving the data table again, the processing engine can additionally receive a cursor reset instruction, and after receiving this instruction, reset the cursor ID to a preset default value. It should be noted that the embodiments of the present invention do not limit the source of the cursor reset instruction. For example, it can be generated when detecting that the user clicks a preset button, or can be sent by the scheduling center to the processing engine, and can be set according to actual application requirements.
[0107] In a possible case, it may further include:
[0108] Step 51: When the processing engine receives the cursor reset instruction, reset the cursor ID to a preset default value.
[0109] Based on the above embodiments, the present invention can additionally set the scheduling center to uniformly configure and start the processing engine to improve the unified scheduling ability of the distributed processing engine.
[0110] The following introduces the above data processing method based on a specific example. Please refer to Figure 7 , Figure 7 which is the execution flowchart of a data processing engine provided by the embodiments of the present invention. The main processes of data processing are as follows:
[0111] i. WAD (Work All Day) engine main process
[0112] 1. Use the XXL-JOB scheduling platform to start the engines of multiple nodes simultaneously;
[0113] 2. Each node can start multiple threads to simultaneously retrieve and process data (to be elaborated later);
[0114] 3. Collect the processing results of multiple threads and merge the result sets;
[0115] 4. If no data is found, it means the data has been processed, and the loop ends;
[0116] 5. If the maximum working time of the task is exceeded, the loop ends;
[0117] 6. Otherwise, keep looping for processing.
[0118] ii. Data retrieval and processing
[0119] 1. Serially retrieve data (to be elaborated later) to ensure that multiple nodes queue up to obtain data in an orderly manner, and then allow parallel processing;
[0120] 2. If no data is retrieved, directly return an empty set;
[0121] 3. If retrieved, note that some downstream processing capabilities are very weak. If the data processed within 120 seconds (but not successfully processed and now retrieved again), it will not be processed and will be temporarily ignored;
[0122] 4. Execute the user-defined business operations.
[0123] iii. Serial retrieval of data
[0124] 1. Find the last stored cursor ID in Redis;
[0125] 2. If the cursor ID cannot be found, reset it to the preset custom starting point (default is 0);
[0126] 3. Retrieve data (in the case of read-write separation, use the write library);
[0127] 4. If data is retrieved, obtain the ID of the last piece of data and mark the new cursor in Redis
[0128] 5. If data cannot be retrieved, use the
last cursor strategy
[0129] and record it in Redis.
[0130] iv. Last cursor strategy
[0131] 1. Ignoring the user-defined business filtering conditions, directly obtain the last piece of data in the database and record it as LastOneId;
[0132] 2. The obtained LastOneId may not be reliable as the data may change at any time. Retrieve again whether there is new valid business data between the cursor ID and LastOneId:
[0133] a) If there is no new data, record LastOneId as the new cursor;
[0134] b) If there is new data, note that at this time, it is not to process the data but to locate the cursor. Record the ID of the first piece of data in ascending order minus 1 as the new cursor.
[0135] The following introduces the data processing system and computer-readable storage medium provided by the embodiments of the present invention. The data processing system and computer-readable storage medium described below can be correspondingly referred to the data processing method described above. Please refer to Figure 8a , Figure 8a which is a structural block diagram of a data processing system provided by an embodiment of the present invention. The system may include: a processing engine 801, a processing thread 802, a first database 803, and a second database 804, where,
[0136] A processing thread 802, which is used to lock a distributed lock after being started by a processing engine 801, and obtain a cursor ID of a specified data table in a first database when it is determined that the lock is successful; the distributed lock and the cursor ID are stored in a second database; determine first data corresponding to the cursor ID and second data of a preset number after the first data in the data table, and retrieve data to be processed from the first data and the second data; data in the data table is set with a data ID; after the retrieval is completed, update the cursor ID to the data ID of the last second data, and release the distributed lock so that other processing threads 802 can lock the distributed lock; if data to be processed is retrieved, process the data to be processed, and send the processing result to the processing engine 801 for merging; relock the distributed lock;
[0137] A processing engine 801, which is used to start the processing thread 802; receive the processing result sent by the processing thread 802 and merge it;
[0138] A first database 803, which is used to store the data table;
[0139] A second database 804, which is used to store the distributed lock and the cursor ID.
[0140] Optionally, the processing thread 802 can also be used to send the data ID to be processed of the data to be processed to the second database 804;
[0141] The second database 804 can also be used to determine whether a data ID to be processed already exists; if so, record the data ID to be processed and send a write success message to the processing thread 802 so that the processing thread 802 executes the step of processing the data to be processed; if not, send a write failure message to the processing thread 802 so that the processing thread 802 does not execute the step of processing the data to be processed.
[0142] Optionally, the second database 804 can also be used to record the write time of the data ID to be processed; determine whether the data ID to be processed has expired according to the write time; if so, clear the data ID to be processed.
[0143] Optionally, the processing thread 802 can also be used to determine whether data to be processed is retrieved; if so, execute the step of updating the cursor ID to the data ID of the last second data; if not, retrieve the data to be processed again from the first data and the second data, and determine whether data to be processed is retrieved; if so, update the cursor ID to the data ID of the data before the first data to be processed; if not, execute the step of updating the cursor ID to the data ID of the last second data.
[0144] Optionally, the second database 804 is a Redis database.
[0145] Optionally, please refer toFigure 8b , Figure 8b is a structural block diagram of another data processing system provided by an embodiment of the present invention. The system may further include: a scheduling center 805, where:
[0146] The processing engine 801 may also be used to receive configuration information sent by the scheduling center 805; the configuration information includes data table parameters, retrieval conditions for data to be processed, and retrieval methods; configure the processing thread 802 according to the configuration information, so that the processing thread 802 searches for data in the data table according to the data table parameters, and retrieves the data to be processed according to the retrieval conditions and retrieval methods for the data to be processed; when receiving a task execution instruction sent by the scheduling center 805, start multiple processing threads 802;
[0147] The scheduling center 805 is used to send configuration information to the processing engine 801; send a task execution instruction to the processing engine 801.
[0148] Optionally, the processing engine 801 may also be used to record the task working time, and determine whether the task working time is less than a preset threshold; if so, continue to record the task working time; if not, stop recording the task working time, and close the processing thread 802.
[0149] Optionally, the processing engine 801 may also be used to reset the cursor ID to a preset default value when receiving a reset cursor instruction.
[0150] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data processing method in any of the above embodiments are implemented.
[0151] Since the embodiments in the computer-readable storage medium part correspond to the embodiments in the data processing method part, for the description of the embodiments in the storage medium part, please refer to the description of the embodiments in the data processing method part, and will not be elaborated here for the time being.
[0152] The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method part.
[0153] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0154] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0155] The above has introduced in detail a data processing method, system, and computer-readable storage medium provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A data processing method, characterized in that, Including: After being started by the processing engine, the processing thread locks the distributed lock and, when determining that the lock is successful, obtains the cursor ID of the specified data table in the first database; the distributed lock and the cursor ID are stored in the second database, and the processing engine is provided with multiple processing threads; Determine the first data corresponding to the cursor ID and the second data of a preset number of entries after the first data in the data table, and retrieve the data to be processed from the first data and the second data; the data in the data table is provided with a data ID; After the retrieval is completed, update the cursor ID to the data ID of the last piece of the second data, and release the distributed lock so that other processing threads can lock the distributed lock; Judge whether the data to be processed is retrieved; If the data to be processed is retrieved, process the data to be processed, send the processing result to the processing engine for merging, and re-lock the distributed lock; If the data to be processed is not retrieved, re-lock the distributed lock.
2. The data processing method according to claim 1, wherein Before processing the data to be processed, it further includes: The processing thread sends the data ID of the data to be processed to the second database; The second database judges whether the data ID of the data to be processed already exists; If not, record the data ID of the data to be processed, and send a write success message to the processing thread so that the processing thread executes the step of processing the data to be processed; If so, send a write failure message to the processing thread so that the processing thread does not execute the step of processing the data to be processed.
3. The data processing method according to claim 2, wherein After sending the write success message to the processing thread, it further includes: The second database records the write time of the data ID of the data to be processed; Judge whether the data ID of the data to be processed has expired according to the write time; If so, clear the data ID of the data to be processed.
4. The data processing method according to claim 1, characterized in that Before updating the cursor ID to the data ID of the last piece of the second data, it further includes: The processing thread judges whether the data to be processed is retrieved; If so, execute the step of updating the cursor ID to the data ID of the last piece of the second data; If not, re-retrieve the data to be processed from the first data and the second data, and judge whether the data to be processed is retrieved; If so, update the cursor ID to the data ID of the data before the first piece of the data to be processed; If not, execute the step of updating the cursor ID to the data ID of the last piece of the second data.
5. The data processing method according to claim 1, wherein The second database is a Redis database.
6. The data processing method according to any one of claims 1 to 5, characterized in that Before the processing thread is started by the processing engine, it further includes: The processing engine receives the configuration information sent by the scheduling center; the configuration information includes data table parameters, data to be processed retrieval conditions, and retrieval methods; Configure the processing thread according to the configuration information so that the processing thread searches for the data in the data table according to the data table parameters in the first data table, and retrieves the data to be processed according to the data to be processed retrieval conditions and the retrieval methods; When receiving the task execution instruction sent by the dispatching center, start multiple processing threads.
7. The data processing method according to claim 6, wherein After starting multiple processing threads, it further includes: The processing engine records the task working time and determines whether the task working time is less than a preset threshold; If so, continue to record the task working time; If not, stop recording the task working time and close the processing threads.
8. The data processing method according to claim 6, wherein It further includes: When the processing engine receives the reset cursor instruction, reset the cursor ID to a preset default value.
9. A data processing system, characterized in that, It includes: A processing engine, processing threads, a first database, and a second database. The processing engine is provided with multiple processing threads. Among them, The processing thread is used to lock the distributed lock after being started by the processing engine, and obtain the cursor ID of a specified data table in the first database when it is determined that the lock is successful; the distributed lock and the cursor ID are stored in the second database; determine the first data corresponding to the cursor ID and the second data of a preset number of entries after the first data in the data table, and retrieve the data to be processed from the first data and the second data; the data in the data table is provided with a data ID; after the retrieval is completed, update the cursor ID to the data ID of the last piece of the second data, and release the distributed lock so that other processing threads can lock the distributed lock; determine whether the data to be processed is retrieved; if the data to be processed is retrieved, process the data to be processed, send the processing result to the processing engine for merging, and re-lock the distributed lock; if the data to be processed is not retrieved, re-lock the distributed lock; The processing engine is used to start the processing threads; receive and merge the processing results sent by the processing threads; The first database is used to store the data table; The second database is used to store the distributed lock and the cursor ID.
10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium. When the computer-executable instructions are loaded and executed by a processor, the data processing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Data migration method and data migration system
CN110532247A
Data pushing method and device, server, and storage medium
WO2021179170A1