A data processing method and apparatus
By inserting identification tasks into tasks distributed by data producers and setting queue locks, the problem of global orderly loss of concurrent processing tasks in real-time data synchronization between heterogeneous databases is solved, and the accuracy of task processing is improved.
Patent Information
- Application Number
- CN202011463399.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-12-11
AI Technical Summary
In real-time data synchronization scenarios between heterogeneous databases, it is difficult for the prior art to retain globally ordered information of multiple tasks when processing tasks concurrently, resulting in possible consumption errors.
By inserting identification tasks into tasks distributed by data producers and setting queue locks in the message queue, ensure that the data task as a prerequisite is processed after the identification task is processed before and after, thus retaining the global ordered information of the task.
The data task as the prerequisite is effectively prevented from being processed before the data task as the prerequisite, and improves the accuracy of data task processing.
Smart Images

Figure CN114625546B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network technologies, and in particular, to a data processing method and apparatus. Background Art
[0002] In the scenario of real-time data synchronization between heterogeneous databases (such as synchronizing data in relational databases MySQL / Oracle to non-relational databases ES, HDFS, HBase, etc. in real time), it is often necessary to introduce a message queue middleware as a data transfer platform to reduce the coupling of each functional module of the system, eliminate traffic peaks, and improve the overall performance of the system through asynchronous concurrent processing. For example, all data changes in the relational database are recorded in a log file, the content of the log file is parsed in real time, and the obtained multiple parsing results are distributed to multiple message queues, and the non-relational database consumes the data in the message queues in real time. However, this method can only ensure the order of each parsing result in one message queue. For example, parsing results 1, 2, 3, 4, 5, 6 are distributed to two message queues, and the parsing results obtained in message queue 1 are 1, 3, 5; the parsing results obtained in message queue 1 are 2, 4, 6. That is to say, this method can only ensure local order, and all parsing results in multiple message queues lose global order information. Therefore, when there are prerequisites in multiple parsing results, the globally disordered parsing results may lead to consumption errors. For example, if there are two types of parsing results, DML (Data Manipulation Language) and DDL (Data Definition Language), in multiple parsing results, where the DDL parsing result is to add column C2, and the DML parsing result is to add A2 in column C2, in this case, 'add column C2' must be consumed before 'add A2 in column C2', otherwise consumption errors will occur.
[0003] Therefore, there is an urgent need for a data processing method and apparatus that can ensure that global order information of multiple tasks is still retained when processing concurrent tasks, and improve the accuracy of task processing. Summary of the Invention
[0004] Embodiments of the present invention provide a data processing method and apparatus that can ensure that global order information of multiple tasks is still retained when processing concurrent tasks, and improve the accuracy of task processing.
[0005] In a first aspect, embodiments of the present invention provide a data processing method, the method including:
[0006] The data consumer obtains tasks from the message queue corresponding to the data consumer; the message queue corresponding to the data consumer is one of multiple message queues in the message partition, and each message queue corresponds to a data consumer; each data consumer executes concurrently; when the data consumer determines that the task is an identification task, it stops obtaining tasks from the message queue corresponding to the data consumer; when the identification task is distributed by the data producer, it is set before and after the data task serving as a prerequisite according to the distribution order and written into each message queue in the message partition; when the data consumer determines that all data consumers have stopped obtaining tasks from the message partition, it resumes obtaining tasks from the message queue corresponding to the data consumer.
[0007] In the above method, when the data consumer obtains an identification task, it stops continuing to obtain tasks. Since the identification task is set before and after the data task serving as a prerequisite according to the distribution order when distributed by the data producer and written into each message queue in the message partition, all data consumers have obtained the identification task and stopped obtaining tasks before processing the data task serving as a prerequisite. That is to say, in the order of data tasks obtained by the data producer, after all data tasks before the data task serving as a prerequisite are processed, all data consumers stop obtaining and processing tasks and wait for the data task serving as a prerequisite to be processed before continuing. In this way, the order information between the prerequisite task and the conditional task can be preserved in the order of data tasks obtained by the data producer. For example, if the current task to be consumed in the first message queue is a DDL task and the DML task related to the DDL task is in the second message queue, and if the DML task in the current second message queue will be processed before the DDL task; this situation will lead to incorrect processing of the DML task. Therefore, the second message queue is locked before processing the DML task, and after the DDL task in the first message queue is processed, the second message queue is allowed to process the DML task, preserving the global order information of the data tasks and improving the accuracy of data task processing.
[0008] Optionally, when the data consumer determines that the task is an identification task and stops obtaining tasks from the message queue corresponding to the data consumer, it includes: when the data consumer determines that the task is an identification task, it sets a queue lock for the message queue corresponding to the data consumer; the queue lock is used to indicate that the data consumer stops obtaining tasks from the message queue corresponding to the data consumer; when the data consumer determines that all data consumers have stopped obtaining tasks from the message partition, it includes: the data consumer determines that each message queue in the message partition is in a locked state.
[0009] In the above method, when the data consumer determines that the obtained task is an identification task, a queue lock is set for the message queue corresponding to the data consumer, so that the data consumer stops obtaining tasks from the corresponding message queue. In this way, it is ensured that the data task as a prerequisite is processed before the data task as a decision condition. This prevents the processing error caused by the data consumer still obtaining tasks from the corresponding message queue without setting the queue lock, resulting in the data task as a decision condition being processed before the data task as a prerequisite.
[0010] Optionally, after setting the queue lock for the message queue corresponding to the data consumer, it further includes:
[0011] The data consumer increases the count value of the queue lock; the data consumer determines that all message queues in the message partition are in a locked state, including: the data consumer determines that the count value meets the number of the message queues.
[0012] In the above method, after setting the queue lock for the message queue corresponding to the data consumer, the data consumer increases the count value of the queue lock. In this way, the data consumer can know the change in the count value. When the count value reaches the total number of message queues, the queue lock is released, so that the data consumer can continue to obtain and process tasks from the corresponding message queue, ensuring the task processing speed.
[0013] Optionally, the data task as a prerequisite is a DDL (Data Definition Language) task.
[0014] In the above method, if the data task as a prerequisite is a DDL task, it can be ensured that the DML (Data Manipulation Language) task is processed after the corresponding DDL task. This ensures the accuracy of the corresponding data task processing.
[0015] In a second aspect, an embodiment of the present invention provides a data processing method, and the method includes:
[0016] The data producer determines the distribution order of each task, where identification tasks are set before and after the data task as a prerequisite; the identification task is used to instruct the data consumer to stop obtaining tasks from the message queue corresponding to the data consumer when the task identifier is obtained;
[0017] The data producer distributes the tasks to the corresponding message queues in the message partition in sequence according to the distribution order; among them, the identification task will be distributed to each message queue in the message partition; each message queue corresponds to a data consumer.
[0018] In the above method, identification tasks are inserted before and after the data tasks that are prerequisites. When distributing tasks in sequence, the identification tasks are sent to all message queues. In this way, when there is a situation where a data task that is a condition to be determined is processed before a data task that is a prerequisite; before processing the data task that is a condition to be determined, other message queues are locked through the identification task; wait until the data consumer where the data task that is a prerequisite is located obtains the identification task and is also locked; then release the queue locks of all locked message queues. In this way, it is ensured that the data task that is a prerequisite is processed first. The accuracy of task processing is increased.
[0019] Optionally, the data producer distributes the respective tasks to the corresponding message queues in the message partition in sequence according to the distribution order, including: the data producer distributes the tasks of the same type or the same primary key to the same message queue in sequence according to the distribution order.
[0020] In the above method, the data producer distributes the tasks of the same type or the same primary key to the same message queue. In this way, the task order is retained among the tasks of the same type or the same primary key. Further, when multiple tasks of the same type or the same primary key are update tasks, the order of task updates is ensured, so as to ensure the accuracy of the final update result.
[0021] In a third aspect, an embodiment of the present invention provides a data processing device, and the device includes:
[0022] An acquisition module, configured to acquire tasks from the message queue corresponding to the data consumer; the message queue corresponding to the data consumer is one of multiple message queues in the message partition, and each message queue corresponds to a data consumer; each data consumer executes concurrently;
[0023] A determination module, configured to stop acquiring tasks from the message queue corresponding to the data consumer when determining that the task is an identification task; the identification task is set before and after the data task that is a prerequisite according to the distribution order when distributed by the data producer, and is written into each message queue in the message partition;
[0024] The acquisition module is further configured to, when the data consumer determines that all data consumers stop acquiring tasks from the message partition, acquire tasks from the message queue corresponding to the data consumer again.
[0025] In a fourth aspect, an embodiment of the present invention provides a data processing device, and the device includes:
[0026] A determination module, configured to determine the distribution order of each task, where identification tasks are arranged before and after a data task that is a prerequisite; the identification task is used to instruct a data consumer to stop obtaining tasks from the message queue corresponding to the data consumer when the task identifier is obtained.
[0027] A distribution module, configured to sequentially distribute each task to the corresponding message queue in the message partition according to the distribution order; wherein, the identification task will be distributed to each message queue in the message partition; each message queue corresponds to a data consumer.
[0028] In a fifth aspect, an embodiment of the present application further provides a computing device, including: a memory, configured to store a program; a processor, configured to call the program stored in the memory and execute the method described in the various possible designs of the first aspect and the second aspect according to the obtained program.
[0029] In a sixth aspect, an embodiment of the present application further provides a computer-readable non-volatile storage medium, including a computer-readable program, when a computer reads and executes the computer-readable program, the computer is caused to execute the method described in the various possible designs of the first aspect and the second aspect.
[0030] These implementation manners or other implementation manners of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0032] Figure 1 A schematic diagram of a data processing architecture provided by an embodiment of the present invention;
[0033] Figure 2 A flowchart of a data processing method provided by an embodiment of the present invention;
[0034] Figure 3 A schematic diagram of task distribution provided by an embodiment of the present invention;
[0035] Figure 4 A flowchart of a data processing method provided by an embodiment of the present invention;
[0036] Figure 5 A flowchart of a data processing method provided by an embodiment of the present invention;
[0037] Figure 6 Schematic diagram of a data processing device provided by an embodiment of the present invention;
[0038] Figure 7 Schematic diagram of a data processing device provided by an embodiment of the present invention. Detailed implementation manners
[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0040] An embodiment of the present invention provides a system architecture for data processing, as Figure 1 shown. Among them, a data producer 101 generates a plurality of data tasks with a certain order, and determines whether there is a data task as a prerequisite among the plurality of data tasks. If there is a data task as a prerequisite, identification tasks are inserted before and after the data task as a prerequisite; then, the plurality of data tasks with identification tasks inserted are distributed to each message queue in a message partition 102 in order. During the task distribution process, when distributing an identification task, one identification task is sent to each message queue in the message partition. Each data consumer in a consumption partition 103 obtains tasks from the corresponding message queue in the message partition 102 for processing. When the obtained task is an identification task, the data consumer stops obtaining tasks from the corresponding message queue, sets a queue lock for the message queue corresponding to the data consumer, and increases the count value of the queue lock; when the data consumers in the consumption partition 103 determine that all the message queues in the message partition 102 are in a locked state, or when the count value meets the number of each message queue, the queue lock is released; each data consumer in the consumption partition 103 continues to obtain and process tasks in the corresponding message queue.
[0041] Based on this, an embodiment of the present application provides a flow of a data processing method, as Figure 2 shown, including:
[0042] Step 201, a data consumer obtains a task from the message queue corresponding to the data consumer; the message queue corresponding to the data consumer is one of the plurality of message queues in the message partition, and each message queue corresponds to a data consumer; each data consumer executes concurrently;
[0043] Step 202: When the data consumer determines that the task is an identification task, it stops obtaining tasks from the message queue corresponding to the data consumer; when the identification task is distributed by the data producer, it is set before and after the data task serving as a prerequisite according to the distribution order and written into each message queue in the message partition.
[0044] Here, the data task serving as a prerequisite is a condition for the execution of one or more of a plurality of data tasks with a certain order; for example, the data task serving as a prerequisite is to add a column in a certain table, and the one or more data tasks after the data task serving as a prerequisite are to write data or update data under this column. Here, the one or more data tasks after the data task serving as a prerequisite can be data tasks serving as determined conditions. The identification task is used to identify that the data consumer stops obtaining tasks from the message queue corresponding to the data consumer, and it can be a symbol, text, number, etc., and is not specifically limited; for example, @, #, *, 1, 0, etc. Among them, if the order of multiple data tasks is A1, #1, B1, A2, C1, C2, and #1 is the data task serving as a prerequisite; then the data producer inserts the identification task @ before and after the data task serving as a prerequisite to obtain A1, @, #1, @, B1, A2, C1, C2; the data producer distributes each task to each message queue according to the order of the multiple tasks after inserting the identification task @, but when distributing the identification task @, this identification task @ will be sent to each message queue, as Figure 3 shown.
[0045] Step 203: When the data consumer determines that all data consumers have stopped obtaining tasks from the message partition, it resumes obtaining tasks from the message queue corresponding to the data consumer.
[0046] In the above method, when the data consumer obtains the identification task, it stops obtaining tasks continuously. Since the identification tasks are set before and after the data tasks serving as prerequisites in the order of distribution when distributed by the data producer and written into each message queue in the message partition, all data consumers have obtained the identification tasks and stopped obtaining tasks before processing the data tasks serving as prerequisites. That is to say, in the order of data tasks obtained by the data producer, after all data tasks before the data tasks serving as prerequisites are processed, all data consumers stop obtaining and processing tasks and wait for the data tasks serving as prerequisites to be processed before continuing to process. In this way, the order information between the prerequisite tasks and the conditional tasks can be preserved in the order of data tasks obtained by the data producer. For example, if the current task to be consumed in the first message queue is a DDL task and the DML task related to this DDL task is in the second message queue, and if the DML task in the current second message queue will be processed before the DDL task; this situation will lead to incorrect processing of the DML task. Therefore, the second message queue is locked before processing the DML task, and after the DDL task in the first message queue is processed, the second message queue is allowed to process the DML task, preserving the global order information of the data tasks and improving the accuracy of data task processing.
[0047] An embodiment of the present application provides a method for locking a message queue. When the data consumer determines that the task is an identification task, it stops obtaining tasks from the message queue corresponding to the data consumer, including: when the data consumer determines that the task is an identification task, it sets a queue lock for the message queue corresponding to the data consumer; the queue lock is used to instruct the data consumer to stop obtaining tasks from the message queue corresponding to the data consumer; when the data consumer determines that all data consumers have stopped obtaining tasks from the message partition, including: the data consumer determines that all message queues in the message partition are in a locked state. That is to say, when the data consumer obtains an identification task, a queue lock is set for the consumption queue corresponding to the data consumer, so that the data consumer stops obtaining tasks from the message queue corresponding to the data consumer. Since the data producer distributes tasks in order and distributes the identification task to all message queues in the message partition when distributing the identification task. Thus, when all message queues in the message partition are in a locked state, it means that the tasks before the data task as a prerequisite have been processed, and the tasks after the data task as a prerequisite have not been processed; at this time, when the queue lock is released, the data consumer corresponding to the message queue where the data task as a prerequisite is located in the message partition will obtain the data task as a prerequisite, while other data consumers still obtain the identification task and continue to set the queue lock for the corresponding message queue; in this way, it is prevented that other data consumers obtain the data task corresponding to the data task as a prerequisite as a conditional data task, resulting in the conditional data task being processed before the prerequisite data task, causing an error in processing. In Figure 3 In the example of
[0048]
[0049]
[0050] Table 1
[0051] Among them, in the first round:
[0052] The data consumer 1 obtains the data task A1 and processes it;
[0053] The data consumer 2 obtains the identification task @ and sets a queue lock for the corresponding message queue 2, so that the message queue 2 is in a locked state, and the data consumer 2 stops obtaining tasks from the message queue 2;
[0054] The data consumer 3 obtains the identification task @, and sets a queue lock for its corresponding message queue 3, making the message queue 3 in a locked state, and the data consumer 3 stops obtaining tasks from the message queue 3;
[0055] The second round:
[0056] The data consumer 1 obtains the identification task @, and sets a queue lock for its corresponding message queue 1, making the message queue 1 in a locked state, and the data consumer 1 stops obtaining tasks from the message queue 1. During this process, the message queue 2 corresponding to the data consumer 2 and the message queue 3 corresponding to the data consumer 3 are both in a locked state;
[0057] The data consumer determines that each data consumer has stopped obtaining tasks from the message partition, unlocks and releases each message queue in the message partition, and the data consumers corresponding to each message queue continue to obtain and process tasks;
[0058] The third round:
[0059] The data consumer 1 obtains the identification task @, and sets a queue lock for its corresponding message queue 1, making the message queue 1 in a locked state, and the data consumer 1 stops obtaining tasks from the message queue 1;
[0060] The data consumer 2 obtains the data task #1 as a prerequisite and processes it;
[0061] The data consumer 3 obtains the identification task @, and sets a queue lock for its corresponding message queue 3, making the message queue 3 in a locked state, and the data consumer 1 stops obtaining tasks from the message queue 3;
[0062] The fourth round:
[0063] The data consumer 2 obtains the identification task @, and sets a queue lock for its corresponding message queue 2, making the message queue 2 in a locked state, and the data consumer 2 stops obtaining tasks from the message queue 2. During this process, the message queue 1 corresponding to the data consumer 1 and the message queue 3 corresponding to the data consumer 3 are both in a locked state;
[0064] The data consumer determines that each data consumer has stopped obtaining tasks from the message partition, unlocks and releases each message queue in the message partition, and the data consumers corresponding to each message queue continue to obtain and process tasks;
[0065] The fifth round:
[0066] The data consumer 1 obtains the data task A2 and processes it;
[0067] The data consumer 2 obtains the data task C1 and processes it;
[0068] Data consumer 3 obtains data task B1 and processes it;
[0069] The sixth round:
[0070] The message queue 1 corresponding to data consumer 1 is empty, and no data task can be obtained, so the task processing is completed;
[0071] Data consumer 2 obtains data task C2 and processes it;
[0072] The message queue 3 corresponding to data consumer 3 is empty, and no data task can be obtained, so the task processing is completed;
[0073] In this way, in the task processing process of the above data consumers, the data task processing order is A1, #1, [B1, A2, C1], C2. If A2 is the data task as the conditioned task of the prerequisite data task #1, the above method can ensure that A2 is processed after #1, effectively preventing data task processing errors and improving the accuracy. It should be noted here that the above first round, second round, etc. are only for convenience of description and do not limit conditions such as the cycle of data consumer task processing.
[0074] The embodiment of the present application provides a data processing method. After setting a queue lock for the message queue corresponding to the data consumer, it further includes: the data consumer increases the count value of the queue lock; the data consumer determines that each message queue in the message partition is in a locked state, including: the data consumer determines that the count value meets the number of each message queue. That is to say, in the first round of the previous example, data consumer 1 obtains data task A1 and processes it; data consumer 2 obtains identification task @ and sets a queue lock for its corresponding message queue 2, and the count value is incremented by 1 here, and the count value is 1 at this time. Data consumer 3 obtains identification task @ and sets a queue lock for its corresponding message queue 3, and the count value is incremented by 1 again here, and the count value is 2 at this time. In this way, when the count value is the total number of message queues in the current message partition, in this example, the total number of message queues is 3, that is, when the count value is 3, the queue locks of each message queue in the message partition are released.
[0075] The embodiment of the present application provides a prerequisite data task, and the prerequisite data task is a DDL data definition language type task. That is to say, the prerequisite data task is a DDL data definition language type task, and the conditioned data task can be a DML data manipulation language.
[0076] Based on the above process, the embodiment of the present application provides a process of a data processing method, as Figure 4 shown, including:
[0077] Step 401: The data producer determines the distribution order of each task. Among them, identification tasks are set before and after the data task that is a prerequisite. The identification task is used to instruct the data consumer to stop obtaining tasks from the message queue corresponding to the data consumer when the task identifier is obtained.
[0078] Step 402: The data producer distributes each task to the corresponding message queue in the message partition in sequence according to the distribution order. Among them, the identification task will be distributed to each message queue in the message partition; each message queue corresponds to a data consumer.
[0079] In the above method, identification tasks are inserted before and after the data task that is a prerequisite. When distributing tasks in order, the identification tasks are sent to all message queues. In this way, when there is a situation where a data task that is a conditional data task is processed before the data task that is a prerequisite data task; before processing the conditional data task, other message queues are locked through the identification task; wait until the data consumer where the prerequisite data task is located obtains the identification task and is also locked; then release the queue locks of all locked message queues. In this way, it is ensured that the prerequisite data task is processed first. The accuracy of task processing is increased.
[0080] An embodiment of the present application provides a data task distribution method. The data producer distributes each task to the corresponding message queue in the message partition in sequence according to the distribution order, including: the data producer distributes each task of the same type or the same primary key to the same message queue in sequence. That is to say, when the data producer distributes tasks in order, each task of the same type or the same primary key is distributed to the same message queue. In the previous example, if C1 and C2 are data tasks of the same type or the same primary key, C1 is to write the content of the first row and first column of Table 1 as D, and C2 is to write E in the first row and first column of Table 1; that is, the data of C2 is equivalent to updating the data of C1. If each task of the same type or the same primary key is not distributed to the same message queue, resulting in C2 being processed before C1, then the final result of the content of the first row and first column of Table 1 will be the old result, affecting data accuracy. Therefore, distributing each task of the same type or the same primary key to the same message queue can ensure the processing order of each task of the same type or the same primary key and ensure the correctness of the final result.
[0081] Based on the above process, an embodiment of the present application provides a process of a data processing method, as Figure 5 shown, including:
[0082] Step 501: The data producer generates a plurality of data tasks with order.
[0083] Step 502: The data producer determines that among multiple data tasks with an order, there is a data task that serves as a prerequisite.
[0084] Step 503: The data producer inserts identification tasks before and after the data task that serves as a prerequisite to obtain multiple tasks with an order.
[0085] Step 504: The data producer distributes tasks in the order of the multiple tasks, distributing tasks of the same type or with the same primary key to the same message queue; when the distributed task is an identification task, the identification task is sent to all message queues in the message partition.
[0086] Step 505: The data consumer obtains tasks from the corresponding message queue.
[0087] Step 506: The data consumer determines the type of the obtained task. When the task is an identification task, step 507 is executed; when the task is a data task, step 509 is executed.
[0088] Step 507: The data consumer sets a queue lock for the message queue, increments the count value by 1, and stops obtaining tasks from the message queue in the message partition.
[0089] Step 508: When it is determined that the count value is equal to the number of current message queues, unlock each message queue; when the data consumer determines that all data consumers have stopped obtaining tasks from the message partition, it starts obtaining tasks from the message queue corresponding to the data consumer again.
[0090] Step 509: The data consumer processes the data task.
[0091] Step 510: The processing of the multiple tasks is completed.
[0092] Based on the same concept, an embodiment of the present invention provides a data processing device. Figure 6 The following is a schematic diagram of a data processing device provided by an embodiment of the present application, as Figure 6 shown, including:
[0093] An obtaining module 601, configured to obtain tasks from the message queue corresponding to the data consumer; the message queue corresponding to the data consumer is one of multiple message queues in the message partition, and each message queue corresponds to a data consumer; each data consumer executes concurrently;
[0094] A determining module 602, configured to stop obtaining tasks from the message queue corresponding to the data consumer when it is determined that the task is an identification task; the identification task is set before and after the data task that serves as a prerequisite in the distribution order when distributed by the data producer, and is written into each message queue in the message partition;
[0095] The obtaining module 601 is further configured to, when the data consumers determine that all data consumers have stopped obtaining tasks from the message partition, obtain tasks from the message queue corresponding to the data consumers again.
[0096] Optionally, the determining module 602 is specifically configured to, when determining that the task is an identification task, set a queue lock for the message queue corresponding to the data consumer; the queue lock is used to instruct the data consumer to stop obtaining tasks from the message queue corresponding to the data consumer;
[0097] The determining module 602 is further configured to determine that all message queues in the message partition are in a locked state.
[0098] Optionally, the determining module 602 is further configured to increase the count value of the queue lock;
[0099] The determining module 602 is further configured to determine that the count value meets the number of the message queues.
[0100] Optionally, the data task serving as a prerequisite is a DDL (Data Definition Language) task.
[0101] Based on the same concept, an embodiment of the present invention provides a data processing device, Figure 7 which is a schematic diagram of a data processing device provided in an embodiment of the present application, as Figure 7 shown, and includes:
[0102] A determining module 701, configured to determine the distribution order of each task, wherein identification tasks are arranged before and after the data task serving as a prerequisite; the identification task is used to instruct the data consumer to stop obtaining tasks from the message queue corresponding to the data consumer when obtaining the task identifier;
[0103] A distributing module 702, configured to sequentially distribute each task to the corresponding message queue in the message partition according to the distribution order; wherein, the identification task is distributed to each message queue in the message partition; each message queue corresponds to a data consumer.
[0104] Optionally, the distributing module 702 is specifically configured to sequentially distribute each task of the same type or the same primary key to the same message queue according to the distribution order.
[0105] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0106] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0109] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A data processing method, characterized in that, it includes: A data consumer obtains tasks from the message queue corresponding to the data consumer; the message queue corresponding to the data consumer is one of multiple message queues in a message partition, and each message queue corresponds to a data consumer; each data consumer executes concurrently; When the data consumer determines that the task is an identification task, it stops obtaining tasks from the message queue corresponding to the data consumer; when the identification task is distributed by the data producer, it is set before and after the data task as a prerequisite according to the distribution order, and is written into each message queue in the message partition; When the data consumer determines that all data consumers have stopped obtaining tasks from the message partition, it resumes obtaining tasks from the message queue corresponding to the data consumer.
2. The method according to claim 1, characterized in that, When the data consumer determines that the task is an identification task and stops obtaining tasks from the message queue corresponding to the data consumer, it includes: When the data consumer determines that the task is an identification task, it sets a queue lock for the message queue corresponding to the data consumer; the queue lock is used to indicate that the data consumer stops obtaining tasks from the message queue corresponding to the data consumer; When the data consumer determines that all data consumers have stopped obtaining tasks from the message partition, it includes: The data consumer determines that each message queue in the message partition is in a locked state.
3. The method according to claim 2, characterized in that, After setting the queue lock for the message queue corresponding to the data consumer, it further includes: The data consumer increases the count value of the queue lock; When the data consumer determines that each message queue in the message partition is in a locked state, it includes: The data consumer determines that the count value meets the number of each message queue.
4. The method according to any one of claims 1-3, characterized in that, The data task as a prerequisite is a DDL (Data Definition Language) task.
5. A data processing method, characterized in that, it includes: The data producer determines the distribution order of each task, wherein identification tasks are set before and after the data task as a prerequisite; the identification task is used to indicate that when the data consumer obtains the task identifier, it stops obtaining tasks from the message queue corresponding to the data consumer; The data producer distributes each task to the corresponding message queue in the message partition in sequence according to the distribution order; wherein, the identification task is distributed to each message queue in the message partition; each message queue corresponds to a data consumer.
6. The method according to claim 5, characterized in that, When the data producer distributes each task to the corresponding message queue in the message partition in sequence according to the distribution order, it includes: The data producer distributes each task of the same type or the same primary key to the same message queue in sequence according to the distribution order.
7. A data processing device, characterized in that, it includes: An acquisition module, configured to acquire tasks from the message queue corresponding to the data consumer; the message queue corresponding to the data consumer is one of multiple message queues in the message partition, and each message queue corresponds to a data consumer; each data consumer executes concurrently; A determination module, configured to stop acquiring tasks from the message queue corresponding to the data consumer when determining that the task is an identification task; the identification task is set before and after the data task as a prerequisite according to the distribution order when distributed by the data producer, and is written into each message queue in the message partition; The acquisition module is further configured to, when the data consumer determines that all data consumers have stopped acquiring tasks from the message partition, acquire tasks from the message queue corresponding to the data consumer again.
8. A data processing device Characterized in that It includes: A determination module, configured to determine the distribution order of each task, wherein identification tasks are set before and after the data task as a prerequisite; the identification task is used to instruct the data consumer to stop acquiring tasks from the message queue corresponding to the data consumer when obtaining the task identifier; A distribution module, configured to sequentially distribute each task to the corresponding message queue in the message partition according to the distribution order; wherein, the identification task will be distributed to each message queue in the message partition; each message queue corresponds to a data consumer.
9. A computer-readable storage medium Characterized in that The computer-readable storage medium stores a program, and when the program runs on a computer, it causes the computer to implement the method described in any one of claims 1 to 4, 5 or 6.
10. A computer device Characterized in that It includes: A memory, configured to store a computer program; A processor, configured to call the computer program stored in the memory and execute the method described in any one of claims 1 to 4, 5 or 6 according to the obtained program.
Citation Information
Patent Citations
Method, system and computer program product for sequencing asynchronous messages in a distributed and parallel environment
CN104428754A
Method, system and computer program products for sequencing asynchronous messages in a distributed and parallel environment
EP2693337A1