A method and device for processing distributed data

By building task control tables and selecting highly available distributed processing strategies, the problems of low fault tolerance and poor resource management flexibility of distributed data processing systems when computing node exceptions are solved, and more efficient and flexible distributed data processing is achieved.

CN119557111BActive Publication Date: 2025-05-13CHINA SECURITIES DEPOSITORY & CLEARING CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510121886.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The existing distributed data processing system has low fault tolerance when computing node exceptions and poor resource management flexibility, resulting in system processing exceptions and waste of resources.

Method used

By obtaining multiple data shards of pending data, building a task for each data shard and storing it in a task control table, determining the target strategy from a variety of highly available distributed processing strategies, using distributed nodes and target strategies to preempt the target task state as pending, processing data shards through multiple target tasks, and using exception monitoring strategies to monitor the task operation status.

Benefits of technology

It improves the scenario adaptability, flexibility and scalability of distributed data processing, enhances processing efficiency and high availability, and reduces the risk of system handling exceptions and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557111B_ABST
    Figure CN119557111B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for processing distributed data, and relates to the field of big data technology. A specific implementation of the method includes: obtaining multiple data slices of data to be processed, constructing corresponding tasks for the data slices, and storing the tasks in a task control table; determining a target high-availability strategy that matches the data to be processed from a variety of high-availability distributed processing strategies; using distributed nodes and target high-availability strategies, preempting multiple target tasks in the task control table whose task status is to be processed, processing corresponding data slices through multiple target tasks, and using a monitoring strategy corresponding to the target high-availability strategy to monitor the task running status of distributed nodes and perform task processing operations. The embodiment of the present invention improves the scenario adaptability, flexibility and scalability of distributed data processing, and improves the efficiency and high availability of processing distributed data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular to a method and device for processing distributed data. Background Art

[0002] With the widespread use of Internet application systems, the requirements for system processing capabilities are becoming increasingly higher. For application scenarios that require streaming data processing (such as processing financial data), there are higher requirements for the real-time and reliability of the system's processing of streaming data.

[0003] Currently, a distributed structure is commonly used to process streaming data. It mainly connects multiple computing nodes through a network and controls the data processing of each computing node through a control center to complete the preset data processing tasks in parallel. The existing method will cause the entire system to process abnormally if any computing node has an abnormality, resulting in poor system fault tolerance. In addition, since the number of computing nodes is fixed, it leads to problems such as lack of resources or waste of resources and poor flexibility. Summary of the Invention

[0004] In view of this, an embodiment of the present invention provides a method and apparatus for processing distributed data, which can obtain multiple data shards of data to be processed, construct corresponding tasks for the data shards, and store the tasks in a task control table; determine a target high-availability strategy that matches the data to be processed from a variety of high-availability distributed processing strategies; utilize distributed nodes and target high-availability strategies to seize multiple target tasks in the task control table whose task status is to be processed, process the corresponding data shards through the multiple target tasks, and utilize the abnormal monitoring strategy corresponding to the target high-availability strategy to monitor the task operation status of the distributed nodes and perform task processing operations accordingly. The embodiment of the present invention improves the scenario adaptability, flexibility, and scalability of distributed data processing, and improves the efficiency and high availability of processing distributed data.

[0005] To achieve the above-mentioned purpose, according to one aspect of an embodiment of the present invention, a method for processing distributed data is provided, which is characterized by comprising: obtaining multiple data shards of the data to be processed, constructing a corresponding task for each of the data shards, and storing the tasks in a task control table; determining a target high-availability strategy that matches the data to be processed from a plurality of high-availability distributed processing strategies; utilizing distributed nodes and the target high-availability strategy to seize multiple target tasks with a task status of pending processing in the task control table, and processing the corresponding data shards through the multiple target tasks; utilizing the abnormal monitoring strategy corresponding to the target high-availability strategy to monitor the task running status of the distributed nodes and perform corresponding task processing operations according to the task running status.

[0006] Optionally, in the case where the target high availability strategy is a preemptive high availability strategy, the distributed nodes and the target high availability strategy are used to preempt the target tasks whose task status in the task control table is pending, including: for any distributed node, starting the main control thread of the node, and using the main control thread to generate multiple processing threads for processing tasks according to the number of shards of the data shards; using the main control thread to obtain each of the target tasks to be processed, and storing the target tasks to local storage, and performing a reordering operation on each target task in the local storage, so that multiple processing threads perform the operation of preempting the target tasks based on the reordered tasks.

[0007] Optionally, in the case where the target high availability strategy is a preemptive high availability strategy, the abnormal monitoring strategy corresponding to the target high availability strategy is used to monitor the task running status of the distributed nodes and perform corresponding task processing operations according to the task running status, including: registering multiple distributed nodes in the monitoring center; using the monitoring center to perform the following operations based on the abnormal monitoring strategy: monitoring each of the distributed nodes to create different heartbeat files at set time intervals; when it is monitored that the heartbeat file of any distributed node is not updated at a set number of set time intervals, determining that there is an abnormality in the task running status of the distributed node; for the distributed node with an abnormality in the task running status, obtaining the current task processed by the thread of the distributed node with the abnormality, and changing the task status of the current task to pending in the task control table.

[0008] Optionally, for the case where the target high availability strategy is a master-slave high availability strategy, the use of distributed nodes and the target high availability strategy to preempt the target task whose task status is pending in the task control table includes: determining multiple master nodes from multiple distributed nodes; obtaining the task identifier of the pending task and the number of master nodes; determining each pending task identifier processed by each master node based on a first calculation relationship between the number of nodes and the task identifier; using the main control thread of each master node to construct multiple processing threads for processing tasks, and using the processing threads to preempt the target task corresponding to the pending task identifier in the task control table.

[0009] Optionally, multiple master nodes are determined from the multiple distributed nodes, including: multiple distributed nodes connect to the distributed coordination system at startup, and create a unique directory identifier in the set storage location of the distributed coordination system, and the distributed node that successfully creates the unique target identifier is used as the master node; the abnormal monitoring strategy corresponding to the target high availability strategy is used to monitor the task running status of the distributed nodes and perform corresponding task processing operations according to the task running status, including: for the case where the target high availability strategy is a master-slave high availability strategy, the distributed coordination system is used to monitor the heartbeat data of multiple master nodes, and when it is determined that the heartbeat data of any master node is abnormal, the task status of the master node with the abnormality is set to pending, a new master node is selected from multiple backup nodes, and the new master node is used to replace the master node with the abnormality to continue executing the pending target task corresponding to the master node with the abnormality.

[0010] Optionally, in the case where the target high availability strategy is a multi-active high availability strategy, the use of distributed nodes and the target high availability strategy to preempt the target task whose task status is pending in the task control table includes: dividing the multiple distributed nodes into multiple node groups; obtaining the task identifier of the pending task and the group number of the node group; determining the identifiers of each pending task processed by each node group based on a second calculation relationship between the task identifier and the group number; for each node group, selecting a master node from the distributed nodes included in the node group, using the main control thread of the master node to construct multiple processing threads for processing tasks, and using the processing thread to preempt the target task corresponding to the task identifier in the task control table.

[0011] Optionally, after selecting the master node from the distributed nodes included in each of the node groups, it further includes: determining one or more standby nodes in each of the node groups; for each of the standby nodes, executing a processing thread for constructing a processing task, and using the processing thread to preempt the target task corresponding to the task identifier in the task control table; wherein the processing progress of the target task of the standby node is later than the processing progress of the target task of the master node; using the abnormality monitoring strategy corresponding to the target high availability strategy to monitor the task running status of the distributed nodes and perform corresponding task processing operations according to the task running status, including: for the case where the target high availability strategy is a multi-active high availability strategy, using a distributed coordination system to monitor the heartbeat data of the master node in each node group, and when it is determined that there is an abnormality in the heartbeat data of the master node in any node group, determining that there is an abnormality in the task running status of the master node of the node group; determining a new master node from the standby nodes of the node group; and using the new master node to continue to execute the unfinished tasks of the master node, and outputting the result data to the downstream based on the data output breakpoint where the master node abnormality occurs.

[0012] Optionally, the processing of the corresponding data shards by the target task further includes: using distributed nodes to write the processing results and the first processing progress obtained by processing the data of multiple data shards into a message queue, and storing the processing results in a database; when it is determined that there is a storage abnormality in the process of storing the processing results in the database, determining the second processing progress before the processing results are stored according to the data source; according to the second processing progress, determining the unprocessed data after the second processing progress, storing the processing results obtained by processing the unprocessed data in the database, and updating the second processing progress; when the updated second processing progress is consistent with the first processing progress, determining that the data processing meets transaction consistency.

[0013] Optionally, the use of distributed nodes to write the processing results obtained by processing the data of multiple data shards into the message queue further includes: after any distributed node determines that the writing of the processing results into the message queue is completed, broadcasting a write task completion message to each downstream node; wherein, the task completion message includes the identifier of each distributed node that performs the write operation; after the downstream node determines that the upstream distributed node has completed writing the data to the message queue after receiving the task completion message sent by each distributed node identifier.

[0014] To achieve the above-mentioned purpose, according to a second aspect of an embodiment of the present invention, there is provided an apparatus for processing distributed data, comprising: a task generation module, a task preemption module, and a status determination module; wherein,

[0015] The task generation module is used to obtain multiple data slices of the data to be processed, construct a corresponding task for each of the data slices, and store the task in the task control table;

[0016] The preemptive task module is used to determine a target high-availability strategy that matches the data to be processed from a plurality of high-availability distributed processing strategies; using the distributed nodes and the target high-availability strategy, preempt multiple target tasks in the task control table whose task status is pending, and process the corresponding data shards through the multiple target tasks;

[0017] The state determination module is used for the state determination module to use the abnormal monitoring strategy corresponding to the target high availability strategy to monitor the task running status of the distributed node and perform corresponding task processing operations according to the task running status.

[0018] Optionally, in the case where the target high availability strategy is a preemptive high availability strategy, the device for processing distributed data is used to use distributed nodes and the target high availability strategy to preempt the target task whose task status is pending in the task control table, including: for any distributed node, starting the main control thread of the node, using the main control thread to generate multiple processing threads for processing tasks according to the number of shards of the data shards; using the main control thread to obtain each of the target tasks to be processed, and storing the target tasks to local storage, and performing a reordering operation on each target task in the local storage, so that multiple processing threads perform the operation of preempting the target task based on the reordered tasks.

[0019] Optionally, in the case where the target high availability strategy is a preemptive high availability strategy, the device for processing distributed data is used to utilize the abnormal monitoring strategy corresponding to the target high availability strategy to monitor the task running status of the distributed nodes and perform corresponding task processing operations according to the task running status, including: registering multiple distributed nodes in a monitoring center; utilizing the monitoring center to perform the following operations based on the abnormal monitoring strategy: monitoring each of the distributed nodes to create different heartbeat files at set time intervals; when it is monitored that the heartbeat file of any distributed node is not updated at a set number of set time intervals, determining that there is an abnormality in the task running status of the distributed node; for the distributed node with an abnormality in the task running status, obtaining the current task processed by the thread of the distributed node with the abnormality, and changing the task status of the current task to pending in the task control table.

[0020] Optionally, for the case where the target high availability strategy is a master-slave high availability strategy, the device for processing distributed data is used to use distributed nodes and the target high availability strategy to preempt the target task whose task status is pending in the task control table, including: determining multiple master nodes from multiple distributed nodes; obtaining the task identifier of the pending task and the number of master nodes; determining each pending task identifier processed by each master node based on a first calculation relationship between the number of nodes and the task identifier; using the main control thread of each master node to construct multiple processing threads for processing tasks, and using the processing threads to preempt the target task corresponding to the pending task identifier in the task control table.

[0021] Optionally, the device for processing distributed data is used to determine multiple master nodes from multiple distributed nodes, including: multiple distributed nodes connect to the distributed coordination system at startup, and create a unique directory identifier in the set storage location of the distributed coordination system, and use the distributed node that successfully creates the unique target identifier as the master node; the abnormal monitoring strategy corresponding to the target high availability strategy is used to monitor the task running status of the distributed nodes and perform corresponding task processing operations according to the task running status, including: for the case where the target high availability strategy is a master-slave high availability strategy, the distributed coordination system is used to monitor the heartbeat data of multiple master nodes, and when it is determined that the heartbeat data of any master node is abnormal, the task status of the master node with the abnormality is set to pending, a new master node is selected from multiple backup nodes, and the new master node is used to replace the master node with the abnormality to continue executing the pending target task corresponding to the master node with the abnormality.

[0022] Optionally, in the case where the target high availability strategy is a multi-active high availability strategy, the device for processing distributed data is used to use distributed nodes and the target high availability strategy to preempt the target task whose task status is pending in the task control table, including: dividing multiple distributed nodes into multiple node groups; obtaining the task identifier of the pending task and the group number of the node group; determining each pending task identifier processed by each node group based on a second calculation relationship between the task identifier and the group number; for each node group, selecting a master node from the distributed nodes included in the node group, using the main control thread of the master node to construct multiple processing threads for processing tasks, and using the processing thread to preempt the target task corresponding to the task identifier in the task control table.

[0023] Optionally, the device for processing distributed data, after selecting a master node from the distributed nodes included in each of the node groups, further includes: determining one or more standby nodes in each of the node groups; for each of the standby nodes, executing a processing thread for constructing a processing task, and using the processing thread to preempt the target task corresponding to the task identifier in the task control table; wherein the processing progress of the target task of the standby node is later than the processing progress of the target task of the master node; using the abnormality monitoring strategy corresponding to the target high availability strategy to monitor the task running status of the distributed nodes and perform corresponding task processing operations according to the task running status, including: for the case where the target high availability strategy is a multi-active high availability strategy, using a distributed coordination system to monitor the heartbeat data of the master node in each node group, and when it is determined that there is an abnormality in the heartbeat data of the master node in any node group, determining that there is an abnormality in the task running status of the master node of the node group; determining a new master node from the standby nodes of the node group; and using the new master node to continue to execute the unfinished tasks of the master node, and outputting result data to the downstream based on the data output breakpoint where the master node abnormality occurs.

[0024] Optionally, the device for processing distributed data, used to process the corresponding data shards through the target task, further includes: using distributed nodes to write the processing results and the first processing progress obtained by processing the data of multiple data shards into the message queue, and storing the processing results in the database; when it is determined that there is a storage abnormality in the process of storing the processing results in the database, determining the second processing progress before the processing results are stored according to the data source; according to the second processing progress, determining the unprocessed data after the second processing progress, storing the processing results obtained by processing the unprocessed data in the database, and updating the second processing progress; when the updated second processing progress is consistent with the first processing progress, determining that the data processing meets transaction consistency.

[0025] Optionally, the device for processing distributed data is used to use distributed nodes to write the processing results obtained by processing the data of multiple data shards into a message queue, and further includes: after any distributed node determines that the writing of the processing results into the message queue is completed, broadcasting a write task completion message to each downstream node; wherein, the task completion message includes the identifier of each distributed node that performs the write operation; after the downstream node determines that the upstream distributed node has completed writing the data to the message queue after receiving the task completion message sent by each distributed node identifier.

[0026] To achieve the above-mentioned purpose, according to the third aspect of an embodiment of the present invention, an electronic device for processing distributed data is provided, characterized in that it includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement any of the methods described above for processing distributed data.

[0027] To achieve the above-mentioned purpose, according to a fourth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, characterized in that when the program is executed by a processor, any of the methods described above for processing distributed data is implemented.

[0028] To achieve the above-mentioned purpose, according to a fifth aspect of an embodiment of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned methods for processing distributed data.

[0029] One embodiment of the above invention has the following advantages or beneficial effects: it is able to obtain multiple data shards of data to be processed, construct corresponding tasks for the data shards, and store the tasks in a task control table; determine a target high-availability strategy that matches the data to be processed from a variety of high-availability distributed processing strategies; utilize distributed nodes and target high-availability strategies to seize multiple target tasks in the task control table whose task status is to be processed, process the corresponding data shards through multiple target tasks, and utilize the abnormal monitoring strategy corresponding to the target high-availability strategy to monitor the task running status of the distributed nodes and perform task processing operations accordingly. The embodiment of the present invention improves the scenario adaptability, flexibility, and scalability of distributed data processing through different high-availability distributed processing strategies; and improves the efficiency, reliability, fault tolerance, and high availability of processing distributed data.

[0030] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings are provided for a better understanding of the present invention and are not intended to limit the present invention.

[0032] Figure 1 This is a flow chart of a method for processing distributed data provided by one embodiment of the present invention;

[0033] Figure 2A This is a schematic diagram of a process for generating a task control table provided by an embodiment of the present invention;

[0034] FIG2B is a schematic diagram of a task status change provided by an embodiment of the present invention;

[0035] Figure 3A This is a flowchart of a process for processing tasks using a preemptive high availability strategy provided by an embodiment of the present invention;

[0036] Figure 3B This is a flowchart of processing tasks in a preemptive high availability strategy provided by an embodiment of the present invention;

[0037] Figure 3C This is another flowchart of handling exceptions in a preemptive high availability strategy provided by an embodiment of the present invention;

[0038] Figure 3D This is another flowchart of handling exceptions in a preemptive high availability strategy provided by an embodiment of the present invention;

[0039] Figure 3E This is a flowchart of a process for processing tasks using a master-slave high-availability strategy provided by an embodiment of the present invention;

[0040] Figure 3F This is a flow chart of an embodiment of the present invention providing a method for handling exceptions using a master-slave high availability strategy;

[0041] Figure 3G This is a flowchart of a process for processing tasks using a multi-active high-availability strategy provided by an embodiment of the present invention;

[0042] Figure 3H This is a flowchart of another method for processing tasks using a multi-active high-availability strategy, provided by an embodiment of the present invention;

[0043] Figure 3I This is a flow chart of a method for handling exceptions using a multi-active high-availability strategy, provided by an embodiment of the present invention;

[0044] Figure 3J This is a flow chart of a data fault tolerance method provided by one embodiment of the present invention;

[0045] Figure 3K This is a flow chart of another data fault tolerance method provided by one embodiment of the present invention;

[0046] Figure 3L This is a schematic diagram of a process for processing a task completion message provided by an embodiment of the present invention;

[0047] Figure 4 This is a schematic structural diagram of a device for processing distributed data provided by one embodiment of the present invention;

[0048] FIG5 is a diagram of an exemplary system architecture in which embodiments of the present invention may be applied;

[0049] Figure 6 It is a schematic diagram of the structure of a computer system of a terminal device or server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, in which various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0051] It should be noted that in the technical solution of the present invention, the collection, use, storage, sharing and transfer of user personal information in the financial data involved comply with the provisions of relevant laws and regulations, and require notification to the user and obtaining the user's consent or authorization. When applicable, the user's personal information is de-identified and / or anonymized and / or encrypted.

[0052] Distributed data processing systems are used in diverse scenarios, with varying data sources, formats, and stability. To meet the data processing needs of these scenarios, different programs often need to be developed. To ensure system fault tolerance in distributed scenarios, high-availability mechanisms are often required when developing distributed data processing programs. High-availability scenarios for different data scenarios must also be addressed separately. Developing multiple corresponding programs increases development costs and complicates subsequent maintenance and management.

[0053] In view of this, if Figure 1 As shown, an embodiment of the present invention provides a method for processing distributed data, which may include the following steps:

[0054] Step S101: Acquire multiple data slices of the data to be processed, construct a corresponding task for each of the data slices, and store the task in a task control table.

[0055] Specifically, the data to be processed (for example, financial data) often has a large amount of data. Therefore, before processing the data to be processed, the operation of splitting the data to be processed into multiple data shards is performed. Specifically, the data sharding information for the data can be configured in the data table of the database (for example, the sharding definition table), or the data sharding information for the data shards can be configured through the configuration file; for the scenario of processing data in the database, the data shards can be logical shards of the data table or physical shards provided by the database; for the scenario of processing data in the message queue, the data shards can be shards defined when the message queue creates a message topic; the present invention improves the flexibility and scalability of distributed data processing by configuring the information of executing data sharding for different types of data sources (databases or message queues), and obtaining multiple data shards of the data to be processed by using the configuration information.

[0056] Further, if Figure 2A As shown, in an embodiment of the present invention, an initialization program is called upon system startup to construct a corresponding task for each data shard, and the task is stored in a task control table. The task control table may include data shard information corresponding to the data shard (e.g., a shard definition table) and a task identifier for the task that processes the data shard. That is, a corresponding task is constructed for each data shard, and the task is stored in the task control table. It is understood that the task status of each task in the task control table before processing is "pending."

[0057] like Figure 2B As shown, the life cycle of a task from creation, running to completion mainly includes the following task states: pending, paused, running, completed, and abnormally ended; among them, "pending" means that the task has not started to be processed; "paused" means that the task is prohibited from being preempted, or the preempted task is running and the processing thread suspends running after receiving a suspend instruction; "running" means that the task has been preempted and an idle thread is created / reused for processing; if there is no idle thread, it waits for execution in the blocking queue; "completed" means that the task has been preempted and the task has been successfully executed; or the preempted task is running and the task has been completed after receiving a forced completion instruction; or the task has not been preempted, the task has not started processing, and the task has been completed after receiving a forced completion instruction; "abnormally ended" means that the task has been preempted and the task has been completed abnormally; or the task has been preempted and the task is running and the task has been completed after receiving a cancel execution instruction; or the task has not been preempted, the task has not started processing, and the task has been completed after receiving a cancel execution instruction; Figure 2B Shows the flow relationship between different task states.

[0058] Furthermore, an embodiment of the present invention provides a control instruction table for controlling the task status of a task. After the system is started, the main control thread is used to scan the corresponding control instruction table at a set time interval to obtain the task status control instruction. The main control thread is used to load the common business static parameter data in the process, manage the task threads, perform heartbeat processing, etc. The life cycle of the main control thread runs through the entire business processing process. The task thread is used to respond to data processing instructions and execute business program logic. The task thread is created by the main control thread and ends after the thread itself is executed. A process is the smallest unit of resource allocation, and a thread is the actual running unit in a process. A program forms a process after startup. If the process contains multiple threads, the program is a multi-threaded program.

[0059] Furthermore, control instructions are categorized as follows: task pause, task resume, forced task completion, and task cancel. After the main control thread obtains the control instruction from the database, it will control the currently executing task thread. The corresponding task thread will change state after processing the batch of data currently being processed. Specifically, the corresponding impact of task state control instructions on task states is shown in Table 1:

[0060]

[0061] The embodiment of the present invention constructs a task control table for the task of processing the data slice, thereby utilizing the task control table to manage the task status during the processing of the task, thereby improving the accuracy, flexibility and efficiency of task management.

[0062] Step S102: Determine a target high availability strategy that matches the data to be processed from a plurality of high availability distributed processing strategies; utilize distributed nodes and the target high availability strategy to seize multiple target tasks whose task status is to be processed in the task control table, and process the corresponding data shards through the multiple target tasks.

[0063] Step S103: using the abnormal monitoring strategy corresponding to the target high availability strategy, monitoring the task running status of the distributed nodes and performing corresponding task processing operations according to the task running status.

[0064] Specifically, in an embodiment of the present invention, a variety of high-availability distributed processing strategies are provided, including: preemptive high-availability strategy, active-standby high-availability strategy, multi-active high-availability strategy, etc.; for a system, the target high-availability strategy to be used can be configured through a configuration file, that is, the target high-availability strategy can be any of the high-availability strategies such as preemptive high-availability strategy, active-standby high-availability strategy, multi-active high-availability strategy, etc. Different high-availability strategies can be configured for different application scenarios and application systems without changing hardware devices such as distributed nodes (i.e., different high-availability strategies reuse existing system hardware), thereby improving the scalability, flexibility, and versatility of the system in processing distributed data.

[0065] Furthermore, different high-availability distributed processing strategies have corresponding exception monitoring strategies. Using the exception monitoring strategy corresponding to the target high-availability strategy, the task execution status of distributed nodes is monitored and corresponding task processing operations are performed based on the task execution status. This improves the data reliability and high availability of distributed data processing. This overcomes the problem of existing methods where an exception in any computing node causes an exception in the entire system processing, resulting in poor system fault tolerance.

[0066] In the case of a preemptive high availability strategy, each process is equivalent and can preempt tasks with a pending status according to the task control table. In each distributed node, a master control thread is started after system startup. The master control thread creates threads based on the number of processing threads (see Formula 1). Each thread preempts tasks, ensuring balanced preemption across all business processes (each process is evenly allocated tasks).

[0067] Formula 1: Number of processing threads = min (number of data shards / total number of business processes, maximum number of configured threads), where min represents the minimum of the two values.

[0068] Figure 3A FIG. 1 shows a flow chart of a process for processing a task using a preemptive high availability strategy provided by an embodiment of the present invention; FIG. Figure 3A As shown in the figure, the distributed nodes include node 1 and node 2, and the task status table includes 6 tasks. The task identifiers (task names) of the 6 tasks are 1 to 6 respectively, and the task status of each task is "pending". Assuming that the number of all processes of node 1 and node 2 is 2, and the maximum number of threads is 5 (where the maximum number of threads is configurable), min (6 / 2, 5) is calculated according to formula 1 to get 3, that is, the distributed node identifier is the main control thread of node 1 or node 2 (that is, Figure 3A The main thread in the creates three processing threads ( Figure 3AjobA-Thread-1, jobA-Thread-2, jobA-Thread-3 in .

[0069] Furthermore, in the preemptive high availability strategy, when executing the preemptive task operation, the corresponding task record can be locked through the database lock mechanism. The first one to be preempted will be processed first. In order to avoid all distributed nodes (single instance after the distributed deployment of this system) grabbing the same task and causing performance degradation, the main control program of the main control thread of the distributed node can obtain all pending task data to local storage (such as memory), and then use the shuffling algorithm to rearrange the tasks, and then each processing thread will preempt them (so that multiple processing threads can execute the operation of preempting the target task based on the reordered tasks), such as Figure 3A As shown in the figure, for the task order of "123456", the order obtained by node 1 after reordering (shuffling algorithm) is "125463", and the order obtained by node 2 after reordering (shuffling algorithm) is "534126". By reordering and preempting the tasks of each distributed node, the efficiency of preemption and the uniform distribution of processing data are improved.

[0070] Figure 3B FIG. 1 shows a flow chart of processing tasks in a preemptive high availability strategy provided by an embodiment of the present invention; FIG. Figure 3B As shown, three processing threads each preempt three pending tasks (i.e., target tasks) from all six pending tasks. These tasks are processed by multiple processing threads. After each processing thread preempts a task, it updates the status in the task control table from "pending" to "processing" (or running). After completing the current task's data, it proceeds to the next task. If a processing process fails, another process can preempt the task, obtain the original program's processing progress, and continue computing and outputting. That is, in the case where the target high availability strategy is a preemptive high availability strategy, the distributed nodes and the target high availability strategy are used to preempt the target tasks whose task status is pending in the task control table, including: for any distributed node, starting the main control thread of the node, and using the main control thread to generate multiple processing threads for processing tasks according to the number of shards of the data shards; using the main control thread to obtain each of the target tasks to be processed, and storing the target tasks to local storage; performing a reordering operation on each target task in the local storage, so that multiple processing threads perform the operation of preempting the target task based on the reordered tasks.

[0071] Furthermore, in an embodiment of the present invention, a corresponding fault tolerance method is provided for the preemptive high availability strategy. Figure 3C This is another flowchart of handling exceptions in a preemptive high availability strategy provided by an embodiment of the present invention; specifically, Figure 3C As shown, an embodiment of the present invention can set up a monitoring center, and use the monitoring center to monitor each distributed node. Each distributed node constructs a heartbeat file in the device to which it belongs according to a set time interval (the time point and content of the heartbeat file constructed each time are different); further, a separately deployed monitoring center is used to check the heartbeat file. If it is found that the heartbeat file of any distributed node (assuming it is node 2) has not been updated within a set number (for example, 4) of heartbeat intervals (set time intervals), the monitoring center remotely checks the process identifier (PID) of the distributed node to determine whether the process of node 2 is alive. If it is not alive, it is determined that there is an abnormality in node 2, and the monitoring center is used to change the status of the tasks being processed by node 2 from "running" to "pending", so that other nodes can preempt the "pending" tasks by scanning the task control table. After other processes scan the tasks to be processed, they will create new task threads for processing. When the number of threads reaches the configured maximum number of threads, it will stop. When a task is completed and the thread is released, it will continue to look for pending tasks to preempt; as Figure 3D Another flow chart of handling exceptions in a preemptive high availability strategy provided by an embodiment of the present invention is shown; Figure 3D As shown, after node 2 fails and exits, node 1 takes over the two tasks 3 and 4 executed by node 2 and creates task threads for processing the corresponding data shards for tasks 3 and 4. Since the number of threads on node 1 reaches the maximum number of threads (for example, 5), node 1 suspends preemptive threads and waits for tasks to complete before preempting task 5 for processing. That is, when the target high availability strategy is a preemptive high availability strategy, the abnormality monitoring strategy corresponding to the target high availability strategy is used to monitor the task execution status of distributed nodes and perform corresponding task processing operations based on the task execution status, including: registering multiple distributed nodes with a monitoring center; using the monitoring center to perform the following operations based on the abnormality monitoring strategy: monitoring each distributed node to create different heartbeat files at set time intervals; if the heartbeat file of any distributed node is not updated within a set number of set time intervals, determining that the task execution status of the distributed node is abnormal; for the distributed node with the abnormal task execution status, obtaining the current task processed by the thread of the abnormal distributed node and changing the task status of the current task to pending in the task control table.

[0072] In an embodiment of the present invention, for the preemptive high availability strategy, when a distributed node fails to write a heartbeat file, but the node is still processing data tasks, the monitoring center considers that the node is faulty. Further, in order to prevent two nodes from processing a data shard at the same time, for a certain distributed node, it can be set that after the number of heartbeat write failures reaches the number specified in the configuration file, the node process will actively terminate and exit.

[0073] When the target high availability strategy is a master-slave high availability strategy, the distributed nodes will be divided into master nodes and slave nodes when they are started. The master node processes data, and the slave node is in a standby state. If an exception (failure) occurs in the master node, a new master node can be selected from the slave nodes to replace the abnormal (failed) master node to continue data processing.

[0074] Furthermore, a corresponding number of master nodes M can be configured in the configuration file for a master-standby high-availability strategy, and a number of processes greater than M can be deployed and started during system deployment. The number of started processes is used to compete for master status in a distributed coordination system (e.g., Zookeeper). Upon system startup, all distributed nodes connect to the distributed coordination system and attempt to create subpaths numbered N (1 <= N <= M) under a specific path. Successful creation indicates that the distributed node becomes master node N. If creation fails and the number of subpaths has not reached M, the distributed node continues to attempt. If creation fails and the number of subpaths reaches M, the distributed node automatically becomes a backup node and simultaneously starts a listener to monitor all subnodes under this path. Upon detecting a path deletion, the distributed node attempts to create a path again to become the master node. Specifically, determining multiple master nodes from the multiple distributed nodes involves: upon startup, the multiple distributed nodes connect to the distributed coordination system and create a unique directory identifier in a designated storage location within the distributed coordination system. The distributed node that successfully creates the unique target identifier is designated as the master node.

[0075] Figure 3E FIG. 1 shows a flow chart of a process for processing a task using a master-slave high availability strategy according to an embodiment of the present invention; FIG. Figure 3E As shown in the example, for example, if the number of configured master nodes is 2 and the number of processes started is 3, each process connects to the distributed coordination system to compete for mastership (master acquisition). Successful masterships are represented by master nodes 1 and 2, while those that fail become standby nodes. Furthermore, in an active-standby high-availability strategy, if the master node fails, the standby node takes over. Furthermore, each master node calculates the target tasks it needs to process based on its node ID and creates a corresponding number of processing threads to process the data shards corresponding to the tasks. The master node uses the following formula to calculate the tasks to be processed and their IDs.

[0076] Calculation formula (i.e., the first calculation relationship): N%M=(node ​​ID - 1) task N. Here, the task ID is N and the total number of nodes is M.

[0077] exist Figure 3E In the diagram shown, the node IDs are "1" and "2." For example, through calculation, node 2 identifies pending tasks as "1, 3, and 5," while node 1 identifies pending tasks as "2, 4, and 6." The standby node is in a standby state (not processing tasks). The master node then further executes multiple processing threads constructed using the master control thread of each master node to process tasks, and uses these processing threads to preempt the target tasks corresponding to the task identifiers in the task control table. That is, for the case where the target high availability strategy is a master-slave high availability strategy, the use of distributed nodes and the target high availability strategy to preempt the target task whose task status is pending in the task control table includes: determining multiple master nodes from multiple distributed nodes; obtaining the task identifier of the task to be processed and the number of master nodes; determining the processing task of each master node and the task identifier corresponding to the processing task based on a first calculation relationship between the number of nodes and the task identifier, and determining the processing task of each master node and the task identifier corresponding to the processing task; using the main control thread of each master node to construct multiple processing threads for processing tasks, and using the processing thread to preempt the target task corresponding to the task identifier in the task control table.

[0078] Furthermore, Figure 3F FIG. 1 shows a flow chart of an embodiment of the present invention providing a method for handling an exception by using a master-slave high availability strategy; FIG. Figure 3F As shown, upon detecting an abnormal heartbeat from any master node (node ​​1), the distributed coordination system determines that the master node has experienced an anomaly (failure), deletes the path information corresponding to the master node, and notifies all monitoring backup nodes, causing each backup node to retry attempting to become the master node. After selecting a new master node (i.e., new node 1), the new master node replaces the failed master node and continues processing the corresponding task. Specifically, the system utilizes the abnormality monitoring strategy corresponding to the target high availability strategy to monitor the task execution status of distributed nodes and executes corresponding task processing operations based on the task execution status, including: utilizing the distributed coordination system to monitor the heartbeat data of multiple master nodes; upon determining that the heartbeat data of any master node is abnormal, setting the task status of the abnormal master node to pending processing; selecting a new master node from the multiple backup nodes; replacing the abnormal master node with the new master node; and continuing to execute the pending target task corresponding to the abnormal master node.

[0079] In an embodiment of the present invention, under the master-slave high availability strategy, in order to prevent multiple master nodes from processing the same data shard, all master nodes will monitor the connection with the distributed coordination system (Zookeeper). When the master node finds that the connection between itself and Zookeeper is interrupted, it will actively terminate and exit its own process.

[0080] In the case where the target high availability strategy is a multi-active high availability strategy, a node group ID (group ID) is configured for each node when it is started. Only one master node is selected for the process with the same node group ID, and the other nodes are all standby nodes. That is, the standby node of each node group can only serve as the standby node of the master node of its own group. The method for becoming a master node in each node group can also be obtained through a distributed coordination system. The specific method is consistent with the method for becoming a master node described in the target high availability strategy and will not be repeated here. Furthermore, the business node is divided into a master node and a standby node at startup. The master node and the standby node perform data processing at the same time, but the processing results of the standby node are not output to the outside. If the master node fails, a new master node is selected from the standby nodes to replace the failed node to continue data processing.

[0081] Furthermore, for the multi-active high availability strategy, the tasks (data shards) processed by each node group are the same, and the allocation of tasks (data shards) between different node groups is similar to the method of the active-standby high availability strategy. Specifically, the task identifier (task ID) is taken as the remainder of the task identifier and the total number of groups. The task identifier obtained by the remainder being equal to group ID-1 (i.e., the second calculation relationship) is determined to be the task processed by the node group.

[0082] Figure 3G FIG. 1 shows a flow chart of processing a task using a multi-active high-availability strategy provided by an embodiment of the present invention; FIG. Figure 3G As shown, there are two node groups: group 1 and group 2. Group 1 includes the master node and the backup node, and group 2 includes the master node and the backup node. Through the second calculation relationship between the task identifier and the number of groups, for example, it is determined that the pending task identifiers processed by group 2 are "1, 3, 5", and the pending task identifiers processed by group 1 are "2, 4, 6". Figure 3GIn the example, the group IDs are "1" and "2". That is, in the case where the target high availability strategy is a multi-active high availability strategy, the use of distributed nodes and the target high availability strategy to preempt the target task whose task status is pending in the task control table includes: dividing the multiple distributed nodes into multiple node groups; obtaining the task identifier of the task to be processed and the number of node groups; determining the identifiers of the tasks to be processed by each node group based on the second calculation relationship between the task identifier and the number of groups; for each node group, selecting a master node from the distributed nodes included in the node group, using the main control thread of the master node to construct multiple processing threads for processing tasks, and using the processing threads to preempt the target task corresponding to the task identifier in the task control table.

[0083] Figure 3H FIG. 4 shows another flow chart of processing tasks using a multi-active high-availability strategy provided by an embodiment of the present invention; FIG. Figure 3H As shown, the backup node in a node group performs the same operations on the data shards as the master node processes the corresponding tasks. The difference between the master and backup nodes within the same node group is that only the master node outputs the data results obtained from processing data to other units (such as downstream nodes, data sources, etc.), while the backup node does not output the data results obtained from processing data. Furthermore, the backup node monitors the processing progress of the master node to ensure that its own processing progress does not exceed that of the master node (the backup node's target task processing progress is lagging behind that of the master node).

[0084] Figure 3I FIG. 4 shows a flow chart of a method for handling an exception by using a multi-active high availability strategy according to an embodiment of the present invention; FIG. Figure 3I As shown, for a certain node group, when Zookeeper determines that an exception (failure) has occurred in the master node by monitoring the heartbeat data, the standby node obtains the notification of the master node exception by monitoring Zookeeper, and all the standby nodes in the node group execute the operation of striving to become the master node (seizing the master position). The standby node that successfully seizes the master position starts to output data to the downstream from the breakpoint where the master node fails (that is, it uses the new master node to continue to execute the unfinished tasks of the master node, and outputs the result data to the downstream based on the data output breakpoint where the master node exception occurs).

[0085] That is, after selecting the master node from the distributed nodes included in each of the node groups, it further includes: determining one or more standby nodes in each of the node groups; for each of the standby nodes, executing a processing thread for constructing a processing task, and using the processing thread to preempt the target task corresponding to the task identifier in the task control table, wherein the processing progress of the standby node target task is later than the processing progress of the target task of the master node; using the abnormality monitoring strategy corresponding to the target high availability strategy to monitor the task running status of the distributed nodes and perform corresponding task processing operations according to the task running status, including: using a distributed coordination system to monitor the heartbeat data of the master node in each node group, and when it is determined that the heartbeat data of the master node in any node group is abnormal, determining that the task running status of the master node of the node group is abnormal; determining a new master node from the standby nodes of the node group; and using the new master node to continue to execute the unfinished tasks of the master node, and outputting the result data to the downstream based on the data output breakpoint where the master node abnormality occurs.

[0086] In some data processing scenarios, after reading data shards, the calculation results are determined and stored simultaneously in a database and a message queue. Because the database and message queue are stored in different storage locations, data inconsistencies may occur due to data storage anomalies. Embodiments of the present invention provide a fault-tolerant method for handling data anomalies in this scenario, improving the reliability and consistency of data processing.

[0087] Figure 3J FIG. 4 shows a flow chart of a data fault tolerance process provided by an embodiment of the present invention; FIG. Figure 3J As shown, Task 1 reads a batch of data (100 data items, represented by numbers 600-699), then uses a processing thread (jobA-Thread-1) to process this batch of data to obtain the data processing results. The data results and data processing progress message (for example, the last successfully saved result item is 699, and Task 1's first processing progress item is 699) are first stored in the message queue. After the data is successfully stored in the message queue, the data processing results of 600-699 are further stored in the database. If the system exits abnormally while storing the data in the database, the results and progress saved in the database will both be the previous successful progress item 599 (i.e., the second processing progress item). This indicates that the data in the message queue and database are inconsistent, and the data processing progress is inconsistent.

[0088] Furthermore, in the scenario of storing data in a database, data can be read from any location. For example, three processing threads can simultaneously write 100 pieces of data into a data shard, with data IDs 100-199, 200-299, and 300-399 respectively. Since the order of writing to the database is arbitrary, the data of 200-299 may be written first, while the data of 100-199 has not yet been written (i.e., it has fallen into the database). At this time, if the data is read, only the data of 200-299 can be read. If the reading method is to read in sequence, the data of 300-399 will be read directly after reading 200-299, resulting in the omission of the data of 100-199. For this scenario, in an embodiment of the present invention, a field is added to the business data table: processing task identifier. This field will bring the task ID of the task processing thread into the database write operation encapsulated by the framework. For example, the task ID corresponding to data 100-199 is A1, the task ID corresponding to data 200-299 is A2, the task ID corresponding to data 300-399 is A3, and so on; that is, the order of distinguishing the written data is distinguished by the processing task ID. Furthermore, when processing data fragments, the downstream task processing thread distinguishes the data according to the processing task and only reads one upstream processed data at a time. For example, after the downstream task processing thread obtains the information that A2's data ID is 200-299, A1's data ID is 100-199, and A3's data ID is 300-399, when reading data from the database, it reads the data according to the order of distinguishing the written data by the processing task ID; because the output of each processing thread itself must be orderly. By splitting the progress to each task processing thread, this design can ensure that all processing threads are processed only once. The management method for database writing provided by the embodiment of the present invention improves the high availability of data and overcomes the possible problem of data processing omissions.

[0089] Figure 3K FIG. 1 is a flow chart of another data fault tolerance method provided by an embodiment of the present invention; Figure 3KAs shown, in the case of data inconsistency, when other distributed nodes take over the task (preemptive) or the standby node continues to process the task after becoming the master node (master-standby), the first processing progress of 699 can be read from the message queue, and the second processing progress of 599 can be read from the database (that is, the second processing progress before the processing result is stored); further, since there is a sequence relationship between storing the message queue and storing it in the database, the breakpoint resumption is started according to the second processing progress (599) of the subsequent database, and then the data of 600-699 is read to continue processing. When outputting the data processing result, the result is not output to the message queue, but only saved to the database, and the second processing progress of the database is updated until the second processing progress of the database is consistent with the first processing progress of the message queue to ensure the data. That is, the processing of the corresponding data shards by the target task further includes: using distributed nodes to write the processing results and the first processing progress obtained by processing the data of multiple data shards into the message queue, and storing the processing results in the database; when it is determined that there is a storage abnormality in the process of storing the processing results in the database, determining the second processing progress before the processing results are stored according to the data source; according to the second processing progress, determining the unprocessed data after the second processing progress, storing the processing results obtained by processing the unprocessed data in the database, and updating the second processing progress; when the updated second processing progress is consistent with the first processing progress, determining that the data processing meets the transaction consistency.

[0090] Furthermore, with respect to data processing of streaming data, since the reception and reading of data are dynamic, existing methods for reading batch data cannot determine whether the reading and processing of streaming data is completed. An embodiment of the present invention provides a method for a downstream system to determine whether data writing of an upstream system is completed. Figure 3L FIG. 1 shows a flow chart of a process for processing a task completion message provided by an embodiment of the present invention; FIG. Figure 3LAs shown: After task A1 processes the data shards (for example, shard 1 to shard 3) and writes the processing results into the message queue, it broadcasts its own task end message to the downstream. The task end message can also include which nodes are used to process job A to which A1 belongs (for example, it is processed by two nodes). Downstream tasks B1, B2, and B3 (which obtain the processing end messages for different data shards) receive the end message broadcast by A1, obtain the number of nodes and node identifiers indicated by the end message, and can determine whether the end messages of tasks from all nodes have been received to determine that job A has ended. Furthermore, after determining that all upstream end messages have been received, if the data source of the downstream output is also a message queue, the current task is further used to broadcast its own end message to the downstream, so that the end message is iteratively transmitted to the downstream in sequence. As a developer, there is no need to develop logic to determine the end of data streaming processing, which improves the efficiency and reliability of data processing. That is, the use of distributed nodes to write the processing results obtained by processing the data of multiple data shards into the message queue further includes: after any distributed node determines that the writing of the processing results into the message queue is completed, broadcasting a write task completion message to each downstream node; wherein, the task completion message includes the identifier of each distributed node that performs the write operation; after the downstream node determines that the upstream distributed node has completed writing the data to the message queue after receiving the task completion message sent by each distributed node identifier.

[0091] Specifically, during the task execution process, the task status corresponding to the task in the task control table is updated according to the execution status of the task. Specifically, the description of the task status is consistent with the description of step S101 and will not be repeated here.

[0092] Furthermore, when the task statuses indicated by the task control table are all completed, it is determined that the processing of each data slice of the to-be-processed data is completed.

[0093] Furthermore, embodiments of the present invention provide a universal development interface that supports both streaming and batch data processing. This embodiment of the present invention can be constructed and published as a Java package. Developers using this universal development interface can reference the Java package to implement the following development interfaces: Initialization interface (optional): implements developer-defined initialization logic; Data reading interface: specifies reading data from a data source (database or message queue) through the development interface; Data processing interface: performs developer-defined logical processing on the read data; Data output interface: outputs the processed data to a specified data source (database or message queue); End determination interface (optional): This interface does not need to be implemented in batch processing scenarios. This interface does not need to be implemented in streaming processing scenarios that read data from a message queue. However, it is required for streaming processing scenarios that read data from a database. Processing completion interface (optional): performs logical operations required after all data processing is complete. After implementing these interfaces, the program can be published as a Java program that can be directly run in a distributed manner, thereby implementing the processing flow of task initialization, data reading, data processing, result output, task completion determination, and post-task completion processing. The data processing to task completion determination process is a loop. By providing a variety of general development interfaces, developers can directly perform development based on the general development interfaces for specific application scenarios of processing distributed data (such as securities trading data, bank transaction data, etc.), thereby improving the development efficiency of distributed data processing and reducing development complexity.

[0094] In the embodiments of the present invention, a job represents a Java program implemented using the development interface provided in the embodiments of the present invention. A job is typically used to read, process, and output data. In distributed scenarios, a job can be split into multiple tasks for execution, each of which is executed by a task thread within one or more processes. The job is considered complete when all tasks are processed. When the data processed by a job is split into multiple data shards for parallel execution, a task is used to execute the processing of any data shard.

[0095] In an embodiment of the present invention, an interface based on the http protocol is further provided for an external system to call a REST interface and a database interface (an interface based on a database table for an external system to control the running status of a task through a control instruction table); wherein the functions of the REST interface may include: 1. Job start processing: start the corresponding job according to the java program name, and obtain the data shard information. If it is determined that the task control table has not yet been generated for the data shard information, perform the initialization action to generate the task control table. If the task control table already exists, start creating a task thread to process the corresponding data shard. 2. Job reprocessing: Before starting the job, regardless of whether the task control table is ready, execute the step of clearing the task information of the corresponding job in the task control table and regenerating the task control table, and then start creating a task thread for processing. 3. Check the job status: return the job status of the set job. The relationship between the job status and the task status is shown in Table 2:

[0096]

[0097] When the task statuses indicated in the task control table are all successfully completed, it is determined that the processing of each data shard of the data to be processed is completed; the embodiment of the present invention improves the scenario adaptability, flexibility and scalability of distributed data processing, improves the efficiency and reliability of processing distributed data, and provides reliability and data consistency of distributed data processing through fault-tolerant processing solutions for different abnormal scenarios.

[0098] like Figure 4 As shown, an embodiment of the present invention provides a device 400 for processing distributed data, including: a task generation module 401, a task preemption module 402 and a status determination module 403; wherein,

[0099] The task generation module 401 is used to obtain multiple data slices of the data to be processed, construct a corresponding task for each of the data slices, and store the task in a task control table;

[0100] The preemption task module 402 is configured to determine a target high-availability strategy that matches the data to be processed from a plurality of high-availability distributed processing strategies; utilize the distributed nodes and the target high-availability strategy to preempt a plurality of target tasks in the task control table whose task status is pending, and process the corresponding data shards through the plurality of target tasks;

[0101] The state determination module 403 is configured to monitor the task running status of the distributed nodes using the abnormality monitoring strategy corresponding to the target high availability strategy and perform corresponding task processing operations according to the task running status.

[0102] An embodiment of the present invention also provides an electronic device for processing distributed data, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for processing distributed data provided in any of the above embodiments.

[0103] An embodiment of the present invention further provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method for processing distributed data provided by any of the above embodiments.

[0104] An embodiment of the present invention further provides a computer program product, including a computer program, which implements any of the above-mentioned methods for processing distributed data when executed by a processor.

[0105] Figure 5 An exemplary system architecture 500 is shown to which the method or apparatus for processing distributed data according to an embodiment of the present invention can be applied.

[0106] like Figure 5 As shown, system architecture 500 may include terminal devices 501, 502, 503, a network 504, and a server 505. Network 504 is used to provide a medium for communication links between terminal devices 501, 502, 503 and server 505. Network 504 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0107] Users can use terminal devices 501, 502, 503 to interact with server 505 via network 504 to receive or send requests, etc. Various client applications can be installed on terminal devices 501, 502, 503, such as e-commerce client applications, web browser applications, financial product clients, etc.

[0108] The terminal devices 501 , 502 , and 503 may be various electronic devices having a display screen and supporting various client applications, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0109] Server 505 can be a server that provides various services, such as a background management server that provides support for client applications used by users using terminal devices 501, 502, and 503. The background management server can process received data requests to be processed and feed back the data processing results to the terminal devices. The multiple distributed nodes provided in the embodiments of the present invention can run on the server.

[0110] It should be noted that the method for processing distributed data provided in the embodiment of the present invention is generally executed by the server 505 , and accordingly, the device for processing distributed data is generally set in the server 505 .

[0111] It should be understood that Figure 5 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0112] Reference below Figure 6 , which shows a schematic structural diagram of a computer system 600 of a terminal device suitable for implementing an embodiment of the present invention. Figure 6 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0113] like Figure 6 As shown, computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of system 600 are also stored in RAM 603. CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0114] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, mouse, and the like; an output section 607 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 608 including devices such as a hard disk; and a communication section 609 including a network interface card such as a LAN card or a modem. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. Removable media 611, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 610 as needed, so that computer programs read from the removable media can be installed in the storage section 608 as needed.

[0115] In particular, according to embodiments disclosed herein, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed herein include a computer program product comprising a computer program embodied on a computer-readable medium, the computer program containing program code for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609 and / or installed from removable media 611. When executed by central processing unit (CPU) 601, the computer program performs the aforementioned functions defined in the system of the present invention.

[0116] It should be noted that the computer-readable medium described in the present invention may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.

[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0118] The modules and / or units involved in the embodiments of the present invention may be implemented in software or in hardware. The modules and / or units described may also be provided in a processor. For example, they may be described as follows: a processor including a task generation module, a task preemption module, and a status determination module. The names of these modules do not, in some cases, constitute a limitation on the modules themselves. For example, the task generation module may also be described as a "module for constructing a corresponding task for each of the data slices."

[0119] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes: obtaining multiple data slices of the data to be processed, constructing corresponding tasks for the data slices, and storing the tasks in a task control table; determining a target high-availability strategy that matches the data to be processed from a variety of high-availability distributed processing strategies; using distributed nodes and target high-availability strategies, preempting multiple target tasks in the task control table whose task status is to be processed, processing the corresponding data slices through multiple target tasks, and using the monitoring strategy corresponding to the target high-availability strategy to monitor the task running status of the distributed nodes and perform task processing operations. The embodiments of the present invention improve the scenario adaptability, flexibility and scalability of distributed data processing, and improve the efficiency and high availability of processing distributed data.

[0120] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for processing distributed data, characterized in that: include: Acquire multiple data slices of the data to be processed, construct a corresponding task for each of the data slices, and store the task in a task control table; Determine a target high availability strategy that matches the data to be processed from a plurality of high availability distributed processing strategies; Using distributed nodes and the target high availability strategy, multiple target tasks whose task status is pending in the task control table are seized, and corresponding data slices are processed by the multiple target tasks; Using the abnormal monitoring strategy corresponding to the target high availability strategy, monitor the task running status of the distributed node and perform corresponding task processing operations according to the task running status; Wherein, in case the target high availability strategy is a multi-active high availability strategy, The method of using the distributed nodes and the target high availability strategy to seize the target task whose task status is pending in the task control table includes: Dividing the plurality of distributed nodes into a plurality of node groups; Get the task ID of the task to be processed and the number of node groups; Determine, based on a second calculation relationship between the task identifier and the number of groups, identifiers of each to-be-processed task to be processed by each of the node groups; For each node group, a master node is selected from the distributed nodes included in the node group, a main control thread of the master node is used to construct multiple processing threads for processing tasks, and the processing threads are used to preempt the target task corresponding to the task identifier in the task control table.

2. The method according to claim 1, characterized in that In the case where the target high availability strategy is a preemptive high availability strategy, The method of using the distributed nodes and the target high availability strategy to seize the target task whose task status is pending in the task control table includes: For any distributed node, start the main control thread of the node, and use the main control thread to generate multiple processing threads for processing tasks according to the number of shards of the data shards; The main control thread is used to obtain each of the target tasks to be processed, and the target tasks are stored in a local storage, and a reordering operation is performed on each of the target tasks in the local storage, so that multiple processing threads can preempt the target tasks based on the reordered tasks.

3. The method according to claim 1, characterized in that In the case where the target high availability strategy is a preemptive high availability strategy, The using of the abnormal monitoring strategy corresponding to the target high availability strategy to monitor the task running status of the distributed node and performing the corresponding task processing operation according to the task running status includes: Registering a plurality of the distributed nodes in a monitoring center; The monitoring center is used to perform the following operations based on the abnormal monitoring strategy: Monitor each of the distributed nodes to create different heartbeat files at set time intervals; If it is monitored that the heartbeat file of any distributed node is not updated within a set number of set time intervals, it is determined that there is an abnormality in the task operation of the distributed node; For a distributed node with an abnormal task running condition, a current task processed by a thread of the distributed node with the abnormal task is obtained, and a task state of the current task is changed to pending in the task control table.

4. The method according to claim 1, characterized in that: In the case where the target high availability strategy is a master-standby high availability strategy, The method of using the distributed nodes and the target high availability strategy to seize the target task whose task status is pending in the task control table includes: Determine a plurality of master nodes from the plurality of distributed nodes; Get the task ID of the task to be processed and the number of nodes of the master node; Determine, based on a first calculation relationship between the number of nodes and the task identifier, each identifier of the task to be processed by each of the master nodes; A plurality of processing threads for processing tasks are constructed by using the main control thread of each master node, and the processing threads are used to preempt the target task corresponding to the to-be-processed task identifier in the task control table.

5. The method according to claim 4, characterized in that Determining a plurality of master nodes from the plurality of distributed nodes comprises: The plurality of distributed nodes connect to the distributed coordination system when starting up, and create a unique directory identifier in a set storage location of the distributed coordination system, and the distributed node that successfully creates the unique directory identifier is used as the master node; The using of the abnormal monitoring strategy corresponding to the target high availability strategy to monitor the task running status of the distributed node and performing the corresponding task processing operation according to the task running status includes: In the case where the target high availability strategy is a master-standby high availability strategy, the distributed coordination system is used to monitor the heartbeat data of multiple master nodes. When it is determined that the heartbeat data of any master node is abnormal, the task status processed by the abnormal master node is set to pending, and a new master node is selected from multiple standby nodes. The new master node is used to replace the abnormal master node, and the pending target task corresponding to the abnormal master node continues to be executed.

6. The method according to claim 1, characterized in that After selecting a master node from the distributed nodes included in each of the node groups, the method further includes: Determine one or more standby nodes in each of the node groups; For each of the standby nodes, executing the step of constructing a processing thread for processing tasks, and using the processing thread to preempt the target task corresponding to the task identifier in the task control table; wherein the processing progress of the target task of the standby node is later than the processing progress of the target task of the master node; The using of the abnormal monitoring strategy corresponding to the target high availability strategy to monitor the task running status of the distributed node and performing the corresponding task processing operation according to the task running status includes: In the case where the target high availability strategy is a multi-active high availability strategy, the heartbeat data of the master nodes in each node group is monitored by using a distributed coordination system, and when it is determined that the heartbeat data of the master node in any node group is abnormal, it is determined that the task running status of the master node of the node group is abnormal; A new master node is determined from the backup nodes of the node group; and the new master node is used to continue to execute the unfinished tasks of the master node, and the result data is output to the downstream based on the data output breakpoint where the master node abnormally occurs.

7. The method according to claim 1, characterized in that The processing of the corresponding data slices by the target task further includes: Using distributed nodes, writing processing results and the first processing progress obtained by processing the data of the plurality of data slices into a message queue, and storing the processing results in a database; In the case where it is determined that a storage anomaly exists in the process of storing the processing result in the database, determining a second processing progress before the processing result is stored according to the data source; According to the second processing schedule, determining unprocessed data after the second processing schedule, storing processing results obtained by processing the unprocessed data in the database, and updating the second processing schedule; When the updated second processing progress is consistent with the first processing progress, it is determined that the data processing satisfies transaction consistency.

8. The method according to claim 7, characterized in that The step of using the distributed nodes to write the processing results obtained by processing the data of the plurality of data slices into a message queue further comprises: After any distributed node determines that the processing result is written into the message queue, it broadcasts a writing task completion message to each downstream node; wherein the task completion message includes the identifier of each distributed node that performs the writing operation; After the downstream node determines that the task completion message sent by each of the distributed node identifiers is received, the downstream node determines that the upstream distributed node completes writing the data to the message queue.

9. A device for processing distributed data, characterized in that: include: Generate task module, seize task module and determine status module; among them, The task generation module is used to obtain multiple data slices of the data to be processed, construct a corresponding task for each of the data slices, and store the task in the task control table; The preemption task module is used to determine a target high-availability strategy that matches the data to be processed from a plurality of high-availability distributed processing strategies; using distributed nodes and the target high-availability strategy, preempt multiple target tasks whose task status is to be processed in the task control table, and process corresponding data slices through the multiple target tasks; Among them, in the case where the target high availability strategy is a multi-active high availability strategy, the use of distributed nodes and the target high availability strategy to preempt the target task whose task status is pending in the task control table includes: dividing a plurality of the distributed nodes into a plurality of node groups; obtaining the task identifier of the task to be processed and the number of node groups; based on a second calculation relationship between the task identifier and the number of groups, determining each pending task identifier processed by each node group; for each node group, selecting a master node from the distributed nodes included in the node group, using the main control thread of the master node to construct multiple processing threads for processing tasks, and using the processing thread to preempt the target task corresponding to the task identifier in the task control table; The state determination module is used to monitor the task running status of the distributed node and perform corresponding task processing operations according to the task running status by utilizing the abnormal monitoring strategy corresponding to the target high availability strategy.

10. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.

11. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

12. A computer program product, comprising a computer program, characterized in that When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Task scheduling method and system

    CN111427670A

  • Distributed task scheduling management method and device

    CN111767122A

  • Distributed data processing method based on dense computing and related equipment

    CN118796478A