Data processing method and device, computer equipment and computer readable storage medium

CN116048788BActive Publication Date: 2026-10-09SHENZHEN KEMAI TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211685798.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2026-10-09
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

[0003]传统的基于流计算技术的数据处理方法,要求与数据处理任务匹配的数据处理流程能够通过一个流式处理链路进行表征,存在应用场景受限的缺点

Benefits of technology

[0053] The aforementioned data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire a data processing flow matching the data processing task; determine a set of task nodes that implement the data processing task, and the data interaction relationships between the task nodes included in the task node set; associate each task node with at least a part of the data processing flow; determine the respective streaming processing link for each task node; the streaming processing link matches at least a part of the data processing flow associated with the task node; the at least a part of the data processing flow satisfies the streaming processing conditions; and, according to the data interaction relationships between the task nodes, process the data to be processed for the data processing task based on each streaming processing link to obtain the data processing result. In the above data processing process, at least two task nodes are configured according to the data processing flow matched with the data task. Each task node can interact with data, and each task node is associated with at least a part of the data processing flow. In this way, it is equivalent to splitting complex tasks that do not meet the conditions for streaming processing into multiple sub-flows that meet the conditions for streaming processing. These sub-flows are then bucketed to the corresponding task nodes for processing. Thus, each task node can adopt the streaming processing method. This can improve the data processing efficiency by applying stream computing technology, enhance the flexibility of the data processing method, and help expand the application scenarios of the data processing method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116048788B_ABST
    Figure CN116048788B_ABST
Patent Text Reader

Abstract

The application relates to a data processing method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring a data processing flow matched with a data processing task; determining a task node set for realizing the data processing task and a data interaction relationship between each task node contained in the task node set; each task node is associated with at least part of the data processing flow meeting a stream processing condition; the stream processing link of each task node is determined; the stream processing link is matched with at least part of the data processing flow associated with the task node; according to the data interaction relationship between each task node, the data processing task is processed based on each stream processing link, and a data processing result is obtained. The above method can improve the flexibility of the data processing method and is beneficial to expanding the application scenario of the data processing method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology

[0002] With the advent of the big data era, more and more data-intensive businesses have emerged, including financial services, telecommunications data management, and so on. Data-intensive businesses are characterized by a wide range of sources, large quantities, and frequent changes in data. In order to improve data processing efficiency, stream computing technology has emerged.

[0003] Traditional data processing methods based on stream computing technology require that the data processing flow matching the data processing task can be represented through a streaming processing link, which has the disadvantage of limiting application scenarios. Summary of the Invention

[0004] Therefore, it is necessary to provide a data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve and expand application scenarios in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a data processing method. The method includes:

[0006] Obtain the data processing flow that matches the data processing task;

[0007] The set of task nodes that implement the data processing task is determined, as well as the data interaction relationship between the task nodes contained in the set of task nodes; each task node is associated with at least a part of the data processing flow; the at least a part of the data processing flow satisfies the streaming processing condition.

[0008] Determine the respective streaming processing link for each of the task nodes; the streaming processing link matches the at least a portion of the data processing flow associated with the task node;

[0009] Based on the data interaction relationships between the task nodes, the data to be processed in the data processing task is processed according to the streaming processing links to obtain the data processing results.

[0010] In one embodiment, the acquisition of a data processing flow matching the data processing task includes:

[0011] Obtain the data to be processed for the data processing task and determine the data type of the data to be processed; the data type has at least two types;

[0012] Based on the data processing task, determine the respective data processing flow for each data type of data to be processed;

[0013] The process of determining the set of task nodes for executing the data processing flow includes:

[0014] Determine the task node corresponding to each of the data processing flows to obtain a task node set including each of the task nodes.

[0015] In one embodiment, determining the respective streaming processing link for each of the task nodes includes:

[0016] For each task node, at least two target processing logics are determined from a plurality of candidate data processing logics; the at least two target processing logics are matched with the at least part of the data processing flow associated with the task node;

[0017] Determine the streaming processing chain that includes each of the target processing logics.

[0018] In one embodiment, determining the streaming processing chain including each of the target processing logics includes:

[0019] Based on at least a portion of the data processing flow, determine the link relationships between each of the target processing logics;

[0020] Based on the link relationship, configure the subscription relationship between each target processing logic, and create a streaming processing link including each target processing logic; in the streaming processing link, the next target processing logic subscribes to the output data of the previous target processing logic.

[0021] In one embodiment, the method further includes:

[0022] In response to a task update event for the data processing task, obtain task update information;

[0023] Determine the update process that matches the task update information, and perform node maintenance on each task node based on the update process.

[0024] In one embodiment, the node maintenance of each task node based on the update process includes:

[0025] If there is no target update process corresponding to the target processing process in the update process, then delete the target processing link that matches the target processing process and cancel the target task node associated with the target processing link;

[0026] If the update process includes a target update process corresponding to the target processing process, then based on the target update process, the target processing link associated with the target task node that matches the target processing process is updated;

[0027] If the update process includes a new processing process, then based on the new processing process, a new task node matching the new processing process is configured on each of the task nodes.

[0028] Secondly, this application provides a data processing apparatus. The apparatus includes:

[0029] The acquisition module is used to acquire the data processing flow that matches the data processing task;

[0030] The task node set determination module is used to determine the task node set that implements the data processing task, and the data interaction relationship between each task node included in the task node set; each task node is associated with at least a part of the data processing process; the at least a part of the data processing process satisfies the streaming processing conditions.

[0031] A streaming processing link determination module is used to determine the respective streaming processing link of each of the task nodes; the streaming processing link matches at least a portion of the data processing flow associated with the task node;

[0032] The data processing module is used to process the data to be processed in the data processing task according to the data interaction relationship between each task node and based on each streaming processing link to obtain the data processing result.

[0033] In one embodiment, the acquisition module is specifically used for:

[0034] Obtain the data to be processed for the data processing task, and determine the data type of the data to be processed; the data type has at least two types; based on the data processing task, determine the respective data processing flow for each data type of data to be processed.

[0035] The task node set determination module is specifically used for:

[0036] Determine the task node corresponding to each of the data processing flows to obtain a task node set including each of the task nodes.

[0037] In one embodiment, the streaming processing link determination module includes:

[0038] The target processing logic determination unit is used to determine at least two target processing logics from a plurality of candidate data processing logics for each task node; the at least two target processing logics are matched with the at least part of the data processing flow associated with the task node;

[0039] A streaming processing link determination unit is used to determine the streaming processing link that includes each of the target processing logics.

[0040] In one embodiment, the streaming processing link determination unit is specifically used for:

[0041] Based on at least a portion of the data processing flow, determine the link relationships between each of the target processing logics;

[0042] Based on the link relationship, configure the subscription relationship between each target processing logic, and create a streaming processing link including each target processing logic; in the streaming processing link, the next target processing logic subscribes to the output data of the previous target processing logic.

[0043] In one embodiment, the data processing apparatus further includes:

[0044] The task update information acquisition module is used to acquire task update information in response to a task update event for the data processing task.

[0045] The node maintenance module is used to determine the update process that matches the task update information, and to perform node maintenance on each task node based on the update process.

[0046] In one embodiment, the node maintenance module is specifically used for:

[0047] If there is no target update process corresponding to the target processing process in the update process, then delete the target processing link that matches the target processing process and cancel the target task node associated with the target processing link;

[0048] If the update process includes a target update process corresponding to the target processing process, then based on the target update process, the target processing link associated with the target task node that matches the target processing process is updated;

[0049] If the update process includes a new processing process, then based on the new processing process, a new task node matching the new processing process is configured on each of the task nodes.

[0050] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the aforementioned data processing method.

[0051] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the above-described data processing method.

[0052] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the aforementioned data processing method.

[0053] The aforementioned data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire a data processing flow matching the data processing task; determine a set of task nodes that implement the data processing task, and the data interaction relationships between the task nodes included in the task node set; associate each task node with at least a part of the data processing flow; determine the respective streaming processing link for each task node; the streaming processing link matches at least a part of the data processing flow associated with the task node; the at least a part of the data processing flow satisfies the streaming processing conditions; and, according to the data interaction relationships between the task nodes, process the data to be processed for the data processing task based on each streaming processing link to obtain the data processing result. In the above data processing process, at least two task nodes are configured according to the data processing flow matched with the data task. Each task node can interact with data, and each task node is associated with at least a part of the data processing flow. In this way, it is equivalent to splitting complex tasks that do not meet the conditions for streaming processing into multiple sub-flows that meet the conditions for streaming processing. These sub-flows are then bucketed to the corresponding task nodes for processing. Thus, each task node can adopt the streaming processing method. This can improve the data processing efficiency by applying stream computing technology, enhance the flexibility of the data processing method, and help expand the application scenarios of the data processing method. Attached Figure Description

[0054] Figure 1 This is a diagram illustrating the application environment of the data processing methods in some embodiments;

[0055] Figure 2 This is a flowchart illustrating the data processing method in some embodiments;

[0056] Figure 3 These are schematic diagrams illustrating different forms of data processing flows in some embodiments;

[0057] Figure 4 This is a schematic diagram of the implementation architecture of the data processing method in some embodiments;

[0058] Figure 5 This is a schematic diagram of the dynamic update logic of the rule chain in some embodiments;

[0059] Figure 6 This is a flowchart illustrating the data processing method in some other embodiments;

[0060] Figure 7 This is a structural block diagram of the data processing apparatus in some embodiments;

[0061] Figure 8 This is a diagram showing the internal structure of a computer device in some embodiments. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0063] The data processing method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, the business terminal 102 can communicate with the server terminal 104 via a network. The data storage system can store the data that the server terminal 104 needs to process. The data storage system can be integrated into a server within the server terminal 104, or it can be located in the cloud or on other servers. The number of business terminals 102 can be one or multiple, including but not limited to mobile phones, computers, smart voice interaction devices, smart home appliances, and in-vehicle terminals. The server terminal 104 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be implemented using terminals, including but not limited to mobile phones, computers, smart voice interaction devices, smart home appliances, and in-vehicle terminals.

[0064] Specifically, server 104 can obtain a data processing flow matching the data processing task defined by business terminal 102; determine the set of task nodes that implement the data processing task and the data interaction relationship between the task nodes included in the task node set; each task node is associated with at least a part of the data processing flow; the at least a part of the data processing flow satisfies the streaming processing conditions; determine the streaming processing link of each task node; the streaming processing link matches the at least a part of the data processing flow associated with the task node; and process the data to be processed in the data processing task based on each streaming processing link according to the data interaction relationship between each task node to obtain the data processing result.

[0065] In some embodiments, such as Figure 2 As shown, a data processing method is provided, which can be applied to... Figure 1 Taking server-side error 104 as an example, the following steps are included:

[0066] Step S202: Obtain the data processing flow that matches the data processing task.

[0067] Data processing tasks refer to the operational tasks of processing raw data to achieve certain objectives. These tasks typically correspond to business requirements. For example, data processing tasks may include checking the source and target databases before data migration, migrating data stored in sub-partitions of the source database to the corresponding target partition in the target database, and data filtering and processing tasks, etc. A data processing flow refers to the process of processing raw data.

[0068] Specifically, for data processing tasks corresponding to business needs, these tasks can be broken down into multiple data processing steps, thereby determining the data processing flow that includes each step. This data processing flow can take various forms, for example, it could be... Figure 3 The chained processing flow without branches shown in (A) can also be Figure 3 The chained processing flow with branches shown in (B) can also be... Figure 3 (C) illustrates a non-chained processing flow containing loops. Taking a chained processing flow without branches as an example, for instance, a data processing flow matching a task involving storing target data within raw data could include data receiving, data filtering, data transformation, and data storage. Similarly, for a task involving searching related data within raw data, a matching data processing flow could include data receiving, data searching, data extraction, and data output. Furthermore, developers can determine the data processing task based on business requirements, thereby determining the matching data processing flow. Alternatively, the server can determine the matching data processing flow by parsing the data processing task.

[0069] It should be noted that since the data to be processed in the same data processing task can be of various types, a data processing task corresponding to the same task objective can include multiple data processing flows. For example, for standard data that meets the data format requirements, the format conversion step can be omitted.

[0070] Step S204: Determine the set of task nodes that implement the data processing task, and the data interaction relationships between the task nodes contained in the set of task nodes.

[0071] Here, a task node set refers to a collection of task nodes that execute data processing tasks, with each task node associated with at least a part of the data processing flow. Specifically, for a data processing task that matches multiple data processing flows, the server can configure corresponding task nodes for each data processing flow, thus associating each task node with a corresponding data processing flow; for a data processing task that matches only one data processing flow, the server can split the data processing flow into multiple sub-flows, and configure corresponding task nodes for each sub-flow, thus each task node represents a part of the key data processing flow.

[0072] Furthermore, at least a portion of the data processing flow associated with each task node satisfies the streaming processing conditions. Specifically, if at least a portion of the data processing flow is a chain-like processing flow without branches, then this processing flow can be implemented using streaming processing logic, meaning that the processing flow satisfies the streaming processing conditions. For example, Figure 3 In (A), "A1→A2→A3" and "A4→A5" all meet the conditions for streaming processing. Figure 3 In (B), “B1→B2”, “B3→A4→A5”, “B6→B7”, “B8→B9”, etc., all meet the conditions for streaming processing. Figure 3 In (C), "C1→C2", "C3→C4", and "C2→C5→C6" all satisfy the streaming processing conditions. That is, the server can, based on the relationships between the various processing steps in the data processing flow, break down a data processing flow that originally did not meet the streaming processing conditions into multiple sub-flows that do. On this basis, corresponding task nodes are configured for each sub-flow, obtaining a set of task nodes to implement the data processing tasks. Then, based on the relationships between the various sub-flows, the data interaction relationships between the task nodes can be determined. Figure 3 For example, if the sub-process corresponding to node 1 is “C1→C2”, the sub-process corresponding to node 2 is “C3→C4”, and the sub-process corresponding to node 3 is “C2→C5→C6”, then the data interaction relationship between the nodes is “node 1→node 2→node 3”.

[0073] Step S206: Determine the streaming processing link for each task node.

[0074] In this context, a streaming processing link refers to a data processing link that meets the conditions for streaming processing. This streaming processing link can include at least two linked processing logics, and it matches at least a portion of the data processing flow associated with a task node. Specifically, for each task node, the server can determine the streaming processing link for that task node by identifying the link formed by the processing logic corresponding to each processing step in the sub-flow associated with that task node. Alternatively, the server can merge or transform each processing step based on its supported processing logic to determine the target processing logic corresponding to each processing step, thereby determining the streaming processing link containing each target processing logic.

[0075] Step S208: According to the data interaction relationship between each task node, the data to be processed in the data processing task is processed based on each streaming processing link to obtain the data processing result.

[0076] The data interaction relationships between task nodes determine the data flow direction between them, and each task node's own streaming processing link determines the processing performed on the data flowing into that node. Specifically, the server can process the data to be processed by the data processing task according to the data interaction relationships between task nodes and based on each streaming processing link to obtain the data processing results.

[0077] The above data processing method involves: acquiring a data processing flow matching the data processing task; determining the set of task nodes implementing the data processing task and the data interaction relationships between the task nodes within the set; associating each task node with at least a portion of the data processing flow; ensuring that this portion of the data processing flow satisfies streaming processing conditions; determining the streaming processing link for each task node; ensuring that this streaming processing link matches the at least portion of the data processing flow associated with the task node; and processing the data to be processed for the data processing task based on the data interaction relationships between the task nodes and each streaming processing link to obtain the data processing result. In this data processing process, at least two task nodes are configured according to the data processing flow matching the data task. These task nodes can interact with each other, and each task node is associated with at least a portion of the data processing flow. This effectively splits complex tasks that do not meet streaming processing conditions into multiple sub-processes that satisfy the streaming processing conditions, which are then bucketed to the corresponding task nodes for processing. Thus, each task node can adopt streaming processing, improving data processing efficiency through stream computing technology while enhancing the flexibility of the data processing method and expanding its application scenarios.

[0078] In one embodiment, step S202 includes: acquiring the data to be processed for the data processing task and determining the data type of the data to be processed; based on the data processing task, determining the respective data processing flow for each data type of data to be processed. In this embodiment, determining the task node set for executing the data processing flow includes: determining the task node corresponding to each data processing flow, and obtaining a task node set including each task node.

[0079] The data types must be at least two. Different data types could be due to different data compositions and / or different data formats. For example, the data type could be a message in a Message Queue (MQ) or data in a database. It's understandable that different data types will lead to different data processing flows, even with the same task objective. Based on this, the server can obtain the data to be processed for the data processing task and determine its data type. Then, based on the data processing task, it determines the data processing flow for each data type. Next, the server determines the corresponding task nodes for each data processing flow, obtaining a task node set including all task nodes.

[0080] In this embodiment, for each type of data to be processed, a corresponding data processing flow is matched, and the data is bucketed to different task nodes for data processing, which helps to improve data processing efficiency.

[0081] In one embodiment, step S206 includes: for each task node, determining at least two target processing logics from a plurality of candidate data processing logics; and determining a streaming processing link that includes each target processing logic.

[0082] Specifically, at least two target processing logics match at least a portion of the data processing flow associated with the task node. Specifically, the server can pre-create multiple different candidate data processing logics, such as data reception, data sharding, routing, data filtering, and data transformation. Then, for each task node, the server can determine at least two target processing logics that match the at least a portion of the data processing flow associated with that task node from among the multiple candidate data processing logics. Next, the server determines the link relationships between the target processing logics based on the at least a portion of the data processing flow associated with that task node, thereby determining the streaming processing link containing each target processing logic.

[0083] Furthermore, among the multiple candidate data processing logics, there are multiple sets of alternative processing logics. For example, candidate data processing logics a1, a2, and a3 are alternative processing logics. In practical applications, the server usually needs to process multiple data processing tasks simultaneously, and there may be situations where multiple task nodes call the same data processing logic. Based on this, the server can obtain the current load of each candidate data processing logic and, according to the current load, prioritize the candidate data processing logic with the lower load as the target processing logic to further improve efficiency.

[0084] In addition, such as Figure 4 As shown, each candidate data processing logic can be carried out by its corresponding candidate data processing module. For example, the general data stream receiving module (Source) carries the data receiving logic. By providing a data acquisition interface, it supports obtaining real-time data streams from Kafka and other interfaces such as HTTP. The data sharding module (Sharding) carries the data sharding logic. It can obtain the node information of each task node and coordinate with the routing module to perform data routing, which is effective in cluster mode. The routing module (Router) carries the data routing logic. It calculates which task node should process the data based on the routing key given in the event information from the data stream. The data filtering module (Filter) carries the data filtering logic. It is used to filter the data stream, perform data cleaning in real-world scenarios, and provide various data filtering options. The system includes a conditional or matching interface; a data transformation module (Map) to carry data transformation logic, used for transforming streaming data, such as deleting, adding, or converting fields; a clock module (Clock) to serve as a time reference module in time processing, usually appearing together with a time window module, providing centralized and unified time calibration and comparison, for example, providing both wall time and event time; a time window module (Window) to provide aggregation operations within a certain time window for streaming data using streaming event time or wall time as calculation factors (e.g., statistics on the number of events in the past 10 seconds), the time window can support both scrolling and sliding window modes; and a general data aggregation module (Sink) to perform data storage and data statistics.

[0085] In this embodiment, the target processing logic is first determined from multiple candidate data processing logics, and then the streaming processing link containing each target processing logic is determined. Developers can dynamically add or delete candidate data processing logics according to actual scenario requirements to further improve the flexibility of the data processing method.

[0086] In a specific implementation, determining the streaming processing link that includes each target processing logic includes: determining the link relationship between each target processing logic based on at least a part of the data processing flow; configuring the subscription relationship between each target processing logic based on the link relationship; and creating a streaming processing link that includes each target processing logic.

[0087] In this streaming processing chain, the next target processing logic subscribes to the output data of the previous target processing logic. Specifically, the server can determine the link relationship between each target processing logic based on at least a part of the data processing flow, and then configure the subscription relationship between each target processing logic based on the link relationship. By having the next target processing logic subscribe to the output data of the previous target processing logic, a streaming processing chain including each target processing logic is created.

[0088] In this embodiment, the creation of the streaming processing link can be completed with simple configuration, which is simple and conducive to improving data processing efficiency.

[0089] In one embodiment, the data processing method further includes: in response to a task update event for a data processing task, obtaining task update information; determining an update process that matches the task update information; and performing node maintenance on each task node based on the update process.

[0090] Specifically, during task processing, business requirements may change. For example, for data filtering, the filtering conditions may be updated; for querying, the query range may be modified; or, with system upgrades, a certain data type may no longer be processed; and so on. Based on this, the server responds to task update events for data processing tasks by obtaining task update information. This task update information is specific to the data processing task, carrying the update items and update details for those items. Then, based on this task update information and the original data processing flow, the server determines the update flow that matches the task update information and performs node maintenance for each task node based on the update flow.

[0091] Furthermore, the specific methods for maintaining task nodes include at least one of adding nodes, deleting nodes, or updating nodes. Updating a node refers to updating the streaming processing link corresponding to the task node. In a specific application, maintaining each task node based on the update process includes: if there is no target update process corresponding to the target processing process in the update process, then delete the target processing link matching the target processing process and deregister the target task node associated with that target processing link; if the update process includes a target update process corresponding to the target processing process, then update the target processing link associated with the target task node that matches the target processing process based on the target update process; if the update process includes a new processing process, then configure a new task node matching the new processing process on each task node based on the new processing process.

[0092] Specifically, one of the task nodes is designated as the target task node. If the update process does not contain a target update process corresponding to the target processing process associated with that target task node, it means the updated data processing task does not need to match the target processing process. In this case, the target processing link matching that target processing process is deleted, and the target task node associated with that target processing link is deregistered. If the update process includes a target update process corresponding to the target processing process, it means the updated data processing task still needs to match the target processing process, but it needs to be updated based on the target processing process, such as adding or deleting certain processing logic. In this case, the server updates the target processing links associated with the target task node that match the target processing process based on the target update process. If the update process includes new processing processes based on the sub-processes associated with each task node, then based on the new processing processes, new task nodes matching the new processing processes are configured on each task node.

[0093] by Figure 5 For example, data processing chains (hereinafter referred to as rule chains) corresponding to data processing tasks can be stored in the database, and the Akka Actor model can be used to organize and implement each rule chain. Specifically, such as Figure 5As shown, a top-level actor (link) is declared based on the AkkaActor model to monitor the status of each rule link. During the initial loading process, all rule links (including Actor-1 to Actor-n) are loaded into their respective task nodes. Updated rule links corresponding to the update process can be synchronously stored in the database, replacing the original rule links for the same data processing task. Furthermore, the server checks links periodically (e.g., every minute), retrieves the latest rule links from the database, compares them with the current application's rule links, and removes or updates rules that have been removed. Further, rule links can be cached locally on the server, and the local cache can be compared with the rule links in the database to achieve synchronous updates of the local cache. For example, the database and local cached rule routing tables can be compared; for the same rule link, if the database update time and the cache update time are inconsistent, the rule is updated, a stop signal is sent to the rule actor, the cache is updated to "updatable," the rule actor is monitored for termination, and finally, the stopped rule actor is reset. Figure 5 In the process, at checkpoint 1, rule Actor-n was updated; at checkpoint 2, rule Actor-2 was removed; and at checkpoint 3, rule Actor-m was added. It can be understood that adding a new rule chain means adding a task node corresponding to that rule chain.

[0094] In the above embodiments, the data processing links of each task node can be dynamically updated in response to task update events for data processing tasks, which is beneficial to further improve the flexibility of data processing methods.

[0095] In some embodiments, such as Figure 6 As shown, the data processing methods include:

[0096] Step S601: Obtain the data to be processed for the data processing task and determine the data type of the data to be processed;

[0097] Among them, there are at least two types of data;

[0098] Step S602: Based on the data processing task, determine the data processing flow for each data type of data to be processed.

[0099] Step S603: Determine the task nodes corresponding to each data processing flow, obtain a task node set including each task node, and determine the data interaction relationship between each task node.

[0100] The data processing flow meets the requirements for streaming processing.

[0101] Step S604: For each task node, determine at least two target processing logics from multiple candidate data processing logics;

[0102] Among them, at least two target processing logics are matched with the data processing flow associated with the task node;

[0103] Step S605: Determine the link relationships between each target processing logic based on at least a portion of the data processing flow;

[0104] Step S606: Configure the subscription relationship between each target processing logic based on the link relationship, and create a streaming processing link including each target processing logic;

[0105] In this streaming processing chain, the next target processing logic subscribes to the output data of the previous target processing logic;

[0106] Step S607: In response to a task update event for a data processing task, obtain task update information;

[0107] Step S608: Determine the update process that matches the task update information, and perform node maintenance on each task node based on the update process.

[0108] To facilitate understanding, the following example illustrates the data processing method in detail, using a data processing task that requires designing a streaming program to filter text and store the data in any location.

[0109] like Figure 3 As shown, the server can be configured with one or more task management nodes to provide registration support for task nodes, enabling mutual discovery among multiple task nodes to form a processing cluster. Data interaction may or may not occur between task nodes, and each task node can be configured with multiple data processing modules to carry candidate data processing logic. Each data processing module provides a common interface for task nodes to call. After the data to be processed by each task node, the processed data can flow to other systems or databases. The types and uses of each data processing module are described above and will not be repeated here. Specifically, the server organizes and implements each streaming processing link based on the Akka Actor model. The Actor model emphasizes asynchronous message communication to avoid mutual interference. In this application, messages flow unidirectionally between modules within the same task node (data flow can be understood as water flow), satisfying the conditions for streaming processing, thus enabling efficient data processing.

[0110] Suppose you need to design a streaming program to filter text and store the data in any location, you can configure the streaming chain in the task node using the following steps:

[0111] Step 1: Based on the Akka actor model, declare a top-level actor to organize the various modules mentioned above.

[0112] Step 2: First, obtain the real-time data stream by using the "General Data Stream Receiving Module (Source)" (hereinafter referred to as Source). Since the module has already defined the interface, you only need to obtain the real-time data stream through the interface. The current module exists in the form of an actor and maintains a subscription registry internally. If an actor needs to obtain data flowing through the current module, it only needs to register with the current module. The current module will then send the flowing data, such as real-time text, to the subscribers.

[0113] Step 3: According to business requirements, it is necessary to receive data from the Source and implement the filtering function. This can be done by using the "Data Filtering Module (Filter)" and registering with the Source to receive real-time data. The Data Filtering Module is configured with multiple Boolean interfaces. As long as the interface returns true, it means that the data meets the conditions and can flow into the next module. The data that meets the conditions is sent to the subscribers of the Filter through the subscription response, while the data that does not meet the conditions is discarded.

[0114] Step 4: Since there is a need to store data, the "Data Transformation Module (Map)" can be registered with the Filter to obtain the output data of the Filter and perform data transformation. Then, the transformed data can be sent to the subscriber of the Map, the "General Data Collection Module (Sink)," to realize data storage.

[0115] It is understandable that, following the above approach, larger and more complex processing flows can be achieved by adding other modules or even configuring data interaction relationships between different task nodes. The aforementioned data processing method can achieve asynchronous decoupling based on the Akka Actor model, improving data processing efficiency. Task nodes support dynamic addition and deletion, and for a single task node, logical processing modules can be dynamically combined according to specific business needs (for example, a user can omit the clock module and time window calculation module, simply implementing other modules, thus enabling real-time filtering calculations; conversely, it forms a data aggregation function within a time window). Through freely organized streaming processing links, multiple module links can also be formed to handle rules with minor differences in the same scenario (such as simultaneously filtering different data), and the rules will not affect each other, maintaining independent operation within the task node. Frequent rule updates are also supported, offering high flexibility and applicability to different business scenarios.

[0116] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0117] Based on the same inventive concept, this application also provides a data processing apparatus for implementing the data processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more data processing apparatus embodiments provided below can be found in the limitations of the data processing method described above, and will not be repeated here.

[0118] In some embodiments, such as Figure 7 As shown, a data processing apparatus 700 is provided, including: an acquisition module 702, a task node set determination module 704, a streaming processing link determination module 706, and a data processing module 708, wherein:

[0119] The acquisition module 702 is used to acquire the data processing flow that matches the data processing task;

[0120] The task node set determination module 704 is used to determine the task node set that implements the data processing task, and the data interaction relationship between the task nodes contained in the task node set; each task node is associated with at least a part of the data processing process; at least a part of the data processing process satisfies the streaming processing conditions.

[0121] The streaming processing link determination module 706 is used to determine the streaming processing link for each task node; the streaming processing link matches at least a portion of the data processing flow associated with the task node.

[0122] The data processing module 708 is used to process the data to be processed by the data processing task according to the data interaction relationship between each task node and based on each streaming processing link to obtain the data processing result.

[0123] In one embodiment, the acquisition module 702 is specifically used to: acquire the data to be processed for the data processing task, determine the data type of the data to be processed; there are at least two types of data types; and based on the data processing task, determine the respective data processing flow for each data type of data to be processed. In this embodiment, the task node set determination module 704 is specifically used to: determine the task node corresponding to each data processing flow, and obtain a task node set including each task node.

[0124] In one embodiment, the streaming processing link determination module 706 includes: a target processing logic determination unit, configured to determine at least two target processing logics from a plurality of candidate data processing logics for each task node; the at least two target processing logics are matched with at least a portion of the data processing flow associated with the task node; and a streaming processing link determination unit, configured to determine a streaming processing link including each target processing logic.

[0125] In one embodiment, the streaming processing link determination unit is specifically used to: determine the link relationship between each target processing logic based on at least a part of the data processing flow; configure the subscription relationship between each target processing logic based on the link relationship, and create a streaming processing link including each target processing logic; and subscribe to the output data of the previous target processing logic in the streaming processing link.

[0126] In one embodiment, the data processing apparatus 700 further includes: a task update information acquisition module, configured to acquire task update information in response to a task update event for a data processing task; and a node maintenance module, configured to determine an update process that matches the task update information and perform node maintenance on each task node based on the update process.

[0127] In one embodiment, the node maintenance module is specifically used to: if there is no target update process corresponding to the target processing process in the update process, delete the target processing link matching the target processing process and deregister the target task node associated with the target processing link; if the update process includes a target update process corresponding to the target processing process, update the target processing link associated with the target task node that matches the target processing process based on the target update process; if the update process includes a new processing process, configure the new task node matching the new processing process on each task node based on the new processing process.

[0128] Each module in the aforementioned data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0129] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores the data involved in the aforementioned methods. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a data processing method.

[0130] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0131] In some embodiments, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the data processing method described above.

[0132] In some embodiments, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the data processing method described above.

[0133] In some embodiments, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the data processing method described above.

[0134] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0135] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0136] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A data processing method, characterized in that, The method includes: Obtain a data processing flow that matches the data processing task; the data processing flow does not meet the conditions for streaming processing. The set of task nodes that implement the data processing task is determined, as well as the data interaction relationship between the task nodes contained in the set of task nodes; each task node is associated with at least a part of the data processing flow; the at least a part of the data processing flow satisfies the streaming processing condition. For each task node, at least two target processing logics are determined from a plurality of candidate data processing logics; the at least two target processing logics are matched with the at least part of the data processing flow associated with the task node; Based on at least a portion of the data processing flow, determine the link relationships between each of the target processing logics; Based on the link relationship, configure the subscription relationship between each target processing logic, and create a streaming processing link including each target processing logic; in the streaming processing link, the next target processing logic subscribes to the output data of the previous target processing logic; Based on the data interaction relationships between the task nodes, the data to be processed in the data processing task is processed according to the streaming processing links to obtain the data processing results.

2. The method according to claim 1, characterized in that, The data processing flow for acquiring data that matches the data processing task includes: Obtain the data to be processed for the data processing task and determine the data type of the data to be processed; the data type has at least two types; Based on the data processing task, determine the respective data processing flow for each data type of data to be processed; The process of determining the set of task nodes that implement the data processing task includes: Determine the task node corresponding to each of the data processing flows to obtain a task node set including each of the task nodes.

3. The method according to claim 1 or 2, characterized in that, The method further includes: In response to a task update event for the data processing task, obtain task update information; Determine the update process that matches the task update information, and perform node maintenance on each task node based on the update process.

4. The method according to claim 3, characterized in that, The node maintenance of each task node based on the update process includes: If there is no target update process corresponding to the target processing process in the update process, then delete the target processing link that matches the target processing process and cancel the target task node associated with the target processing link; If the update process includes a target update process corresponding to the target processing process, then based on the target update process, the target processing link associated with the target task node that matches the target processing process is updated; If the update process includes a new processing process, then based on the new processing process, a new task node matching the new processing process is configured on each of the task nodes.

5. A data processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire a data processing flow that matches the data processing task; the data processing flow does not meet the conditions for streaming processing. The task node set determination module is used to determine the task node set that implements the data processing task, and the data interaction relationship between each task node included in the task node set; each task node is associated with at least a part of the data processing process; the at least a part of the data processing process satisfies the streaming processing conditions. The target processing logic determination unit is used to determine at least two target processing logics from a plurality of candidate data processing logics for each task node; the at least two target processing logics are matched with the at least part of the data processing flow associated with the task node; A streaming processing link determination unit is configured to determine the link relationship between each of the target processing logics based on the at least part of the data processing flow, and to configure the subscription relationship between each of the target processing logics based on the link relationship, thereby creating a streaming processing link including each of the target processing logics; wherein the next target processing logic in the streaming processing link subscribes to the output data of the previous target processing logic; The data processing module is used to process the data to be processed in the data processing task according to the data interaction relationship between each task node and based on each streaming processing link to obtain the data processing result.

6. The apparatus according to claim 5, characterized in that, The acquisition module is specifically used for: acquiring the data to be processed for the data processing task, determining the data type of the data to be processed; and, based on the data processing task, determining the data processing flow for each data type of the data to be processed; the data types are at least two. The task node set determination module is specifically used to: determine the task node corresponding to each of the data processing flows, and obtain a task node set including each of the task nodes.

7. The apparatus according to claim 5 or 6, characterized in that, The device further includes: The task update information acquisition module is used to acquire task update information in response to a task update event for the data processing task. The node maintenance module is used to determine the update process that matches the task update information, and to perform node maintenance on each task node based on the update process.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Data processing method, computer equipment and storage medium

    CN110795215A