Data processing methods, apparatus, equipment, storage media and program products
By identifying anomalous subgraphs in the dynamic graph and processing them with graph neural networks, the problem that graph updates are difficult to reflect the real-time propagation of system anomalies is solved, enabling fast and accurate anomaly detection and repair, and improving the system's autonomous maintenance efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2025-08-27
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies have shortcomings in terms of real-time performance and intelligence. The map updates are difficult to reflect the impact of real-time systems on the propagation of anomalies, and the system cannot be adjusted quickly and accurately, resulting in delayed processing of erroneous data in regulatory reporting scenarios.
By identifying anomalous subgraphs from dynamic graphs, updating the initial graph using edge weights, dynamically adjusting edge weights based on historical anomalous frequencies, identifying anomalous nodes, and determining root cause nodes, graph neural networks are used to process dynamic graphs to improve the efficiency of anomaly detection and repair.
It improves the timeliness and accuracy of dynamic graphs, enabling real-time reflection of the impact of system load on anomaly propagation, automatic identification and repair of anomaly sources, reduced manual intervention, and improved anomaly detection and repair efficiency.
Smart Images

Figure CN122133004A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of big data, artificial intelligence, and fintech, and specifically to a data processing method, apparatus, device, storage medium, and program product. Background Technology
[0002] Dynamic graph iterative update methods mostly adopt a batch processing approach, capturing incremental data through data snapshot comparison, and using dynamic update technology of graph structure to complete the graph maintenance, which can meet the data iterative update needs of a certain scale.
[0003] In the process of realizing the inventive concept of this application, it was found that the relevant technologies still have shortcomings in terms of real-time performance and intelligence. The map update is difficult to reflect the impact of the real-time system on the propagation of anomalies, and the system cannot be adjusted quickly and accurately. It cannot meet the real-time update requirements. For example, in the regulatory reporting scenario of relevant agencies, once an error is found in the batch window of extraction, transformation and loading at the end of the processing, the batch long transaction cannot be rolled back and rerun because the map cannot be updated in time. The error data can only be reported, which leads to errors. Summary of the Invention
[0004] In view of the above problems, this application provides a data processing method, apparatus, device, storage medium and program product.
[0005] According to a first aspect of this application, a data processing method is provided, comprising: determining an anomalous subgraph from a dynamic graph, wherein the dynamic graph is obtained by updating an initial graph according to edge weights, the initial graph including node data of multiple nodes and edge data of at least one edge, the nodes being storage nodes or computing nodes, the edges being used to connect two nodes with a dependency relationship, the edge weights being determined according to the historical anomalous frequency of the edges in a first predetermined time period, the anomalous subgraph including node data of anomalous nodes with an anomalous probability greater than or equal to a first predetermined threshold, node data of other nodes associated with the anomalous nodes, and edge data of the edges connecting the anomalous nodes and other nodes respectively; and determining a root cause node from the anomalous nodes and other nodes according to the anomalous contribution degree of the anomalous nodes and other nodes determined by the anomalous subgraph.
[0006] A second aspect of this application provides a data processing apparatus, comprising: an anomaly subgraph determination module, configured to determine an anomaly subgraph from a dynamic graph, wherein the dynamic graph is obtained by updating an initial graph according to edge weights, the initial graph includes node data of multiple nodes and edge data of at least one edge, the nodes being storage nodes or computing nodes, the edges connecting two nodes with a dependency relationship, the edge weights being determined based on the historical anomaly frequency of the edges in a first predetermined time period, the anomaly subgraph including node data of anomaly nodes with an anomaly probability greater than or equal to a first predetermined threshold, node data of other nodes associated with the anomaly nodes, and edge data of the edges connecting the anomaly nodes and other nodes respectively; and a root cause node determination module, configured to determine root cause nodes from the anomaly nodes and other nodes based on the anomaly contribution rates of the anomaly nodes and other nodes determined from the anomaly subgraph.
[0007] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0008] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0009] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0010] According to the data processing method, apparatus, equipment, medium, and product provided in this application, since edge weights can be determined based on the historical anomaly frequency of edges in a first predetermined time period, there is no need to manually adjust edge weights. Edge weights can be dynamically adjusted based on historical anomaly frequency, reflecting the impact of system load on anomaly propagation in real time. Because the dynamic graph is obtained by updating the initial graph based on edge weights, the timeliness and accuracy of the dynamic graph are improved. The anomaly sub-graph is determined from the dynamic graph. The anomaly sub-graph can include node data of anomaly nodes with anomaly probabilities greater than or equal to a first predetermined threshold, node data of other nodes associated with the anomaly nodes, and edge data of the edges connecting the anomaly nodes and other nodes. Based on the anomaly contribution of the anomaly nodes and other nodes determined by the anomaly sub-graph, the root cause node can be determined from the anomaly nodes and other nodes, thereby realizing the identification of the anomaly source from multiple anomaly nodes. Furthermore, since the anomaly sub-graph retains the full-link context of the root cause node, the anomaly propagation path can be viewed without manual exclusion, improving the efficiency of anomaly detection. Attached Figure Description
[0011] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0012] Figure 1 The illustration shows application scenarios of data processing methods, apparatus, devices, media, and program products according to embodiments of this application.
[0013] Figure 2 A flowchart of a data processing method according to an embodiment of this application is shown.
[0014] Figure 3 A flowchart of a method for obtaining edge weights according to an embodiment of this application is shown.
[0015] Figure 4 A structural block diagram of a data processing apparatus according to an embodiment of this application is shown.
[0016] Figure 5 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of this application is shown. Detailed Implementation
[0017] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0020] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0021] Data lineage graphs are an important tool in modern data management. By describing the flow and dependencies of data within a system, they provide support for scenarios such as data governance, data traceability, and compliance analysis. Iterative update methods for data lineage graphs refer to continuously maintaining their integrity and real-time performance through incremental capture and dynamic updates as the data environment changes.
[0022] For lineage maps, traditional methods use fixed weights, which fail to reflect the impact of real-time system load on anomaly propagation; related technologies manually adjust dynamic coefficients, which cannot quickly and accurately adjust the system; in addition, the detection and repair of anomaly nodes largely rely on manual operation, reducing the system's autonomous maintenance efficiency.
[0023] In view of this, embodiments of this application provide a data processing method to determine an abnormal sub-graph from a dynamic graph. The dynamic graph is obtained by updating an initial graph according to edge weights. The initial graph includes node data of multiple nodes and edge data of at least one edge. Nodes are storage nodes or computing nodes, and edges are used to connect two nodes with dependencies. Edge weights are determined based on the historical abnormal frequency of the edges in a first predetermined time period. The abnormal sub-graph includes node data of abnormal nodes with an abnormal probability greater than or equal to a first predetermined threshold, node data of other nodes associated with the abnormal nodes, and edge data of the edges connecting the abnormal nodes and other nodes. Based on the abnormal contribution of the abnormal nodes and other nodes determined by the abnormal sub-graph, root cause nodes are determined from the abnormal nodes and other nodes.
[0024] It should be noted that the data processing method and data processing device provided in this application can be used in the field of big data, or in any field other than big data, such as the field of artificial intelligence, the field of fintech, etc. Therefore, the application field of the data processing method and data processing device provided in this application is not limited.
[0025] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this application all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.
[0026] Figure 1 The illustration shows application scenarios of data processing methods, apparatus, devices, media, and program products according to embodiments of this application.
[0027] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0028] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication terminal applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email terminals, social media platform software, etc. (for example only).
[0029] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0030] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0031] It should be noted that the data processing method provided in the embodiments of this application can generally be executed by server 105. Correspondingly, the data processing device provided in the embodiments of this application can generally be located in server 105. The data processing method provided in the embodiments of this application can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the data processing device provided in the embodiments of this application can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0032] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0033] The following will be based on Figure 1 The described scene, through Figures 2-3 The data processing method according to the embodiments of this application will be described in detail.
[0034] Figure 2 A flowchart of a data processing method according to an embodiment of this application is shown.
[0035] like Figure 2 As shown, the data processing method of this embodiment includes operations S210 to S220.
[0036] In operation S210, anomaly subgraphs are determined from the dynamic graph. The dynamic graph is obtained by updating the initial graph based on the edge weights, which are determined based on the historical anomaly frequency of the edges in the first predetermined time period.
[0037] In operation S220, the root cause node is determined from the abnormal nodes and other nodes based on the abnormal contribution of the abnormal nodes and other nodes determined by the abnormal sub-graph.
[0038] According to embodiments of this application, the initial graph may include node data for each of multiple nodes and edge data for each of at least one edge. Nodes may be storage nodes or computation nodes. Storage nodes may store data at the physical layer level, such as libraries, tables, partitions, files, and message queues, or data at the logical layer level, such as fields, column families, features, and metrics. Computation nodes may store jobs or tasks, such as offline jobs, streaming tasks, and machine learning training tasks, or store storage services, such as microservices and algorithm models.
[0039] According to embodiments of this application, database tables, extracted task logs, Structured Query Language (SQL) scripts, Application Programming Interface (API) logs, etc., can be used as data sources to parse SQL syntax dependencies, capture real-time computing task logs, store data entities as nodes and lineage relationships as edges in a directed graph, and construct an initial graph.
[0040] According to embodiments of this application, anomaly subgraphs can be determined from a dynamic graph, which is obtained by updating an initial graph based on edge weights. The edge weights are determined based on the historical anomaly frequency of the edges during a first predetermined time period. Edges are used to connect two nodes with a dependency relationship. The anomaly subgraph may include node data of anomaly nodes with an anomaly probability greater than or equal to a first predetermined threshold, node data of other nodes associated with the anomaly nodes, and edge data of the edges connecting the anomaly nodes and other nodes respectively.
[0041] According to embodiments of this application, the root cause node can be determined from the abnormal nodes and other nodes based on the abnormal contribution degree of the abnormal nodes and other nodes determined by the abnormal sub-graph. If the abnormal contribution degree is greater than a preset threshold, the root cause node can be determined from the abnormal nodes and other nodes.
[0042] According to embodiments of this application, since edge weights can be determined based on the historical anomaly frequency of edges within a first predetermined time period, there is no need to manually adjust edge weights. The edge weights can be dynamically adjusted in conjunction with the historical anomaly frequency, reflecting the impact of system load on anomaly propagation in real time. Because the dynamic graph is obtained by updating the initial graph based on edge weights, the timeliness and accuracy of the dynamic graph are improved. The anomaly sub-graph is determined from the dynamic graph. The anomaly sub-graph can include node data of anomaly nodes with an anomaly probability greater than or equal to a first predetermined threshold, node data of other nodes associated with the anomaly nodes, and edge data of the edges connecting the anomaly nodes and other nodes. Based on the anomaly contribution of the anomaly nodes and other nodes determined by the anomaly sub-graph, the root cause node can be identified from the anomaly nodes and other nodes, thereby enabling the identification of the anomaly source from multiple anomaly nodes. Furthermore, since the anomaly sub-graph retains the full-link context of the root cause node, the anomaly propagation path can be viewed without manual exclusion, improving the efficiency of anomaly detection.
[0043] Figure 3 A flowchart of a method for obtaining edge weights according to an embodiment of this application is shown.
[0044] like Figure 3 As shown, this embodiment includes operations S310 to S330.
[0045] In operation S310, for any one of the edges, the computational complexity of the edge is obtained based on the edge's computational type weight and the sum of the in-degree and out-degree of the two nodes connected by the edge.
[0046] In operation S320, the historical anomaly frequency of the edge is obtained based on the total number of runs in the first predetermined time period and the historical anomaly count of the two nodes connected by the edge.
[0047] In operation S330, the edge weights are obtained based on the computational complexity of the edges and the frequency of historical anomalies.
[0048] According to embodiments of this application, the computation type may include basic computation or window functions, and weights can be assigned according to the importance of the computation type. For any edge in at least one context, the computational complexity of the edge is obtained based on the weight of the edge's computation type and the sum of the in-degrees of the two nodes connected to the edge. For example, if the edge's computation type is basic computation, its weight is 0.2, and the sum of the in-degrees of the two nodes connected to the edge is 5, then the computational complexity of the edge is 1.
[0049] According to an embodiment of this application, the historical anomaly frequency of an edge can be obtained based on the total number of runs in a first predetermined time period and the historical anomaly count of the two nodes connected by the edge. For example, if the edge has a total of 10 runs in the first predetermined time period and the historical anomaly count of the two nodes connected by the edge is 5, then the historical anomaly frequency of the edge is the ratio of the historical anomaly count to the total number of runs, and this ratio is 0.5.
[0050] According to an embodiment of this application, the edge weight can be obtained based on the computational complexity of the edge and the frequency of historical anomalies, as shown in formula (1).
[0051] (1);
[0052] in, This represents the edge weight between node i and node j. and Values can be assigned to indicate the degree of importance. This can represent the computational complexity of the edge between node i and node j. This represents the historical frequency of anomalies. In formula (1), and You can set it according to your actual needs; there are no restrictions here.
[0053] According to embodiments of this application, by utilizing historical anomaly frequency and the computational complexity of edges, edge weights can be updated, making the prediction of anomaly propagation paths more closely match the actual system state and improving the accuracy of predicting anomaly nodes.
[0054] According to an embodiment of this application, determining an anomalous sub-graph from a dynamic graph includes: determining anomalous nodes from a plurality of nodes included in the dynamic graph based on prediction results obtained by processing the dynamic graph using a graph neural network; and determining anomalous sub-graphs from the dynamic graph based on the anomalous nodes.
[0055] According to embodiments of this application, abnormal nodes can be identified from multiple nodes in a dynamic graph based on prediction results obtained by processing the dynamic graph using a graph neural network. All downstream nodes, including the abnormal nodes, and the connecting edges can be identified as abnormal subgraphs from the dynamic graph.
[0056] According to embodiments of this application, graph neural networks can capture the spatiotemporal evolution rules of node features and topological structures in dynamic graphs, identify abnormal patterns that are difficult to capture in traditional static methods, and improve the detection rate of abnormal nodes.
[0057] According to an embodiment of this application, based on the prediction results obtained by processing a dynamic graph using a graph neural network, abnormal nodes are determined from multiple nodes included in the dynamic graph. This includes: obtaining a target dynamic graph related to regulatory reporting, the target dynamic graph including multiple nodes, the categories of which include table fields, monitoring indicators, and regulatory reports; determining the node features of multiple different categories of nodes and the edge features of at least one edge based on the node data of each node and the edge data of each edge of each of the multiple different categories of nodes included in the target dynamic graph; inputting the node features of each node and the edge features of each edge of each of the multiple different categories of nodes into a graph neural network to obtain the abnormal probability of each of the multiple different categories of nodes; for any node among the multiple different categories of nodes, if the abnormal probability of the node is greater than or equal to a first predetermined threshold, the node is designated as an abnormal node. According to an embodiment of this application, for example, if the Extract, Load, Transform (ETL) batch processing window is 01:00–04:00, and the indicator is abnormal at 03:30, it cannot be rerun due to insufficient time. Abnormal source table data quality can lead to a 0.3% underestimation of the capital adequacy ratio. The regulatory system will automatically compare this at 08:30 AM, resulting in rectification suggestions. The supplementary data entry and manual review take a total of 36 hours, and the delayed disclosure is recorded in the regulatory file. However, the method for identifying abnormal nodes described in this application can rewrite the dynamic graph and automatically recalculate indicators when the probability of an anomaly is greater than or equal to a first predetermined threshold, reducing the service level risk to zero. Specifically, it can obtain a target dynamic graph related to the regulatory reporting scenario. Nodes in the target dynamic graph can include different categories of nodes related to field-level quality probes, table fields, regulatory indicators, messages, and scheduling tasks. Edges can include real-time timestamp attributes and edge propagation directions. Based on the node data of multiple different categories of nodes and the edge data of at least one edge included in the target dynamic graph, the node features of each of the multiple different categories of nodes and the edge features of at least one edge are determined. These node features and edge features are then input into a graph neural network to obtain the anomaly probabilities of each of the multiple different categories of nodes. For any node among the multiple different categories of nodes, if the node's anomaly probability is greater than or equal to a first predetermined threshold, the node can be identified as an anomaly node. By tracking anomaly fields, indicators, and reports in real time, the problems of traditional batch processing are avoided, preventing the embarrassment of discovering anomalies only at 3:30 AM and not having enough time to rerun.
[0058] According to embodiments of this application, based on the node data of each of the multiple nodes and the edge data of at least one edge included in the dynamic graph, the node features of each of the multiple nodes and the edge features of at least one edge can be determined. Inputting the node features of each of the multiple nodes and the edge features of at least one edge into a graph neural network yields the anomaly probabilities of each of the multiple nodes. A node is identified as an anomalous node if its anomaly probability is greater than or equal to a first predetermined threshold. For example, if the first predetermined threshold is set to 0.7, nodes with anomaly probabilities greater than 0.7 can be marked as anomalous nodes.
[0059] According to the embodiments of this application, the node features and edge features in the dynamic graph can be processed based on the graph artifact network to obtain the anomaly probability of each of the multiple nodes, accurately quantify node anomalies, and if the anomaly probability of a node is greater than or equal to a first predetermined threshold, the node can be regarded as an anomaly node, thereby improving the interpretability of node anomalies.
[0060] According to embodiments of this application, the data processing method further includes: performing repair using a repair strategy that matches the anomaly type of the root cause node.
[0061] According to embodiments of this application, an anomaly type label can be assigned to the root cause node. Common anomaly types include parameter errors, version incompatibility, downstream timeouts, certificate expiration, and domain name system resolution failures. Based on the mapping table between anomaly types and remediation strategies, after locating the root cause node, the corresponding remediation strategy can be automatically matched and then remediated.
[0062] According to the embodiments of this application, by matching the anomaly type of the root cause node with a predefined repair strategy library, accurate mapping is achieved, which improves repair efficiency, enables timely loading of matching repair strategies, and reduces the delay time of matching degradation strategies.
[0063] According to embodiments of this application, node features include at least one of the following: out-degree, in-degree, out-in-degree, dependency depth, edge weight associated with the node, anomaly type label, or historical anomaly frequency during a first predetermined time period. Dependency depth characterizes the node's level in the dynamic graph, and anomaly type label characterizes the anomaly type of the node. Edge features include at least one of the following: edge weight, edge weight change, or edge propagation direction. Edge weight change characterizes the change in edge weight during a second predetermined time period.
[0064] According to embodiments of this application, node characteristics may include out-degree, in-degree, sum of in-degree and out-degree, dependency depth, edge weights associated with the node, anomaly type labels, or historical anomaly frequency during a first predetermined time period. Dependency depth characterizes the node's hierarchy in the dynamic graph, and the anomaly type label characterizes the type of anomaly the node exhibits. Edge characteristics may include edge weights, edge weight changes, or edge propagation directions. Edge weight changes characterize the change in edge weights during a second predetermined time period. Anomaly type labels may further include data quality anomaly labels, transformation logic anomaly labels, and performance indicator anomaly labels.
[0065] According to embodiments of this application, common error types in the data lake process may include data quality anomaly tags, transformation logic anomaly tags, source access anomaly tags, lineage dependency anomaly tags, resource performance anomaly tags, and release and rollback anomaly tags. Data quality anomaly tags may include a null value rate exceeding a threshold for key fields, duplicate business primary keys, and invalid date formats. Transformation logic anomaly tags may include query statement syntax errors, User Datagram Protocol (UDP) null pointer exceptions, and partition fields not being pushed down. Source access anomaly tags may include source database connection failures and upstream table structure drift. Lineage dependency anomaly tags may include upstream task failures and downstream task idle runs. Resource performance anomaly tags may include insufficient queue resources, task inability to complete, or disk damage resulting in unreadable files. Release and rollback anomaly tags may include version number conflicts during gray-scale rollout and missing permissions in scripts. According to embodiments of this application, by introducing multiple different types of data, the node features in the dynamic graph are enriched, which helps improve prediction accuracy when using the dynamic graph to detect abnormal nodes.
[0066] According to embodiments of this application, determining a root cause node from abnormal nodes and other nodes based on their respective abnormal contribution degrees determined by the abnormal subgraph includes: for any node in the abnormal subgraph, whether the node is an abnormal node or another node, determining at least one connecting node associated with the node and the edges between the node and the at least one connecting node; for any connecting node among the at least one connecting node, determining the intermediate abnormal contribution degree of the connecting node to the node based on the abnormal contribution degree of the connecting node, the out-degree of the connecting node, the abnormal probability of the node, and the edge weight of the edge used for the edge between the connecting nodes; determining the abnormal contribution degree of the node based on the intermediate abnormal contribution degree of the at least one connecting node to the node; and identifying the node corresponding to an abnormal contribution degree greater than or equal to a second predetermined threshold among a plurality of abnormal contribution degrees as the root cause node.
[0067] According to an embodiment of this application, for any node in the abnormal subgraph, whether the node is an abnormal node or another node, at least one connecting node associated with the node and the edge between the node and at least one connecting node are determined from the abnormal subgraph; for any connecting node among the at least one connecting node, the intermediate abnormal contribution of the connecting node to the node is determined based on the abnormal contribution degree of the connecting node, the out-degree of the connecting node, the abnormal probability of the node, and the edge weight of the edge used for the edge between the connecting node and the connecting node; the abnormal contribution degree of the node is determined based on the intermediate abnormal contribution degree of at least one connecting node to the node, as shown in formula (2).
[0068] (2);
[0069] in, The abnormal contribution of node u. This represents the damping coefficient, which can be 0.85. Let v be the edge weight of the edge vu, where v represents the node connected to node u via the edge. Indicates the out-degree of node v. This represents the anomaly probability of node u. This represents the abnormal contribution of node v.
[0070] For example, calculate the anomalous contribution of the first node in an anomalous subgraph. The first node is connected to the second node via a second edge, and the first node is connected to the third node via a third edge. For the second node, its intermediate anomalous contribution to the first node can be determined based on its anomalous contribution, out-degree, anomalous probability, and edge weight. Similarly, for the third node, its intermediate anomalous contribution to the first node can be determined based on its anomalous contribution, out-degree, anomalous probability, and edge weight. Finally, the anomalous contribution of the first node can be determined based on the intermediate anomalous contributions of the second and third nodes.
[0071] According to embodiments of this application, a node corresponding to an abnormal contribution degree greater than or equal to a second predetermined threshold among a plurality of abnormal contribution degrees can be used as a root cause node.
[0072] According to an embodiment of this application, the abnormal contribution of a node can be determined by calculating the sum of the intermediate abnormal contributions of other nodes connected to the node. If the abnormal contribution is greater than or equal to a second predetermined threshold, the root cause node can be located. Since the abnormal sub-graph preserves the full-link context of the root cause node, the abnormal propagation path can be viewed without manual exclusion, thus improving the efficiency of abnormal detection.
[0073] According to embodiments of this application, repair is performed using a repair strategy that matches the anomaly type of the root cause node, including: when the anomaly type indicates that the root cause node is a computational logic anomaly, repairing the current version code based on the current version code associated with the root cause node and the reference version code adjacent to the current version code; and when the anomaly type indicates that the root cause node is a data quality anomaly, executing a data completion task associated with the root cause node.
[0074] According to embodiments of this application, when the root cause of the anomaly is a computational logic anomaly, the differences between the most recent version of the code and the current version can be displayed, a rollback command can be generated, and the user can manually decide whether to automatically roll back. In the event of a rollback, the current version of the code is repaired.
[0075] According to embodiments of this application, when the root cause of the anomaly is data quality anomaly, a data completion task can be triggered. For example, this could involve re-running the incremental data from the previous day or inserting data consistency checkpoints.
[0076] According to the embodiments of this application, by accurately matching the anomaly type with the repair strategy, a closed loop from root cause to repair is achieved, thereby improving the processing efficiency of repairing root cause nodes.
[0077] According to an embodiment of this application, the data processing method further includes: performing isolation operations on abnormal nodes.
[0078] According to the embodiments of this application, by performing isolation operations on abnormal nodes, the abnormal nodes can be cut off from data calls and other functions with upstream and downstream, and requests or other tasks can be stopped from being sent to the abnormal nodes, thereby saving resources such as memory.
[0079] According to an embodiment of this application, performing isolation operations for abnormal nodes includes: when the abnormal node is a data node, setting the access permission identifier of the data partition associated with the abnormal node to a read-only state; switching the access path from the access path of the data partition to the access path of the backup data partition associated with the data partition; and when the abnormal node is a compute node, suspending the execution of downstream tasks associated with the abnormal node.
[0080] According to an embodiment of this application, when the abnormal node is a data node, the access permission identifier of the data partition associated with the abnormal node can be set to read-only, and the access path can be switched from the access path of the data partition to the access path of the backup data partition associated with the data partition; when the abnormal node is a compute node, the processing of other tasks including the abnormal node and its downstream can be suspended.
[0081] According to the embodiments of this application, different operations can be taken according to the type of abnormal node, which improves flexibility, avoids the suspension of the entire downstream task due to data node errors, and improves processing efficiency.
[0082] Based on the above data processing method, this application also provides a data processing apparatus. The following will be combined with... Figure 4 The device is described in detail.
[0083] Figure 4 A structural block diagram of a data processing apparatus according to an embodiment of this application is shown.
[0084] like Figure 4 As shown, the data processing device 400 in this embodiment includes an anomaly sub-map determination module 410 and a root cause node determination module 420.
[0085] The abnormal subgraph determination module 410 is used to determine the abnormal subgraph from the dynamic graph. The dynamic graph is obtained by updating the initial graph according to the edge weights. The initial graph includes the node data of multiple nodes and the edge data of at least one edge. The nodes are storage nodes or computing nodes. The edges are used to connect two nodes with a dependency relationship. The edge weights are determined according to the historical abnormal frequency of the edges in a first predetermined time period. The abnormal subgraph includes the node data of abnormal nodes with an abnormal probability greater than or equal to a first predetermined threshold, the node data of other nodes associated with the abnormal nodes, and the edge data of the edges that connect the abnormal nodes and other nodes respectively.
[0086] The root cause node determination module 420 is used to determine the root cause node from the abnormal nodes and other nodes based on the abnormal contribution of the abnormal nodes determined by the abnormal sub-graph and other nodes respectively.
[0087] According to the data processing method, apparatus, equipment, medium, and product provided in this application, anomaly sub-graphs can be determined from a dynamic graph. The dynamic graph is obtained by updating the initial graph based on edge weights. The edge weights are determined based on the historical anomaly frequency of the edges in a first predetermined time period, without the need for manual adjustment of the edge weights. The historical anomaly sub-graph can include node data of anomaly nodes with anomaly probabilities greater than or equal to a first predetermined threshold, node data of other nodes associated with the anomaly nodes, and edge data of the edges connecting the anomaly nodes and other nodes respectively. Based on the anomaly contribution of the anomaly nodes and other nodes determined by the anomaly sub-graph, the root cause node can be determined from the anomaly nodes and other nodes without the need for manual adjustment of the edge weights. The edge weights can be dynamically adjusted in combination with the historical anomaly frequency, which can reflect the impact of system load on anomaly propagation in real time, thereby automatically predicting anomaly nodes, reducing the deviation caused by manual intervention, and improving the accuracy of prediction.
[0088] According to an embodiment of this application, the anomaly subgraph determination module 410 includes: a computational complexity submodule, a historical anomaly frequency submodule, and an edge weight acquisition submodule.
[0089] The computational complexity submodule is used to calculate the computational complexity of any edge, based on the edge's computational type weight and the sum of the in-degree and out-degree of the two nodes connected by the edge.
[0090] The historical anomaly frequency submodule is used to obtain the historical anomaly frequency of any edge for at least one edge, based on the total number of runs in a first predetermined time period and the number of historical anomalies of the two nodes connected to the edge.
[0091] The edge weighting submodule is used to obtain edge weights based on the computational complexity of the edges and the frequency of historical anomalies.
[0092] According to an embodiment of this application, the abnormal sub-map determination module 410 includes: an abnormal node determination sub-module and an abnormal sub-map determination sub-module.
[0093] The abnormal node identification submodule is used to identify abnormal nodes from multiple nodes in the dynamic graph based on the prediction results obtained by processing the dynamic graph using a graph neural network.
[0094] The Abnormal Subgraph Determination Submodule is used to determine abnormal subgraphs from the dynamic graph based on abnormal nodes.
[0095] According to an embodiment of this application, the abnormal node determination submodule includes: a feature determination unit, an abnormal probability unit, and an abnormal node unit.
[0096] The acquisition unit is used to acquire the target dynamic graph related to regulatory reporting. The target dynamic graph includes multiple nodes, and the categories of target nodes include table fields, monitoring indicators, and regulatory reports.
[0097] The feature determination unit is used to determine the node features of the nodes of the different categories and the edge features of the at least one edge based on the node data of the nodes of the different categories and the edge data of the at least one edge included in the target dynamic graph.
[0098] The anomaly probability unit is used to input the node features of multiple different categories of nodes and the edge features of at least one edge into the graph neural network to obtain the anomaly probabilities of multiple different categories of nodes.
[0099] An abnormal node unit is used to identify any node among multiple different types of nodes as an abnormal node if the abnormal probability of the node is greater than or equal to a first predetermined threshold.
[0100] Node characteristics include at least one of the following: out-degree, in-degree, out-in-degree, dependency depth, edge weight associated with the node, anomaly type label or historical anomaly frequency in a first predetermined period. Dependency depth represents the node’s level in the dynamic graph, and anomaly type label represents the anomaly type of the node. Anomaly type labels include data quality anomaly labels, transformation logic anomaly labels and performance indicator anomaly labels.
[0101] The root cause node determination module 420 includes: edge units, intermediate anomaly contribution units, anomaly contribution units, and root cause node units.
[0102] An edge element is used to determine, for any node in an anomalous subgraph (whether the node is an anomalous node or another node), at least one connected node associated with the node and the edges between the node and at least one connected node.
[0103] The intermediate anomaly contribution unit is used to determine the intermediate anomaly contribution of a node to a node for any of the at least one connected nodes, based on the node's anomaly contribution, the node's out-degree, the node's anomaly probability, and the edge weight of the edge between the connected nodes.
[0104] Anomaly contribution unit is used to determine the abnormal contribution of a node based on the intermediate abnormal contribution of at least one connected node to the node.
[0105] The root cause node unit is used to identify the node corresponding to the abnormal contribution degree that is greater than or equal to a second predetermined threshold among a plurality of abnormal contribution degrees as the root cause node.
[0106] The data processing apparatus 400 in this embodiment includes a repair module.
[0107] The repair module is used to perform repairs using a repair strategy that matches the anomaly type of the root cause node.
[0108] The repair module includes: a first repair submodule and a second repair submodule.
[0109] The first repair submodule is used to repair the current version code when the root cause node is identified as a computational logic exception. This is done by using the current version code associated with the root cause node and the reference version code adjacent to the current version code.
[0110] The second repair submodule is used to perform data completion tasks associated with the root cause node when the anomaly type characterizes the root cause node as a data quality anomaly.
[0111] The data processing apparatus 400 in this embodiment includes an isolation module.
[0112] The isolation module is used to perform isolation operations on abnormal nodes.
[0113] The isolation module includes: a settings submodule, a toggle submodule, and a pause submodule.
[0114] The configuration submodule is used to set the access permission flag of the data partition associated with the abnormal node to read-only when the abnormal node is a data node.
[0115] The switching submodule is used to switch the access path from the data partition's access path to the access path of the backup data partition associated with the data partition.
[0116] The pause submodule is used to pause the execution of downstream tasks associated with the abnormal node when the abnormal node is a compute node.
[0117] According to embodiments of this application, any plurality of modules in the anomaly sub-map determination module 410 and the root cause node determination module 420 can be merged into one module, or any one of these modules can be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules can be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the anomaly sub-map determination module 410 and the root cause node determination module 420 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any appropriate combination of any of these three implementation methods. Alternatively, at least one of the anomaly sub-map determination module 410 and the root cause node determination module 420 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0118] Figure 5 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of this application is shown.
[0119] like Figure 5As shown, an electronic device 500 according to an embodiment of this application includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.
[0120] RAM 503 stores various programs and data required for the operation of electronic device 500. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Processor 501 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 502 and / or RAM 503. It should be noted that programs may also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in one or more memories.
[0121] According to embodiments of this application, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to a bus 504. The electronic device 500 may also include one or more of the following components connected to the input / output (I / O) interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.
[0122] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the data processing method according to the embodiments of this application.
[0123] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 502 and / or RAM 503 and / or one or more memories other than ROM 502 and RAM 503 described above.
[0124] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the data processing methods provided in the embodiments of this application.
[0125] When the computer program is executed by the processor 501, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0126] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 509, and / or installed from a removable medium 511. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0127] In such an embodiment, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by processor 501, it performs the functions defined in the system of this application embodiment. According to embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0128] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0130] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
Claims
1. A data processing method, characterized in that, include: An anomalous subgraph is determined from a dynamic graph, which is obtained by updating an initial graph based on edge weights. The initial graph includes node data for multiple nodes and edge data for at least one edge. The nodes are storage nodes or computing nodes, and the edges are used to connect two nodes with a dependency relationship. The edge weights are determined based on the historical anomalous frequency of the edges in a first predetermined time period. The anomalous subgraph includes node data for anomalous nodes with an anomalous probability greater than or equal to a first predetermined threshold, node data for other nodes associated with the anomalous nodes, and edge data for the edges connecting the anomalous nodes and the other nodes. Based on the abnormal contribution of the abnormal nodes and other nodes determined by the abnormal sub-graph, the root cause node is determined from the abnormal nodes and other nodes.
2. The method according to claim 1, characterized in that, The edge weights are obtained in the following way: For any one of the at least one edges, The computational complexity of the edge is obtained based on the weight of the edge's computational type and the sum of the in-degree and out-degree of the two nodes connected by the edge. The historical anomaly frequency of the edge is obtained based on the total number of runs during the first predetermined time period and the historical anomaly counts of the two nodes connected by the edge. The edge weight is obtained based on the computational complexity of the edge and the frequency of historical anomalies.
3. The method according to claim 1 or 2, characterized in that, The step of determining anomalous sub-maps from dynamic maps includes: Based on the prediction results obtained by processing the dynamic graph using a graph neural network, the abnormal node is determined from the multiple nodes included in the dynamic graph; The abnormal sub-graph is determined from the dynamic graph based on the abnormal node.
4. The method according to claim 3, characterized in that, The step of determining the abnormal node from multiple nodes included in the dynamic graph based on the prediction result obtained by processing the dynamic graph using a graph neural network includes: Obtain a dynamic target graph related to regulatory reporting. The dynamic target graph includes multiple nodes, and the categories of the target nodes include table fields, monitoring indicators, and regulatory reports. Based on the node data of each of the multiple different categories of nodes and the edge data of each of the at least one edge included in the target dynamic graph, determine the node features of each of the multiple different categories of nodes and the edge features of each of the at least one edge. The node features of the nodes of the multiple different categories and the edge features of the at least one edge are input into the graph neural network to obtain the anomaly probabilities of the nodes of the multiple different categories. For any node among the plurality of different categories of nodes, if the abnormal probability of the node is greater than or equal to the first predetermined threshold, the node is designated as the abnormal node.
5. The method according to claim 4, characterized in that, The node features include at least one of the following: out-degree, in-degree, out-in-degree, dependency depth, edge weight associated with the node, anomaly type label, or historical anomaly frequency during the first predetermined time period. The dependency depth represents the level of the node in the dynamic graph, and the anomaly type label represents the anomaly type of the node. The anomaly type label includes data quality anomaly label, transformation logic anomaly label, and performance indicator anomaly label. The edge features include at least one of the following: edge weight, edge weight change, or edge propagation direction, wherein the edge weight change represents the change in edge weight during a second predetermined time period.
6. The method according to claim 1 or 2, characterized in that, The step of determining the root cause node from the abnormal nodes and other nodes based on the abnormal contribution values of the abnormal nodes and other nodes determined by the abnormal sub-graph includes: For any node in the anomalous subgraph, the node is either the anomalous node or one of the other nodes. From the anomalous subgraph, determine at least one connecting node associated with the node and the edges between the node and each of the at least one connecting node; For any of the at least one connected nodes, the intermediate abnormal contribution of the connected node to the node is determined based on the abnormal contribution of the connected node, the out-degree of the connected node, the abnormal probability of the node, and the edge weight of the edge used to connect the node and the connected node. The abnormal contribution of a node is determined based on the intermediate abnormal contribution of each of the at least one connected node to the node. The node corresponding to the abnormal contribution value that is greater than or equal to the second predetermined threshold among the multiple abnormal contribution values is taken as the root cause node.
7. The method according to claim 1 or 2, characterized in that, Also includes: Repair is performed using a repair strategy that matches the anomaly type of the root cause node.
8. The method according to claim 7, characterized in that, The repair strategy, which uses a repair strategy that matches the anomaly type of the root cause node, includes: When the anomaly type indicates that the root cause node is a computational logic anomaly, the current version code is repaired based on the current version code associated with the root cause node and the reference version code adjacent to the current version code; If the anomaly type indicates that the root cause node is a data quality anomaly, a data completion task associated with the root cause node is executed.
9. The method according to claim 1 or 2, characterized in that, Also includes: Perform isolation operations on the abnormal node.
10. The method according to claim 9, characterized in that, The isolation operation performed on the abnormal node includes: In the case where the abnormal node is the data node. Set the access permission flag of the data partition associated with the abnormal node to read-only; Switch the access path from the access path of the data partition to the access path of the backup data partition associated with the data partition; If the abnormal node is the computing node, the execution of downstream tasks associated with the abnormal node is suspended.
11. A data processing apparatus, characterized in that, The device includes: An abnormal subgraph determination module is used to determine an abnormal subgraph from a dynamic graph. The dynamic graph is obtained by updating an initial graph based on edge weights. The initial graph includes node data of multiple nodes and edge data of at least one edge. The nodes are storage nodes or computing nodes. The edges are used to connect two nodes with a dependency relationship. The edge weights are determined based on the historical abnormal frequency of the edges in a first predetermined time period. The abnormal subgraph includes node data of abnormal nodes with an abnormal probability greater than or equal to a first predetermined threshold, node data of other nodes associated with the abnormal nodes, and edge data of the edges connected to the abnormal nodes and the other nodes respectively. The root cause node determination module is used to determine the root cause node from the abnormal nodes and the other nodes based on the abnormal contribution degree of the abnormal nodes and other nodes determined by the abnormal sub-graph.
12. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 10.
14. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 10.