Data governance early warning processing method and device based on blood relationship analysis, equipment and storage medium

By establishing a data blood relationship model and real-time monitoring of data flow, combined with automated early warning and processing mechanisms, the problem of insufficient accuracy and real-time data blood relationship analysis in the existing technology is solved, and efficient and accurate data governance is achieved, and risks and error risks are reduced.

CN120180019APending Publication Date: 2025-06-20WUHAN HONGXIN TECH SERVICE CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510129234.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing data blood relationship analysis methods are insufficient in terms of accuracy and real-timeness. They cannot fully accurately track and represent the complex relationships between data, and cannot reflect changes in data quality or blood relationships in real time, resulting in the inability to issue early warnings in a timely manner.

Method used

By analyzing the metadata information of the target data, a blood relationship model between data is established, including the complete link between data from generation to transmission and consumption, and the blood relationship between data is continuously stored and dynamically updated based on the graph database, reflecting the changing state of data flow in real time. At the same time, combined with preset warning rules, data blood relationship analysis, statistical methods and machine learning, automatic triggering of early warning and data quality detection can be achieved, abnormal data can be obtained and traced.

Benefits of technology

It significantly improves the accuracy and efficiency of data governance, can quickly locate the source of data problems, reduce repair time, and reduce risks. The automated early warning and processing mechanism reduces manual intervention, reduces the risks of errors and underreports, and enhances the security and credibility of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180019A_ABST
    Figure CN120180019A_ABST
Patent Text Reader

Abstract

The invention discloses a data governance early warning processing method based on blood relationship analysis. The method comprises the following steps: establishing a blood relationship model between data; continuously storing and dynamically updating the blood relationship between the data based on the graph database; a preset early warning rule is bound with the target data, and when the target data is abnormal, early warning is automatically triggered; detecting the data quality, the blood relationship link and the flow trend to obtain abnormal data; performing abnormal data tracing, obtaining a relation graph of corresponding unprocessed nodes from a graph database, and obtaining a tracing result; and based on a traceability result, performing early warning task processing, performing data restoration and processing, and automatically performing data quality reverse checking. The invention further discloses a blood relationship analysis-based data governance early warning processing device, corresponding equipment and a storage medium. According to the method provided by the invention, the data quality problem can be monitored and early warned more accurately, the data problem can be quickly responded and processed, and the potential risk and loss are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data governance, and more specifically, to a data governance early warning processing method, device, equipment and storage medium based on blood relationship analysis. Background Technique

[0002] In the context of the big data era, data has characteristics such as a large volume of information, complex relationships, diverse types, strong timeliness, complex processing processes, and many data flow links. Data problems are becoming increasingly prominent. When data problems occur, it is necessary to locate the root cause of the problem and solve the problem at once. Existing technologies often use methods based on data blood relationship analysis to locate the source of problem data. Data blood relationship refers to the relationship between data generated during the processing and flow of data. The method based on data blood relationship analysis provides a means to explore data relationships and is used to track the data flow path.

[0003] However, there are some problems with existing data blood relationship analysis methods: the accuracy problem of data blood relationship analysis. Current data blood relationship analysis tools may not be able to completely and accurately track and represent the complex relationships between data. In large distributed systems, the data flow path may be very complex and often changing, which may lead to deviations in the results of blood relationship analysis; the real-time problem of early warning. Many existing early warning systems cannot reflect changes in data quality or blood relationship in real time, which may cause early warnings not to be issued in time when problems occur, thus increasing the difficulty and risk of data governance. Summary of the Invention

[0004] In view of at least one defect or improvement requirement of the existing technology, the present invention provides a data governance early warning processing method, device, equipment and storage medium based on blood relationship analysis, which can solve at least one of the problems existing in the above background technology.

[0005] To achieve the above object, according to the first aspect of the present invention, a data governance early warning processing method based on blood relationship analysis is provided. The method includes:

[0006] Analyze the metadata information of the target data and establish a blood relationship model between the data. The blood relationship model includes the complete link of the data from generation to transmission and consumption;

[0007] Based on the blood relationship model, continuously store and dynamically update the blood relationship between the data based on the graph database, and reflect the changing state of data flow in real time;

[0008] Bind the preset early warning rules to the target data, and automatically trigger an early warning when the target data is abnormal;

[0009] Detect data quality, lineage links, and traffic trends based on data lineage analysis, statistical methods, and machine learning to obtain abnormal data;

[0010] Based on the obtained abnormal data, perform abnormal data tracing, obtain the relationship graph of the corresponding unprocessed nodes from the graph database, perform recursive search, and obtain the tracing result;

[0011] Based on the tracing result, execute the warning task disposal, perform data repair and processing, submit the processing result, and automatically execute data quality backcheck.

[0012] Furthermore, for the above data governance warning processing method based on lineage analysis, analyze the metadata information of the target data and establish a data lineage relationship model. The data lineage relationship model includes the complete link of data from generation to transmission and consumption, specifically including:

[0013] Parse the meta-information of the data based on Antlr4;

[0014] Based on data mapping, obtain the complete link of data from generation, transmission to consumption, including data flow direction and data derivation information;

[0015] Based on event-driven, obtain the incremental update of the data lineage relationship model.

[0016] Furthermore, for the above data governance warning processing method based on lineage analysis, continuously store and dynamically update the data lineage relationship between the data based on the graph database, including constructing the data lineage relationship model into a graph data structure and recording the change history of the data lineage relationship.

[0017] Furthermore, for the above data governance warning processing method based on lineage analysis, bind the preset warning rules to the target data to achieve automatic warning when the target data is abnormal, specifically including:

[0018] The preset warning rules include static rules, dynamic rules, and hybrid rules;

[0019] The static rules define the warning criteria based on traditional advance or logical conditions; the dynamic rules identify abnormal data by using statistical analysis, machine learning, and graph algorithm techniques; the hybrid rules are a combination of static rules and dynamic rules through composite conditions.

[0020] Furthermore, for the above data governance warning processing method based on lineage analysis, detect data quality, lineage links, and traffic trends based on data lineage analysis, statistical methods, and machine learning to obtain abnormal data, specifically including:

[0021] The detection of data quality detects the key quality indicators of data; the blood relationship link detection detects abnormal data in the blood relationship link by constructing a data blood relationship graph and combining graph algorithms; the traffic trend detection analyzes the abnormal changes in the dynamic changes of data traffic to identify abnormal traffic points or trends.

[0022] Further, for the above data governance warning processing method based on blood relationship analysis, based on the obtained abnormal data, trace the source of the abnormal data, obtain the relationship graph of the corresponding unprocessed nodes from the graph database, perform recursive search, and obtain the traceability result, which specifically includes:

[0023] Obtain the blood relationship link graph related to the current abnormal task or attribute from the graph database;

[0024] Starting from the current abnormal data node, recursively search for upstream and downstream nodes layer by layer to confirm whether there are data processing problems in the upstream data and downstream data in the data stream;

[0025] Generate a warning task based on the obtained abnormal source, and give repair suggestions based on the abnormal type.

[0026] Further, for the above data governance warning processing method based on blood relationship analysis, based on the traceability result, execute the warning task disposal, perform data repair and processing, submit the processing result, and automatically execute the data quality backcheck, which specifically includes:

[0027] Perform data repair and processing based on the analysis result, repair suggestions and blood relationship of the abnormal data;

[0028] After the task disposal is completed, submit the processing result and automatically execute the data quality backcheck;

[0029] Based on the blood relationship chain in the graph database, judge whether all relevant data nodes of the abnormal data have been repaired;

[0030] If it is found that one or more abnormal data have not been repaired, the system will mark them as pending processing and reassign tasks.

[0031] According to the second aspect of the present invention, there is also provided a data governance warning processing device based on blood relationship analysis, including:

[0032] A modeling module for analyzing the metadata information of target data and establishing a blood relationship model between data, where the blood relationship model includes the complete link of data from generation to transmission and consumption;

[0033] A storage module for continuously storing and dynamically updating the blood relationship between data based on the graph database according to the blood relationship model, and reflecting the changing state of data flow in real time;

[0034] A rule module, used to bind preset warning rules to target data, and automatically trigger a warning when the target data is abnormal;

[0035] An auditing module, used to detect data quality, lineage links, and traffic trends based on data lineage analysis, statistical methods, and machine learning to obtain abnormal data;

[0036] An analysis module, used to trace the source of the abnormal data based on the obtained abnormal data, obtain the relationship graph of the corresponding unprocessed nodes from the graph database, perform recursive search, and obtain the tracing result;

[0037] A processing module, used to execute warning task handling based on the tracing result, perform data repair and processing, submit the processing result, and automatically execute data quality backcheck.

[0038] According to the third aspect of the present invention, there is also provided a data governance warning processing device based on lineage analysis, which includes at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the processing unit executes the steps of the method described in any one of the above.

[0039] According to the fourth aspect of the present invention, there is also provided a storage medium, which stores a computer program executable by a data governance warning processing device based on lineage analysis. When the computer program runs on the data governance warning processing device based on lineage analysis, the data governance warning processing device based on lineage analysis executes the steps of the method described in any one of the above.

[0040] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0041] The data governance warning processing method provided by the present invention significantly improves the accuracy and efficiency of data governance by establishing a data lineage relationship model and real-time monitoring of data flow, can quickly locate the source of data problems, reduce repair time, reduce risks, and the automated warning and processing mechanism reduces manual intervention and reduces the risks of errors and missed reports. At the same time, it enhances the security and credibility of data, providing an efficient and accurate solution for data governance in the big data environment. Description of the Drawings

[0042] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0043] Figure 1 It is a schematic flowchart of the data governance early warning processing method based on blood relationship analysis provided by the embodiment of the present application;

[0044] Figure 2 It is a schematic structural diagram of the data governance early warning processing device based on blood relationship analysis provided by the embodiment of the present application. Detailed implementation manners

[0045] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following further details the present invention in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0046] The terms "first", "second", "third", etc. in the specification, claims and drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0047] Figure 1 It is a schematic flowchart of the data governance early warning processing method based on blood relationship analysis provided by the embodiment of the present application. As Figure 1 shown, the data governance early warning processing method based on blood relationship analysis provided by the embodiment of the present application includes the following steps:

[0048] S101 Analyze the metadata information of the target data and establish a blood relationship model between the data. The blood relationship model includes the complete link of the data from generation to transmission and consumption;

[0049] S102 Based on the blood relationship model, continuously store and dynamically update the blood relationship between the data based on the graph database, and reflect the changing state of data flow in real time;

[0050] S103 Bind the preset warning rules to the target data, and automatically trigger a warning when the target data is abnormal;

[0051] S104 Detect the data quality, blood relationship link and traffic trend based on data blood relationship analysis, statistical methods and machine learning, and obtain abnormal data;

[0052] S105 Based on the obtained abnormal data, conduct traceability of the abnormal data, obtain the relationship graph of the corresponding unprocessed nodes from the graph database, and perform recursive search to obtain the traceability result;

[0053] S106 Based on the traceability result, execute the warning task disposal, perform data repair and processing, submit the processing result, and automatically execute the data quality backcheck.

[0054] Specifically, for the data governance warning processing method provided by the embodiments of the present application, first, by analyzing the metadata information of the data, the system analyzes the metadata information of the target data, including table structure, field definition, etc., through metadata parsing technology, and establishes a blood relationship model between the data. The blood relationship describes the complete link of the data from generation to transmission and consumption, covering information such as data flow and data derivation. By combining Antlr4 to complete the blood relationship analysis of data tables and data fields, and used to construct a Data Lineage Graph, the system can intuitively display the path of data flow, helping to trace and analyze the source and destination of the data. For example, in an e-commerce platform, the whole process of order data from generation to final storage, including various links such as order creation, payment, and delivery, will be detailedly recorded in the blood relationship model.

[0055] The system continuously stores and dynamically updates the blood relationship, and reflects the latest change status of the data flow in real time. It can be realized by using a big data processing framework and a graph database (Neo4j). Using a graph database (such as Neo4j) to store the blood relationship model, the system can continuously store and dynamically update the blood relationship between the data. Whenever a data processing task (such as an ETL job, data synchronization, etc.) is completed, the platform will automatically update the blood relationship graph according to the execution result of the task. For example, if order table A generates order table B after a series of conversions, the system will update the blood relationship graph, add the dependency relationship from table A to table B, and reflect this change in real time.

[0056] Based on the data characteristics and abnormal scenarios, the system binds the preset warning rules to the target data, similar to imposing controls on the data, ensuring that when the target data is abnormal, the system can automatically trigger a warning, thereby enabling timely response and effective execution of data governance tasks. These rules can be static, such as triggering a warning when the null value rate exceeds 5%; they can also be dynamic, such as using time series analysis to detect abnormal fluctuations; or they can be hybrid rules that combine algorithm calculation results with business logic. For example, if the order amount suddenly increases and the payment process is interrupted, the system will trigger a high-priority alarm.

[0057] By continuously monitoring the data flow in real time and combining lineage analysis, statistical methods, and machine learning techniques, the system comprehensively detects data quality, lineage links, and traffic trends. For example, the system may detect that the null value rate of a certain order table suddenly increases, or the traffic of the order payment process abnormally decreases. These abnormal data will be marked by the system and further analyzed.

[0058] When the system discovers abnormal data, it retrieves the lineage link graph related to the current abnormal task or attribute from the graph database and conducts recursive searches until the source of the data quality problem is found. For example, if a large number of null values appear in the payment status field of a certain order table, the system will trace back to the source of these orders to determine which link caused the data quality problem.

[0059] After locating the root cause of the anomaly, the system generates corresponding warning tasks and notifies the data governance responsible person to repair or further process the abnormal data. After the task is disposed of, a detailed disposal record needs to be filled out and the processing result submitted. After the task is submitted, the system will automatically perform a data quality backcheck to ensure that the traceability processing of the abnormal data has been fully followed up. For example, if it is found that the data transmission of a certain payment service provider is interrupted, resulting in the loss of order payment status, the system will prompt the responsible person to repair the data transmission link and recheck the data quality after the repair to ensure that all affected orders have been correctly updated.

[0060] The data governance warning processing method based on lineage analysis provided by the embodiments of this application significantly improves the accuracy and efficiency of data governance by establishing a data lineage relationship model and continuously monitoring data flow. It can quickly locate the source of data problems, reduce the repair time, lower risks, and the automated warning and processing mechanism reduces manual intervention, lowering the risks of errors and missed reports. At the same time, it enhances the security and credibility of the data, providing an efficient and accurate solution for data governance in the big data environment.

[0061] Optionally, the data governance warning processing method provided by the embodiments of the present application analyzes the metadata information of the target data to establish a blood relationship model between the data. The blood relationship model includes the complete link of data from generation to transmission and consumption, specifically including:

[0062] Parse the metadata of the data based on Antlr4;

[0063] Based on data mapping, obtain the complete link of data from generation, transmission to consumption, including data flow direction and data derivation information;

[0064] Based on event-driven, obtain incremental updates of the blood relationship model.

[0065] Specifically, parse the metadata of the data (such as table structure, field definition) based on Antlr4. Describe the complete link of data from generation, transmission to consumption, including information such as data flow direction and data derivation. Trigger incremental updates of blood relationship modeling through event-driven (such as data processing task changes). Among them, Antlr4 is a grammar parser generator that can identify and extract metadata such as table structure and field definition in SQL statements, providing basic data for establishing a blood relationship model.

[0066] The task management module and data processing task module of the data governance platform trigger incremental updates of the blood relationship diagram through the event-driven mechanism. Whenever a new data processing task (such as an ETL job, data synchronization, etc.) is completed, the platform will automatically update the blood relationship diagram according to the execution result of the task (such as a newly added data table or field change). This supports not only full updates but also incremental updates.

[0067] In one embodiment, assume that in a certain ETL task, table A generates table B after a certain transformation. The system will update the blood relationship diagram according to the execution of this task, adding a dependency relationship from table A to table B.

[0068] Based on the task configuration of the data governance platform and using Antlr4 to parse table operations (such as creating tables, inserting data, updating data, etc.) in SQL scripts or ETL tasks, extract the relationships between tables.

[0069] Use Antlr4 to parse operations on fields in SQL (such as field mapping, calculation formula, field transformation, etc.); identify the dependency relationships between fields. The value of a certain field may be obtained through the calculation results of other fields (such as addition, multiplication, etc.), or merged with fields in other tables; analyze complex SQL operations, such as field references in JOIN, GROUP BY, WHERE clauses, to understand the source, transformation process, and final destination of the fields.

[0070] The data governance early warning processing method provided by the embodiments of this application can accurately trace the data flow path, update the blood relationship in real time, and quickly respond to data changes, thereby providing strong support for data governance.

[0071] Optionally, for the data governance early warning processing method provided by the embodiments of this application, the blood relationship between the data is continuously stored and dynamically updated based on a graph database, including constructing the blood relationship model into a graph data structure and recording the change history of the blood relationship.

[0072] Specifically, a graph database (such as Neo4j) is used to store the blood relationship as a graph data structure. Through the relationship between nodes and edges, the graph database can intuitively and efficiently store and query data blood information. Among them, nodes represent data entities (such as data tables, fields, tasks, databases, etc.). Each node has its own attributes to describe the relevant information of the data entity (such as table name, field name, data type, etc.); edges represent the dependency relationship between data.

[0073] The system captures events of data asset changes and triggers real-time updates of the blood relationship map through an event-driven mechanism, such as the integration of Kafka and the graph database. Each update of the blood relationship creates a new version and records the differences before and after the update to achieve the management of historical versions. The query language of the graph database (such as Cypher for Neo4j) is used for the retrieval and graph calculation of the blood relationship, supporting in-depth analysis and impact assessment of data blood. When the system is upgraded and transformed, through the impact analysis of dependent data, the affected range downstream can be quickly located, reducing the risks brought by system upgrade and transformation. At the same time, by analyzing the upstream and downstream flows, a complete data flow chain can be obtained to assist in the addition of subsequent system functions.

[0074] To trace the change history of the blood relationship and provide a basis for subsequent traceability and problem troubleshooting, the system needs to manage the version of the blood relationship map. Each time the blood relationship map is updated, a new version is created and the differences before and after the update are recorded. The graph database supports a version control mechanism, allowing queries of historical versions (the system can query the blood relationship of a certain table in the past period of time, or trace back the source and changes of a certain field in historical versions); each update of the blood relationship generates a change log, recording the addition, deletion, or modification of nodes and edges in the blood relationship map. These change logs provide a transparent audit trail for problem traceability and anomaly detection.

[0075] In one embodiment, assume that in a data platform, an ETL task is executed, and data table A is transformed through a series of processes to generate data table B. Through the blood relationship storage module:

[0076] The relationship between Table A and Table B will be connected by edges in the graph database to form a data lineage link. If a certain field in Table A changes (such as adding a new field or modifying the calculation rule of the field), this change will trigger an update of the lineage graph, adding or modifying relevant nodes and edges. The system will create a version for each update of the lineage graph, recording the lineage relationship between Table A and Table B and the update information of relevant fields. Users can query historical versions to view the status of the lineage relationship between Table A and Table B at a certain point in the past, ensuring the integrity of data traceability.

[0077] The data governance warning processing method based on lineage analysis provided by the embodiments of this application realizes real-time update, storage, and query optimization of data lineage relationships, improves the accuracy and efficiency of data governance, reduces the cost of solving data problems, and enhances the data traceability and security.

[0078] Optionally, for the data governance warning processing method based on lineage analysis provided by the embodiments of this application, binding a preset warning rule to the target data to automatically trigger a warning when the target data is abnormal specifically includes:

[0079] The preset warning rules include static rules, dynamic rules, and hybrid rules;

[0080] The static rule defines the warning standard based on traditional thresholds or logical conditions; the dynamic rule identifies abnormal data by using statistical analysis, machine learning, and graph algorithm techniques; the hybrid rule combines the static rule and the dynamic rule through composite conditions.

[0081] Specifically, in this step, the system binds the preset warning rule to the target data according to the data characteristics and abnormal scenarios, similar to conducting surveillance on the data. This process ensures that when the target data is abnormal, the system can automatically trigger a warning, thereby achieving timely response and effective execution of data governance tasks. The warning rules can be divided into static rules, dynamic rules, and hybrid rules according to the complexity and nature of the problem.

[0082] The static rule uses traditional thresholds or logical conditions to define the warning standard, and is applicable to scenarios such as data quality detection and simple business rule detection. The static rule is usually a fixed condition defined based on business experience or historical data. For example, for null value rate detection, when the null value rate of a certain field exceeds a predetermined threshold (such as 5%); or for duplicate data detection, if the number of records in a certain table exceeds the specified threshold (such as the duplication rate is greater than 2%); or for data range detection, when the value of a certain field exceeds a reasonable range (such as negative numbers, overly large values, etc.).

[0083] Dynamic rules are applicable to complex data anomaly detection scenarios, especially when the data change trend is uncertain or the data flow has certain regularity. Dynamic rules identify abnormal patterns, trends, and structures by introducing technologies such as statistical analysis, machine learning, and graph algorithms. Use time series analysis (such as ARIMA, LSTM) to detect abnormal fluctuations; or use clustering algorithms (such as DBSCAN) to discover abnormal clustering points; or identify structural anomalies in the lineage based on graph algorithms (such as loop detection). For example, use graph algorithms to detect whether there is a data loop in the data transfer process. For example, an ETL task may form a closed loop due to the dependency relationship between data tables, resulting in a data processing dead loop.

[0084] Hybrid rules combine static rules and dynamic rules, and improve the accuracy and real-time performance of early warnings through composite condition design. For example, a sudden increase in data traffic + interruption of a specific business process triggers a high-priority alarm; or if the distribution pattern of data quality anomaly points is consistent with the characteristics of business peak hours, the early warning sensitivity is reduced.

[0085] Optionally, the data governance early warning processing method provided by the embodiments of the present application detects data quality, lineage links, and traffic trends based on data lineage analysis, statistical methods, and machine learning to obtain abnormal data, specifically including:

[0086] The detection of data quality detects the key quality indicators of the data; the lineage link detection is to detect abnormal data in the lineage link by constructing a data lineage graph and combining graph algorithms; the traffic trend detection is to perform anomaly analysis on the dynamic changes of data traffic to identify abnormal traffic points or trends.

[0087] Specifically, data quality anomaly detection aims to ensure that key quality indicators such as the accuracy, integrity, consistency, and uniqueness of the data are kept within a controllable range. First, configure the acquisition logic, calculate indicators such as the null value rate, duplication rate, and primary key conflict rate, and perform deviation detection on the data through the Z-score method based on the Gaussian distribution. For each data field, calculate its historical mean and standard deviation, and use the Z-score formula to determine whether the current data point exceeds the normal range (for example, when the Z-score exceeds 3, the data point is considered abnormal); use DBSCAN, K-Means clustering algorithms to analyze the data, and divide the data into normal clusters and abnormal clusters. If a data point belongs to an abnormal cluster, it means that the data point has a significant difference from other data in the feature space and needs further verification.

[0088] For example, in an e-commerce platform, the order amount should be roughly distributed within a certain range. If a certain order amount is extremely high or low, it may be incorrect data. Use the DBSCAN clustering algorithm to perform clustering analysis on the order amount to identify isolated points; if DBSCAN marks some order amounts as abnormal points (i.e., these data cannot be assigned to any normal cluster), the system will mark these data as abnormal.

[0089] Abnormal analysis of data lineage links can detect possible structural abnormalities during the data flow process, such as link interruptions, circular dependencies, isolated nodes, etc. By constructing a data lineage graph and combining graph algorithms, problems in the lineage links can be effectively identified.

[0090] Abnormalities in the lineage links include: broken link abnormalities, detecting upstream and downstream link interruptions that result in data loss; loop abnormalities, identifying invalid circular dependencies that may cause deadlocks or performance issues; isolated node abnormalities, finding nodes with no incoming or outgoing edges, indicating that the data is not being used or the source is lost. By loop detection, closed loops in the data lineage graph are found to identify invalid circular dependencies and prevent problems such as deadlocks. By shortest path analysis, the data transmission path is evaluated to detect upstream and downstream link interruptions and analyze the reasons for link breaks or path abnormalities. By node centrality, by calculating the degree of nodes, it is determined which data nodes are key nodes to help discover potential abnormalities.

[0091] For example, if a circular dependency is detected in a data processing task (e.g., data A depends on data B, and data B depends on data A), the system will mark this loop as abnormal; identify nodes in the lineage graph with no incoming or outgoing edges, usually these nodes represent isolated nodes with no data flowing in or out.

[0092] Data traffic anomaly detection aims to monitor the dynamic changes in data traffic in real time and identify potential abnormal traffic points or trends. By using time series models and clustering algorithms, sudden fluctuations or unexpected behavior patterns in data traffic can be effectively discovered. Methods for identifying abnormal traffic points or trends can include: time series analysis, using models such as ARIMA and LSTM to perform time series modeling based on historical traffic, predicting future traffic and identifying deviations; unsupervised clustering, using K-Means and DBSCAN to classify traffic into normal and abnormal categories to discover sudden traffic changes; trend analysis, detecting abnormal patterns such as traffic peaks, sudden drops, or long-term stability at low values.

[0093] Use ARIMA (Autoregressive Integrated Moving Average Model) to establish a prediction model based on the time series of historical data traffic, predict future traffic, compare the actual traffic with the predicted traffic, and if a significant deviation occurs, it is judged as abnormal.

[0094] For example, page access fluctuates at different times, but generally follows a certain trend. A sudden surge or decrease in the number of accesses may be a signal of system failure or malicious attack; use the ARIMA model to model historical access data and predict future access volume; compare the difference between the actual access volume and the predicted value, and calculate the abnormal deviation; if the actual access volume exceeds the expected range, trigger an early warning

[0095] Optionally, for the data governance early warning processing method provided by the embodiments of the present application, based on the obtained abnormal data, trace the source of the abnormal data, obtain the relationship graph of the corresponding unprocessed nodes from the graph database, perform recursive search, and obtain the traceability result, specifically including:

[0096] Obtain the blood relationship link graph related to the current abnormal task or attribute from the graph database;

[0097] Starting from the current abnormal data node, recursively search for upstream and downstream nodes layer by layer to confirm whether there are data processing problems in the upstream data and downstream data in the data stream;

[0098] Generate an early warning task based on the obtained abnormal source, and give a repair suggestion based on the abnormal type.

[0099] Specifically, when the system discovers an abnormal task or an abnormal attribute node, it will then obtain the relationship graph of the corresponding unprocessed nodes from the graph database and perform a recursive search until the source of the data quality problem is found. First, obtain the abnormal node relationship graph and the blood relationship link graph related to the current abnormal task or attribute from the graph database; then recursively search for the abnormal source node and quickly locate the problem node or link using blood relationship analysis; finally, generate an early warning task, generate an early warning task based on the abnormal source, and give a repair suggestion based on the abnormal type.

[0100] Extract the blood relationship link graph related to the current abnormal task or abnormal attribute node from the graph database (such as Neo4j). Find the currently marked abnormal nodes or tasks in the system. Usually, these nodes represent abnormal data, abnormal calculation results, or intermediate points of data transmission interruption. Based on the current abnormal node, query the directly or indirectly related blood relationship link from the graph database. Each node in the blood relationship link represents a data entity (such as a data table, a data field, a data processing process, etc.), and each edge represents the flow or dependency relationship of the data; starting from the current abnormal node, recursively search for upstream or downstream nodes layer by layer to confirm whether there are upstream data problems or downstream data processing problems in the data stream.

[0101] When the root cause of the abnormality is located, the system will generate a corresponding early warning task to notify the data governance responsible person to repair or further process the abnormal data. The task will include the following information:

[0102] Exception type: Indicate data missing, data inconsistency, data type error, data dependency interruption, etc.

[0103] Exception source: Clearly point out the root node to help the responsible person quickly locate the problem.

[0104] Scope of influence: Describe other data nodes or processes affected by the exception source.

[0105] Repair suggestions: Based on the exception type, provide repair solutions, such as resetting data, modifying calculation formulas, repairing data transmission links, etc.

[0106] Optionally, for the data governance warning processing method provided by the embodiments of the present application, based on the tracing result, execute warning task handling, perform data repair and processing, submit the processing result, and automatically execute data quality backcheck, specifically including:

[0107] Perform data repair and processing based on the analysis result, repair suggestions, and blood relationship of the exception data;

[0108] After the task handling is completed, submit the processing result and automatically execute data quality backcheck;

[0109] Based on the blood relationship chain of the graph database, determine whether all related data nodes of the exception data have been repaired;

[0110] If it is found that one or more exception data have not been repaired, the system will mark them as pending and reassign tasks.

[0111] Specifically, after the relevant data governance responsible person receives the warning task, they can perform data repair and processing based on the analysis result, repair suggestions, and blood relationship chain of the exception data. After the task handling is completed, a detailed handling record needs to be filled in and the processing result needs to be submitted. Traverse all the processed warning tasks in the warning data list. Each task contains the exception data and its related warning information, including exception type, data source, data quality rules, and trigger thresholds, etc. These information provide a basis for subsequent tracing and processing.

[0112] After the task is submitted, the system will automatically execute data quality backcheck to ensure that the tracing and processing of the exception data have been fully followed up. The exception data is backchecked through data quality rules, and at the same time, the upstream and downstream data flow relationships of the exception data will be found to check whether the data has been processed. If the exception data has been repaired, continue to check all its associated data (including directly and indirectly dependent data nodes).

[0113] Through the blood relationship chain of the graph database, the system will determine whether all related data nodes of the abnormal data have been repaired. If it is confirmed that all related data nodes have been processed, a complete data quality warning traceability processing result is generated and recorded in the warning report. This report not only includes the processing of the abnormal data itself, but also lists in detail the processing progress of the entire data bloodline chain.

[0114] If it is found that some data has not been repaired, the system will mark it as pending and reassign the task to ensure the integrity of the entire data chain. For those abnormal tasks with unprocessed associated data, the system will reject the task. At this time, the person responsible for the relevant data needs to further investigate and repair the unprocessed abnormal data.

[0115] The data governance early warning processing method based on blood relationship analysis provided in the embodiment of the present application can achieve the following effects:

[0116] Improve the accuracy of data governance: This method uses lineage analysis technology to fully understand the data generation, flow and consumption process, which can more accurately monitor and warn of data quality issues and avoid the negative impact of data anomalies on business and decision-making.

[0117] Real-time monitoring of data flow: Real-time monitoring and analysis can continuously monitor data flow, detect abnormal situations in time and issue early warnings, which helps to quickly respond to and handle data problems and reduce potential risks and losses.

[0118] Improve data tracing and source tracking: Based on the blood relationship model, the flow and changes of data can be traced and traced, helping to quickly locate problems and reduce the time and energy costs required to repair and restore data.

[0119] Improve data governance efficiency: Automated data warning and processing mechanisms can reduce the need for manual intervention, improve the efficiency and automation of data governance, and reduce the risk of human errors and underreporting.

[0120] Strengthen data security: Through the early warning processing mechanism, data quality problems can be discovered and resolved in a timely manner, data leakage, tampering and abuse can be prevented, and the security and credibility of data can be enhanced.

[0121] This application also provides a data governance early warning processing device based on blood relationship analysis. Figure 2 A schematic diagram of the structure of a data governance early warning processing device based on blood relationship analysis provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the device comprises:

[0122] A modeling module is used to analyze the metadata information of the target data and establish a blood relationship model between the data. The blood relationship model includes the complete link from data generation to transmission and consumption;

[0123] A storage module, configured to continuously store and dynamically update the lineage relationship between the data based on the lineage relationship model in a graph database, and to reflect the changing state of data flow in real time;

[0124] A rule module, configured to bind preset warning rules to target data and automatically trigger a warning when the target data is abnormal;

[0125] An auditing module, configured to detect data quality, lineage links, and traffic trends based on data lineage analysis, statistical methods, and machine learning to obtain abnormal data;

[0126] An analysis module, configured to trace the source of the abnormal data based on the obtained abnormal data, obtain a relationship graph of corresponding unprocessed nodes from the graph database, perform recursive search, and obtain a tracing result;

[0127] A processing module, configured to execute warning task handling based on the tracing result, perform data repair and processing, submit a processing result, and automatically perform data quality backcheck.

[0128] The present application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented. Among them, the computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, micro drives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nano-systems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0129] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0130] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0131] In several embodiments provided in this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0132] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0133] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0134] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned memory includes: USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical disks, etc., which can store program codes.

[0135] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory. The memory can include: flash drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc.

[0136] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

[0137] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0138] It is easy for those skilled in the art to understand that the above are only preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A data governance early warning processing method based on blood relationship analysis, characterized in that: The following steps are involved: Analyze the metadata information of the target data and establish a blood relationship model between the data, wherein the blood relationship model includes a complete link from data generation to transmission and consumption; Based on the blood relationship model, the blood relationship between the data is continuously stored and dynamically updated based on the graph database, reflecting the changing state of data flow in real time; Bind the preset warning rules to the target data to automatically trigger the warning when the target data is abnormal; Based on data lineage analysis, statistical methods and machine learning, data quality, lineage links and traffic trends are tested to obtain abnormal data; Based on the acquired abnormal data, tracing the abnormal data source, acquiring a relationship graph of corresponding unprocessed nodes from the graph database, performing a recursive search, and acquiring a tracing result; Based on the traceability results, the early warning task disposal is executed, data repair and processing are performed, the processing results are submitted, and data quality reverse checking is automatically performed.

2. The data governance early warning processing method based on blood relationship analysis as claimed in claim 1, characterized in that: The metadata information of the target data is analyzed to establish a blood relationship model between the data. The blood relationship model includes a complete link from data generation to transmission and consumption, specifically including: Meta information of parsed data based on Antlr4; Based on data mapping, the complete chain of data generation, transmission and consumption is obtained, including data flow and data-derived information; Get incremental updates of the blood relationship model based on event-driven.

3. The data governance early warning processing method based on blood relationship analysis as claimed in claim 2 is characterized in that: The continuous storage and dynamic updating of the blood relationship between the data based on the graph database includes constructing the blood relationship model into a graph data structure to record the change history of the blood relationship.

4. The data governance early warning processing method based on blood relationship analysis as claimed in claim 3 is characterized in that: The preset warning rules are bound to the target data to automatically trigger the warning when the target data is abnormal, specifically including: The preset warning rules include static rules, dynamic rules and mixed rules; The static rules define warning standards based on traditional prepayments or logical conditions; the dynamic rules identify abnormal data by using statistical analysis, machine learning and graph algorithm technology; the hybrid rules combine static rules and dynamic rules through complex conditions.

5. The data governance early warning processing method based on blood relationship analysis as claimed in claim 4 is characterized in that: The data quality, lineage links and traffic trends are detected based on data lineage analysis, statistical methods and machine learning to obtain abnormal data, specifically including: The data quality detection is to detect the key quality indicators of the data; the lineage link detection is to construct a data lineage graph and combine it with a graph algorithm to detect abnormal data in the lineage link; the traffic trend detection is to perform an abnormal analysis on the dynamic changes of data traffic and identify abnormal traffic points or trends.

6. The data governance early warning processing method based on blood relationship analysis as claimed in claim 5, characterized in that: The abnormal data is traced based on the acquired abnormal data, a relationship graph of corresponding unprocessed nodes is obtained from the graph database, and a recursive search is performed to obtain a traceability result, which specifically includes: Acquire a bloodline link graph related to the current abnormal task or attribute from the graph database; Starting from the current abnormal data node, recursively search the upstream and downstream nodes layer by layer to confirm whether there are data processing problems in the upstream and downstream data in the data flow; Generate early warning tasks based on the acquired anomaly sources and provide repair suggestions based on the anomaly types.

7. The data governance early warning processing method based on blood relationship analysis as claimed in claim 6 is characterized in that: Based on the traceability results, the early warning task processing is executed, data repair and processing are performed, processing results are submitted, and data quality reverse checking is automatically performed, specifically including: Perform data repair and processing based on the analysis results, repair suggestions and blood relationships of abnormal data; After the task is processed, the processing results are submitted and data quality is automatically checked; Based on the blood relationship chain of the graph database, determine whether all related data nodes of the abnormal data have been repaired; If one or more abnormal data are found to have not been repaired, the system will mark them as pending and reassign the task.

8. A data governance early warning processing device based on blood relationship analysis, characterized in that: include: A modeling module is used to analyze the metadata information of the target data and establish a blood relationship model between the data. The blood relationship model includes the complete link from data generation to transmission and consumption; A storage module, used for continuously storing and dynamically updating the blood relationship between the data based on the blood relationship model and the graph database, so as to reflect the changing state of data flow in real time; The rule module is used to bind the preset warning rules to the target data, so as to automatically trigger the warning when the target data is abnormal; The audit module is used to detect data quality, lineage links, and traffic trends based on data lineage analysis, statistical methods, and machine learning to obtain abnormal data; An analysis module is used to trace the abnormal data based on the acquired abnormal data, obtain the relationship graph of the corresponding unprocessed nodes from the graph database, perform recursive search, and obtain the tracing result; The processing module is used to execute early warning task disposal, perform data repair and processing, submit processing results, and automatically perform data quality reverse check based on the tracing results.

9. Data governance and early warning processing equipment based on blood relationship analysis, characterized in that: The method comprises at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the processing unit executes the steps of the method according to any one of claims 1 to 7.

10. A storage medium, characterized in that It stores a computer program that can be executed by a data governance and early warning processing device based on bloodline analysis. When the computer program runs on the data governance and early warning processing device based on bloodline analysis, the data governance and early warning processing device based on bloodline analysis executes the steps of any one of the methods described in claims 1 to 7.

Citation Information

Cited By

  • Data management method and system for digital government affairs

    CN120508554A

  • Intelligent early warning method and system based on multiple Agents

    CN121077946A

  • A multi-agent based intelligent early warning method, system, device and storage medium

    CN121077946B

  • Service flow auditing method, system and device and storage medium

    CN121561691A