Iterative updating method for data consanguinity map
By capturing incremental data, building directed acyclic graphs, hierarchical modeling, distributed computing and stream processing technologies, as well as machine learning and automated repair mechanisms in the data blood relationship map iterative update method, the problems of insufficient real-time, lack of intelligence and poor resource optimization in the existing technology are solved, and efficient, intelligent and resource optimization data blood relationship map updates are achieved.
Patent Information
- Application Number
- CN202510398258.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing data blood relationship map iterative update methods are insufficient in real-time, intelligence and resource optimization, and it is difficult to meet the real-time update requirements of high-frequency data changes. It lacks the ability to predict data-dependent path changes, and the abnormal node detection and repair efficiency is low.
By capturing data increments and metadata changes, a directed acyclic graph of data blood relationship is constructed, and the graph is hierarchical modeled, and iterative updates are performed based on the hierarchical modeling results. Use distributed computing and stream processing technology to optimize incremental data processing and storage, use machine learning models to predict blood relationship path changes, and detect and repair abnormal nodes through automated mechanisms.
It realizes efficient iterative update of the data blood relationship map, has the ability to capture, quickly process and accurately update in real time, significantly improves the update efficiency and the independent optimization capabilities of the system, and ensures the logical integrity and stability of the data blood relationship.
Smart Images

Figure CN119917515A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to an iterative updating method for data lineage graphs. Background Art
[0002] Data lineage graph is an important tool for modern data management. It provides support for scenarios such as data governance, data traceability, and compliance analysis by describing the flow and dependencies of data in the system. The iterative update method for data lineage graph refers to the continuous maintenance of the integrity and real-time nature of the data lineage graph through incremental capture and dynamic update technology when the data environment changes.
[0003] At present, the existing data lineage graph iterative update methods mostly adopt a batch-based approach, capturing incremental data through data snapshot comparison, and using the dynamic update technology of the graph structure to complete the graph maintenance. Some technologies also combine topological sorting to achieve hierarchical updates of dependencies, thereby improving the efficiency of updates. At the same time, existing technologies support the processing and storage of large-scale data in a distributed environment, have good scalability, and can meet the needs of iterative updates of data of a certain scale.
[0004] However, existing technologies still have shortcomings in real-time, intelligence and resource optimization. For example, existing technologies do not respond well to high-frequency data changes and are unable to meet real-time update requirements. They lack the ability to predict changes in data dependency paths, and the degree of optimization of update operations is limited. In addition, the detection and repair of abnormal nodes mostly rely on manual operations, which reduces the efficiency of the system's autonomous maintenance. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides an iterative update method for data lineage maps, which solves the problems of insufficient real-time performance in iterative update of data lineage maps, lack of prediction ability for data change paths, and low efficiency in abnormal node detection and repair.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an iterative update method for a data kinship map, comprising the following steps: Capture data increments and metadata changes; Construct a directed acyclic graph of data lineage relationships; Perform hierarchical modeling on the data lineage graph; Iteratively update the blood relationship map based on the hierarchical modeling results; Optimize the processing and storage of incremental data through distributed computing; Use stream processing technology to achieve real-time updates of bloodline graphs; Use machine learning models to predict changes in bloodline paths and detect and repair abnormal nodes.
[0007] Preferably, capturing data increments and metadata changes includes: Get data snapshots, including new and old snapshots of data; Identify incremental portions of data by comparing snapshots; Perform timestamp filtering, hash value verification, and version comparison on incremental data to extract new, changed, and deleted data; Capture metadata changes, including additions, modifications, and deletions at the table and field levels; Standardize the recognition results into incremental data event format for subsequent processing.
[0008] Preferably, the construction of the directed acyclic graph of data kinship includes: Define data nodes, including table-level and field-level data entities; Define data flow relationships, including input relationships and output relationships; Establish the node set and edge set of the graph based on data dependency; Ensure that the constructed graph structure satisfies the acyclic property.
[0009] Preferably, the hierarchical modeling of the data blood relationship graph includes: Divide the nodes in the graph into levels through topological sorting; Calculate node priority based on node access frequency and change frequency; Nodes with higher priority are classified as core layers, and nodes with lower priority are classified as edge layers; Mark the dependencies of nodes at each layer to ensure that they are executed first by layer during iterative updates.
[0010] Preferably, the hierarchical modeling results perform iterative updates on the blood relationship map, including: For core layer nodes, incremental update operations are performed in real time; For edge layer nodes, incremental update operations are performed in batches according to time windows; Detect data dependency conflicts during incremental updates and mark conflicting nodes and paths; Call the rule engine to automatically repair conflicting nodes and paths.
[0011] Preferably, the optimizing the processing and storage of incremental data by distributed computing includes: Use a distributed graph database to store bloodline graphs; The data lineage graph is partitioned and stored according to the access frequency and priority of the nodes; Use a distributed computing framework to perform parallel processing on incremental data; The processing results are synchronized to the graph database through the batch write interface.
[0012] Preferably, the use of stream processing technology to achieve real-time bloodline map update includes: Monitor data change events through the stream processing engine; Parse the captured event stream into incremental update requests; Parse data dependencies in real-time event streams and generate incremental relationships; The blood relationship map in the graph database is updated in real time based on the analysis results.
[0013] Preferably, the method of using a machine learning model to predict changes in bloodline pathways includes: Extract the characteristics of historical lineage change data, including access frequency, dependency, and change frequency; Use time series models to predict the path of data lineage changes; Generates priority update incremental data requests based on prediction results.
[0014] Preferably, the detecting and repairing of abnormal nodes includes: Use anomaly detection algorithms to identify abnormal nodes in the bloodline graph; Generate repair strategies for detected abnormal nodes; Execute the repair strategy and rebuild the dependency relationship of abnormal nodes; Record the repair operation log for subsequent audit.
[0015] Preferably, the incremental data capturing method includes: For large-scale distributed data sources, incremental data is captured in parallel by data sharding; Check the consistency of data sources across systems during the capture process and identify possible synchronization errors; Perform fault tolerance processing for identified erroneous data, including retry mechanisms or data compensation strategies; The captured incremental data is marked as traceable batch tasks for subsequent iterative updates.
[0016] The present invention provides an iterative update method for a data lineage map, which has the following beneficial effects: 1. The present invention realizes efficient iterative update of data lineage graph through data sharding, distributed computing, stream processing engine and other technologies. It has the ability of real-time capture, rapid processing and accurate update in dynamic data environment, and can adapt to the multi-level requirements of large-scale distributed data sources and complex data dependencies.
[0017] 2. The present invention introduces a machine learning model to predict changes in lineage paths, combines anomaly detection and automatic repair mechanisms, and constructs an intelligent graph update framework, which significantly improves the update efficiency and the system's autonomous optimization capabilities, while ensuring the logical integrity and stability of data lineage.
[0018] 3. The present invention optimizes the allocation of update resources through node hierarchical modeling and priority division, and achieves data consistency and resource utilization in high-concurrency scenarios through strategies such as batch management and transaction control. It has good scalability and fault tolerance, and is adaptable to multi-scenario business needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] Please see attached Figure 1 The embodiment of the present invention provides an iterative update method for a data lineage map, comprising the following steps: S1. Capturing data increments and metadata changes. By capturing data increments and metadata changes, efficient and accurate data increment and metadata change capture is achieved, providing reliable input for subsequent iterative updates of the bloodline map, while reducing the processing cost of the full amount of data; S2. Construct a directed acyclic graph of data lineage relationships. The constructed data lineage relationship graph clearly describes the data flow and dependency relationships, provides an accurate basic data structure for hierarchical modeling and iterative updates, and improves the accuracy and stability of graph updates. S3. Hierarchical modeling of the data lineage relationship graph significantly improves the efficiency of updating the lineage graph. The core layer is updated first to ensure the timeliness of key nodes, and the edge layer batch update reduces the system load; S4. Perform iterative updates on the bloodline map based on the hierarchical modeling results. The iterative update method optimizes the update efficiency through the hierarchical mechanism while ensuring the consistency of the data bloodline map, thereby reducing the computational overhead of full reconstruction. S5. Optimize the processing and storage of incremental data through distributed computing. The distributed computing framework effectively improves the efficiency of incremental data processing. Combined with distributed storage, it optimizes the access performance and scalability of the graph, ensuring the stability of the system in large-scale data scenarios. S6. Use stream processing technology to achieve real-time update of the bloodline map. Stream processing technology enables real-time update of the bloodline map, ensuring that data changes can be reflected in the map in a timely manner to meet the needs of high real-time scenarios; S7. Use machine learning models to predict changes in bloodline paths and detect and repair abnormal nodes. The machine learning model improves the accuracy of bloodline map change predictions and enables automatic detection and repair of abnormal nodes, thereby improving the intelligence level and maintenance efficiency of the system.
[0022] Please see attached Figure 1 In a preferred embodiment of the present invention, S1, said capturing data increments and metadata changes comprises: Get data snapshots, including new and old snapshots of data. The essence of a data snapshot is a complete record of the data status. By comparing the new and old snapshots, data changes can be identified, thereby avoiding the computational cost of rebuilding the full data. Data snapshots are uniformly identified using a version management system. For example, each snapshot is assigned a unique version number and timestamp for subsequent incremental calculations. Versioning management of snapshots improves data traceability and provides efficient input for iterative updates. By comparing snapshots to identify incremental parts of the data, the differences between snapshots are calculated using set operations:
[0023] in, is the new snapshot dataset, For the old snapshot dataset, is an incremental data set; Field-level comparison: If the data values in two snapshots are consistent but the field attributes (such as type) have changed, the field change event is recorded separately; Result marking: Mark the identified incremental data as new, changed or deleted events for subsequent hierarchical modeling and iterative updates; Difference calculation compares the content and structure of new and old data sets to quickly identify the changed data parts and reduce the resource overhead of processing the full amount of data.
[0024] Difference records are stored in a key-value structure to ensure that data changes can be efficiently mapped to specific graph nodes; Timestamp filtering, hash value verification and version comparison are performed on incremental data to extract new data, changed data and deleted data. The combination of timestamp filtering and hash verification ensures accurate capture of data changes. The version comparison mechanism further classifies incremental data types and provides structured input data for hierarchical modeling and iterative updates. Through the joint use of timestamps, hash values and version comparisons, new, changed and deleted data can be accurately extracted, providing an efficient input mechanism for subsequent processing of incremental data and reducing redundant calculations. Capture metadata changes, including addition, modification, and deletion of table-level and field-level information. Table-level and field-level metadata capture is based on structured comparison of snapshots, which can accurately locate metadata changes and ensure that metadata changes are synchronized with the update of the lineage map. Capturing metadata changes provides complete contextual information, laying a structured foundation for the construction and update of the lineage map, while supporting complex metadata management in a dynamic data environment. Standardize the recognition results into incremental data event format for subsequent processing. Standardized event format design: Encapsulate all recognized incremental data and metadata changes into a unified event format: Event type: add, change, delete; Data types: table-level data, field-level data; Event content: changed data value, field attribute or table attribute; Event time: the timestamp of the time when the event occurred; Event storage and transmission: Standardized events are stored in queues or databases and transmitted to subsequent hierarchical modeling and iterative update modules through stream processing engines (such as Kafka); The design of the standardized event format ensures that different types of incremental data can be processed uniformly, reducing the complexity caused by heterogeneous data structures and improving the scalability of subsequent modules; Standardized incremental events provide a unified input interface for subsequent iterative updates, ensuring the simplicity, accuracy, and scalability of the incremental data processing process.
[0025] Please see attached Figure 1 In a preferred embodiment of the present invention, S2, the directed acyclic graph of constructing data kinship relationships includes: Define data nodes, including table-level and field-level data entities. The establishment of node hierarchical relationships enables the graph to display high-level table relationships and describe in detail the specific dependencies between fields, ensuring the granularity and hierarchy of lineage information. Defining table-level and field-level nodes achieves multi-level coverage of kinship descriptions, provides a comprehensive view of fields, and lays the foundation for accurate modeling and visualization of data flows; Define the data flow relationship, including input relationship and output relationship. The input relationship and output relationship together constitute a complete description of the data flow. By accurately recording the source and destination of the data, the full-link tracking capability of the lineage information is ensured. Defining data flow relationships enables full tracking and transparency of the data processing process, providing complete data flow information for subsequent dependency analysis and conflict detection; The node set and edge set of the graph are established based on the data dependency relationship. The construction of the node set and edge set is essentially to map the lineage information into a graph structure, providing a panoramic view of data dependency in the form of a graph, which is convenient for subsequent modeling and analysis. The establishment of node sets and edge sets forms the core data structure of the bloodline graph, which can efficiently support the storage, query and visualization of bloodline relationships; Ensure that the constructed graph structure meets the acyclic property. The data lineage graph needs to meet the characteristics of a directed acyclic graph (DAG) to ensure the logical consistency and parsability of data dependencies. Loop detection and removal ensure that the graph remains correct during dynamic updates; Through loop detection and removal, the loop-free nature of the bloodline graph is ensured, the uncertainty of data flow is avoided, and the stability and availability of the graph are improved.
[0026] Please see attached Figure 1 In a preferred embodiment of the present invention, S3, said hierarchical modeling of the data blood relationship graph includes: The nodes in the graph are divided into levels through topological sorting. Topological sorting algorithm: A topological sorting algorithm based on depth-first search (DFS) is used to sort the data lineage relationship graph. The nodes are divided into levels. The specific steps include: Initialization: for all nodes , , Set the unvisited state; Recursive search: Starting from any unvisited node Departure, press Point to recursively traverse its successor node; Adding nodes: After completing the node After traversing the successor nodes of Add to the result list; Arrange in reverse order: The final list is the result of topological sorting; Level division: For each node in the topological sorting result list, the level is calculated according to the number of its predecessor nodes;
[0027] in For Node The level of Representation Node The set of predecessor nodes of ; Topological sorting uses the characteristics of directed acyclic graphs to parse the pre-dependencies of nodes layer by layer to ensure the correct dependency order when the graph is updated. The hierarchical division provides the basis for subsequent priority calculation and hierarchical modeling; The topological sorting method realizes the dependency analysis and hierarchical division of data nodes, ensures the logic of the node update sequence, and lays the foundation for the accurate execution of iterative updates. The node priority is calculated based on the access frequency and change frequency of the node. The node priority takes into account the business importance of the node (reflected by the access frequency) and data stability (reflected by the change frequency). The weight parameter provides flexible adjustment capabilities to adapt to different scenario requirements. Priority calculation enables ranking of data nodes by importance, provides an accurate basis for the division of the core layer and edge layer, and ensures the optimality of resource allocation and update strategies; Nodes with higher priority are divided into the core layer, and nodes with lower priority are divided into the edge layer. Nodes in the core layer are usually data entities with high frequency of access or frequent changes. Their priority update has a direct impact on the overall performance and stability of the system; nodes in the edge layer can be updated later to save computing resources; By dividing the core layer and the edge layer, the update strategy is hierarchical and prioritized, which significantly improves the efficiency of iterative updates and the timeliness of key data. Mark the dependencies of nodes at each layer to ensure that the iterative updates are executed in priority by layer. The strategy of executing in priority by layer can ensure that key data nodes are updated first, while avoiding data inconsistency caused by incorrect update order of dependency paths; The marking of hierarchical dependencies and the priority update strategy optimize the execution logic of the update order, reduce the risk of update conflicts, and improve the update efficiency and stability of the system.
[0028] Please see attached Figure 1 In a preferred embodiment of the present invention, S4, the hierarchical modeling result performs iterative update on the blood relationship map, comprising: For core layer nodes, incremental update operations are performed in real time. The real-time update of core layer nodes ensures the synchronization of key data entities in the lineage map, can reflect high-frequency changes in a timely manner, and support the query and analysis needs of real-time data lineage; The real-time incremental update operation for core layer nodes significantly improves the timeliness of the bloodline map, providing technical support for business scenarios with high real-time requirements (such as exception tracking and data tracing); For edge layer nodes, incremental update operations are performed in batches according to time windows. The time window mechanism executes low-frequency update operations of the edge layer in batches, avoiding the consumption of system resources by frequent update requests while maintaining the availability and consistency of the data lineage map. Incremental updates of edge layer nodes are performed in batches according to time windows, balancing system performance and update timeliness, and reducing the interference of low-priority node updates on high-priority nodes. Detect data dependency conflicts during incremental updates and mark conflicting nodes and paths. Data dependency conflicts are usually caused by asynchronous node updates or changes in incremental data dependencies. The dependency conflict detection mechanism ensures the correctness of the blood relationship through hash verification and version comparison. The data dependency conflict detection and marking mechanism avoids logical errors caused by inconsistent updates of the bloodline graph, improving the integrity of data relationships and the stability of the system; Call the rule engine to automatically repair conflicting nodes and paths. The rule engine quickly resolves dependency conflicts through predefined rules and automated operations, ensuring the consistency and integrity of the bloodline map during the update process. Calling the rule engine to automatically repair conflicting nodes and paths greatly reduces manual intervention, improves the automation level and maintenance efficiency of bloodline map updates, and enhances the reliability of the system.
[0029] Please see attached Figure 1 In a preferred embodiment of the present invention, S5, said optimizing the processing and storage of incremental data by distributed computing comprises: Use a distributed graph database to store the bloodline graph. The distributed graph database stores the nodes and edges of the bloodline graph on different physical servers, which improves the storage capacity and concurrency performance while reducing the risk of single point failure. Using a distributed graph database to store bloodline graphs supports the storage requirements of large-scale bloodline graphs, while improving query efficiency and system fault tolerance. The data lineage map is partitioned and stored according to the access frequency and priority of the nodes. Partitioned storage reduces access latency by storing the core nodes with high frequency access in partitions with higher performance. Partitioned storage of low frequency access nodes improves the overall utilization efficiency of storage resources. Partitioning storage according to node priority optimizes the storage structure of the bloodline graph, increases the access speed of high-priority nodes, and reduces the overall storage and access costs of the system; Use a distributed computing framework to perform parallel processing on incremental data. The distributed computing framework significantly improves computing efficiency by processing large-scale incremental data in pieces and using the computing power of multiple nodes to parallelly calculate blood relationships. The distributed computing framework is used to process incremental data in parallel, which significantly improves the computing performance of updating large-scale data lineage graphs and supports real-time and high-frequency update scenarios. The processing results are synchronized to the graph database through the batch write interface. Batch write reduces the number of write operations to the database by merging write requests for updated data, thereby reducing write latency and improving database processing performance. By synchronizing incremental data through the batch write interface, the update speed of the bloodline map is optimized, the stability and consistency of the update operation are ensured, and the writing requirements in high-concurrency scenarios are met.
[0030] Please see attached Figure 1 In a preferred embodiment of the present invention, S6, the method of using stream processing technology to implement real-time bloodline map update includes: The stream processing engine monitors data change events. The core of the stream processing engine is to capture real-time change events. By monitoring the change log of the data source, the change content is transmitted to the lineage map update process in real time to achieve real-time response to data changes. By monitoring data change events through the stream processing engine, data changes can be captured with millisecond delays, meeting the needs of dynamic updates of data lineage graphs in high-real-time scenarios; Parse the captured event stream into incremental update requests. By parsing the event stream, convert the low-level log change information into high-level lineage map update instructions to ensure that the data lineage map is synchronized with the actual data status. Capturing event streams and parsing them into incremental update requests enables standardized and modular processing of data lineage graph updates, improving the scalability and compatibility of the system. Parse the data dependency relationships in the real-time event stream and generate incremental relationships. The dependency relationship analysis of the real-time event stream ensures the integrity and correctness of the data lineage graph. By dynamically parsing the relationship between upstream and downstream nodes, it solves the problem of dependency link breakage caused by data changes. Real-time analysis of dependency relationships in event streams to generate incremental relationships effectively ensures the logical consistency of the lineage graph and provides technical support for lineage maintenance in a dynamic data environment; The bloodline graph in the graph database is updated in real time based on the analysis results. The real-time update strategy based on the analysis results ensures the dynamic response capability of the bloodline graph to the event stream, and the transaction management mechanism enhances the reliability of incremental updates. The blood relationship map in the graph database is updated in real time based on the analysis results, which improves the efficiency of dynamic maintenance of the map and supports real-time query and analysis needs in high-frequency change scenarios.
[0031] Please see attached Figure 1In a preferred embodiment of the present invention, S7, the predicting of bloodline path changes using a machine learning model comprises: Extract the characteristics of historical lineage change data, including access frequency, dependency, and change frequency. By extracting the access behavior, dependency structure, and historical change records of nodes, a feature vector that fully describes the dynamic behavior of nodes is constructed to provide high-quality input data for machine learning models. Extracting the features of historical bloodline change data provides detailed input for predicting the change trend of bloodline paths, ensuring the accuracy and interpretability of the prediction results; Use the time series model to predict the data lineage change path. The LSTM model can capture the long-term and short-term dependencies in the time series features, and predict the future change path by analyzing the historical change trend, thus improving the accuracy and timeliness of the prediction results. The use of time series models to predict the path of data lineage changes effectively improves the system's ability to perceive future changes and provides a scientific basis for proactively adjusting update strategies. Generate priority update incremental data requests based on prediction results. By combining prediction probability and node priority, dynamically adjust the execution order of incremental update tasks to ensure that nodes with high impact or high possibility of change are processed first, optimizing system resource allocation. Prioritized incremental data requests are generated based on the prediction results, which enables intelligent sorting of update tasks and significantly improves the efficiency of incremental updates and the timeliness of important nodes.
[0032] Please see attached Figure 1 In a preferred embodiment of the present invention, S8, detecting and repairing abnormal nodes includes: Use anomaly detection algorithms to identify abnormal nodes in the bloodline graph. The anomaly detection algorithm analyzes the structure and attribute characteristics of the nodes in the bloodline graph to identify abnormal nodes that may cause data relationship errors or failures. The anomaly detection algorithm is used to automatically identify abnormal nodes in the bloodline map, which improves the efficiency of problem discovery and reduces the time cost of manual troubleshooting; Generate a repair strategy for the detected abnormal nodes. The repair strategy generates automated repair instructions based on the abnormal characteristics of the nodes and predefined rules to ensure that the problem repair is targeted and complete. The generated repair strategy provides a standardized anomaly repair solution, significantly reducing the complexity of manual intervention while improving the accuracy and efficiency of repair; Execute the repair strategy and rebuild the dependencies of abnormal nodes. The execution of the repair strategy restores the logical integrity and availability of the graph by dynamically adjusting the node attributes and dependencies of the bloodline graph. The execution of the repair strategy realizes the automated abnormal repair operation, ensures the integrity and correctness of the bloodline map, and avoids the impact of abnormal nodes on data flow analysis; Record the repair operation log for subsequent auditing. The repair log provides transparency of the operation, making it easier for system maintenance personnel to analyze the cause of the abnormality and verify the correctness of the repair results. Recording repair operation logs provides data support for system operation and maintenance and auditing, which helps to further optimize anomaly detection and repair strategies.
[0033] Please see attached Figure 1 In a preferred embodiment of the present invention, S9, the incremental data capture method includes: For large-scale distributed data sources, incremental data is captured in parallel by data sharding. The sharding parallel capture method for large-scale distributed data sources improves the speed and scalability of data capture and supports real-time extraction of incremental data in massive data scenarios. During the capture process, the consistency of cross-system data sources is detected to identify possible synchronization errors. Cross-system consistency detection quickly identifies shard data anomalies caused by synchronization delays, task interruptions, or configuration errors by comparing data snapshots and verifications. During the capture process, the consistency of cross-system data sources is detected, and possible synchronization errors are identified in a timely manner, ensuring the accuracy and integrity of incremental data; For the identified erroneous data, fault-tolerant processing is performed, including a retry mechanism or a data compensation strategy. The retry mechanism and the data compensation strategy ensure the reliability and integrity of the data capture task by dynamically recovering and reacquiring the lost data shards. The captured incremental data is marked as a traceable batch task for subsequent iterative updates. The batch task marking mechanism associates the incremental data with the task metadata to achieve the traceability of the incremental data and transparent management of the task status. Marking the captured incremental data as traceable batch tasks simplifies the subsequent incremental data management process and improves the transparency and operability of data processing.
[0034] In order to better understand the present invention, the above contents are described in detail below in conjunction with specific embodiments.
[0035] Example 1: A method for iteratively updating a bloodline graph without using stream processing technology Specific implementation Use batch processing to capture incremental data and identify data changes by regularly generating data snapshots for comparison.
[0036] Graph updates are performed in batch mode, and full topological sorting is performed each time to ensure the correctness of dependencies.
[0037] Data anomaly nodes are completed through manual analysis and manual repair.
[0038] Effect It is suitable for static data environments or low-frequency data update scenarios, but the update delay is high and it is difficult to support dynamic data environments with high real-time requirements.
[0039] Lack of automated support for anomaly detection and repair leads to high maintenance costs.
[0040] Embodiment 2: Real-time bloodline graph update method using stream processing technology Specific implementation Use stream processing technologies such as Apache Kafka and Apache Flink to capture data change events, parse incremental data in real time, and update the lineage graph.
[0041] The graph update order is optimized through hierarchical modeling, core layer nodes are updated in real time, and edge layer nodes are updated in batches according to time windows.
[0042] During the incremental update process, abnormal nodes are automatically detected and marked, and abnormal paths are manually repaired.
[0043] Effect Improved real-time performance, able to quickly respond to high-frequency data change scenarios.
[0044] Automated anomaly detection has partially improved the efficiency of problem identification, but repairs still rely on manual intervention, resulting in the failure to further reduce maintenance costs.
[0045] Example 3: Intelligent bloodline map updating method combining machine learning and automated repair mechanism Specific implementation Combined with the stream processing technology of Example 2, the data lineage path changes are predicted through a machine learning model, the update order is dynamically adjusted, and key nodes are updated first.
[0046] The rule engine is called to automatically repair the detected abnormal nodes, and the repair operation is completed through dependency reconstruction and attribute adjustment.
[0047] Record repair operation logs for subsequent auditing and system optimization.
[0048] Effect It achieves efficient real-time updates and reduces maintenance costs through path prediction and automatic abnormal repair mechanisms.
[0049] It has a higher level of intelligence and adaptability, and is suitable for highly dynamic, multi-source heterogeneous data scenarios.
[0050] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An iterative updating method for a data lineage graph, characterized in that: The following steps are involved: Capture data increments and metadata changes; Construct a directed acyclic graph of data lineage relationships; Perform hierarchical modeling on the data lineage graph; Iteratively update the blood relationship map based on the hierarchical modeling results; Optimize the processing and storage of incremental data through distributed computing; Use stream processing technology to achieve real-time updates of bloodline graphs; Use machine learning models to predict changes in bloodline paths and detect and repair abnormal nodes.
2. The iterative update method for data lineage graph according to claim 1, characterized in that: The capturing of data increments and metadata changes includes: Get data snapshots, including new and old snapshots of data; Identify incremental portions of data by comparing snapshots; Perform timestamp filtering, hash value verification, and version comparison on incremental data to extract new, changed, and deleted data; Capture metadata changes, including additions, modifications, and deletions at the table and field levels; Standardize the recognition results into incremental data event format for subsequent processing.
3. The iterative updating method for data lineage graph according to claim 1, characterized in that: The directed acyclic graph for constructing data kinship relationships includes: Define data nodes, including table-level and field-level data entities; Define data flow relationships, including input relationships and output relationships; Establish the node set and edge set of the graph based on data dependency; Ensure that the constructed graph structure satisfies the acyclic property.
4. The iterative updating method for data lineage graph according to claim 1, characterized in that: The hierarchical modeling of the data blood relationship graph includes: Divide the nodes in the graph into levels through topological sorting; Calculate node priority based on node access frequency and change frequency; Nodes with higher priority are classified as core layers, and nodes with lower priority are classified as edge layers; Mark the dependencies of nodes at each layer to ensure that they are executed first by layer during iterative updates.
5. The iterative updating method for data lineage graph according to claim 1, characterized in that: The hierarchical modeling results perform iterative updates on the blood relationship map, including: For core layer nodes, incremental update operations are performed in real time; For edge layer nodes, incremental update operations are performed in batches according to time windows; Detect data dependency conflicts during incremental updates and mark conflicting nodes and paths; Call the rule engine to automatically repair conflicting nodes and paths.
6. The iterative updating method for data lineage graph according to claim 1, characterized in that: Optimizing the processing and storage of incremental data through distributed computing includes: Use a distributed graph database to store bloodline graphs; The data lineage graph is partitioned and stored according to the access frequency and priority of the nodes; Use a distributed computing framework to perform parallel processing on incremental data; The processing results are synchronized to the graph database through the batch write interface.
7. The iterative updating method for data lineage graph according to claim 1, characterized in that: The method of using stream processing technology to implement real-time bloodline graph update includes: Monitor data change events through the stream processing engine; Parse the captured event stream into incremental update requests; Parse data dependencies in real-time event streams and generate incremental relationships; The blood relationship map in the graph database is updated in real time based on the analysis results.
8. The iterative updating method for data lineage graph according to claim 1, characterized in that: The use of a machine learning model to predict changes in bloodline pathways includes: Extract the characteristics of historical lineage change data, including access frequency, dependency, and change frequency; Use time series models to predict the path of data lineage changes; Generates priority update incremental data requests based on prediction results.
9. The iterative updating method for data lineage graph according to claim 8, characterized in that: The detecting and repairing of abnormal nodes includes: Use anomaly detection algorithms to identify abnormal nodes in the bloodline graph; Generate repair strategies for detected abnormal nodes; Execute the repair strategy and rebuild the dependency relationship of abnormal nodes; Record the repair operation log for subsequent audit.
10. The iterative updating method for data lineage graph according to claim 2, characterized in that: The incremental data capturing method comprises: For large-scale distributed data sources, incremental data is captured in parallel by data sharding; Check the consistency of data sources across systems during the capture process and identify possible synchronization errors; Perform fault tolerance processing for identified erroneous data, including retry mechanisms or data compensation strategies; The captured incremental data is marked as traceable batch tasks for subsequent iterative updates.
Citation Information
Patent Citations
Power data traceability method and system based on data blood relationship graph
CN114491081A
Method and device for constructing metadata consanguinity map and related equipment
CN114510611A
Distributed data blood relationship construction and display method
CN116662441A
Data blood relationship analysis method, device and equipment and readable storage medium
CN118394829A
Data consanguinity traceability analysis method based on power grid data center
CN119127917A
Cited By
Local knowledge base automatic construction system based on multi-source acquisition and distributed computing
CN120723742A
Metadata self-description and blood relationship tracking method and system based on Flink
CN121029820A
A flink-based metadata self-description and bloodline tracking method and system
CN121029820B
Multi-source index data intelligent storage management method and system
CN121412211A
Data value evaluation and blood relationship map construction method based on artificial intelligence
CN122489549A