Data dynamic blood relationship extraction method, device, equipment, medium and product
By monitoring and constructing dynamic data lineage in real time within a big data scheduling platform, the problems of low efficiency, insufficient accuracy, and poor flexibility in existing technologies have been solved. This enables efficient and accurate data traceability and governance, making it suitable for industries such as finance and healthcare.
Patent Information
- Application Number
- CN202510911076.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-21
AI Technical Summary
Existing data lineage extraction technology has problems such as low extraction efficiency, insufficient accuracy and poor flexibility in big data environments. In particular, lineage capture fails in dynamic scheduling scenarios, and the lineage gaps in heterogeneous systems and the shortcomings in metadata management collaboration are prominent, making it difficult to meet the real-time requirements of data governance.
The big data scheduling platform monitors the running status of data processing tasks in real time, collects metadata and constructs dynamic data lineage relationships, uses graph databases for storage and management, and combines artificial intelligence and heterogeneous system semantic mapping to realize the automated construction and updating of dynamic lineage relationships.
It improves the efficiency and accuracy of data lineage extraction, enhances the system's flexibility and adaptability, provides precise lineage support, and supports data governance and traceability, making it particularly suitable for the financial and healthcare industries.
Smart Images

Figure CN120821739A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data processing, and specifically relates to a method, device, equipment, medium and product for extracting dynamic blood relationships from data. Background Art
[0002] In today's big data era, enterprise-level big data architectures have evolved into a complex ecosystem encompassing diverse components, including ETL (Extract-Transform-Load) tools, real-time computing frameworks, data warehouses, and integrated data lake and warehouse platforms. Data frequently circulates across processes such as Kafka stream processing (data streams streamed through the Kafka platform), Spark batch processing (batch processing is one of the most common and fundamental data processing models in Spark; batch processing involves breaking datasets into small chunks / batches and processing each batch independently), and Hive (a Hadoop-based data warehouse framework that uses simple SQL statements to query and analyze data stored in HDFS and converts these statements into MapReduce programs for data processing). This has exposed multiple technical bottlenecks in traditional data lineage extraction technologies, including the following:
[0003] (1) The problem of lineage discontinuity in heterogeneous systems is prominent. That is, existing data lineage extraction methods have a semantic gap in lineage mapping between Hadoop ecosystems (such as Hive or Pig) and cloud-native architectures (such as Snowflake or Databricks). For example, the lineage traces of state storage in Spark / Flink stream processing are often lost due to the lack of operator state serialization mechanisms, resulting in a break in the lineage chain between real-time data and offline data.
[0004] (2) Lineage capture fails in dynamic scheduling scenarios. Existing big data scheduling platforms generally use workflow engines such as Airflow or Oozie. Their DAG (Directed Acyclic Graph) tasks have dynamic characteristics such as parameterized execution (such as dynamically generating task nodes by date partition) and conditional branching (such as triggering circuit breakers when data quality does not meet standards). Traditional static lineage analysis cannot capture the instantiation lineage of parameterized nodes. When the scheduling cycle is upgraded from T+1 days to minutes, the lineage lag problem causes data traceability errors exceeding 40%.
[0005] (3) The metadata management system has coordination shortcomings. While existing mainstream metadata management platforms (such as Atlas or Amundsen) support lineage storage, metadata synchronization with the scheduling platform suffers from issues such as poor timeliness (batch synchronization interval ≥ 1 hour) and insufficient field-level lineage granularity (only table-level lineage is recorded). When the data center performs field-level lineage analysis, manual correlation between scheduling logs and calculation code is required, making it difficult to meet the real-time requirements of data governance.
[0006] The aforementioned technical pain points stem from three contradictions within big data scheduling platforms: the conflict between dynamic scheduling logic and static lineage models, the semantic barriers of heterogeneous computing engines, and the mismatch between unstructured computing logic and structured lineage representation. Therefore, a new data lineage extraction solution is urgently needed that can adapt to the dynamic characteristics of big data scheduling platforms and cover all computing scenarios. Summary of the Invention
[0007] The purpose of the present invention is to provide a method, device, computer equipment, computer-readable storage medium and computer program product for extracting dynamic blood relationships from data, so as to solve the problems of low extraction efficiency, insufficient accuracy and poor flexibility in existing data blood relationship extraction technologies.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] In a first aspect, a method for extracting dynamic blood relationships from data is provided, comprising:
[0010] In the big data scheduling platform, metadata is collected for the data sources, data processing components, and data targets involved in each data processing task submitted to the platform, and metadata including data source metadata, component metadata, and target metadata is obtained, wherein the data source metadata includes the type, structure information, and field information of the data source, the component metadata includes the type and parameter configuration information of the data processing component, and the target metadata includes the storage location and structure information of the data target;
[0011] Utilizing the task scheduling mechanism of the big data scheduling platform, the running status of the data processing task is monitored in real time during task execution to capture key events in the data processing process, wherein the key events include data reading events, data conversion events, and data writing events;
[0012] Based on the metadata and the event information of the key events, a data dynamic lineage relationship is constructed, specifically including: for the data reading event, determining the corresponding data source, and establishing an input relationship between the data target and the data source; for the data conversion event, analyzing the field mapping relationship and data processing logic of the corresponding data during the conversion process, and constructing an intermediate lineage relationship during the data processing process; for the data writing event, determining the destination of the corresponding data, and establishing an output relationship between the data source and the data target; associating and integrating the input relationship, the intermediate lineage relationship, and the output relationship to form a complete data dynamic lineage relationship diagram;
[0013] The dynamic blood relationship of the data is stored in a graph database for query and management.
[0014] Based on the above invention content, a new solution for extracting dynamic data lineage relationships based on a big data scheduling platform is provided, that is, on the one hand, metadata is collected for the data sources, data processing components and data targets involved in each data processing task submitted to the platform in the big data scheduling platform to obtain metadata including data source metadata, component metadata and target metadata; on the other hand, the task scheduling mechanism of the big data scheduling platform is utilized to monitor the running status of the data processing task in real time during the task execution to capture key events in the data processing process; finally, based on the metadata and event information of the key events, dynamic data lineage relationships are constructed and stored in a graph database for query and management, which can solve the problems of low extraction efficiency, insufficient accuracy and poor flexibility in the existing technology, break through the bottlenecks of the existing technology in terms of dynamicity, heterogeneity and semantic analysis, and provide precise lineage support for data governance. It is especially suitable for industry scenarios such as finance or medical care that have strict requirements on data traceability, and is convenient for practical application and promotion.
[0015] In one possible design, data lineage extraction logic is embedded in the task lifecycle of the big data scheduling platform in any one of the following ways (A1) to (A4) or any combination thereof to achieve synchronous triggering of the scheduling process and lineage collection:
[0016] (A1) adding lineage collection parameters to the task definition of the big data scheduling platform, wherein the lineage collection parameters include a task input table identifier, a task output table identifier, and a field mapping tag;
[0017] (A2) Utilize the hook function mechanism of the scheduling engine to parse the task's SQL / script file before task execution to obtain static lineage, and / or obtain dynamic data flow through runtime tracing during task execution;
[0018] (A3) Designing semantic automatic mapping capabilities for heterogeneous systems: Build a semantic mapping dictionary across engine operators to support automatic conversion of operational semantics across different engines, and / or add a lineage parsing module for unstructured data to leverage NLP techniques to extract data flow relationships from engine execution logs;
[0019] (A4) Design a lineage collection strategy for edge computing scenarios: Design a local lineage caching mechanism for edge nodes to support temporary storage of lineage data during network outages and batch synchronization after network recovery, and / or optimize the lineage transmission protocol between edge nodes and central nodes to reduce data transmission volume and latency.
[0020] In one possible design, the running status of the data processing task is monitored in real time during task execution to capture key events in the data processing process, including:
[0021] During the execution of the task, the running status of the data processing task is monitored in real time, and key events in the data processing process are captured by log collection points embedded in the data processing framework or by using interceptors, where the key events include data reading events, data conversion events and data writing events.
[0022] In one possible design, in the process of building dynamic data lineage relationships, the method also includes: integrating the task dependencies of the big data scheduling platform with the data lineage relationships, building a lineage topology diagram with dual associations between tasks and data, and supporting full-link lineage tracking of data flow from ETL tasks to report tasks.
[0023] In one possible design, the task dependency and data lineage relationships of the big data scheduling platform are integrated to construct a lineage topology diagram with dual associations between tasks and data, and support lineage tracking of the entire data flow from ETL tasks to reporting tasks, including any one of the following methods (B1) to (B4) or any combination thereof:
[0024] (B1) parsing task metadata contained in a DAG node of a task dependency relationship of the big data scheduling platform to extract field-level lineage, wherein the task metadata includes an SQL script, a task input table, and a task output table;
[0025] (B2) Utilize scheduling dependencies to automatically associate cross-task lineage chains;
[0026] (B3) Using artificial intelligence technology to intelligently infer implicit blood relationships: Introducing business knowledge graphs as prior knowledge to enhance the model's understanding of blood relationships; constructing and training an artificial intelligence model for inferring implicit blood relationships; using the generated explicit blood relationships and corresponding features as training samples, performing supervised learning training on the artificial intelligence model to enable the model to learn the mapping relationship between data features and blood relationships;
[0027] (B4) Design a multi-level strategy for dynamic lineage updating: when the data processing task is created, possible lineage change paths are generated in advance based on the implicit lineage relationship inferred intelligently, and a lineage snapshot is automatically generated when the data processing task is retried or the parameters are dynamically adjusted.
[0028] In one possible design, after constructing the dynamic data lineage relationship, the method further includes iteratively designing a lineage relationship incremental update mechanism for tasks in the big data scheduling platform and including the following methods (C1) and / or (C2):
[0029] (C1) Compare the SQL / script files of the new and old tasks and only update the lineage relationship of the changed parts;
[0030] (C2) Using the task execution log of the big data scheduling platform to record dynamic change information during the data processing process.
[0031] In a second aspect, a data dynamic blood relationship extraction device is provided, comprising a metadata collection unit, a key event capture unit, a blood relationship construction unit, and a blood relationship storage unit;
[0032] The metadata collection unit is used to collect metadata about the data sources, data processing components, and data targets involved in each data processing task submitted to the platform in the big data scheduling platform, and obtain metadata including data source metadata, component metadata, and target metadata, wherein the data source metadata includes the type, structure information, and field information of the data source, the component metadata includes the type and parameter configuration information of the data processing component, and the target metadata includes the storage location and structure information of the data target;
[0033] The key event capture unit is used to utilize the task scheduling mechanism of the big data scheduling platform to monitor the running status of the data processing task in real time during the task execution process to capture key events in the data processing process, wherein the key events include data reading events, data conversion events and data writing events;
[0034] The blood relationship construction unit is communicatively connected to the metadata collection unit and the key event capture unit respectively, and is used to construct a data dynamic blood relationship based on the metadata and the event information of the key event, specifically including: for the data reading event, determining the corresponding data source, and establishing an input relationship between the data target and the data source; for the data conversion event, analyzing the field mapping relationship and data processing logic of the corresponding data in the conversion process, and constructing an intermediate blood relationship in the data processing process; for the data writing event, determining the destination of the corresponding data, and establishing an output relationship between the data source and the data target; associating and integrating the input relationship, the intermediate blood relationship and the output relationship to form a complete data dynamic blood relationship diagram;
[0035] The blood relationship storage unit is communicatively connected to the blood relationship construction unit and is used to store the data dynamic blood relationship in a graph database for query and management.
[0036] In a third aspect, the present invention provides a computer device comprising a memory, a processor and a transceiver which are communicatively connected in sequence, wherein the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the data dynamic blood relationship extraction method as described in the first aspect or any possible design of the first aspect.
[0037] In a fourth aspect, the present invention provides a computer-readable storage medium having instructions stored thereon. When the instructions are run on a computer, the method for extracting dynamic blood relationships from data as described in the first aspect or any possible design of the first aspect is executed.
[0038] In a fifth aspect, the present invention provides a computer program product, comprising a computer program or instructions, which, when executed by a computer, implements the data dynamic blood relationship extraction method as described in the first aspect or any possible design of the first aspect.
[0039] Beneficial effects of the above scheme:
[0040] (1) The present invention creatively provides a new solution for extracting dynamic blood relationship of data based on a big data scheduling platform, namely, on the one hand, metadata of data sources, data processing components and data targets involved in each data processing task submitted to the platform are collected in the big data scheduling platform to obtain metadata including data source metadata, component metadata and target metadata; on the other hand, the task scheduling mechanism of the big data scheduling platform is utilized to monitor the running status of the data processing task in real time during the task execution to capture key events in the data processing process; finally, based on the metadata and event information of the key events, the dynamic blood relationship of data is constructed and stored in a graph database for query and management, thereby solving the problems of low extraction efficiency, insufficient accuracy and poor flexibility existing in the prior art, breaking through the bottlenecks of the prior art in terms of dynamicity, heterogeneity and semantic analysis, and providing accurate blood relationship support for data governance, which is particularly suitable for industry scenarios such as finance or medical care that have strict requirements on data traceability;
[0041] (2) It can improve the efficiency of data kinship extraction, that is, through automated data collection and analysis processing, it greatly reduces manual intervention, can quickly process large amounts of data, and improve processing efficiency;
[0042] (3) It can improve the accuracy of data kinship extraction, that is, the algorithm that integrates rule reasoning and machine learning can better cope with complex data processing scenarios, accurately identify the kinship between data, and reduce the error rate;
[0043] (4) It can enhance the flexibility and adaptability of the system, that is, the multi-dimensional data collection and dynamic update mechanism enables the system to adapt to the dynamic changes of the data processing flow in the big data scheduling platform and update the blood relationship information in a timely manner;
[0044] (5) It can provide good visualization effects, that is, intuitive graphical display, which makes it easy for users to quickly understand the source and destination of data, and provides strong support for data governance and decision-making;
[0045] (6) It helps improve data quality and data governance. That is, accurate kinship information can help users promptly identify problems in the data processing process and trace the source of the data, thereby improving data quality and strengthening data governance.
[0046] (7) Through this solution, enterprises can achieve full-chain visibility, manageability and traceability of data assets, significantly improve data governance efficiency and compliance level, and facilitate practical application and promotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 A flow chart of the method for extracting dynamic blood relationships from data provided in an embodiment of the present application.
[0049] Figure 2 This is an example diagram of the full-link data lineage visualization results provided in the embodiment of this application.
[0050] Figure 3 A schematic diagram of the structure of the data dynamic blood relationship extraction device provided in an embodiment of the present application.
[0051] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the present invention will be briefly introduced below in conjunction with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the structures of the drawings is only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these embodiments without creative work. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.
[0053] It should be understood that although the terms first, second, etc. may be used herein to describe various objects, these objects should not be limited by these terms. These terms are merely used to distinguish one object from another. For example, a first object can be referred to as a second object, and similarly, a second object can be referred to as a first object without departing from the scope of the exemplary embodiments of the present invention.
[0054] It should be understood that the term "and / or" that may appear in this document is merely a description of the association relationship between associated objects, indicating that there may be three relationships. For example, A and / or B can indicate three situations: A exists alone, B exists alone, or A and B exist at the same time. For another example, A, B and / or C can indicate the existence of any one of A, B and C or any combination of them. The term " / and" that may appear in this document describes another type of association object relationship, indicating that there may be two relationships. For example, A / and B can indicate two situations: A exists alone or A and B exist at the same time. In addition, the character " / " that may appear in this document generally indicates that the previous and next associated objects are in an "or" relationship.
[0055] Example
[0056] like Figure 1 As shown, the data dynamic blood relationship extraction method provided in the first aspect of this embodiment can be, but is not limited to, executed by a computer device with certain computing resources, such as a cloud platform server or other electronic device. Figure 1 As shown, the data dynamic blood relationship extraction method may include, but is not limited to, the following steps S1 to S4.
[0057] S1. In the big data scheduling platform, metadata is collected for the data sources, data processing components, and data targets involved in each data processing task submitted to the platform, and metadata including data source metadata, component metadata, and target metadata is obtained, wherein the data source metadata includes but is not limited to the type, structure information, and field information of the data source, the component metadata includes but is not limited to the type and parameter configuration information of the data processing component, and the target metadata includes but is not limited to the storage location and structure information of the data target.
[0058] In step S1, in order to achieve the real-time performance of "scheduling is collection" through the deep coupling design of the scheduling platform and lineage extraction, preferably, the data lineage extraction logic is embedded in the task life cycle of the big data scheduling platform in any one of the following ways (A1) to (A4) or any combination thereof to achieve the synchronous triggering of the scheduling process and lineage collection: (A1) adding lineage collection parameters to the task definition of the big data scheduling platform, wherein the lineage collection parameters include but are not limited to task input table identifier, task output table identifier and field mapping tag, etc.; (A2) utilizing the hook function mechanism of the scheduling engine to parse the SQL (Structured Query Language) / script file of the task before task execution to obtain static lineage, and / or obtain dynamic data flow through the runtime tracking function during task execution; (A3) designing the semantic automatic mapping capability of heterogeneous systems: constructing a semantic mapping dictionary of cross-engine operators to support automatic conversion of operational semantics of different engines, and / or adding a lineage parsing module for unstructured data to utilize NLP (Natural Language Processing (Natural Language Processing) technology extracts data flow relationships from engine execution logs; (A4) Designs a lineage collection strategy for edge computing scenarios: Designs a local lineage caching mechanism for edge nodes to support temporary storage of lineage data during network outages and batch synchronization after network restoration, and / or optimizes the lineage transmission protocol between edge nodes and central nodes to reduce data transmission volume and latency. Examples of the operational semantics of the aforementioned different engines are Spark's join and Flink's connect.
[0059] After step S1, in order to improve the data quality of the metadata, it is necessary to pre-process the collected metadata, such as performing data cleaning and format standardization, so as to ensure the accuracy and consistency of the metadata.
[0060] S2. Utilize the task scheduling mechanism of the big data scheduling platform to monitor the running status of the data processing task in real time during task execution to capture key events in the data processing process, wherein the key events include but are not limited to data reading events, data conversion events, and data writing events.
[0061] In step S2, the event information of the key event contains key information of the data in the flow process, such as the identifier of the data source, the identifier of the data target, and the type of data processing operation, etc., and is therefore closely related to the data and needs to be captured during the data processing process. Specifically, the running status of the data processing task is monitored in real time during the task execution to capture key events in the data processing process, including but not limited to: monitoring the running status of the data processing task in real time during the task execution, and capturing key events in the data processing process through log collection points embedded in the data processing framework or using interceptors, wherein the key events include but are not limited to data reading events, data conversion events, and data writing events.
[0062] S3. Construct a dynamic data lineage relationship based on the metadata and the event information of the key events, specifically including: for the data reading event, determining the corresponding data source, and establishing an input relationship between the data target and the data source; for the data conversion event, analyzing the field mapping relationship and data processing logic of the corresponding data during the conversion process, and constructing an intermediate lineage relationship during the data processing process; for the data writing event, determining the destination of the corresponding data, and establishing an output relationship between the data source and the data target; associating and integrating the input relationship, the intermediate lineage relationship and the output relationship to form a complete data dynamic lineage relationship diagram.
[0063] In the step S3, the lineage topology of the data dynamic lineage relationship is constructed based on the scheduling dependency graph. Preferably, in the process of constructing the data dynamic lineage relationship, the method further includes: fusing the task dependency relationship of the big data scheduling platform with the data lineage relationship, constructing a lineage topology graph with dual association between tasks and data, and supporting full-link lineage tracking of data flow from ETL tasks to report tasks, including any one of the following methods (B1) to (B4) or any combination thereof: (B1) parsing the task metadata contained in the DAG node of the task dependency relationship of the big data scheduling platform to extract field-level lineage, wherein the task metadata includes but is not limited to SQL scripts, task input tables, and task output tables; (B2 ) Using scheduling dependencies to automatically associate cross-task bloodline chains; (B3) Using artificial intelligence technology to intelligently infer implicit bloodline relationships: Introduce business knowledge graphs as prior knowledge to enhance the model's understanding of bloodline relationships; Build and train an artificial intelligence model for inferring implicit bloodline relationships; Use the generated explicit bloodline relationships and corresponding features as training samples to perform supervised learning training on the artificial intelligence model so that the model can learn the mapping relationship between data features and bloodline relationships; (B4) Design a multi-level strategy for dynamic bloodline updates: When the data processing task is created, possible bloodline change paths are generated in advance based on the intelligently inferred implicit bloodline relationships, and bloodline snapshots are automatically generated when the data processing task is retried or parameters are dynamically adjusted. The aforementioned scheduling dependency can be exemplified by the output table of task A being the input table of task B. The hyperparameters of the aforementioned artificial intelligence model can be optimized and determined using existing optimization algorithms (such as the Bayesian optimization algorithm) to reduce manual configuration and improve the efficiency and accuracy of data bloodline relationship extraction.
[0064] After step S3, it is necessary to perform incremental updates and consistency maintenance of dynamic lineage, that is, preferably, after constructing the dynamic lineage relationship of the data, the method also includes a lineage relationship incremental update mechanism designed for task iterations (such as version updates and parameter changes, etc.) in the big data scheduling platform and including the following methods (C1) and / or (C2): (C1) comparing the differences in SQL / script files between new and old tasks, and only updating the lineage relationship of the changed part (such as the addition and deletion of fields); (C2) using the task execution log of the big data scheduling platform to record dynamic change information during the data processing process (such as changes in the input table caused by parameters during runtime).
[0065] S4. Store the dynamic blood relationship of the data in a graph database for query and management.
[0066] In step S4, the nodes in the graph database represent entities such as data sources, data processing components or data targets, and the edges represent the blood relationships between entities. At the same time, a metadata center can be established to uniformly manage the collected metadata and the constructed blood relationships to achieve metadata updates and blood relationship maintenance. After step S4, users can also use the user-friendly query interface and display interface to easily query the dynamic blood relationship of the data, that is, users can query according to conditions such as data source, data target or time range, and obtain a data blood relationship diagram displayed in an intuitive graphical manner, including information such as the source of the data, processing path and destination, such as Figure 2 The full-link data lineage visualization results are shown.
[0067] Based on the data dynamic lineage relationship extraction method described in the aforementioned steps S1 to S4, a new solution for extracting data dynamic lineage relationships based on a big data scheduling platform is provided. That is, on the one hand, metadata is collected on the big data scheduling platform for the data sources, data processing components, and data targets involved in each data processing task submitted to the platform to obtain metadata including data source metadata, component metadata, and target metadata. On the other hand, the task scheduling mechanism of the big data scheduling platform is utilized to monitor the running status of the data processing task in real time during task execution to capture key events in the data processing process. Finally, based on the metadata and event information of the key events, dynamic data lineage relationships are constructed and stored in a graph database for query and management. This can solve the problems of low extraction efficiency, insufficient accuracy, and poor flexibility existing in the existing technology, break through the bottlenecks of the existing technology in terms of dynamicity, heterogeneity, and semantic analysis, and provide precise lineage support for data governance. It is particularly suitable for industry scenarios such as finance or healthcare that have strict requirements on data traceability. In addition, through this method, enterprises can achieve full-link visibility, controllability, and traceability of data assets, significantly improve data governance efficiency and compliance levels, and facilitate practical application and promotion.
[0068] like Figure 3 As shown, the second aspect of this embodiment provides a virtual device for implementing the data dynamic blood relationship extraction method described in the first aspect, including a metadata collection unit, a key event capture unit, a blood relationship construction unit and a blood relationship storage unit;
[0069] The metadata collection unit is used to collect metadata about the data sources, data processing components, and data targets involved in each data processing task submitted to the platform in the big data scheduling platform, and obtain metadata including data source metadata, component metadata, and target metadata, wherein the data source metadata includes the type, structure information, and field information of the data source, the component metadata includes the type and parameter configuration information of the data processing component, and the target metadata includes the storage location and structure information of the data target;
[0070] The key event capture unit is used to utilize the task scheduling mechanism of the big data scheduling platform to monitor the running status of the data processing task in real time during the task execution process to capture key events in the data processing process, wherein the key events include data reading events, data conversion events and data writing events;
[0071] The blood relationship construction unit is communicatively connected to the metadata collection unit and the key event capture unit respectively, and is used to construct a data dynamic blood relationship based on the metadata and the event information of the key event, specifically including: for the data reading event, determining the corresponding data source, and establishing an input relationship between the data target and the data source; for the data conversion event, analyzing the field mapping relationship and data processing logic of the corresponding data in the conversion process, and constructing an intermediate blood relationship in the data processing process; for the data writing event, determining the destination of the corresponding data, and establishing an output relationship between the data source and the data target; associating and integrating the input relationship, the intermediate blood relationship and the output relationship to form a complete data dynamic blood relationship diagram;
[0072] The blood relationship storage unit is communicatively connected to the blood relationship construction unit and is used to store the data dynamic blood relationship in a graph database for query and management.
[0073] The working process, working details and technical effects of the aforementioned device provided in the second aspect of this embodiment can be found in the data dynamic blood relationship extraction method described in the first aspect, and will not be repeated here.
[0074] like Figure 4As shown, the third aspect of this embodiment provides a computer device for executing the data dynamic blood relationship extraction method as described in the first aspect, including a memory, a processor and a transceiver that are sequentially connected in communication, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the data dynamic blood relationship extraction method as described in the first aspect. For example, the memory may include, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a flash memory (Flash Memory), a first-in-first-out memory (FIFO) and / or a first-in-last-out memory (FILO), etc.; the processor may include, but is not limited to, a microprocessor of the STM32F105 series. In addition, the computer device may also include, but is not limited to, a power module, a display screen and other necessary components.
[0075] The working process, working details and technical effects of the aforementioned computer device provided in the third aspect of this embodiment can be found in the data dynamic blood relationship extraction method described in the first aspect, and will not be repeated here.
[0076] A fourth aspect of this embodiment provides a computer-readable storage medium storing instructions containing the method for extracting dynamic blood relationships from data as described in the first aspect, that is, the computer-readable storage medium stores instructions that, when executed on a computer, execute the method for extracting dynamic blood relationships from data as described in the first aspect. The computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, computer-readable storage media such as a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash drive, and / or a memory stick. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
[0077] The working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment can be found in the data dynamic blood relationship extraction method described in the first aspect, and will not be repeated here.
[0078] A fifth aspect of this embodiment provides a computer program product, including a computer program or instructions, which, when executed by a computer, implements the method for extracting dynamic blood relationships from data as described in the first aspect. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
[0079] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A method for extracting dynamic blood relationship from data, characterized in that: include: In the big data scheduling platform, metadata is collected for the data sources, data processing components, and data targets involved in each data processing task submitted to the platform, and metadata including data source metadata, component metadata, and target metadata is obtained, wherein the data source metadata includes the type, structure information, and field information of the data source, the component metadata includes the type and parameter configuration information of the data processing component, and the target metadata includes the storage location and structure information of the data target; Utilizing the task scheduling mechanism of the big data scheduling platform, the running status of the data processing task is monitored in real time during task execution to capture key events in the data processing process, wherein the key events include data reading events, data conversion events, and data writing events; Based on the metadata and the event information of the key events, a data dynamic lineage relationship is constructed, specifically including: for the data reading event, determining the corresponding data source, and establishing an input relationship between the data target and the data source; for the data conversion event, analyzing the field mapping relationship and data processing logic of the corresponding data during the conversion process, and constructing an intermediate lineage relationship during the data processing process; for the data writing event, determining the destination of the corresponding data, and establishing an output relationship between the data source and the data target; associating and integrating the input relationship, the intermediate lineage relationship, and the output relationship to form a complete data dynamic lineage relationship diagram; The dynamic blood relationship of the data is stored in a graph database for query and management.
2. The method for extracting dynamic blood relationships from data according to claim 1, characterized in that: In the task lifecycle of the big data scheduling platform, data lineage extraction logic is embedded in any one of the following methods (A1) to (A4) or any combination thereof to achieve synchronous triggering of the scheduling process and lineage collection: (A1) adding lineage collection parameters to the task definition of the big data scheduling platform, wherein the lineage collection parameters include a task input table identifier, a task output table identifier, and a field mapping tag; (A2) Utilize the hook function mechanism of the scheduling engine to parse the task's SQL / script file before task execution to obtain static lineage, and / or obtain dynamic data flow through runtime tracing during task execution; (A3) Designing semantic automatic mapping capabilities for heterogeneous systems: Build a semantic mapping dictionary across engine operators to support automatic conversion of operational semantics across different engines, and / or add a lineage parsing module for unstructured data to leverage NLP techniques to extract data flow relationships from engine execution logs; (A4) Design a lineage collection strategy for edge computing scenarios: Design a local lineage caching mechanism for edge nodes to support temporary storage of lineage data during network outages and batch synchronization after network recovery, and / or optimize the lineage transmission protocol between edge nodes and central nodes to reduce data transmission volume and latency.
3. The method for extracting dynamic blood relationships from data according to claim 1, characterized in that: During the execution of the task, the running status of the data processing task is monitored in real time to capture key events in the data processing process, including: During the execution of the task, the running status of the data processing task is monitored in real time, and key events in the data processing process are captured by log collection points embedded in the data processing framework or by using interceptors, where the key events include data reading events, data conversion events and data writing events.
4. The method for extracting dynamic blood relationships from data according to claim 1, characterized in that: In the process of building dynamic data lineage relationships, the method also includes: integrating the task dependencies of the big data scheduling platform with the data lineage relationships, building a lineage topology diagram with dual associations between tasks and data, and supporting full-link lineage tracking of data flow from ETL tasks to report tasks.
5. The method for extracting dynamic blood relationship from data according to claim 1, characterized in that: The task dependency and data lineage relationships of the big data scheduling platform are integrated to construct a lineage topology diagram with dual associations between tasks and data, and support lineage tracking of the entire data flow from ETL tasks to reporting tasks, including any one of the following methods (B1) to (B4) or any combination thereof: (B1) parsing task metadata contained in a DAG node of a task dependency relationship of the big data scheduling platform to extract field-level lineage, wherein the task metadata includes an SQL script, a task input table, and a task output table; (B2) Utilize scheduling dependencies to automatically associate cross-task lineage chains; (B3) Using artificial intelligence technology to intelligently infer implicit blood relationships: Introducing business knowledge graphs as prior knowledge to enhance the model's understanding of blood relationships; constructing and training an artificial intelligence model for inferring implicit blood relationships; using the generated explicit blood relationships and corresponding features as training samples, performing supervised learning training on the artificial intelligence model to enable the model to learn the mapping relationship between data features and blood relationships; (B4) Design a multi-level strategy for dynamic lineage updating: when the data processing task is created, possible lineage change paths are generated in advance based on the implicit lineage relationship inferred intelligently, and a lineage snapshot is automatically generated when the data processing task is retried or the parameters are dynamically adjusted.
6. The method for extracting dynamic blood relationships from data according to claim 1, characterized in that: After constructing the dynamic data lineage relationship, the method further includes iteratively designing a lineage relationship incremental update mechanism for tasks in the big data scheduling platform and including the following methods (C1) and / or (C2): (C1) Compare the SQL / script files of the new and old tasks and only update the lineage relationship of the changed parts; (C2) Using the task execution log of the big data scheduling platform to record dynamic change information during the data processing process.
7. A data dynamic blood relationship extraction device, characterized in that: It includes metadata collection unit, key event capture unit, blood relationship construction unit and blood relationship storage unit; The metadata collection unit is used to collect metadata about the data sources, data processing components, and data targets involved in each data processing task submitted to the platform in the big data scheduling platform, and obtain metadata including data source metadata, component metadata, and target metadata, wherein the data source metadata includes the type, structure information, and field information of the data source, the component metadata includes the type and parameter configuration information of the data processing component, and the target metadata includes the storage location and structure information of the data target; The key event capture unit is used to utilize the task scheduling mechanism of the big data scheduling platform to monitor the running status of the data processing task in real time during the task execution process to capture key events in the data processing process, wherein the key events include data reading events, data conversion events and data writing events; The blood relationship construction unit is communicatively connected to the metadata collection unit and the key event capture unit respectively, and is used to construct a data dynamic blood relationship based on the metadata and the event information of the key event, specifically including: for the data reading event, determining the corresponding data source, and establishing an input relationship between the data target and the data source; for the data conversion event, analyzing the field mapping relationship and data processing logic of the corresponding data in the conversion process, and constructing an intermediate blood relationship in the data processing process; for the data writing event, determining the destination of the corresponding data, and establishing an output relationship between the data source and the data target; associating and integrating the input relationship, the intermediate blood relationship and the output relationship to form a complete data dynamic blood relationship diagram; The blood relationship storage unit is communicatively connected to the blood relationship construction unit and is used to store the data dynamic blood relationship in a graph database for query and management.
8. A computer device, characterized in that: It includes a memory, a processor and a transceiver which are communicatively connected in sequence, wherein the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the data dynamic blood relationship extraction method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are run on a computer, the data dynamic blood relationship extraction method as described in any one of claims 1 to 6 is executed.
10. A computer program product comprising a computer program or instructions, characterized in that When executed by a computer, the computer program or the instruction implements the data dynamic blood relationship extraction method as described in any one of claims 1 to 6.