Data blood relationship identification method, system and device, medium and program product
By parsing configuration information, calculating metadata similarity, and performing large-scale model inference, a target lineage map is constructed, solving the accuracy and completeness issues of cross-platform lineage identification and enabling efficient tracking of complex data links.
Patent Information
- Application Number
- CN202510989610.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-31
AI Technical Summary
Existing bloodline identification technologies suffer from low accuracy and completeness in cross-platform environments, low data tracking efficiency, and difficulty in handling complex cross-platform data links.
By extracting entity nodes from various data platforms, parsing configuration information to obtain the first lineage relationship, calculating metadata similarity to obtain the second lineage relationship, and using large model reasoning to obtain the third lineage relationship, a target lineage graph is constructed to identify and update changed nodes.
It significantly improves the accuracy and completeness of cross-platform data tracking, solves the problem of broken cross-platform lineage, and enables efficient identification and tracking of complex data links.
Smart Images

Figure CN120873095A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data governance technology, and in particular to data lineage identification methods, data lineage identification systems, electronic devices, storage media, and computer program products. Background Technology
[0002] Existing lineage identification technologies perform well in data pipelines that strictly adhere to fixed processes. However, with the increasing number of business scenarios using data, data is frequently consumed by multiple different platforms after it is generated. In such cases, tracing data lineage requires traversing multiple platforms to determine the complete data path. Differences in data formats, processing logic, and storage methods between different platforms necessitate frequent data mapping and conversion, complicating cross-platform data links, leading to broken lineage records, low accuracy and completeness in lineage identification, and ultimately, low data tracing efficiency. Summary of the Invention
[0003] The main objective of this application is to provide a data lineage identification method, data lineage identification system, electronic device, storage medium, and computer program product, aiming to solve the technical problems of low accuracy and completeness of lineage identification and low data tracking efficiency when existing lineage identification technologies are used for lineage identification across platforms.
[0004] To achieve the above objectives, this application proposes a data lineage identification method, the method comprising:
[0005] Entity nodes are extracted from each data platform, and the configuration information of each data platform for processing the data corresponding to the entity nodes is parsed to obtain the first lineage relationship between the entity nodes.
[0006] For the first breakpoint entity nodes not covered by the first bloodline relationship, calculate the metadata similarity results between the first breakpoint entity nodes, and obtain the second bloodline relationship based on the metadata similarity results;
[0007] Cluster the second breakpoint entity nodes that are not covered by the second bloodline relationship to obtain multiple entity node sets, and obtain the third bloodline relationship between entity nodes in each entity node set through large model reasoning.
[0008] In one embodiment, the step of clustering the second breakpoint entity nodes not covered by the second bloodline relationship to obtain multiple entity node sets, and obtaining the third bloodline relationship between entity nodes in each entity node set through large model inference includes:
[0009] Extract the metadata features of the second breakpoint entity node and vectorize the metadata features;
[0010] Based on the vectorized target features, the second breakpoint entity nodes are grouped to obtain multiple breakpoint entity node sets.
[0011] Candidate entity node pairs are determined from the set of entity nodes at the same breakpoint. The metadata features of the first lineage relationship, the second lineage relationship, and the candidate entity node pairs are input into the large model to generate the third lineage relationship.
[0012] In one embodiment, the step of grouping the second breakpoint entity nodes based on the vectorized target features to obtain a set of multiple breakpoint entity nodes includes:
[0013] The second breakpoint entity nodes are clustered using a preset lightweight model to obtain multiple breakpoint entity node sets, wherein the cosine similarity between any nodes in the same breakpoint entity node set is greater than a preset association threshold.
[0014] In one embodiment, the step of calculating the metadata similarity results between the first breakpoint entity nodes includes:
[0015] Extract the metadata information of the table entity node and the file entity node from the first breakpoint entity node to form a table set and a file set respectively;
[0016] Calculate the similarity of metadata information in the table set and the file set to obtain the metadata similarity result of the first breakpoint entity node.
[0017] In one embodiment, the step of parsing the configuration information of each data platform for processing the data corresponding to the entity nodes to obtain the first lineage relationship between the entity nodes includes:
[0018] Analyze the scheduling dependencies in the configuration information of each data platform;
[0019] Based on the scheduling dependencies, extract the read / write operation identifiers between jobs and tables, the physical storage mapping relationship between fields and corresponding tables, the read / write operation identifiers between jobs and files, and the association identifiers between tables;
[0020] A first lineage relationship is generated based on the read / write operation identifiers between jobs and tables, the physical storage mapping relationship, the read / write operation identifiers between jobs and files, and the association identifier.
[0021] In one embodiment, the method further includes:
[0022] A target lineage map is constructed based on the first lineage relationship, the second lineage relationship, and the third lineage relationship between the entity nodes and the entity nodes.
[0023] Based on the target lineage map, the changed entity nodes are identified and the change type is marked;
[0024] Based on the change type, update the lineage relationship of the changed entity node in the target lineage graph.
[0025] Furthermore, to achieve the above objectives, this application also proposes a data lineage identification system, which includes:
[0026] The parsing module is used to extract entity nodes from various data platforms, parse the configuration information of each data platform for processing the data corresponding to the entity nodes, and obtain the first lineage relationship between entity nodes.
[0027] The calculation module is used to calculate the metadata similarity result between the first breakpoint entity nodes that are not covered by the first bloodline relationship, and to obtain the second bloodline relationship based on the metadata similarity result;
[0028] The reasoning module is used to cluster the second breakpoint entity nodes that are not covered by the second bloodline relationship to obtain multiple entity node sets, and to obtain the third bloodline relationship between entity nodes in each entity node set through large model reasoning.
[0029] In addition, to achieve the above objectives, this application also proposes an electronic device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the data lineage identification method as described above.
[0030] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the data lineage identification method described above.
[0031] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the data lineage identification method described above.
[0032] One or more technical solutions proposed in this application have at least the following technical effects:
[0033] The technical solution of this application extracts entity nodes from various data platforms, parses the configuration information of each data platform for processing the corresponding data of entity nodes, and obtains the first lineage relationship between entity nodes. However, due to the differences between different platforms, the first lineage relationship between some entity nodes cannot be directly parsed. Therefore, for the first breakpoint entity nodes not covered by the first lineage relationship, the metadata similarity results of the first breakpoint entity nodes in each data platform are calculated. Based on the metadata similarity results, the second lineage relationship is obtained, that is, the simple implicit lineage of missing metadata is supplemented by the metadata similarity results, and the local break caused by the differences between platforms is repaired. The second breakpoint entity nodes not covered by the second lineage relationship are clustered to obtain multiple entity node combinations. Then, the reasoning ability of the large model is used to obtain the third lineage relationship between entity nodes in each entity node set, and to handle complex cross-platform conversion scenarios. The entire data lineage identification method enables data scattered on heterogeneous platforms to be accurately and completely identified in terms of data lineage, which significantly improves data tracking efficiency.
[0034] Specifically, this application ensures the complete capture of fixed process lineage relationships across data platforms through configuration parsing, obtaining the first lineage relationship; automatically handles breakpoints caused by differences in data formats, processing logic, and storage methods between data platforms through metadata similarity calculation, obtaining the second lineage relationship; and finally, obtains the third lineage relationship through complex breakpoints requiring deep semantic transformation via large model reasoning. The three-layer lineage relationship progressively covers the identification of lineage relationships from explicit to implicit entity nodes, adapting to complex data links across platforms, and ultimately significantly improving the efficiency of data tracking and the accuracy and completeness of lineage identification in cross-platform environments. Attached Figure Description
[0035] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a flowchart illustrating an embodiment of the data lineage identification method of this application;
[0038] Figure 2 A schematic diagram of an entity node set provided for yet another embodiment of the data lineage identification method of this application;
[0039] Figure 3A schematic diagram of the target kinship map provided in another embodiment of the data kinship identification method of this application;
[0040] Figure 4 This is a schematic diagram of the module structure of the data lineage identification system of this application;
[0041] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the data lineage identification method of this application.
[0042] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0043] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0044] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0045] As business data usage scenarios continue to increase, data is frequently consumed by multiple different platforms after it is generated, ultimately leading to increasingly complex cross-platform data chains and making it difficult to establish a consistent cross-platform data lineage. Existing lineage identification technologies mainly rely on code, scripts, logs, and syntax tree parsing to standardize data development and processing workflows (such as Spark jobs, which define the actual execution process of all data processing in the application). However, existing methods have significant shortcomings in the following scenarios:
[0046] Existing solutions lack a unified metadata description standard when transferring data across platforms via files. For example, key information such as data source, processing logic, and field mapping rules are not recorded during file transfer, causing the lineage relationship to fail to be automatically linked at cross-platform boundaries. In other words, existing methods suffer from metadata gaps and lineage breaks in unstructured file exchange scenarios.
[0047] Existing solutions for pedigree identification across data platforms rely on platform-specific integration interfaces or code specifications, lacking universality. However, in practical applications, there are significant differences in data formats, processing logic, and metadata storage methods among different data processing platforms. The metadata description methods for entities such as files, database tables, and message queues are not uniform, making it difficult to construct a unified pedigree graph. Furthermore, the processing logic of file conversion tools (such as ETL tools and scripts) is not explicitly recorded, making it impossible to trace field-level mapping relationships. In other words, existing methods face difficulties in pedigree identification across data platforms.
[0048] In this application, entity nodes are extracted from various data platforms, and the configuration information for processing the corresponding data of each entity node is parsed to obtain the first lineage relationship between entity nodes. However, due to the differences between different platforms, the first lineage relationship between some entity nodes cannot be directly parsed. Therefore, for the first breakpoint entity nodes not covered by the first lineage relationship, the metadata similarity results of the first breakpoint entity nodes in each data platform are calculated. Based on the metadata similarity results, the second lineage relationship is obtained, that is, the simple implicit lineage of missing metadata is supplemented by the metadata similarity results, and the local break caused by the differences between platforms is repaired. The second breakpoint entity nodes not covered by the second lineage relationship are clustered to obtain multiple entity node combinations. Then, the reasoning ability of the large model is used to obtain the third lineage relationship between entity nodes in each entity node set, handling complex cross-platform conversion scenarios. The entire data lineage identification method enables accurate and complete data lineage identification of data scattered on heterogeneous platforms, significantly improving data tracking efficiency.
[0049] It should be noted that the executing entity of this embodiment can be a data lineage identification system, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or processor capable of performing the above functions. The following description uses a data lineage identification system as an example to illustrate this embodiment and the subsequent embodiments.
[0050] Based on this, embodiments of this application provide a data lineage identification method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the data lineage identification method of this application.
[0051] In this embodiment, the data lineage identification method includes steps S10 to S30:
[0052] Step S10: Extract entity nodes from each data platform, parse the configuration information of each data platform for processing the data corresponding to the entity nodes, and obtain the first lineage relationship between entity nodes.
[0053] It should be noted that entity nodes are extracted from various data platforms. Entity nodes can include tables, jobs, fields, and files. Jobs can be job code or job scripts. Data platforms can include data warehouses, TwinTransfer, application libraries, etc.
[0054] The configuration information for processing data corresponding to entity nodes in each data platform is analyzed to obtain the first lineage relationship between entity nodes. This first lineage relationship can include read / write relationships between jobs and tables, ownership relationships between fields and tables, read / write relationships between jobs and files, and influence relationships between tables. It should also be noted that the first lineage relationship between entity nodes can be between different entity nodes within the same data platform, or between different entity nodes on different data platforms.
[0055] For example, in a heterogeneous platform environment in a real-world scenario, the downstream of the source system's business tables may involve various scenarios, such as direct extraction into the data warehouse, transmission via dual-platform systems, or batch processing via DGP (Data Governance Platform). However, the data transmission methods differ significantly across these scenarios, making it impossible to establish a unified approach to connect upstream and downstream relationships. Therefore, it is necessary to address specific situations by employing methods such as parsing and scheduling jobs and reading task flow configurations. Currently, the downstream of the source system's business tables mainly involves the following three exchange scenarios: ① Direct extraction into the data warehouse: identified through EX (Extract Job) and LD (Load Job). ② Source system business tables passing through transmission platforms such as dual-platform systems or exchange platforms: identified through FR (File Receive Job), TR (Transformation Job), and FS (File Send Job). ③ Processing on the DGP platform: based on DGP task flow configuration information.
[0056] In one implementation, step S10 includes:
[0057] Analyze the scheduling dependencies in the configuration information of each data platform;
[0058] Extract read / write operation identifiers between jobs and tables, physical storage mapping relationships between fields and corresponding tables, read / write operation identifiers between jobs and files, and association identifiers between tables based on scheduling dependencies;
[0059] The first lineage relationship is generated based on the read / write operation identifiers between jobs and tables, the physical storage mapping relationship, the read / write operation identifiers between jobs and files, and the association identifiers.
[0060] It should be noted that the scheduling dependencies (which can be the logical order of job execution) are extracted from the configuration information. This leads to the extraction of identifiers for read / write operations between jobs and tables (clarifying the job's consumption or output actions on the data table), physical storage mappings between fields and tables (revealing the actual location of fields in the storage medium), read / write operation identifiers between jobs and files (recording file-level data flow paths), and association identifiers between tables (table inheritance relationships or data transfer relationships within tables). Based on these four types of identifiers, the data lineage identification system obtains the first lineage relationship, which includes job-table read / write relationships, field-table ownership relationships, job-file transfer relationships, and table-table influence relationships.
[0061] In this embodiment, by parsing the scheduling dependencies in the configuration information of each data platform, the read and write operation identifiers between jobs and tables, the physical storage mapping relationship between fields and tables, the read and write operation identifiers between jobs and files, and the association identifiers between tables are extracted. The explicit lineage relationship between cross-platform entities is directly established, which is the first lineage relationship. This solves the problem of missing metadata leading to lineage breakage in the unstructured file exchange scenario of existing methods, and significantly improves the efficiency of cross-platform data tracking.
[0062] Step S20: For the first breakpoint entity nodes not covered by the first bloodline relationship, calculate the metadata similarity results between the first breakpoint entity nodes, and obtain the second bloodline relationship based on the metadata similarity results;
[0063] It should be noted that for the first breakpoint entity node not covered by the first lineage relationship (such as in the scenario of unstructured file transfer), the second lineage relationship is obtained by calculating the metadata similarity of the first breakpoint entity node. The metadata of the first breakpoint entity node includes file size and encoding format, field names, table structure definition, storage path, etc. The second lineage relationship can be the format mapping relationship between tables and files, such as the storage path of Hive tables mapped to HDFS files, cross-platform transmission path relationship, and content equivalence relationship of unstructured files, such as CSV files and JSON files under different paths containing the same business data.
[0064] In one embodiment, step S20 includes:
[0065] Extract the metadata information of the table entity node and the file entity node from the first breakpoint entity node, and form a table set and a file set respectively;
[0066] Calculate the similarity of metadata information in the table set and the file set to obtain the metadata similarity result of the first breakpoint entity node.
[0067] It should be noted that after identifying the first breakpoint entity node, metadata information needs to be extracted from both the table entity node and the file entity node to form a structured set. For table entities, the metadata information includes field names, data types, data volume, and partition structure; for file entities, it covers file path mode, storage format, number of records, and compression method. Based on the table set and the file set, a cross-platform metadata similarity calculation is performed to obtain the metadata similarity result of the first breakpoint entity node. For example, Jaccard similarity can be used, with values ranging from 0 to 1; a higher value indicates greater similarity between the two sets. The result is calculated by dividing the number of elements in the intersection of the sets by the number of elements in the union of the sets.
[0068] For example, after processing by the data lineage identification method based on parsing configuration, the lineage relationship between files and downstream systems can be largely established. However, because the file transfer system does not record the file's source information, the lineage relationship between the table entity nodes and file entity nodes in the source system remains disconnected. To address this missing lineage, since file transfer does not change the data information, the fields of the table entity nodes and file entity nodes in the source system are either inclusive or equal. Inclusion means that the fields in the file entity node are a subset of the fields in the table entity node; equality means that the fields in the file entity node are completely identical to the fields in the table entity node. If both inclusion and equality exist, the one with the highest similarity is selected. Identifying this lineage relationship can be measured by calculating the metadata similarity between entity nodes. Specific steps include: extracting the field information from the tables in the source system and the files in the exchange platform one by one, forming a structure like... Figure 2 The set shown is in the form of, where, in Figure 2 This includes set A (e.g., a set of tables) and set B (e.g., a set of files). Set A contains multiple tables such as T1, T2, T3, and T4, each containing many fields. For example, T1 contains C1, C2, C3, C4, and C5; T2 contains C2, C6, C3, C4, and C5; T3 contains C2, C7, C3, C4, and C5; and T4 contains C2, C6, C3, C4, and C5. Set A also contains multiple files such as T6, T8, T9, and T7, each containing many fields. For example, T6 contains C1, C2, C3, C4, and C5; T8 contains C2, C6, C3, C4, and C5; T9 contains C2, C6, C3, C4, and C5; and T7 contains C2, C7, C3, C4, and C5.
[0069] Then calculate the similarity for each pair of files, find the file entity nodes with broken lineages in batches, and find the source system table entity nodes that may be related to these file entity nodes based on the calculated metadata similarity.
[0070] In this embodiment, by extracting the metadata information of tables and files from the first breakpoint entity node, a set is obtained and the similarity of the metadata information in the set is calculated. This solves the problem of metadata isolation caused by differences in data formats across heterogeneous platforms. When data is transmitted across platforms and the lineage is broken due to loss of source information or differences in storage methods, the field set of the table entity and the field set of the file entity are matched for similarity. Thus, without the need for manual mapping, entity nodes that appear heterogeneous but have the same data source can be identified, and the lineage link broken due to platform boundaries can be reconstructed, significantly improving the efficiency of cross-platform data tracking.
[0071] Step S30: Cluster the second breakpoint entity nodes that are not covered by the second bloodline relationship to obtain multiple entity node sets, and obtain the third bloodline relationship between entity nodes in each entity node set through large model reasoning.
[0072] It should be noted that for second-breakpoint entity nodes not covered by the second lineage relationship (such as cross-platform field renaming), a set of entity nodes with potential semantic relationships is identified among the uncovered second-breakpoint entity nodes. This set of entity nodes and its context are input into the large model, and a third lineage relationship is inferred and output. This resolves the lineage path breakage caused by complex situations such as data format conversion, different processing logic, and different storage methods across heterogeneous data platforms (i.e., cross-data platforms). The third lineage relationship can be a cross-platform field semantic mapping relationship (e.g., a source system field mapped to a data warehouse field after calculation), a path change relationship (e.g., field A in file path 1 in the storage platform is redirected to file path 2 due to a path change), or a processing relationship (e.g., field B obtained by de-identifying field A).
[0073] In one embodiment, step S30 includes steps A10 to A30:
[0074] Step A10: Extract the metadata features of the second breakpoint entity node and vectorize the metadata features;
[0075] It should be noted that for the second breakpoint entity nodes not covered by the second lineage relationship, their metadata features are extracted and vectorized. Metadata features refer to structured information describing the attributes of entity nodes. For table entities, this includes elements such as the semantic structure of field names, data value distribution patterns, and storage strategy rules; for file entities, it includes parameters such as content feature encoding, path pattern features, and format compression tags. Vectorization refers to converting these heterogeneous metadata features into numerical vectors of a unified dimension. For example, the field name "cust_id" is mapped to a 128-dimensional semantic vector, and the file path rule " / data / dt=${date}" is encoded into a path variable matrix.
[0076] Step A20: Based on the vectorized target features, the second breakpoint entity nodes are grouped to obtain multiple breakpoint entity node sets;
[0077] It should be noted that, based on the vectorized features generated in step A10, the entity nodes at the second breakpoint are clustered. By calculating distances in the vector space (such as cosine similarity or Euclidean distance), entity nodes with similar feature vectors are identified and grouped into clusters. For example, file entities with similar path feature vectors or table fields with similar semantic vectors are grouped together. Entity grouping essentially involves automatically discovering potentially related sets of entities through an algorithm. The grouping results are output as multiple sets of breakpoint entity nodes, each set representing a group of candidate entities with related relationships. For instance, in one breakpoint entity node set, all entities are related to the customer.
[0078] Step A30: Determine candidate entity node pairs from the set of entity nodes at the same breakpoint, and input the metadata features of the first lineage relationship, the second lineage relationship, and the candidate entity node pairs into the large model to generate the third lineage relationship.
[0079] It should be noted that candidate entity node pairs are selected from the set of breakpoint entity nodes within the same group. For example, the table entity and file entity with the closest distance within the group are selected to form a pair. The vectorized features of the candidate pairs, the topological path of the first lineage relationship, and the similarity results of the second lineage relationship are used as context and input into the large language model for inference. The large model generates the third lineage relationship based on its pre-trained business semantic understanding capabilities and prompting rules.
[0080] For example, the effectiveness of boundary lineage identification methods based on pairwise metadata similarity calculations relies on clear upstream and downstream platform identifiers, clear missing boundary types, and constraints such as a small range of similar assets. However, in complex scenarios involving missing lineages across multiple layers (such as data flowing through multiple heterogeneous systems, ambiguous intermediate node types, or large missing path spans), traditional methods face bottlenecks due to their difficulty in capturing cross-layer semantic relationships and dynamic feature changes. To address this, the technical solution in this application overcomes the limitations of single similarity calculations through strategies such as feature enhancement, group clustering, and global semantic association of large models, achieving automatic completion of lineages across multiple layers. Specifically, metadata features (entity node descriptions, labels, fields, etc.) of the second breakpoint entity node are extracted, and a unified feature set is constructed through data cleaning, text segmentation, and feature encoding. Text fields such as entity node descriptions, labels, and fields are converted into dense vectors using methods such as Word2Vec and FastText, and vectorization of all metadata features is completed through operations such as vector weighted concatenation and dimensionality reduction. Lightweight models, such as those incorporating the HDBSCAN clustering algorithm, automatically identify potential clusters of relationships between entity nodes, i.e., sets of entity nodes at breakpoints. Entity nodes within the same cluster may have lineage relationships. Entity nodes within the same cluster are then paired, and cosine similarity is calculated. Based on similarity priority, candidate entity node pairs are initially generated. Based on the candidate entity node pairs in each cluster, combined with existing lineage relationships (i.e., first and second lineage relationships) and information about the entity nodes at the breakpoints—i.e., the metadata features of the candidate entity node pairs—and leveraging the understanding capabilities of the larger model, inputting validation rule prompts, automatically completes and validates the broken lineage relationships to obtain the third lineage relationship.
[0081] In this embodiment, a large model is used for semantic reasoning to cover lineage relationships that cannot be obtained by parsing configuration, so that broken paths caused by format differences or logical changes between heterogeneous data platforms can be reconstructed, thereby improving the accuracy and completeness of cross-data platform lineage identification.
[0082] In another implementation, step A20 includes:
[0083] Clustering of the second breakpoint entity nodes using a preset lightweight model yields multiple breakpoint entity node sets, where the cosine similarity between any nodes within the same breakpoint entity node set is greater than a preset association threshold.
[0084] It should be noted that the second breakpoint entity nodes are clustered using a pre-set lightweight model, such as a clustering model. Here, the lightweight model refers to a clustering algorithm with high computational efficiency and low resource consumption (such as HDBSCAN or MiniBatch K-Means) that aims to automatically identify and aggregate a set of entity nodes with potential lineage relationships.
[0085] The second breakpoint entity node refers to an isolated data entity not covered by the first or second lineage relationship. Multiple breakpoint entity node sets are obtained by calculating cosine similarity based on the vectorized metadata features of the entity nodes (such as field semantic embedding vectors, file path encoding, and table structure signatures). The clustering process essentially merges entities with adjacent feature vectors into the same logical group, automatically dividing the system into multiple breakpoint entity node sets through algorithms.
[0086] Specifically, the cosine similarity between any nodes within the same set of entity nodes at the same breakpoint is greater than a preset association threshold. This preset association threshold can be set to a semantic vector similarity of 0.7. Cosine similarity determines the degree of semantic and structural closeness between pairs of entity nodes. When entities cluster into high-density groups in the vector space and their mutual similarity meets the threshold, they can be considered to have a latent kinship basis.
[0087] In this embodiment, the second breakpoint entity nodes are clustered and grouped using a preset lightweight model. Multiple breakpoint entity node sets are obtained by filtering based on the cosine similarity of node vectorization features. Through clustering, the set of entity nodes with potential lineage relationships is automatically identified in the vector space. This breaks through the dependence of traditional methods on the uniformity of metadata specifications, and aggregates entity nodes such as tables and files that are discrete on different platforms into related sets, ensuring that entity nodes within the same group have lineage relationships, and improving the integrity and efficiency of data lineage tracing in complex heterogeneous environments.
[0088] This embodiment provides a data lineage identification method. Entity nodes are extracted from various data platforms, and the configuration information for processing the corresponding data of each entity node on each platform is parsed to obtain the first lineage relationship between entity nodes. However, due to differences between different platforms, the first lineage relationship between some entity nodes cannot be directly parsed. Therefore, for first breakpoint entity nodes not covered by the first lineage relationship, the metadata similarity results of the first breakpoint entity nodes in each data platform are calculated. Based on the metadata similarity results, the second lineage relationship is obtained, that is, the simple implicit lineage missing in metadata is supplemented by the metadata similarity results, repairing the local breaks caused by differences between platforms. The second breakpoint entity nodes not covered by the second lineage relationship are clustered to obtain multiple entity node combinations. Then, the reasoning capability of a large model is used to obtain the third lineage relationship between entity nodes in each entity node set, handling complex cross-platform conversion scenarios. The entire data lineage identification method enables accurate and complete data lineage identification of data scattered across heterogeneous platforms, significantly improving data tracking efficiency. Specifically, this application ensures the complete capture of fixed process lineage relationships across data platforms through configuration parsing, obtaining the first lineage relationship; automatically handles breakpoints caused by differences in data formats, processing logic, and storage methods between data platforms through metadata similarity calculation, obtaining the second lineage relationship; and finally, obtains the third lineage relationship through complex breakpoints requiring deep semantic transformation via large model reasoning. The three-layer lineage relationship progressively covers the identification of lineage relationships from explicit to implicit entity nodes, adapting to complex data links across platforms, and ultimately significantly improving the efficiency of data tracking and the accuracy and completeness of lineage identification in cross-platform environments.
[0089] Based on the above embodiments of this application, in another embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter.
[0090] Existing solutions lack the ability to adaptively track dynamically changing lineage relationships, making it difficult to meet the needs of complex business scenarios. In actual business operations, data flow paths and processing logic may change frequently (such as file path migration, field renaming, and ETL task updates), while existing technologies rely on static lineage identification configurations, requiring manual maintenance or periodic metadata scanning, and cannot perceive dynamic changes in real time; version management is lacking: it is impossible to record historical versions of lineage relationships, making it difficult to trace the scope of impact of data changes.
[0091] In one embodiment of this application, the data lineage identification method includes:
[0092] Based on the first, second, and third lineage relationships between entity nodes, the target lineage map is constructed.
[0093] Based on the target lineage map, the changed entity nodes are identified and the change type is marked;
[0094] Based on the change type, update the lineage relationship of the changed entity node in the target lineage graph.
[0095] It should be noted that a complete target lineage graph is constructed based on entity nodes and their first, second, and third lineage relationships. As a data tracking tool, the target lineage graph records the entire link of cross-platform data flow through the association relationship between nodes and edges.
[0096] For example, the target lineage graph stores (entity node 1, relationship, entity node 2) and (entity node, entity node attribute, attribute value). Data integration, disambiguation, and processing operations are performed on the target lineage graph to establish connections between heterogeneous data platforms.
[0097] When actual business operations change, the data lineage identification system automatically identifies the affected changed entity nodes and marks the change type (including structural changes such as field additions and deletions, logical changes such as ETL task updates, path changes such as storage location relocation, etc.). For example, this identification process is achieved by real-time monitoring of metadata change events (such as DDL statement logs and job scheduling records).
[0098] Based on the type of change marked, the data lineage identification system updates the affected lineage relationships in the target lineage graph in a targeted manner: for example, when a field renaming is detected, the connection edges of the old field node are automatically removed, and the historical lineage relationship is seamlessly transferred to the new field node; if a file path migration occurs, the read and write relationship between the file entity and the downstream job is dynamically updated through path mapping rules to realize the historical version management of lineage relationships.
[0099] For example, by comparing the lineage relationships of entity nodes on day T and day T-1, the entity node initiating the change is dynamically identified, and the change type of the changed entity node is marked. Then, by analyzing various change types and combining them with existing lineage relationships, the entity nodes and lineage relationships that need to be updated are detected. The changed entity nodes and lineage relationships are then passed to the downstream links to complete the update of the entire data link.
[0100] In this embodiment, when data format conversion, platform migration, or processing logic update occurs in a business scenario, the changed entity nodes are automatically identified and the change type is marked based on the target lineage graph. Then, the lineage relationship is updated according to the change type. This ensures that the data mapping relationship between heterogeneous platforms is always synchronized with the current business state, thereby eliminating the lag of traditional static lineage relationships that require manual maintenance. It ensures the continuous traceability of end-to-end data paths in complex cross-platform environments, solves the problem of lack of real-time cross-platform lineage due to frequent changes in data structure, and improves data governance efficiency.
[0101] For example, to help understand the target kinship map obtained by combining this embodiment with the above embodiments, please refer to... Figure 3 , Figure 3 A schematic diagram of a target lineage graph is provided. Specifically, 100, 200, 300, and 400 represent entity nodes in each data platform, and the connecting lines represent the relationships between the entity nodes, such as ownership relationships, influence relationships, and read / write relationships. Thus, the target lineage graph is formed through the data structure of nodes (entity nodes) and edges (connections between entity nodes).
[0102] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the data lineage identification method of this application. Any simple variations based on this technical concept, such as the interaction and combination of various embodiments, are all within the protection scope of this application.
[0103] This application also provides a data lineage identification system; please refer to [reference needed]. Figure 4 The data lineage identification system includes:
[0104] The parsing module 10 is used to extract entity nodes from each data platform, parse the configuration information of each data platform for processing the data corresponding to the entity nodes, and obtain the first lineage relationship between entity nodes.
[0105] The calculation module 20 is used to calculate the metadata similarity result between the first breakpoint entity nodes that are not covered by the first blood relationship, and to obtain the second blood relationship based on the metadata similarity result;
[0106] The reasoning module 30 is used to cluster the second breakpoint entity nodes that are not covered by the second bloodline relationship to obtain multiple entity node sets, and to obtain the third bloodline relationship between entity nodes in each entity node set through large model reasoning.
[0107] The reasoning module 30 is also used to extract the metadata features of the second breakpoint entity node and vectorize the metadata features;
[0108] Based on the vectorized target features, the second breakpoint entity nodes are grouped to obtain multiple breakpoint entity node sets.
[0109] Candidate entity node pairs are determined from the set of entity nodes at the same breakpoint. The metadata features of the first lineage relationship, the second lineage relationship, and the candidate entity node pairs are input into the large model to generate the third lineage relationship.
[0110] The reasoning module 30 is also used to cluster the second breakpoint entity nodes using a preset lightweight model to obtain multiple breakpoint entity node sets, wherein the cosine similarity between any nodes in the same breakpoint entity node set is greater than a preset association threshold.
[0111] The calculation module 20 is also used to extract the metadata information of the table entity node and the file entity node in the first breakpoint entity node, and form a table set and a file set respectively;
[0112] Calculate the similarity of metadata information in the table set and the file set to obtain the metadata similarity result of the first breakpoint entity node.
[0113] The parsing module 10 is also used to parse the scheduling dependencies in the configuration information of each data platform;
[0114] Based on the scheduling dependencies, extract the read / write operation identifiers between jobs and tables, the physical storage mapping relationship between fields and corresponding tables, the read / write operation identifiers between jobs and files, and the association identifiers between tables;
[0115] A first lineage relationship is generated based on the read / write operation identifiers between jobs and tables, the physical storage mapping relationship, the read / write operation identifiers between jobs and files, and the association identifier.
[0116] The data lineage identification system further includes an update module, used to construct a target lineage map based on the first lineage relationship, the second lineage relationship, and the third lineage relationship between the entity node and the entity node;
[0117] Based on the target lineage map, the changed entity nodes are identified and the change type is marked;
[0118] Based on the change type, update the lineage relationship of the changed entity node in the target lineage graph.
[0119] The data kinship identification system provided in this application, employing the data kinship identification method described in the above embodiments, can solve the technical problems of low accuracy and completeness of kinship identification and low data tracking efficiency in existing kinship identification technologies when performing kinship identification across platforms. Compared with the prior art, the beneficial effects of the data kinship identification system provided in this application are the same as those of the data kinship identification method provided in the above embodiments, and other technical features of the data kinship identification system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0120] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the data lineage identification method described in the first embodiment above.
[0121] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0122] like Figure 5 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0123] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0124] The electronic device provided in this application, employing the data lineage identification method described in the above embodiments, can solve the technical problems of low accuracy and completeness of lineage identification and low data tracking efficiency in existing lineage identification technologies when performing lineage identification across platforms. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the data lineage identification method provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0125] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0126] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0127] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the data lineage identification method in the above embodiments.
[0128] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0129] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0130] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by an electronic device, the electronic device causes the following: extracts entity nodes from various data platforms, parses the configuration information of each data platform for processing the data corresponding to the entity nodes, and obtains the first lineage relationship between the entity nodes; calculates the metadata similarity result between the first breakpoint entity nodes not covered by the first lineage relationship, and obtains the second lineage relationship based on the metadata similarity result; clusters the second breakpoint entity nodes not covered by the second lineage relationship to obtain multiple entity node sets, and obtains the third lineage relationship between the entity nodes in each entity node set through large model inference.
[0131] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0133] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0134] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described data lineage identification method. This solves the technical problems of low accuracy and completeness of lineage identification and low data tracking efficiency in existing lineage identification technologies when performing lineage identification across platforms. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the data lineage identification method provided in the above embodiments, and will not be repeated here.
[0135] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the data lineage identification method described above.
[0136] The computer program product provided in this application can solve the technical problems of low accuracy and completeness of bloodline identification and low data tracking efficiency when existing bloodline identification technologies are used for cross-platform bloodline identification. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the data bloodline identification method provided in the above embodiments, and will not be repeated here.
[0137] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for identifying bloodlines based on data, characterized in that, The data lineage identification method includes: Entity nodes are extracted from each data platform, and the configuration information of each data platform for processing the data corresponding to the entity nodes is parsed to obtain the first lineage relationship between the entity nodes. For the first breakpoint entity nodes not covered by the first bloodline relationship, calculate the metadata similarity results between the first breakpoint entity nodes, and obtain the second bloodline relationship based on the metadata similarity results; Cluster the second breakpoint entity nodes that are not covered by the second bloodline relationship to obtain multiple entity node sets, and obtain the third bloodline relationship between entity nodes in each entity node set through large model reasoning.
2. The data lineage identification method as described in claim 1, characterized in that, The step of clustering the second breakpoint entity nodes not covered by the second bloodline relationship to obtain multiple entity node sets, and obtaining the third bloodline relationship between entity nodes in each entity node set through large model inference includes: Extract the metadata features of the second breakpoint entity node and vectorize the metadata features; Based on the vectorized target features, the second breakpoint entity nodes are grouped to obtain multiple breakpoint entity node sets. Candidate entity node pairs are determined from the set of entity nodes at the same breakpoint. The metadata features of the first lineage relationship, the second lineage relationship, and the candidate entity node pairs are input into the large model to generate the third lineage relationship.
3. The data lineage identification method as described in claim 2, characterized in that, The step of grouping the second breakpoint entity nodes based on the vectorized target features to obtain multiple breakpoint entity node sets includes: The second breakpoint entity nodes are clustered using a preset lightweight model to obtain multiple breakpoint entity node sets, wherein the cosine similarity between any nodes in the same breakpoint entity node set is greater than a preset association threshold.
4. The data lineage identification method as described in claim 1, characterized in that, The step of calculating the metadata similarity results between the first breakpoint entity nodes includes: Extract the metadata information of the table entity node and the file entity node from the first breakpoint entity node to form a table set and a file set respectively; Calculate the similarity of metadata information in the table set and the file set to obtain the metadata similarity result of the first breakpoint entity node.
5. The data lineage identification method as described in claim 1, characterized in that, The step of parsing the configuration information of each data platform for processing the data corresponding to the entity nodes to obtain the first lineage relationship between entity nodes includes: Analyze the scheduling dependencies in the configuration information of each data platform; Based on the scheduling dependencies, extract the read / write operation identifiers between jobs and tables, the physical storage mapping relationship between fields and corresponding tables, the read / write operation identifiers between jobs and files, and the association identifiers between tables; A first lineage relationship is generated based on the read / write operation identifiers between jobs and tables, the physical storage mapping relationship, the read / write operation identifiers between jobs and files, and the association identifier.
6. The data lineage identification method as described in claim 1, characterized in that, The method further includes: A target lineage map is constructed based on the first lineage relationship, the second lineage relationship, and the third lineage relationship between the entity nodes and the entity nodes. Based on the target lineage map, the changed entity nodes are identified and the change type is marked; Based on the change type, update the lineage relationship of the changed entity node in the target lineage graph.
7. A data lineage identification system, characterized in that, The data lineage identification system includes: The parsing module is used to extract entity nodes from various data platforms, parse the configuration information of each data platform for processing the data corresponding to the entity nodes, and obtain the first lineage relationship between entity nodes. The calculation module is used to calculate the metadata similarity result between the first breakpoint entity nodes that are not covered by the first bloodline relationship, and to obtain the second bloodline relationship based on the metadata similarity result; The reasoning module is used to cluster the second breakpoint entity nodes that are not covered by the second bloodline relationship to obtain multiple entity node sets, and to obtain the third bloodline relationship between entity nodes in each entity node set through large model reasoning.
8. An electronic device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the data lineage identification method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the data lineage identification method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the data lineage identification method as described in any one of claims 1 to 6.
Citation Information
Cited By
Methods and systems for constructing end-to-end data lineages across heterogeneous platforms
CN122412396A