Data blood relationship analysis visualization method and device, equipment and medium
By obtaining metadata in the financial insurance field and building a blood relationship model, using the graph algorithm to generate an accurate blood relationship map, the complexity and dependence misunderstanding problems in blood relationship map technology are solved, and the transparency and stability of the project are improved.
Patent Information
- Application Number
- CN202510450407.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-29
AI Technical Summary
The existing blood relationship map technology has complexity and misunderstandings in the financial and insurance field, affecting the application and optimization of projects.
By obtaining the metadata of the target data from multiple data sources, determining the metadata type and attributes, building a blood relationship model, and using the preset graph algorithm to perform data blood relationship analysis, generating and displaying an accurate blood relationship diagram.
It improves the transparency and understanding of the project, reduces development costs, improves the overall quality and stability of the project, and guides code decoupling and architectural optimization.
Smart Images

Figure CN120386893A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data technology and is applied to the processing scenarios of fintech and healthcare management services. In particular, it relates to a method, apparatus, device, and medium for visualizing data lineage analysis. Background Art
[0002] The Blood Graph, as a charting tool widely used in software development, is mainly used to display the dependency relationships between classes in code. Especially in JavaScript projects, it is of great significance for understanding the interactions between modules and analyzing the project architecture. In addition, the concept of the Blood Graph has been extended to data management to describe the flow and transformation relationships of data throughout its life cycle, namely data lineage, data origin, or data pedigree. This kind of chart helps developers better understand and manage the data flow by intuitively showing the upstream and downstream sources and destinations of the data.
[0003] In the field of finance and insurance, the application of the Blood Graph also has far-reaching significance. Financial institutions and insurance companies process a vast amount of data, including customer information, transaction records, risk assessment reports, etc. These data flow and transform in the business processes, forming a complex data network. The Blood Graph has become a key tool for sorting out and managing this data here. It can help financial institutions and insurance companies trace the source of the data, understand the data processing process, and master the final destination of the data. This is crucial for ensuring the accuracy, integrity, and compliance of the data, especially in meeting regulatory requirements, conducting risk assessments, and providing customer services.
[0004] The advantage of the Blood Graph lies in its visualization feature, which can significantly improve the understandability and maintainability of the project. In the field of finance and insurance, this means that more complex data flows and business processes can become more intuitive and understandable, facilitating communication and collaboration between business personnel and technical personnel. At the same time, the Blood Graph also helps with code decoupling, reducing direct dependencies between modules, and thus improving the testability of the code. However, the application of the Blood Graph in the field of finance and insurance also comes with a series of challenges. Overemphasis on the Blood Graph may lead to over-complication of project design and increase unnecessary development costs. In addition, inaccurate or misleading Blood Graphs may cause developers to misunderstand the relationships between modules, thereby affecting the overall quality and stability of the project.
[0005] In summary, the core problems faced by the current Blood Graph technology lie in its complexity and the possible dependency misunderstandings it may cause. These problems limit the wide application and in-depth optimization of the Blood Graph in actual projects. Summary of the Invention
[0006] The purpose of the embodiments of the present application is to propose a method, device, equipment and medium for visualizing data lineage analysis, so as to solve the complexity problems faced by existing lineage graph technologies and the possible problems of dependency misunderstandings caused thereby.
[0007] In a first aspect, a method for visualizing data lineage analysis is provided, and the following technical solution is adopted:
[0008] Obtain the metadata of the target data from multiple data sources, and determine the metadata type and metadata attributes of the metadata; obtain the lineage relationship model of the target data; based on the lineage relationship model, metadata type, and metadata attributes, use a preset graph algorithm to perform data lineage analysis on the target data to obtain the lineage analysis result of the target data; based on the lineage analysis result, establish a lineage relationship graph of the target data and display the lineage relationship graph on a preset interactive interface.
[0009] In a second aspect, a device for visualizing data lineage analysis is provided, and the following technical solution is adopted:
[0010] A first acquisition module, configured to obtain the metadata of the target data from multiple data sources, and determine the metadata type and metadata attributes of the metadata;
[0011] A second acquisition module, configured to obtain the lineage relationship model of the target data;
[0012] An analysis module, configured to perform data lineage analysis on the target data based on the lineage relationship model, metadata type, and metadata attributes, using a preset graph algorithm, to obtain the lineage analysis result of the target data;
[0013] A building module, configured to establish a lineage relationship graph of the target data based on the lineage analysis result, and display the lineage relationship graph on a preset interactive interface.
[0014] In a third aspect, a computer device is provided, including a memory and a processor. A computer-readable instruction is stored in the memory, and when the processor executes the computer-readable instruction, the steps of the data lineage analysis visualization method as described above are implemented.
[0015] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer-readable instruction, and the computer-readable instruction can be executed by at least one processor, so that at least one processor executes the steps of the data lineage analysis visualization method as described above.
[0016] In the solution implemented by the above data lineage analysis visualization method, device, equipment, and medium, by accurately obtaining the target data metadata of multiple data sources and clarifying their types and attributes, a solid foundation is provided for the accurate analysis of data lineage. By introducing a preset graph algorithm and combining it with the lineage relationship model of the target data, in-depth analysis of data lineage is achieved, effectively solving the problems of complexity and dependency misunderstanding in lineage graph technology. This solution can automatically generate an accurate lineage graph and visually display it in a preset interactive interface, greatly improving the transparency and understandability of the project. This not only helps developers efficiently understand and manage data streams, but also effectively guides code decoupling and architecture optimization, reducing development costs and improving the overall quality and stability of the project. Brief Description of the Drawings
[0017] To more clearly illustrate the solution in this application, the following will briefly introduce the drawings required for the description of the embodiments of this application. Obviously, the drawings below are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 is an exemplary system architecture diagram to which this application can be applied;
[0019] Figure 2 is a flowchart of a data lineage analysis visualization method provided by this application;
[0020] Figure 3 is a structural diagram of a data lineage analysis visualization device provided by this application;
[0021] Figure 4 is a structural diagram of a computer device provided by this application. Detailed Embodiments
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the description of the embodiments of this application in this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the description and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the description and claims of this application or the above drawings are used to distinguish different objects and are not used to describe a specific order.
[0023] Reference to "embodiments" in this document means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment each time, nor are they independent or alternative embodiments mutually exclusive of other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0024] To enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0025] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0026] A user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.
[0027] The terminal device 101 may be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 may also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop portable computer, a desktop computer, etc.
[0028] The server 103 may be a server providing various services, such as a background server that provides support for the pages displayed on the terminal device 101.
[0029] It should be noted that the data lineage analysis visualization method provided by the embodiments of the present application is generally executed by the server. Correspondingly, the data lineage analysis visualization device is generally disposed in the server.
[0030] It should be understood, Figure 1The numbers of the terminal devices, networks, and servers in [it] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0031] Continuing to refer to Figure 2 , a flowchart of an embodiment of the data lineage analysis visualization method according to the present application is shown. The data lineage analysis visualization method includes the following steps:
[0032] Step S201 obtains the metadata of the target data from multiple data sources, and determines the metadata type and metadata attributes of the metadata.
[0033] Among them, the multiple data sources refer to various types of data sets obtained from different channels or systems. These data sources can include databases, file servers, API interfaces, etc. For example, a project may simultaneously obtain data from user behavior logs, transaction records, and external APIs, and these data sources together constitute the "multiple data sources".
[0034] Among them, the target data refers to the data set that needs to be concerned about or operated on in a specific analysis or processing task. The target data can be filtered and integrated from multiple data sources to meet specific business requirements or analysis purposes. For example, in financial insurance analysis, the insurance claim records of customers are the target data, which are used to analyze customers' claim behaviors, claim frequencies, claim amounts, and claim reasons, etc.
[0035] Among them, the metadata refers to the data that describes the target data, including information such as the source, format, structure, and meaning of the target data. Metadata is the key to understanding and managing data. For example, the metadata of a data table may include the table name, field names, field types, data lengths, etc., and these information help to understand the content and structure of the data.
[0036] Among them, the metadata type refers to the classification or type of metadata, such as data tables, fields, views, etc. The metadata type defines the basic attributes and uses of the metadata. For example, in database management, the metadata type may include table metadata, field metadata, etc., which respectively describe the relevant information of database tables and fields.
[0037] Among them, the metadata attribute refers to the specific characteristics or attributes possessed by the metadata, such as name, description, data type, source, destination, etc. The metadata attribute is used to describe the characteristics and constraint conditions of the metadata in detail.
[0038] Step S202 obtains the lineage relationship model of the target data.
[0039] Among them, the blood relationship model is a model that describes the dependency and flow relationship between target data, and is used to display the source, destination, and conversion process of the target data. The blood relationship model helps to understand the full life cycle and flow path of the target data. For example, in a data warehouse, the blood relationship model may display the complete flow process of data from the original data source to the final report.
[0040] Step S203: Based on the blood relationship model, metadata type, and metadata attributes, use a preset graph algorithm to perform data blood relationship analysis on the target data to obtain the blood relationship analysis result of the target data.
[0041] Among them, the graph algorithm is an algorithm applied in a graph structure and is used to solve graph theory problems, such as graph traversal, search, shortest path, etc. The graph algorithm plays an important role in blood relationship analysis and can be used to discover the association and dependency relationships between data. For example, the depth-first search algorithm can be used to traverse the blood relationship graph to find all upstream sources of the data.
[0042] Among them, data blood relationship analysis is a process of analyzing the dependency and flow relationship between data, aiming to understand the source, destination, and conversion process of the data. Data blood relationship analysis helps to improve data quality and data management capabilities.
[0043] Among them, the blood relationship analysis result is the result obtained after data blood relationship analysis, including information such as the source, destination, and conversion process of the data. The blood relationship analysis result can be presented in the form of charts, reports, etc., which is convenient for understanding and analysis. For example, the blood relationship analysis result may display the complete flow path of a data field from the original data source to the final report.
[0044] Step S204: Based on the blood relationship analysis result, establish a blood relationship graph of the target data and display the blood relationship graph in a preset interactive interface.
[0045] Among them, the blood relationship graph is a chart that graphically displays the data blood relationship and is used to intuitively present the dependency and flow relationship between target data. The blood relationship graph helps to quickly understand the full life cycle and flow path of the data. For example, in the data governance platform of an insurance company, the blood relationship graph can display the complete process of customer data from initial collection (such as the insurance application information filled in by customers), through the processing and conversion of different systems (such as underwriting systems, claims systems), to the final use in risk assessment, customer service, or regulatory reports, etc.
[0046] Among them, the interaction interface is a graphical interface for users to interact with the computer system, used to receive user input, display system output, and provide operation feedback. The interaction interface is an important bridge for communication between users and the system. For example, in a data lineage analysis system, the interaction interface may provide functions such as data selection, analysis parameter setting, and result display.
[0047] Embodiments of the present application can provide a solid foundation for the accurate analysis of data lineage by precisely obtaining the target data metadata of multiple data sources and clarifying its type and attributes. By introducing a preset graph algorithm and combining it with the lineage model of the target data, in-depth analysis of data lineage is achieved, effectively solving the problems of complexity and dependency misunderstanding in lineage graph technology. This solution can automatically generate an accurate lineage graph and visually display it in a preset interaction interface, greatly improving the transparency and understandability of the project. This not only helps developers efficiently understand and manage data streams, but also effectively guides code decoupling and architecture optimization, reducing development costs and improving the overall quality and stability of the project.
[0048] In some alternative implementation manners of this embodiment, step 201 of obtaining the metadata of the target data from multiple data sources specifically includes the following steps:
[0049] Obtain a preset metadata extraction tool; use the metadata extraction tool to extract the initial metadata of the target data from multiple data sources, and preprocess the initial metadata to obtain the preprocessed metadata.
[0050] Among them, the metadata extraction tool is a specially designed software component, and its main function is to automatically identify and extract the metadata of the target data from various data sources. These tools are built with predefined parsing rules and data models, and can parse data source files with different formats and structures, such as database table structure files, configuration files, API documents, etc., and extract key information as metadata from them. The metadata extraction tool generates structured initial metadata by parsing the table structure, field definition, data type, relationship description, etc. in the data source.
[0051] Among them, the initial metadata refers to the metadata directly obtained from the data source through the metadata extraction tool and preliminarily processed. These metadata maintain the original information and structure in the data source when extracted, but have not undergone further cleaning, transformation, or standardization processing.
[0052] In one example, a financial insurance company needs to integrate and analyze data from multiple data sources, including insurance business systems, customer management systems, risk assessment platforms, and external data providers. These data sources contain a large amount of insurance policy information, customer identity information, risk assessment results, and market dynamic data. To effectively extract and process this data, a preset metadata extraction tool can be obtained first. This tool is specifically optimized for the characteristics of financial insurance data and can identify and extract key metadata such as policy numbers, applicant information, insured information, insurance amounts, insurance terms, risk levels, and claim records. Using this metadata extraction tool, the initial metadata of the target data is extracted from the above-mentioned multiple data sources. For example, metadata such as policy numbers, insurance amounts, and insurance terms are extracted from the insurance business system; metadata such as customer names, ID numbers, and contact information are extracted from the customer management system; metadata such as risk levels and claim probabilities are extracted from the risk assessment platform. After extraction, these initial metadata are preprocessed to obtain preprocessed metadata. The preprocessing process can include operations such as removing redundant information (such as duplicate policy records), unifying data formats, and filling in missing values (such as using interpolation to fill in missing risk level data). These processing steps ensure the quality and consistency of the metadata and provide a reliable basis for subsequent data analysis.
[0053] In one example, in the field of medical and health management, it is often necessary to process data from different data sources, such as electronic medical record systems, medical device monitoring systems, and drug management systems. To integrate and analyze this data, a preset metadata extraction tool can also be obtained. This tool is specifically designed for medical data and can efficiently extract key metadata such as patient information, diagnosis results, and drug usage records. Using this tool, the initial metadata of the target data is extracted from the above-mentioned data sources and these initial metadata are preprocessed. For example, metadata such as the age, gender, and diagnosis results of patients are extracted from the electronic medical record system; vital sign monitoring data is extracted from the medical device monitoring system. The preprocessing steps also include removing redundant information and unifying data formats.
[0054] In the embodiments of the present application, the initial metadata of the target data is extracted from multiple data sources through a metadata extraction tool, avoiding the cumbersome and error-prone nature of manual extraction and significantly improving the efficiency and accuracy of data lineage analysis. By preprocessing the initial metadata, such as data cleaning, format unification, and redundancy removal, the quality and usability of the metadata are further improved, laying a solid foundation for subsequent data lineage analysis.
[0055] In some optional implementation manners of this embodiment, in step S201, determining the metadata type and metadata attributes of the metadata specifically includes the following steps:
[0056] Obtain a predefined metadata model, match the metadata model with the metadata to determine the metadata type of the metadata; obtain a predefined list of metadata attributes, where the list of metadata attributes describes the attributes that different types of metadata have; match the metadata with the list of metadata attributes to determine the metadata attributes of the metadata.
[0057] Among them, the metadata model is a model that defines the structure and attributes of metadata, and is used to standardize and manage metadata. For example, in database design, the metadata model can define the attributes and relationships of metadata elements such as tables, fields, and indexes.
[0058] Among them, the list of metadata attributes is a list that enumerates the metadata attributes and is used to describe the attributes that different types of metadata possess. The list of metadata attributes helps to understand and manage the attribute information of metadata.
[0059] In an example, a financial insurance company needs to process various types of data including policy information, customer information, claim records, risk assessment reports, etc. To effectively manage this data, a set of metadata models can be predefined first. This model details the structure, meaning, and relationships among various types of data in the financial insurance domain. For example, policy information includes fields such as policy number, policyholder, insured, insurance amount, insurance period, etc., and customer information includes fields such as customer name, ID number, contact information, etc. When the initial metadata of the target data is obtained from multiple data sources (such as insurance business systems, customer management systems, claim processing systems, risk assessment platforms, etc.), the metadata is preprocessed to ensure data quality and consistency. After the preprocessing is completed, the preprocessed metadata is matched with the predefined metadata model. By comparing information such as the field names, data types, data structures, and business meanings of the metadata, the type of each piece of metadata can be accurately determined. For example, a certain piece of metadata is identified as policy information, and another is identified as customer information, etc. Further, to more precisely describe the metadata, a set of metadata attribute lists is also predefined. This list details the attributes that different types of metadata may have. For example, the attributes of policy information may include policy status, insurance type, and the attributes of customer information may include customer level, customer status. By matching the preprocessed metadata with the list of metadata attributes, the specific attributes of each piece of metadata can be determined.
[0060] In one example, in the field of healthcare management, the accuracy and integrity of data are directly related to the diagnosis and treatment of patients. Similarly, a set of metadata models and attribute lists applicable to the medical field can be predefined. For example, the patient information metadata model may include fields such as name, gender, age, medical record number, etc.; the examination result metadata model may include information such as examination items, examination results, examination time, etc. After obtaining the metadata of the target data from multiple data sources, these metadata are also matched with the predefined metadata models to determine their types. At the same time, through the matching of the metadata attribute list, the attribute description of each piece of metadata can be further refined to determine the metadata attributes of the metadata.
[0061] In the embodiments of the present application, by obtaining the predefined metadata models and matching them with the metadata obtained from multiple data sources, the type of each piece of metadata can be accurately identified, such as customer information, transaction records, etc., providing a clear data classification basis for subsequent data lineage analysis. At the same time, using the predefined metadata attribute list can further refine the attribute description of each piece of metadata, helping to more deeply understand the structure and meaning of the data, and providing rich information support for subsequent data processing and analysis. Through the matching process with the metadata models and attribute lists, the accurate identification and attribute determination of the target data metadata are realized, laying a solid foundation for subsequent data lineage analysis, effectively improving the accuracy and readability of the data lineage graph, helping developers better understand and manage the data flow, and enhancing the comprehensibility and maintainability of the project.
[0062] In some alternative implementation manners of this embodiment, step S202, obtaining the lineage model of the target data, specifically includes the following steps:
[0063] Using a preset data management tool, obtain the association relationships between the target data; based on the metadata and the association relationships, construct an initial lineage model of the target data; obtain the data architecture document of the target data, and extract the table-level association relationships of the target data from the data architecture document; obtain the design document of the target data, and extract the field-level association relationships of the metadata from the design document; based on the association relationships, table-level association relationships, and field-level association relationships, adjust the initial lineage model to obtain the lineage model of the target data.
[0064] Among them, the data management tool: a software tool for managing data, including functions such as data collection, storage, processing, and analysis. The data management tool is an important support for data management and analysis. For example, in a data warehouse, the data management tools may include ETL tools, data query tools, data analysis tools, etc.
[0065] Among them, the association relationship refers to a certain connection or dependency relationship existing between target data, such as parent-child relationship, reference relationship, etc. The association relationship is the basis for understanding the data structure and data flow. For example, in database design, the foreign key relationship between tables is an association relationship, which defines the data dependency and flow path between tables.
[0066] Among them, the initial blood relationship model is the model initially formed when constructing the blood relationship model, usually constructed based on metadata and association relationships. The initial blood relationship model provides a basis for subsequent adjustment and optimization.
[0067] Among them, the data architecture document is a document that describes the data architecture, including the definitions and relationships of elements such as data tables, fields, indexes, etc. The data architecture document is an important reference for understanding and managing the data architecture.
[0068] Among them, the table-level association relationship refers to the association relationship existing between tables in the data architecture, such as foreign key relationship, union query relationship, etc. The table-level association relationship describes the data dependency and flow path between data tables. For example, in the field of finance and insurance, a foreign key relationship is established between the policy table and the customer table through the customer ID field, indicating the association between the policy and the customer, that is, each policy belongs to a specific customer.
[0069] Among them, the design document is a document that describes the system design and data design, including information such as system architecture, data model, algorithm flow, etc. The design document is the basis for understanding and implementing the system.
[0070] Among them, the field-level association relationship refers to the association relationship existing between fields in a data table, such as calculation relationship, reference relationship, etc. between fields. The field-level association relationship describes the data dependency and conversion process between data fields.
[0071] In one example, a financial insurance company needs to process various types of data including policy information, customer information, claim records, risk assessment reports, etc. To build a lineage model among these data, first, a preset data management tool is used to automatically capture the association relationships among target data from multiple data sources (such as insurance business systems, customer management systems, claim processing systems, risk assessment platforms, etc.). These association relationships can include the association between policy information and customer information (such as a policy being associated with a specific customer), the association between claim records and policy information (such as a claim record being generated based on a specific policy), and the association between risk assessment reports and customer information or policy information, etc. Based on the captured association relationships and the metadata extracted from the data sources (including metadata types such as policies, customers, claim records, etc., and metadata attributes such as policy numbers, customer names, claim amounts, etc.), an initial lineage model is constructed. However, this initial model may only contain the basic associations between data and is not sufficient to describe the detailed structure of the data and the association relationships between fields. To further improve the model, the data architecture document of the target data is obtained. From the data architecture document, table-level association relationships are extracted. For example, it is found that there is a foreign key association between the policy information table and the customer information table, which links the two through the customer ID field. This table-level association relationship reveals the organizational structure of the data at the database level. In addition, the design document of the target data is obtained, and field-level association relationships are extracted from it. For example, in the claim record table, it is found that the calculation of the "claim amount" field may depend on the "insured amount" field and the "claim ratio" field in the policy information table. This field-level association relationship reveals the dependency relationship of the data at the business logic level. Based on these additional table-level and field-level association relationships, the initial lineage model is adjusted to obtain a lineage model that not only contains the basic associations between data but also details the association relationships of the data at the database level and the business logic level.
[0072] Embodiments of this application can efficiently capture the association relationships among target data through a data management tool, laying a solid foundation for building an initial lineage model. Further, by combining the data architecture document and the design document of the target data, table-level and field-level association relationships are respectively extracted. This meticulous operation ensures that the data associations at all levels in the lineage model are accurately presented. On this basis, the initial lineage model is adjusted in multiple dimensions, fully considering the complexity and diversity of the data, making the finally obtained lineage model closer to the actual business scenario. This modeling method that combines various data sources and document information not only enhances the reliability of the model but also significantly improves the traceability and understandability of data lineage, providing strong support for subsequent data management and analysis work.
[0073] In some alternative implementation manners of this embodiment, in step S203, based on the blood relationship model, metadata type, and metadata attributes, a preset graph algorithm is used to perform data blood relationship analysis on the target data to obtain the blood relationship analysis result of the target data, which specifically includes the following steps:
[0074] Obtain a parser and mapping rules according to the metadata type and metadata attributes; based on the mapping rules, use the parser to convert the metadata into node information in the blood relationship model; based on the mapping rules, use the parser to convert the association relationship, table-level association relationship, and field-level association relationship into edge information in the blood relationship model; generate a data graph of the target data based on the node information and edge information; use a preset graph algorithm to analyze the data graph to obtain the blood relationship analysis result of the target data.
[0075] Among them, the parser is used to extract key information from the preprocessed metadata according to the type and attributes of the metadata in the data blood relationship analysis. This information is then mapped to the nodes and edges of the blood relationship model, thereby constructing a data graph.
[0076] Among them, the mapping rules are a set of rules that define the data conversion logic. In the data blood relationship analysis, the mapping rules define how to map the metadata and its association relationships to the nodes and edges of the blood relationship model. These rules ensure the consistency and accuracy of the data during the conversion process.
[0077] Among them, the node information is the specific description of the nodes in the blood relationship model, including the identification, type, attributes, etc. of the nodes. In the data blood relationship analysis, the node information represents the abstract representation of data elements (such as database tables, fields, data streams, etc.). The node information is extracted from the metadata through the parser and mapping rules and is used to construct the data graph.
[0078] Among them, the edge information is the specific description of the edges in the blood relationship model, which is used to represent the association relationship between nodes. In the data blood relationship analysis, the edge information represents the dependency, transfer, or conversion relationship between data elements. The edge information is also extracted from the metadata through the parser and mapping rules and is used to construct the data graph.
[0079] Among them, the data graph is a graphical data structure used to represent data elements and their association relationships. In the data blood relationship analysis, the data graph is constructed based on the node information and edge information, intuitively showing the source, destination, and conversion process of the data.
[0080] In one example, the metadata of target data can be obtained from multiple data sources. These data sources may include an insurance business system database, a customer management system, a claims processing record repository, and an external market data interface, etc. Through data extraction and collation, the types of metadata are determined, such as policy information, customer information, claims records, and market data, as well as the corresponding attributes, such as policy number, customer ID, claim amount, and transaction time, etc. Subsequently, a lineage model related to the target data is obtained. This model serves as the basic framework for subsequent data processing and defines the possible types of association relationships between data. According to the metadata types and attributes, a suitable parser is selected from a pre-built parser library. The selection of the parser is based on its ability to process specific types of data and its ability to convert the data into node information recognizable in the lineage model. At the same time, the corresponding mapping rules are determined. These rules define how to map the metadata and its association relationships into nodes and edges in the lineage model. Based on the mapping rules, the selected parser is used to convert the metadata into node information in the lineage model. For example, policy information is converted into a policy node, and customer information is converted into a customer node. At the same time, the association relationships, table-level association relationships, and field-level association relationships are converted into edge information. For example, the association relationship between a policy and a customer is converted into an edge connecting the policy node and the customer node, and the field-level association relationship in the claims record referring to the policy information is converted into a specific field edge connecting the claims record node and the policy node. Based on the node information and edge information, a data graph of the target data is generated. This data graph intuitively shows the association relationships and data flows between data. By using preset graph algorithms, such as depth-first search and breadth-first search, the data graph is analyzed. By traversing the data graph through the algorithms, key lineage relationships such as the association path between policy information and customer information and the dependency relationship between claims records and policy information are identified. These analysis results are of great significance for understanding data flow, tracing data sources, ensuring data accuracy and compliance.
[0081] In the embodiments of the present application, by intelligently selecting matching parsers and mapping rules for different types of metadata and attributes, it is ensured that the metadata is accurately converted into node information in the lineage model. At the same time, based on the detailed mapping rules, the parser is used to effectively convert complex association relationships, table-level associations, and field-level association relationships into edge information in the model. This conversion process greatly enriches the details and levels of the lineage model. Based on these accurate node and edge information, a detailed data graph is constructed. This graph not only intuitively reflects the overall structure and internal connections of the target data but also lays a solid foundation for subsequent graph algorithm analysis. By using the preset graph algorithms to deeply analyze the data graph, the potential lineage relationships between data can be efficiently mined, further improving the accuracy and depth of lineage analysis.
[0082] In some alternative implementation manners of this embodiment, step S204, based on the blood relationship analysis result, establish a blood relationship graph of the target data, specifically including the following steps:
[0083] Obtain a preset graphical tool; create relationship graph nodes through the graphical tool based on the node information in the blood relationship analysis result; create relationship graph edges through the graphical tool based on the edge information in the blood relationship analysis result; obtain a preset layout algorithm, and based on the relationship graph nodes, relationship graph edges, and layout algorithm, render through the graphical tool to obtain the blood relationship graph of the target data.
[0084] Among them, the graphical tool is a software tool for creating, editing, and displaying graphical data structures. In data blood relationship analysis, the graphical tool is used to create a blood relationship graph according to the node information and edge information in the blood relationship analysis result.
[0085] Among them, the relationship graph node is a graphical element representing a data element in the blood relationship graph. In data blood relationship analysis, the relationship graph node is created based on the node information and is used to display information such as the identifier, type, and attributes of the data element. The relationship graph node is connected to other nodes through edges, jointly constituting the main structure of the data blood relationship graph.
[0086] Among them, the relationship graph edge is a graphical element representing the association relationship between data elements in the blood relationship graph. In data blood relationship analysis, the relationship graph edge is created based on the edge information and is used to display the dependency, flow, or conversion relationship between data elements. The relationship graph edge constructs the network structure of the data blood relationship graph by connecting relationship graph nodes.
[0087] Among them, the layout algorithm is an algorithm for determining the position and arrangement method of graphical elements in space. In data blood relationship analysis, the layout algorithm is used to optimize the presentation effect of the blood relationship graph, making the distribution of nodes and edges more reasonable and clear.
[0088] In one example, the metadata of target data can be obtained from multiple data sources. These data sources include the transaction database of the insurance business system, the user information database of the customer management system, the record database of the claim handling system, and the external market data interface, etc. Through data extraction and collation, the types of metadata are determined, such as policy transaction records, customer information, claim records, and market data, as well as the corresponding attributes, such as policy numbers, customer IDs, claim amounts, insurance product types, transaction times, etc. Next, the lineage model of the target data is obtained. This model describes the possible association relationships and data flows between data, and is the basis for subsequent lineage analysis. Based on the lineage model, metadata types, and metadata attributes, a preset graph algorithm (such as depth-first search, breadth-first search, etc.) is used to perform data lineage analysis on the target data. The aim is to discover the potential connections between data, such as the associations between policy transaction records, the correspondence between customer information and policy records, the dependency relationships between claim records and policies and customer information, etc. Through the analysis, the lineage analysis results of the target data are obtained, which include node information (such as data entities) and edge information (such as the association relationships between data). Subsequently, preset graphical tools are obtained. These tools have powerful graph rendering and interaction functions and can support the creation and display of complex relationship graphs. Based on the node information in the lineage analysis results, the nodes of the relationship graph are created through the graphical tools. These nodes represent data entities, such as policies, customers, claim records, etc. Similarly, based on the edge information, the edges of the relationship graph are created, and these edges represent the association relationships between data entities, such as the association between a policy and a customer, the dependency of a claim record on a policy, etc. To optimize the graph display effect, a preset layout algorithm is obtained, such as the force-directed layout algorithm or the hierarchical layout algorithm. These algorithms can automatically adjust the positions of nodes and edges, making the relationship graph clearer and easier to understand. The force-directed layout algorithm simulates the gravitational and repulsive forces in physical mechanics to keep a reasonable distance and layout between nodes. The hierarchical layout algorithm, on the other hand, makes the relationship graph more hierarchical and structural by presenting data entities and association relationships in layers. Finally, based on the relationship graph nodes, relationship graph edges, and layout algorithm, the lineage relationship graph of the target data is rendered through the graphical tool. This graph visually shows the association relationships and data flows between data, providing strong support for the data management, compliance auditing, and business decision-making of financial insurance companies.
[0089] In the embodiments of the present application, by obtaining a graphing tool, the lineage analysis results can be efficiently converted into an intuitive and easy-to-understand relationship graph. Based on the node information in the lineage analysis results, the graphing tool automatically creates relationship graph nodes to ensure that each data entity is clearly represented in the graph. At the same time, relationship graph edges are created according to the edge information to accurately depict the association relationships between data entities. In addition, a preset layout algorithm is adopted, combined with the relationship graph nodes and edges, and the graphing tool renders a lineage relationship graph with a clear structure and reasonable layout. In this process, the optimized application of the layout algorithm makes the distribution of nodes and edges in the relationship graph more balanced, avoiding graph chaos and visual interference, thus greatly improving the readability and practicality of the lineage relationship graph.
[0090] In some optional implementation manners of this embodiment, in step S204, after the lineage relationship graph is displayed on the preset interaction interface, the following steps are further included:
[0091] Obtain a preset task scheduler, a schedule, and a task script file; based on the task script file, use the task scheduler to trigger a lineage relationship update task of the lineage relationship graph according to the schedule to update the lineage relationship graph.
[0092] Among them, the task scheduler is a software tool for managing and scheduling computer tasks. In data lineage analysis, the task scheduler is used to trigger the update task of the lineage relationship graph according to the schedule. These tasks can include obtaining metadata from data sources, building a lineage relationship model, analyzing data lineage, and generating and updating the lineage relationship graph, etc. The task scheduler ensures the automation and timed execution of data lineage analysis, improving the efficiency and accuracy of the analysis.
[0093] Among them, the schedule is a tool for planning and managing the execution time of tasks. In data lineage analysis, the schedule defines the execution frequency and time points of the lineage relationship graph update task. The schedule can be adjusted and optimized according to business requirements and data change situations to ensure the timeliness and effectiveness of data lineage analysis.
[0094] Among them, the task script file is a text file containing automated task execution instructions. In data lineage analysis, the task script file defines the specific execution steps and parameters of the lineage relationship graph update task. These script files are usually read and executed by the task scheduler to achieve the automation and timing of data lineage analysis. The task script file can contain various types of instructions, such as data acquisition, model building, analysis calculation, and graph generation, etc.
[0095] Among them, the blood relationship update task is an automated task for updating and maintaining the data blood relationship diagram. In data blood relationship analysis, the blood relationship update task is triggered and executed according to a schedule, and calls components such as relevant parsers, mapping rules, and graphical tools to update the blood relationship diagram. These tasks ensure the timeliness and accuracy of the data blood relationship diagram.
[0096] In one example, to achieve the automatic update of the blood relationship diagram of target data, a preset task scheduler, schedule, and task script file can be obtained. The task script file defines the specific steps for updating the blood relationship diagram, including obtaining the latest metadata from the data source, re-performing blood relationship analysis, and updating the blood relationship diagram, etc. According to the schedule, the task scheduler automatically triggers the update task of the blood relationship diagram. For example, after the daily transactions end, the task scheduler runs the task script file, obtains the latest transaction data metadata from the data source, re-performs blood relationship analysis, and updates the blood relationship diagram on the interaction interface.
[0097] The embodiments of this application can achieve an automated update mechanism for the blood relationship diagram of target data by integrating a preset task scheduler, schedule, and task script file. The task scheduler triggers the update task at regular intervals according to the preset schedule, while the task script file details each operation in the update process. This mechanism ensures that the blood relationship diagram can reflect the latest data flow and transformation relationships in real time. As the data continues to grow and change, the blood relationship diagram can automatically capture these dynamics and maintain its accuracy and timeliness. The automated update not only reduces the workload of manually maintaining the blood relationship diagram but also reduces the risk of inaccurate or misleading blood relationship diagrams caused by human errors. In addition, the regularly updated blood relationship diagram helps developers promptly identify and solve problems in data flow, improving the overall quality and stability of the project.
[0098] It should be emphasized that to further ensure the above metadata, blood relationship analysis results, and blood relationship diagram, the above metadata, blood relationship analysis results, and blood relationship diagram can also be stored in the nodes of a blockchain.
[0099] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.
[0100] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0101] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0102] Further referring to Figure 3 As an implementation of the method shown above Figure 2 In an embodiment of a data lineage analysis visualization device provided by the present application, this device embodiment corresponds to the method embodiment shown in Figure 2 and can be specifically applied to various electronic devices.
[0103] As shown in Figure 3 the data lineage analysis visualization device 400 in this embodiment includes: a first acquisition module 401, a second acquisition module 402, an analysis module 403, and a construction module 404. Among them:
[0104] The first acquisition module 401 is used to acquire the metadata of the target data from multiple data sources and determine the metadata type and metadata attributes of the metadata;
[0105] The second acquisition module 402 is used to acquire the lineage relationship model of the target data;
[0106] The analysis module 403 is used to perform data lineage analysis on the target data based on the lineage relationship model, metadata type, and metadata attributes, and adopt a preset graph algorithm to obtain the lineage analysis result of the target data;
[0107] The construction module 404 is used to establish a lineage relationship graph of the target data based on the lineage analysis result and display the lineage relationship graph on a preset interaction interface.
[0108] Embodiments of the present application can provide a solid foundation for the accurate analysis of data lineage by precisely obtaining the target data metadata of multiple data sources and clarifying its type and attributes. By introducing a preset graph algorithm and combining it with the lineage relationship model of the target data, in-depth analysis of data lineage is achieved, effectively solving the problems of complexity and dependency misunderstanding in lineage graph technology. This solution can automatically generate an accurate lineage graph and visually display it in a preset interactive interface, greatly improving the transparency and understandability of the project. This not only helps developers efficiently understand and manage data flows, but also effectively guides code decoupling and architecture optimization, reducing development costs and improving the overall quality and stability of the project.
[0109] In one embodiment, the first acquisition module 401 includes:
[0110] The first acquisition sub-module is used to acquire a preset metadata extraction tool;
[0111] The first extraction sub-module is used to use the metadata extraction tool to extract the initial metadata of the target data from multiple data sources, and preprocess the initial metadata to obtain the preprocessed metadata.
[0112] Embodiments of the present application extract the initial metadata of the target data from multiple data sources through a metadata extraction tool, avoiding the cumbersome and error-prone nature of manual extraction, and can significantly improve the efficiency and accuracy of data lineage analysis. By preprocessing the initial metadata, such as data cleaning, format unification, redundancy removal, etc., the quality and usability of the metadata are further improved, laying a solid foundation for subsequent data lineage analysis.
[0113] In one embodiment, the first acquisition module 401 includes:
[0114] The first matching sub-module is used to acquire a predefined metadata model, match the metadata model with the metadata, and determine the metadata type of the metadata;
[0115] The second acquisition sub-module is used to acquire a predefined metadata attribute list, and the metadata attribute list describes the attributes of different types of metadata;
[0116] The second matching sub-module is used to match the metadata with the metadata attribute list to determine the metadata attributes of the metadata.
[0117] In the embodiments of the present application, by obtaining a predefined metadata model and matching it with the metadata obtained from multiple data sources, the type of each piece of metadata can be accurately identified, such as customer information, transaction records, etc., providing a clear data classification basis for subsequent data lineage analysis. At the same time, using the predefined metadata attribute list can further refine the attribute description of each piece of metadata, helping to more deeply understand the structure and meaning of the data, and providing rich information support for subsequent data processing and analysis. Through the matching process with the metadata model and attribute list, the accurate identification and attribute determination of the target data metadata are realized, laying a solid foundation for subsequent data lineage analysis, effectively improving the accuracy and readability of the data lineage diagram, helping developers better understand and manage the data flow, and enhancing the comprehensibility and maintainability of the project.
[0118] In one embodiment, the second acquisition module 402 includes:
[0119] A third acquisition sub-module, configured to use a preset metadata management tool to obtain the association relationship between target data;
[0120] A construction sub-module, configured to construct an initial lineage model of the target data based on the metadata and the association relationship;
[0121] A second extraction sub-module, configured to obtain the data architecture document of the target data and extract the table-level association relationship of the target data from the data architecture document;
[0122] A third extraction sub-module, configured to obtain the design document of the target data and extract the field-level association relationship of the metadata from the design document;
[0123] An adjustment sub-module, configured to adjust the initial lineage model based on the association relationship, the table-level association relationship, and the field-level association relationship to obtain the lineage model of the target data.
[0124] In the embodiments of the present application, the association relationship between target data can be efficiently captured through a data management tool, laying a solid foundation for constructing an initial lineage model. Further, by combining the data architecture document and the design document of the target data, the table-level and field-level association relationships are respectively extracted. This meticulous operation ensures that the data associations at all levels in the lineage model are accurately presented. On this basis, the initial lineage model is adjusted in multiple dimensions, fully considering the complexity and diversity of the data, making the finally obtained lineage model closer to the actual business scenario. This modeling method that combines multiple data sources and document information not only enhances the reliability of the model but also significantly improves the traceability and comprehensibility of the data lineage, providing strong support for subsequent data management and analysis work.
[0125] In one embodiment, the analysis module 403 includes:
[0126] A fourth acquisition sub-module, configured to acquire a parser and a mapping rule according to the metadata type and metadata attribute;
[0127] A first conversion sub-module, configured to convert the metadata into node information in the lineage model by using the parser based on the mapping rule;
[0128] A second conversion sub-module, configured to convert the association relationship, table-level association relationship, and field-level association relationship into edge information in the lineage model by using the parser based on the mapping rule;
[0129] A generation sub-module, configured to generate a metadata graph of the target data based on the node information and the edge information;
[0130] An analysis sub-module, configured to analyze the metadata graph by using a preset graph algorithm to obtain a lineage analysis result of the target data.
[0131] In the embodiment of the present application, by intelligently selecting a matching parser and mapping rule for different types of metadata and attributes, it is ensured that the metadata is accurately converted into node information in the lineage model. At the same time, based on the detailed mapping rule, the parser is used to effectively convert the complex association relationship, table-level association, and field-level association relationship into edge information in the model. This conversion process greatly enriches the details and levels of the lineage model. Based on these accurate node and edge information, a detailed data graph is constructed. This graph not only intuitively reflects the overall structure and internal connection of the target data, but also lays a solid foundation for the subsequent graph algorithm analysis. By deeply analyzing the data graph by using the preset graph algorithm, the potential lineage relationship between data can be efficiently mined, further improving the accuracy and depth of the lineage analysis.
[0132] In one embodiment, the establishment module 404 includes:
[0133] A fifth acquisition sub-module, configured to acquire a preset graphical tool;
[0134] A first creation sub-module, configured to create a relationship graph node through the graphical tool based on the node information in the lineage analysis result;
[0135] A second creation sub-module, configured to create a relationship graph edge through the graphical tool based on the edge information in the lineage analysis result;
[0136] A rendering sub-module, configured to acquire a preset layout algorithm, and render through the graphical tool based on the relationship graph node, the relationship graph edge, and the layout algorithm to obtain a lineage relationship graph of the target data.
[0137] Embodiments of the present application can obtain a graphical tool to efficiently convert the blood relationship analysis results into an intuitive and understandable relationship diagram. Based on the node information in the blood relationship analysis results, the graphical tool automatically creates relationship diagram nodes to ensure that each data entity is clearly represented in the diagram. At the same time, relationship diagram edges are created according to the edge information to accurately depict the association relationships between data entities. In addition, a preset layout algorithm is adopted, combined with the relationship diagram nodes and edges, and the graphical tool renders a blood relationship diagram with a clear structure and reasonable layout. In this process, the optimized application of the layout algorithm makes the distribution of nodes and edges in the relationship diagram more balanced, avoiding graph chaos and visual interference, thus greatly improving the readability and practicality of the blood relationship diagram.
[0138] In one embodiment, the data blood relationship analysis visualization device 400 further includes:
[0139] A third acquisition module, configured to acquire a preset task scheduler, a schedule, and a task script file;
[0140] A trigger module, configured to trigger a blood relationship update task of the blood relationship diagram according to the schedule by using the task scheduler based on the task script file to update the blood relationship diagram.
[0141] Embodiments of the present application can implement an automatic update mechanism for the blood relationship diagram of target data by integrating a preset task scheduler, a schedule, and a task script file. The task scheduler triggers the update task at regular intervals according to the preset schedule, and the task script file details each operation in the update process. This mechanism ensures that the blood relationship diagram can reflect the latest data flow and transformation relationships in real time. As the data grows and changes continuously, the blood relationship diagram can automatically capture these dynamics and maintain its accuracy and timeliness. The automatic update not only reduces the workload of manually maintaining the blood relationship diagram but also reduces the risk of inaccurate or misleading blood relationship diagrams caused by human errors. In addition, the regularly updated blood relationship diagram helps developers identify and solve problems in data flow in a timely manner, improving the overall quality and stability of the project.
[0142] To solve the above technical problems, embodiments of the present application also provide a computer device. For details, please refer to Figure 4 , Figure 4 which is the basic structural block diagram of the computer device in this embodiment.
[0143] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that communicate with each other via a system bus. It should be noted that only the computer device 6 with the memory 61, the processor 62, and the network interface 63 is shown in the figure. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of this technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0144] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can interact with the user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.
[0145] The memory 61 includes at least one type of readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 61 can be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 can also be an external storage device of the computer device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 6. Of course, the memory 61 can also include both the internal storage unit and the external storage device of the computer device 6. In this embodiment, the memory 61 is generally used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions for the data lineage analysis visualization method. In addition, the memory 61 can also be used to temporarily store various data that have been output or will be output.
[0146] The processor 62 may be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor 62 is generally used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to run the computer-readable instructions stored in the memory 61 or process data, such as running the computer-readable instructions of the data lineage analysis visualization method.
[0147] The network interface 63 may include a wireless network interface or a wired network interface, and the network interface 63 is generally used to establish a communication connection between the computer device 6 and other electronic devices.
[0148] The embodiments of the present application can provide a solid foundation for the accurate analysis of data lineage by accurately obtaining the target data metadata of multiple data sources and clarifying its type and attributes. By introducing a preset graph algorithm and combining it with the lineage relationship model of the target data, in-depth analysis of data lineage is achieved, effectively solving the problems of complexity and dependency misunderstanding in the lineage graph technology. This solution can automatically generate an accurate lineage graph and visually display it in a preset interactive interface, greatly improving the transparency and understandability of the project. This not only helps developers efficiently understand and manage data streams, but also effectively guides code decoupling and architecture optimization, reducing development costs and improving the overall quality and stability of the project.
[0149] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that at least one processor executes the steps of the data lineage analysis visualization method as described above.
[0150] The embodiments of the present application can provide a solid foundation for the accurate analysis of data lineage by accurately obtaining the target data metadata of multiple data sources and clarifying its type and attributes. By introducing a preset graph algorithm and combining it with the lineage relationship model of the target data, in-depth analysis of data lineage is achieved, effectively solving the problems of complexity and dependency misunderstanding in the lineage graph technology. This solution can automatically generate an accurate lineage graph and visually display it in a preset interactive interface, greatly improving the transparency and understandability of the project. This not only helps developers efficiently understand and manage data streams, but also effectively guides code decoupling and architecture optimization, reducing development costs and improving the overall quality and stability of the project.
[0151] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0152] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements for some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is similarly within the scope of the patent protection of the present application. The non-company enterprise software tools or components that appear in the embodiments of the present application are only for illustrative purposes and do not represent actual use.
Claims
1. A method for visualizing data lineage analysis, characterized in that, Including the following steps: Obtain the metadata of the target data from multiple data sources, and determine the metadata type and metadata attributes of the metadata; Obtain the lineage model of the target data; Based on the lineage model, the metadata type, and the metadata attributes, use a preset graph algorithm to perform data lineage analysis on the target data to obtain the lineage analysis result of the target data; Based on the lineage analysis result, establish a lineage graph of the target data, and display the lineage graph on a preset interactive interface.
2. The method according to claim 1, characterized in that, The step of determining the metadata type and metadata attributes of the metadata specifically includes: Obtain a predefined metadata model, match the metadata model with the metadata, and determine the metadata type of the metadata; Obtain a predefined list of metadata attributes, where the list of metadata attributes describes the attributes of different types of metadata; Match the metadata with the list of metadata attributes to determine the metadata attributes of the metadata.
3. The method according to claim 1, wherein The step of obtaining the lineage model of the target data specifically includes: Use a preset data management tool to obtain the association relationships between the target data; Based on the metadata and the association relationships, construct an initial lineage model of the target data; Obtain the data architecture document of the target data, and extract the table-level association relationships of the target data from the data architecture document; Obtain the design document of the target data, and extract the field-level association relationships of the metadata from the design document; Based on the association relationships, the table-level association relationships, and the field-level association relationships, adjust the initial lineage model to obtain the lineage model of the target data.
4. The method according to claim 3, characterized in that, The step of using a preset graph algorithm to perform data lineage analysis on the target data based on the lineage model, the metadata type, and the metadata attributes to obtain the lineage analysis result of the target data specifically includes: According to the metadata type and the metadata attributes, obtain a parser and a mapping rule; Based on the mapping rule, use the parser to convert the metadata into node information in the lineage model; Based on the mapping rule, use the parser to convert the association relationships, the table-level association relationships, and the field-level association relationships into edge information in the lineage model; Based on the node information and the edge information, generate a data graph of the target data; Use a preset graph algorithm to analyze the data graph to obtain the lineage analysis result of the target data.
5. The method according to claim 1, wherein The step of establishing a lineage graph of the target data based on the lineage analysis result specifically includes: Obtain a preset graphical tool; Based on the node information in the lineage analysis result, create relationship graph nodes through the graphical tool; Based on the edge information in the lineage analysis result, create relationship graph edges through the graphical tool. Obtain a preset layout algorithm, and based on the relationship graph nodes, the relationship graph edges, and the layout algorithm, render the lineage graph of the target data through the graphical tool.
6. The method according to claim 1, wherein The step of obtaining the metadata of the target data from multiple data sources specifically includes: Obtain a preset metadata extraction tool; Use the metadata extraction tool to extract the initial metadata of the target data from multiple data sources, and preprocess the initial metadata to obtain the preprocessed metadata.
7. The method according to claim 1, characterized in that After the step of displaying the lineage graph on the preset interaction interface, it further includes: Obtain a preset task scheduler, a time schedule, and a task script file; Based on the task script file, use the task scheduler to trigger the lineage update task of the lineage graph according to the time schedule to update the lineage graph.
8. A data lineage analysis visualization device, characterized in that It includes: A first acquisition module for obtaining the metadata of the target data from multiple data sources and determining the metadata type and metadata attributes of the metadata; A second acquisition module for obtaining the lineage model of the target data; An analysis module for performing data lineage analysis on the target data based on the lineage model, the metadata type, and the metadata attributes using a preset graph algorithm to obtain the lineage analysis result of the target data; A construction module for constructing the lineage graph of the target data based on the lineage analysis result and displaying the lineage graph on a preset interaction interface.
9. A computer device, characterized in that, It includes a memory and a processor. Computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the data lineage analysis visualization method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by the processor, the steps of the data lineage analysis visualization method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Metadata tag generation method and device, equipment, medium and program product
CN120929448A
Intelligent tracking method and system for complex data blood relationship
CN121092531A
Intelligent tracking method and system for complex data blood relationship
CN121092531B