A method and system for managing blood relationships of data

CN117033410BActive Publication Date: 2026-07-21DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310920275.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-25
Publication Date
2026-07-21
Estimated Expiration
2043-07-25

AI Technical Summary

Technical Problem

Existing kinship analysis methods suffer from low accuracy, limited analytical scope, and inability to determine the cause of failures, leading to data processing chain collapse and hindering the execution of other data tasks.

Method used

By connecting to different types of data sources through a unified interface, initial data is collected and the lineage relationship is extracted. Parsing and data collection methods are used to verify the compliance of transaction data offline and convert it into a preset format for display.

Benefits of technology

It improved the efficiency and accuracy of kinship analysis, accurately located fault processes, ensured the stable operation of the big data system, and realized the compliant circulation and secure use of transaction data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117033410B_ABST
    Figure CN117033410B_ABST
Patent Text Reader

Abstract

The disclosure provides a kind of management method and device of transaction data's blood relationship, it is related to big data analysis technical field.The specific implementation of the method includes: receiving the blood relationship analysis request of transaction data;Wherein, blood relationship analysis request includes the data source type of one or more data sources;According to data source type, determine the collection scheme of transaction data;Scan each data source, according to the collection scheme corresponding to data source, offline collection blood relationship analysis request initial data;Wherein, initial data includes data source identification;According to data source identification, initial data is shunted, extract the blood relationship of transaction data, and the blood relationship is shown.The implementation can realize the compliance circulation of transaction data, improve the analysis efficiency, accuracy and completeness of blood relationship, and accurately locate fault fast recovery, avoid the hindrance brought by local anomaly, guarantee the stable operation of big data system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of big data analytics, and in particular to a method and system for managing data lineage. Background Technology

[0002] The concept of "bloodline relationship" refers to the link relationship between data. It can represent the entire flow process of data from generation, processing, creation to extinction. This enables the development of big data to provide assistance for data use, helping data users to efficiently manage massive, complex, and chaotic data, thereby effectively supervising data, controlling data risks, and improving the value of data use.

[0003] Existing lineage analysis processes typically employ log data analysis and open-source tool analysis to process and analyze lineage relationships in log data; or they utilize open-source tools such as Atlas and Nifi to extract data based on pre-defined metadata and analyze the lineage relationships between the data.

[0004] However, log data is typically diverse in type, massive in volume, and complex, resulting in significant noise and low completeness in the analyzed data, leading to low accuracy in lineage relationship analysis. Furthermore, open-source tools support a limited range of data sources, and some missing data is simply ignored, resulting in incomplete information. Metadata information is also often too fragmented, limiting the scope and reducing the completeness of lineage relationship analysis. Moreover, when analysis anomalies or data processing failures occur, existing analysis methods are unable to determine the cause of the failure, causing prolonged analysis disruptions, collapsing the entire data processing chain, and hindering the execution of other data tasks. Summary of the Invention

[0005] In view of this, the present disclosure provides a method and system for managing the lineage of transaction data, which can solve the problems of low accuracy of lineage relationships; limited analysis scope and low completeness; inability to determine the cause of failure; resulting in long-term analysis obstruction; collapse of the entire data processing chain; and hindering the execution of other data tasks.

[0006] To achieve the above objectives, according to one aspect of this disclosure, a method for managing the lineage of transaction data is provided, comprising:

[0007] Receive a lineage analysis request for transaction data; wherein, the lineage analysis request includes data source types of one or more data sources;

[0008] Determine the data collection plan for the transaction data based on the data source type;

[0009] Scan each of the data sources and collect the initial data for the blood relationship analysis request offline according to the collection scheme corresponding to the data source; wherein, the initial data includes the data source identifier;

[0010] The initial data is split according to the data source identifier, the lineage of the transaction data is extracted, and the lineage is displayed.

[0011] According to another aspect of this disclosure, a management system for the lineage of transaction data is provided, comprising:

[0012] A receiving module is used to receive a lineage analysis request for transaction data; wherein, the lineage analysis request includes one or more data source types;

[0013] The data processing module is used to determine the data collection scheme for the transaction data based on the data source type.

[0014] The data acquisition module is used to scan each of the data sources and, according to the acquisition scheme corresponding to the data source, to collect the initial data of the blood relationship analysis request offline; wherein, the initial data includes the data source identifier;

[0015] The display module is used to split the initial data according to the data source identifier, extract the lineage relationship of the transaction data, and display the lineage relationship.

[0016] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0017] Processor; and

[0018] Stored program memory,

[0019] The program includes instructions that, when executed by the processor, cause the processor to perform a method for managing the lineage of the transaction data.

[0020] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute a method for managing the lineage of transaction data.

[0021] One or more technical solutions provided in this application embodiment connect to different types of data sources through a unified interface, collect initial data and extract lineage relationships in a distributed manner, and display them in the required form. This can achieve compliant flow of transaction data, improve the efficiency and accuracy of lineage relationship analysis, accurately locate fault processes for rapid recovery, avoid obstacles caused by local anomalies, ensure the stable operation of big data systems, and thus guarantee the technical effects of compliant management of lineage relationships and secure use of big data. Attached Figure Description

[0022] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0023] Figure 1 A flowchart illustrating a method for managing the lineage of transaction data according to an exemplary embodiment of this disclosure is shown;

[0024] Figure 2 A flowchart illustrating a method for determining a data acquisition scheme according to an exemplary embodiment of the present disclosure is shown;

[0025] Figure 3 A flowchart illustrating a method for acquiring initial data according to an exemplary embodiment of the present disclosure is shown;

[0026] Figure 4 A flowchart illustrating a method for extracting blood relations according to exemplary embodiments of the present disclosure is shown;

[0027] Figure 5 A flowchart of a method for locating abnormal faults according to an exemplary embodiment of the present disclosure is shown;

[0028] Figure 6 A flowchart illustrating a method for querying blood relations according to an exemplary embodiment of the present disclosure is shown;

[0029] Figure 7 A schematic block diagram of a management system for the kinship of transaction data according to an exemplary embodiment of the present disclosure is shown;

[0030] Figure 8 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0031] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0032] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0033] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0035] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0036] Hive is a data warehouse tool based on the Hadoop distributed architecture, used for data extraction, transformation, and loading (ETL).

[0037] Spark is a fast, general-purpose, and scalable big data processing framework that provides efficient data processing and analysis capabilities, supporting tasks such as large-scale data processing, machine learning, graph computing, and stream processing.

[0038] Flink is an open-source big data processing framework that also supports high-performance, scalable, and fault-tolerant large-scale data processing, enabling various computing modes such as streaming, batch, and iterative processing.

[0039] GP: GreenPlum is a relational database built on an open-source platform and employing a massively parallel processing architecture, capable of handling large-scale data analysis tasks.

[0040] Data lineage analysis is further enhancing the value of big data. However, with the diversification of data sources and the explosive growth of data volume, lineage analysis based on log data suffers from complex formats, massive data volumes, and severe information loss, causing the computational costs required for lineage processing to skyrocket while accuracy remains low. Lineage analysis based on open-source tools relies on the accurate construction of metadata and has limited analytical capabilities, supporting only specific types of data sources. It also struggles with missing or confidential fields (such as "****"), leading to analysis delays and hindering the execution of other data tasks. Furthermore, when real-time analysis encounters data anomalies or analysis failures, the inability to accurately pinpoint the cause not only hinders analysis efficiency but also interrupts the operation of the big data system. All of these factors contribute to a surge in the cost of lineage analysis and management, an inability to control data risks in transaction data, and poor security in the use of big data.

[0041] The disclosed method for managing the lineage of transaction data connects to different types of source databases or data platforms through a unified interface. It verifies the compliance of transaction data offline, converts it into a preset format, and then extracts the lineage relationships for display in the required form. It supports the compliance conversion of any type of data source and special fields, realizing the compliant flow of transaction data, improving the efficiency and accuracy of lineage relationship analysis, accurately locating fault processes for rapid recovery, avoiding obstacles caused by local anomalies, ensuring the stable operation of the big data system, and thus guaranteeing the compliant management of lineage relationships and the secure use of big data.

[0042] The present disclosure is described below with reference to the accompanying drawings.

[0043] Figure 1 A flowchart illustrating a method for managing the kinship of transaction data according to an exemplary embodiment of this disclosure is shown, such as... Figure 1 As shown, the method for managing the lineage of transaction data disclosed herein includes the following steps:

[0044] In this embodiment of the disclosure, the method for managing the lineage of transaction data is executed by a lineage analysis server, which includes a data collection interface and a display interface.

[0045] Step S101: Receive a lineage analysis request for transaction data; wherein the lineage analysis request includes the data source type and database identifier of one or more data sources.

[0046] In this embodiment, the data source types include relational databases and data platforms. Relational databases use SQL statements to perform queries and other processing, such as offline big data engines like Hive, Spark, Flink, and Gp. Data platforms do not use SQL statements to perform queries and other processing. The lineage analysis server receives one or more lineage analysis requests for transaction data sent by the requesting terminal. The lineage analysis request is selected by the user of the requesting terminal according to actual needs; for example, the lineage analysis request may be to analyze the closeness of association between transaction users based on transaction location.

[0047] Furthermore, the lineage analysis request includes the data source address, access permissions, analysis source, analysis table, and / or analysis fields for each data source. Different data sources can be configured with different access permissions to restrict the transaction data that the lineage analyzer can access, ensuring data security and preventing data leakage. For example, access permissions include high-level and low-level. High-level access allows access to all data from the data source, while low-level access only allows access to a portion of the data source.

[0048] Step S102: Determine the data collection scheme for the transaction data based on the data source type.

[0049] In this embodiment, the lineage analysis server determines the collection scheme for the collection interface based on the different data source types. The collection scheme, as an executable roadmap, guides the collection interface in performing query and collection operations. The collection scheme includes collection methods, including parsing collection and event tracking collection. Parsing collection is suitable for relational databases; it generates a collection scheme by parsing the request statement of the lineage analysis request, enabling the collection interface to execute the collection scheme and obtain the initial data from the relational database. Event tracking collection is suitable for data platforms; it generates a collection scheme based on the data platform's calling interface, enabling the collection interface to connect to the calling interface and receive the initial data collected by the calling interface using multiple pre-set event tracking points on the data platform.

[0050] In the embodiments disclosed herein, such as Figure 2 As shown, the method for determining the data acquisition scheme disclosed herein includes the following steps:

[0051] Step S201: Obtain the data source type of the blood relationship analysis request.

[0052] Step S202: Determine whether the data source type is a relational database. If yes, proceed to step S203; otherwise, proceed to step S205.

[0053] Step S203: Determine the acquisition method as parsing acquisition.

[0054] In this embodiment of the disclosure, the relational database uses SQL statements, and when processing the lineage analysis request of the relational database, a parsing-based acquisition method is adopted.

[0055] Step S204: Parse the request statement of the blood relationship analysis request to generate a collection plan including collection type, collection time and collection content.

[0056] In this embodiment of the disclosure, the lineage analysis server parses the request statement, extracts the information useful for data collection, and transforms it into a collection plan. That is, the lineage analysis server parses the SQL statement, extracts the useful information, and transforms it into a collection plan, including collection type, collection time, collection content, output template, etc.

[0057] Furthermore, the data collection types include incremental collection and full collection. Incremental collection only collects the changed parts of the transaction data, which can reduce the costs of data collection, transmission, and processing. It is suitable for data sources with large data volume changes. The changed parts of the transaction data can be operations such as adding, deleting, and modifying transaction data. This can be determined by whether the metadata of the data source has changed. Changes in metadata indicate the existence of incremental data, while unchanged metadata indicates the absence of incremental data. The metadata includes file names, column numbers, project names, fields, access permissions, etc. For example, the metadata for day one includes table_a, table_b, table_c, and table_e, while the metadata for day two includes table_a, table_b, table_c, and table_d. Incremental collection for day one involves adding table_a, table_b, table_c, and table_e, while incremental collection for day two involves deleting table_e and adding table_d. Full collection, on the other hand, collects all transaction data each time, providing a complete representation of the transaction data. This is suitable for static data sources or small-scale data sources. For example, full collection for day one would include table_a, table_b, table_c, and table_e, while full collection for day two would include table_a, table_b, table_c, and table_d.

[0058] The data collected includes: ① field keywords and field attribute values; ② data content keywords and field attribute values; ③ field keywords, field attribute values, and field correspondences; ④ data content keywords, field attribute values, and field correspondences; ⑤ source data of the data source; and ⑥ table data of the data table. The output template includes the output data source, output data table, output fields, and field attribute values.

[0059] Furthermore, for example, the request statement is to collect incremental data of transaction locations from the transaction information table every day. The corresponding collection type is incremental collection, the collection time is daily, the collection content includes the attribute values ​​of "transaction location" and "transaction location", and the output fields are the attribute values ​​of "output location" and "transaction location".

[0060] Step S205: Determine the data collection method as embedded point collection.

[0061] In this embodiment, multiple data points are pre-set in the data platform. When processing the lineage analysis request of the data platform, the data collection method using data points is adopted. These data points can be in the form of code functions, sending the function's output data to the calling interface; alternatively, data points can be selectively set according to actual data collection needs, such as data source, table name, field keywords, and data content keywords.

[0062] Step S206: Determine the calling interface address based on the platform identifier of the bloodline analysis request.

[0063] In this embodiment of the disclosure, when data is collected using the method of tracking points, it is necessary to call the calling interface of the data platform to obtain the data of the tracking points. Therefore, the calling interface address of the calling interface connected to the collection interface is determined according to the correspondence between the data platform identifier and the calling interface address.

[0064] Step S207: Generate a collection scheme including the call interface address, collection type, collection time, and collection content.

[0065] In the embodiments of this disclosure, the method for determining the acquisition scheme of this disclosure selects different acquisition methods and determines different acquisition schemes according to different data source types, which can meet the acquisition needs of different types of data sources, so that the subsequent acquisition interface can collect initial data from each data source. It has strong scalability and wide adaptability.

[0066] Step S103: Scan each of the data sources and collect the initial data of the blood relationship analysis request offline according to the collection scheme corresponding to the data source; wherein, the initial data includes the data source identifier.

[0067] In this embodiment of the disclosure, the data source identifier can be a database identifier or a platform identifier. The collection interface of the bloodline analysis server collects the initial data of the database according to the collection scheme, or the interface receiving the initial data returned by the interface according to the collection scheme, and converts and stores it in a unified format to facilitate the subsequent analysis of bloodline relationships.

[0068] In the embodiments disclosed herein, such as Figure 3 As shown, the initial data acquisition method of this disclosure includes the following steps:

[0069] In this embodiment of the disclosure, the initial data is collected by the acquisition interface using an offline scanning method, and the initial data is stored in a unified database, thereby reducing the consumption of network bandwidth and server resources and improving the efficiency of data acquisition and processing.

[0070] Step S301: Obtain the acquisition scheme.

[0071] Step S302: Determine whether the acquisition method is parsing acquisition. If yes, proceed to step S303; otherwise, proceed to step S307.

[0072] Step S303: Determine whether the acquisition interface has access permissions to the data source. If yes, proceed to step S304; otherwise, proceed to step S306.

[0073] In this embodiment of the disclosure, the access permissions of the data source can be selectively set according to actual data confidentiality requirements.

[0074] Step S304: Access the data source address using the acquisition interface.

[0075] In this embodiment of the disclosure, the data acquisition interface establishes a connection with the relational database based on the data source address of the data source.

[0076] Step S305: Collect and output data according to the output template based on the collection type, collection time, and collection content, and proceed to step S309.

[0077] In this embodiment of the disclosure, the data acquisition interface traverses each data source and collects output data according to the output template, including the attribute values ​​of each output field, the source data of each data source, and the table data of each data table.

[0078] Step S3051: When the collected content consists of field keywords and field attribute values, collect the field attribute values ​​corresponding to the field keywords as the output data.

[0079] In this embodiment of the disclosure, for example, the output template includes output fields and field attribute values. The field keyword is "transaction location". The field attribute value of "transaction location" is directly collected as the field attribute value of the output field "output location".

[0080] Step S3052: When the collected content consists of data content keywords and field attribute values, according to a preset ratio, a number of comparison data entries equal to the preset ratio are filtered from the data source. The data content keywords are matched with the comparison data to determine the target field containing the data content keywords. The field attribute values ​​of the target field are collected as the output data.

[0081] In this embodiment of the disclosure, for example, the output template includes output fields and field attribute values. The data content keyword is the first four digits of the ATM number, "****". The preset ratio is 3%. The total number of data entries in the data source is 1252. 37 comparison data entries are filtered. "****" is matched with the 37 comparison data entries to determine the target field "transaction location" where the data content keyword exists. The attribute value of the target field "transaction location" is collected as the field attribute value of the output field "output location".

[0082] By matching keywords in the data content, for missing fields in the data source, the field attribution can be determined through data content matching, the target field corresponding to the data content keywords can be obtained, and the field attribute values ​​can be extracted as output data, ensuring the accuracy and completeness of data collection.

[0083] Step S3053: When the collected content consists of field keywords, field attribute values, and field correspondences, collect the field attribute values ​​corresponding to the field keywords to obtain first data; based on the field correspondences, collect the field attribute values ​​of the corresponding fields of the field keywords, and combine them with the first data to obtain the output data.

[0084] In this embodiment of the disclosure, for example, the output template includes multiple output fields and their field attribute values. The first data is the output field "output location" and its field attribute values. The field correspondence is transaction location-username. The field attribute values ​​of the username corresponding to the transaction location are collected to obtain the "output username" field and its field attribute values. Combined with the first data, this is the output data.

[0085] Furthermore, the corresponding fields can also be determined by matching data content keywords to ensure data integrity.

[0086] Step S3054: When the collected content consists of data content keywords, field attribute values, and field correspondences, collect the field attribute values ​​corresponding to the target field to obtain second data; based on the field correspondences, collect the field attribute values ​​of the corresponding fields of the target field, and combine them with the second data to obtain the output data.

[0087] In this embodiment of the disclosure, for example, the output template includes multiple output fields and their field attribute values. The second data is the output field "output location" and its field attribute values ​​corresponding to the data content keyword "****". The field correspondence is transaction location-username. The field attribute values ​​of the username corresponding to the transaction location are collected to obtain the "output username" field and its field attribute values. Combined with the second data, this is the output data.

[0088] Step S3055: If the collected content is source data of the data source, collect incremental data or full data corresponding to the data source identifier of the data source as the output data.

[0089] Step S3056: If the collected content is table data of a data table, collect incremental data or full data corresponding to the table name of the data table as the output data.

[0090] Step S306: Reject the blood relationship analysis request.

[0091] In this embodiment of the disclosure, if the data acquisition interface does not have access to the data source, the bloodline analysis server rejects the bloodline analysis request.

[0092] Step S307: The acquisition interface is connected to the calling interface according to the calling interface address.

[0093] In this embodiment of the disclosure, the data acquisition interface establishes a connection with the relational database based on the interface address of the calling interface.

[0094] Step S308: Receive the output data returned by the calling interface according to the collection type, the collection time, and the collection content, and proceed to step S309.

[0095] In this embodiment of the disclosure, the calling interface receives the output data returned by each data point according to the collection type, collection time and collection content, and sends it to the collection interface.

[0096] Step S309: Convert the output data into a preset format to obtain initial data, and store the initial data in the analysis database.

[0097] In this embodiment, the analysis database is located in a lineage analysis server. The preset format can be selectively set according to actual needs, such as JSON format. Initial data includes the converted database identifier or platform identifier, data collection start time, data collection end time, output data, etc.

[0098] In this embodiment of the disclosure, the initial data collection method of this disclosure directly collects the attribute values ​​of each field, or matches the missing target field according to the data content and collects the attribute values ​​of the target field, converts them into a preset format, obtains initial data, and stores it. This can improve the accuracy and completeness of the collected data, reduce the computing power cost required for bloodline processing, improve the accuracy and completeness of bloodline extraction, and enhance analysis efficiency. This ensures the smooth operation of analysis tasks and other data tasks, and reduces data management costs and computing resource costs.

[0099] Step S104: Divide the initial data according to the data source identifier, extract the lineage relationship of the transaction data, and display the lineage relationship.

[0100] In this embodiment, the lineage analysis server distributes the initial data to different relation parsers based on the database identifier or platform identifier. Each relation parser extracts the lineage relationships, generates various lineage relationship display styles, and displays them through the terminal. This facilitates global control and understanding of the lineage relationships of transaction data, timely location and repair of faults, and improves the efficiency of lineage relationship analysis while ensuring the normal operation of various data tasks.

[0101] In the embodiments disclosed herein, such as Figure 4 As shown, the method for extracting blood relations disclosed herein includes the following steps:

[0102] Step S401: Obtain the initial data.

[0103] Step S402: Based on the database identifier or the platform identifier, the initial data is distributed to different relation resolvers.

[0104] In this embodiment of the disclosure, the lineage analysis server includes multiple relation resolvers, each corresponding to a different data source, and the resolver identifier of the relation resolver is the same as the database identifier or platform identifier.

[0105] In step S403, in response to the analysis target of the bloodline analysis request, the relationship parser extracts the initial data of the analysis source, the analysis table, and / or the analysis field from the initial data.

[0106] In this embodiment of the disclosure, the relation resolver includes multiple parsing threads. Different parsing threads can respond to different analysis targets. For example, if the analysis target is the lineage relationship between the data table name and the analysis fields, the parsing thread extracts the initial data of the analysis table and the analysis fields; if the analysis target is the lineage relationship between the analysis fields, the parsing thread extracts the initial data of each analysis field; if the analysis target is the lineage relationship of the data source, the parsing thread extracts the initial data of the analysis source; or if the analysis target is the lineage relationship of the data table, the parsing thread extracts the initial data of the analysis table.

[0107] Furthermore, the relation resolver also includes an aggregation thread to aggregate the lineage relationships of multiple resolution threads, resulting in an aggregated lineage relationship. For example, a resolution thread might extract the lineage relationship between "username-transaction location" and "username-user address," while the aggregation thread extracts the lineage relationship between "transaction location-user address." Similarly, a resolution thread might extract the lineage relationship between "data source" and "data platform," while the aggregation thread extracts the lineage relationship between "data source-data platform." It should be noted that the relation resolver can extract various types of data according to actual analytical needs to analyze the lineage relationships between data sources, between data sources and data tables, between data tables and fields, between fields, and between a single data source, a single data table, or a single field.

[0108] Furthermore, the parsing thread can perform operations such as cleaning the initial data before extracting the initial data from the analysis source, analysis table, and / or analysis fields.

[0109] Step S404: Extract the lineage of the initial data of the analysis source, the analysis table, and / or the analysis field.

[0110] In this embodiment of the disclosure, the extraction of lineage relationships for each parsing thread may include the following scenarios in response to different analysis objectives:

[0111] Step S4041: The analysis fields include username and transaction location. Taking the field attribute value of the username as the center, the lineage relationship between the field attribute value of the transaction location corresponding to the username is extracted.

[0112] Step S4042: Extract the lineage relationship between the table name of the analysis table and the analysis fields.

[0113] In this embodiment of the disclosure, the analysis field is derived from the output field. Since the output data has undergone processing such as data content keyword matching, even if the data source field is missing, it will not affect the accuracy and completeness of the blood relationship analysis.

[0114] Step S4043: The analysis fields include transaction account, transaction amount, and transaction time. Based on the field attribute values ​​of transaction amount and transaction time corresponding to the transaction account, the lineage of the transaction process is extracted.

[0115] In this embodiment of the disclosure, the relation parser stores the bloodline relationships extracted by each parsing thread in the form of key-value pairs to the disk of the bloodline analysis server. The key of the key-value pair is the analysis target, and the value is the extracted bloodline relationship.

[0116] Furthermore, based on the key of the key-value pair, the offset of the corresponding key is constructed and stored in the memory of the lineage analysis server. Thus, during the query, the offset of the key-value pair can be determined from memory first, and then the value of the key-value pair can be read from the disk to obtain the lineage relationship. This effectively utilizes the storage characteristics of memory and disk, reducing the storage and access pressure on the lineage analysis server.

[0117] Furthermore, before storing the bloodline key-value pairs, each bloodline is deduplicated to reduce the storage cost of the bloodline analysis server.

[0118] In this embodiment of the disclosure, the lineage relationship extraction method is used to distribute the initial data to each relation parser according to the data source identifier, and then allocate it to each parsing thread to extract the lineage relationship according to the analysis target. This allows the lineage dependency relationship between the data source, data table, fields, etc., to be obtained, avoiding reliance on incomplete and inaccurate metadata extraction, and ensuring the accuracy, completeness and reliability of the lineage relationship, so as to truly know the flow of data.

[0119] In this embodiment of the disclosure, a bloodline diagram or bloodline table can also be established and displayed through a request terminal to further improve the utilization rate of bloodline relationships, so as to facilitate timely location of data faults, ensure the normal operation of analysis tasks and data tasks, prevent the collapse of computing resources, and ensure the stable and secure operation of the system.

[0120] Furthermore, such as Figure 5 As shown, the fault location method of this disclosure includes the following steps:

[0121] Step S501: Generate a blood relationship diagram based on the blood relationship analysis request.

[0122] In this embodiment of the disclosure, a lineage diagram refers to a graphical way of displaying lineage relationships, which can be in various forms such as tree diagrams and flowcharts. A lineage diagram includes multiple tree nodes, process nodes, and edges between nodes. Nodes can be data sources, data tables, or fields, etc., and edges represent data transmission and lineage relationships between nodes, thereby allowing for an intuitive understanding of the flow of data.

[0123] Furthermore, the root node of the tree diagram can be a data source, the child nodes can be data tables, and the next level node of the child node can be various fields, making it convenient to view the lineage relationship between nodes level by level.

[0124] Furthermore, blood relations can also be displayed as a blood relation table, where blood relations are presented in tabular form, with rows or columns representing data sources, data tables, or fields, and the corresponding table values ​​representing the blood relations between rows or columns.

[0125] Step S502: Display the bloodline diagram through the requesting terminal; wherein the bloodline diagram includes abnormal nodes.

[0126] In this embodiment of the disclosure, the bloodline diagram can be generated synchronously during the parsing process of the relationship parser, thereby enabling timely detection and analysis of faults for repair.

[0127] Step S503: In response to touch on the abnormal node, display the node data of the abnormal node.

[0128] In this embodiment of the disclosure, the user can touch the abnormal node to obtain the node data of the abnormal node, locate the problematic data or problematic process, and repair it.

[0129] In this embodiment of the disclosure, by using the fault location method of this disclosure and displaying the lineage relationship, users can grasp the data flow process, understand and analyze the data propagation path and dependencies, making the data lineage more intuitive and the evolution process easier to understand, facilitating the comprehensive tracing and in-depth utilization of the data lineage, improving the data understandability and analysis efficiency, and at the same time, timely locating and repairing fault problems, preventing analysis blockage, ensuring the normal execution of other data tasks, ensuring the server can flexibly respond to format analysis requests while the big data system operates stably, and making it convenient for users to obtain various lineage relationships.

[0130] In this embodiment of the disclosure, the kinship analysis server can also respond to user query requests and display the stored kinship relationships to the user, avoiding redundant analysis of kinship relationships, reducing analysis costs, and thereby improving the effective utilization of computing resources, such as... Figure 6 As shown, the method for determining blood relations disclosed herein includes the following steps:

[0131] Step S601: Receive one or more blood relationship query requests from the requesting terminals; wherein the blood relationship query request includes a query target.

[0132] Step S602: Based on the query target, search the memory of the bloodline server to determine the offset of the analysis target corresponding to the query target.

[0133] In this embodiment of the disclosure, the key-value pairs of blood relations can be stored in the memory or disk of the blood relation analysis server.

[0134] Step S603: Locate the storage location corresponding to the offset in the disk of the blood relationship server.

[0135] Step S604: Read the key-value pairs from the storage location to obtain the target blood relationship in the blood relationship query request.

[0136] Step S605: Perform a security verification on the target blood relationship and determine whether the verification result is successful. If yes, proceed to step S606; otherwise, proceed to step S607.

[0137] In this embodiment, the lineage analysis server pre-stores the confidentiality levels of various data. The security verification result is determined by judging the confidentiality levels of the upstream and downstream of the target lineage relationship. If the confidentiality level of the upstream lineage relationship is higher than that of the downstream lineage relationship, the security verification result is determined to be successful. If the confidentiality level of the upstream lineage relationship is lower than that of the downstream lineage relationship, it indicates that the downstream data used confidential data from the upstream data, posing a data leakage risk, and the security verification result is determined to be unsuccessful.

[0138] Furthermore, the security verification in step S605 can be omitted, and the security verification can be placed in the initial data collection method of this disclosure, for example, between steps S304-S305 and / or between steps S307-S308, so as to avoid the risk of data leakage during the data collection stage.

[0139] Step S606: Display the target blood relationship through the requesting terminal.

[0140] In this embodiment of the disclosure, users can also perform operations such as filtering and screening on the analysis source, analysis table, and analysis fields of blood relations to obtain blood relations that meet actual usage needs.

[0141] Step S607: Reject the blood relationship query request.

[0142] In the embodiments of this disclosure, the blood relationship query method can respond to various blood relationship query requests, thereby effectively utilizing historical data to quickly respond to user query, filtering, and screening requests, making it easier for users to view and improving the user experience.

[0143] In this embodiment of the disclosure, the method for managing the lineage of transaction data can obtain various data sources, data tables, fields, and their interrelationships, including parent-child relationships between data sources or data tables, dependency relationships between data sources or data tables, and concrete relationships between data tables and fields. In the face of a strict regulatory environment, the lineage of upstream and downstream data can be quickly controlled, and lineage relationships can be audited to identify those with potential leakage risks. This achieves secure data management and avoids data leakage risks. Simultaneously, it supports lineage analysis of various types of data sources, with high analysis and query efficiency, preventing analysis blockages, accurately locating and recovering from faults, ensuring the efficient operation of big data, improving the efficiency and accuracy of lineage analysis, avoiding obstacles caused by local anomalies, and ensuring the stable operation of the big data system.

[0144] Figure 7 This is a schematic diagram of a management system for the lineage of transaction data according to an embodiment of this disclosure, such as... Figure 7 As shown, the transaction data lineage management system 700 disclosed herein includes: a data acquisition layer 701, an analysis layer 702, a management layer 703, and an application layer 704, wherein:

[0145] The acquisition layer 701 includes an acquisition interface 7011, which is equipped with a receiving module, a data processing module, and an acquisition module. The receiving module is used to receive a lineage analysis request for transaction data, wherein the lineage analysis request includes one or more data source types. The data processing module is used to determine the acquisition scheme for the transaction data based on the data source type. The acquisition module is used to scan each of the data sources and, according to the acquisition scheme corresponding to the data source, offline acquire the initial data of the lineage analysis request, wherein the initial data includes a data source identifier.

[0146] The analysis layer 702 includes multiple relation parsers 7021, each relation parser 7021 including multiple parsing threads 70211 and aggregation threads 70212. The parsing thread 70211 includes an extraction module, which is used to split the initial data according to the data source identifier and extract the lineage relationship of the transaction data.

[0147] The management layer 703 includes memory and disk. The disk is used to store lineage key-value pairs, and the memory is used to store the offset of the key of the lineage key-value pair in order to locate the storage location of the lineage key-value pair on the disk.

[0148] The application layer 704 includes a graph database and a WebUI (Website User Interface). In response to a lineage query request, the target lineage can be read from the management layer 703, rendered using the graph database and WebUI, and displayed through the requesting terminal.

[0149] In this embodiment of the disclosure, the bloodline management system for transaction data enables rapid collection, extraction, and display of bloodline relationships, reduces network bandwidth and server resource consumption, improves analysis and processing efficiency, flexibly supports various types of data sources, improves missing fields, enhances the accuracy and completeness of bloodline relationship analysis, and the visualized bloodline relationship display facilitates data analysis for users, provides users with accurate decision-making basis, and enhances the value of data utilization.

[0150] Exemplary embodiments of this disclosure also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of this disclosure.

[0151] Exemplary embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to embodiments of this disclosure.

[0152] Exemplary embodiments of this disclosure also provide a computer program product, including a computer program, wherein, when executed by a processor of a computer, the computer program is used to cause the computer to perform a method according to an embodiment of this disclosure.

[0153] refer to Figure 8 The present invention describes a structural block diagram of an electronic device 800 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0154] like Figure 8As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0155] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, output unit 807, storage unit 808, and communication unit 809. Input unit 806 can be any type of device capable of inputting information to electronic device 800. Input unit 806 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 807 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 804 may include, but is not limited to, disk and optical disk. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMa8 devices, cellular communication devices, and / or the like.

[0156] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above. For example, in some embodiments, Figures 1 to 6 The method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 800 via ROM 802 and / or communication unit 809. In some embodiments, computing unit 801 can be configured to execute by any other suitable means (e.g., by means of firmware). Figures 1 to 6 The method.

[0157] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0158] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0159] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0161] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0162] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

Claims

1. A method for managing the lineage of transaction data, characterized in that, include: Receive a lineage analysis request for transaction data; wherein, the lineage analysis request includes data source types of one or more data sources; Determine the data collection plan for the transaction data based on the data source type; Scan each of the data sources and collect the initial data for the blood relationship analysis request offline according to the collection scheme corresponding to the data source; wherein, the initial data includes the data source identifier; The initial data is split according to the data source identifier, the lineage of the transaction data is extracted, and the lineage is displayed. The data source types include relational databases and data platforms; determining the transaction data collection scheme based on the data source type includes: Obtain the data source type of the bloodline analysis request, and determine whether the data source type is a relational database; When the data source type is a relational database, the collection method is determined to be parsing collection; The request statement of the blood relationship analysis request is parsed to generate a collection plan including the collection type, collection time and collection content; Also includes: When the data source type is a non-relational database, the data collection method is determined to be data collection via embedded points. Based on the platform identifier of the bloodline analysis request, the calling interface address is determined, and a collection scheme including the calling interface address, collection type, collection time, and collection content is generated.

2. The management method as described in claim 1, characterized in that, The step of collecting initial data for the kinship analysis request offline according to the collection scheme corresponding to the data source includes: Determine whether the acquisition method of the acquisition scheme is parsing acquisition. If the acquisition method is parsing acquisition, access the data source address using the acquisition interface. Based on the collection type, collection time, and collection content, data is collected and output according to the output template. The output data is then converted into a preset format to obtain the initial data.

3. The management method as described in claim 2, characterized in that, The step of collecting and outputting data according to the collection type, collection time, and collection content, and in accordance with the output template, includes: When the collected content consists of field keywords and field attribute values, the field attribute values ​​corresponding to the field keywords are collected as the output data; or, When the collected content consists of data content keywords and field attribute values, according to a preset ratio, a number of comparison data entries equal to the preset ratio are filtered from the data source. The data content keywords are matched with the comparison data to determine the target field containing the data content keywords. The field attribute values ​​of the target field are then collected as the output data.

4. The management method as described in claim 2, characterized in that, When the data acquisition method is embedded point acquisition, it also includes: Based on the API call address, the data collection interface is connected to the API call interface; The acquisition interface receives the output data returned by the calling interface according to the acquisition type, the acquisition time, and the acquisition content, and converts it into a preset format to obtain the initial data.

5. The management method as described in claim 1, characterized in that, The bloodline analysis request further includes an analysis target, which includes an analysis source, an analysis table, and / or an analysis field; the step of streamlining the initial data according to the data source identifier and extracting the bloodline relationship of the transaction data includes: Based on the database identifier or platform identifier, the initial data is distributed to different relation resolvers; In response to the analysis objective, the relation parser extracts the initial data of the analysis source, the analysis table, and / or the analysis field from the initial data; Extract the lineage of the initial data of the analysis source, the analysis table, and / or the analysis field.

6. The management method as described in claim 5, characterized in that, The extraction of the lineage of the initial data of the analysis source, the analysis table, and / or the analysis field includes: The analysis fields include username and transaction location. Taking the field attribute value of the username as the center, the lineage relationship between the field attribute value of the transaction location corresponding to the username is extracted. or, Extract the lineage relationship between the table name and the analysis fields of the analysis table; or, The analysis fields include transaction account, transaction amount, and transaction time. Based on the field attribute values ​​of transaction amount and transaction time corresponding to the transaction account, the lineage of the transaction process is extracted.

7. The management method as described in claim 5, characterized in that, Also includes: Receive one or more bloodline query requests from requesting terminals; wherein, the bloodline query request includes a query target; Based on the query target, search the memory of the bloodline server to determine the offset of the analysis target corresponding to the query target; The offset is used to locate the storage location corresponding to the offset in the disk of the blood relationship server, and the key-value pair of the storage location is read to obtain the target blood relationship of the blood relationship query request. The target blood relationship is displayed through the requesting terminal.

8. A management system for the lineage of transaction data, characterized in that, include: A receiving module is used to receive a lineage analysis request for transaction data; wherein, the lineage analysis request includes one or more data source types; The data processing module is used to determine the data collection scheme for the transaction data based on the data source type. The data acquisition module is used to scan each of the data sources and, according to the acquisition scheme corresponding to the data source, to collect the initial data of the blood relationship analysis request offline; wherein, the initial data includes the data source identifier; The display module is used to split the initial data according to the data source identifier, extract the lineage relationship of the transaction data, and display the lineage relationship; The data source types include relational databases and data platforms; determining the transaction data collection scheme based on the data source type includes: Obtain the data source type of the bloodline analysis request, and determine whether the data source type is a relational database; When the data source type is a relational database, the collection method is determined to be parsing collection; The request statement of the blood relationship analysis request is parsed to generate a collection plan including the collection type, collection time and collection content; Also includes: When the data source type is a non-relational database, the data collection method is determined to be data collection via embedded points. Based on the platform identifier of the bloodline analysis request, the calling interface address is determined, and a collection scheme including the calling interface address, collection type, collection time, and collection content is generated.

9. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform a method for managing the lineage of transaction data according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method for managing the lineage of transaction data according to any one of claims 1-7.

Citation Information

Patent Citations

  • Metadata management method and apparatus, computer device and storage medium

    CN112182045A

  • Data blood relationship analysis method, computer device and storage medium

    CN112559493A

  • Metadata acquisition method and device, equipment and medium

    CN115858548A