Data cleaning system and method based on blood relationship network

The data cleaning system using a graph database to manage data lineage networks automatically identifies and removes dirty data, improving data quality and analysis accuracy while enhancing traceability and compliance.

CN120316098APending Publication Date: 2025-07-15SHANGHAI QIYU INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510163318.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing data blood tracing method is difficult to apply when processing massive data in the face of distributed computing and storage frameworks, resulting in complex data processing processes and lack of automation and dynamic update capabilities, which affects data quality and analysis accuracy.

Method used

By building a blood relationship network based on the graph database, nodes represent data tables or fields, edges represent relationships between nodes, use preset attribute policies and extension policies to identify and delete dirty data, and dynamically update data dependencies.

Benefits of technology

It realizes automatic and accurate data cleaning, improves data quality and analysis accuracy, enhances data traceability and compliance, reduces manual intervention and errors, and improves data governance efficiency and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316098A_ABST
    Figure CN120316098A_ABST
Patent Text Reader

Abstract

The invention relates to a data cleaning system and method based on a blood relationship network, electronic equipment, a computer readable medium and a computer program product. The method comprises the steps that a blood relationship network of data is constructed based on a graph database, in the blood relationship network, nodes represent data tables or fields, and edges represent blood relationships among the nodes; determining a target node to be cleaned in the blood relationship network; obtaining all driving edges of the target node; judging the driving edges one by one through a preset attribute strategy and an expansion strategy; and when the driving edge does not meet a preset attribute strategy and an expansion strategy, deleting the driving edge to clean data in the blood relationship network. According to the method, dirty data can be automatically and accurately identified and deleted, data quality and analysis accuracy are improved, data traceability and compliance are enhanced, manual intervention and errors are reduced, and data governance efficiency and reliability are comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer information processing. Specifically, it relates to a data cleaning system, method, electronic device, computer-readable medium, and computer program product based on a blood relationship network. Background Art

[0002] Data Lineage refers to the tracking and description of data throughout its entire flow from its source to its destination. It reflects the path of data generation, transmission, processing, transformation, and use, as well as all change records. Data Lineage is typically used to understand where data comes from (source), how data is processed and transformed (process), and where data ultimately goes (destination).

[0003] Early data lineage tools mainly relied on manual documentation and static charts, and depended on the logging and code parsing of ETL tools for lineage tracking, lacking automation and dynamic update capabilities. With the rise of big data technologies (such as Hadoop, Spark, etc.), the demand for data lineage has been increasing. As the importance of data-driven decision-making has become increasingly prominent, the application scope of data lineage has expanded from traditional data governance to fields such as real-time data analysis, data compliance, data security, and privacy protection. Data lineage is not only used to track data flows but also to support scenarios such as data quality management, data impact analysis, and root cause analysis. Although there are currently data lineage tools integrated into big data platforms to support more complex computing logics and the tracking of data flows, as enterprises begin to use distributed computing and storage frameworks to process massive amounts of data, this has further led to the complication of data processing processes, making traditional lineage tracking methods no longer applicable.

[0004] Therefore, there is a need for a new data cleaning system, method, electronic device, computer-readable medium, and computer program product based on a blood relationship network.

[0005] The above information disclosed in the Background Art section is only used to enhance the understanding of the background of this application. Therefore, it may include information that does not constitute prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] In view of this, this application provides a data cleaning system, method, electronic device, computer-readable medium, and computer program product based on a blood relationship network, which can automatically and accurately identify and delete dirty data, improve data quality and analysis accuracy, while enhancing data traceability and compliance, reducing manual intervention and errors, and comprehensively improving data governance efficiency and reliability.

[0007] Other features and advantages of this application will become apparent through the following detailed description, or will be partially learned through the practice of this application.

[0008] According to one aspect of the present application, a data cleaning method based on a blood relationship network is proposed. The method includes: constructing a blood relationship network of data based on a graph database. In the blood relationship network, nodes represent data tables or fields, and edges represent the blood relationships between nodes; determining target nodes to be cleaned in the blood relationship network; obtaining all driving edges of the target nodes; judging each of the driving edges one by one through a preset attribute policy and an extension policy; when the driving edges do not meet the preset attribute policy and extension policy, deleting the driving edges to clean the data in the blood relationship network.

[0009] Optionally, it further includes: when updated data is added to the blood relationship network, determining nodes in the blood relationship network according to the updated data; generating temporary edges of the nodes; judging the temporary edges according to the attribute policy and the extension policy; when the temporary edges meet the preset attribute policy and extension policy, updating the updated data to the blood relationship network; when the temporary edges do not meet the preset attribute policy and extension policy, deleting the temporary edges.

[0010] Optionally, generating the temporary edges of the nodes includes: obtaining operation information of the updated data; generating the temporary edges of the nodes according to the operation information.

[0011] Optionally, judging the temporary edges according to the attribute policy and the extension policy includes: obtaining all driving edges corresponding to the nodes; judging the temporary edges based on the attribute policy and the extension policy of all the driving edges.

[0012] Optionally, constructing a blood relationship network of data based on a graph database includes: obtaining data and its corresponding operation information from a data source, where the operation information includes an operation time, an operation object, and an operation content; performing blood relationship registration on the data to generate nodes in the blood relationship network; parsing the operation information to generate edges between nodes in the blood relationship network, where the edges include out-edges and in-edges.

[0013] Optionally, performing blood relationship registration on the data to generate nodes in the blood relationship network includes: when the data corresponds to an existing node in the blood relationship network, taking the data as updated data; when the data does not correspond to an existing node in the blood relationship network, registering the data as a new node in the blood relationship network.

[0014] Optionally, parsing the operation information to generate edges between nodes in the blood relationship network includes: parsing the operation information to obtain the processing relationship between an operation and data; taking the edge corresponding to the operation that can change the data as an out-edge; taking the next-level operation that the data flows to as an in-edge.

[0015] Optionally, parse the operation information to obtain the processing relationship between operations and data, including: parsing the operation information based on the SQL language to obtain the processing relationship between operations and data.

[0016] Optionally, obtain all the driving edges of the target node, including: obtaining all the incoming edges corresponding to the target node and using them as the driving edges.

[0017] Optionally, discriminate each of the driving edges one by one through a preset attribute policy and extension policy, including: matching and discriminating the generation date of the driving edge with the system scheduling period; and / or matching and discriminating the task information of the driving edge with the system scheduling period.

[0018] According to one aspect of the present application, a data cleaning system based on a blood relationship network is proposed. The system includes: a network module for constructing a blood relationship network of data based on a graph database, where in the blood relationship network, nodes represent data tables or fields, and edges represent the blood relationship between nodes; a node module for determining a target node to be cleaned in the blood relationship network; an edge module for obtaining all the driving edges of the target node; a discrimination module for discriminating each of the driving edges one by one through a preset attribute policy and extension policy; and a cleaning module for deleting the driving edges to clean the data in the blood relationship network when the driving edges do not meet the preset attribute policy and extension policy.

[0019] According to one aspect of the present application, an electronic device is proposed. The electronic device includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.

[0020] According to one aspect of the present application, a computer-readable medium is proposed, on which a computer program is stored, and when the program is executed by a processor, the method as described above is implemented.

[0021] According to one aspect of the present application, a computer program product is proposed, including: a computer program / instructions, and when the computer program / instructions are executed by a processor, the method as described above is implemented.

[0022] A data cleaning system, method, electronic device, computer-readable medium, and computer program product based on a blood relationship network according to the present application construct a blood relationship network of data based on a graph database. In the blood relationship network, nodes represent data tables or fields, and edges represent the blood relationships between nodes. Determine target nodes to be cleaned in the blood relationship network; obtain all driving edges of the target nodes; discriminate each driving edge one by one through a preset attribute policy and extension policy; when the driving edge does not meet the preset attribute policy and extension policy, delete the driving edge to clean the data in the blood relationship network, which can automatically and accurately identify and delete dirty data, improve data quality and analysis accuracy, enhance data traceability and compliance at the same time, reduce manual intervention and errors, and comprehensively improve data governance efficiency and reliability.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other objectives, features, and advantages of the present application will become more apparent. The following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is a flowchart of a data cleaning method based on a blood relationship network shown according to an exemplary embodiment.

[0026] Figure 2 is a schematic diagram of a data cleaning method based on a blood relationship network shown according to another exemplary embodiment.

[0027] Figure 3 is a flowchart of a data cleaning method based on a blood relationship network shown according to another exemplary embodiment.

[0028] Figure 4 is a schematic diagram of a data cleaning method based on a blood relationship network shown according to another exemplary embodiment.

[0029] Figure 5 is a schematic diagram of a data cleaning method based on a blood relationship network shown according to another exemplary embodiment.

[0030] Figure 6 is a block diagram of a data cleaning system based on a blood relationship network shown according to another exemplary embodiment.

[0031] Figure 7It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0032] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repeated description will be omitted.

[0033] Figure 1 It is a flowchart of a data cleaning method based on a blood relationship network shown according to an exemplary embodiment. The data cleaning method 10 based on the blood relationship network at least includes steps S102 to S110.

[0034] As Figure 1 shown, in S102, a blood relationship network of data is constructed based on a graph database, wherein, in the blood relationship network, nodes represent data tables or fields, and edges represent the blood relationships between the nodes.

[0035] In one embodiment, for example, data and its corresponding operation information can be obtained from a data source, the operation information includes an operation time, an operation object, and an operation content; blood relationship registration is performed on the data to generate nodes in the blood relationship network; the operation information is parsed to generate edges between the nodes in the blood relationship network, wherein the edges include: out-edges and in-edges.

[0036] More specifically, in the process of data blood relationship tracking, it is necessary to obtain data and its corresponding operation information from a data source. These operation information may include: an operation time (the generation or modification time of the data), an operation object (the target data table or field of the operation), and an operation content (how the operation is performed, such as an SQL query or an ETL step). The data blood relationship network is a network constructed by a graph database, where the nodes represent data tables or fields, and the edges represent the dependencies or blood relationships between these data entities. This network shows the flow and transformation process of data from the source to the final destination.

[0037] Among them, a graph database is a database specially used for processing graph-structured data, and its core elements are nodes and edges. Different from traditional relational databases, graph databases are particularly suitable for representing and storing complex relationships, and are usually used in scenarios such as data blood relationship tracking and social network analysis.

[0038] Figure 2 It is a schematic diagram of a data cleaning method based on a blood relationship network shown according to another exemplary embodiment. Figure 2Taking an actual scenario as an example, the process of building a lineage network in a production system is specifically illustrated. Data is obtained from an external system, and the data can be registered for lineage or used as updated data to supplement existing nodes.

[0039] Among them, for example, when the data corresponds to an existing node in the lineage network, the data is used as updated data; when the data does not correspond to an existing node in the lineage network, the data is registered as a new node in the lineage network. The data can also be processed by a driving model and sent to a distribution module.

[0040] Operation information can also be obtained from the external system. The operation information is the operation information on the data. The operation is processed through a reporting interface and then transferred to an SQL parsing module. More specifically, for example, the operation information can be parsed based on the SQL language to obtain the processing relationship between the operation and the data.

[0041] The processed operation data and data are converted into nodes and edges. The information related to the nodes and edges is sent to the distribution module, and then after being assembled by an assembly module, it is transferred to the lineage network in the storage module for storage.

[0042] In a specific embodiment, for example, the operation information can also be parsed to obtain the processing relationship between the operation and the data; the edge corresponding to the operation that can change the data is used as the out-edge; the next-level operation to which the data flows is used as the in-edge. When an operation (such as insert, update, delete, etc.) changes the data, these operations will be represented as out-edges between nodes, and the next-level processing operation of the data will become the in-edge.

[0043] In one embodiment, the interface of the external query service can directly call the data in the lineage network in the storage module for use.

[0044] Suppose there is a data table Table_A, and a new data table Table_B is generated through SQL operations (such as JOIN or SELECT).

[0045] In the lineage network: both Table_A and Table_B will be stored as nodes in the graph database. The data transfer relationship from Table_A to Table_B will be represented by an edge, and the specific conversion logic will be stored as an attribute of this edge in the network.

[0046] When a data engineer uses an SQL query to select certain fields from Table_A to generate Table_B, the operation information may be:

[0047] Operation time: September 1, 2023.

[0048] Operation object: Table_A.

[0049] Operation content: SELECT col1, col2 FROM Table_A.

[0050] If the data in Table_A is operated to generate Table_B, then in the lineage network: both Table_A and Table_B exist as nodes, and the relationship between Table_A and Table_B can be represented by an out-edge and an in-edge, that is, the data of Table_A flows to Table_B.

[0051] In S104, determine the target node to be cleaned in the lineage network.

[0052] In S106, obtain all the driving edges of the target node. Obtain all the in-edges corresponding to the target node and use them as the driving edges. Data cleaning refers to identifying and deleting inaccurate or invalid data dependencies to improve the accuracy of data analysis. The lineage network can identify dirty data through the attributes of nodes and edges (such as operation time, task information, etc.).

[0053] Among them, the driving edge is the in-edge connecting the target node, indicating the upstream source of the data. By matching the attributes of the driving edge (such as generation date, scheduling period, etc.), it can be judged whether these edges are dirty data.

[0054] In S108, discriminate each driving edge one by one through a preset attribute policy and extension policy. For example, the generation date of the driving edge can be matched and discriminated with the system scheduling period; for another example, the task information of the driving edge can be matched and discriminated with the system scheduling period.

[0055] More specifically, in an actual application, assuming that the data of Table_C is generated by Table_A and Table_B, if it is found that the data of Table_A is expired or non-compliant within the scheduling period, the system will judge this out-edge as dirty data and delete it, thereby cleaning the data relationship that does not meet the preset conditions.

[0056] More specifically, when it is determined that a certain driving edge does not meet the preset attribute policy and extension policy, this edge can be deleted, thereby cleaning the dirty data in the data lineage network. Doing so can ensure that there are only valid data dependencies in the data stream and prevent dirty data from affecting the downstream analysis or processing results.

[0057] If the system finds that a certain data record of Table_A does not match the scheduling period, the out-edge between Table_A and Table_C can be deleted to prevent this invalid data from continuing to affect the data analysis of Table_C.

[0058] In S110, when the driving edge does not meet the preset attribute policy and extension policy, the driving edge is deleted to clean the data in the lineage network. Figure 3 It is a schematic diagram of a data cleaning method based on a lineage network shown according to another exemplary embodiment. Figure 3 Take a specific application to illustrate the data cleaning process. In data cleaning, in the preprocessing of the storage module, the target lineage node (such as table X or field) can be used as the driving node, and all incoming edges (i.e., driving edges) of this node can be searched backward. Analyze these driving edges and make judgments based on the attributes and extension policies on the edges. The judgment criteria include scheduling frequency, data production date, etc. If the attribute of a certain edge does not meet the preset conditions (such as data expiration), then delete this edge.

[0059] As Figure 3 shown, make judgments based on the scheduling frequency and lineage production date: for example, the normal lineage date is October 31st, but the data date of a certain edge is October 29th. If the task information of this edge shows that its scheduling frequency is once a day, then it is determined that this edge (such as from table B to table X) belongs to an expired lineage relationship, and this edge can be deleted.

[0060] Furthermore, in the graph database, perform a deletion operation to delete the driving edges identified as dirty data. The purpose of this is to break those invalid data dependency relationships that should not exist, thereby eliminating the negative impact on the downstream data analysis results. After deleting some dirty data edges, if some nodes become isolated nodes (i.e., have no valid upstream or downstream connections), then these isolated nodes can be further deleted to further streamline the lineage graph.

[0061] In this application, the data cleaning method based on the graph database can track the source and transformation path of data and precisely manage the data flow by constructing and maintaining the lineage network. By parsing the operation information to generate nodes and edges, the system can dynamically update and maintain the data dependency relationships. And through the judgment and deletion of the driving edges, the automated data cleaning is achieved, ensuring the accuracy and reliability of the data analysis results.

[0062] According to the data cleaning method based on the blood relationship network of the present application, a blood relationship network of data is constructed based on a graph database. Among them, in the blood relationship network, nodes represent data tables or fields, and edges represent the blood relationship between nodes; determining a target node to be cleaned in the blood relationship network; obtaining all driving edges of the target node; judging each driving edge one by one through a preset attribute policy and an extension policy; when the driving edge does not meet the preset attribute policy and extension policy, deleting the driving edge to clean the data in the blood relationship network can automatically and accurately identify and delete dirty data, improve data quality and analysis accuracy, enhance data traceability and compliance at the same time, reduce manual intervention and errors, and comprehensively improve data governance efficiency and reliability.

[0063] It should be clearly understood that this application describes how to form and use specific examples, but the principles of this application are not limited to any details of these examples. On the contrary, based on the teachings of the content disclosed in this application, these principles can be applied to many other embodiments.

[0064] Figure 4 It is a flowchart of a data cleaning method based on a blood relationship network shown according to another exemplary embodiment. Figure 4 The shown process 40 is for Figure 1 a supplementary description of the shown process.

[0065] As Figure 4 shown, in S402, when updated data is added to the blood relationship network, nodes in the blood relationship network are determined according to the updated data. Newly added data tables, fields, or their corresponding data records, these data are added to the existing data blood relationship network through a certain operation. Each data table or field can be regarded as a node in the blood relationship network. When there is new data update, it is first necessary to determine which node this data belongs to. If the node already exists, the status of the node is updated; if it is new data, a new node is created in the blood relationship network.

[0066] Suppose a new data table Table_X is generated through a certain ETL process. When updated data is added to the blood relationship network, it is first necessary to check whether there is already a node of Table_X in the network. If it exists, the node is updated; if it does not exist, a new Table_X node is added to the graph.

[0067] In S404, temporary edges of the node are generated. For example, operation information of the updated data is obtained; temporary edges of the node are generated according to the operation information. In the data blood relationship network, edges represent the dependency relationship between data. Temporary edges refer to the data dependency relationships temporarily generated after determining data nodes. These edges are not immediately written into the graph but need to be verified.

[0068] Suppose there is a data conversion operation from Table_A to Table_X, such as the SQL statement SELECT * FROM Table_A to generate Table_X. The system will generate a temporary edge for Table_X, indicating that Table_A is the data source of Table_X. However, this edge is temporary and needs further verification.

[0069] In S406, the temporary edge is discriminated according to the attribute policy and the extension policy. For example, all the driving edges corresponding to the node are obtained; the temporary edge is discriminated based on the attribute policy and the extension policy of all the driving edges.

[0070] Among them, the attribute policy can be, for example, a judgment policy based on the inherent attributes of the edge, such as the time of data generation, data source, operation type, etc. These attributes determine the validity of the data relationship.

[0071] The extension policy can be, for example, a more complex discrimination mechanism, which may include policy judgments based on system context or external rules. For example, the edge is evaluated according to rules such as scheduling frequency and data expiration period.

[0072] In S408, when the temporary edge meets the preset attribute policy and extension policy, the updated data is updated to the lineage network. When the temporary edge passes the verification of the attribute policy and the extension policy, this temporary edge can be converted into a formal lineage edge and updated to the lineage network. At this time, the states of the node and the edge are formally confirmed and written into the graph database to become part of the network.

[0073] Continuing with the above example, if the temporary edge of Table_X passes the verification of the scheduling frequency, the system will convert the temporary edge of Table_A -> Table_X into a formal edge and update it to the lineage network, indicating that the data of Table_X depends on Table_A.

[0074] In S410, when the temporary edge does not meet the preset attribute policy and extension policy, the temporary edge is deleted. Continuing with the above example, suppose the data generation time of Table_X is October 1, 2023, and the system's scheduling frequency is once a day. If the attribute of a certain temporary edge shows that its data generation time does not conform to the scheduling frequency (such as the data generation time is September 28), then this edge can be determined to be unqualified according to the attribute and extension policies.

[0075] Figure 5 It is a schematic diagram of a data cleaning method based on a lineage network shown according to another exemplary embodiment. As Figure 5As shown, in daily applications, external lineage data input is obtained, the target node is determined based on the lineage data, the edge is determined based on the target node, and then the attribute data on the edge is obtained. The attribute attributes may include date attributes and scheduling frequency, for example, and the local storage task scheduling frequency is called to make a judgment in combination with the date attribute. If it does not meet the policy, the edge is deleted to remove the dirty data relationship. For data that meets the policy normally, the edge can be inserted.

[0076] In this application, the data cleaning method ensures that the data nodes and edges newly added to the bloodline network are valid by generating temporary edges, verifying attributes and expansion strategies. If the temporary edge passes the verification, the system will update the node and edge to the graph database to form a complete bloodline network; if it fails, the temporary edge will be deleted to prevent dirty data from contaminating the entire bloodline relationship graph. This method can dynamically maintain the integrity and accuracy of the bloodline network and improve the quality of data governance.

[0077] Those skilled in the art will appreciate that all or part of the steps for implementing the above embodiments are implemented as a computer program executed by a CPU. When the computer program is executed by the CPU, the above functions defined by the above method provided in the present application are executed. The program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0078] In addition, it should be noted that the above figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, and are not intended to be limiting. It is easy to understand that the processes shown in the above figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.

[0079] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.

[0080] Figure 6 FIG. 1 is a block diagram of a data cleaning system based on a blood relationship network according to an exemplary embodiment. Figure 6 As shown, the data cleaning system 60 based on the blood relationship network includes: a network module 602, a node module 604, an edge module 606, a determination module 608, and a cleaning module 610. The data cleaning system 60 based on the blood relationship network may also include: an update module 612.

[0081] The network module 602 is used to construct a blood relationship network of data based on the graph database, wherein in the blood relationship network, nodes represent data tables or fields, and edges represent blood relationships between nodes;

[0082] The node module 604 is used to determine target nodes to be cleaned in the lineage network;

[0083] The edge module 606 is used to obtain all driving edges of the target nodes;

[0084] The discrimination module 608 is used to discriminate each of the driving edges one by one through a preset attribute policy and extension policy;

[0085] The cleaning module 610 is used to delete the driving edges to clean the data in the lineage network when the driving edges do not meet the preset attribute policy and extension policy.

[0086] The update module 612 is used to determine nodes in the lineage network according to the updated data when the updated data is added to the lineage network; generate temporary edges of the nodes; discriminate the temporary edges according to the attribute policy and extension policy; update the updated data to the lineage network when the temporary edges meet the preset attribute policy and extension policy; and delete the temporary edges when the temporary edges do not meet the preset attribute policy and extension policy.

[0087] According to the data cleaning system based on the lineage network of the present application, by constructing the lineage network of data based on a graph database, wherein in the lineage network, nodes represent data tables or fields, and edges represent the lineage relationships between nodes; determining target nodes to be cleaned in the lineage network; obtaining all driving edges of the target nodes; discriminating each of the driving edges one by one through a preset attribute policy and extension policy; and deleting the driving edges to clean the data in the lineage network when the driving edges do not meet the preset attribute policy and extension policy, it is possible to automatically and accurately identify and delete dirty data, improve data quality and analysis accuracy, enhance data traceability and compliance at the same time, reduce manual intervention and errors, and comprehensively improve data governance efficiency and reliability.

[0088] As Figure 7 shown, an embodiment of the present application provides an electronic device, including a processor 710, a memory 720, and a bus. Among them, the processor 710 and the memory 720 complete communication with each other through the bus 740;

[0089] The memory 720 is used to store a computer program;

[0090] When the processor 710 is used to execute the program stored on the memory 720, it implements the data cleaning method based on the lineage network in any of the above embodiments.

[0091] The communication interface 720 is used for communication between the above electronic device and other devices.

[0092] The memory 720 may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory 720 may also be at least one storage device located far from the aforementioned processor 710.

[0093] If the above method in this application is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the process of the above method embodiments in this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc.

[0094] The embodiments of this application provide a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the data cleaning method based on the blood relationship network in any of the above embodiments. For example, a blood relationship network of data can be constructed based on a graph database. In the blood relationship network, nodes represent data tables or fields, and edges represent the blood relationship between nodes; determine the target nodes to be cleaned in the blood relationship network; obtain all the driving edges of the target nodes; judge each driving edge one by one through a preset attribute policy and extension policy; when the driving edge does not meet the preset attribute policy and extension policy, delete the driving edge to clean the data in the blood relationship network.

[0095] The above specifically shows and describes the exemplary embodiments of this application. It should be understood that this application is not limited to the detailed structure, setting method or implementation method described here; on the contrary, this application is intended to cover various modifications and equivalent settings included in the spirit and scope of the appended claims.

Claims

1. A data cleaning system based on a blood relationship network, characterized in that, including: a network module, configured to construct a lineage network of data based on a graph database, wherein in the lineage network, nodes represent data tables or fields, and edges represent the lineage relationships between the nodes; a node module, configured to determine target nodes to be cleaned in the lineage network; an edge module, configured to obtain all driving edges of the target nodes; a discrimination module, configured to discriminate each of the driving edges one by one through a preset attribute policy and an extension policy; a cleaning module, configured to delete the driving edges to clean the data in the lineage network when the driving edges do not meet the preset attribute policy and extension policy.

2. A data cleaning method based on a blood relationship network, characterized in that, including: constructing a lineage network of data based on a graph database, wherein in the lineage network, nodes represent data tables or fields, and edges represent the lineage relationships between the nodes; determining target nodes to be cleaned in the lineage network; obtaining all driving edges of the target nodes; discriminating each of the driving edges one by one through a preset attribute policy and an extension policy; deleting the driving edges to clean the data in the lineage network when the driving edges do not meet the preset attribute policy and extension policy.

3. The method according to claim 2, wherein further including: when updated data is added to the lineage network, determining nodes in the lineage network according to the updated data; generating temporary edges of the nodes; discriminating the temporary edges according to the attribute policy and the extension policy; updating the updated data to the lineage network when the temporary edges meet the preset attribute policy and extension policy; deleting the temporary edges when the temporary edges do not meet the preset attribute policy and extension policy.

4. The method according to claim 3, wherein generating temporary edges of the nodes includes: obtaining operation information of the updated data; generating temporary edges of the nodes according to the operation information.

5. The method according to claim 3, wherein, discriminating the temporary edges according to the attribute policy and the extension policy includes: obtaining all driving edges corresponding to the nodes; discriminating the temporary edges based on the attribute policy and the extension policy of all the driving edges.

6. The method according to claim 2, wherein constructing a lineage network of data based on a graph database includes: obtaining data and its corresponding operation information from a data source, where the operation information includes an operation time, an operation object, and an operation content; performing lineage registration on the data to generate nodes in the lineage network; parsing the operation information to generate edges between nodes in the lineage network, where the edges include out-edges and in-edges.

7. The method according to claim 6, wherein performing lineage registration on the data to generate nodes in the lineage network includes: when the data corresponds to existing nodes in the lineage network, regarding the data as updated data; when the data does not correspond to existing nodes in the lineage network, registering the data as a new node in the lineage network.

8. The method according to claim 6, characterized in that, parsing the operation information to generate edges between nodes in the lineage network includes: parsing the operation information to obtain the processing relationship between the operation and the data; regarding the edges corresponding to the operations that can change the data as out-edges; regarding the next-level operation that the data flows to as an in-edge.

9. The method according to claim 8, wherein parsing the operation information to obtain the processing relationship between the operation and the data includes: Parse the operation information based on the SQL language to obtain the processing relationship between operations and data.

10. The method according to claim 2, characterized in that, Obtain all the driving edges of the target node, including: Obtain all the incoming edges corresponding to the target node and use them as the driving edges.

11. The method according to claim 2, characterized in that, Discriminate each of the driving edges one by one through a preset attribute policy and extension policy, including: Match and discriminate the generation date of the driving edge with the system scheduling period; and / or Match and discriminate the task information of the driving edge with the system scheduling period.

12. An electronic device, characterized in that, Include: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 2 to 11.

13. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 2 to 11.

14. A computer program product, characterized in that, Include computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 2 to 11 are implemented.

Citation Information

Cited By

  • Multi-source heterogeneous data adaptive conversion system and method driven by pluggable resolver

    CN121833693A