Data management method and device based on blood relationship
By conducting blood relationship analysis and path tracking on the target data tables in the data cascading system, errors and missing problems during data transmission are solved to ensure data integrity.
Patent Information
- Application Number
- CN202111165388.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-09-30
AI Technical Summary
In a data cascading system, errors or missing data may occur during transmission, and it is difficult for the prior art to effectively verify data integrity.
By performing blood ties analysis on the target data table, determine its original data table and its blood ties path, obtain the verification data table and perform matching verification to ensure the integrity of the data transmission process.
It realizes timely verification of data in the data cascading system, discovers and corrects errors or missing during transmission, and improves the accuracy of data transmission.
Smart Images

Figure CN113901049B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and in particular to a method and device for data management based on blood relationship. Background Art
[0002] In the data cascade system, data flows between data nodes at all levels. Massive and complex data is transformed through various processing and fusion processes, gradually converging from low-level data nodes to high-level data nodes.
[0003] In related technologies, data aggregation in a data cascade system may cause data errors or data loss during data transmission from low-level data nodes to high-level data nodes due to processing or transmission anomalies in some data nodes. Summary of the Invention
[0004] In view of this, the present application provides a data management method and device based on blood relationship, which are used to verify the data stored in the data cascade system, so as to discover abnormal data stored in the data cascade system.
[0005] Specifically, this application is implemented through the following technical solutions:
[0006] According to the first aspect of the present application, a data management method based on blood relationship is proposed, which is applied to any data node in a data cascade system, including:
[0007] Performing a lineage analysis on a target data table to be verified in any data node to determine an original data table corresponding to the target data table and an original data node to which the original data table belongs;
[0008] Obtaining a target lineage path corresponding to the original data table, where the starting point of the target lineage path is the original data node and the end point is another data node at the same level as any one of the data nodes;
[0009] Acquire, according to the target lineage path, a verification data table in which the original data table is aggregated in the other data nodes;
[0010] If the data in the target data table matches the data in the verification data table, it is determined that the data in the target data table and the verification data table are correct; if the data in the target data table does not match the data in the verification data table, it is determined that there are data errors in the target data table and / or the verification data table.
[0011] According to a second aspect of the present application, a data management device based on blood relationship is proposed, which is applied to any data node in a data cascade system, including:
[0012] a lineage analysis unit, configured to perform lineage analysis on a target data table to be verified in any data node, so as to determine an original data table corresponding to the target data table and an original data node to which the original data table belongs;
[0013] A lineage path acquisition unit, configured to acquire a target lineage path corresponding to the original data table, wherein the starting point of the target lineage path is the original data node and the end point is another data node at the same level as any one of the data nodes;
[0014] A verification data table acquisition unit, configured to acquire, according to the target lineage path, a verification data table in which the original data table is aggregated from the other data nodes;
[0015] A matching unit is configured to determine that the data in the target data table and the data in the verification data table are correct if the data in the target data table matches the data in the verification data table; and to determine that there are data errors in the target data table and / or the verification data table if the data in the target data table does not match the data in the verification data table.
[0016] According to a third aspect of the present application, an electronic device is provided, including:
[0017] processor;
[0018] a memory for storing processor-executable instructions;
[0019] The processor implements the method described in the embodiment of the first aspect above by running the executable instructions.
[0020] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in the embodiment of the first aspect are implemented.
[0021] It can be seen from the technical solution provided by the above application that the application performs a lineage relationship analysis on the target data table to be verified to determine its corresponding original data table, and searches the lineage path of the original data table to determine the verification data table obtained by aggregating each original data table in other data nodes at the level of the target data table. The target data table is matched and verified through the verification data table, so as to timely determine whether the data in the target data table may be missing or erroneous during the aggregation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0023] Figure 1is a network architecture diagram of a data cascade system according to an exemplary embodiment of the present application;
[0024] Figure 2 is a flow chart of a data management method based on blood relationship according to an exemplary embodiment of the present application;
[0025] Figure 3 This is a multi-party interaction flow chart of a data management method based on blood relationship according to an exemplary embodiment of the present application;
[0026] Figure 4 This is a bloodline path topology diagram shown according to an exemplary embodiment of the present application;
[0027] Figure 5 is a multi-party interaction flow chart illustrating another method for data management based on blood relationship according to an exemplary embodiment of the present application;
[0028] Figure 6 is a schematic diagram of an electronic device for data management based on blood relationship according to an exemplary embodiment of the present application;
[0029] Figure 7 It is a block diagram of a data management device based on blood relationship according to an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0030] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0031] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0032] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0033] Next, the embodiments of the present application are described in detail.
[0034] See also Figure 1 , is a network architecture diagram of a data cascade system shown in this application. Figure 1 As shown, the data cascade system has multiple levels of data nodes distributed in a tree-like topology, including primary database node 1-1, primary data center node 1, primary database node 1-2, secondary database node 1-1, secondary data center node 1, secondary database node 1-2, secondary database node 2-1, secondary data center node 2, secondary database node 2-2, tertiary database node 1-1, tertiary data center node 1, tertiary database node 1-2, tertiary database node 2-1, tertiary data center node 2, and tertiary database node 2-2. Each data node is used to schedule and store data tables. When any data node receives a data table from a data node at a lower level, it can store the data table, process the data, and then send it to the data node at the upper level.
[0035] Data lineage analysis is a core function in metadata management and data governance. Data lineage refers to the upstream and downstream links between data, similar to the blood relationships in human society, that form between data during its creation, processing, and dissolution. It describes the dependencies between different data tables or fields within different tables. By performing lineage analysis on data tables in a data cascade system, the data source of all data involved in that table can be found.
[0036] Figure 2 FIG. 1 is a flow chart showing a method for managing data based on blood relationship according to an exemplary embodiment of the present application. Figure 2 As shown, the method is applied to any data node in the data cascade system and may include the following steps:
[0037] Step 202: performing a lineage analysis on the target data table to be verified in any data node to determine the original data table corresponding to the target data table and the original data node to which the original data table belongs.
[0038] In the technical solution of the present application, the lineage analysis for the target data table can be obtained by calculating the similarity between the table name and table structure information of each data table in the data cascade system, wherein the table structure information may include field name, field type, field length, etc.; or it can also be analyzed through a variety of tools such as DRUID or spark's logicplan, and a corresponding data model is established to form a lineage relationship map of the data. In fact, the data lineage analysis methods in the prior art can all be used in the technical solution of the present application, and the present application does not impose any restrictions on this. By performing a lineage analysis on the target data table, the source of the data table can be traced back to determine the original data table corresponding to the target data table in the data cascade system. In addition to the table-level lineage analysis, any data node to which the target data table belongs can also perform a lineage analysis on each field in the target data table to determine the data source of each field in the target data table, and the present application does not impose any restrictions on this.
[0039] Step 204: Obtain a target lineage path corresponding to the original data table, where the starting point of the target lineage path is the original data node and the end point is another data node at the same level as any one of the data nodes.
[0040] In addition to being uploaded to any of the data nodes to aggregate the target data table to be verified, each original data table is also transmitted through other paths at the same time. That is to say, by analyzing the transmission path of the original data table in the data cascade system, other data tables aggregated from the original data tables of the target data table to be detected can be queried.
[0041] In one embodiment, after determining the original data table corresponding to the target data table, any data node to which the target data table belongs can send query requests for the original data table to the original data nodes to which each original data table belongs, so that each original data node can find its lineage path when converging to the upper level based on the information of the original data table, and return the found lineage path information to the any data node. The lineage path information includes multiple lineage paths with different destination directions starting from the original data node. After receiving the lineage path information returned by the original data node, any data node needs to filter the lineage paths therein, and obtain the lineage paths of the multiple lineage paths with the lineage path information as the end points of other data nodes at the same level as its own node as the target lineage path. Figure 1Taking the data cascade system shown as an example, if table b in tertiary database node 1-1 is one of the original data tables corresponding to table a in primary data center node 1 to be verified, then after receiving a query request from primary data center node 1, tertiary database node 1-1 can query the lineage path of table b based on the data information in table b. This includes three lineage paths: tertiary database node 1-1 - tertiary data center node 1 - secondary data center node 1 - primary data center node 1; tertiary database node 1-1 - secondary database node 1-1 - secondary data center node 1; and tertiary database node 1-1 - secondary database node 1-1 - primary database node 1-1. After receiving these three lineage paths returned by tertiary database node 1-1, primary data center node 1 can filter them and select the lineage path whose endpoint is another data node at the same level as itself, namely, tertiary database node 1-1 - secondary database node 1-1 - primary database node 1-1, as the target lineage path.
[0042] In another embodiment, the data node to which the original data table belongs may be at a lower level and have weak data management capabilities, and therefore cannot store data lineage information or analyze the lineage path of the original data table, and therefore needs to rely on the lineage analysis capabilities of other data nodes. After determining the original data table corresponding to the target data table, any data node to which the target data table belongs can send a lineage path query request for the original data table to other data nodes of the level to which it belongs, so that other data nodes of the level can retrieve the lineage paths of their respective local data tables, and compare the lineage path information of each local data table with the information of the original data table. If there is a lineage path of any local data table that includes content corresponding to each original data table, it can be determined that the local data table is obtained by aggregating the various original data tables, that is, the lineage path of the local data table is the target lineage path of the original data table. Still with Figure 1Taking the data cascade system shown as an example, when the target data table to be detected is data table a in the first-level data center node 1, and the original data table corresponding to data table a is table b in the third-level database node 1-1 and table c in the third-level database node 2-1, the first-level data center node 1 can send a lineage path query request about table b and table c to the first-level database node 1-1 and the first-level database node 1-2 at the same level. If the first-level database node 1-1 can find that there is a lineage path of a local data table corresponding to table b and table c by retrieving the lineage path of the local data table, then the lineage path is used as the target lineage path corresponding to the original data table. By retrieving the lineage path of the original data table from other data nodes at the same level as any data node to which the target data table belongs, the processing pressure of the original data nodes at the lower level can be reduced, and the situation where the data nodes at the lower level cannot perform lineage analysis on the data and cannot determine the verification data table corresponding to the target data table to be verified can be avoided.
[0043] Generally speaking, data lineage can be divided into two dimensions: table-level data lineage and field-level data lineage. Specifically, similar to the above-mentioned lineage analysis of the fields in the target data table, in addition to determining the association between tables from the perspective of the table, the lineage path of the original data table can also be determined from the perspective of the field. The lineage path query request sent by any data node to which the target data table belongs can contain the field information of the original data table. In the case of a lineage path query by the original data node, the original data node can find the lineage path of the corresponding field based on the field information; in the case of a lineage path query by other data nodes belonging to the same level as any data node to which the target data table belongs, the other data nodes can retrieve the lineage path of the field in the local data table to determine whether the lineage path of each field in the local data table corresponds to the target lineage path of the field information. In addition to the above-mentioned lineage path query directly based on the field information, the present application can also compare the fields in the data table corresponding to the table-level lineage path with the fields required by the target data table on the basis of the table-level lineage path query, and filter out the field-level lineage path from the table-level lineage path. For example, if only field A' exists in the original data table in the third-level database node 1-1, which corresponds to the target data table to be detected, and when querying the table-level lineage path, it is found that in one lineage path, the original data table is transmitted to the third-level data center node 1 to generate data table f, and in another table-level lineage path, the original data table is transmitted to the third-level data center node 1 to generate data table e. The original data node can compare the fields transmitted to data table f and data table e with field A' respectively. If the fields transmitted to data table e do not contain field A', it can be determined that the lineage path corresponding to data table e is not the field-level lineage path required by the target data table, and it can be excluded. By obtaining the field-level lineage path, the transmission path of the fields related to the target data table to be detected can be queried in a targeted manner, without considering the transmission of other irrelevant fields in the original data table, thereby achieving accurate screening of the lineage path.
[0044] Step 106: Obtain the verification data table that aggregates the original data table in the other data nodes according to the target lineage path.
[0045] In the case where the original data node analyzes the lineage path of the original data table, any data node to which the target data table belongs can determine the target lineage path based on the lineage path information returned by the received original data table. It can then determine the verification data table that aggregates the various original data tables in other data nodes at the end point of the target lineage path based on the target lineage path, and send a data synchronization request for the verification data table to the other data node, so that the other data node synchronizes the data in the verification data table to any data node to which the target data table to be verified belongs.
[0046] When the lineage path of the original data table is analyzed by other data nodes belonging to the same level as any data node, any data node to which the target data table belongs, upon receiving the target lineage path sent by the other data node, can use the local data table corresponding to the target lineage path as a verification data table and send a data synchronization request regarding the verification data table to the other data node, so that the other data node synchronizes the data in the verification data table to any data node to which the target data table to be verified belongs. Furthermore, after searching the lineage path of a field in the local data table to determine that there is a target lineage path corresponding to the field information, the other data table can directly and proactively synchronize the local data table corresponding to the target lineage path as a verification data table to any data node to which the target data table to be verified belongs.
[0047] Step 108: If the data in the target data table matches the data in the verification data table, it is determined that the data in the target data table and the verification data table are correct; if the data in the target data table does not match the data in the verification data table, it is determined that there are data errors in the target data table and / or the verification data table.
[0048] In one embodiment, upon receiving the verification data table, any data node to which the target data table to be verified belongs may arrange the data in both the verification data table and the target data table according to the primary key, so as to match the data in the verification data table with the data in the corresponding fields of the target data table based on the arrangement of the primary keys. If the data in the target data table and the verification data table match, the likelihood of both the target data table and the verification data table being simultaneously incorrect is negligible, and thus the data in both the target data table and the verification data table can be determined to be correct. If the data in the target data table and the verification data table do not match, then an error has occurred in at least one of the data in the target data table and the verification data table, and thus it can be determined that there is a data error in the target data table and / or the verification data table.
[0049] Furthermore, in Figure 1In the data cascade system shown, the data nodes of each level include at least one data center node, and each data center node is associated with at least one database node of this level. The database node is used to transmit the data table to the corresponding data center node and / or the upper-level database node, and the data center node is used to transmit the data table to the corresponding database node and / or the upper-level data center node. The database node can send data to the data center node associated with it and / or the upper-level database node associated with it. The data center node is used to summarize and manage the data of the database nodes associated with this level, that is, even if the data table is uploaded through multiple paths during the upload and aggregation process, it is not necessary to transmit it to multiple database nodes. Therefore, when any data node to which the target data to be detected belongs is any data center node, the other nodes to which the verification data table belongs are the database nodes corresponding to any data center node; when any data node is any database node, the other nodes are the data center nodes corresponding to any database node.
[0050] It can be seen from the technical solution provided by the present application that the present application performs a lineage analysis on the target data table to be verified to determine its corresponding original data table, and searches for the lineage path of the original data table to determine the verification data table obtained by aggregating each original data table in other data nodes at the same level as the target data table. By matching and verifying the target data table through the verification data table, it is possible to promptly determine whether the data in the target data table may be missing or erroneous during the aggregation process. Figure 3 Detailed description is given. Figure 3 A multi-party interaction flow chart of a data management method based on blood relationship is shown according to an exemplary embodiment of the present application. Figure 3 As shown, the interaction process of the primary data center node 1, the primary database node 1-1, the tertiary database node 1-1, and the tertiary database node 2-1 includes the following steps:
[0051] Step 301 : the primary data center node 1 performs a lineage analysis on the target data table a to be verified in the primary data center node 1 .
[0052] When it is necessary to manage and verify data table a in the first-level data center node 1, and data center node 1 can process data table a through, for example, the data lineage analysis method in the relevant technology, it is determined that the data in data table a comes from data table b in the third-level database node 1-1 and data table c in the third-level database node 2-1, that is, table b and table c are the original data tables of data table a, and the third-level database node 1-1 and the third-level database node 2-1 are the corresponding original data nodes.
[0053] In step 302a, the primary data center node 1 sends a lineage path query request about the data table b to the tertiary database node 1-1.
[0054] In step 302b, the primary data center node 1 sends a lineage path query request about the data table c to the tertiary database node 2-1.
[0055] After determining that data table b and data table c are original data tables of data table a to be detected, the primary data center node 1 may generate a query request about data table b and data table c.
[0056] If it is determined through field-level data lineage analysis that field A in data table a originates from field A' in data table b, and field B in data table a originates from field B' in data table c, then when querying the lineage path of the original data table, the lineage path can also be targeted at the field. Node 1 in the first-level data can send a lineage path query request containing field information related to field A' to node 1-1 in the third-level database, and send a lineage path query request containing field information related to field B' to node 2-1 in the third-level database.
[0057] In step 303a, the third-level database node 1-1 queries the lineage path of the data table b.
[0058] In step 303b, the third-level database node 2-1 queries the lineage path of the data table c.
[0059] After receiving the lineage path query request sent by the first-level data center node 1, the third-level database nodes 1-1 and 2-1 can query the lineage path of the data table or field when it converges to the upper level based on the data table information or field information contained in the query request.
[0060] For example, Table 1 is the lineage path information table obtained by querying the field information of field A' in data table b by the third-level database node 1-1, and Table 2 is the lineage path information table obtained by querying the field information of field B' in data table c by the third-level database node 2-1.
[0061]
[0062] Table 1
[0063]
[0064]
[0065] Table 2
[0066] By performing a lineage path query on the original data table, the upload path of the original data table in the data cascade system can be determined. Furthermore, those skilled in the art will appreciate that, although not shown in the table, performing a lineage path query on the data table not only determines the nodes that the data table passes through during transmission, but also provides information about the corresponding data tables generated at each data node that the data table passes through.
[0067] Furthermore, in addition to the lineage path information table, the lineage path of the original data can also be displayed in other forms. Figure 4 The figure shows a lineage path topology diagram according to an exemplary embodiment of the present application. As shown in the figure, for the third-level database node 1-1 and the third-level database node 2-1, there are two lineage paths respectively, where the solid line is the lineage path for generating the target data table to be tested, and the dotted line is the target lineage path.
[0068] In step 304a, the first-level data center node 1 receives the lineage path information about the data table b sent by the third-level database node 1-1.
[0069] In step 304b, the first-level data center node 1 receives the lineage path information about the data table c sent by the third-level database node 2-1.
[0070] After determining the lineage path information according to the query information, the third-level database node 1-1 and the third-level database node 2-1 can send the lineage path information tables such as Table 1 and Table 2 above to the first-level data center node 1.
[0071] Step 305: Level 1 data center node 1 determines the target lineage path and the corresponding verification data table.
[0072] After receiving the lineage path information, the first-level data center node 1 can analyze the multiple lineage paths contained in the lineage path information, take the lineage path whose end point is other data nodes at the same level as the first-level data center node as the target lineage path, and determine the corresponding verification data table based on the target lineage path of each original data table.
[0073] For example, Tables 1 and 2 each contain two lineage paths. By analyzing the endpoints of each lineage path, we can determine that lineage path ② in Table 1 is the target lineage path corresponding to data table b, and lineage path ② in Table 2 is the target lineage path corresponding to data table c. Data tables b and c converge on primary database node 1-1. Based on the information contained in the lineage path about the corresponding data tables generated at each data node passed through by data tables b and c, we can determine that data table d, where data tables b and c converge on primary database node 1-1, can be determined as the verification data table.
[0074] Step 306: Level 1 data center node 1 sends a data synchronization request to level 1 database node 1-1.
[0075] After determining that the verified data table is data table d, the primary data center node 1 generates a data synchronization request for data table d and sends the data synchronization request to the primary database node 1 - 1 .
[0076] Step 307: Level 1 data center node 1 receives the verification data table sent by database node 1-1.
[0077] After receiving the data synchronization request regarding the data table d sent by the data center node 1 , the primary database 1 - 1 may return the data table d to the primary data center node 1 .
[0078] Step 308: Level 1 data center node 1 matches the verification data table with the target data table to be verified.
[0079] After receiving Data Table D, Level 1 Data Center Node 1 can compare Data Table D with the data in Data Table A to be verified. If the data in Data Table D matches the data in Data Table A, it can be determined that the data in both Data Table D and Data Table A are correct. If the data in Data Table D does not match the data in Data Table A, it can be determined that the mismatch is due to an error in the data in at least one of Data Table D and Data Table A.
[0080] Furthermore, since the original data node is at a relatively low level, its data lineage storage and analysis capabilities may be relatively weak, so the lineage path of the original data table can be queried through other data nodes. Figure 4 FIG. 1 is a multi-party interaction flow chart of another data management method based on blood relationship according to an exemplary embodiment of the present application. Figure 5 As shown, the interaction process between the primary data center node 1 and the primary database node 1-1 includes the following steps:
[0081] Step 501 : the primary data center node 1 performs a lineage analysis on the target data table a to be verified in the primary data center node 1 .
[0082] When it is necessary to manage and verify data table a in the first-level data center node 1, and data center node 1 can process data table a through, for example, the data lineage analysis method in the relevant technology, it is determined that the data in data table a comes from data table b in the third-level database node 1-1 and data table c in the third-level database node 2-1, that is, table b and table c are the original data tables of data table a, and the third-level database node 1-1 and the third-level database node 2-1 are the corresponding original data nodes.
[0083] In step 502, the primary data center node 1 sends a lineage path query request about the data table b and the data table c to the primary database node 1-1.
[0084] After determining that Table B and Table C are the original data tables of Table A to be tested, Level 1 data center node 1 can generate a query request for Table B and Table C. It then sends a lineage path query request for Table B and Table C to Level 1 database node 1-1, which is at the same level and associated with the original data node.
[0085] Step 503: the primary database node 1-1 searches the lineage of each of its own local data tables.
[0086] After receiving the lineage path query request sent by the data center node 1, the first-level database node 1-1 can retrieve the lineage path of the local data table based on the data table information of data table b and data table c contained in the query request, and determine whether the lineage path of the local data table is associated with the data table information of data table b and data table c, that is, whether the data source of the local data table includes data table b and data table c.
[0087] Step 504: Level 1 data center node 1 receives the verification data table sent by level 1 database node 1-1.
[0088] In one embodiment, primary database node 1-1 may first determine the local data table lineage path associated with the data table information of data table b and data table c as the target lineage path and send the target lineage path to primary data center node 1. Primary data center node 1 then determines the local data table corresponding to the target lineage path as the verification data table and sends a data synchronization request regarding the verification data table to primary database node 1-1. After receiving the data synchronization request, primary database node 1-1 then sends the verification data table to primary data center node 1.
[0089] In another embodiment, after determining that there is a lineage path associated with the data table information of data table b and data table c, the first-level database node 1-1 can directly use the local data table corresponding to the lineage path as a verification data table, and directly send the verification data table to the first-level data center node 1.
[0090] Step 505 , the primary data center node 1 matches the verification data table with the target data table to be verified.
[0091] After receiving the local data table, primary data center node 1 can compare the local data table with the data in Data Table A to be verified. If the data in the local data table matches the data in Data Table A, it can be determined that the data in both the local data table and Data Table A are correct. If the data in the local data table does not match the data in Data Table A, it can be determined that the mismatch is due to an error in the data in at least one of the local data table and Data Table A.
[0092] Corresponding to the above method embodiment, this specification also provides an embodiment of a device.
[0093] Figure 6 This is a structural diagram of an electronic device for data management based on blood relationship according to an exemplary embodiment of the present application. Figure 6 At the hardware level, the electronic device includes a processor 602, an internal bus 604, a network interface 606, a memory 608, and a non-volatile memory 610. Of course, it may also include hardware required for other services. The processor 602 reads the corresponding computer program from the non-volatile memory 610 into the memory 608 and then runs it. Of course, in addition to software implementation, this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware. In other words, the execution body of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0094] Figure 7 FIG1 is a block diagram of a data management device based on blood relationship according to an exemplary embodiment of the present application. Figure 7 The device includes a blood relationship analysis unit 702, a blood relationship path acquisition unit 704, a verification data table acquisition unit 706, and a matching unit 708, wherein:
[0095] The lineage analysis unit 702 is configured to perform lineage analysis on the target data table to be verified in any data node to determine the original data table corresponding to the target data table and the original data node to which the original data table belongs.
[0096] The lineage path acquisition unit 704 is configured to acquire a target lineage path corresponding to the original data table, where the starting point of the target lineage path is the original data node and the end point is another data node at the same level as the any data node.
[0097] The verification data table acquisition unit 706 is configured to acquire, according to the target lineage path, a verification data table in which the original data table is aggregated from the other data nodes.
[0098] The matching unit 708 is configured to determine that the data in the target data table and the data in the verification data table are correct if the target data table matches the data in the verification data table; if the data in the target data table and the data in the verification data table do not match, determine that there is a data error in the target data table and / or the verification data table.
[0099] Optionally, obtaining the target lineage path corresponding to the original data table includes: sending a query request about the original data table to the original data node, so that the original data node searches for the lineage path of the original data table; receiving lineage path information returned by the original data node, the lineage path information including multiple lineage paths starting from the original data node; obtaining the target lineage path from the multiple lineage paths whose end points are other data nodes at the same level as any of the data nodes.
[0100] Optionally, sending a query request about the original data table to the original data node so that the original data node searches for the lineage path of the original data table includes: sending a query request containing field information of the original data table to the original data node so that the original data node searches for the lineage path of the field corresponding to the field information based on the field information.
[0101] Optionally, obtaining the verification data table that aggregates the original data table in the other data nodes according to the target lineage path includes: determining the verification data table that aggregates the original data table in the other data nodes according to the target lineage path; sending a data synchronization request regarding the verification data table to the other data nodes; and receiving the verification data table sent by the other data nodes.
[0102] Optionally, obtaining the target lineage path of the original data table includes: sending a query request about the original data table to other data nodes in the level where any data node is located, so that the other data nodes retrieve the lineage path of the local data table and determine the target lineage path corresponding to the original data table; receiving the target lineage path sent by the other data nodes.
[0103] Optionally, the sending of a query request regarding the original data table to other data nodes in the layer where any one of the data nodes is located, so that the other data nodes retrieve the lineage path of the local data table and determine the target lineage path corresponding to the original data table, includes: sending a query request containing field information of the original data table to other data nodes in the layer where any one of the data nodes is located, so that the other data nodes retrieve the lineage path of the field in the local data table and determine the target lineage path corresponding to the field information.
[0104] Optionally, obtaining the verification data table that aggregates the original data table in the other data nodes based on the target lineage path includes: determining the local data table corresponding to the target lineage path as the verification data table; sending a data synchronization request regarding the verification data table to the other data nodes; and receiving the verification data table sent by the other data nodes.
[0105] Optionally, the data cascade system has multiple levels of data nodes distributed according to a tree topology structure, wherein the data nodes at each level include at least one data center node and at least one database node corresponding to each of the data center nodes, the database node is used to transfer data tables to the corresponding data center node and / or the upper-level database node, and the data center node is used to transfer data tables to the corresponding database node and / or the upper-level data center node; when any of the data nodes is any of the data center nodes, the other nodes are the database nodes corresponding to the any of the data center nodes; when any of the data nodes is any of the database nodes, the other nodes are the data center nodes corresponding to the any of the database nodes.
[0106] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0107] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0108] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is further provided, such as a memory including instructions. The instructions may be executed by a processor of a kinship-based data management device to implement any of the methods described in the above embodiments. For example, the method may include:
[0109] Perform a lineage analysis on the target data table to be verified in any of the data nodes to determine the original data table corresponding to the target data table and the original data node to which the original data table belongs; obtain a target lineage path corresponding to the original data table, where the starting point of the target lineage path is the original data node and the end point is other data nodes at the same level as any of the data nodes; obtain a verification data table in the other data nodes that aggregates the original data table according to the target lineage path; if the target data table matches the data in the verification data table, determine that the data in the target data table and the verification data table are both correct; if the target data table does not match the data in the verification data table, determine that there are data errors in the target data table and / or the verification data table.
[0110] The non-temporary computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc., and this application does not limit this.
[0111] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A data management method based on blood relationship, characterized in that: Applied to any data node in a data cascade system, the method includes: Performing a lineage analysis on a target data table to be verified in any data node to determine an original data table corresponding to the target data table and an original data node to which the original data table belongs; Obtaining a target lineage path corresponding to the original data table, where the starting point of the target lineage path is the original data node and the end point is another data node at the same level as any one of the data nodes; Acquire, according to the target lineage path, a verification data table in which the original data table is aggregated in the other data nodes; If the data in the target data table matches the data in the verification data table, it is determined that the data in the target data table and the verification data table are both correct; if the data in the target data table does not match the data in the verification data table, it is determined that there are data errors in the target data table and / or the verification data table; Wherein, obtaining the target bloodline path corresponding to the original data table includes: Sending a query request about the original data table to other data nodes in the hierarchy where the any data node is located, so that the other data nodes retrieve the lineage path of the local data table and determine a target lineage path corresponding to the original data table; Receive the target lineage path sent by the other data nodes.
2. The method according to claim 1, characterized in that The step of obtaining the target bloodline path corresponding to the original data table includes: Sending a query request about the original data table to the original data node, so that the original data node searches for a lineage path of the original data table; Receiving lineage path information returned by the original data node, the lineage path information including multiple lineage paths starting from the original data node; Obtain a target lineage path from the plurality of lineage paths, the destination of which is another data node at the same level as the any data node.
3. The method according to claim 2, characterized in that The sending of a query request about the original data table to the original data node so that the original data node searches for a lineage path of the original data table includes: A query request containing the field information of the original data table is sent to the original data node, so that the original data node searches for a lineage path of a field corresponding to the field information according to the field information.
4. The method according to claim 2, characterized in that The step of obtaining the verification data table containing the original data table from the other data nodes according to the target lineage path includes: Determine, according to the target lineage path, a verification data table in the other data nodes that aggregates the original data table; Sending a data synchronization request regarding the verification data table to the other data nodes; Receive the verification data table sent by the other data nodes.
5. The method according to claim 1, characterized in that: The sending of a query request about the original data table to other data nodes in the hierarchy where any one of the data nodes is located, so that the other data nodes retrieve the lineage path of the local data table and determine the target lineage path corresponding to the original data table, includes: Send a query request containing the field information of the original data table to other data nodes in the level where any data node is located, so that the other data nodes retrieve the lineage path of the field in the local data table and determine the target lineage path corresponding to the field information.
6. The method according to claim 1, characterized in that The step of obtaining the verification data table containing the original data table from the other data nodes according to the target lineage path includes: Determine the local data table corresponding to the target bloodline path as the verification data table; Sending a data synchronization request regarding the verification data table to the other data nodes; Receive the verification data table sent by the other data nodes.
7. The method according to claim 1, characterized in that: The data cascade system is distributed with multiple levels of data nodes in a tree topology, wherein the data nodes at each level include at least one data center node and at least one database node corresponding to each of the data center nodes, the database node is used to transmit data tables to the corresponding data center node and / or the upper-level database node, and the data center node is used to transmit data tables to the corresponding database node and / or the upper-level data center node; When any data node is any data center node, other nodes at the same level as any data center node are database nodes corresponding to any data center node; When any data node is any database node, other nodes at the same level as the any database node are data center nodes corresponding to the any database node.
8. A data management device based on blood relationship, characterized in that: Applied to any data node in a data cascade system, the device comprises: a lineage analysis unit, configured to perform lineage analysis on a target data table to be verified in any data node, so as to determine an original data table corresponding to the target data table and an original data node to which the original data table belongs; A lineage path acquisition unit, configured to acquire a target lineage path corresponding to the original data table, wherein the starting point of the target lineage path is the original data node and the end point is another data node at the same level as any one of the data nodes; A verification data table acquisition unit, configured to acquire, according to the target lineage path, a verification data table in which the original data table is aggregated from the other data nodes; a matching unit, configured to determine that the data in the target data table and the verification data table are both correct if the data in the target data table matches the data in the verification data table; and to determine that there are data errors in the target data table and / or the verification data table if the data in the target data table does not match the data in the verification data table; Wherein, obtaining the target bloodline path corresponding to the original data table includes: Sending a query request about the original data table to other data nodes in the hierarchy where the any data node is located, so that the other data nodes retrieve the lineage path of the local data table and determine a target lineage path corresponding to the original data table; Receive the target lineage path sent by the other data nodes.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 7 by running the executable instructions.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Data blood relationship analysis method, device, equipment and system and readable storage medium
CN109582660A
Directory reporting method and device
CN113342816A