Data processing method, device and equipment
By obtaining and verifying the node data information in the data lineage, the problems of complex data flow and difficult traceability are solved, the reliability of data lineage is guaranteed, and the integrity and security of the data flow process are ensured.
Patent Information
- Application Number
- CN202011297031.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-18
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2040-11-18
AI Technical Summary
In a big data environment, the data flow process is complex and difficult to track and trace. The existing data lineage update mechanism is difficult to verify its effectiveness, making it difficult to ensure data security.
By obtaining the data information of each node in the data lineage to be verified, including the data volume and data type, verifying the data volume of the corresponding node, generating lineage completion information or alarm information, the reliability of data lineage is ensured.
The reliability of data lineage is guaranteed, and the integrity and security of the data flow process are ensured through data volume auditing and path auditing.
Smart Images

Figure CN114519060B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a data processing method, device and equipment. Background Art
[0002] In a big data environment, data rapidly flows and merges within systems (or between systems). The process from data source to data requester is more complex than in traditional data environments, making data more difficult to trace. Therefore, establishing data lineage is crucial for protecting data security. Data lineage describes the flow of data from source to target data by studying the processes of data generation, circulation, use, and fusion.
[0003] However, in existing networks, data flow is complex, and "nodes" like database tables and business systems are not stable. Nodes can increase or decrease with business changes, and the flow of data can also shift. To address this, data lineage systems generally establish update mechanisms, but their effectiveness can be difficult to verify. Summary of the Invention
[0004] The purpose of the present invention is to provide a data processing method, device and equipment to achieve the purpose of ensuring the reliability of data lineage.
[0005] To achieve the above objectives, an embodiment of the present invention provides a data processing method, comprising:
[0006] Obtaining data information of each node in the data lineage to be checked, wherein the data information includes data volume and data type;
[0007] According to each data type, verify the data amount of the corresponding first node and second node to obtain a first verification result;
[0008] The second node is a data source node of the first data in the first node, and the first data belongs to a data type currently being verified.
[0009] Optionally, verifying the data amount of the corresponding first node and second node according to each data type includes:
[0010] determining whether the data volume of the first data is the same as the data volume of the second data;
[0011] The second data is data transmitted from the second node to the first node, and the second data and the first data are of the same data type.
[0012] Optionally, the method further includes:
[0013] If the first verification result indicates that the data volume of the first data is different from the data volume of the second data, obtaining audit data of the first node;
[0014] Parsing the audit data of the first node;
[0015] According to the result of the analysis, if the third node includes other nodes except the second node, generating first lineage completion information; if the third node only includes the second node, generating first warning information;
[0016] The third node is a data source node of the first data.
[0017] Optionally, after obtaining the data information of each node in the data lineage to be verified, the method further includes:
[0018] Verify whether data of various data types in each node have a data source node to obtain a second verification result.
[0019] Optionally, the method further includes:
[0020] If the second verification result indicates that the third data in the fourth node does not have a data source node, obtaining the audit data of the fourth node;
[0021] parsing the audit data of the fourth node;
[0022] According to the analysis result, if the fifth node exists, then second bloodline completion information is generated; if the fifth node does not exist, then second warning information is generated;
[0023] The fifth node is a data source node of the third data.
[0024] Optionally, the method further includes:
[0025] generating a mark of the data lineage according to the first verification result and / or the second verification result;
[0026] Wherein, the marking includes:
[0027] The first mark is used to indicate that there is a break in the data lineage; or
[0028] The second mark is used to indicate that the data has complete lineage.
[0029] Optionally, the method further includes:
[0030] If the mark of the data lineage is the second mark, a third warning message is generated.
[0031] To achieve the above-mentioned object, a data processing device according to an embodiment of the present invention includes:
[0032] A first acquisition module is used to acquire data information of each node in the data lineage to be checked, wherein the data information includes data volume and data type;
[0033] A first processing module is configured to verify the data volume of the corresponding first node and second node according to each data type to obtain a first verification result;
[0034] The second node is a data source node of the first data in the first node, and the first data belongs to a data type currently being verified.
[0035] Optionally, the first processing module is further configured to:
[0036] determining whether the data volume of the first data is the same as the data volume of the second data;
[0037] The second data is data transmitted from the second node to the first node, and the second data and the first data are of the same data type.
[0038] Optionally, the device further comprises:
[0039] a second acquisition module, configured to acquire the audit data of the first node if the first verification result indicates that the data volume of the first data is different from the data volume of the second data;
[0040] a second processing module, configured to generate, based on the audit data of the first node, first lineage completion information if the third node includes other nodes except the second node; and generate first alarm information if the third node only includes the second node;
[0041] The third node is a data source node of the first data.
[0042] Optionally, the device further comprises:
[0043] The third processing module is used to verify whether data of various data types in each node has a data source node, and obtain a second verification result.
[0044] Optionally, the device further comprises:
[0045] a third acquisition module, configured to acquire the audit data of the fourth node if the second verification result indicates that the third data in the fourth node does not have a data source node;
[0046] a fourth processing module, configured to generate second lineage completion information based on the audit data of the fourth node if the fifth node exists; and generate second warning information if the fifth node does not exist;
[0047] The fifth node is a data source node of the third data.
[0048] Optionally, the device further comprises:
[0049] a fifth processing module, configured to generate a marker of the data lineage according to the first verification result and / or the second verification result;
[0050] Wherein, the marking includes:
[0051] The first mark is used to indicate that there is a break in the data lineage; or
[0052] The second mark is used to indicate that the data has complete lineage.
[0053] Optionally, the device further comprises:
[0054] The sixth processing module is configured to generate a third warning message if the mark of the data lineage is the second mark.
[0055] To achieve the above-mentioned object, a data processing device according to an embodiment of the present invention includes a processor, wherein the processor is configured to:
[0056] Obtaining data information of each node in the data lineage to be checked, wherein the data information includes data volume and data type;
[0057] According to each data type, verify the data amount of the corresponding first node and second node to obtain a first verification result;
[0058] The second node is a data source node of the first data in the first node, and the first data belongs to a data type currently being verified.
[0059] Optionally, the processor is further configured to:
[0060] determining whether the data volume of the first data is the same as the data volume of the second data;
[0061] The second data is data transmitted from the second node to the first node, and the second data and the first data are of the same data type.
[0062] Optionally, the processor is further configured to:
[0063] If the first verification result indicates that the data volume of the first data is different from the data volume of the second data, obtaining audit data of the first node;
[0064] According to the audit data of the first node, if the third node includes other nodes except the second node, generating first lineage completion information; if the third node only includes the second node, generating first alarm information;
[0065] The third node is a data source node of the first data.
[0066] Optionally, the processor is further configured to:
[0067] Verify whether data of various data types in each node have a data source node to obtain a second verification result.
[0068] Optionally, the processor is further configured to:
[0069] If the second verification result indicates that the third data in the fourth node does not have a data source node, obtaining the audit data of the fourth node;
[0070] According to the audit data of the fourth node, if the fifth node exists, generating second lineage completion information; if the fifth node does not exist, generating second warning information;
[0071] The fifth node is a data source node of the third data.
[0072] Optionally, the processor is further configured to:
[0073] generating a mark of the data lineage according to the first verification result and / or the second verification result;
[0074] Wherein, the marking includes:
[0075] The first mark is used to indicate that there is a break in the data lineage; or
[0076] The second mark is used to indicate that the data has complete lineage.
[0077] Optionally, the processor is further configured to:
[0078] If the mark of the data lineage is the second mark, a third warning message is generated.
[0079] To achieve the above-mentioned objectives, an embodiment of the present invention provides a data processing device, comprising a transceiver, a processor, a memory, and a program or instruction stored in the memory and executable on the processor; when the processor executes the program or instruction, the data processing method described above is implemented.
[0080] To achieve the above objectives, an embodiment of the present invention provides a readable storage medium having a program or instruction stored thereon, which implements the steps in the above-mentioned data processing method when the program or instruction is executed by a processor.
[0081] The beneficial effects of the above technical solution of the present invention are as follows:
[0082] The method of the embodiment of the present invention, for the data lineage to be verified, first obtains the data information of each node therein, and the data information includes the data volume and data type. Then, based on each data type, the data volume of the corresponding first node and the second node is further verified to complete the data volume audit of the data lineage to be verified, thereby achieving the purpose of ensuring the reliability of data lineage. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 is a flow chart of a data processing method according to an embodiment of the present invention;
[0084] Figure 2 This is one of the schematic diagrams of data lineage;
[0085] Figure 3 This is the second diagram of data lineage;
[0086] Figure 4 is a structural diagram of a data processing device according to an embodiment of the present invention;
[0087] Figure 5 is a structural diagram of a data processing device according to an embodiment of the present invention;
[0088] Figure 6 This is a structural diagram of a data processing device according to another embodiment of the present invention. DETAILED DESCRIPTION
[0089] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0090] It should be understood that references throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present invention. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout this specification do not necessarily refer to the same embodiment. Furthermore, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0091] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the following processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0092] Additionally, the terms "system" and "network" are often used interchangeably herein.
[0093] In the embodiments provided herein, it should be understood that "B corresponding to A" means that B is associated with A and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information.
[0094] like Figure 1 As shown, a data processing method according to an embodiment of the present invention includes:
[0095] Step 101: Obtain data information of each node in the data lineage to be checked, wherein the data information includes data volume and data type;
[0096] Here, the node includes the intra-system or inter-system node involved in the data lineage, which can be a database table, business system, etc. In this step, the data information of each node in the data lineage to be verified is obtained for subsequent further verification.
[0097] Step 102: Verify the data volume of the corresponding first node and second node according to each data type to obtain a first verification result;
[0098] The second node is a data source node of the first data in the first node, and the first data belongs to a data type currently being verified.
[0099] Here, the first data in the first node belongs to the data type currently being verified, and the second node is the data source node of the first data, which can be one or more nodes. Of course, both the first node and the second node are nodes in the data lineage to be verified. In this step, after obtaining the data information of each node in the data lineage to be verified in step 101, the data volume of the corresponding first and second nodes is verified based on each data type, thereby obtaining a first verification result.
[0100] In this way, the data processing method of the embodiment of the present invention is performed by the data processing device to execute the above steps 101 and 102. For the data lineage to be verified, the data information of each node is first obtained, and the data information includes the data volume and data type. Then, based on each data type, the data volume of the corresponding first node and the second node is further verified to complete the data volume audit of the data lineage to be verified, thereby achieving the purpose of ensuring the reliability of data lineage.
[0101] It should be noted that in this embodiment, a node scanning unit can scan data information for each node in the data lineage to be verified. The node scanning unit can be implemented through a scanning program or a function such as a database count. If the data lineage to be verified is established for data a, b, c...l, k, then data a, b, c...l, k are scanned at each node. Here, data a, b, c...l, k are data classified according to different data types.
[0102] After the scan is completed, the data information of each node in the data lineage to be tested can be formed into a data asset inventory directory. Each node in the directory is numbered, and the data assets of each node will be distinguished according to the data type, and the storage capacity is the data volume.
[0103] Optionally, in this embodiment, in step 102, verifying the data amount of the corresponding first node and second node according to each data type includes:
[0104] determining whether the data volume of the first data is the same as the data volume of the second data;
[0105] The second data is data transmitted from the second node to the first node, and the second data and the first data are of the same data type.
[0106] In this way, the obtained first verification result will be used to indicate whether the data volume of the first data is the same as the data volume of the second data, which can also be understood as whether there is a link break in the first data of the first node.
[0107] It should be known that, through step 101, the data information of each node in the data lineage to be checked is known. Therefore, optionally, on the one hand, for the data type currently being checked, all target nodes of all data of this data type can be determined, and then the respective source nodes for each target node can be determined to perform data quantity verification; on the other hand, for each node, the data type corresponding to its data can be used as the data type currently being checked, and then the source node can be determined to perform quantity verification.
[0108] Assuming that the data lineage to be verified includes 5 nodes, the data asset inventory directory formed after scanning is shown in Table 1 below:
[0109]
[0110] Table 1
[0111] Among them, taking the data a in node 5 as an example, Figure 2The diagram of the data lineage to be verified shows that its corresponding source node is node 2. The data volumes of data a in node 5 and node 2 are verified. The result is that the storage volume of data a in node 5 (15,000) is greater than the storage volume of data a in node 2 (10,000), indicating a link break in node 5. Regarding data b in node 5, its corresponding source node is node 4. The data volumes of data b in node 5 and node 4 are verified. The result is that the storage volume of data b in node 5 (8,000) is equal to the storage volume of data b in node 4 (8,000), indicating a link break in node 5.
[0112] In the case where the amount of the first data of the first node is not equal to the amount of the second data of the second node, optionally, this embodiment further includes:
[0113] If the first verification result indicates that the data volume of the first data is different from the data volume of the second data, obtaining audit data of the first node;
[0114] According to the audit data, if the third node includes other nodes except the second node, generating first lineage completion information; if the third node only includes the second node, generating first warning information;
[0115] The third node is a data source node of the first data.
[0116] In this way, the generated first lineage completion information can be pushed to the target object (such as the data administrator) for subsequent completion and update of the data lineage. The generated first alarm information can also be pushed to the target object (such as the data administrator) for subsequent investigation.
[0117] Among them, obtaining the audit data of the first node can be achieved through a data collection tool. Optionally, for nodes within the system, the data collection tool can be a plug-in tool such as a hook implemented based on a database or data system to obtain the flow of data between data tables, and can also obtain the amount of data returned by each operation result to obtain audit data. For inter-system lineage, the data collection tool is a traffic collection tool, such as a data packet analyzer (packetbeat), which captures the entire traffic to obtain audit data. The captured traffic contains the IP information of the previous node, and the subsequent analysis of the audit data can parse and count information such as the amount of data from each IP from the traffic.
[0118] Taking the data a in the above-mentioned node 5 as an example, since the storage capacity of data a in node 5 (15,000) is greater than the storage capacity of data a in node 2 (10,000), the traffic data of node 5 will be further obtained through data collection tools. The collection method is generally to monitor the database / data platform component port and capture and parse the traffic (common tools such as packetbeat). Complete audit data can be parsed from the traffic data, such as import operations, export operations, IP addresses of execution operations, users, execution operation instructions, target data for operations, etc. According to the audit data, the source node of data a is parsed from the traffic data, in addition to node 2, also includes node 3. The flow link between node 3 and node 5 is used as the lineage completion information of data a and is pushed to the data administrator. Of course, if the audit data finds that there is no other source node of data a, the "node 5 data a broken link" can be pushed to the data administrator as an alarm message for subsequent investigation. In addition, considering the situation where the source of data in the data lineage is unknown, the method of the embodiment of the present invention, optionally, after step 101, also includes:
[0119] Verify whether data of various data types in each node have a data source node to obtain a second verification result.
[0120] In this way, for the nodes in the data lineage to be verified, an audit will be conducted to check whether the data of various data types in each node have the data source node, further ensuring the reliability of the data lineage.
[0121] The second verification result is used to indicate whether the data of the target data type exists in the data source node in the current node, which can also be understood as whether the data of the target data type in the node is broken. The target data type is the data type of the data in the current node.
[0122] Taking the data c in the above node 5 as an example, Figure 2 The schematic diagram of the data lineage to be verified does not have its corresponding source node. Therefore, the verification result is that the data c in node 5 does not have a source node and there is a broken link in the data c in node 5.
[0123] Similar to the completion of the data volume, optionally, this embodiment further includes:
[0124] If the second verification result indicates that the third data in the fourth node does not have a data source node, obtaining the audit data of the fourth node;
[0125] According to the audit data of the fourth node, if the fifth node exists, generating second lineage completion information; if the fifth node does not exist, generating second warning information;
[0126] The fifth node is a data source node of the third data.
[0127] In this way, the generated second lineage completion information can be pushed to the target object (such as a data administrator) for subsequent completion and update of the data lineage. The generated second alarm information can also be pushed to the target object (such as a data administrator) for subsequent investigation.
[0128] The method for obtaining the audit data of the fourth node is the same as that for obtaining the audit data of the first node, which will not be described in detail here.
[0129] Taking the data c in the above-mentioned node 5 as an example, since it is currently known that there is no source node for the data c in node 5, the traffic data of node 5 is further obtained through data collection tools. The collection method is generally to monitor the database / data platform component port and capture and parse the traffic (common tools such as packetbeat). Complete audit data can be parsed from the traffic data, such as import operations, export operations, IP addresses that perform operations, users, executed operation instructions, target data for operations, etc. Through traffic analysis, it is found that a certain IP address has performed an import operation (such as load, import, etc.) on the data c of node 5, then the node corresponding to the IP is the possible source of data c. For example, the traffic data of node 5 after parsing (only a format example) is as follows:
[0130] {"@timestamp":"2018-09-20T01:00:42.059Z","@metadata":{"beat":"packetbeat","type":"doc","version":"6.3 .1"},"client_ip":"10.2.41.119","bytes_out":0,"server":"","status":"Error","direction":"in","beat":{"v version":"6.3.1","name":"gateway.bdsmdemo.cmcc","hostname":"gateway.bdsmdemo.cmcc"},"host":{"name":"ga teway.bdsmdemo.cmcc"},"client_server":"","query":"impusername / password@TEST_ORCL / test_dbfile=bak_file path full=y","client_proc":"","path":"","user":"sys","ip":"10.2.41.203","bytes_in":0,"client_port":50757,"responset ime":7,"method":"IMP","proc":"","oracle_version":314,"oracle":{"error_code":0,"error_message":"ORA-00001:unique constraint(SYS.SYS_C0011126)violated\n","affected_rows":0,"insert_id":0,"num_rows":0,"num_fields":0,"iserror":true},"port":1521,"type":"oracle"}
[0131] The above traffic data shows that user "sys" performed remote import operations such as imp on 10.2.41.203 (node 5) on 10.2.41.119 (node 7), and test_db (node 5's database) is the storage location of data c. At this time, lineage completion information (the flow relationship between node 7 and node 5) can be generated and pushed to the data administrator to update the data lineage, such as Figure 3When hot-pressing, if the source of data c cannot be determined through audit data, the "node 5 data c broken link" alarm can be pushed to the data administrator for subsequent investigation.
[0132] In this embodiment, considering the data volume and path audit of the data lineage to be checked, the method may further include:
[0133] generating a mark of the data lineage according to the first verification result and / or the second verification result;
[0134] Wherein, the marking includes:
[0135] The first mark is used to indicate that there is a break in the data lineage; or
[0136] The second mark is used to indicate that the data has complete lineage.
[0137] In this way, the integrity of data lineage can be marked by either the result of data volume or path audit; or the integrity of data lineage can be marked by the results of both data volume and path audit.
[0138] Of course, for the data lineage to be verified, if there is a broken link at any node, then the data lineage is broken; if there is no broken link at all nodes, then the data lineage is complete.
[0139] In addition, optionally, the method further comprises:
[0140] If the mark of the data lineage is the first mark, a third warning message is generated.
[0141] Here, when the first mark generated by the verification result, that is, when there is a break in the data lineage, a third alarm message is generated, and the third alarm message can be pushed to the target object (such as the data administrator) for subsequent investigation.
[0142] To sum up, the method of the embodiment of the present invention, for the data lineage to be verified, first obtains the data information of each node therein, and the data information includes the data volume and data type. Then, based on each data type, the data volume of the corresponding first node and the second node is verified to complete the data volume audit of the data lineage to be verified, thereby achieving the purpose of ensuring the reliability of data lineage.
[0143] like Figure 4 As shown, a data processing device according to an embodiment of the present invention includes:
[0144] An acquisition module 410 is configured to acquire data information of each node in the data lineage to be checked, wherein the data information includes data volume and data type;
[0145] A first processing module 420 is configured to verify the data volume of the corresponding first node and second node according to each data type to obtain a first verification result;
[0146] The second node is a data source node of the first data in the first node, and the first data belongs to a data type currently being verified.
[0147] Optionally, the first processing module is further configured to:
[0148] determining whether the data volume of the first data is the same as the data volume of the second data;
[0149] The second data is data transmitted from the second node to the first node, and the second data and the first data are of the same data type.
[0150] Optionally, the device further comprises:
[0151] a second acquisition module, configured to acquire the audit data of the first node if the first verification result indicates that the data volume of the first data is different from the data volume of the second data;
[0152] a second processing module, configured to generate, based on the audit data of the first node, first lineage completion information if the third node includes other nodes except the second node; and generate first alarm information if the third node only includes the second node;
[0153] The third node is a data source node of the first data.
[0154] Optionally, the device further comprises:
[0155] The third processing module is used to verify whether data of various data types in each node has a data source node, and obtain a second verification result.
[0156] Optionally, the device further comprises:
[0157] a third acquisition module, configured to acquire the audit data of the fourth node if the second verification result indicates that the third data in the fourth node does not have a data source node;
[0158] a fourth processing module, configured to generate second lineage completion information based on the audit data of the fourth node if the fifth node exists; and generate second warning information if the fifth node does not exist;
[0159] The fifth node is a data source node of the third data.
[0160] Optionally, the device further comprises:
[0161] a fifth processing module, configured to generate a marker of the data lineage according to the first verification result and / or the second verification result;
[0162] Wherein, the marking includes:
[0163] The first mark is used to indicate that there is a break in the data lineage; or
[0164] The second mark is used to indicate that the data has complete lineage.
[0165] Optionally, the device further comprises:
[0166] The sixth processing module is configured to generate a third warning message if the mark of the data lineage is the second mark.
[0167] For the data lineage to be checked, the device will first obtain the data information of each node, which includes the data volume and data type. Then, based on each data type, it will further verify the data volume of the corresponding first node and second node to complete the data volume audit of the data lineage to be checked, so as to achieve the purpose of ensuring the reliability of data lineage.
[0168] It should be noted that the device applies the above-mentioned data processing method, and the implementation method of the above-mentioned method embodiment is applicable to the device and can also achieve the same technical effect.
[0169] like Figure 5 As shown, a data processing device 500 according to an embodiment of the present invention includes a processor 510, wherein the processor 510 is configured to:
[0170] Obtaining data information of each node in the data lineage to be checked, wherein the data information includes data volume and data type;
[0171] According to each data type, verify the data amount of the corresponding first node and second node to obtain a first verification result;
[0172] The second node is a data source node of the first data in the first node, and the first data belongs to a data type currently being verified.
[0173] Optionally, the processor is further configured to:
[0174] determining whether the data volume of the first data is the same as the data volume of the second data;
[0175] The second data is data transmitted from the second node to the first node, and the second data and the first data are of the same data type.
[0176] Optionally, the processor is further configured to:
[0177] If the first verification result indicates that the data volume of the first data is different from the data volume of the second data, obtaining audit data of the first node;
[0178] According to the audit data of the first node, if the third node includes other nodes except the second node, generating first lineage completion information; if the third node only includes the second node, generating first alarm information;
[0179] The third node is a data source node of the first data.
[0180] Optionally, the processor is further configured to:
[0181] Verify whether data of various data types in each node have a data source node to obtain a second verification result.
[0182] Optionally, the processor is further configured to:
[0183] If the second verification result indicates that the third data in the fourth node does not have a data source node, obtaining the audit data of the fourth node;
[0184] According to the audit data of the fourth node, if the fifth node exists, generating second lineage completion information; if the fifth node does not exist, generating second warning information;
[0185] The fifth node is a data source node of the third data.
[0186] Optionally, the processor is further configured to:
[0187] generating a mark of the data lineage according to the first verification result and / or the second verification result;
[0188] Wherein, the marking includes:
[0189] The first mark is used to indicate that there is a break in the data lineage; or
[0190] The second mark is used to indicate that the data has complete lineage.
[0191] Optionally, the processor is further configured to:
[0192] If the mark of the data lineage is the second mark, a third warning message is generated.
[0193] The data processing device of this embodiment first obtains the data information of each node for the data lineage to be verified, and the data information includes the data volume and data type. Then, based on each data type, the data volume of the corresponding first node and the second node is verified to complete the data volume audit of the data lineage to be verified, thereby achieving the purpose of ensuring the reliability of data lineage.
[0194] It should be noted that the device applies the above-mentioned data processing method, and the implementation method of the above-mentioned method embodiment is applicable to the device and can also achieve the same technical effect.
[0195] A data processing device according to another embodiment of the present invention, Figure 6 As shown, it includes a transceiver 610, a processor 600, a memory 620, and a program or instruction stored in the memory 620 and executable on the processor 600; the processor 600 implements the above-mentioned data processing method when executing the program or instruction.
[0196] The transceiver 610 is configured to receive and send data under the control of the processor 600 .
[0197] Among them, Figure 6 In the embodiment of the present invention, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits such as one or more processors represented by processor 600 and memory represented by memory 620. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore not further described herein. The bus interface provides an interface. The transceiver 610 can be multiple components, namely, a transmitter and a receiver, providing a unit for communicating with various other devices over a transmission medium.
[0198] The processor 600 is responsible for managing the bus architecture and general processing, and the memory 620 can store data used by the processor 600 when performing operations.
[0199] A readable storage medium according to an embodiment of the present invention stores a program or instruction thereon. When the program or instruction is executed by a processor, the steps in the data processing method described above are implemented and the same technical effect can be achieved. To avoid repetition, they will not be described here.
[0200] The processor is the processor in the data processing device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.
[0201] It should be further noted that many functional components described in this specification are referred to as modules in order to more particularly emphasize the independence of their implementation methods.
[0202] In embodiments of the present invention, modules can be implemented in software so that they can be executed by various types of processors. For example, an identified executable code module can include one or more physical or logical blocks of computer instructions, for example, which can be constructed as objects, procedures, or functions. Nevertheless, the executable code of the identified module does not need to be physically located together, but can include different instructions stored in different locations, which, when logically combined together, constitute the module and achieve the specified purpose of the module.
[0203] In fact, executable code module can be a single instruction or many instructions, and can even be distributed on a plurality of different code segments, distributed in the middle of different programs, and distributed across a plurality of memory devices.Similarly, operating data can be identified in the module, and can be implemented and organized in the data structure of any appropriate type according to any appropriate form.Described operating data can be collected as a single data set, or can be distributed in different locations (including on different storage devices), and can only be present on a system or network as an electronic signal at least in part.
[0204] When a module can be implemented using software, given the current state of hardware technology, those skilled in the art can build corresponding hardware circuits to implement the corresponding functions of the module, regardless of cost. The hardware circuits may include conventional very large scale integration (VLSI) circuits or gate arrays, as well as existing semiconductors such as logic chips and transistors, or other discrete components. Modules may also be implemented using programmable hardware devices, such as field programmable gate arrays, programmable array logic, or programmable logic devices.
[0205] The above exemplary embodiments are described with reference to the accompanying drawings. Many different forms and embodiments are possible without departing from the spirit and teachings of the present invention. Therefore, the present invention should not be construed as limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided so that this disclosure will be complete and perfect and will convey the scope of the invention to those skilled in the art. In the drawings, component sizes and relative sizes may be exaggerated for clarity. The terminology used herein is for purposes of describing specific exemplary embodiments only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include plural forms, unless the context clearly indicates otherwise. It will be further understood that the terms "comprising" and / or "including," when used in this specification, indicate the presence of stated features, integers, steps, operations, components, and / or elements, but do not preclude the presence or addition of one or more other features, integers, steps, operations, components, elements, and / or groups thereof. Unless otherwise indicated, when stated, a range of values includes the upper and lower limits of that range and any subranges therebetween.
[0206] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A data processing method, characterized in that: include: Obtaining data information of each node in the data lineage to be checked, wherein the data information includes data volume and data type; According to each data type, verify the data amount of the corresponding first node and second node to obtain a first verification result; The second node is a data source node of the first data in the first node, and the first data belongs to the data type currently being verified; After obtaining the data information of each node in the data lineage to be checked, the method further includes: Verifying whether data of various data types in each node exist in a data source node, and obtaining a second verification result; generating second bloodline completion information and / or second warning information according to the second verification result; The method further comprises: If the first verification result indicates that the data volume of the first data is different from the data volume of the second data, obtaining audit data of the first node, where the second data is data transmitted from the second node to the first node, and the second data and the first data are of the same data type; Based on the audit data of the first node, if the third node includes other nodes except the second node, generating first lineage completion information; if the third node only includes the second node, generating first alarm information; wherein the third node is the data source node of the first data; The method further comprises: If the second verification result indicates that the third data in the fourth node does not have a data source node, obtaining the audit data of the fourth node; According to the audit data of the fourth node, if the fifth node exists, the second bloodline completion information is generated; if the fifth node does not exist, the second alarm information is generated; wherein the fifth node is the data source node of the third data.
2. The method according to claim 1, characterized in that The checking of the data amount of the corresponding first node and second node according to each data type includes: It is determined whether the data volume of the first data is the same as the data volume of the second data.
3. The method according to claim 1, characterized in that Also includes: generating a mark of the data lineage according to the first verification result and / or the second verification result; Wherein, the marking includes: The first mark is used to indicate that there is a break in the data lineage; or The second mark is used to indicate that the data has complete lineage.
4. The method according to claim 3, characterized in that Also includes: If the mark of the data lineage is the second mark, a third warning message is generated.
5. A data processing device, characterized in that: include: An acquisition module, configured to acquire data information of each node in the data lineage to be inspected, wherein the data information includes data volume and data type; A first processing module is configured to verify the data volume of the corresponding first node and second node according to each data type to obtain a first verification result; The second node is a data source node of the first data in the first node, and the first data belongs to the data type currently being verified; Wherein, the device further comprises: A third processing module is used to verify whether data of various data types in each node exist in a data source node, and obtain a second verification result; an information generating module, configured to generate second blood relationship completion information and / or second warning information according to the second verification result; Wherein, the device further comprises: a second acquisition module, configured to acquire audit data of the first node if the first verification result indicates that the data volume of the first data is different from the data volume of second data, where the second data is data transmitted from the second node to the first node and the second data is of the same data type as the first data; a second processing module, configured to generate, based on the audit data of the first node, first lineage completion information if a third node includes other nodes in addition to the second node; and generate first alarm information if the third node includes only the second node; wherein the third node is a data source node of the first data; The device further comprises: a third acquisition module, configured to acquire the audit data of the fourth node if the second verification result indicates that the third data in the fourth node does not have a data source node; The fourth processing module is used to generate second bloodline completion information based on the audit data of the fourth node if a fifth node exists; if the fifth node does not exist, generate second alarm information; wherein the fifth node is the data source node of the third data.
6. A data processing device, characterized in that: include: A processor configured to: Obtaining data information of each node in the data lineage to be checked, wherein the data information includes data volume and data type; According to each data type, verify the data amount of the corresponding first node and second node to obtain a first verification result; The second node is a data source node of the first data in the first node, and the first data belongs to the data type currently being verified; The processor is further configured to: Verifying whether data of various data types in each node exist in a data source node, and obtaining a second verification result; generating second bloodline completion information and / or second warning information according to the second verification result; The processor is further configured to: If the first verification result indicates that the data volume of the first data is different from the data volume of the second data, obtaining audit data of the first node, where the second data is data transmitted from the second node to the first node, and the second data and the first data are of the same data type; Based on the audit data of the first node, if the third node includes other nodes except the second node, generating first lineage completion information; if the third node only includes the second node, generating first alarm information; wherein the third node is the data source node of the first data; The processor is further configured to: If the second verification result indicates that the third data in the fourth node does not have a data source node, obtaining the audit data of the fourth node; According to the audit data of the fourth node, if the fifth node exists, the second bloodline completion information is generated; if the fifth node does not exist, the second alarm information is generated; wherein the fifth node is the data source node of the third data.
7. A data processing device comprising: A transceiver, a processor, a memory, and a program or instruction stored in the memory and executable on the processor; wherein the processor implements the data processing method according to any one of claims 1 to 4 when executing the program or instruction.
8. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or instruction is executed by a processor, the steps in the data processing method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Data clearing method and device
CN106997369A
Electric power standing book data checking method and device based on blood relationship
CN108763304A