Method and device for constructing test data

By determining the change type of the change object and the relationship type of the downstream blood link in the data processing link, accurately locate the nodes to be verified, and construct the test data, the problem of inaccurate range of the nodes to be verified in the traditional method is solved, and more efficient test data construction and verification is achieved.

CN120216487APending Publication Date: 2025-06-27WEBANK (CHINA)
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510276175.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When the logic of data processing is changed, the downstream object of the traditional method changes the object and the downstream object is noisy, resulting in inaccurate ranges that require re-verification, which increases the analysis cost of staff.

Method used

The change object collection is determined based on the data processing links of the i-th and i-1th versions, and the nodes to be verified are accurately positioned based on the change type of the change object and the relationship type of the downstream blood link in the blood relationship knowledge graph, and the test data is constructed based on the dependent objects of the change object.

Benefits of technology

The range of nodes to be verified is streamlined, the number of nodes to be verified is reduced, the cost of building test data is saved, and the accuracy of verification data is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216487A_ABST
    Figure CN120216487A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for constructing test data, and the method comprises the steps: determining a change object set based on a data processing link of an ith version and a data processing link of an (i-1) th version; determining the change type of the change object according to the difference of the change object between the ith version and the (i-1) th version; according to the blood relationship knowledge graph of the ith version, determining a downstream blood relationship link of the changed object in the blood relationship knowledge graph, and determining a relationship type between any adjacent nodes in the downstream blood relationship link; determining a node to be verified in the downstream blood relationship link according to the change type of the change object and / or the relationship type between any adjacent nodes in the downstream blood relationship link; and constructing test data of the node to be verified in the ith version according to the data of the dependent object of the changed object in the (i-1) th version. By adopting the method, the node to be verified can be accurately determined, and the analysis cost of workers is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer information technology, and in particular, to a method and device for constructing test data. Background Art

[0002] In the era of big data, more and more data is applied in various business scenarios, and data needs to go through multiple processing processes in each business scenario. When the logic in the business scenario changes, the data processing logic will also change accordingly. Since the change in the processing logic of the upstream object will directly or indirectly affect the values of the associated downstream objects, all the changed downstream objects need to be re-verified by testers. Therefore, how to accurately locate the objects that need to be verified is crucial.

[0003] In the traditional method, according to the blood relationship knowledge graph, the downstream objects of the changed object in the graph are circled as the influence range of the changed object, and both the changed object and the objects in the circled influence range need to be re-verified. However, in actual business, the downstream objects of the changed object in the graph may not necessarily be affected by the changed object. Therefore, there is a lot of noise in the range that needs to be re-verified determined by the traditional method, and secondary analysis by staff needs to be introduced, resulting in a high cost. Summary of the Invention

[0004] This application provides a method and device for constructing test data, which are used to accurately determine the nodes to be verified and reduce the analysis cost of staff.

[0005] In a first aspect, an embodiment of this application provides a method for constructing test data. This method can be executed by a device for constructing test data. The method includes: determining a set of changed objects based on the data processing link of the i-th version and the data processing link of the i-1-th version; any data processing link includes multiple nodes with upstream and downstream link relationships; for any changed object in the set of changed objects, determining the change type of the changed object according to the difference between the changed object in the i-th version and the i-1-th version; according to the blood relationship knowledge graph of the i-th version, determining the downstream blood relationship link of the changed object in the blood relationship knowledge graph, and determining the relationship type between any adjacent nodes in the downstream blood relationship link; determining the nodes to be verified in the downstream blood relationship link according to the change type of the changed object and / or the relationship type between any adjacent nodes in the downstream blood relationship link; constructing test data for the nodes to be verified in the i-th version according to the data of the dependent object of the changed object in the i-1-th version.

[0006] Using the above method, according to the change type of the change object and / or the relationship type between adjacent nodes, the nodes to be verified in the downstream lineage link of the change object can be determined, and then, according to the data of the dependent objects of the change object, the test data for each node to be verified can be constructed. Compared with the traditional method in which all the nodes in the downstream lineage link of the change object are designated as nodes to be verified, the scope of the nodes to be verified in this application is streamlined, the number of nodes to be verified is reasonably simplified, and thus the cost of constructing the test data for the nodes to be verified is saved.

[0007] In a possible implementation manner, constructing the test data for the nodes to be verified in the i-th version according to the data of the dependent objects of the change object in the (i - 1)-th version includes: determining the dependent objects of the change object according to the lineage knowledge graph; according to the data of the dependent objects in the (i - 1)-th version, parallelly executing the data processing logic between the dependent objects of the change object and each node to be verified in the i-th version, so as to obtain the test data for each node to be verified; the test data for any node to be verified is the data after the operation of the upper-layer node of the node to be verified.

[0008] Using the above method, executing the data processing logic between the change object and each node to be verified according to the data of the dependent objects of the change object in the (i - 1)-th version to obtain the test data for each node to be verified. Compared with the traditional method in which the test data for each node to be verified is constructed according to the data of the first-layer objects in the (i - 1)-th version in the entire data link, the data processing logic to be executed for constructing the test data for the nodes to be verified is shortened. In addition, the embodiments of this application construct the test data for each node to be verified in a parallel manner, saving time costs and improving the efficiency of constructing the test data.

[0009] In a possible implementation manner, constructing the test data for the nodes to be verified in the i-th version according to the data of the dependent objects of the change object in the (i - 1)-th version includes: determining the dependent objects of the change object according to the lineage knowledge graph; according to the data of the dependent objects of the change object in the (i - 1)-th version, executing the data processing logic between the dependent objects and the change object in the i-th version to obtain the data of the change object in the i-th version; according to the data of the change object in the i-th version, sequentially executing the data processing logic of each node to be verified in the i-th version, so as to obtain the test data corresponding to each node to be verified.

[0010] Using the above method, the embodiments of this application provide another method for constructing the test data for the nodes to be verified, that is, sequentially executing the data processing logic of each node to be verified according to the data of the change object in the i-th version to obtain the test data for each node to be verified.

[0011] In a possible implementation, the dependent objects of the changed object are determined in the following manner: Determine n upper-layer objects on which the changed object depends according to the blood relationship knowledge graph. If the n upper-layer objects all have data in the (i - 1)-th version and have not changed in the i-th version, then the n upper-layer objects are the dependent objects of the changed object; if m upper-layer objects among the n upper-layer objects do not have data in the (i - 1)-th version, or have data in the (i - 1)-th version but have changed in the i-th version, then determine the dependent objects of the m upper-layer objects that have not changed in the i-th version according to the blood relationship knowledge graph. The dependent objects of the m upper-layer objects that have not changed in the i-th version and the upper-layer objects among the n upper-layer objects other than the m upper-layer objects are the dependent objects of the changed object, where m and n are positive integers, and m is less than or equal to n.

[0012] In a possible implementation, after constructing the test data of the node to be verified in the i-th version, the method further includes: determining the data processing logic for executing the node to be verified in the i-th version according to the test data of the node to be verified, and determining the data corresponding to the node to be verified; comparing the data corresponding to the node to be verified with the data of the node to be verified in the (i - 1)-th version to determine the difference content; if the difference content meets the expectation, then saving the data corresponding to the node to be verified as the data of the node to be verified in the i-th version.

[0013] Using the above method, compare the data after execution of the node to be verified in the i-th version with the data in the (i - 1)-th version. If the difference content obtained from the comparison meets the expectation, then save the data after execution in the i-th version in the database. In this way, it can be viewed by relevant staff or used as a reference for the data of the next version.

[0014] In a possible implementation, determining the node to be verified in the downstream blood relationship link according to the change type of the changed object and / or the relationship type between any adjacent nodes in the downstream blood relationship link includes: if the change type is HIVE data structure change and the relationship type between consecutive adjacent nodes in the downstream blood relationship link is a pass-through relationship, then the consecutive adjacent nodes in the downstream blood relationship link are non-nodes to be verified; if the change type is HIVE data structure change and the relationship type between adjacent nodes in the downstream blood relationship link is conditional fixed-value judgment or grouping and screening, then the adjacent nodes in the downstream blood relationship link are non-nodes to be verified; if the change type is TDSQL data structure change and the data type of any node in the HIVE data table in the downstream blood relationship link is compatible with the data type of the changed object in the TDSQL data table, then the node is a non-node to be verified.

[0015] In a possible implementation, the method further includes: if the data table where the node in the downstream lineage link is a task that has not been enabled within a preset duration, the node is a non-verification node; if the data table where the node in the downstream lineage link is a temporary table, the node is a non-verification node; if the data table where the node in the downstream lineage link is a data cleaning table, the node is a non-verification node.

[0016] In a second aspect, an embodiment of the present application provides a device for constructing test data. The device includes a determination module, configured to determine a set of changed objects based on the data processing link of the i-th version and the data processing link of the (i - 1)-th version; any data processing link includes multiple nodes with an upstream and downstream link relationship; the determination module is further configured to, for any changed object in the set of changed objects, determine the change type of the changed object according to the difference between the changed object in the i-th version and the (i - 1)-th version; according to the lineage knowledge graph of the i-th version, determine the downstream lineage link of the changed object in the lineage knowledge graph, and determine the relationship type between any adjacent nodes in the downstream lineage link; according to the change type of the changed object and / or the relationship type between any adjacent nodes in the downstream lineage link, determine the nodes to be verified in the downstream lineage link; a construction module, configured to construct test data for the nodes to be verified in the i-th version according to the data of the dependent object of the changed object in the (i - 1)-th version.

[0017] In a possible implementation, the determination module is further configured to determine the dependent object of the changed object according to the lineage knowledge graph; the construction module is specifically configured to, according to the data of the dependent object in the (i - 1)-th version, execute the data processing logic between the dependent object of the changed object and each node to be verified in the i-th version in parallel, so as to obtain the test data of each node to be verified; the test data of any node to be verified is the data after the upper-layer node of the node to be verified runs.

[0018] In a possible implementation, the construction module is further specifically configured to, according to the data of the dependent object of the changed object in the (i - 1)-th version, execute the data processing logic between the dependent object and the changed object in the i-th version to obtain the data of the changed object in the i-th version; according to the data of the changed object in the i-th version, sequentially execute the data processing logic of each node to be verified in the i-th version, so as to obtain the test data corresponding to each node to be verified.

[0019] In a possible implementation manner, the determining module is specifically configured to determine n upper-layer objects on which the changed object depends according to the blood relationship knowledge graph. If the n upper-layer objects all have data in the (i - 1)-th version and have not changed in the i-th version, then the n upper-layer objects are the dependent objects of the changed object; if m upper-layer objects among the n upper-layer objects do not have data in the (i - 1)-th version, or have data in the (i - 1)-th version but have changed in the i-th version, then determine the dependent objects of the m upper-layer objects that have not changed in the i-th version according to the blood relationship knowledge graph, and the dependent objects of the m upper-layer objects that have not changed in the i-th version and the upper-layer objects among the n upper-layer objects except the m upper-layer objects are the dependent objects of the changed object, where m and n are positive integers, and m is less than or equal to n.

[0020] In a possible implementation manner, the determining module is further configured to determine, according to the test data of the node to be verified, to execute the data processing logic of the node to be verified in the i-th version, and determine the data corresponding to the node to be verified; the apparatus further includes a comparison module, and the comparison module is configured to compare the data corresponding to the node to be verified with the data of the node to be verified in the (i - 1)-th version to determine the difference content; the apparatus further includes a saving module, and the saving module is configured to save the data corresponding to the node to be verified as the data of the node to be verified in the i-th version if the difference content meets the expectation.

[0021] In a possible implementation manner, the determining module is specifically configured to, if the change type is a HIVE data structure change and the relationship type between consecutive adjacent nodes in the downstream blood relationship link is a pass-through relationship, then the consecutive adjacent nodes in the downstream blood relationship link are non-nodes to be verified; if the change type is a HIVE data structure change and the relationship type between adjacent nodes in the downstream blood relationship link is a conditional fixed-value judgment or a grouping and screening, then the adjacent nodes in the downstream blood relationship link are non-nodes to be verified; if the change type is a TDSQL data structure change and the data type of any node in the downstream blood relationship link in the HIVE data table is compatible with the data type of the changed object in the TDSQL data table, then the node is a non-node to be verified.

[0022] In a possible implementation manner, if the data table where the node in the downstream blood relationship link is located is a task that has not been enabled within a preset duration, then the node is a non-node to be verified; if the data table where the node in the downstream blood relationship link is located is a temporary table, then the node is a non-node to be verified; if the data table where the node in the downstream blood relationship link is located is a data cleaning table, then the node is a non-node to be verified.

[0023] In a third aspect, an embodiment of the present application provides a device for constructing test data, including a memory and a processor. The memory is used to store computer programs or instructions, and the processor is used to call the computer programs or instructions stored in the memory to execute the method in any possible implementation manner of the first aspect.

[0024] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which instructions are stored. When a computer reads and executes the instructions, the computer is caused to execute the method in any possible implementation manner of the first aspect.

[0025] In a fifth aspect, an embodiment of the present application provides a computer program product, in which instructions are stored. When a computer reads and executes the instructions, the computer is caused to execute the method in any possible implementation manner of the above first aspect. Description of the Drawings

[0026] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 It is a schematic flowchart corresponding to a method for constructing test data provided by an embodiment of the present application;

[0028] Figure 2 It is a schematic diagram of a field blood relationship knowledge graph provided by an embodiment of the present application;

[0029] Figure 3 It is a schematic diagram of the downstream blood relationship link of the DTL_NUM field provided by an embodiment of the present application;

[0030] Figure 4 It is a schematic diagram of the downstream blood relationship link of the ORG_NAME field provided by an embodiment of the present application;

[0031] Figure 5 It is a schematic diagram of the internal modules of a device 5000 for constructing test data provided by an embodiment of the present application;

[0032] Figure 6 It is a schematic diagram of the structure of a device 6000 for constructing test data provided by an embodiment of the present application. Detailed Embodiments

[0033] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0034] With the progress of technology and the advent of the big data era, in various business scenarios, especially in the banking business scenario, data often needs to be processed through multiple layers. When the data processing logic in the business scenario changes, it will inevitably lead to changes in some data, and the changes in upstream data will directly or indirectly affect the values of downstream objects. For the changed objects and the affected objects, testers need to verify whether the data after the object changes meets the expectations.

[0035] In the traditional method, based on the lineage knowledge graph, the downstream objects of the changed object in the graph are all circled as the influence scope of the changed object. However, in actual business, the downstream objects of the changed object in the graph do not necessarily be affected by the changed object. Based on this, the embodiments of this application provide a method for constructing test data to accurately determine the nodes to be verified and reduce the analysis cost of staff.

[0036] Figure 1 It is a schematic flowchart corresponding to a method for constructing test data provided by an embodiment of this application. This flowchart can be executed by a device for constructing test data, such as Figure 1 shown. This process includes the following steps:

[0037] Step 101, the device for constructing test data determines a set of changed objects based on the data processing link of the i-th version and the data processing link of the (i - 1)-th version. Any data processing link includes multiple nodes with upstream and downstream link relationships.

[0038] Exemplarily, the device for constructing test data can be a computer device, a server, or other intelligent devices. The data processing link is determined based on the lineage knowledge graph, and the lineage knowledge graph can be a field lineage knowledge graph or a table lineage knowledge graph. Exemplarily, the field lineage knowledge graph represents the history and relationships of data fields evolving and moving in different systems, processes, or conversions within its organization. It tracks the origin, modification, and destination of data fields to provide a clear understanding of how data is created, changed, and used.

[0039] Figure 2A schematic diagram of a field lineage knowledge graph provided by an embodiment of the present application. Circles represent nodes corresponding to fields. Nodes can represent the tables and databases to which the fields belong. Nodes with a relational connection are connected by a line with an arrow. The line between nodes represents the relationship between nodes. Figure 2 The relationships between nodes include a pass-through relationship, a condition - association judgment, a condition - fixed value judgment, and a many - to - one relationship. Figure 2 The provided field lineage knowledge graph is only an example. The present application does not limit the specific field lineage knowledge graph nor the type of relationships between nodes.

[0040] The processing of data in a business scenario changes with the version. Each version has a corresponding field lineage knowledge graph. The field lineage knowledge graph represents the data processing link, and the data processing link includes multiple nodes with upstream and downstream link relationships. According to the data processing link of the i - th version and the data processing link of the (i - 1) - th version, the set of changed objects in the i - th version compared to the (i - 1) - th version can be determined.

[0041] Optionally, based on the release materials of the version, HIVE data tables, distributed database (Tencent Distributed SQL, TDSQL) data tables, and processing scripts can be obtained. Other content can also be obtained, and not too many examples are given here. By comparing the HIVE data tables, TDSQL data tables, and processing scripts of the i - th version with those of the (i - 1) - th version, the set of changed objects in the i - th version compared to the (i - 1) - th version can be determined.

[0042] HIVE is a data warehouse tool based on Hadoop, used for data extraction, transformation, loading, and storage, query, and analysis of large - scale data. TDSQL is an enterprise - level distributed database product developed based on the MYSQL database engine, with enhanced distributed processing capabilities. In an embodiment of the present application, the HIVE data table is downstream of the TDSQL data table. The content of the HIVE data table is obtained by extracting data from the TDSQL data table. It can be understood that the TDSQL data table is the data table corresponding to the business layer, and data extraction from the business - layer data table runs in HIVE to obtain the HIVE data table.

[0043] The set of changed objects can be changed fields or changed data tables. Changed data tables are newly added data tables or taken - off - line data tables in the i - th version. Since a table is composed of multiple fields, such a change in a newly added data table can also be understood as a change in fields. In an embodiment of the present application, the changed fields are used as an example of the changed objects.

[0044] Step 102, for any change object in the set of change objects, the device for constructing test data determines the change type of the change object according to the differences between the change object in the i-th version and the (i - 1)-th version; determines the downstream blood relationship link of the change object in the blood relationship knowledge graph according to the blood relationship knowledge graph of the i-th version, and determines the relationship type between any adjacent nodes in the downstream blood relationship link; determines the nodes to be verified in the downstream blood relationship link according to the change type of the change object and / or the relationship type between any adjacent nodes in the downstream blood relationship link.

[0045] Specifically, for any change object in the set of change objects, the change type of the change object is determined according to the differences between the change object in the i-th version and the (i - 1)-th version. The change types include HIVE data structure change, TDSQL data structure change, enumeration new value change, and processing logic change. For any change object, the change type of the change object is determined based on the HIVE data tables, TDSQL data tables, and processing scripts of the i-th version and the (i - 1)-th version. Exemplarily, Table A is the HIVE data table of the i-th version, and Table a is the HIVE data table of the (i - 1)-th version. Tables A and a are as follows.

[0046] Table A

[0047]

[0048]

[0049] Table a

[0050]

[0051] It can be found from Tables A and a that Table A has added a new field ACTIVITY_RULE, and the DOMAIN_ID field has added an enumeration, and the type of the DTL_NUM field has changed. That is, the change form of the HIVE data table A is (1) HIVE data structure change - adding the field ACTIVITY_RULE, (2) HIVE data structure change - changing the type of the DTL_NUM field, and (3) enumeration new value change of the DOMAIN_ID field.

[0052] Furthermore, Table B is the TDSQL data table of the i-th version, and Table b is the TDSQL data table of the (i - 1)-th version. Tables B and b are as follows.

[0053] Table B

[0054] Field Name Type Description Whether Enumeration Type Enumeration Value ORG Char(12) Organization Number Yes A, B ACCT_NO Bigint Account Number No ATTR_TYPE Char(3) Attribute Category Yes A, B, C, D ATTR_VALUE Char(500) Attribute Value Yes A, B, C, D ORG_NAME Char(500) Organization Name No

[0055] Table b

[0056] Field Name Type Description Whether Enumeration Type Enumeration Value ORG Char(12) Organization Number Yes A, B ACCT_NO Bigint Account Number No ATTR_TYPE Char(3) Attribute Category Yes A, B, C ATTR_VALUE Char(500) Attribute Value Yes A, B, C

[0057] It can be found from Table B and Table b that there are two fields in Table B with newly added enumerated values, and a new field ORG_NAME is added. That is, the change form of the TDSQL data table B is (1) TDSQL data structure change - new field ORG_NAME added, (2) change of enumerated new values in the ATTR_TYPE field, and (3) change of enumerated new values in the ATTR_VALUE field. In addition, by comparing the processing script of the i-th version with that of the i-1-th version, the change objects in the processing script can be determined, which will not be elaborated here.

[0058] According to the lineage knowledge graph of the i-th version, determine the downstream lineage link of the change object in the lineage knowledge graph, and determine the relationship type between any adjacent nodes in the downstream lineage link. Continuing with the above example, taking the DTL_NUM field as the change object, determine the downstream lineage link of the DTL_NUM field in the lineage knowledge graph and the relationship type between any adjacent nodes in the downstream lineage link. Figure 3 This is a schematic diagram of the downstream lineage link of a DTL_NUM field provided by an embodiment of the present application.

[0059] Furthermore, according to the change type of the change object and / or the relationship type between any adjacent nodes in the downstream lineage link, determine the nodes to be verified in the downstream lineage link, including: if the change type is a HIVE data structure change and the relationship type between consecutive adjacent nodes in the downstream lineage link is a pass-through relationship, then the consecutive adjacent nodes in the downstream lineage link are non-nodes to be verified; if the change type is a HIVE data structure change and the relationship type between adjacent nodes in the downstream lineage link is a conditional fixed-value judgment or grouping screening, then the adjacent nodes in the downstream lineage link are non-nodes to be verified.

[0060] From Figure 3It can be seen that the relationship type between consecutive adjacent nodes from A2.TOTAL_DTL_NUM to A10.DTL_NUM is a pass-through relationship. Therefore, the nodes from A2.TOTAL_DTL_NUM to A10.DTL_NUM are all non-verification-required nodes. The relationship type between A2.TOTAL_DTL_NUM and D1.TXN_NO is a condition-fixed value judgment. Therefore, the D1.TXN_NO node is a non-verification-required node. However, the relationship type of B1.CRT_NUM is a many-to-one relationship and needs to be verified. The relationship type of C1.CTN_TYPE is a conditional association judgment type and needs to be verified. Since the relationship type of A1.DTL_NUM -> A2.TOTAL_DTL_NUM is a one-to-one relationship and there is a function process, A1.DTL_NUM and A2.TOTAL_DTL_NUM need to be verified. The change object of A.DTL_NUM also needs to be verified. Therefore, the verification-required nodes are A.DTL_NUM, A1.DTL_NUM, A2.TOTAL_DTL_NUM, B1.CRT_NUM, and C1.CTN_TYPE.

[0061] If the change type is a TDSQL data structure change, and the data type of any node in the downstream lineage link in the HIVE data table is compatible with the data type of the change object in the TDSQL data table, then the node is a non-verification-required node. Exemplarily, taking the ORG_NAME field as the change object as an example, Figure 4 This is a schematic diagram of the downstream lineage link of an ORG_NAME field provided by an embodiment of the present application. From Figure 4 It can be seen that if B.ORG_NAME is of character type, B1.ORG_NAME is of String type, table B is a TDSQL data table, and table B1 is a HIVE data table. Since the Sting type is compatible with the character type, the node B1.ORG_NAME is a non-verification-required node. Since B1.ORG_NAME is a non-verification-required node, B2.ORG_NAME to B10.ORG_NAME are all non-verification-required nodes.

[0062] Optionally, if the data table where the node in the downstream lineage link is located is a task that has not been enabled within a preset duration, then the node is a non-verification-required node; if the data table where the node in the downstream lineage link is located is a temporary table, then the node is a non-verification-required node; if the data table where the node in the downstream lineage link is located is a data cleaning table, then the node is a non-verification-required node. That is to say, if the data table where the node in the downstream lineage link of the change object is located is a task table, a temporary table, or a data cleaning table that has not been enabled within a preset duration, then the node in the downstream lineage link is a non-verification-required node.

[0063] It should be noted that when the change type of the change object is the change of enumeration new value and the change of processing logic, the nodes in the downstream blood relationship link of the change object are all nodes to be verified.

[0064] Step 103, the device for constructing test data constructs the test data of the nodes to be verified in the i-th version according to the data of the dependent object of the change object in the (i - 1)-th version.

[0065] There are two methods for constructing the test data of the nodes to be verified in the i-th version according to the data of the dependent object of the change object in the (i - 1)-th version. One is the parallel processing method, and the other is the serial processing method. Among them, the parallel processing method includes: determining the dependent object of the change object according to the blood relationship knowledge graph; according to the data of the dependent object in the (i - 1)-th version, parallelly execute the data processing logic between the dependent object of the change object in the i-th version and each node to be verified, so as to obtain the test data of each node to be verified. The test data of any node to be verified is the data after the upper-layer node of the node to be verified runs.

[0066] Exemplarily, taking the change object as A.DTL_NUM as an example, determine the dependent object of A.DTL_NUM according to the blood relationship knowledge graph. If the dependent object is X.DTL_NUM, obtain the data of the X.DTL_NUM field in the (i - 1)-th version, which is the test data of A.DTL_NUM; the test data of the A1.DTL_NUM field is the data after A.DTL_NUM runs in the i-th version, and the test data of other nodes to be verified in the downstream is the data after the upper-layer node of the node to be verified runs.

[0067] Since the dependent object of the object to be changed is an unchanged object, that is, the data of the dependent object of the object to be changed in the (i-1)-th version is the same as that in the i-th version. Therefore, according to the data of the X.DTL_NUM field in the (i-1)-th version, execute the logic of X.DTL_NUM -> A.DTL_NUM to obtain the data of the A.DTL_NUM field after running in the i-th version, which is the test data of the A1.DTL_NUM field; according to the data of the X.DTL_NUM field in the (i-1)-th version, execute the logic of X.DTL_NUM -> A.DTL_NUM -> A1.DTL_NUM to obtain the data of the A1.DTL_NUM field after running in the i-th version, which is the test data of the A2.TOTAL_DTL_NUM field; according to the data of the X.DTL_NUM field in the (i-1)-th version, execute the logic of X.DTL_NUM -> A.DTL_NUM -> A1.DTL_NUM -> A2.TOTAL_DTL_NUM to obtain the data of the A2.TOTAL_DTL_NUM field after running in the i-th version, which is the test data of the C1.CTN_TYPE field. The test data of other fields is also determined by the same parallel method and will not be elaborated here. Since the test data of the nodes to be verified is obtained based on the data of the dependent objects of the objects to be changed in the (i-1)-th version and the data processing logic between the dependent objects and the nodes to be verified in the i-th version, the determination of the test data of each node to be verified can be processed in parallel.

[0068] Optionally, another serial processing method includes: determining the dependent object of the object to be changed according to the blood relationship knowledge graph; according to the data of the dependent object of the object to be changed in the (i-1)-th version, execute the data processing logic between the dependent object and the object to be changed in the i-th version to obtain the data of the object to be changed in the i-th version; according to the data of the object to be changed in the i-th version, sequentially execute the data processing logic of each node to be verified in the i-th version, so as to obtain the test data corresponding to each node to be verified.

[0069] Still taking the change object A.DTL_NUM as an example, determine the dependent object of A.DTL_NUM according to the lineage knowledge graph, which is X.DTL_NUM. Obtain the data of the X.DTL_NUM field in the (i - 1)-th version, and execute the logic of X.DTL_NUM -> A.DTL_NUM based on the data of the X.DTL_NUM field in the (i - 1)-th version to obtain the data of the A.DTL_NUM field of the change object in the i-th version. Further, execute the logic of A.DTL_NUM -> A1.DTL_NUM based on the data of the A.DTL_NUM field in the i-th version to obtain the data after running of A1.DTL_NUM in the i-th version, which is the test data of A2.TOTAL_DTL_NUM; execute the logic of A1.DTL_NUM -> A2.TOTALDTL_NUM based on the data of the A1.DTL_NUM field in the i-th version to obtain the data after running of A2.TOTAL_DTL_NUM in the i-th version, which is the test data of A3.DTL_NUM, C1.CTN_TYPE, D1.TXN_NO, and B1.CRT_NUM. The test data of other fields is also determined according to the same serial method and will not be elaborated here. Since the test data of the nodes to be verified are all obtained based on the data of the upper-layer nodes they depend on in the i-th version, it is possible to process the test data of each node to be verified serially.

[0070] Regarding the dependent object of the change object, it is determined in the following way: Determine the n upper-layer objects that the change object depends on according to the lineage knowledge graph. If data exists for all n upper-layer objects in the (i - 1)-th version and no changes occur in the i-th version, then the n upper-layer objects are the dependent objects of the change object; if data does not exist for m upper-layer objects out of the n upper-layer objects in the (i - 1)-th version, or if data exists for the m upper-layer objects in the (i - 1)-th version but changes occur in the i-th version, then determine the dependent objects of the m upper-layer objects that have not changed in the i-th version according to the lineage knowledge graph. The dependent objects of the m upper-layer objects that have not changed in the i-th version and the upper-layer objects other than the m upper-layer objects among the n upper-layer objects are the dependent objects of the change object. Here, m and n are positive integers, and m is less than or equal to n.

[0071] Specifically, a change object can directly depend on one or more objects above the change object. If the dependent object has data in the (i - 1)-th version and the dependent object remains unchanged in the i-th version compared to the (i - 1)-th version, then one or more objects above can be directly determined as the dependent objects of the change object; if the one or more objects directly depended on by the change object above do not have data in the (i - 1)-th version, indicating that the dependent object is a newly added object, then trace back the dependent objects of the object without data upward according to the blood relationship knowledge graph until an object with data and unchanged is found, which is the dependent object of the change object.

[0072] After constructing the test data of the node to be verified in the i-th version, the method further includes: determining the data processing logic for executing the node to be verified in the i-th version according to the test data of the node to be verified, determining the data corresponding to the node to be verified, comparing the data corresponding to the node to be verified with the data of the node to be verified in the (i - 1)-th version, and determining the difference content. If the difference content meets the expectation, then save the data corresponding to the node to be verified as the data of the node to be verified in the i-th version. Specifically, the data after the upper-layer node of the node to be verified runs in the i-th version is the test data of the node to be verified. Execute the data processing logic of the node to be verified to obtain the data after the node to be verified runs in the i-th version, automatically compare the data after the node to be verified runs in the i-th version with the data in the (i - 1)-th version, and determine whether the difference content meets the expectation. If it meets the expectation, save the data after running in the i-th version in the database; if it does not meet the expectation, an alarm is issued, and the staff can check whether the test data is abnormal or re-run the data processing logic of the node to be verified according to the test data. It can be understood that the embodiments of the present application also need to execute the data processing logic of the change object according to the test data of the change object to determine the data after the change object runs in the i-th version.

[0073] It should be noted that the test data also has a time attribute. For example, if a certain bank needs to count the total deposit amount in the fourth quarter of 2024, then it is necessary to obtain the deposit data within the time range from September to December 2024. Some test data requires data within a certain day or a certain month range, which will not be elaborated here. Therefore, when saving the data of the node to be verified in the i-th version in the database, the time attribute of the data can also be saved in the database.

[0074] Regarding the blood relationship knowledge graph, the blood relationship knowledge graph is updated according to the version update. The nodes in the field blood relationship knowledge graph are the fields of the data table, and the connections between the nodes in the field blood relationship knowledge graph represent the relationships between fields. By executing the data flow through HIVE, the finer-grained processing relationships and conditional relationships of fields are obtained. The relationships between fields include pass-through, one-to-one relationship, many-to-one relationship, condition-fixed value judgment, condition-correlation judgment, and condition-grouping screening. Specifically, the pass-through relationship means that the fields are directly mapped, and there is no function expression for the fields. For example:

[0075] insert overwrite table tableA

[0076] select

[0077] emp.user_name

[0078] from

[0079] emp where emp.age>20;

[0080] As shown above, emp.user_name -> tableA.user_name is a pass-through relationship.

[0081] The one-to-one relationship means that the fields are directly mapped, but there is a function expression in the field. For example: insert overwrite table tableA select

[0082] max(emp.age) as max_age

[0083] from

[0084] emp where emp.age>20;

[0085] As shown above, emp.age -> tableA.max_age is a single-field processing relationship.

[0086] The many-to-one relationship means that multiple fields are mapped to the same target field. For example: insert overwrite table tableA select

[0087] emd.num + dept.num as total_num

[0088] from

[0089] emp

[0090] left join dept d on d.dept_id = emp.dept_id where emp.age > 20;

[0091] As shown above, emp.num -> tableA.total_num is a multi-field processing relationship.

[0092] Condition-fixed value judgment means including a conditional expression that is not a field. For example: insert overwrite tabletableA select

[0093] user_name

[0094] from

[0095] emp

[0096] where emp.age > 20;

[0097] As shown above, emp.age -> tableA.user_name is a fixed value judgment relationship

[0098] Condition - association judgment means including a conditional expression whose value is a field, or matching the join on condition. For example:

[0099] insert overwrite table tableA

[0100] select

[0101] emp.user_name

[0102] from

[0103] emp

[0104] left join dept d

[0105] on d.dept_id = emp.dept_id

[0106] As shown above, dept.id -> tableA.user_name is a condition - association judgment relationship, and emp.id -> tableA.user_name is a condition - association judgment relationship

[0107] Condition - grouping and filtering means regular matching, matching the group by condition. For example:

[0108] insert overwrite table tableB

[0109] select

[0110] dept_id,

[0111] emp.user_name

[0112] from

[0113] emp

[0114] group by dept_id

[0115] As shown above, emp.dept_id -> tableB.dept_id is a conditional-grouping filtering relationship, and emp.dept_id -> tableB.user_name is a conditional-grouping filtering relationship.

[0116] Figure 5 It is a schematic diagram of the internal modules of a device 5000 for constructing test data provided by an embodiment of the present application. As Figure 5 shown, the device may include: a determination module 501, a construction module 502, a comparison module 503, and a storage module 504.

[0117] Among them, the determination module 501 is used to determine a set of change objects based on the data processing link of the i-th version and the data processing link of the (i - 1)-th version; any data processing link includes multiple nodes with upstream and downstream link relationships; the determination module 501 is further used to, for any change object in the set of change objects, determine the change type of the change object according to the difference between the change object in the i-th version and the (i - 1)-th version; according to the lineage knowledge graph of the i-th version, determine the downstream lineage link of the change object in the lineage knowledge graph, and determine the relationship type between any adjacent nodes in the downstream lineage link; according to the change type of the change object and / or the relationship type between any adjacent nodes in the downstream lineage link, determine the nodes to be verified in the downstream lineage link; the construction module 502 is used to construct the test data of the nodes to be verified in the i-th version according to the data of the dependent object of the change object in the (i - 1)-th version.

[0118] In a possible implementation manner, the determination module 501 is further used to determine the dependent object of the change object according to the lineage knowledge graph; the construction module 502 is specifically used to, according to the data of the dependent object in the (i - 1)-th version, execute the data processing logic between the dependent object of the change object in the i-th version and each node to be verified in parallel, so as to obtain the test data of each node to be verified; the test data of any node to be verified is the data after the upper-layer node of the node to be verified runs.

[0119] In a possible implementation, the building module 502 is further specifically configured to execute the data processing logic between the dependent object and the changed object in the i-th version according to the data of the dependent object of the changed object in the (i-1)-th version, so as to obtain the data of the changed object in the i-th version; and execute the data processing logic of each verification node in the i-th version in sequence according to the data of the changed object in the i-th version, so as to obtain the test data corresponding to each verification node.

[0120] In a possible implementation, the determining module 501 is specifically configured to determine n upper-layer objects on which the changed object depends according to the blood relationship knowledge graph. If the n upper-layer objects all have data in the (i-1)-th version and have not changed in the i-th version, then the n upper-layer objects are the dependent objects of the changed object; if m upper-layer objects among the n upper-layer objects do not have data in the (i-1)-th version, or have data in the (i-1)-th version but have changed in the i-th version, then determine the dependent objects of the m upper-layer objects that have not changed in the i-th version according to the blood relationship knowledge graph. The dependent objects of the m upper-layer objects that have not changed in the i-th version and the upper-layer objects other than the m upper-layer objects among the n upper-layer objects are the dependent objects of the changed object, where m and n are positive integers and m is less than or equal to n.

[0121] In a possible implementation, the determining module 501 is further configured to determine to execute the data processing logic of the verification node in the i-th version according to the test data of the verification node, and determine the data corresponding to the verification node; the apparatus further includes a comparison module 503, and the comparison module is configured to compare the data corresponding to the verification node with the data of the verification node in the (i-1)-th version to determine the difference content; the apparatus further includes a storage module 504, and the storage module is configured to save the data corresponding to the verification node as the data of the verification node in the i-th version if the difference content meets the expectation.

[0122] In a possible implementation, the determining module 501 is specifically configured to: if the change type is a HIVE data structure change and the relationship type between consecutive adjacent nodes in the downstream lineage link is a pass-through relationship, then the consecutive adjacent nodes in the downstream lineage link are non-verification nodes; if the change type is a HIVE data structure change and the relationship type between adjacent nodes in the downstream lineage link is a conditional fixed-value judgment or grouping and filtering, then the adjacent nodes in the downstream lineage link are non-verification nodes; if the change type is a TDSQL data structure change and the data type of any node in the downstream lineage link in the HIVE data table is compatible with the data type of the change object in the TDSQL data table, then the node is a non-verification node.

[0123] In a possible implementation, if the data table where the node in the downstream lineage link is located is a task that has not been enabled within a preset time period, then the node is a non-verification node; if the data table where the node in the downstream lineage link is located is a temporary table, then the node is a non-verification node; if the data table where the node in the downstream lineage link is located is a data cleaning table, then the node is a non-verification node.

[0124] Figure 6 The structure diagram of a device 6000 for constructing test data provided by an embodiment of the present application is as follows. Figure 6 As shown, it includes at least one processor 601 and a memory 602 connected to at least one processor 601. In the embodiment of the present application, the specific connection medium between the processor 601 and the memory 602 is not limited. Figure 6 Taking the connection between the processor 601 and the memory 602 through a bus as an example. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0125] In the embodiment of the present application, the memory 602 stores instructions executable by at least one processor 601. By executing the instructions stored in the memory 602, the at least one processor 601 can implement the steps of the above method for constructing test data.

[0126] Among them, the processor 601 is the control center of the computer device. It can connect various parts of the computer device through various interfaces and circuits. By running or executing the instructions stored in the memory 602 and calling the data stored in the memory 602, resource settings can be performed. Optionally, the processor 601 may include one or more processing units. The processor 601 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 601. In some embodiments, the processor 601 and the memory 602 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.

[0127] The processor 601 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0128] The memory 602, being a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 602 can include at least one type of storage medium. For example, it can include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical disks, and so on. The memory 602 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 602 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0129] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0130] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0131] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means embodying the functionality specified in the flowchart(s) Figure 1 one or more flowcharts and / or block diagrams Figure 1 specified in block(s) or block diagrams.

[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functionality specified in the flowchart(s) Figure 1 one or more flowcharts and / or block diagrams Figure 1 specified in block(s) or block diagrams.

[0133] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to cover these changes and modifications.

Claims

1. A method for constructing test data, characterized in that: Based on the data processing link of the i-th version and the data processing link of the i-1-th version, a set of changed objects is determined; any data processing link includes multiple nodes having upstream and downstream link relationships; For any change object in the change object set, determining a change type of the change object according to a difference between the i-th version and the (i-1)-th version of the change object; According to the i-th version of the lineage knowledge graph, determine the downstream lineage link of the changed object in the lineage knowledge graph, and determine the relationship type between any adjacent nodes in the downstream lineage link; Determine the node to be verified in the downstream lineage link according to the change type of the change object and / or the relationship type between any adjacent nodes in the downstream lineage link; According to the data of the dependent object of the changed object in the (i-1)th version, the test data of the node to be verified in the (i)th version is constructed.

2. The method according to claim 1, characterized in that: Constructing test data of the node to be verified in the i-th version according to the data of the dependent object of the changed object in the i-th version, including: Determine the dependent objects of the changed object according to the blood relationship knowledge graph; According to the data of the dependent object in the i-1th version, the data processing logic between the dependent object of the changed object in the i-th version and each node to be verified is executed in parallel, so as to obtain the test data of each node to be verified; the test data of any node to be verified is the data after the upper layer node of the node to be verified is run.

3. The method according to claim 1, characterized in that Constructing test data of the node to be verified in the i-th version according to the data of the dependent object of the changed object in the i-th version, including: Determine the dependent objects of the changed object according to the blood relationship knowledge graph; According to the data of the dependent object of the changed object in the (i-1)th version, execute the data processing logic between the dependent object and the changed object in the (i)th version to obtain the data of the changed object in the (i)th version; According to the data of the changed object in the i-th version, the data processing logic of each node to be verified in the i-th version is executed in sequence, so as to obtain the test data corresponding to each node to be verified.

4. The method according to claim 1, characterized in that The dependent objects of the changed object are determined in the following way: Determine n upper-layer objects that the changed object depends on according to the lineage knowledge graph. If the n upper-layer objects all have data in the i-1th version and have not been changed in the i-th version, then the n upper-layer objects are dependent objects of the changed object; If m of the n upper-level objects do not have data in the i-1th version, or data exist in the i-1th version but are changed in the i-th version, then the dependent objects of the m upper-level objects that have not been changed in the i-th version are determined according to the lineage knowledge graph, and the dependent objects of the m upper-level objects that have not been changed in the i-th version and the upper-level objects of the n upper-level objects except the m upper-level objects are the dependent objects of the changed object, m and n are positive integers, and m is less than or equal to n.

5. The method according to any one of claims 1 to 4, characterized in that: After constructing the test data of the node to be verified in the i-th version, the method further includes: Determine, according to the test data of the node to be verified, to execute the data processing logic of the node to be verified in the i-th version, and determine the data corresponding to the node to be verified; Compare the data corresponding to the node to be verified with the data of the node to be verified in the (i-1)th version to determine the difference; If the difference content is consistent with expectations, the data corresponding to the node to be verified is saved as the data of the node to be verified in the i-th version.

6. The method according to any one of claims 1 to 4, characterized in that: Determining the node to be verified in the downstream lineage link according to the change type of the change object and / or the relationship type between any adjacent nodes in the downstream lineage link includes: If the change type is a HIVE data structure change, and the relationship type between the continuous adjacent nodes in the downstream lineage link is a transparent transmission relationship, then the continuous adjacent nodes in the downstream lineage link are non-to-be-verified nodes; If the change type is HIVE data structure change, and the relationship type between adjacent nodes in the downstream lineage link is conditional solid value judgment or group screening, then the adjacent nodes in the downstream lineage link are non-to-be-verified nodes; If the change type is a TDSQL data structure change, and the data type of any node in the downstream lineage link in the HIVE data table is compatible with the data type of the change object in the TDSQL data table, then the node is a non-to-be-verified node.

7. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: If the data table where the node in the downstream lineage link is located is a task that has not been enabled within the preset time period, the node is a non-to-be-verified node; If the data table where the node in the downstream bloodline link is located is a temporary table, the node is a non-to-be-verified node; If the data table where the node in the downstream bloodline link is located is a data cleaning table, the node is a non-to-be-verified node.

8. A device for constructing test data, characterized in that: include: A determination module, used to determine a set of changed objects based on the data processing link of the i-th version and the data processing link of the i-1-th version; any data processing link includes a plurality of nodes having an upstream and downstream link relationship; The determination module is further configured to determine, for any change object in the change object set, a change type of the change object according to a difference between the i-th version and the (i-1)-th version of the change object; According to the i-th version of the lineage knowledge graph, determine the downstream lineage link of the changed object in the lineage knowledge graph, and determine the relationship type between any adjacent nodes in the downstream lineage link; Determine the node to be verified in the downstream lineage link according to the change type of the change object and / or the relationship type between any adjacent nodes in the downstream lineage link; A construction module is used to construct test data of the node to be verified in the i-th version according to the data of the dependent object of the changed object in the i-1-th version.

9. A device for constructing test data, characterized in that: include: Memory, used to store computer programs or instructions; A processor, configured to call a computer program or instruction stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when a computer reads and executes the instructions, the computer executes the method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The computer program product stores instructions, and when a computer reads and executes the instructions, the computer executes the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Test data construction method and device, equipment and storage medium

    CN120973694A

  • A test data construction method, device and equipment and a storage medium

    CN120973694B

  • Flight passenger data changing method and device

    CN121277958A

  • Oil and gas reservoir representation data reconstruction method fusing data consanguinity and knowledge graph

    CN122388046A