A testing method, device, equipment, and storage medium for a data mapping file
By standardizing the format and analyzing the data mapping files and generating test reports, the problem of inconsistent format of the data mapping table and missing information is solved, and the testing efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202111515291.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-13
AI Technical Summary
In system background construction projects, the inconsistent format of the data mapping table, information omissions and design errors lead to difficulty in testing, and manual proofreading is large, costly and low coverage.
By obtaining the data mapping file, format normalization is performed based on the data mapping template, blood relationship analysis is performed to determine the blood line of the field, and the fields are tested to generate a test report.
Automatic testing of data mapping files is realized, which improves testing efficiency and adequacy, and reduces the cost of manual proofreading and data error rate.
Smart Images

Figure CN114185791B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of computer technology, and in particular, to a method, device, equipment and storage medium for testing a data mapping file. Background Art
[0002] For projects involving the construction of the background in a system, most of the development work content is the transformation and processing of data. In order to ensure the quality of data processing, project team members discuss and compile a data mapping table during the requirement or development stage and incorporate it into the requirement document as a reference basis for later development.
[0003] Due to the large amount of data in the data processing mapping table, generally in the scale of hundreds or thousands of tables, and each table has multiple fields, it is generally compiled by multiple people simultaneously. As a result, the data processing mapping tables delivered to the test personnel have problems such as inconsistent formats, information omission, and incorrect design, resulting in problems such as unconnected fields in the associated system, difficult test analysis, and numerous test execution problems when the project is developed and put into production.
[0004] Currently, it is mainly necessary to manually check the data processing design document (such as the data mapping file) and the development document one by one according to the requirements, which has certain requirements for the test ability of the test personnel. This is also the reason why most of the current document tests still stay at the proofreading of operation documents and involve less data processing design documents; moreover, the amount of data in the design document is large, the manual proofreading workload is large and the cost is high, the test progress is slow and the proofreading coverage is low. Summary of the Invention
[0005] The embodiments of the present invention provide a method, device, equipment and storage medium for testing a data mapping file, so as to be able to automatically test the data mapping file, improve the test efficiency and test sufficiency, and reduce the problems of high cost and high data error rate brought by manual proofreading.
[0006] In a first aspect, the embodiments of the present invention provide a method for testing a data mapping file, including:
[0007] Obtain a data mapping file, where the data mapping file includes a plurality of data mapping tables, and the data mapping tables are used to represent the field mapping relationship between two data tables;
[0008] Perform format normalization processing on each of the data mapping tables based on a data mapping template;
[0009] Perform blood relationship parsing on each of the data mapping tables after the normalization processing to determine the field blood relationship line of the data mapping file, and the field blood relationship line is used to represent the field mapping relationship between a plurality of data tables;
[0010] Test the fields in the field blood relationship line to generate a test report of the data mapping file.
[0011] In a second aspect, an embodiment of the present invention further provides a test device for a data mapping file, the device comprising:
[0012] An acquisition module, configured to acquire a data mapping file, where the data mapping file includes a plurality of data mapping tables; the data mapping table is used to represent the field mapping relationship between two data tables;
[0013] A processing module, configured to perform format normalization processing on each of the data mapping tables based on a data mapping template;
[0014] A determination module, configured to perform lineage parsing on each of the data mapping tables after the normalization processing to determine the field lineage of the data mapping file, where the field lineage is used to represent the field mapping relationship between a plurality of data tables;
[0015] A test module, configured to test the fields in the field lineage to generate a test report of the data mapping file.
[0016] Further, the performing format normalization processing on each of the data mapping tables based on a data mapping template includes:
[0017] For each data mapping table, obtain the column name of each column of data in the data mapping table;
[0018] Based on a common column name dictionary, determine the canonical column name corresponding to the column name, where the common column name dictionary includes the column names corresponding to the canonical column names, and the canonical column names include: source field name, source data table name, target field name, and target data table name;
[0019] Adjust the column sorting in the data mapping table so that the column numbers of the canonical column names in the data mapping template and in the data mapping table are the same.
[0020] Further, performing field parsing on each of the data mapping tables after the normalization processing to determine the field lineage includes:
[0021] For each adjusted data mapping table, obtain the source field name and the source data table name of the data mapping table to form a source data key-value pair;
[0022] Obtain the target field name and the target data table name of the data mapping table to form a target data key-value pair;
[0023] Based on the source data key-value pair and the target data key-value pair, determine the data chain mapping relationship of the data mapping table;
[0024] Based on the data chain mapping relationships of each of the data mapping tables, determine at least one field lineage.
[0025] Further, test the fields in the field bloodline to generate a test report for the data mapping file, including:
[0026] Determine the abnormal fields in the field bloodline, where the abnormal fields include at least one of the following: orphan fields, invalid fields, and error fields;
[0027] Obtain the attribute information of the abnormal fields, where the attribute information includes at least one of the following: the abnormal type of the abnormal field, the table name of the data table to which the abnormal field belongs, the source data table name corresponding to the abnormal field, and the target data table name corresponding to the abnormal field;
[0028] Generate a test report for the data mapping file based on the attribute information.
[0029] Further, determine the orphan table fields in the field bloodline, including:
[0030] For each field bloodline in the data mapping file, determine the starting field of the field bloodline and the starting data mapping table corresponding to the starting field;
[0031] If the starting field is the target field of the starting data mapping table and the source field of the starting data mapping table is empty, then determine the starting field as an orphan field.
[0032] Further, determine the invalid fields in the field bloodline, including:
[0033] For each field bloodline in the data mapping file, determine the starting field and the ending field of the field bloodline;
[0034] If the starting field or the starting data table corresponding to the starting field is not in the first interface set, then determine the starting field as an invalid field;
[0035] If the ending field or the ending data table corresponding to the ending field is not in the second interface set, then determine the ending field as an invalid field.
[0036] Further, determine the error fields in the field bloodline, including:
[0037] For each field bloodline in the data mapping file, determine the fields whose field content does not meet the preset requirements as error fields;
[0038] Among them, the preset requirements are determined based on the business requirements of the data mapping table or set by the user himself.
[0039] In a third aspect, an embodiment of the present invention further provides a terminal device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the test method for the data mapping file as described in any one of the embodiments of the present invention.
[0040] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the test method for the data mapping file as described in any one of the embodiments of the present invention.
[0041] In the embodiment of the present invention, by obtaining a data mapping file, the data mapping file includes a plurality of data mapping tables; performing format normalization processing on each of the data mapping tables based on a data mapping template; performing lineage parsing on each of the normalized data mapping tables to determine the field lineage of the data mapping file, where the field lineage is used to represent the mapping relationship between the source fields and target fields of each of the normalized data mapping tables; and testing the fields in the field lineage to generate a test report for the data mapping file, it is possible to automatically test the data mapping file, improve the test efficiency and test sufficiency, and reduce the problems of high cost and high data error rate caused by manual proofreading. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0043] Figure 1 is a flowchart of a test method for a data mapping file in Embodiment 1 of the present invention;
[0044] Figure 2 is a flowchart of a test method for a data mapping file in Embodiment 2 of the present invention;
[0045] Figure 3 is a schematic structural diagram of a test device for a data mapping file in Embodiment 3 of the present invention;
[0046] Figure 4 is a schematic structural diagram of a terminal device in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention and not for limiting the present invention. In addition, it should be noted that, for the sake of description, only the parts related to the present invention rather than all the structures are shown in the drawings.
[0048] It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.
[0049] Embodiment 1
[0050] Figure 1 The figure is a flowchart of a test method for a data mapping file provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of automatically testing a data mapping file. This method can be executed by a test device for a data mapping file in an embodiment of the present invention, and the device can be implemented in software and / or hardware.
[0051] A data mapping file is a normative file for data processing. For projects developed in the background, data processing and other secondary developments are often carried out on existing basic system files. Since project files often need to be jointly developed by a team, in order to keep the data types and names of variables, function names, interface names, etc. in the project files developed by each team member consistent, it is necessary to determine a data mapping file based on the data mapping relationship between the basic system file and the project file before project development, as a reference basis for later development. The data mapping file is used to record the mapping relationship of fields in the data tables of the basic system file and the project file.
[0052] Due to the large amount of data mapping, a data mapping file generally includes hundreds or thousands of data mapping tables, and each table contains multiple fields, which requires cooperation among multiple people. Therefore, it is necessary to ensure the standardization of the data mapping file first to ensure that the project files developed are uniformly recognizable.
[0053] As Figure 1 shown, the method specifically includes the following steps:
[0054] S110, obtain a data mapping file, where the data mapping file contains multiple data mapping tables, and the data mapping tables reflect the field mapping relationship between two data tables.
[0055] Among them, the data mapping file contains multiple data mapping tables, and the data mapping tables are used to reflect the mapping relationship between the fields of two data tables. The data table is a table composed of data such as variable names, function names, and interface names included in the program file. The data table is a two-dimensional table composed of rows and columns, and the field is a column in the data table. The data mapping table includes four fields: the target data table name, the target field name, the source data table name, and the source field name. Exemplarily, if the field a in data table A is mapped to the field b in data table B, the source data table name of the data mapping table is A, the source field name is a, the target data table name is B, and the target field name is b.
[0056] Exemplarily, the method for obtaining the data mapping file can be user input or can be called from the database by means of calling. The embodiments of the present invention do not limit this.
[0057] S120, perform format normalization processing on each data mapping table based on the data mapping template.
[0058] Among them, the data mapping template refers to a standard data mapping table and is used as a standard file for evaluating whether the designed data mapping table meets the specification requirements.
[0059] Specifically, based on the form requirements for the table and fields in the data mapping template, perform normalization processing on each data mapping table included in the data mapping file. The normalization processing may include at least one of the following: adjustment of the field order, supplementation of default fields, and deletion of redundant fields.
[0060] It should be noted that the format normalization processing normalizes the form problems of the data mapping table and cannot modify the field content in the data mapping table.
[0061] Optionally, before performing format normalization processing on each data mapping table, it further includes: determining whether the data mapping file can be subjected to format normalization processing. If not, generate an error report so that the tester can manually repair each data mapping table.
[0062] S130, perform blood relationship analysis on each normalized data mapping table to determine the field blood relationship line of the data mapping file. The field blood relationship line is used to represent the field mapping relationship between multiple data tables.
[0063] Specifically, for the normalized data mapping table, determine the mapping relationship between the fields reflected in the data mapping table, and perform blood relationship analysis on the field mapping relationship between the data tables reflected in each data mapping table to determine the field blood relationship line. One or more field blood relationship lines can be determined from the multiple data mapping tables in the data mapping file.
[0064] Exemplarily, the first data mapping table reflects that field a in data table A is mapped to field b in data table B, and the second data mapping table reflects that field b in data table B is mapped to field c in data table C. Then the field lineage is a - b - c.
[0065] S140. Test the fields in the field lineage to generate a test report for the data mapping file.
[0066] Specifically, test the field lineage to determine whether there are any abnormalities in each field in the field lineage, and generate a test report for the data mapping file for the abnormal fields. The test report is used to indicate the non - standard parts in the data mapping file and provide a basis for modification to testers and developers.
[0067] Exemplarily, the test report of the data mapping file may include the abnormal fields in the data mapping file, the types of the abnormal fields, or the reasons for the abnormalities.
[0068] The technical solution of this embodiment, by obtaining a data mapping file, where the data mapping file contains multiple data mapping tables; performing format normalization processing on each of the data mapping tables based on a data mapping template; performing lineage parsing on each of the normalized data mapping tables to determine the field lineage of the data mapping file, where the field lineage is used to represent the mapping relationship between the source fields and target fields of each of the normalized data mapping tables; testing the fields in the field lineage to generate a test report for the data mapping file, can automatically test the data mapping file, improve the test efficiency and test sufficiency, and reduce the problems of high cost and high data error rate caused by manual proofreading.
[0069] Optionally, performing format normalization processing on each of the data mapping tables based on a data mapping template includes:
[0070] For each data mapping table, obtain the column names of each column of data in the data mapping table;
[0071] Based on a common column name dictionary, determine the standard column names corresponding to the column names, where the common column name dictionary contains the column names corresponding to the standard column names, and the standard column names include: source field name, source data table name, target field name, and target data table name;
[0072] Based on the column numbers corresponding to the standard column names, adjust the column sorting in the data mapping table so that the arrangement order of the column names in the data mapping template is the same as that of the standard column names in the data mapping table.
[0073] Among them, the source data table name is the name of the data table to which the field corresponding to the source field name belongs; the target data table name is the name of the data table to which the field corresponding to the target field name belongs.
[0074] Specifically, the names and arrangement orders of the fields in the data mapping table and the data mapping template may be inconsistent, resulting in the inability to recognize the column names in the data mapping table during the testing process. Therefore, it is necessary to normalize the data mapping table so that the column names in the data mapping table are consistent with the data mapping template.
[0075] First, if the column names of the data mapping table are different from the standard column names of the data mapping template, the standard column names corresponding to the column names of the data mapping table are found through the common column name dictionary, and the column sorting in the data mapping table is adjusted based on the column numbers corresponding to the standard column names, so that the arrangement orders of the column names of the data mapping template and the standard column names of the data mapping table are the same.
[0076] Exemplarily, the column names of the data mapping table are a, A, B, b respectively, and the corresponding standard column names are source field name, source data table name, target data table name and target field name; in the data mapping template, the arrangement order of the standard column names is source field name, source data table name, target field name and target data table name, then the arrangement order of the column names of the data mapping table is adjusted to a, A, b, B.
[0077] Optionally, field parsing is performed on each normalized data mapping table to determine the field bloodline, including:
[0078] For each adjusted data mapping table, the source field name and source data table name of the data mapping table are obtained to form a source data key-value pair;
[0079] The target field name and target data table name of the data mapping table are obtained to form a target data key-value pair;
[0080] Based on the source data key-value pair and the target data key-value pair, the data chain mapping relationship of the data mapping table is determined;
[0081] Based on the data chain mapping relationships of each data mapping table, at least one field bloodline is determined.
[0082] Exemplarily, if the source field name, source data table name, target data table name and target field name of the first data mapping table are a, A, b, B respectively, the source data key-value pair is [a, A], and the target data key-value pair is [b, B], and the data chain mapping relationship of the first data mapping table is determined to be [a, A]—[b, B]; if the source field name, source data table name, target data table name and target field name of the second data mapping table are b, B, c, C respectively, the source data key-value pair is [b, B], and the target data key-value pair is [c, C], and the data chain mapping relationship of the first data mapping table is determined to be [b, B]—[c, C], then the field bloodline is [a, A]—[b, B]—[c, C].
[0083] Embodiment 2
[0084] Figure 2 This is a flowchart of a test method for a data mapping file in the second embodiment of the present invention. This embodiment is optimized based on the above embodiment. In this embodiment, the fields in the field bloodline are tested to generate a test report for the data mapping file, including: determining the abnormal fields in the field bloodline, where the abnormal fields include at least one of the following: orphan fields, invalid fields, and error fields; obtaining the attribute information of the abnormal fields, where the attribute information includes at least one of the following: the type of abnormality and the table name to which the abnormal field belongs; generating a field test report based on the attribute information.
[0085] As Figure 2 shown, the method of this embodiment specifically includes the following steps:
[0086] S210, obtain a data mapping file, where the data mapping file contains multiple data mapping tables, and the data mapping tables are used to represent the field mapping relationship between two data tables.
[0087] S220, perform format normalization processing on each of the data mapping tables based on a data mapping template.
[0088] S230, perform bloodline parsing on each of the data mapping tables after the normalization processing to determine the field bloodline of the data mapping file, where the field bloodline is used to represent the field mapping relationship between multiple data tables.
[0089] S240, determine the abnormal fields in the field bloodline, where the abnormal fields include at least one of the following: orphan fields, invalid fields, and error fields.
[0090] Among them, an orphan field refers to a starting field that cannot be traced back to the source of the field bloodline, that is, a field is a starting field in the field bloodline, but the source field of this field is empty. An invalid field refers to a starting field in the field bloodline or the starting table to which the starting field belongs is not in the first preset interface set, or the ending field in the field bloodline or the ending table to which the ending field belongs is not in the second preset interface set. An error field refers to a field in the field bloodline that does not meet the preset requirements.
[0091] S250, obtain the attribute information of the abnormal fields, where the attribute information includes at least one of the following: the type of abnormality of the abnormal field, the table name of the data table to which the abnormal field belongs, the name of the source data table corresponding to the abnormal field, and the name of the target data table corresponding to the abnormal field.
[0092] Specifically, after determining the abnormal fields, obtain the attribute information of the abnormal fields. The attribute information may include: the abnormal type of the abnormal fields and the table name of the data table to which the abnormal fields belong. The type of the abnormal fields is used to explain the reason for the abnormality of the abnormal fields. The table name of the data table to which the abnormal fields belong is used to explain the data table corresponding to the abnormal fields, which is convenient for testers or technical developers to find the abnormal fields in the data table for review. The source data table name corresponding to the abnormal fields and the target data table name of the abnormal fields are used to explain the data mapping table corresponding to the abnormal fields, which is convenient for testers or technical developers to correct the abnormal fields.
[0093] S260, generate a test report for the data mapping file based on the attribute information.
[0094] Specifically, generate a test report for the data mapping file according to the obtained attribute information of the abnormal fields. The test report records the abnormal fields, the types of the abnormal fields, the data mapping tables and data tables to which the abnormal fields belong. Testers or technical developers can correct and check for omissions in the fields in the data mapping tables in the data mapping file according to the test report.
[0095] The technical solution of this embodiment is to obtain a data mapping file, where the data mapping file contains multiple data mapping tables; the data mapping tables are used to represent the field mapping relationship between two data tables; perform format normalization processing on each of the data mapping tables based on a data mapping template; perform blood relationship parsing on each of the normalized data mapping tables to determine the field blood relationship line of the data mapping file, and the field blood relationship line is used to represent the field mapping relationship between multiple data tables; determine the abnormal fields in the field blood relationship line, and the abnormal fields include at least one of the following: orphan fields, invalid fields, and error fields; obtain the attribute information of the abnormal fields, and the attribute information includes at least one of the following: the abnormal type of the abnormal fields, the table name of the data table to which the abnormal fields belong, the source data table name corresponding to the abnormal fields, and the target data table name corresponding to the abnormal fields; generate a test report for the data mapping file based on the attribute information, which can automatically test the data mapping file, improve the test efficiency and test sufficiency, reduce the problems of high cost and high data error rate caused by manual proofreading, and give the attribute information of the abnormal fields in the test report, which is convenient for testers or technical developers to correct and check for omissions in the fields in the data mapping tables in the data mapping file.
[0096] Optionally, determining the orphan table fields in the field blood relationship line includes:
[0097] For each field blood relationship line in the data mapping file, determine the starting field of the field blood relationship line and the starting data mapping table corresponding to the starting field;
[0098] If the starting field is the target field of the starting data mapping table and the source field of the starting data mapping table is empty, the starting field is determined as an orphan field.
[0099] Exemplarily, if the field bloodline is a—b—c, the data mapping table where the starting field a is located is the starting data mapping table, and the starting field a is in the target field of the starting data mapping table, but the source field in the starting data mapping table is empty. Therefore, the source field corresponding to the starting field a cannot be traced and queried, and the starting field is determined as an orphan field. If the data mapping table where the starting field a is located is the starting data mapping table and the starting field a is in the source field of the starting data mapping table, the data table corresponding to the starting field a is the source data table of the data mapping file, and this field is not an orphan field.
[0100] Optionally, determining invalid fields in the field bloodline includes:
[0101] For each field bloodline in the data mapping file, determine the starting field and the ending field of the field bloodline;
[0102] If the starting field or the starting data table corresponding to the starting field is not in the first interface set, the starting field is determined as an invalid field;
[0103] If the ending field or the ending data table corresponding to the ending field is not in the second interface set, the ending field is determined as an invalid field.
[0104] Among them, the fields and data tables stored in the first interface set can be used for interaction with the associated source system interface, and the fields and data tables stored in the second interface set can be applied to interact with the associated business system interface.
[0105] Exemplarily, if the field bloodline is [a, A]—[b, B]—[c, C], the starting data table corresponding to the starting field a is data table A, and the ending data table corresponding to the ending field c is data table C. If the starting field a or the starting data table A corresponding to the starting field is not in the first interface set, the starting field a is determined as an invalid field; if the ending field c or the ending data table C corresponding to the ending field c is not in the second interface set, the ending field is determined as an invalid field.
[0106] Optionally, for each field bloodline in the data mapping file, determine the fields whose field content does not meet the preset requirements as error fields;
[0107] Among them, the preset requirements are determined based on the business requirements of the data mapping table or set by the user himself.
[0108] Specifically, due to different business requirements, different requirements for fields are also different. Therefore, the preset requirements for fields can be determined according to the business requirements of the data mapping table, or the preset requirements for fields can also be set by the forecaster himself. The field content can include field length, field type, and constituent elements of the field (such as Arabic numerals, uppercase and lowercase letters, whether special characters are included, etc.). Fields that do not meet the preset requirements are determined as error fields.
[0109] Embodiment III
[0110] Figure 3 The following is a schematic structural diagram of a test device for a data mapping file provided by Embodiment III of the present invention. This embodiment is applicable to the situation where a data mapping file is automatically tested. The device can be implemented in a software and / or hardware manner, and the device can be integrated in any device that provides the function of testing a data mapping file, such as Figure 3 As shown, the test device for the data mapping file specifically includes: an acquisition module 310, a processing module 320, a determination module 330, and a test module 340.
[0111] Among them, the acquisition module 310 is used to acquire a data mapping file, and the data mapping file contains multiple data mapping tables; the data mapping table is used to represent the field mapping relationship between two data tables;
[0112] The processing module 320 is used to perform format normalization processing on each of the data mapping tables based on a data mapping template;
[0113] The determination module 330 is used to perform blood relationship analysis on each of the data mapping tables after the normalization processing to determine the field blood relationship line of the data mapping file, and the field blood relationship line is used to represent the field mapping relationship between multiple data tables;
[0114] The test module 340 is used to test the fields in the field blood relationship line to generate a test report of the data mapping file.
[0115] Optionally, the processing module 320 is specifically used for:
[0116] For each data mapping table, obtain the column name of each column of data in the data mapping table;
[0117] Based on a common column name dictionary, determine the corresponding standard column name for the column name. The common column name dictionary contains the column names corresponding to the standard column names, and the standard column names include: source field name, source data table name, target field name, and target data table name;
[0118] Adjust the column sorting in the data mapping table so that the column numbers of the standard column names in the data mapping template and in the data mapping table are the same.
[0119] Optionally, the determining module 330 is specifically configured to:
[0120] For each adjusted data mapping table, obtain the source field name and source data table name of the data mapping table to form source data key-value pairs;
[0121] Obtain the target field name and target data table name of the data mapping table to form target data key-value pairs;
[0122] Based on the source data key-value pairs and the target data key-value pairs, determine the data chain mapping relationship of the data mapping table;
[0123] Based on the data chain mapping relationships of the data mapping tables, determine at least one field bloodline.
[0124] Optionally, the testing module 340 includes:
[0125] A determining unit, configured to determine abnormal fields in the field bloodline, where the abnormal fields include at least one of the following: orphan fields, invalid fields, and error fields;
[0126] An obtaining unit, configured to obtain attribute information of the abnormal fields, where the attribute information includes at least one of the following: the abnormal type of the abnormal field, the table name of the data table to which the abnormal field belongs, the source data table name corresponding to the abnormal field, and the target data table name corresponding to the abnormal field;
[0127] A generating unit, configured to generate a test report of the data mapping file based on the attribute information.
[0128] Optionally, the determining unit is specifically configured to:
[0129] For each field bloodline in the data mapping file, determine the starting field of the field bloodline and the starting data mapping table corresponding to the starting field;
[0130] If the starting field is the target field of the starting data mapping table and the source field of the starting data mapping table is empty, determine the starting field as an orphan field.
[0131] Optionally, the determining unit is specifically configured to:
[0132] For each field bloodline in the data mapping file, determine the starting field and the ending field of the field bloodline;
[0133] If the starting field or the starting data table corresponding to the starting field is not in the first interface set, determine the starting field as an invalid field;
[0134] If the termination field or the start / termination data table corresponding to the termination field is not in the second interface set, the termination field is determined as an invalid field.
[0135] Optionally, the determining unit is further configured to:
[0136] For each field bloodline in the data mapping file, determine a field with field content not meeting the preset requirements as an error field;
[0137] Wherein, the preset requirements are determined based on the business requirements of the data mapping table or set by the user himself.
[0138] The above product can execute the test method of the data mapping file provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0139] Embodiment 4
[0140] Figure 4 As shown in the structural block diagram of a terminal device provided in Embodiment 4 of the present invention, Figure 4 the terminal device includes a processor 410, a memory 420, an input device 430, and an output device 440; the number of processors 410 in the terminal device can be one or more, Figure 4 taking one processor 410 as an example; the processor 410, the memory 420, the input device 430, and the output device 440 in the terminal device can be connected through a bus or other means, Figure 4 taking the connection through the bus as an example.
[0141] The memory 420, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the test method of the data mapping file in the embodiment of the present invention (for example, the acquisition module 310, the processing module 320, the determination module 330, and the test module 340 in the test device of the data mapping file). The processor 410 executes various functional applications and data processing of the terminal device by running the software programs, instructions, and modules stored in the memory 420, that is, implements the above test method of the data mapping file.
[0142] The memory 420 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 420 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 420 may further include a memory remotely provided with respect to the processor 410, and these remote memories may be connected to the terminal device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0143] The input device 430 may be used to receive input digital or character information and generate key signal inputs related to user settings and function controls of the terminal device. The output device 440 may include a display device such as a display screen.
[0144] Embodiment Five
[0145] Embodiment Five of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the test method of the data mapping file provided in all the invention embodiments of the present application: obtaining a data mapping file, where the data mapping file contains a plurality of data mapping tables; the data mapping table is used to represent the field mapping relationship between two data tables; performing format normalization processing on each of the data mapping tables based on a data mapping template; performing lineage parsing on each of the normalized data mapping tables to determine the field lineage of the data mapping file, where the field lineage is used to represent the field mapping relationship between a plurality of data tables; testing the fields in the field lineage to generate a test report of the data mapping file.
[0146] One or more arbitrary combinations of computer-readable media may be adopted. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, be but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program may be used by or in combination with an instruction execution system, apparatus, or device.
[0147] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0148] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0149] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0150] Note that the above is only a preferred embodiment of the present invention and the applied technical principles. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments may be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A test method for a data mapping file, characterized in that Including: Obtain a data mapping file, which contains multiple data mapping tables; the data mapping tables are used to represent the field mapping relationships between two data tables; Perform format normalization processing on each of the data mapping tables based on a data mapping template; Perform lineage analysis on each of the normalized data mapping tables to determine the field lineage of the data mapping file, where the field lineage is used to represent the field mapping relationships between multiple data tables; Test the fields in the field lineage to generate a test report for the data mapping file; wherein, the test report records abnormal fields, the types of abnormal fields, the data mapping tables and data tables to which the abnormal fields belong; Among them, performing lineage analysis on each of the normalized data mapping tables to determine the field lineage includes: For each adjusted data mapping table, obtain the source field name and source data table name of the data mapping table to form a source data key-value pair; Obtain the target field name and target data table name of the data mapping table to form a target data key-value pair; Based on the source data key-value pair and the target data key-value pair, determine the data chain mapping relationship of the data mapping table; Based on the data chain mapping relationships of each of the data mapping tables, determine at least one field lineage.
2. The method according to claim 1, wherein The performing format normalization processing on each of the data mapping tables based on a data mapping template includes: For each data mapping table, obtain the column name of each column of data in the data mapping table; Based on a common column name dictionary, determine the corresponding standard column name for the column name, where the common column name dictionary contains the column names corresponding to the standard column names, and the standard column names include: source field name, source data table name, target field name, and target data table name; Adjust the column sorting in the data mapping table so that the column numbers of the standard column names in the data mapping template and in the data mapping table are the same.
3. The method according to claim 1, wherein Testing the fields in the field lineage to generate a test report for the data mapping file includes: Determine the abnormal fields in the field lineage, where the abnormal fields include at least one of the following: orphan fields, invalid fields, and error fields; Obtain the attribute information of the abnormal fields, where the attribute information includes at least one of the following: the abnormal type of the abnormal field, the table name of the data table to which the abnormal field belongs, the source data table name corresponding to the abnormal field, and the target data table name corresponding to the abnormal field; Generate a test report for the data mapping file based on the attribute information.
4. The method according to claim 3, characterized in that Determining the orphan table fields in the field lineage includes: For each field lineage in the data mapping file, determine the starting field of the field lineage and the starting data mapping table corresponding to the starting field; If the starting field is the target field of the starting data mapping table and the source field of the starting data mapping table is empty, then determine the starting field as an orphan field.
5. The method according to claim 3, characterized in that, Determining the invalid fields in the field lineage includes: For each field lineage in the data mapping file, determine the starting field and the ending field of the field lineage; If the starting field or the starting data table corresponding to the starting field is not in the first interface set, the starting field is determined as an invalid field; If the ending field or the starting and ending data table corresponding to the ending field is not in the second interface set, the ending field is determined as an invalid field.
6. The method according to claim 3, wherein Determining the error fields in the field bloodline includes: For each field bloodline in the data mapping file, determining the fields whose field content does not meet the preset requirements as error fields; Among them, the preset requirements are determined based on the business requirements of the data mapping table or set by the user himself.
7. A test device for a data mapping file, characterized in that, Including: An acquisition module for acquiring a data mapping file, the data mapping file containing multiple data mapping tables; the data mapping table is used to represent the field mapping relationship between two data tables; A processing module for performing format normalization processing on each of the data mapping tables based on a data mapping template; A determination module for performing bloodline analysis on each of the data mapping tables after normalization processing to determine the field bloodline of the data mapping file, the field bloodline being used to represent the field mapping relationship between multiple data tables; A test module for testing the fields in the field bloodline to generate a test report of the data mapping file; wherein, the test report records the abnormal fields, the types of the abnormal fields, the data mapping tables and data tables where the abnormal fields are located; Among them, the determination module is specifically used for: For each adjusted data mapping table, obtaining the source field name and source data table name of the data mapping table to form a source data key-value pair; Obtaining the target field name and target data table name of the data mapping table to form a target data key-value pair; Based on the source data key-value pair and the target data key-value pair, determining the data chain mapping relationship of the data mapping table; Determining at least one field bloodline based on the data chain mapping relationships of the data mapping tables.
8. A terminal device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the test method of the data mapping file as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the test method of the data mapping file as described in any one of claims 1-6.
Citation Information
Patent Citations
Blood relationship analysis method of structured query language and tool thereof
CN110232056A
Standard method and system for logging discrete data
CN112861508A
Metadata-based big data platform construction method, system, equipment and medium
CN113051263A
Database query SQL (Structured Query Language) field blood relationship generation method
CN115080599A