Data processing method and data processing device
By establishing multiple environments in the data center and analyzing the conversion pre-store program, generating and comparing the blood relationship diagram of the data table field, the difficulty and error risks of revising the data table field logic in the existing technology are solved, the system is efficient and accurate, and the company's digital transformation is supported.
Patent Information
- Application Number
- CN202311536004.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2025-05-16
AI Technical Summary
When revising or querying data table field logic, the existing technology faces the difficulty and error risks of verification caused by the large number of pre-stored programs, logical associations, fields with different meanings of the same name and wildcards, which affects the speed of digital transformation in the data center.
Provide a data processing method, by establishing a development environment, data quality verification environment and formal environment in a data center, analyzing and converting pre-stored programs to eliminate wildcards, generating and comparing data table fields ties of each environment, and performing notification functions to ensure system correctness.
Real-time comparison of blood relationship differences in data table fields in different environments has been achieved, which improves the system accuracy and timeliness of revising pre-stored programs, ensures that data changes meet demand, and improves the speed of the company's digital transformation.
Smart Images

Figure CN120011357A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data processing method and a data processing device, in particular to a data processing method and a data processing device capable of comparing the blood relationship differences of data table fields in different environments in real time. Background Art
[0002] Data centers usually include data tables and stored procedures. Data tables can be used as data carriers to store data. The data association logic and conversion rules in data tables are determined by stored procedures. Data tables contain multiple data table fields. When software engineers in data centers need to add, modify or delete data table fields, in addition to adjusting the data table form specifications, they must also find all related stored procedures and perform programming operations. However, there are a large number of stored procedures in data centers, and the logic between different stored procedures is also related. Therefore, when software engineers in data centers program any stored procedure, the blood relationship of the data table fields in the entire data center will change drastically. In addition, a data center may have multiple data tables with data fields with the same name. Even in stored procedures, wildcards "*" are often used to omit the program code of the data field name. However, problems such as data table fields or wildcards with the same name but different meanings have greatly increased the difficulty for data center software engineers to add, modify or query the logic of the fields in stored procedures. The existing verification method is to perform sampling verification on the data before and after the stored program is modified. However, this verification method is time-consuming and cannot confirm complete correctness. Moreover, many errors often rely on data users to discover anomalies and feedback to the software engineers in the data center. However, incorrect data will cause companies to suffer potential and unassessable losses, so how to correctly modify the logic of stored programs is very important. As more and more stored programs are created, the difficulty of correctly modifying stored programs increases, which in turn affects the speed at which companies promote digital transformation. Therefore, the existing technology really needs to be improved. Summary of the invention
[0003] In order to solve the above problems, the present invention provides a data processing method and a data processing device that can compare the blood relationship differences of data table fields in different environments in real time to solve the above problems.
[0004] The present invention provides a data processing method for a data center, comprising: obtaining, by the data center, stored programs of a development environment, a data quality verification environment, and a formal environment, and field information of all data tables; for each stored program, analyzing the stored program and determining whether the stored program contains a wildcard; for each stored program, converting and restoring the stored program into a restored stored program with complete field information based on the wildcard; generating, according to the restored stored programs of the development environment, the data quality verification environment, and the formal environment, blood relationship diagrams of the data table fields of the development environment, the data quality verification environment, and the formal environment; and comparing the blood relationship diagrams of the data table fields of the development environment, the data quality verification environment, and the formal environment to generate a comparison result and execute a notification function accordingly.
[0005] The present invention provides a data processing device for a data center, comprising: a storage device for storing instructions; and a processing circuit configured to execute the instructions, wherein the instructions include: obtaining from the data center the stored programs of a development environment, a data quality verification environment, and a formal environment, as well as field information of all data tables; for each stored program, analyzing the stored program and determining whether the stored program contains a wildcard; for each stored program, converting and restoring the stored program into a restored stored program with complete field information based on the wildcard; generating a blood relationship diagram of the data table fields of the development environment, the data quality verification environment, and the formal environment according to the restored stored programs of the development environment, the data quality verification environment, and the formal environment; and comparing the blood relationship diagrams of the data table fields of the development environment, the data quality verification environment, and the formal environment to generate a comparison result and execute a notification function accordingly. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 Schematic diagram of a data processing method in an embodiment of the present invention.
[0007] Figure 2 The figure is a flow chart of restoring a stored program into a restored stored program with complete field information based on wildcards in an embodiment of the present invention.
[0008] Figure 3 The diagram is a schematic diagram of converting a wildcard into complete field information in an embodiment of the present invention.
[0009] Figure 4 Schematic diagram of a wildcard not used to represent field information in an embodiment of the present invention.
[0010] Figure 5It is a schematic diagram of a directed graph of data table fields of a restored pre-stored program in an embodiment of the present invention.
[0011] Figure 6 It is a schematic diagram of a source directed subgraph in a blood relationship diagram of a data table field in an embodiment of the present invention.
[0012] Figure 7 It is a schematic diagram of a target directed subgraph in a blood relationship graph of data table fields in an embodiment of the present invention.
[0013] Figure 8 It is a comparative schematic diagram of the blood relationship diagram of the data table fields of the development environment and the data quality verification environment in the embodiment of the present invention.
[0014] Fig. 9 It is a schematic diagram for comparing the blood relationship diagram of data table fields in the data quality verification environment and the formal environment in an embodiment of the present invention.
[0015] Component number description
[0016] 10 Process
[0017] 300, 400 query statements
[0018] 302 First Clause
[0019] 304 Second Clause
[0020] 502, 506, 508, 512, 514 nodes
[0021] 504, 510 Directed edges
[0022] 602 Source directed subgraph
[0023] 604, 604 Input Menu
[0024] 702 Target directed subgraph
[0025] S100, S102, S104, S106, S1060, S1062, S1064,
[0026] S1066, S1068, S1070, S108, S110, S112 steps DETAILED DESCRIPTION
[0028] The embodiment of the present invention can be applied to a data center or a database system. In order to prevent errors caused by adding or modifying stored programs, the embodiment of the present invention can establish a development environment, a data quality assurance system environment, and a formal environment in the data center. The development environment can be used to add or modify stored programs and confirm that the stored programs can be executed normally after the addition or modification without errors in the program coding. In the development environment, developers or software engineers can load original data and stored programs and add or modify the stored programs according to requirements. After adding or modifying the stored programs, the new version of the data and stored programs in the development environment that have been added or modified can be copied to the data quality assurance environment. The data quality assurance environment can provide data users with confirmation whether the data changes meet the requirements after the stored programs are added or modified. In the data quality assurance environment, data users are provided with various testing methods to perform tests to determine whether the data changes meet the requirements. When any problems, flaws or defects are found during testing, developers or software engineers can process the problems found in the data quality verification environment in the development environment, and then copy the processed data and stored programs to the data quality verification environment to conduct relevant tests again until the expected results are obtained or no problems occur. The formal environment is the real operating data center environment. Therefore, when the data quality verification environment completes the data user test and confirms that the data changes meet the requirements, the data center developers or software engineers can upload the modified new version of the stored program to the formal environment.
[0029] Please refer to Figure 1 , Figure 1 FIG. 1 is a schematic diagram of a process 10 of a data processing method in an embodiment of the present invention. The process 10 includes the following steps:
[0030] Step S100: Start.
[0031] Step S102: The data center obtains the pre-stored programs of the development environment, data quality verification environment, and formal environment and the field information of all data tables.
[0032] Step S104: For each stored program, analyze the stored program and determine whether the stored program contains a wildcard.
[0033] Step S106: For each stored procedure, convert and restore the stored procedure into a restored stored procedure with complete field information based on the wildcard.
[0034] Step S108: Generate blood relationship diagrams of data table fields of the development environment, data quality verification environment and formal environment respectively according to the restored pre-stored programs of the development environment, data quality verification environment and formal environment.
[0035] Step S110: Compare the blood relationship diagrams of the data table fields in the development environment, the data quality verification environment, and the formal environment to generate a comparison result and execute a notification function accordingly.
[0036] Step S112: End.
[0037] According to process 10, in step S102, the embodiment of the present invention can obtain the stored programs of the development environment, the stored programs of the data quality verification environment, and all the stored programs of the formal environment and the field information of all data tables from the data center.
[0038] Because software engineers often use wildcards to perform the action of selecting all fields of a data table when writing the program code of a programming pre-stored program. However, the use of wildcards will result in the inability to know that all field information represented by the wildcards has fields that are ignored and omitted when parsing the pre-stored program, thereby parsing incorrect data table field blood relations. In step 104, for each pre-stored program in the development environment, data quality verification environment, and formal environment, the program code of the pre-stored program can be analyzed and it can be judged whether there is a wildcard in the program code of the pre-stored program. For example, the wildcard can be represented by an asterisk "*", but it is not limited thereto. When it is judged that there is a wildcard in the program code of the pre-stored program, step 106 is executed. When it is judged that there is no wildcard in the program code of the pre-stored program, it is not necessary to modify the pre-stored program and directly output the pre-stored program as the pre-stored program after restoration.
[0039] In step 106, for each stored procedure in the development environment, the data quality verification environment, and the formal environment, the stored procedure can be converted and restored into a restored stored procedure with complete field information based on the wildcard. For detailed operation methods of converting and restoring the stored procedure into a restored stored procedure with complete field information based on the wildcard, please refer to Figure 2 . In step 1060, for each stored procedure, when it is determined that a wildcard exists in the program code of the stored procedure, the context information of the wildcard in the stored procedure can be further analyzed. Then, in step 1062, it can be determined whether the wildcard is in a query statement based on the context information of the wildcard in the stored procedure. When it is determined that the wildcard is in a query statement, step 1064 is executed. When it is determined that the wildcard is not in a query statement, step 1070 is executed.
[0040] For example, the query statement may be a SELECT query statement (select query statement), an UPDATE query statement, a WHERE query statement, or an OREDR query statement, but is not limited thereto. Figure 3 , the syntax of a SELECT query statement 300 may be as follows Figure 3 As shown in the upper part, the SELECT query statement 300 includes a first clause 302 (also referred to as a SELECT clause) that begins with a SELECT command and a second clause 304 (also referred to as a FROM clause) that begins with a FROM command. The second clause 304 follows the first clause 302. The "SELECT" in the first clause 302 represents a SELECT command, and the "FROM" in the second clause 304 represents a FROM command. The first clause 302 that begins with a SELECT command is used to indicate the field information to be returned in this query, and the second clause 304 that begins with a FROM command is used to indicate which data table (hereinafter referred to as the designated data table) the query is to be performed on. In one embodiment, in step S1062, it can be determined whether there is program code that conforms to the query statement format in the context of the wildcard. For example, if Figure 3 As shown in the upper part, when it is determined that there is a wildcard "*" in the program code of the stored program and there is a first clause 302 starting with a SELECT command and a second clause 304 starting with a FROM command in the context of the wildcard "*", it can be determined that the combination of the first clause 302 and the second clause 304 is a SELECT query statement 300. Therefore, it can be determined that the wildcard "*" is located in the SELECT query statement 300.
[0041] In another embodiment, please continue to refer to Figure 3 When it is determined that there is a wildcard "*" in the program code of the stored program, there is a SELECT command before the wildcard "*", and there is a FROM command before the wildcard "*", it can be further determined whether there are more than one space after the SELECT command, more than one space after the wildcard "*", and more than one space after the FROM command. When it is determined that there are more than one space after the SELECT command, more than one space after the wildcard "*", and more than one space after the SELECT command, it is determined that the wildcard "*" is located in a query statement.
[0042] In another embodiment, please continue to refer to Figure 3 When it is determined that there is a wildcard "*" in the program code of the stored program, there is a SELECT command before the wildcard "*", and there is a FROM command before the wildcard "*", it can be further determined that there is a wildcard "*" between the SELECT command and the wildcard "*" (i.e. Figure 3 The upper part (a) shows whether there are more than one space and other symbols or commands. The SELECT command and the FROM command are case insensitive. The symbols may include any cursor control symbols, such as \n, \r, or \t. The commands may include pre-stored program-specific commands for filtering column data, such as DISTINCT (case insensitive) and TOP (case insensitive). It can be determined between the wildcard "*" and the FROM command (i.e. Figure 3 The upper part (b) shows whether there is more than one space and other symbols or commands. You can judge whether there is more than one space and other symbols or commands between the FROM command and the subsequent data table name (i.e. Figure 3 The upper part (c) shows whether there are more than one space and other symbols or commands. If the results of the above judgments are all yes, it is determined whether the wildcard "*" is located in a query statement.
[0043] In addition, in step 1062, when it is determined that the wildcard is in a query statement, it can be further determined whether the wildcard "*" in the select query statement is related to the field information of the data table. When it is determined that the wildcard is not related to the field information, it is not necessary to modify the stored program, and then step 1070 is executed to output the stored program as the restored stored program. For example, Figure 4 The selection query statement 400 shown, when it is determined that the wildcard "*" is in the first clause at the beginning of the SELECT command, can further determine whether the wildcard "*" is related to the field information in the selection query statement. When it is determined that the wildcard "*" is used to indicate a multiplication operation in the selection query statement. That is to say, the "PricePerUnit*Volume AS Amount" described in the SELECT clause of the selection query statement indicates that the value Amount of each column of data is the result of multiplying the domain value PricePerUnit of each column of data by the domain value Volume. Therefore, it is determined that the wildcard "*" is used for multiplication operation in the selection query statement and is related to the field information, and there is no need to modify the stored program, and then step 107 is executed to output the stored program as a restored stored program.
[0044] In step 1064, after determining that the wildcard is in the query statement, the specified data table to be queried by the query statement can be determined based on the recorded content of the query statement. Figure 3 In the second clause 304 of the SELECT query statement 300, the content following the (C) portion records the Inventory data table. When it is determined that the wildcard "*" is located in the SELECT query statement 300, it can be determined that the designated data table to be queried by the query statement is the Inventory data table according to the Inventory data table recorded in the second clause 304 of the SELECT query statement 300.
[0045] In step 1066, the field information of the designated data table is compared with the field information of all data tables obtained by the data center to determine all field information corresponding to the designated data table. It can be analyzed and queried whether the field information of all data tables obtained by the data center includes all field information of the designated data table. If so, read all field information corresponding to the designated data table. Assume that the field information of all data tables obtained by the data center records all fields of the Inventory data table including the Name field, the PricePerUnit field, and the Volume field. When step 1064 determines that the designated data table is the Inventory data table, the field information of all data tables obtained by the data center can be queried to determine that all field information corresponding to the designated data table (Inventory data table) includes the Name field, the PricePerUnit field, and the Volume field.
[0046] In step 1068, all the field information corresponding to the designated data table is used to replace the wildcard in the stored procedure to generate a restored stored procedure with complete field information. For example, all the field information corresponding to the designated data table can be used to replace the wildcard in the stored procedure according to the program code writing rules of the stored procedure to generate a restored stored procedure with complete field information. For example, please continue Figure 3 ,like Figure 3 As shown in the lower half of , the Name field, the PricePerUnit field, and the Volume field can be concatenated to create a string and this substring can replace the wildcard "*" in the first clause 302 of the SELECT query statement 300 to generate a restored stored procedure with complete field information. In this way, all field information corresponding to the Inventory data table (i.e., the Name field, the PricePerUnit field, and the Volume field) can be clearly presented in the SELECT query statement 300, and the field blood relationship information can be correctly provided during the subsequent blood relationship analysis of the data table fields. In other words, since it is impossible to know all the field information represented by the wildcard when parsing the stored procedure, it is impossible to obtain the blood relationship of the data table fields simply by parsing the stored procedure. In this embodiment, the stored procedure is converted and restored into a restored stored procedure with complete field information based on the wildcard in step S106. In the subsequent blood relationship analysis of the data table fields, no fields will be ignored or omitted, resulting in incorrect blood relationship analysis of the data table fields.
[0047] Through step S106, all the restored stored programs of the development environment, the data quality verification environment, and the formal environment can be converted and restored into restored stored programs with complete field information. Then, in step S108, the blood relationship diagrams of the data table fields of the development environment, the data quality verification environment, and the formal environment are generated respectively according to the stored programs of the development environment, the stored programs of the data quality verification environment, and the restored stored programs of the formal environment. In one embodiment, the blood relationship diagram of the data table fields of the development environment is generated according to all the restored stored programs of the development environment. Each restored stored program of the development environment can be parsed to determine the directed graph of the data table fields of each restored stored program. For example, the stored program parsing engine can be used to parse the restored stored program to generate the directed graph of the data table fields. For example, the SQL parser parsing engine of the Python program software can be used to parse each restored stored program and generate the directed graph of the data table fields corresponding to the restored stored program. Each stored program may include at least one source data table field and at least one target data table field. Each stored procedure may include only at least one source data table field. Each stored procedure may also include only at least one target data table field. The directed graph of the data table fields of the stored procedure may include nodes for representing the source data table fields and / or the target data table fields. The directed graph of the data table fields of the stored procedure may include directed edges for representing the stored procedure, and the direction of the directed edges is from the source data table field to the target data table field. Each data table field of the stored procedure may correspond to at least one directed graph.
[0048] For example, see Figure 5 ,like Figure 5 As shown in the upper left part, the first stored procedure after restoration is INT_CUSTOMER_SATISFACTION_SP.STOREDPROCEDURE.SQL. A data table field directed graph of the first stored procedure after restoration includes a node 502, a node 506 and a directed edge 504. Node 502 is used to represent a target data table field of the first stored procedure after restoration. Figure 5As shown in the upper left part, the target data table field of the first restored stored procedure is the data table field TP1_INT_CUSTOMER_SATISFACTION.CUSTOMER. Node 506 is used to represent a source data table field of the first restored stored procedure, wherein the source data table of the first restored stored procedure is the data table field ODS_CUSTOMER_SATISFACTION.CUSTOMER. Directed edge 504 is used to represent the aforementioned first restored stored procedure, wherein the direction of directed edge 504 is from node 506 of the source data table to node 502 of the target data table field. The second restored stored procedure is ODS_CUSTOMER_SATISFACTION_SP.STOREDPROCEDURE.SQL, and a data table field directed graph of the second restored stored procedure includes node 508, node 512, and directed edge 510. Node 508 is used to represent a target data table field of the second restored stored program, wherein the target data table field of the second restored stored program is the data table field ODS_CUSTOMER_SATISFACTION.CUSTOMER. Node 512 is used to represent a source data table field of the second restored stored program, wherein the source data table field of the second restored stored program is the data table field TP1_ODS_CUSTOMER_SATISFACTION.CUSTOMER. Directed edge 510 is used to represent the second restored stored program. The direction of directed edge 510 is from node 512 of the source data table to node 508 of the target data table.
[0049] In step S108, after parsing the directed graph of the data table fields of all restored stored programs of the development environment, a lineage graph of a data table field of the development environment can be generated by merging the directed graphs of the data table fields of all restored stored programs of the development environment. The lineage graph of the data table fields can represent the relationship between the data table fields. For example, the directed graphs of the data table fields of all restored stored programs of the development environment can be compared. When it is determined that a source data table field of a restored stored program and a target data table field of another restored stored program are the same data table field, the node used to represent the source data table field of the restored stored program and the node used to represent the target data table field of another restored stored program can be merged to form a corresponding data table field lineage graph. For example, when it is determined that the source data table field of the first restored stored program and the target data table field of the second restored stored program are the same data table field (i.e., data table field ODS_CUSTOMER_SATISFACTION_CUSTOMER_SATISFACTION.CUSTOMER), nodes 506 and 508 can be merged to form node 514, so as to combine the two directed graphs to form a part of the data table field lineage relationship graph. In this case, node 514 is used to represent the source data table field of the first restored stored program and the target data table field of the second restored stored program, and the directed graph of the data table field of the first restored stored program and the directed graph of the data table field of the second restored stored program are connected and merged. Through the above-mentioned merging method, the directed graphs of the data table fields in all the restored stored programs of the development environment can be merged and converted into the lineage relationship graph of the data table fields of the development environment. The lineage relationship graph of the data table fields may include at least one source directed subgraph and / or at least one target directed subgraph. The source directed subgraph can be a directed subgraph with a data table field as the end point. The target directed subgraph can be a directed subgraph with a data table field as the starting point. For example, please refer to Figure 6 , Figure 6 FIG. 1 is a schematic diagram of a source directed subgraph in a blood relationship diagram of a data table field in an embodiment of the present invention. Figure 6 As shown, Figure 6 The middle part shows a source directed subgraph 602 corresponding to the data table field INT_CUSTOMER_SATISFACTION.CUSTOMER. The node used to represent the data table field INT_CUSTOMER_SATISFACTION.CUSTOME in the source directed subgraph 602 is the end node. The source directed subgraph 602 includes parent nodes related to the data table field INT_CUSTOMER_SATISFACTION.CUSTOME. For example, please refer to Figure 7 , Figure 7It is a schematic diagram of a target directed subgraph in a blood relationship graph of data table fields in an embodiment of the present invention. Figure 7 The middle part shows a target directed subgraph 702 corresponding to the data table field INT_CUSTOMER_SATISFACTION.CUSTOMER. The node used to represent the data table field INT_CUSTOMER_SATISFACTION.CUSTOMER in the target directed subgraph 702 is a starting node. The target directed subgraph 702 includes child nodes related to the data table field INT_CUSTOMER_SATISFACTION.CUSTOMER.
[0050] Similarly, according to the aforementioned method of generating the blood relationship graph of the data table fields of the development environment, the blood relationship graph of the data table fields of the data quality verification environment can be generated according to the restored stored programs of the data quality verification environment. Each restored stored program of the data quality verification environment can be analyzed to determine the directed graph of all data table fields of each restored stored program, and after parsing the directed graph of the data table fields of all restored stored programs of the data quality verification environment, the directed graphs of the data table fields of all restored stored programs of the data quality verification environment can be merged to generate the blood relationship graph of all data table fields of the data quality verification environment. Similarly, the blood relationship graph of the data table fields of the formal environment can be generated according to the restored stored programs of the formal environment. Each restored stored program of the formal environment can be analyzed to determine the directed graph of all data table fields of each restored stored program, and after parsing the directed graph of the data table fields of all restored stored programs of the formal environment, the directed graphs of the data table fields of all restored stored programs of the formal environment can be merged to generate a data table blood relationship graph of the formal environment.
[0051] In one embodiment, after the blood relationship diagram of the data table fields of the development environment, the data quality verification environment, and the formal environment is generated, an input menu can be generated for the user to input and select. Figure 6 and Figure 7 ,like Figure 6 and Figure 7Input menus 604 and 704 are shown. The user can select the desired environment type, data table type (schema), data table name, domain name and other fields through the input menu. For example, in the environment type field, click the item "DEV" to select the development environment, in the data table type field, click the item "INT" to select the INT data table, in the data table name field, click the item "INT_CUSTOMER_SATISFACTION" to select the data table INT_CUSTOMER_SATISFACTION, and in the domain name field, click the item "CUSTOMER" to select the field CUSTOMER. This embodiment can visually display the input menu for the user to view and input the items they want to select, and visually display the relevant information of the blood relationship diagram of the selected data table field. Figure 6 and Figure 7 As shown, this embodiment can visualize the input menu for the user to view and input the desired item, and visualize the content of the source directed subgraph of the selected data table field and the corresponding restored stored program, source data table field, target data table field, etc. In this way, it can provide users (such as administrators or engineers of the notification data center) with a quick and clear understanding of the data structure relationship of the restored stored program data table fields in the corresponding environment.
[0052] In step 110, the data table lineage diagrams of the development environment, the data quality verification environment, and the formal environment can be compared to generate a comparison result and perform a notification function accordingly. For example, the lineage diagrams of the data table fields of at least two environments among the development environment, the data quality verification environment, and the formal environment are compared to generate a comparison result. For example, the lineage diagrams of the data table fields of any two environments among the development environment, the data quality verification environment, and the formal environment are compared to generate a comparison result. For example, the lineage diagrams of the data table fields of the development environment and the data quality verification environment are compared to generate a comparison result. The lineage diagrams of the data table fields of the data quality verification environment and the formal environment are compared to generate a comparison result. The lineage diagrams of the data table fields of the development environment and the formal environment are compared to generate a comparison result. In addition, specific data table fields can also be selected for comparison to generate a comparison result. For example, the source directed subgraph or the target directed subgraph corresponding to at least one data table field of the data table lineage diagram of any two environments among the development environment, the data quality verification environment, and the formal environment are compared to generate a comparison result. For example, common types of data tables include DM data tables, INT data tables, ODS data tables, and STAGE data tables. A data table whose name is prefixed with DM is called a DM data table, such as data table DM_XXXX. A data table whose name is prefixed with INT is called an INT data table, and so on. Among them, the DM data table is usually located at the end point of the data table blood relationship. Since the DM data table is usually the end point data table in the blood relationship, if there are differences and inconsistencies between two environments, they can be quickly and easily compared. In one embodiment, the source directed subgraph corresponding to the fields of the DM data table in the blood relationship diagram of the data table fields of any two environments among the development environment, the data quality verification environment, and the formal environment can be compared to produce a comparison result.
[0053] In step S110, a notification function may be executed according to the aforementioned comparison result. When the comparison result shows that there is a difference in the blood relationship diagram of the data table fields of any two environments, the embodiment of the present invention may generate and send a notification signal to remind the administrator or engineer of the data center to execute the notification function. For example, Figure 8 A source directed subgraph corresponding to the data table field INT_CUSTOMER_X.CUSTOMER in the blood relationship diagram of the data table fields of the development environment and the data quality verification environment is respectively displayed. Fig. 9 The source directed subgraph corresponding to the data table field INT_CUSTOMER_X.CUSTOMER in the blood relationship diagram of the data table fields in the data quality verification environment and the formal environment is shown respectively. Figure 8 and Fig. 9As shown in the figure, in the development environment, data quality verification environment and formal environment, the data table field INT_CUSTOMER_X.CUSTOMER is the end point of the source directed subgraph. After comparing the data quality verification environment with the development environment, the comparison result shows that the data quality verification environment and the development environment have the same source directed subgraph, while the source directed subgraph of the formal environment is different from that of the development environment and the data quality verification environment. Figure 8 As shown in the figure, when the software engineer selects the development environment and the data quality verification environment for comparison, the comparison result shows that the source directed subgraph corresponding to the data table field INT_CUSTOMER_X.CUSTOMER of the development environment and the source directed subgraph corresponding to the data table field INT_CUSTOMER_X.CUSTOMER of the data quality verification environment are the same. Fig. 9 As shown in the figure, the software engineer selects the formal environment and the data quality verification environment for comparison. The comparison result shows that the source directed subgraph corresponding to the data table field INT_CUSTOMER_X.CUSTOMER in the formal environment is different from the source directed subgraph corresponding to the data table field INT_CUSTOMER_X.CUSTOMER in the data quality verification environment. Fig. 9 As shown, the source data table fields of the restored stored procedure INT_CUSTOMER_X_SP.STOREDPROCEDURE.SQL corresponding to level 0 are inconsistent. More specifically, in the formal environment, the source data table field of the restored stored procedure INT_CUSTOMER_X_SP.STOREDPROCEDURE.SQL corresponding to level 0 is the data table field TP1_INT_CUSTOMER_X.CUSTOMER. However, in the data quality verification environment, the source data table field of the restored stored procedure INT_CUSTOMER_X_SP.STOREDPROCEDURE.SQL corresponding to level 0 is the data table field ODS_CUSTOMER_X.CUSTOMER. This also means that the update operation of the data in the quality verification environment and the restored stored procedure when it is deployed to the formal environment may be abnormal or the update operation has not yet been performed. In this case, the comparison result shows that there are differences in the source directed subgraphs in the blood relationship graph of the data table fields of the formal environment data and the quality verification environment. Based on the comparison result showing that there are differences between the formal environment data and the quality verification environment, a notification signal can be generated and sent to notify the administrator or engineer of the data center to perform the notification function. In this way, the administrator or engineer of the data center can easily and immediately identify whether there is an error based on the notification signal.
[0054] Therefore, during the application service development process, when the software engineer adds or modifies the restored stored program in the development environment, he can compare the lineage diagram of the data table fields of the development environment with the lineage diagram of the data table fields of the data quality verification environment to determine whether there is a difference between the two, so as to confirm the correctness of the overall system after the addition or modification of the restored stored program and greatly improve the timeliness of application service development. During the application service development process, when the restored stored program after the addition or modification is synchronously updated to the formal environment, the software engineer can compare the lineage diagram of the data table of the formal environment with the data quality verification environment to determine whether there is a difference between the two, so as to confirm whether the operation of adding or modifying the stored program to the formal environment is correctly executed. When an error occurs in the boarding operation, the embodiment of the present invention will be able to quickly find the affected stored program. In one embodiment, a monitoring cycle can be set, such as daily, weekly, or every certain time, so that the data center executes the steps of process 10 every monitoring cycle to compare the blood relationship diagrams of the data table fields in the development environment, the data quality verification environment, and the formal environment, and transmits the comparison results back to the data center, thereby realizing an automated testing and notification mechanism for the software engineers in the data center and avoiding manual input errors by personnel to achieve complete test automation. At the same time, automated testing will also greatly improve the efficiency of the production testing process.
[0055] A person skilled in the art may combine, modify or change the above-described embodiments according to the spirit of the present invention, but is not limited thereto. All of the above statements, steps, and / or processes (including recommended steps) may be implemented by hardware, software, firmware (i.e., a combination of hardware devices and computer instructions, where the data in the hardware devices are read-only software data), electronic systems, or a combination of the above devices. The hardware may include analog, digital and hybrid circuits (i.e., microcircuits, microchips or silicon chips). For example, the hardware may be an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic component, a coupled hardware component, or a combination of the above hardware. In other embodiments, the hardware may include a general purpose processor, a microprocessor, a controller, a digital signal processor (DSP), or a combination of the above hardware. The software may be a combination of program codes, a combination of instructions and / or a combination of functions (functionality), which is stored in a storage device, such as a computer readable recording medium or a non-transitory computer-readable medium. For example, a computer-readable recording medium may include a read-only memory (ROM), a flash memory (Flash Memory), a random-access memory (RAM), a subscriber identity module (SIM), a hard disk, a floppy disk, or a compact disk read-only memory (CD-ROM / DVD-ROM / BD-ROM), but is not limited thereto. An embodiment of the present invention may include a data processing device applied to a data center, the data processing device including a processing circuit and a storage device. The process steps and embodiments of the present invention may be compiled into a program code or instruction form and stored in the storage device of the data processing device. The processing circuit of the data processing device can be used to read and execute the program code or instructions stored in the storage device to implement all the aforementioned steps and functions.
[0056] In summary, the embodiment of the present invention converts and restores the stored program into a restored stored program with complete field information based on wildcards. In the subsequent parsing of the lineage of the data table fields, no fields will be ignored or omitted, resulting in the parsing of incorrect data table field lineage. In addition, the data processing method of the embodiment of the present invention can obtain the data table lineage diagram of each environment. When software engineers develop application services for the data center, they can compare the differences in data table lineages in different environments in real time, which will effectively improve the system correctness and timeliness when adding and revising stored programs. At the same time, it can also confirm whether the added and revised stored programs are indeed synchronized with the formal environment, thereby effectively improving and optimizing the speed of the company's digital transformation.
[0057] The above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.
Claims
1. A data processing method, characterized in that: For a data center, including: Obtaining from the data center the pre-stored programs of a development environment, a data quality verification environment, and a formal environment and field information of all data tables; For each pre-stored program, analyzing the pre-stored program and determining whether the pre-stored program contains a wildcard; For each stored procedure, converting and restoring the stored procedure into a restored stored procedure with complete field information based on the wildcard; Generate a blood relationship diagram of the data table fields of the development environment, the data quality verification environment and the formal environment respectively according to the restored pre-stored programs of the development environment, the data quality verification environment and the formal environment; and The blood relationship diagrams of the data table fields of the development environment, the data quality verification environment and the formal environment are compared to generate a comparison result and execute a notification function accordingly.
2. The data processing method according to claim 1, characterized in that: The step of converting and restoring the stored program into a restored stored program with complete field information based on the wildcard comprises: Analyzing context information of the wildcard in a stored procedure to determine whether the wildcard is in a select query statement; When it is determined that the wildcard is in the select query statement, a specified data table to be queried in the select query statement is determined, wherein the select query statement includes a first clause and a second clause, the second clause is subsequent to the first clause, and the wildcard is in the first clause of the select query statement, wherein the step includes determining the specified data table to be queried in the select query statement from the second clause of the select query statement; Comparing the field information of the designated data table with the field information of all the acquired data tables to determine all the field information corresponding to the designated data table; and The wildcard in the stored procedure is replaced by all the field information corresponding to the designated data table to generate a restored stored procedure with complete field information.
3. The data processing method according to claim 1, characterized in that: The step of converting and restoring each stored program into a restored stored program with complete field information based on the wildcard comprises: For each pre-stored program, when it is determined that the pre-stored program does not have a wildcard, the pre-stored program is output as a restored pre-stored program.
4. The data processing method according to claim 1, characterized in that: The steps of generating the blood relationship diagrams of the data table fields of the development environment, the data quality verification environment and the formal environment respectively according to the restored pre-stored programs of the development environment, the data quality verification environment and the formal environment include: For each of the development environment, the data quality verification environment, and the formal environment, parsing each restored stored program to determine a directed graph of all data table fields of each restored stored program; and The directed graphs of the data table fields of all restored stored procedures in each environment are merged to generate a lineage graph of a data table field of each environment.
5. The data processing method according to claim 4, characterized in that: The step of merging the directed graphs of the data table fields of all restored stored programs in each environment to generate a blood relationship graph of the data table fields of each environment comprises: When it is determined that a source data table field of a first stored program after restoration and a target data table field of a second stored program after restoration are the same data table field, a first node of the source data table used to represent the first stored program after restoration and a second node of the target data table used to represent the second stored program after restoration are merged.
6. The data processing method according to claim 1, characterized in that: The steps of comparing the blood relationship diagrams of the data table fields of the development environment, the data quality verification environment, and the formal environment to generate the comparison result and execute the notification function accordingly include: Comparing a source directed subgraph or a target directed subgraph corresponding to at least one data table field of a blood relationship graph of data table fields of at least two environments among the development environment, the data quality verification environment, and the formal environment to generate the comparison result; and When the comparison result shows that there are differences in the blood relationship diagrams of the data table fields of any two environments among the development environment, the data quality verification environment and the formal environment, a notification signal is generated and sent to execute the notification function.
7. A data processing device, characterized in that: For a data center, including: a storage device for storing instructions; and a processing circuit configured to execute the instructions, wherein the instructions include: Obtaining from the data center the pre-stored programs of a development environment, a data quality verification environment, and a formal environment and field information of all data tables; For each pre-stored program, analyzing the pre-stored program and determining whether the pre-stored program contains a wildcard; For each stored procedure, converting and restoring the stored procedure into a restored stored procedure with complete field information based on the wildcard; Generate a blood relationship diagram of the data table fields of the development environment, the data quality verification environment and the formal environment respectively according to the restored pre-stored programs of the development environment, the data quality verification environment and the formal environment; and The blood relationship diagrams of the data table fields of the development environment, the data quality verification environment and the formal environment are compared to generate a comparison result and execute a notification function accordingly.
8. The data processing device according to claim 7, characterized in that: The instructions also include: Analyzing context information of the wildcard in a stored procedure to determine whether the wildcard is in a select query statement; When it is determined that the wildcard is in the select query statement, a specified data table to be queried in the select query statement is determined, wherein the select query statement includes a first clause and a second clause, the second clause is subsequent to the first clause, and the instruction includes determining the specified data table to be queried in the select query statement from the second clause of the select query statement; Comparing the field information of the designated data table with the field information of all the acquired data tables to determine all the field information corresponding to the designated data table; and The wildcard in the stored procedure is replaced by all the field information corresponding to the designated data table to generate a restored stored procedure with complete field information.
9. The data processing device according to claim 7, characterized in that: The instructions also include: For each pre-stored program, when it is determined that the pre-stored program does not have a wildcard, the pre-stored program is output as a restored pre-stored program.
10. The data processing device according to claim 7, characterized in that: The instructions also include: For each of the development environment, the data quality verification environment, and the formal environment, parsing each restored stored program to determine a directed graph of all data table fields of each restored stored program; and The directed graphs of the data table fields of all restored stored procedures in each environment are merged to generate a lineage graph of a data table field of each environment.
11. The data processing device according to claim 10, characterized in that: The instructions also include: When it is determined that a source data table field of a first stored program after restoration and a target data table field of a second stored program after restoration are the same data table field, a first node of the source data table used to represent the first stored program after restoration and a second node of the target data table used to represent the second stored program after restoration are merged.
12. The data processing device according to claim 11, characterized in that: The instructions also include: Comparing a source directed subgraph or a target directed subgraph corresponding to at least one data table field of a blood relationship graph of data table fields of at least two environments among the development environment, the data quality verification environment, and the formal environment to generate the comparison result; and When the comparison result shows that there are differences in the blood relationship diagrams of the data table fields of any two environments among the development environment, the data quality verification environment and the formal environment, a notification signal is generated and sent to execute the notification function.