A data quality inspection rule matching method, storage medium and system

By collecting field metadata and data quality inspection rules in the business system, calculating the correlation degree and replacing rule information, the problem that field metadata cannot match data quality inspection rules is solved, and a comprehensive quality inspection of business data is achieved.

CN115328902BActive Publication Date: 2025-05-16INFORMATION CENT OF YUNNAN POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211049853.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-05-16
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

In the prior art, the correlation between field metadata and data quality inspection rules fails to meet the standards, resulting in the field metadata not being able to match the data quality inspection rules, and thus the quality inspection of the corresponding business data cannot be carried out.

Method used

By collecting field metadata in the business system and preset data quality inspection rules, the correlation between each field metadata and data quality inspection rules is calculated. If the standards are not met, the candidate field metadata with the same text similarity and data type are identified, the field information and conditional parameters of the data quality inspection rules are replaced, and a new data quality inspection rule is generated to match the field metadata to be matched.

Benefits of technology

It realizes effective matching of the metadata of the field to be matched that the data quality inspection rules are not matched, ensuring that the business data described can be quality-checked, and improving the coverage of data quality inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115328902B_ABST
    Figure CN115328902B_ABST
Patent Text Reader

Abstract

The present invention provides a data quality inspection rule matching method, storage medium and system. The method comprises: collecting multiple field metadata and multiple data quality inspection rules, calculating the correlation between each field metadata and each data quality inspection rule, matching the field metadata with the correlation reaching the standard with the data quality inspection rule, identifying the candidate field metadata that has matched the data quality inspection rule and the to-be-matched field metadata that has not matched the data quality inspection rule, if there is candidate field metadata whose text similarity with the to-be-matched field metadata is greater than a preset threshold and whose data type is consistent, then replacing the parameter information contained in the data quality inspection rule selected by the user with the data information of the to-be-matched field metadata, replacing the condition parameter with the new condition parameter input by the user, obtaining the new data quality inspection rule and matching the to-be-matched field metadata with the new data quality inspection rule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data quality inspection rule matching method, storage medium and system. Background Art

[0002] The power grid system generates a large amount of business data when it is running. These business data can reflect the operating status of the power grid system and need to be collected and stored in the business system. At present, data quality inspection rules are usually used to check the quality of business data in the business system. If the quality inspection result of business data is abnormal, the staff needs to monitor the power grid operation business corresponding to the abnormal business data.

[0003] Currently, multiple data quality check rules used for quality checks are usually pre-set. Therefore, in the process of selecting data quality check rules for quality checks on business data, the correlation between the field metadata used to describe the business data and the data quality check rules must be calculated first. If the correlation meets the standard, the field metadata is matched with the data quality check rules for quality checks. If there is a field metadata whose correlation with each data quality check rule does not meet the standard, the field metadata cannot match the data quality check rule, so the quality check of the business data described by the field metadata cannot be performed. Summary of the invention

[0004] The technical problem to be solved by the present invention is how to improve the situation where field metadata cannot match data quality inspection rules.

[0005] In order to solve the above technical problems, the present invention provides a data quality inspection rule matching method, comprising the following steps:

[0006] A. Collect multiple field metadata used to describe business data and the name information, source information and data type information of each field metadata from the business system;

[0007] B. Obtain multiple preset data quality check rules and the field name information, field source information and condition parameters contained in each data quality check rule;

[0008] C. Based on the name information and source information of each field metadata and the field name information and field source information contained in each data quality check rule, determine whether the correlation between each field metadata and each data quality check rule meets the standard;

[0009] D. Make the metadata of the fields that meet the relevance standards match the data quality check rules;

[0010] E. Identifying candidate field metadata that has matched the data quality check rule and to-be-matched field metadata that has not matched the data quality check rule among the multiple field metadata;

[0011] F. For each field metadata to be matched, perform the following steps F1, F2, F3, and F4:

[0012] ——F1. Determine whether there is candidate field metadata whose text similarity with the to-be-matched field metadata is greater than a preset threshold and whose data type is consistent. If so, display the candidate field metadata and the data quality check rules it matches for the user to select;

[0013] ——F2. Obtain the candidate field metadata selected by the user and the data quality check rule it matches, and replace the field name information and field source information contained in the data quality check rule with the name information and source information of the field metadata to be matched;

[0014] ——F3. Obtain the new condition parameters input by the user, replace the condition parameters contained in the data quality check rule selected by the user with the new condition parameters input by the user, and obtain a new data quality check rule;

[0015] ——F4. Match the metadata of the field to be matched with the new data quality check rule.

[0016] Preferably, in step D, if there is a data quality check rule, the text similarity between its field name information and the name information of the metadata of this field reaches a first preset value, and the text similarity between its field source information and the source information of the metadata of this field reaches a second preset value, then the correlation between the data quality check rule and the metadata of this field meets the standard.

[0017] Preferably, in step F1, the text similarity between the metadata of the field to be matched and the metadata of each candidate field is calculated based on the name information of the metadata of the field to be matched and the name information of the metadata of each candidate field, and it is determined whether there is candidate field metadata whose text similarity with the metadata of the field to be matched is greater than a preset threshold. If so, it is determined whether the data type of the metadata of the field to be matched is consistent with that of the candidate field metadata based on the data type information of the metadata of the field to be matched and the data type information of the metadata of the candidate field.

[0018] Preferably, in step F1, if there are multiple candidate field metadata whose text similarity with the to-be-matched field metadata is greater than a preset threshold and whose data types are consistent, these multiple candidate field metadata are sorted from large to small according to text similarity for user selection.

[0019] Preferably, in step F1, candidate field metadata with text similarity ranking higher than a predetermined ranking is selected for display.

[0020] Preferably, in step F2, the data quality check rule is first decomposed into a select clause including field name information, a from clause including field source information and a where clause including conditional parameters using a SQL engine, and then the field name information in the select clause is replaced with the name information of the field metadata to be matched, and the field source information in the from clause is replaced with the source information of the field metadata to be matched; in step F3, the conditional parameters in the where clause are replaced with new conditional parameters input by the user, and then the replaced select clause, from clause and where clause are combined to obtain a new data quality check rule.

[0021] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the data quality check rule matching method as described above are implemented.

[0022] The present invention also provides a data quality check rule matching system, comprising a computer-readable storage medium and a processor connected to each other, wherein the computer-readable storage medium is as described above.

[0023] The present invention has the following beneficial effects: for the metadata of the to-be-matched field that does not match the data quality check rule, it is determined whether there is candidate field metadata whose text similarity with the metadata of the to-be-matched field is greater than a preset threshold and whose data type is consistent; if so, it means that the template of the data quality check rule matched by the candidate field metadata is suitable for the metadata of the to-be-matched field, so the candidate field metadata and the data quality check rule matched by it are displayed for the user to select, and then the candidate field metadata selected by the user and the data quality check rule matched by it are obtained, the field name information and the field source information contained in the data quality check rule are replaced with the name information and the source information of the to-be-matched field metadata, and the new condition parameter input by the user is obtained, and the data quality check rule selected by the user is replaced with the name information and the source information of the to-be-matched field metadata. The condition parameters contained in the check rule are replaced with the new condition parameters input by the user, that is, based on the template of the original data quality check rule, the field name information, field source information and condition parameters of the data quality check rule are changed according to the name information, source information and new condition parameters of the metadata of the field to be matched, so as to obtain a new data quality check rule. Since the field name information, field source information and condition parameters of the new data quality check rule are changed from the data information of the metadata of the field to be matched, the correlation between the metadata of the field to be matched and the new data quality check rule will meet the standard, so the metadata of the field to be matched is matched with the new data quality check rule, and the new data quality check rule can be used to perform quality inspection on the business data described by the metadata of the field to be matched. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flowchart of the data quality check rule matching method. DETAILED DESCRIPTION

[0025] The present invention is further described in detail below in conjunction with specific implementation methods.

[0026] This embodiment provides a data quality check rule matching system, which includes a computer-readable storage medium and a processor connected to each other. The computer-readable storage medium stores a computer program. When the computer program is executed by the processor, the following is implemented: Figure 1 The data quality inspection rule matching method shown includes the following steps A, B, C, D, E, and F.

[0027] A. Collect multiple field metadata used to describe business data and the name information, source information and data type information of each field metadata from the business system.

[0028] In this embodiment, the business system stores a large amount of business data generated by the power grid system during operation, and these business data can reflect the operating status of the power grid system. The data quality inspection rule matching system collects multiple field metadata used to describe business data from the business system, as well as data information such as name information, source information, and data type information of each field metadata. For example, the name information of field metadata one is name, the source information is t_userinfo, and the data type information is text; the name information of field metadata two is name, the source information is t_admininfo, and the data type information is text; the name information of field metadata three is age, the source information is t_userinfo, and the data type information is numeric;.

[0029] B. Obtain multiple preset data quality check rules and the field name information, field source information and condition parameters contained in each data quality check rule.

[0030] In order to perform quality checks on business data in a business system, multiple data quality check rules are usually preset, and each data quality check rule contains parameter information such as field name information, field source information, and condition parameters. For example, there is a preset data quality check rule 1 "Select name from t_userinfo where len(name)>8", whose field name information is "name", field source information is "t_userinfo", and condition parameter is "8"; there is a preset data quality check rule 2 "Select date from t_userinfo where date is null", whose field name information is "date", field source information is "t_userinfo", and condition parameter is "null". The data quality check rule matching system obtains multiple preset data quality check rules, and obtains the field name information, field source information, and condition parameters contained in each data quality check rule.

[0031] C. Based on the name information and source information of each field metadata and the field name information and field source information contained in each data quality check rule, determine whether the correlation between each field metadata and each data quality check rule meets the standard;

[0032] The system uses the Levenshtein Distance algorithm to calculate the text similarity between the name information of each field metadata and the field name information of each data quality check rule, and determines whether the text similarity reaches a first preset value (80%), and calculates the text similarity between the source information of each field metadata and the field source information of each data quality check rule, and determines whether the text similarity reaches a second preset value (100%). When the text similarity between the name information of the field metadata and the field name information of the data quality check rule reaches the first preset value, and the text similarity between the source information of the field metadata and the field source information of the data quality check rule reaches the second preset value, it is judged that the correlation between the field metadata and the data quality check rule meets the standard, otherwise it is judged as not meeting the standard. For example, for field metadata 1, field metadata 2, field metadata 3, data quality check rule 1, and data quality check rule 2, the system needs to calculate the correlation between field metadata 1 and data quality check rule 1, the correlation between field metadata 1 and data quality check rule 2, the correlation between field metadata 2 and data quality check rule 1, the correlation between field metadata 2 and data quality check rule 2, the correlation between field metadata 3 and data quality check rule 1, and the correlation between field metadata 3 and data quality check rule 2, and then determine whether these correlations meet the standards, as follows:

[0033] The system calculates the text similarity between the name information "name" of field metadata one and the field name information "name" of data quality check rule one, and the calculation result is a text similarity of 100%, which reaches the first preset value (80%). Then, the system calculates the text similarity between the source information "t_userinfo" of field metadata one and the field source information "t_userinfo" of data quality check rule one, and the calculation result is a text similarity of 100%, which reaches the second preset value (100%). In this case, the system determines that the correlation between field metadata one and data quality check rule one meets the standard.

[0034] The system calculates the text similarity between the name information "name" of field metadata one and the field name information "date" of data quality check rule two. The calculation result is a text similarity of 50%, which does not reach the first preset value (80%). Then, the system calculates the text similarity between the source information "t_userinfo" of field metadata one and the field source information "t_userinfo" of data quality check rule two. The calculation result is a text similarity of 100%, which reaches the second preset value (100%). In this case, the system determines that the correlation between field metadata one and data quality check rule two does not meet the standard.

[0035] The system calculates the text similarity between the name information "name" of field metadata two and the field name information "name" of data quality check rule one, and the calculation result is a text similarity of 100%, which does not reach the first preset value (80%). Then, the system calculates the text similarity between the source information "t_admininfo" of field metadata two and the field source information "t_userinfo" of data quality check rule one, and the calculation result is a text similarity of 50%, which does not reach the second preset value (100%). In this case, the system determines that the correlation between field metadata two and data quality check rule one does not meet the standard.

[0036] The system calculates the text similarity between the name information "name" of field metadata two and the field name information "date" of data quality check rule two, and the calculation result is a text similarity of 50%, which does not reach the first preset value (80%). Then, the system calculates the text similarity between the source information "t_admininfo" of field metadata two and the field source information "t_userinfo" of data quality check rule two, and the calculation result is a text similarity of 50%, which does not reach the second preset value (100%). In this case, the system determines that the correlation between field metadata two and data quality check rule two does not meet the standard.

[0037] The system calculates the text similarity between the name information "age" of field metadata three and the field name information "name" of data quality check rule one. The calculation result is a text similarity of 30%, which does not reach the first preset value (80%). Then, the system calculates the text similarity between the source information "t_userinfo" of field metadata three and the field source information "t_userinfo" of data quality check rule one. The calculation result is a text similarity of 100%, which reaches the second preset value (100%). In this case, the system determines that the correlation between field metadata three and data quality check rule two does not meet the standard.

[0038] The system calculates the text similarity between the name information "age" of field metadata three and the field name information "date" of data quality check rule two. The calculation result is a text similarity of 30%, which does not reach the first preset value (80%). Then, the system calculates the text similarity between the source information "t_userinfo" of field metadata three and the field source information "t_userinfo" of data quality check rule two. The calculation result is a text similarity of 100%, which reaches the second preset value (100%). In this case, the system determines that the correlation between field metadata three and data quality check rule two does not meet the standard.

[0039] It should be noted that the Levenshtein Distance algorithm is also called the Edit Distance algorithm, which is an edit distance algorithm. It calculates the minimum number of edit operations required to convert two strings from one to another to obtain the edit distance between the two strings. The smaller the edit distance, the greater the text similarity of the two strings. The edit operations include replacing one character with another, inserting a character, and deleting a character.

[0040] D. Make the field metadata that meets the relevance criteria match the data quality check rules.

[0041] After determining whether the correlation between each field metadata and each data quality check rule meets the standard, the system establishes a mapping relationship between the field metadata with the correlation meeting the standard and the data quality check rule, so that the field metadata with the correlation meeting the standard matches the data quality check rule, and does not establish a mapping relationship between the field metadata with the correlation not meeting the standard and the data quality check rule, so that the field metadata with the correlation not meeting the standard does not match the data quality check rule. Specifically: the system determines that the correlation between field metadata one and data quality check rule one meets the standard, the correlation between field metadata one and data quality check rule two does not meet the standard, and the correlation between field metadata two and data quality check rule two does not meet the standard. The correlation between field metadata 2 and data quality check rule 2 does not meet the standard, the correlation between field metadata 3 and data quality check rule 1 does not meet the standard, and the correlation between field metadata 3 and data quality check rule 2 does not meet the standard. Therefore, the system matches field metadata 1 with data quality check rule 1, does not match field metadata 1 with data quality check rule 2, does not match field metadata 2 with data quality check rule 1, does not match field metadata 2 with data quality check rule 2, does not match field metadata 3 with data quality check rule 1, and does not match field metadata 3 with data quality check rule 2. That is, field metadata 1 matches data quality check rule 1 "Select name from t_userinfo where len(name)>8", while field metadata 2 and 3 do not match data quality check rules.

[0042] E. Identify candidate field metadata that has matched the data quality check rule and to-be-matched field metadata that has not matched the data quality check rule among the multiple field metadata.

[0043] In this embodiment, the system records the field metadata one that has matched the data quality check rule one as the candidate field metadata, and records the field metadata two and three that have not matched the data quality check rule as the field metadata to be matched. The system identifies and distinguishes the candidate field metadata one and the field metadata two and three to be matched.

[0044] F. Execute the following steps F1, F2, F3, and F4 for each field metadata to be matched.

[0045] (1) Steps F1, F2, F3, and F4 are performed on the second matching field metadata, as detailed below:

[0046] F1. Determine whether there is candidate field metadata whose text similarity with the to-be-matched field metadata is greater than a preset threshold and whose data type is consistent. If so, display the candidate field metadata and the data quality check rules it matches for the user to select;

[0047] The system first calculates the text similarity between the metadata of the field to be matched and the metadata of each candidate field based on the name information of the metadata of the field to be matched and the name information of the metadata of each candidate field, and determines whether there is candidate field metadata whose text similarity with the metadata of the field to be matched is greater than a preset threshold (for example, 80%). If so, the system then compares the data type information of the metadata of the field to be matched and the data type information of the candidate field metadata to determine whether the data type of the metadata of the field to be matched is consistent with that of the candidate field metadata. If there is candidate field metadata whose text similarity with the metadata of the field to be matched is greater than the preset threshold and whose data type is consistent, the candidate field metadata and its matching data quality inspection rules are displayed for user selection.

[0048] In this embodiment, the name information of the second field metadata to be matched is "name", and the data type information is "text class", and there is one candidate field metadata, specifically candidate field metadata 1, whose field name information is "name" and data type information is "text class". Therefore, according to the name information "name" of the second field metadata to be matched and the name information "name" of the candidate field metadata 1, the text similarity between the second field metadata to be matched and the candidate field metadata 1 is calculated to be 100%, which is greater than the preset threshold (80%), that is, there is a candidate field metadata 1 whose text similarity with the second field metadata to be matched is greater than the preset threshold. Therefore, according to the data type information "text class" of the second field metadata to be matched and the data type information "text class" of the first candidate field metadata, it is compared and judged that the data types of the second field metadata to be matched and the candidate field metadata 1 are consistent, that is, the text similarity between the second field metadata to be matched and the candidate field metadata 1 is greater than the preset threshold and the data types are consistent. In this case, the system displays the candidate field metadata 1 and its matched data quality inspection rule 1 "Select name from t_userinfo where len(name)>8" for user selection.

[0049] In other embodiments, if there are multiple candidate field metadata, and there are also multiple candidate field metadata whose text similarity with the second field metadata to be matched is greater than a preset threshold and whose data type is consistent, then these multiple candidate field metadata are sorted from large to small according to text similarity, and the candidate field metadata with the top three text similarities are selected for display for user selection.

[0050] F2. Obtain the candidate field metadata selected by the user and the data quality check rule it matches, and replace the field name information and field source information contained in the data quality check rule with the name information and source information of the field metadata to be matched.

[0051] In this embodiment, the system displays the candidate field metadata 1 and the data quality check rule 1 "Select name from t_userinfo where len(name)>8" that it matches. If the user thinks that the data quality check rule 1 is suitable after viewing, the user can select the candidate field metadata 1 and the data quality check rule 1 that it matches. The system then obtains the candidate field metadata 1 and the data quality check rule 1 that the user selects, and then uses the Sql block module in the SQL engine to decompose the data quality rule 1 "Select name from t_userinfo where len(name)>8" into the select clause "Select name", the from clause "from t_userinfo" and the where clause "where len(name)>8"; then, the parameter filling module in the SQL engine is used to replace the field name information "name" in the select clause with the name information "anme" of the field metadata 2 to be matched, and replace the field source information "t_userinfo" in the from clause with the source information "t_admininfo" of the field metadata 2 to be matched.

[0052] F3. Obtain the new condition parameters input by the user, replace the condition parameters included in the data quality check rule selected by the user with the new condition parameters input by the user, and obtain a new data quality check rule.

[0053] After selecting candidate field metadata 1 and the data quality check rule 1 that matches it, the user needs to input new condition parameters of the data quality check rule into the system according to experience. For example, if you input "10", the system will use the parameter filling module in the SQL engine to replace the condition parameter "8" in the where clause with the new condition parameter "10" input by the user after obtaining the new condition parameter "10" input by the user. Then, the SQL combination module in the SQL engine will be used to combine the replaced select clause, from clause and where clause to obtain the new data quality check rule "Select name from t_admininfo where len(name)>10".

[0054] It can be seen from steps F2 and F3 that the system changes the field name information "name", field source information "t_userinfo" and condition parameter "8" of data quality check rule one based on the template of the original data quality check rule one "Select name from t_userinfo where len(name)>8" according to the name information "name" of the metadata field two to be matched, the source information "t_admininfo" and the new condition parameter "10" entered by the user, so that the new data quality check rule "Select name from t_admininfo where len(name)>10" can be obtained.

[0055] It should be noted that SQL is the abbreviation of Structured Query Language, which is a computer language used to access, query, update and manage data in relational databases. The SQL engine is one of the important subsystems of the database. It is responsible for accepting SQL statements sent by applications and directing the executor to run the execution plan.

[0056] F4. Match the metadata of the field to be matched with the new data quality check rules.

[0057] After obtaining the new data quality check rule "Select name from t_admininfo where len(name)>10", the correlation between the second metadata of the field to be matched and the new data quality check rule "Select name from t_admininfowhere len(name)>10" will meet the standard, so the system establishes a mapping relationship between the second metadata of the field to be matched and the new data quality check rule "Select name from t_admininfo where len(name)>10", so that the second metadata of the field to be matched is matched with the new data quality check rule "Select name from t_admininfo where len(name)>10". In this way, the new data quality check rule "Select name from t_admininfo where len(name)>10" can be used to perform quality inspection on the business data described by the second metadata of the field to be matched.

[0058] (2) Steps F1, F2, F3, and F4 are executed for the matching field metadata 3, as detailed below:

[0059] First, execute step F1: In this embodiment, the name information of the third field metadata to be matched is "age", and the data type information is "numeric type", and there is one candidate field metadata, specifically candidate field metadata 1, whose field name information is "name", and the data type information is "text type". Therefore, according to the name information "age" of the third field metadata to be matched and the name information "name" of the candidate field metadata 1, the text similarity between the third field metadata to be matched and the candidate field metadata 1 is calculated to be 30%, which is not greater than the preset threshold value (80%), that is, there is no candidate field metadata whose text similarity with the third field metadata to be matched is greater than the preset threshold value, so there is no need to compare and determine whether the data types of the third field metadata to be matched and the candidate field metadata 1 are consistent. It can be known that the text similarity between the third field metadata to be matched and the candidate field metadata 1 is not greater than the preset threshold value and the data type is consistent. In this case, the system will not display the candidate field metadata 1 and its matched data quality inspection rule 1 "Select name from t_userinfo where len(name)>8" for the user to select.

[0060] Then, the system does not need to execute steps F2, F3, and F4.

[0061] The above is only an implementation method of the invention, and does not limit the scope of patent protection. Those skilled in the art can make non-substantial changes or substitutions based on the invention, which still fall within the scope of patent protection.

Claims

1. A data quality inspection rule matching method, characterized in that: The steps include: A. Collect multiple field metadata used to describe business data and the name information, source information and data type information of each field metadata from the business system; B. Obtain multiple preset data quality check rules and the field name information, field source information and condition parameters contained in each data quality check rule; C. Based on the name information and source information of each field metadata and the field name information and field source information contained in each data quality check rule, determine whether the correlation between each field metadata and each data quality check rule meets the standard; D. Make the metadata of the fields that meet the relevance standards match the data quality check rules; E. Identifying candidate field metadata that has matched the data quality check rule and to-be-matched field metadata that has not matched the data quality check rule among the multiple field metadata; F. For each field metadata to be matched, perform the following steps F1, F2, F3, and F4: ——F1. Determine whether there is candidate field metadata whose text similarity with the to-be-matched field metadata is greater than a preset threshold and whose data type is consistent. If so, display the candidate field metadata and the data quality check rules it matches for the user to select; ——F2. Obtain the candidate field metadata selected by the user and the data quality check rule it matches, and replace the field name information and field source information contained in the data quality check rule with the name information and source information of the field metadata to be matched; ——F3. Obtain the new condition parameters input by the user, replace the condition parameters contained in the data quality check rule selected by the user with the new condition parameters input by the user, and obtain a new data quality check rule; ——F4. Match the metadata of the field to be matched with the new data quality check rule.

2. The data quality inspection rule matching method according to claim 1 is characterized in that: In step C, if there is a data quality check rule, the text similarity between its field name information and the name information of the metadata of this field reaches a first preset value, and the text similarity between its field source information and the source information of the metadata of this field reaches a second preset value, then the correlation between the data quality check rule and the metadata of this field meets the standard.

3. The data quality inspection rule matching method according to claim 1 is characterized in that: In the step F1, the text similarity between the metadata of the field to be matched and the metadata of each candidate field is first calculated based on the name information of the metadata of the field to be matched and the name information of the metadata of each candidate field, and it is determined whether there is candidate field metadata whose text similarity with the metadata of the field to be matched is greater than a preset threshold. If so, it is determined whether the data type of the metadata of the field to be matched is consistent with that of the candidate field metadata based on the data type information of the metadata of the field to be matched and the data type information of the metadata of the candidate field.

4. The data quality inspection rule matching method according to claim 1 is characterized in that: In step F1, if there are multiple candidate field metadata whose text similarity with the to-be-matched field metadata is greater than a preset threshold and whose data types are consistent, the multiple candidate field metadata are sorted from large to small according to text similarity for user selection.

5. The data quality inspection rule matching method according to claim 4 is characterized in that: In step F1, candidate field metadata with text similarity ranking higher than a predetermined ranking is selected for display.

6. The data quality inspection rule matching method according to claim 1 is characterized in that: In the step F2, the data quality inspection rule is first decomposed into a select clause including field name information, a from clause including field source information, and a where clause including conditional parameters using a SQL engine, and then the field name information in the select clause is replaced with the name information of the field metadata to be matched, and the field source information in the from clause is replaced with the source information of the field metadata to be matched; in the step F3, the conditional parameters in the where clause are replaced with new conditional parameters input by the user, and then the replaced select clause, from clause, and where clause are combined to obtain a new data quality inspection rule.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the data quality check rule matching method according to any one of claims 1 to 6 are implemented.

8. A data quality check rule matching system, comprising a computer-readable storage medium and a processor connected to each other, characterized in that: The computer readable storage medium as claimed in claim 7.

Citation Information

Patent Citations

  • Data matching method and device, computer readable medium and electronic equipment

    CN111667923A

  • Method and device for generating SQL (Structured Query Language) statement, computer equipment and storage medium

    CN113901075A