A method and device for adjusting a data quality evaluation model, equipment and medium

By determining data quality assessment indicators and metadata characteristics across multiple dimensions and adaptively adjusting the weight values ​​of the data quality assessment model, the problem of insufficient applicability of existing models is solved, and more accurate data assessment is achieved.

CN117093571BActive Publication Date: 2026-04-07INSPUR ZHUOSHU BIG DATA IND DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-21
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing data quality assessment models cannot adapt to the characteristics of different data, resulting in inaccurate assessment results and inability to effectively label training data.

Method used

By determining data quality assessment indicators across multiple dimensions, acquiring metadata table characteristics, analyzing data types and relationships, and adaptively adjusting the weight values ​​of the data quality assessment model, an assessment model that meets the requirements can be generated.

Benefits of technology

It improves the accuracy of the data evaluation model, enabling it to better adapt to different data tables and provide more accurate evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1
    Figure 1
Patent Text Reader

Abstract

This specification discloses an adaptive adjustment method, apparatus, device, and medium for a data quality assessment model, comprising: determining data quality assessment indicators corresponding to a data table to be assessed in multiple dimensions; obtaining metadata tables related to the data table to be assessed; analyzing the features of the data table to be assessed and the metadata tables to obtain feature analysis results; determining the weight values ​​of each data quality assessment indicator based on the feature analysis results; and adaptively adjusting a pre-generated data quality assessment model based on the weight values ​​to obtain a data quality assessment model that meets the requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a method, apparatus, device and medium for adjusting a data quality assessment model. Background Technology

[0002] As a crucial information carrier, data is a new factor of production and a vital productive force, serving as the cornerstone of socio-economic development and a fundamental and strategic resource for modern society. Its practical value lies primarily in two key aspects: firstly, data can help enterprises analyze markets and development trends to enhance their innovation capabilities and core competitiveness; secondly, it can assist government regulation and decision-making to improve the service quality and efficiency of government departments. The prerequisite for the scientific and rational use of data and realizing its immense application value is obtaining high-quality data, which plays a pivotal role in the development of the nation and society, and is a crucial guarantee for socio-economic development. However, data often suffers from quality issues in practice. These problems significantly impact the reliability of the information contained within the data, thereby affecting its actual value.

[0003] In existing technologies, most data quality assessment models are used to evaluate data quality. However, these models are not applicable to all data; their effectiveness depends on the data used to train them. Furthermore, these models do not identify the specific data from which they were trained. Therefore, applying a fixed data quality assessment model may not yield satisfactory results and may fail to produce accurate assessments. Summary of the Invention

[0004] This specification provides one or more embodiments of a method, apparatus, device, and medium for adjusting a data quality assessment model to solve the technical problems raised in the background art.

[0005] One or more embodiments of this specification employ the following technical solutions:

[0006] This specification provides an adaptive adjustment data quality assessment model method through one or more embodiments, including:

[0007] Determine the data quality assessment indicators corresponding to the data table to be evaluated from multiple dimensions; obtain the metadata table related to the data table to be evaluated;

[0008] The features of the data table to be evaluated and the metadata table are analyzed to obtain the feature analysis results;

[0009] Based on the results of the feature analysis, determine the weight values ​​of each data quality assessment indicator;

[0010] Based on the weight values, the pre-generated data quality assessment model is adaptively adjusted to obtain a data quality assessment model that meets the requirements.

[0011] This specification provides an adaptive adjustment data quality assessment model apparatus according to one or more embodiments, the apparatus comprising:

[0012] The evaluation index determination unit determines the data quality evaluation indexes corresponding to the data table to be evaluated from multiple dimensions; the metadata table acquisition unit acquires the metadata table related to the data table to be evaluated.

[0013] The feature analysis unit analyzes the features of the data table to be evaluated and the metadata table to obtain feature analysis results;

[0014] The weight value determination unit determines the weight value of each data quality assessment indicator based on the feature analysis results.

[0015] The adaptive adjustment unit adaptively adjusts the pre-generated data quality assessment model according to the weight values ​​to obtain a data quality assessment model that meets the requirements.

[0016] This specification provides an adaptive adjustment data quality assessment model device according to one or more embodiments, comprising:

[0017] At least one processor; and,

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions that can be executed by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:

[0020] Determine the data quality assessment indicators corresponding to the data table to be evaluated from multiple dimensions; obtain the metadata table related to the data table to be evaluated;

[0021] The features of the data table to be evaluated and the metadata table are analyzed to obtain the feature analysis results;

[0022] Based on the results of the feature analysis, determine the weight values ​​of each data quality assessment indicator;

[0023] Based on the weight values, the pre-generated data quality assessment model is adaptively adjusted to obtain a data quality assessment model that meets the requirements.

[0024] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, which, when executed by a computer, can perform the following:

[0025] Determine the data quality assessment indicators corresponding to the data table to be evaluated from multiple dimensions; obtain the metadata table related to the data table to be evaluated;

[0026] The features of the data table to be evaluated and the metadata table are analyzed to obtain the feature analysis results;

[0027] Based on the results of the feature analysis, determine the weight values ​​of each data quality assessment indicator;

[0028] Based on the weight values, the pre-generated data quality assessment model is adaptively adjusted to obtain a data quality assessment model that meets the requirements.

[0029] The above-described at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:

[0030] This specification describes an embodiment that analyzes the characteristics of the data table and metadata table to be evaluated, obtains feature analysis results, determines the weight values ​​of each data quality evaluation indicator in the data table to be evaluated based on the feature analysis results, and adaptively adjusts the data quality evaluation model according to the weight values ​​of each data quality evaluation indicator to make the data quality evaluation model more suitable for the data table to be evaluated. The adjusted data quality evaluation model can be used to evaluate the data table to be evaluated, and the evaluation results obtained are more accurate. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0032] Figure 1 A flowchart illustrating an adaptive adjustment data quality assessment model method provided for one or more embodiments of this specification;

[0033] Figure 2 A schematic diagram of an adaptive adjustment data quality assessment model device provided for one or more embodiments of this specification;

[0034] Figure 3 This is a schematic diagram of the structure of an adaptive adjustment data quality assessment model device provided in one or more embodiments of this specification. Detailed Implementation

[0035] This specification provides an embodiment of a method, apparatus, device, and medium for adjusting a data quality assessment model.

[0036] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0037] Figure 1 This diagram illustrates a flowchart of an adaptive adjustment method for a data quality assessment model, provided for one or more embodiments of this specification. This process can be executed by an adjustment data quality assessment model system. Certain input parameters or intermediate results in the process may be manually adjusted to help improve accuracy.

[0038] The method flow steps of the embodiments in this specification are as follows:

[0039] S102, determine the data quality assessment indicators corresponding to the data table to be evaluated in multiple dimensions.

[0040] In the embodiments of this specification, multiple dimensions may include completeness, accuracy, consistency, uniqueness, and timeliness. Determining the data quality assessment indicators corresponding to the data table to be evaluated across multiple dimensions may include the following:

[0041] When the dimension is completeness, the proportion of missing data in the data table to be evaluated can be determined. For example, if the data table to be evaluated has 1,000 records, of which 100 records are missing data in a certain field, then the proportion of missing data is 100 / 1000 = 0.1.

[0042] When the dimension is the accuracy, the proportion of the discrepancy data in the data table to be evaluated can be determined. For example, if the data table to be evaluated has 100 records, of which 10 records are different from the actual data records, then the proportion of the discrepancy data is 10 / 100 = 0.1.

[0043] When the dimension is the consistency, the data conflict rate between the data in the data table to be evaluated can be determined. For example, if there are 100 records in the data table to be evaluated, and 10 of them have contradictory data, then the data conflict rate is 10 / 100 = 0.1.

[0044] When the dimension is unique, the proportion of duplicate data in the data table to be evaluated can be determined. For example, if there are 1,000 records in the data table to be evaluated, and 50 of them contain duplicate data, then the proportion of duplicate data is 50 / 1,000 = 0.05.

[0045] When the dimension is timeliness, the timeliness and update frequency of the data in the data table to be evaluated can be determined. For example, if the data in the data table to be evaluated is required to be updated once a day, but is actually only updated once a week, then the timeliness of the data needs to be improved.

[0046] The aforementioned data quality assessment indicators can be used to evaluate the data quality of the data table to be evaluated within their respective dimensions. By quantifying these indicators, data quality can be accurately measured and compared, allowing for an understanding of the data's quality status and the implementation of corresponding measures to improve it.

[0047] S104, Obtain the metadata table related to the data table to be evaluated.

[0048] In the embodiments of this specification, if the data table to be evaluated is stored in a database, the relevant metadata information can be obtained by querying the database's system tables or metadata tables. Alternatively, data analysis tools can be used to obtain the metadata information of the data table to be evaluated. Data analysis tools can analyze the structure, relationships, and dependencies of the data table and generate a metadata table.

[0049] Specifically, the embodiments in this specification can determine the metadata collection method related to the data table to be evaluated based on the specific circumstances and complexity of the data source, so as to achieve the goal of efficiently and accurately obtaining metadata. If the data table to be evaluated is stored in a database, stored procedures and functions can be used to query the database's system tables or metadata tables to obtain relevant metadata information and generate a metadata table. Alternatively, the metadata table can be generated by obtaining DDL statements related to the data table to be evaluated, such as table creation, table attribute modification, and table structure modification, as well as DML statements related to data insertion, update, and deletion, from the database's binlog logs, and using data analysis tools combined with regular expressions and abstract syntax trees to parse these related statements, thereby obtaining the table structure, field data types, field relationships, and data lineage of the data table to be evaluated. Alternatively, after connecting to the database using the JDBC API based on data analysis tools, the ResultSetMetaData interface, which can obtain metadata information related to the connected database, can be used to obtain a result set of metadata information such as the data type, data length, and whether data is missing from the data table to be evaluated, thereby generating a metadata table.

[0050] Using the methods described above, we can obtain the metadata tables related to the data table to be evaluated. These metadata tables can include information such as field definitions, data types, table structure, and data sources. This metadata information is crucial for assessing data quality, data management, and data governance.

[0051] S106, Analyze the features of the data table to be evaluated and the metadata table to obtain feature analysis results.

[0052] In the embodiments of this specification, the characteristics of the data table to be evaluated and the metadata table include the data types in the data table to be evaluated and the relationships between attributes in the metadata table. The data types in the data table to be evaluated can be analyzed to determine the data type of each column in the data table to be evaluated; and the relationships between attributes in the metadata table can be analyzed to determine the relationships between the attributes in each column of the metadata table.

[0053] It should be noted that when analyzing the data type in the data table to be evaluated in the embodiments of this specification, the data type of each column can be determined by analyzing the data in each column of the data table to be evaluated. The data type can be text (string), number, date / time, boolean value, etc.

[0054] It should be noted that when analyzing the relationships between attributes in the metadata table in this embodiment of the specification, the relationships between column attributes can be determined by analyzing the relationships between different column attributes in the metadata table. For example, one column may be a foreign key, primary key, or index of another column, or there may be specific constraints between two columns. This allows for understanding the structure and relationships of the data table, aiding in data quality assessment and data management. For instance, by analyzing relationships, it can be determined whether a column must be consistent with other columns, or the relationship conditions between data tables can be identified, thereby helping to locate conflicting data.

[0055] It should be noted that, for analyzing the data types in the data table to be evaluated, the embodiments in this specification can use stored procedures or functions in the database to query and extract data from the metadata table to determine the data type of each column in the data table to be evaluated. Alternatively, data analysis tools or programming languages, such as Python or SQL, can be used to determine the data type of each column by parsing the structure of the data table and sampling data.

[0056] It should be noted that, for analyzing the relationships between attributes in the metadata table, the embodiments in this specification can use data modeling tools, graph databases, or metadata management systems to analyze the structure and relationships of the data table. These tools can visualize the data, obtain relationship diagrams between data tables, basic field information attributes, and data types based on data relationships, and obtain and analyze the relationships between different columns in the data table to be evaluated through corresponding operations and commands.

[0057] S108, Based on the feature analysis results, determine the weight values ​​of each data quality assessment indicator.

[0058] In the embodiments of this specification, the proportion of each data type can be determined based on the data type of each column in the data table to be evaluated, and a first weight value for each data quality evaluation indicator can be determined based on the proportion of each data type. The proportion of each relationship can be determined based on the relationships between the attributes of each column in the metadata table, and a second weight value for each data quality evaluation indicator can be determined based on the proportion of each relationship. The weight value for each data quality evaluation indicator is then determined based on the first weight value and the second weight value. The data types may include numeric, character, boolean, date, and text types; the relationships may include hierarchical structures, dependencies, and constraints.

[0059] It should be noted that, in the embodiments of this specification, the proportion of each data type can be determined by statistically analyzing the data types of each column in the data table to be evaluated and calculating the proportion of each data type. For example, if the data table to be evaluated has 10 columns of data, of which 3 are numeric, 2 are character, 2 are boolean, 2 are date, and 1 is text, then the proportion of numeric data is 3 / 10 = 0.3, the proportion of character data is 2 / 10 = 0.2, and so on. This is how the proportion of each data type can be determined.

[0060] It should be noted that when determining the first weight value of each data quality assessment indicator based on the occupancy ratio of each data type in the embodiments of this specification, the correlation between each data type and each data quality assessment indicator can be considered, and then the first weight value of each data quality assessment indicator can be determined based on the occupancy ratio of each data type.

[0061] Specifically, the embodiments in this specification can analyze the data quality assessment problems that need to be solved for each data type, and then analyze the data quality assessment indicators related to that data type. For example, for numerical data, data quality assessment problems related to accuracy, range validation, and outlier detection may be solved based on this data type, and will be related to indicators such as the proportion of discrepancies and the data conflict rate; for text data, data quality assessment problems related to completeness and format validation may be solved based on this data type, and will be related to indicators such as the proportion of missing data and the data conflict rate. Each data type can be pre-paired with relevant data quality assessment indicators according to actual needs and business background, forming a correspondence between data type and data quality assessment indicators. Then, based on the proportion of each data type, combined with fuzzy comprehensive evaluation method and quantification function, the first weight value of each data quality assessment indicator can be determined. For example, if a certain data type has a high proportion, the importance of the indicator of the proportion of differential data related to accuracy is also high. The importance of the indicator can be determined by combining the proportion of data types using the fuzzy comprehensive evaluation method. Then, a quantification function based on the large Cauchy distribution and logarithmic function can be used to determine the quantification value of the importance of the indicator and determine the weight value of the indicator of the proportion of differential data related to accuracy.

[0062] This analysis allows us to determine the primary weight of each data quality assessment indicator based on its usage proportion and the importance of the associated evaluation metrics. This more accurately reflects the importance of each data type in data quality assessment, providing guidance for subsequent assessment work.

[0063] The embodiments in this specification determine the proportion of each relationship by analyzing the relationships between the attributes in each column of the metadata table, including hierarchical structure, dependencies, and constraints. For example, if there are hierarchical relationships between 5 columns, dependencies between 3 columns, and constraints between 2 columns, then the proportion of the hierarchical relationships is 5 / (5+3+2) = 0.5, the proportion of dependencies is 3 / (5+3+2) = 0.3, and the proportion of constraints is 2 / (5+3+2) = 0.2. This allows the determination of the proportion of each relationship.

[0064] When determining the second weight value of each data quality assessment indicator based on the occupancy ratio of each relationship in the embodiments of this specification, the relationship between each relationship and each data quality assessment indicator can be considered, and then the second weight value of each data quality assessment indicator can be determined based on the occupancy ratio of each relationship.

[0065] For each relationship, analyze the data quality assessment problems that need to be solved based on that relationship, and then analyze the data quality assessment indicators related to that relationship. For example, for a hierarchical relationship, data quality assessment problems related to relationship integrity and hierarchical structure verification may be solved based on this relationship, and it will be related to indicators such as the proportion of missing data and the data conflict rate. For a dependency relationship, data quality assessment problems related to dependency verification and data consistency may be solved based on this relationship, and it will be related to indicators such as the data conflict rate. Each relationship can be paired with relevant data quality assessment indicators according to actual needs and business background, forming a correspondence between relationship and indicator. Then, based on the proportion of each relationship, combined with fuzzy comprehensive evaluation method and quantification function, determine the second weight value of each data quality assessment indicator. For example, if a certain relationship has a high proportion and the indicator of the proportion of missing data related to the integrity of the relationship is of high importance, the importance of the indicator can be determined by using the fuzzy comprehensive evaluation method in combination with the proportion of the relationship. Then, a quantification function based on the large Cauchy distribution and the logarithmic function can be used to determine the quantification value of the importance of the indicator and determine the weight value of the indicator of the proportion of missing data related to the integrity of the relationship.

[0066] This analysis allows us to determine the second weight value for each data quality assessment indicator based on the proportion of each relationship and the importance of the assessment indicators associated with those relationships. This more accurately reflects the importance of each relationship in the data quality assessment, providing guidance for subsequent assessment work.

[0067] Finally, based on the first and second weight values, the weights of each data quality assessment indicator can be calculated according to actual needs and business context. For example, assuming a data quality assessment indicator has a first weight value of 0.6 and a second weight value of 0.4, then the weight value of that indicator is 0.6 (first weight value) + 0.4 (second weight value). This calculation determines the weight values ​​of each data quality assessment indicator for subsequent data quality assessment work.

[0068] S110, Based on the weight values, adaptively adjust the pre-generated data quality assessment model to obtain a data quality assessment model that meets the requirements.

[0069] In the embodiments of this specification, the weight parameters of each data quality assessment indicator in the pre-generated data quality assessment model can be adaptively adjusted according to the weight values ​​to obtain a data quality assessment model that meets the requirements. Using the weight values ​​obtained in the above steps, the weight parameters of each indicator in the pre-generated data quality assessment model are adjusted according to certain algorithms and rules.

[0070] The following are several adjustment methods:

[0071] Linear adjustment method: Based on the proportional relationship of the weight values, the weight parameters are linearly adjusted to a suitable range. For example, if the weight value of a certain indicator is twice its original value, then the weight parameter of that indicator can be set to twice its original value.

[0072] The proportional adjustment method involves adjusting the weight parameters based on the proportional relationship between the weight values ​​and other indicators, so that the relative weights of each indicator match the proportion of its weight value. For example, if the weight value of a certain indicator is doubled while the weight values ​​of other indicators remain unchanged, then the weight parameter of that indicator can be set to double its original value, while the weight parameters of the other indicators remain unchanged.

[0073] Proportional Adjustment Method: Adjust the weight parameters of each indicator proportionally to the ratio of its weight value to the total weight value. For example, if the weight value of a certain indicator is half of the total weight value, then the weight parameter of that indicator can be set to half of the total weight value.

[0074] Iterative adjustment method: This method involves continuously adjusting the weight parameters iteratively until the desired result is achieved. Optimization algorithms, such as genetic algorithms and particle swarm optimization, can be used to search for the optimal solution in the search space of the weight parameters. During the search process, the weight parameters are continuously updated by evaluating the model's performance until the optimal solution is reached.

[0075] It should be noted that the specific adjustment method can be determined based on the form and requirements of the specific model.

[0076] Furthermore, the embodiments in this specification can evaluate the performance of the model using adjusted weight parameters to assess the data quality assessment model. This can be done by comparing it with actual data or by evaluating it according to actual needs and objectives.

[0077] Furthermore, the embodiments in this specification can be iteratively adjusted. Based on the evaluation results, the weight parameters are iteratively adjusted until a data quality evaluation model that meets the requirements is obtained.

[0078] Figure 2 This is a schematic diagram of an adaptive adjustment data quality assessment model device provided in one or more embodiments of this specification. The device includes: an assessment index determination unit 202, a metadata table acquisition unit 204, a feature analysis unit 206, a weight value determination unit 208, and an adaptive adjustment unit 210.

[0079] The evaluation index determination unit 202 determines the data quality evaluation indexes corresponding to the data table to be evaluated in multiple dimensions.

[0080] Metadata table acquisition unit 204 acquires the metadata table related to the data table to be evaluated;

[0081] The feature analysis unit 206 analyzes the features of the data table to be evaluated and the metadata table to obtain feature analysis results;

[0082] The weight value determination unit 208 determines the weight value of each data quality assessment indicator based on the feature analysis results.

[0083] The adaptive adjustment unit 210 adaptively adjusts the pre-generated data quality assessment model according to the weight values ​​to obtain a data quality assessment model that meets the requirements.

[0084] Figure 3 A schematic diagram of an adaptive adjustment data quality assessment model device provided for one or more embodiments of this specification includes:

[0085] At least one processor; and,

[0086] A memory communicatively connected to the at least one processor; wherein,

[0087] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:

[0088] Determine the data quality assessment indicators corresponding to the data table to be evaluated from multiple dimensions; obtain the metadata table related to the data table to be evaluated;

[0089] The features of the data table to be evaluated and the metadata table are analyzed to obtain the feature analysis results;

[0090] Based on the results of the feature analysis, determine the weight values ​​of each data quality assessment indicator;

[0091] Based on the weight values, the pre-generated data quality assessment model is adaptively adjusted to obtain a data quality assessment model that meets the requirements.

[0092] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, which, when executed by a computer, can perform the following:

[0093] Determine the data quality assessment indicators corresponding to the data table to be evaluated from multiple dimensions; obtain the metadata table related to the data table to be evaluated;

[0094] The features of the data table to be evaluated and the metadata table are analyzed to obtain the feature analysis results;

[0095] Based on the results of the feature analysis, determine the weight values ​​of each data quality assessment indicator;

[0096] Based on the weight values, the pre-generated data quality assessment model is adaptively adjusted to obtain a data quality assessment model that meets the requirements.

[0097] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0098] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0099] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.

Claims

1. An adaptive adjustment data quality assessment model method, characterized in that, The method includes: Determine the data quality assessment indicators for the data tables to be evaluated from multiple dimensions; Obtain the metadata table related to the data table to be evaluated; The features of the data table to be evaluated and the metadata table are analyzed to obtain the feature analysis results; Based on the results of the feature analysis, determine the weight values ​​of each data quality assessment indicator; Based on the weight values, the pre-generated data quality assessment model is adaptively adjusted to obtain a data quality assessment model that meets the requirements. The features of the data table to be evaluated and the metadata table include the data types in the data table to be evaluated and the relationships between attributes in the metadata table; The analysis of the features of the data table to be evaluated and the metadata table to obtain feature analysis results includes: Analyze the data types in the data table to be evaluated to determine the data type of each column in the data table to be evaluated; Analyze the relationships between attributes in the metadata table to determine the relationships between attributes in each column of the metadata table; The step of determining the weight values ​​of each data quality assessment indicator based on the feature analysis results includes: Based on the data type of each column in the data table to be evaluated, determine the proportion of each data type, and determine the first weight value of each data quality evaluation indicator based on the proportion of each data type. Based on the relationships between the attributes in each column of the metadata table, determine the proportion of each relationship, and determine the second weight value of each data quality assessment indicator based on the proportion of each relationship. Based on the first weight value and the second weight value, determine the weight value of each data quality assessment indicator; The step of adaptively adjusting the pre-generated data quality assessment model according to the weight values ​​to obtain a data quality assessment model that meets the requirements includes: Based on the weight values, the weight parameters of each data quality assessment indicator in the pre-generated data quality assessment model are adaptively adjusted to obtain a data quality assessment model that meets the requirements.

2. The method according to claim 1, characterized in that, The multiple dimensions include completeness, accuracy, consistency, uniqueness, and timeliness.

3. The method according to claim 2, characterized in that, The process of determining the data quality assessment indicators corresponding to the data table to be evaluated from multiple dimensions includes: When the dimension is completeness, determine the proportion of missing data in the data table to be evaluated; When the dimension is the accuracy, determine the proportion of the differential data in the data table to be evaluated; When the dimension is the aforementioned consistency, determine the data conflict rate in the data table to be evaluated; When the dimension is unique, determine the proportion of duplicate data in the data table to be evaluated; When the dimension is timeliness, the timeliness and update frequency of the data in the data table to be evaluated are determined.

4. The method according to claim 1, characterized in that, The data types include one or more of the following: numeric, character, boolean, date, and text. Relationships include one or more of the following: hierarchy, dependency, and constraints.

5. An adaptive adjustment data quality assessment model device, characterized in that, The device includes: The evaluation index determination unit determines the data quality evaluation indexes corresponding to the data table to be evaluated from multiple dimensions. Metadata table acquisition unit acquires metadata tables related to the data table to be evaluated; The feature analysis unit analyzes the features of the data table to be evaluated and the metadata table to obtain feature analysis results; The weight value determination unit determines the weight value of each data quality assessment indicator based on the feature analysis results. An adaptive adjustment unit adaptively adjusts the pre-generated data quality assessment model according to the weight values ​​to obtain a data quality assessment model that meets the requirements. The features of the data table to be evaluated and the metadata table include the data types in the data table to be evaluated and the relationships between attributes in the metadata table; The analysis of the features of the data table to be evaluated and the metadata table to obtain feature analysis results includes: Analyze the data types in the data table to be evaluated to determine the data type of each column in the data table to be evaluated; Analyze the relationships between attributes in the metadata table to determine the relationships between attributes in each column of the metadata table; The step of determining the weight values ​​of each data quality assessment indicator based on the feature analysis results includes: Based on the data type of each column in the data table to be evaluated, determine the proportion of each data type, and determine the first weight value of each data quality evaluation indicator based on the proportion of each data type. Based on the relationships between the attributes in each column of the metadata table, determine the proportion of each relationship, and determine the second weight value of each data quality assessment indicator based on the proportion of each relationship. Based on the first weight value and the second weight value, determine the weight value of each data quality assessment indicator; The step of adaptively adjusting the pre-generated data quality assessment model according to the weight values ​​to obtain a data quality assessment model that meets the requirements includes: Based on the weight values, the weight parameters of each data quality assessment indicator in the pre-generated data quality assessment model are adaptively adjusted to obtain a data quality assessment model that meets the requirements.

6. An adaptive adjustment data quality assessment model device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Determine the data quality assessment indicators corresponding to the data table to be evaluated from multiple dimensions; obtain the metadata table related to the data table to be evaluated; The features of the data table to be evaluated and the metadata table are analyzed to obtain the feature analysis results; Based on the results of the feature analysis, determine the weight values ​​of each data quality assessment indicator; Based on the weight values, the pre-generated data quality assessment model is adaptively adjusted to obtain a data quality assessment model that meets the requirements. The features of the data table to be evaluated and the metadata table include the data types in the data table to be evaluated and the relationships between attributes in the metadata table; The analysis of the features of the data table to be evaluated and the metadata table to obtain feature analysis results includes: Analyze the data types in the data table to be evaluated to determine the data type of each column in the data table to be evaluated; Analyze the relationships between attributes in the metadata table to determine the relationships between attributes in each column of the metadata table; The step of determining the weight values ​​of each data quality assessment indicator based on the feature analysis results includes: Based on the data type of each column in the data table to be evaluated, determine the proportion of each data type, and determine the first weight value of each data quality evaluation indicator based on the proportion of each data type. Based on the relationships between the attributes in each column of the metadata table, determine the proportion of each relationship, and determine the second weight value of each data quality assessment indicator based on the proportion of each relationship. Based on the first weight value and the second weight value, determine the weight value of each data quality assessment indicator; The step of adaptively adjusting the pre-generated data quality assessment model according to the weight values ​​to obtain a data quality assessment model that meets the requirements includes: Based on the weight values, the weight parameters of each data quality assessment indicator in the pre-generated data quality assessment model are adaptively adjusted to obtain a data quality assessment model that meets the requirements.

7. A non-volatile computer storage medium, characterized in that, It stores computer-executable instructions, which, when executed by a computer, can achieve the following: Determine the data quality assessment indicators corresponding to the data table to be evaluated from multiple dimensions; obtain the metadata table related to the data table to be evaluated; The features of the data table to be evaluated and the metadata table are analyzed to obtain the feature analysis results; Based on the results of the feature analysis, determine the weight values ​​of each data quality assessment indicator; Based on the weight values, the pre-generated data quality assessment model is adaptively adjusted to obtain a data quality assessment model that meets the requirements. The features of the data table to be evaluated and the metadata table include the data types in the data table to be evaluated and the relationships between attributes in the metadata table; The analysis of the features of the data table to be evaluated and the metadata table to obtain feature analysis results includes: Analyze the data types in the data table to be evaluated to determine the data type of each column in the data table to be evaluated; Analyze the relationships between attributes in the metadata table to determine the relationships between attributes in each column of the metadata table; The step of determining the weight values ​​of each data quality assessment indicator based on the feature analysis results includes: Based on the data type of each column in the data table to be evaluated, determine the proportion of each data type, and determine the first weight value of each data quality evaluation indicator based on the proportion of each data type. Based on the relationships between the attributes in each column of the metadata table, determine the proportion of each relationship, and determine the second weight value of each data quality assessment indicator based on the proportion of each relationship. Based on the first weight value and the second weight value, determine the weight value of each data quality assessment indicator; The step of adaptively adjusting the pre-generated data quality assessment model according to the weight values ​​to obtain a data quality assessment model that meets the requirements includes: Based on the weight values, the weight parameters of each data quality assessment indicator in the pre-generated data quality assessment model are adaptively adjusted to obtain a data quality assessment model that meets the requirements.

Citation Information

Patent Citations

  • Service data set quality evaluation method and device and computer readable medium

    CN111369136A

  • Multi-source data quality control method for power distribution network

    CN113869633A

  • Data quality evaluation method and system based on big data analysis

    CN116467292A