Agent-based multi-dimensional data quality intelligent evaluation method

By building a multi-dimensional data quality evaluation index system and an Agent intelligent model, the problems of incomplete and inefficient data quality evaluation in the existing technology are solved, and comprehensive, flexible and intelligent evaluation of data quality is achieved, and evaluation efficiency and accuracy are improved.

CN120496720APending Publication Date: 2025-08-15THE FIRST RES INST OF MIN OF PUBLIC SECURITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510661612.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing data quality evaluation methods are single, unable to fully reflect the data status, and are inefficient, making it difficult to adapt to changes in different types of data and evaluation needs, and cannot meet the rapid evaluation needs in the big data era.

Method used

Using the multi-dimensional data quality intelligent evaluation method based on Agent, a multi-dimensional data quality evaluation index system is built, including indicators of the data itself, application and security dimensions, and an automated evaluation is used to use the Agent intelligent model to achieve comprehensive, flexible and intelligent evaluation of data quality through data perception, evaluation, decision-making and interactive Agent.

Benefits of technology

It realizes the comprehensiveness and accuracy of data quality assessment, dynamically adapts to data changes, reduces manual intervention, and improves assessment efficiency, especially in medical data assessment, shortens from several days to several hours, and promptly discovers potential problems and provides solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496720A_ABST
    Figure CN120496720A_ABST
Patent Text Reader

Abstract

The invention discloses an Agent-based multi-dimensional data quality intelligent evaluation method. The method comprises the following steps: establishing a multi-dimensional data quality evaluation index system; and an Agent agent model is constructed, the Agent agent model comprises a data perception Agent, an evaluation Agent, a decision-making Agent and an interaction Agent, and automatic evaluation of data quality is realized. According to the invention, comprehensive improvement of data quality evaluation, dynamic adaptability enhancement, intelligence and high efficiency can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data quality assessment, and in particular to an agent-based multidimensional data quality intelligent assessment method. Background Art

[0002] Data is a crucial basis for management decisions across various administrative departments and businesses, such as personnel records in administrative departments, medical records in medical institutions, and credit data in the financial industry. Data quality can significantly impact the effectiveness of management and decision-making.

[0003] Currently, there are two main types of research on data quality assessment. One focuses on the dimensions of data quality assessment, which are mostly based on the accuracy, completeness, consistency, and timeliness of the data itself. However, a single assessment dimension cannot fully reflect the data quality status, resulting in inaccurate assessment of data value and potential for companies to make erroneous decisions based on a one-sided data quality assessment. The other focuses on data quality assessment methods. Existing assessment methods are mostly applicable only to specific scenarios and are difficult to adapt to different types of data and changing assessment requirements. They also typically require a large amount of manual participation, are inefficient, and struggle to process massive amounts of data in real time, failing to meet the demand for rapid data quality assessment in the big data era. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, the present invention aims to provide an agent-based multidimensional data quality intelligent assessment method to achieve comprehensive, flexible, intelligent and efficient assessment of data quality, and provide reliable support for data-driven decision-making.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] An agent-based multidimensional data quality intelligent evaluation method includes the following steps:

[0007] S1. Establish a multidimensional data quality assessment indicator system:

[0008] Build a data quality assessment indicator system from one or more dimensions of data itself, data application, and data security;

[0009] (1) Data quality assessment indicators based on the data itself include:

[0010] Accuracy: used to assess whether the data is true, reliable, and accurately reflects the real world;

[0011] Completeness: used to assess whether the data is complete and whether there are missing or null values;

[0012] Consistency: used to evaluate whether data is consistent across different systems, at different time points, and / or across different data tables;

[0013] Timeliness: used to assess whether the data is updated in a timely manner and whether it can reflect the latest situation;

[0014] Uniqueness: used to evaluate whether the data contains duplicate records or redundant information;

[0015] (2) Data quality assessment indicators based on data application include:

[0016] Usability: used to assess whether the data is easy to access, understand and use;

[0017] Relevance: used to assess whether the data is relevant to business needs and whether it can meet analytical requirements;

[0018] Credibility: used to assess whether the data source is reliable and whether the data processing process is transparent;

[0019] Interpretability: used to evaluate whether the data is easy to understand and interpret and whether it can support decision-making;

[0020] (3) Data quality assessment indicators from the perspective of data security include:

[0021] Security: used to assess whether data is secure and protected from unauthorized access, use, disclosure, damage, modification, or destruction;

[0022] Compliance: used to assess whether data complies with relevant laws, regulations and industry standards;

[0023] S2. Construct an agent model; the agent model is used to obtain the evaluation results of various data quality evaluation indicators, comprehensive evaluation conclusions, decision results and modification suggestions through pre-set programs and pre-learned and trained algorithms; the agent model includes a data perception agent, an evaluation agent, a decision agent and an interaction agent;

[0024] S3, the data perception agent collects data and corresponding metadata information from various data sources in real time and sends it to the evaluation agent;

[0025] S4. The evaluation agent performs data quality analysis and evaluation on the data transmitted by the data perception agent based on the multidimensional data quality evaluation index system established in step S1, the index rules in the rule library, and / or the expert knowledge in the expert library. The evaluation process of each index in the multidimensional data quality evaluation index system includes:

[0026] 1) Accuracy evaluation: The evaluation agent finds the amount of data that does not meet the preset rules or standards according to the preset rules or standards and calculates the error rate, or finds outliers in the data and calculates the outlier ratio;

[0027] 2) Completeness assessment: The assessment agent checks whether there are missing values in the data records and calculates the missing rate, or checks whether data of a specified type is missing;

[0028] 3) Consistency assessment: The assessment agent calculates the consistency rate of data using the associated fields between different data sources, or uses data matching algorithms to verify the data consistency between different data sources or data records;

[0029] 4) Timeliness evaluation: The evaluation agent calculates the timeliness rate based on the difference between the business processing time or data update time and the data storage time, or calculates the average maintenance frequency within a certain evaluation period based on the number of data update maintenance times;

[0030] 5) Uniqueness Assessment: The assessment agent uses a data duplication detection algorithm to find duplicate records in the data and calculate the data duplication rate, or find records in the data with non-unique primary keys and calculate the primary key conflict rate;

[0031] 6) Usability evaluation: The evaluation agent simulates different user roles to access data, checks whether the data can be obtained smoothly, evaluates the compatibility of the data format and the clarity of the data description documents, and calculates the data accessibility score, comprehensibility score and / or ease of use score;

[0032] 7) Relevance Assessment: The assessment agent combines business objectives and data analysis requirements to analyze the degree of relevance between data and business;

[0033] 8) Credibility assessment: Evaluate the source of the Agent query data and / or review the documentation of the data processing process to obtain the credibility assessment results;

[0034] 9) Interpretability evaluation: Evaluate the storage format and visualization of the agent's analysis data, determine the clarity of the naming of data tables or data files, the clarity of the naming of variables or features in data tables or data files, and the suitability and simplicity of the visualization;

[0035] 10) Security Assessment: The Assessment Agent checks the data encryption algorithm and key management and examines the strictness of the data access control policy. It determines whether the data encryption algorithm strength, key update frequency, key distribution mechanism, data access control policy, and data backup and recovery policy meet the preset security requirements or standards for various data types. It then uses text to clearly describe any existing security risks or vulnerabilities.

[0036] 11) Compliance Assessment: The assessment agent assesses whether the data complies with laws and regulations, whether it implements industry standards, whether a data compliance management system has been established, and whether the data has been third-party certified and audited. The specific assessment results are described in text form;

[0037] S5. After completing the evaluation of each dimension, the evaluation agent organizes the evaluation results of each indicator according to the set format, forms an evaluation report containing the evaluation results of each indicator, and sends the evaluation report to the decision agent;

[0038] S6. The decision agent receives the evaluation report from the evaluation agent, combines the evaluation results of various indicators with the preset thresholds and / or business rules, and draws an evaluation conclusion on whether the data quality in each dimension and the overall quality of the data meet the standards. If the data quality meets the standards, the decision agent directly transmits the evaluation conclusion to the interactive agent. If not, the decision agent further generates a decision result and suggestions, and sends the evaluation conclusion, decision result and suggestions to the interactive agent. The decision result and suggestions include specific problems with each indicator in each dimension of the data, as well as modification suggestions corresponding to the problems with each indicator. The modification suggestions corresponding to the problems with each indicator are generated by the decision agent based on the previous learning and training results and preset expert experience.

[0039] S7. The interactive agent interacts with the user and presents the data quality assessment conclusions, decision results and suggestions to the user.

[0040] Furthermore, the data perception agent establishes a connection with the target data source, uses the data acquisition interface and protocol, obtains the target data according to the preset acquisition frequency and rules, and performs preliminary cleaning and format conversion on the acquired data to remove noise data and duplicate data, and then sends the processed data to the evaluation agent.

[0041] Furthermore, in the accuracy evaluation, the error rate formula is:

[0042]

[0043] The evaluation agent also uses the standard deviation, Z-score method or isolation forest algorithm to detect outliers in the data to be evaluated, and then calculates the ratio of the number of detected outliers to the total data volume to obtain the outlier ratio;

[0044] The standard deviation formula is:

[0045]

[0046] Among them, x i is the i-th data value, is the average value of the data, n is the number of data; when a data value x i and mean When the distance is greater than k times σ, it is considered an outlier, where k is a preset value;

[0047] The calculation formula of Z-score is:

[0048]

[0049] Where X is the data value, μ is the mean, and σ is the standard deviation. The absolute value of the Z-score determines the degree of anomaly, and positive or negative only indicates the direction of deviation. When setting the threshold, you need to choose whether to distinguish between positive and negative anomalies based on the actual scenario. If you only focus on excessively high values, the larger the positive Z-score value, the more abnormal the data value. If you need to pay attention to both high and low values, the larger the absolute Z-score value, the more abnormal the data value.

[0050] The Isolation Forest algorithm calculates the anomaly score of the data by constructing a binary decision tree:

[0051]

[0052] n is the number of data, h(x) is the path length of data point x in the isolation forest, E(h(x)) is the mean path length of data x in the tree, and c(n) is the normalization factor; the anomaly score S value is between [0,1], and the larger S is, the more abnormal the data point is.

[0053] Furthermore, in the integrity assessment, the missing rate is calculated as follows:

[0054]

[0055] Furthermore, in the consistency assessment:

[0056] The formula for calculating the consistency rate is:

[0057]

[0058] The data matching algorithm is the cosine similarity algorithm, which is used to calculate the similarity between different data sources. The calculation formula is:

[0059]

[0060] Where A and B are two data vectors. If the data source record is in table form, the statistical value of each feature of the data source is used as a vector element. The closer the calculated cosine similarity is to 1, the higher the consistency of the two data sources on the selected features; the closer it is to 0, the lower the consistency.

[0061] Furthermore, in the timeliness evaluation, the timeliness rate calculation formula is:

[0062]

[0063] Among them, τ is the timely rate, s i is the storage time of the i-th data record, t i is the business processing time or update time of the i-th data record, and n is the total amount of data;

[0064] The formula for calculating the average maintenance frequency is:

[0065]

[0066] Furthermore, in the uniqueness evaluation:

[0067] The calculation formula of the data repetition rate is:

[0068]

[0069] The primary key conflict rate calculation formula is:

[0070]

[0071] Furthermore, in the usability evaluation:

[0072] The accessibility score is calculated as follows:

[0073]

[0074] The calculation formula of the comprehensibility score is:

[0075]

[0076] The calculation formula for the usability score is:

[0077] Usability score = user task completion efficiency score * k1 + user satisfaction score * k2;

[0078] in, k1 and k2 are weights.

[0079] Furthermore, in the correlation evaluation, the evaluation agent uses the Pearson correlation coefficient to calculate the linear correlation between the data and the business indicators. The formula is:

[0080]

[0081] Among them, x i is the i-th record value of the data variable X, y i is the i-th record value of business indicator Y, n is the number of record values, is the average value of n records of data variable X, It is the average value of n record values of business indicator Y; the value range of Pearson correlation coefficient r is [-1,1]. When r=1, it means that there is a completely positive linear correlation between the data variable and the business indicator, that is, when the data variable increases, the business indicator will also increase proportionally; when r=-1, it means that there is a completely negative linear correlation, that is, when the data variable increases, the business indicator will decrease proportionally; when r=0, it means that there is no linear correlation between the two.

[0082] Furthermore, in the credibility assessment, the evaluation agent evaluates the data credibility by evaluating the authority of the data source and / or the completeness of the documentation of the data processing process and the rationality of the processing method;

[0083] The authority of the data source is measured by the authority score. If the data comes from a designated authoritative institution, it is a highly credible source and is assigned a higher authority score. If the data comes from an industry report, it is a moderately credible source and is assigned a moderate authority score. Otherwise, it is a low-credible source and is assigned a lower authority score.

[0084] The integrity of data processing documentation is measured by a completeness score. If the documentation of the data processing process records every step of the data collection, cleaning, conversion, and analysis in detail, and includes information on all key links and parameter settings, then the integrity is high and a high completeness score is assigned. If the documentation records the main steps but lacks some details, then the integrity is fair and a medium completeness score is assigned. If the documentation is very brief and key processes are not recorded, then the integrity is poor and a low completeness score is assigned.

[0085] The rationality of the data processing method is measured by the processing rationality score. If the method used for data processing is a standard method recognized by the industry and is consistent with the data characteristics and research objectives, then the rationality is high and a higher processing rationality score is assigned. If the method used does not belong to the set standard method, but has been reasonably verified and supported by relevant literature, then the rationality is good and a medium processing rationality score is assigned. Otherwise, a lower processing rationality score is assigned.

[0086] Furthermore, in the interpretability evaluation, the evaluation agent evaluates whether the naming of the data table or data file, and / or the names of each variable or feature in the data table or data file are intuitive, whether their meanings are clear and easy to understand, and whether there are relevant documents to explain them. The specific calculation formula for clarity is as follows:

[0087]

[0088] The suitability evaluation of visualization refers to evaluating whether the chart type selected in the data table or data file is suitable for displaying data. The proportion of the number of appropriately selected charts to the total number of charts is calculated to obtain the suitability of visualization:

[0089]

[0090] Visualization clarity and simplicity: Evaluate the elements of the charts obtained by the agent from the data table or data file, determine whether the visualization results are concise and clear, and whether they can clearly convey the main characteristics and trends of the data. Then calculate the ratio of the number of concise and clear charts to the total number of charts to obtain the visualization clarity and simplicity:

[0091]

[0092] The beneficial effects of the present invention are:

[0093] 1. Comprehensiveness improvement: This invention evaluates data quality from multiple dimensions, making the evaluation results more accurate and comprehensive in reflecting the true status of the data;

[0094] 2. Enhanced dynamic adaptability: The present invention utilizes the Agent intelligent body model to dynamically adjust the evaluation strategy and parameters according to the real-time changes of data and the adjustment of business needs, thereby ensuring the accuracy of the evaluation results.

[0095] 3. Intelligent and Efficient: This invention reduces manual intervention and improves evaluation efficiency. This is particularly true for medical data evaluation, where traditional manual evaluation of large amounts of medical records is time-consuming and labor-intensive. Agent-based evaluation methods enable automated and rapid evaluation, reducing evaluation time from days to hours. Furthermore, intelligent analysis can promptly identify potential data quality issues, providing early warnings and solutions to ensure data quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] Figure 1 Schematic diagram of the method flow in Example 1 of the present invention;

[0097] Figure 2 This is a schematic diagram of the method flow in Example 2 of the present invention;

[0098] Figure 3 This is a schematic diagram of the indicator system construction in Example 2 of the present invention;

[0099] Figure 4 Schematic diagram of the method flow in Example 3 of the present invention;

[0100] Figure 5 This is a schematic diagram of constructing an indicator system based on the data itself in Example 3 of the present invention;

[0101] Figure 6 This is a schematic diagram of constructing an indicator system from the data application dimension in Example 3 of the present invention. DETAILED DESCRIPTION

[0102] The present invention will be further described below in conjunction with the accompanying drawings. It should be noted that this embodiment is based on the technical solution and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to this embodiment.

[0103] Example 1

[0104] This embodiment provides an agent-based multi-dimensional data quality intelligent evaluation method, such as Figure 1 As shown, the following steps are included:

[0105] S1. Establish a multidimensional data quality assessment indicator system:

[0106] The method of this embodiment constructs a data quality assessment indicator system from one or more dimensions such as data itself, data application, and data security. Users can add or reduce several dimensions or several indicators in the dimensions according to their actual situation to construct a data quality assessment indicator system that suits their actual needs.

[0107] (1) Data quality assessment indicators based on the data itself include:

[0108] Accuracy: This is used to assess whether data is authentic, reliable, and accurately reflects the real world. Typical accuracy metrics include error rate and outlier ratio. For example, in a financial data quality assessment, the transaction records to be evaluated are compared with authoritative data from the banking system, and the proportion of accurate transaction records in the evaluated transaction records is calculated.

[0109] Completeness: This is used to assess data completeness and the presence of missing or null values. Typical completeness metrics include missing value ratio, field fill rate, and record completeness rate. For example, in a personnel information management system, the percentage of records with complete personal information in the total number of records is counted.

[0110] Consistency: This is used to assess whether data remains consistent across different systems, at different points in time, and / or across different data tables. Typical consistency metrics include data consistency ratio, data conflict rate, and data redundancy. For example, comparing employee basic information recorded by different departments within an enterprise, calculate the percentage of consistent matches across these records.

[0111] Timeliness: This is used to assess whether data is updated promptly and reflects the latest developments. Timeliness indicators generally include data update frequency, data latency, and data validity period. For example, for stock trading data, if the update interval exceeds the specified time, the timeliness score will decrease.

[0112] Uniqueness: This is used to assess whether data contains duplicate records or redundant information. Uniqueness metrics typically include the number of duplicate records, the ratio of unique values, and the primary key conflict rate. For example, in a population information system, the ratio of personnel records with duplicate primary keys to the total personnel records is calculated.

[0113] (2) Data quality assessment indicators based on data application include:

[0114] Usability: This is used to assess whether data is easy to access, understand, and use. Usability metrics generally include data accessibility, data understandability, and data usability. Usability metrics can be measured using scores or grades. For example, if medical data cannot be read by some medical institutions due to format incompatibility, its usability grade will be lowered.

[0115] Relevance: This is used to assess whether data is relevant to business needs and can meet analytical requirements. Relevance metrics typically include data relevance score, data utilization rate, and data value density. For example, a retail company collects a large amount of user browsing data but lacks data on user purchasing behavior. The data analysis team is unable to accurately predict user purchase intentions, and therefore the data is not relevant to business objectives.

[0116] Credibility: This is used to assess the reliability of data sources and the transparency of data processing. Credibility indicators generally include data source credibility, data processing transparency, and data security level. For example, a financial institution may use user credit score data provided by a third-party data provider, but the provider fails to disclose the data source and processing methods, resulting in low data credibility.

[0117] Interpretability: This is used to assess whether data is easy to understand and interpret, and whether it can support decision-making. Interpretability metrics generally include data interpretability scores, data visualization, and data storytelling. For example, patient health data in medical institutions is stored in complex encodings, making it difficult for doctors to understand the data and making rapid diagnostic decisions, resulting in low interpretability.

[0118] (3) Data quality assessment indicators from the perspective of data security include:

[0119] Security: This is used to assess whether data is secure and protected from unauthorized access, use, disclosure, damage, modification, or destruction. Security indicators generally include data encryption level, data access control level, and data security incident rate. For example, if a company's customer data is stored in an unencrypted database and the database is hacked, resulting in a leak of customer data, the data would be considered to have low security.

[0120] Compliance: This is used to assess whether data complies with relevant laws, regulations, and industry standards. Compliance indicators generally include data compliance scores and data privacy protection levels. For example, an e-commerce platform may have violated laws and regulations by not handling user data in accordance with the Personal Information Protection Law, resulting in low data compliance.

[0121] S2. Construct an agent model. This agent model is used to derive evaluation results, comprehensive evaluation conclusions, decision-making results, and modification recommendations for various data quality assessment indicators through pre-set programs and pre-learned and trained algorithms. The agent model includes a data perception agent, an evaluation agent, a decision-making agent, and an interaction agent.

[0122] S3. The Data Perception Agent collects data and corresponding metadata from various data sources in real time. Specifically, the Data Perception Agent establishes connections with target data sources such as databases, file systems, and network interfaces. It utilizes data collection interfaces and protocols to acquire target data according to pre-set collection frequencies and rules. It then performs preliminary cleaning and format conversion on the acquired data to remove noise and duplicate data, and then sends the processed data to the Evaluation Agent.

[0123] S4. The evaluation agent performs data quality analysis and evaluation on the data transmitted by the data perception agent based on the multidimensional data quality evaluation indicator system established in step S1, the indicator rules in the rule library (including metadata standards, such as whether data item fields are required or must meet certain standards. Indicator rules can be configured according to needs), and / or the expert knowledge in the expert library (i.e., thresholds or weights. For example, if the error rate, missing rate, or data anomaly ratio exceeds a certain threshold, the data quality is considered substandard). After receiving the data from the perception agent, the evaluation agent will perform a series of complex and detailed tasks based on the evaluation indicators of different dimensions.

[0124] The evaluation process of each indicator in the multidimensional data quality evaluation indicator system is as follows:

[0125] 1) Accuracy evaluation: The evaluation agent finds the amount of data items that do not meet the preset rules or standards according to the preset rules or standards, and then calculates the error rate;

[0126] The error rate formula is:

[0127]

[0128] In addition, the evaluation agent can use the standard deviation, Z-score method or isolation forest algorithm to detect outliers in the data to be evaluated, and then calculate the ratio of the number of detected outliers to the total data volume to obtain the outlier ratio.

[0129] The standard deviation formula is:

[0130]

[0131] Among them, x i is the i-th data value, x is the average value of the data, and n is the number of data. i When the distance from the mean x is greater than k times σ (k is the threshold), it is considered an outlier.

[0132] The calculation formula of Z-score is:

[0133]

[0134] Where X is the data value, μ is the mean, and σ is the standard deviation. The absolute value of the Z-score determines the degree of anomaly; a positive or negative value only indicates the direction of deviation. When setting the threshold (e.g., |Z| > 3), you should choose whether to distinguish between positive and negative anomalies based on the actual scenario. If you are only concerned with excessively high values (such as financial fraud detection), a larger positive Z-score value indicates a more abnormal data value. If both high and low values are important (such as quality control), a larger absolute Z-score value indicates a more abnormal data value.

[0135] The Isolation Forest algorithm calculates the anomaly score of the data by constructing a binary decision tree:

[0136]

[0137] n is the number of data points, h(x) is the path length of data point x in the isolation forest (i.e., the number of edges from the root node to the leaf node containing x), E(h(x)) is the mean path length of data point x in the tree, and c(n) is the normalization factor. The anomaly score S ranges from 0 to 1, with larger values indicating more abnormal data points.

[0138] 2) Completeness Assessment: The assessment agent checks whether there are missing values in the data records and calculates the missing rate. The missing rate calculation formula is:

[0139]

[0140] 3) Consistency assessment: The assessment agent calculates the consistency rate of data using the associated fields between different data sources, or uses data matching algorithms to verify the data consistency between different data sources or data records.

[0141] The formula for calculating the consistency rate is:

[0142]

[0143] The cosine similarity algorithm can be used as a data matching algorithm to calculate the similarity between different data sources. The calculation formula is:

[0144]

[0145] Where A and B are two data vectors. If the data source records are in tabular format, the statistical value of each feature of the data source (such as mean, sum, or frequency) can be used as a vector element. Generally, the closer the calculated cosine similarity is to 1, the more consistent the two data sources are on the selected features; the closer it is to 0, the lower the consistency. In practical applications, a threshold can be set as needed. Values above the threshold indicate high consistency of the data sources on these features, while values below the threshold indicate poor consistency.

[0146] 4) Timeliness evaluation: The evaluation agent calculates the timeliness rate based on the difference between the business processing time or data update time and the data storage time, or uses the number of data update maintenance times to calculate the average maintenance frequency within a certain evaluation period.

[0147] The timeliness calculation formula is:

[0148]

[0149] Among them, τ is the timely rate, s i is the storage time of the i-th data record, t i is the business processing time or update time of the i-th data record, and n is the total amount of data.

[0150] The formula for calculating the average maintenance frequency is:

[0151]

[0152] 5) Uniqueness Assessment: The assessment agent uses a data duplication detection algorithm to find duplicate records in the data and calculate the data duplication rate, or find records in the data with non-unique primary keys and calculate the primary key conflict rate.

[0153] The formula for calculating the data repetition rate is:

[0154]

[0155] The formula for calculating the primary key conflict rate is:

[0156]

[0157] 6) Usability evaluation: The evaluation agent simulates different user roles to access the data, checks whether the data can be obtained smoothly, evaluates the compatibility of the data format and the clarity of the data description document, and calculates the accessibility score, comprehensibility score and / or ease of use score of the data.

[0158] The accessibility score is calculated as:

[0159]

[0160] The formula for calculating the comprehensibility score is:

[0161]

[0162] The usability score is calculated as follows:

[0163] Usability score = user task completion efficiency score * k1 + user satisfaction score * k2

[0164] in, k1 and k2 are weights.

[0165] 7) Correlation Assessment: The assessment agent analyzes the degree of correlation between data and business based on business goals and data analysis requirements. Specifically, the assessment agent can use the Pearson correlation coefficient to calculate the linear correlation between data and business indicators. The formula is:

[0166]

[0167] Among them, x i is the i-th record value of the data variable X, y i is the i-th record value of business indicator Y, n is the number of record values, is the average value of n records of data variable X, is the average of n recorded values of business indicator Y. The Pearson correlation coefficient r ranges from -1 to 1. When r = 1, there is a perfect positive linear correlation between the data variable and the business indicator. That is, as the data variable increases, the business indicator also increases proportionally. When r = -1, there is a perfect negative linear correlation. That is, as the data variable increases, the business indicator decreases proportionally. When r = 0, there is no linear correlation between the two, but other nonlinear relationships may exist. In real-world business, it is generally believed that when |r| > 0.7, the correlation between data and business is strong; when 0.3 < |r| < 0.7, the correlation is moderate; and when |r| < 0.3, the correlation is weak.

[0168] 8) Credibility Assessment: Evaluate the source of the Agent's query data and / or review the documentation of the data processing process (if any). Specifically, the credibility of the data is assessed by evaluating the authority of the data source and / or the completeness of the documentation of the data processing process and the rationality of the processing method.

[0169] The authority of the data source is measured by the authority score. If the data comes from a designated authoritative institution, such as a government department, a well-known scientific research institution, an industry-leading research company, etc., it is a highly credible source and is assigned a higher authority score; if the data comes from an industry report, it is a medium credible source and is assigned a medium authority score; otherwise, it is a low credible source (such as some ordinary corporate or personal websites) and is assigned a lower authority score.

[0170] The integrity of data processing documents is measured by the integrity score. If the documentation of the data processing process records in detail every step from data collection, cleaning, conversion to analysis, and includes information such as all key links and parameter settings, then the integrity is very high and is assigned a high integrity score. If the document records the main steps but lacks some details, such as the specific threshold settings during data cleaning, then the integrity is average and is assigned a medium integrity score. If the document record is very brief, only mentioning the general processing direction, and no key processes are recorded, then the integrity is poor and is assigned a low integrity score.

[0171] The rationality of the data processing method is measured by the processing rationality score. If the method used for data processing is a standard method recognized by the industry and is consistent with the data characteristics and research objectives, then the rationality is high and a higher processing rationality score is assigned. If the method used does not belong to the set standard method, but has been reasonably verified and supported by relevant literature, then the rationality is good and a medium processing rationality score is assigned. Otherwise, a lower processing rationality score is assigned.

[0172] 9) Interpretability evaluation: Evaluate the storage format and visualization effect of the agent's analysis data, judge the clarity of the naming of data tables or data files, the clarity of the naming of variables or features in data tables or data files, and the suitability, simplicity and clarity of the visualization.

[0173] The evaluation agent evaluates whether the naming of the data table or data file, and / or the names of each variable or feature in the data table or data file are intuitive, whether their meanings are clear and easy to understand, and whether there is relevant documentation to explain them. The specific calculation formula for clarity is as follows:

[0174]

[0175] Applicability of visualization: The evaluation agent evaluates whether the chart type selected in the data table or data file is suitable for displaying the data (this evaluation can be performed based on user feedback). The proportion of appropriately selected charts to the total number of charts is calculated to obtain the suitability of visualization:

[0176]

[0177] Visualization Clarity: The evaluation agent obtains elements such as axis labels and legends from charts in data tables or data files, and determines whether the charts are clear and easy to understand, whether the color combination is reasonable, whether the visualization results are concise and clear, and whether they can clearly convey the main characteristics and trends of the data (this can be evaluated based on user feedback). Then, the ratio of concise and clear charts to the total number of charts is calculated to obtain the visualization clarity:

[0178]

[0179] 10) Security Assessment: The assessment agent detects the data encryption algorithm and key management and reviews the strictness of the data access control policy. It determines whether the data encryption algorithm strength, key update frequency, key distribution mechanism, data access control policy, data backup and recovery policy meet the preset security requirements or standards of various types of data, as well as the historical incidence of data security incidents, and then clearly points out the existing security risks or vulnerabilities in text descriptions.

[0180] 11) Compliance Assessment: The assessment agent assesses whether the data complies with laws and regulations, whether it implements industry standards, whether a data compliance management system is established, and the third-party certification and audit status of the data, and describes the specific assessment results in text form.

[0181] S5. After completing the evaluation of each dimension, the evaluation agent organizes the evaluation results of each indicator according to the specified format, generating an evaluation report containing the evaluation results of each indicator (calculated values or text descriptions). This evaluation report is then sent to the decision agent, providing a key basis for subsequent data quality management decisions. The evaluation agent also regularly updates the evaluation results to address dynamic data changes and ensure that data quality is always under effective monitoring.

[0182] S6. The decision agent receives the evaluation report from the evaluation agent, combines the evaluation results of various indicators with the preset thresholds and / or business rules, and draws a conclusion on whether the data quality in each dimension and the overall quality of the data meet the standards. If the data quality meets the standards, the decision agent will pass the evaluation conclusion directly to the interactive agent. If it does not meet the standards, the decision agent will further generate decision results and suggestions, and send the evaluation conclusions, decision results and suggestions to the interactive agent; the decision results and suggestions include specific problems with the data in each indicator of each dimension, as well as modification suggestions corresponding to the problems with each indicator. The modification suggestions corresponding to the problems of each indicator are generated by the decision agent based on the results of previous learning and training and preset expert experience.

[0183] S7. The interactive agent interacts with the user and presents the data quality assessment conclusions, decision results and suggestions to the user in the form of visual reports, charts, etc.

[0184] Furthermore, the interactive agent also receives user feedback and instructions, and sends instructions such as adjusting evaluation parameters and re-evaluation to other agents according to user needs.

[0185] In this embodiment, after the data perception agent collects and processes data, it sends the data to the evaluation agent through the message queue; after the evaluation agent completes the evaluation of various indicators, it encapsulates the evaluation report in JSON format and passes it to the decision agent through the network interface; the decision agent generates evaluation conclusions, decision results and suggestions based on the evaluation report, and then sends them to the interaction agent in XML format; after receiving the instructions, the interaction agent sends control instructions to other agents through the RESTful API to achieve collaborative work between agents.

[0186] It's important to note that the message queue uses a publish-subscribe model. Data-aware agents, acting as producers, publish data to the message queue, while evaluation agents, acting as consumers, retrieve data from the queue. This model decouples the data collection and evaluation processes, improving the system's asynchronous processing capabilities and scalability. The RESTful API is based on the HTTP protocol and uses standard HTTP methods (such as GET, POST, PUT, and DELETE) for resource operations and interactions. Interaction agents send commands to other agents through the RESTful API, following unified resource location and operation specifications to ensure standardized and compatible communication between agents.

[0187] Example 2

[0188] This embodiment provides an application example of the method of embodiment 1, including the following steps:

[0189] S1. Establish a data quality assessment indicator system:

[0190] This embodiment only evaluates the data quality of the data itself according to the actual application and business needs of specific population management. Figure 2 As shown, this embodiment selects data integrity, accuracy, consistency, timeliness, and uniqueness as evaluation indicators, and adds progressive evaluation indicators. It compares the current period and base period of the evaluation indicators, analyzes the quality changes of data from various provinces and departments within a certain period, promptly informs and pays attention to any anomalies, and gives praise when significant improvements in data quality are found.

[0191] S2. Build Agent Intelligent Model

[0192] ① Data Perception Agent: This agent establishes connections with data sources, such as those in various locations or other departments, and utilizes corresponding data aggregation interfaces and protocols to collect population data and metadata daily. During the aggregation process, the agent performs preliminary data cleaning to remove noise and duplicate records, such as duplicate registration information and invalid medical records. Furthermore, the agent converts data in different formats (for example, unifying different date formats) to conform to internal data processing standards. Once processed, the data is sent to the evaluation agent.

[0193] ② Evaluation Agent: Based on the data quality evaluation index system constructed in step S1, each indicator is evaluated, and finally a comprehensive evaluation result is given based on the evaluation result of each indicator (evaluation score or evaluation conclusion), such as Figure 3 The following is a detailed example of the completeness and accuracy assessment process.

[0194] Integrity Assessment: Regarding data integrity, the reporting status of four types of basic personnel information and four types of extended information (e.g., case information, feature information, relationship information, and execution information) is checked to determine the status of eight sub-indicators of data integrity. For logical verification between data tables, correlation analysis is performed across different data tables to calculate four sub-indicators, such as the missing rate for basic personnel information but not extended information. For example, basic personnel information and business information are correlated to identify cases where business information exists but basic personnel information is missing. This issue is then reported to local authorities to determine whether basic personnel information has been omitted or business information has been misreported. For maintenance anomalies, decreases in four types of information, including basic personnel information, within a certain period are counted to monitor for abnormal changes in data volume. Since these indicators include both judgmental and quantitative indicators, each indicator is first standardized, then weighted using the analytic hierarchy process or entropy weighting method. Finally, a weighted summation method is used to determine the integrity indicator assessment score.

[0195] Accuracy Assessment: Based on relevant industry standards within the rule base, error rate indicators are calculated for key data items (configurable). These include the error rate for key data items in basic personnel information (e.g., ID card number verification errors, address format errors, and irregular mobile phone numbers) and the error rate for key data items in extended AD information (e.g., non-dictionary items for "case category" in case information). Error details are also identified and sent to the decision-making agent, who then decides whether to disseminate the information to local management agencies or other departments.

[0196] At the same time, duplicate registrations of mobile phone numbers are detected to prevent data entry errors. After standardizing and de-dimensionalizing each indicator, the accuracy dimension can be evaluated using the analytic hierarchy process, entropy weight method, or a combination of the two.

[0197] ③ The Decision Agent: After receiving the results from the Evaluation Agent, it conducts an overall assessment of data quality based on pre-set thresholds and business rules. If the data quality meets the standards, the assessment results are passed to the Interaction Agent. If not, the type of data quality issue is analyzed. For example, if accuracy is a serious issue, details of the data errors and data calibration suggestions are sent to local management agencies. Local authorities will investigate the cause, correct the erroneous data, and resubmit it. If completeness issues are significant, a data supplementation plan is developed, such as requiring community staff to provide missing residential information for specific populations. The decision results and recommendations are then sent to the Interaction Agent.

[0198] ④ Interactive Agent: This displays data quality assessment results to relevant personnel in the form of visual reports and charts. These reports detail the indicator scores, assessment conclusions, and any issues for each data source. The Interactive Agent also receives user feedback and instructions. Based on user needs, it sends instructions to the Data Perception Agent and the Evaluation Agent via a RESTful API to adjust assessment parameters, reassess, and so on. For example, if a user discovers that the assessment weighting for a specific population's data is unreasonable, the Interactive Agent will receive feedback and send instructions to the Evaluation Agent to adjust the assessment parameters.

[0199] In this embodiment, after the Data Perception Agent collects data on a specific population, it asynchronously sends the data to the Evaluation Agent via a message queue. After processing the data, the Evaluation Agent encapsulates the evaluation results in JSON format and transmits them to the Decision Agent via a network interface. The Decision Agent makes a decision based on the evaluation results and sends the results and recommendations in XML format to the Interaction Agent. After receiving the instructions, the Interaction Agent sends control commands to the Data Perception Agent and the Evaluation Agent via a RESTful API, enabling inter-agent collaboration and ensuring the smooth operation of the entire specific population data quality assessment process.

[0200] Example 3

[0201] In the financial credit business scenario, data quality directly affects the accuracy of credit decisions and the effectiveness of risk control. This embodiment uses an agent-based multidimensional data quality assessment method to perform a comprehensive and efficient quality assessment of financial credit data.

[0202] 1) Establish a data quality assessment indicator system

[0203] In financial credit, we should consider comprehensively evaluating data quality from two major dimensions: data itself and data application. For example, Figure 4 As shown in the figure, the indicator system built according to the business scenario includes accuracy, completeness, consistency, timeliness, availability, relevance, credibility and explainability.

[0204] S2. Building Agent Model

[0205] ① Data Perception Agent: The Data Perception Agent is responsible for collecting data and metadata from various data sources. Data sources include the bank's internal customer information database and transaction record system, external credit reporting systems, and third-party data platforms (such as the business registration information platform to obtain business customer operating status data). The Data Perception Agent utilizes database connection interfaces and network data scraping tools to acquire data according to pre-set collection frequencies (e.g., collecting customer transaction records daily at midnight, or obtaining credit data from the credit reporting system weekly) and rules (e.g., collecting only data fields related to specific credit operations). After collection, the data undergoes preliminary cleansing to identify and remove duplicate transaction records and malformed data (e.g., non-numeric characters in the amount field). The data is also converted to a format that can be recognized and processed internally by the system (e.g., standardizing the date format across different data sources to "YYYY-MM-DD"). Finally, the processed data is sent to a message queue for acquisition by the Evaluation Agent.

[0206] ② Evaluation Agent: After obtaining data from the message queue, the evaluation agent calculates each sub-index according to the preset index system in the system, and then calculates each dimension based on the index calculation, and finally gives a comprehensive evaluation score, such as Figure 5-6 The following is a specific example.

[0207] Data Accuracy Assessment: A customer's credit history is compared with authoritative data from the central bank's credit reporting system. A data comparison algorithm is used to calculate the error rate of the credit history data. Furthermore, standard deviation and the isolation forest algorithm are used to detect outliers in the customer's financial data. For example, if the standard deviation of a customer's reported income data is excessive and the isolation forest algorithm identifies it as an outlier, the data is flagged as suspicious.

[0208] Data application relevance assessment: Analyze the relevance of a customer's transaction and asset data to business needs such as credit approval and risk assessment, and calculate a data relevance score. For example, for small consumer credit businesses, a customer's daily consumption transaction data is highly relevant; the lack of such data can affect the assessment of the customer's repayment ability.

[0209] ③ Decision Agent: The Decision Agent receives the results from the Evaluation Agent and, based on preset thresholds (e.g., a 5% accuracy error rate threshold and a 10% completeness missing value ratio threshold) and business rules (e.g., higher data quality requirements for high-risk customers), makes an overall assessment of data quality. If the data quality meets the standards, the assessment results are passed to the Interaction Agent. If not, the data quality issue is analyzed. For example, if the accuracy error rate exceeds 5%, the accuracy issue is considered serious and data calibration recommendations are generated, such as re-verifying the source of the customer's credit record. If the completeness missing value ratio exceeds 10%, a data supplementation plan is developed, such as requiring loan officers to supplement the customer's missing financial information. The Decision Agent sends the results and recommendations to the Interaction Agent in XML format.

[0210] ④ Interactive Agent: The Interactive Agent presents the data quality assessment results to users such as credit approval officers and risk management personnel in the form of visual reports and charts. The reports detail the scores of each indicator, the assessment conclusions, and any existing issues. For example, a bar chart can be used to display the data accuracy scores of different customer groups, and a line chart can be used to display the trend of data integrity changes over time. At the same time, the Interactive Agent receives user feedback and instructions. If a user finds that certain assessment parameters are set unreasonably (such as the time threshold for timeliness assessment is set too long), the Interactive Agent sends an instruction to adjust the assessment parameters to the Assessment Agent via the RESTful API. If a user requests a reassessment of a batch of credit data, the Interactive Agent sends a reassessment instruction to the Assessment Agent. At the same time, the Interactive Agent can also convey user feedback and needs to the Data Perception Agent to adjust the frequency and rules of data collection (such as increasing the collection frequency of certain key data fields).

[0211] In this embodiment, after the Data Perception Agent collects data, it asynchronously sends it to the Evaluation Agent via a message queue, ensuring efficient and stable data transmission. After processing the data, the Evaluation Agent encapsulates the evaluation results in JSON format and transmits them to the Decision Agent via a network interface. The Decision Agent makes a decision based on the evaluation results and sends the results and recommendations in XML format to the Interaction Agent. After receiving the instructions, the Interaction Agent sends control commands to the Data Perception Agent and the Evaluation Agent via a RESTful API, enabling inter-agent collaboration and ensuring the smooth operation of the entire financial credit data quality assessment process.

[0212] Those skilled in the art can make various corresponding changes and modifications based on the above technical solutions and concepts, and all of these changes and modifications should be included in the scope of protection of the claims of the present invention.

Claims

1. An agent-based multidimensional data quality intelligent assessment method, characterized in that: The steps include: S1. Establish a multidimensional data quality assessment indicator system: Build a data quality assessment indicator system from one or more dimensions of data itself, data application, and data security; (1) Data quality assessment indicators based on the data itself include: Accuracy: used to assess whether the data is true, reliable, and accurately reflects the real world; Completeness: used to assess whether the data is complete and whether there are missing or null values; Consistency: used to evaluate whether data is consistent across different systems, at different time points, and / or across different data tables; Timeliness: used to assess whether the data is updated in a timely manner and whether it can reflect the latest situation; Uniqueness: used to evaluate whether the data contains duplicate records or redundant information; (2) Data quality assessment indicators based on data application include: Usability: used to assess whether the data is easy to access, understand and use; Relevance: used to assess whether the data is relevant to business needs and whether it can meet analytical requirements; Credibility: used to assess whether the data source is reliable and whether the data processing process is transparent; Interpretability: used to evaluate whether the data is easy to understand and interpret and whether it can support decision-making; (3) Data quality assessment indicators from the perspective of data security include: Security: used to assess whether data is secure and protected from unauthorized access, use, disclosure, damage, modification, or destruction; Compliance: used to assess whether data complies with relevant laws, regulations and industry standards; S2. Construct an agent model; the agent model is used to obtain the evaluation results of various data quality evaluation indicators, comprehensive evaluation conclusions, decision results and modification suggestions through pre-set programs and pre-learned and trained algorithms; the agent model includes a data perception agent, an evaluation agent, a decision agent and an interaction agent; S3, the data perception agent collects data and corresponding metadata information from various data sources in real time and sends it to the evaluation agent; S4. The evaluation agent performs data quality analysis and evaluation on the data transmitted by the data perception agent based on the multidimensional data quality evaluation index system established in step S1, the index rules in the rule library, and / or the expert knowledge in the expert library. The evaluation process of each index in the multidimensional data quality evaluation index system includes: 1) Accuracy evaluation: The evaluation agent finds the amount of data that does not meet the preset rules or standards according to the preset rules or standards and calculates the error rate, or finds outliers in the data and calculates the outlier ratio; 2) Completeness assessment: The assessment agent checks whether there are missing values in the data records and calculates the missing rate, or checks whether data of a specified type is missing; 3) Consistency assessment: The assessment agent calculates the consistency rate of data using the associated fields between different data sources, or uses data matching algorithms to verify the data consistency between different data sources or data records; 4) Timeliness evaluation: The evaluation agent calculates the timeliness rate based on the difference between the business processing time or data update time and the data storage time, or calculates the average maintenance frequency within a certain evaluation period based on the number of data update maintenance times; 5) Uniqueness Assessment: The assessment agent uses a data duplication detection algorithm to find duplicate records in the data and calculate the data duplication rate, or find records in the data with non-unique primary keys and calculate the primary key conflict rate; 6) Usability evaluation: The evaluation agent simulates different user roles to access data, checks whether the data can be obtained smoothly, evaluates the compatibility of the data format and the clarity of the data description documents, and calculates the data accessibility score, comprehensibility score and / or ease of use score; 7) Relevance Assessment: The assessment agent combines business objectives and data analysis requirements to analyze the degree of relevance between data and business; 8) Credibility assessment: Evaluate the source of the Agent query data and / or review the documentation of the data processing process to obtain the credibility assessment results; 9) Interpretability evaluation: Evaluate the storage format and visualization of the agent's analysis data, determine the clarity of the naming of data tables or data files, the clarity of the naming of variables or features in data tables or data files, and the suitability and simplicity of the visualization; 10) Security Assessment: The Assessment Agent checks the data encryption algorithm and key management and examines the strictness of the data access control policy. It determines whether the data encryption algorithm strength, key update frequency, key distribution mechanism, data access control policy, and data backup and recovery policy meet the preset security requirements or standards for various data types. It then uses text to clearly describe any existing security risks or vulnerabilities. 11) Compliance Assessment: The assessment agent assesses whether the data complies with laws and regulations, whether it implements industry standards, whether a data compliance management system has been established, and whether the data has been third-party certified and audited. The specific assessment results are described in text form; S5. After completing the evaluation of each dimension, the evaluation agent organizes the evaluation results of each indicator according to the set format, forms an evaluation report containing the evaluation results of each indicator, and sends the evaluation report to the decision agent; S6. The decision agent receives the evaluation report from the evaluation agent, combines the evaluation results of various indicators with the preset thresholds and / or business rules, and draws an evaluation conclusion on whether the data quality in each dimension and the overall quality of the data meet the standards. If the data quality meets the standards, the decision agent directly transmits the evaluation conclusion to the interactive agent. If not, the decision agent further generates a decision result and suggestions, and sends the evaluation conclusion, decision result and suggestions to the interactive agent. The decision result and suggestions include specific problems with each indicator in each dimension of the data, as well as modification suggestions corresponding to the problems with each indicator. The modification suggestions corresponding to the problems with each indicator are generated by the decision agent based on the previous learning and training results and preset expert experience. S7. The interactive agent interacts with the user and presents the data quality assessment conclusions, decision results and suggestions to the user.

2. The method according to claim 1, characterized in that The data perception agent establishes a connection with the target data source, uses the data acquisition interface and protocol, obtains the target data according to the preset acquisition frequency and rules, performs preliminary cleaning and format conversion on the acquired data, removes noise data and duplicate data, and then sends the processed data to the evaluation agent.

3. The method according to claim 1, characterized in that In accuracy evaluation, the error rate formula is: The evaluation agent also uses the standard deviation, Z-score method or isolation forest algorithm to detect outliers in the data to be evaluated, and then calculates the ratio of the number of detected outliers to the total data volume to obtain the outlier ratio; The standard deviation formula is: Among them, x i is the i-th data value, is the average value of the data, n is the number of data; when a data value x i and mean When the distance is greater than k times σ, it is considered an outlier, where k is a preset value; The calculation formula of Z-score is: Where X is the data value, μ is the mean, and σ is the standard deviation. The absolute value of the Z-score determines the degree of anomaly, and positive or negative only indicates the direction of deviation. When setting the threshold, you need to choose whether to distinguish between positive and negative anomalies based on the actual scenario. If you only focus on excessively high values, the larger the positive Z-score value, the more abnormal the data value. If you need to pay attention to both high and low values, the larger the absolute Z-score value, the more abnormal the data value. The Isolation Forest algorithm calculates the anomaly score of the data by constructing a binary decision tree: n is the number of data, h(x) is the path length of data point x in the isolation forest, E(h(x)) is the mean path length of data x in the tree, and c(n) is the normalization factor; the anomaly score S value is between [0,1], and the larger S is, the more abnormal the data point is.

4. The method according to claim 1, wherein In the completeness assessment, the missing rate is calculated as:

5. The method according to claim 1, wherein In the consistency assessment: The formula for calculating the consistency rate is: The data matching algorithm is the cosine similarity algorithm, which is used to calculate the similarity between different data sources. The calculation formula is: Where A and B are two data vectors. If the data source record is in table form, the statistical value of each feature of the data source is used as a vector element. The closer the calculated cosine similarity is to 1, the higher the consistency of the two data sources on the selected features; the closer it is to 0, the lower the consistency.

6. The method according to claim 1, characterized in that In the timeliness evaluation, the timeliness rate calculation formula is: Among them, τ is the timely rate, s i is the storage time of the i-th data record, t i is the business processing time or update time of the i-th data record, and n is the total amount of data; The formula for calculating the average maintenance frequency is:

7. The method according to claim 1, characterized in that Uniqueness evaluation: The calculation formula of the data repetition rate is: The primary key conflict rate calculation formula is:

8. The method according to claim 1, characterized in that In a usability review: The accessibility score is calculated as follows: The calculation formula of the comprehensibility score is: The calculation formula for the usability score is: Usability score = user task completion efficiency score * k1 + user satisfaction score * k2; in, k1 and k2 are weights.

9. The method according to claim 1, characterized in that In the correlation evaluation, the evaluation agent uses the Pearson correlation coefficient to calculate the linear correlation between the data and the business indicators. The formula is: Among them, x i is the i-th record value of the data variable X, y i is the i-th record value of business indicator Y, n is the number of record values, is the average value of n records of data variable X, It is the average value of n record values of business indicator Y; the value range of Pearson correlation coefficient r is [-1,1]. When r=1, it means that there is a completely positive linear correlation between the data variable and the business indicator, that is, when the data variable increases, the business indicator will also increase proportionally; when r=-1, it means that there is a completely negative linear correlation, that is, when the data variable increases, the business indicator will decrease proportionally; when r=0, it means that there is no linear correlation between the two.

10. The method according to claim 1, characterized in that In the credibility assessment, the evaluation agent evaluates the data credibility by evaluating the authority of the data source and / or the completeness of the documentation of the data processing process and the rationality of the processing method; The authority of the data source is measured by the authority score. If the data comes from a set authoritative institution, it is a highly credible source and is given a higher authority score. If the data comes from an industry report, it is a medium-trusted source and is assigned a medium authority score; Otherwise, it is a low-trust source and is assigned a lower authority score; The integrity of data processing documentation is measured by a completeness score. If the documentation of the data processing process records every step of the data collection, cleaning, conversion, and analysis in detail, and includes information on all key links and parameter settings, then the integrity is high and a high completeness score is assigned. If the documentation records the main steps but lacks some details, then the integrity is fair and a medium completeness score is assigned. If the documentation is very brief and key processes are not recorded, then the integrity is poor and a low completeness score is assigned. The rationality of the data processing method is measured by the processing rationality score. If the method used for data processing is a standard method recognized by the industry and is consistent with the data characteristics and research objectives, then the rationality is high and a higher processing rationality score is assigned. If the method used does not belong to the set standard method, but has been reasonably verified and supported by relevant literature, then the rationality is good and a medium processing rationality score is assigned. Otherwise, a lower processing rationality score is assigned.

11. The method according to claim 1, wherein In the interpretability evaluation, the evaluation agent evaluates whether the naming of the data table or data file, and / or the names of each variable or feature in the data table or data file are intuitive, whether their meanings are clear and easy to understand, and whether there is relevant documentation to explain them. The specific calculation formula for clarity is as follows: The suitability evaluation of visualization refers to evaluating whether the chart type selected in the data table or data file is suitable for displaying data. The proportion of the number of appropriately selected charts to the total number of charts is calculated to obtain the suitability of visualization: Visualization clarity and simplicity: Evaluate the elements of the charts obtained by the agent from the data table or data file, determine whether the visualization results are concise and clear, and whether they can clearly convey the main characteristics and trends of the data. Then calculate the ratio of the number of concise and clear charts to the total number of charts to obtain the visualization clarity and simplicity:

Citation Information

Cited By

  • Three-dimensional visual anaphora data quality evaluation method based on multi-role multi-view reasoning

    CN120833341A

  • Multi-domain task processing method, system and equipment based on knowledge fusion and Agent cooperation and medium

    CN120996215A

  • Data quality management system and method for cyberspace security compliance field

    CN121530629A

  • Substation field safety control system and method, storage medium and computer equipment

    CN121770153A