Data quality evaluation method and system based on semantic analysis
Patent Information
- Application Number
- CN202611311471.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-27
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]本发明提供了基于语义分析的数据质量评估方法及系统,用于解决现有技术难以自动化验证元数据描述与数据内容之间语义一致性的技术问题
第一,本发明采用规则引擎、统计画像与语义比对三层递进式检查架构,从确定性规则匹配到统计特征比对再到语义相似度计算,逐层深入,将复杂度从低到高依次引入,在保证检查覆盖率的同时控制了计算成本。规则引擎捕获结构性错误,统计画像比对发现数据内容与同类资源的分布偏差,语义比对在边界情况下对元数据描述与数据内容进行语义一致性验证,三层互补覆盖了从格式正确性到语义一致性的完整检查维度。
Smart Images

Figure CN122819263A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data quality management technology, specifically to a data quality assessment method and system based on semantic analysis. Background Technology
[0002] In data circulation and utilization platforms, data resources and products undergo quality audits from registration, cataloging, and uploading to final delivery. Currently, metadata quality management primarily relies on rule-based validation and manual review. Rule-based validation checks metadata against predefined field formats and value ranges, but can only detect structural errors and cannot identify semantic inconsistencies. For example, if a metadata description claims data covers all hospitals in a region while the actual data only includes a subset, rule-based validation cannot detect this discrepancy. Manual review relies on the industry experience and expertise of approvers, comparing metadata descriptions against data samples line by line. This is inefficient, subjective, and cannot achieve full coverage in scenarios with large volumes of data resources. When a data provider submits a metadata description for a data resource, how to automatically and interpretably verify the consistency between the description and the underlying actual data content is a technical problem that needs to be solved in this field. Summary of the Invention
[0003] This invention provides a data quality assessment method and system based on semantic analysis, which addresses the technical problem that existing technologies struggle to automatically verify the semantic consistency between metadata descriptions and data content.
[0004] In a first aspect, the present invention provides a data quality assessment method based on semantic analysis, the method comprising: Receive a request to list the target data resource, and determine the metadata description information and data content sample of the target data resource; Obtain the historical quality assessment records of the supplier corresponding to the target data resource, determine the supplier's credit rating based on the historical quality assessment records, and determine the quality inspection strategy based on the credit rating. According to the quality inspection strategy, a multi-level progressive evaluation is performed on the metadata description information and the data content sample to obtain a multi-level evaluation result. The multi-level progressive evaluation includes at least the following in sequence: structured matching inspection based on preset rules, profile comparison based on statistical features, and semantic similarity comparison based on semantic models. Based on the quality inspection strategy, the obtained inspection results are integrated and scored to generate a quality assessment report for the target data resource.
[0005] Secondly, the present invention also provides a data quality assessment system based on semantic analysis, the system comprising: The request receiving and sampling module is used to receive the listing request of the target data resource and determine the metadata description information and data content sample of the target data resource; The reputation assessment and strategy determination module is used to obtain the historical quality assessment records of the supplier corresponding to the target data resource, determine the supplier's reputation level based on the historical quality assessment records, and determine the quality inspection strategy based on the reputation level. The multi-level progressive evaluation module is used to perform multi-level progressive evaluation on the metadata description information and the data content sample according to the quality inspection strategy to obtain multi-level evaluation results. The multi-level progressive evaluation includes at least the following in sequence: structured matching inspection based on preset rules, profile comparison based on statistical features, and semantic similarity comparison based on semantic models. The scoring fusion and report generation module is used to fuse and score the obtained inspection results according to the quality inspection strategy, and generate a quality assessment report of the target data resource.
[0006] One or more technical solutions provided in this invention have at least the following technical effects or advantages: First, this invention employs a three-layer progressive inspection architecture: rule engine, statistical profiling, and semantic comparison. It progresses from deterministic rule matching to statistical feature comparison and then to semantic similarity calculation, gradually increasing complexity to ensure comprehensive inspection coverage while controlling computational costs. The rule engine captures structural errors, statistical profiling identifies distributional deviations between data content and similar resources, and semantic comparison verifies semantic consistency between metadata descriptions and data content in boundary cases. These three complementary layers cover the complete inspection dimensions from format correctness to semantic consistency.
[0007] Second, this invention introduces a progressive trust assessment mechanism, which converts the supplier's historical quality assessment results into a reputation level and dynamically adjusts the inspection depth based on the reputation level. This allows suppliers with good quality performance to obtain more efficient inspection channels, while suppliers with poor quality performance or newly registered suppliers receive more comprehensive inspections. This achieves differentiated allocation of inspection resources and improves audit efficiency while ensuring the overall data quality and credibility.
[0008] Third, this invention enables the system to autonomously accumulate rules and knowledge during operation by means of three continuous evolutionary closed loops: rejection feedback feedback, terminology expansion feedback, and automatic updating of reference groups. This allows the rejection reason correlation analysis to generate new rule candidates, unidentified terms confirmed by humans to be automatically included in the synonym mapping table, and newly approved resource profiles to automatically enrich the reference groups. This gradually reduces the reliance on human intervention. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a flowchart illustrating the data quality assessment method based on semantic analysis provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the data quality assessment method based on semantic analysis provided in this embodiment of the invention. Figure 3 This is a structural diagram of the data quality assessment system based on semantic analysis provided in this embodiment of the invention; The diagram shows: Request reception and sampling module 11, reputation assessment and strategy determination module 12, multi-level progressive assessment module 13, and score fusion and report generation module 14. Detailed Implementation
[0011] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0012] Example 1, as Figure 1 , Figure 2 As shown, this invention provides a data quality assessment method based on semantic analysis, the method comprising: S1: Receive the listing request for the target data resource and determine the metadata description information and data content sample of the target data resource; Step S1 provided in this embodiment of the invention includes: Receive the listing request and extract the resource identifier and supplier identifier from the listing request; Based on the resource identifier, obtain the metadata description information of the target data resource, wherein the metadata description information includes at least the resource name, resource description, industry classification, data type, data source, and field declaration information; Based on the resource identifier, data sampling is performed on the target data resource to obtain the data content sample.
[0013] The specific implementation method is as follows: Receive listing requests for target data resources. These requests are submitted by the data provider through the platform's front-end interface, sent as an HTTP POST request to the request receiving interface of the quality inspection microservice. Extract JSON-formatted request data from the request body of this HTTP POST request, and parse it to obtain the resource identifier and provider identifier. The resource identifier is a unique code for the target data resource within the platform, and the provider identifier is a unique identification code for the party submitting the data resource.
[0014] Based on the resource identifier, the metadata description information of the target data resource is retrieved from the platform's metadata management database. The metadata management database is the platform's core storage component; each data resource generates a metadata record during registration and cataloging, indexed using the resource identifier as the primary key. The metadata description information returned by the query includes at least the resource name, resource description, industry classification, data type, data source, and field declaration information. The resource name is the provider's name for the data resource; the resource description is a free-text description of the data resource's coverage, collection method, and usage scenarios; the industry classification is the label for the industry sector to which the data resource belongs; the data type is the data structure type label for the data resource; the data source is a description of the original collection channel for the data resource; and the field declaration information is a structured list of all field names and their data types contained in the data resource.
[0015] For example, the resource name of a target data resource is "2024 Inpatient Medical Record Homepage Data of Tertiary Hospitals in Hubei Province", the resource description is "covering the inpatient medical record homepage data of 52 tertiary hospitals in Hubei Province throughout 2024, including fields such as diagnosis, surgery, and cost, totaling approximately 1.2 million records", the industry classification is "medical and health", the data type is "structured tabular data", the data source is "exported from the medical record management system of various hospitals", and the field declaration information includes fields such as medical record number, admission date, discharge date, main diagnosis code, and main surgery code.
[0016] Based on the resource identifier, data sampling is performed on the target data resource to obtain data content samples. The specific data sampling method is as follows: The storage path of the actual data file of the target data resource is obtained based on the resource identifier, and records are read from the data file using a data sampling strategy. The data sampling strategy involves reading sample data from the first N records of the data file, where N is a preset sampling quantity. The sampling quantity is dynamically determined based on the total number of records in the target data resource. When the total number of records is greater than a preset record count threshold, the sampling quantity is taken as the preset upper limit; when the total number of records is less than or equal to the preset record count threshold, the sampling quantity is directly taken as the total number of records.
[0017] The record count threshold is determined as follows: the record count threshold is determined based on the distribution of record counts of historical data resources that have been approved in the platform, and the upper limit of the sampling quantity is determined based on the sample sufficiency conditions required for calculating field-level indicators in the statistical profile. Specifically, the total number of records of all data resources that have been approved in the platform in the past quarter is obtained, sorted by the total number of records from smallest to largest, and the median is taken as the record count threshold.
[0018] The upper limit for the sampling quantity is calculated as follows: For each field, the relative error between the number of unique values for that field in the sample and the number of unique values for that field in the full dataset should not exceed a preset error tolerance at the selected confidence level. The error tolerance is set at 0.10, i.e., a 10% relative error range; the confidence level is set at 0.95, ensuring that at a 95% confidence level, the deviation between the number of unique values for each field in the sample and the total number of unique values does not exceed 10%. This guarantees the statistical reliability of the profiling indicators while keeping the sampling quantity within a reasonable load range for platform database read / write operations. The upper limit for the sampling quantity is calculated based on the average field unique value ratio of the platform's historical data resources and the aforementioned error tolerance and confidence level.
[0019] For example, the record count threshold is set to 500,000, and the maximum sampling quantity is set to 1,000. When the total number of records exceeds 500,000, the first 1,000 records are taken as data content samples; when the total number of records does not exceed 500,000, all records are taken as data content samples. For example, if the total number of records in the target data resource is 1.2 million, which exceeds the record count threshold of 500,000, the first 1,000 records are taken as samples.
[0020] The following technical effects were achieved through this step: Resource identifiers and supplier identifiers are automatically extracted from listing requests. Based on the resource identifiers, metadata description information is retrieved and data content samples are collected, unifying information scattered across three sources—requests, databases, and data files—as input data for subsequent evaluation steps. Dynamically determining the sample size ensures sufficient data for statistical profiling while avoiding the full reading of large-scale data resources, achieving a balance between evaluation efficiency and sample representativeness.
[0021] S2: Obtain the historical quality assessment records of the supplier corresponding to the target data resource, determine the supplier's credit rating based on the historical quality assessment records, and determine the quality inspection strategy based on the credit rating; Step S2 provided in this embodiment of the invention includes: Based on the supplier identifier, query the corresponding historical quality assessment records, wherein the historical quality assessment records contain the rating levels corresponding to each quality assessment; Based on the preset score change corresponding to the rating level of each quality assessment in the historical quality assessment record, the current credit score is obtained by weighting the cumulative score with a preset time decay coefficient and adding it to the preset base credit score. The supplier's credit rating is determined based on the preset credit score range in which the current credit score falls. The quality inspection strategy is determined based on the preset correspondence between the reputation level and the quality inspection strategy. In the preset correspondence, different reputation levels correspond to different quality inspection strategies. The quality inspection strategy is composed of a combination of evaluation levels, and the combination of evaluation levels is a subset selected from each level of the multi-level progressive evaluation.
[0022] The specific implementation method is as follows: Based on the supplier identifier, the platform's historical quality assessment database is used to query the corresponding historical quality assessment records for that supplier. The supplier identifier was extracted from the listing request in step S1. The historical quality assessment database uses the supplier identifier as the index key to store the rating level in the quality assessment report generated by the platform each time the supplier lists resources. Each historical quality assessment record includes the rating level and timestamp of that assessment, and the rating levels include four types: A, B, C, and D.
[0023] If no historical quality assessment records are found, meaning the supplier is a newly registered supplier and has never submitted a listing application before, the supplier's current credit score is directly set as the preset base credit score. Newly registered suppliers lack available historical quality data, so the base credit score is used as the initial trust value, and is dynamically adjusted based on actual performance in subsequent quality assessments. For example, the preset base credit score is 60 points, corresponding to a bronze reputation level, and the newly registered supplier undergoes all three levels of progressive assessment.
[0024] When historical quality assessment records are retrieved, the current credit score is calculated by weighting the changes in scores for each assessment level using a preset time decay coefficient and then adding this weighted average to a preset base credit score. The formula for calculating the current credit score is as follows: , Indicates the current credit score. This represents the basic credit score. This represents the change in the preset score corresponding to the rating level in the i-th quality assessment. The time decay coefficient is denoted by n, which represents the total number of times the supplier has listed the product in history. The basic credit score is the initial credit score for a newly registered supplier, set at 60 points. The preset score change is determined based on the rating level: +5 for level A, +2 for level B, -10 for level C, and -25 for level D. The time decay coefficient is a positive number less than 1, ensuring that recent assessments have a greater weight on the current credit score than earlier assessments. The value of the time decay coefficient ensures that the further back in time the assessment, the smaller the contribution of the score change to the current credit score, reflecting the time dynamics of supplier quality performance. The time decay coefficient is set between 0.9 and 0.99, ensuring that recent assessments have a significantly higher weight than earlier assessments, keeping the credit score sensitive to changes in supplier quality performance without decaying too rapidly to mean that only the most recent one or two assessments affect the credit score, thus guaranteeing a certain degree of stability when the number of supplier assessments is low. For example, a time decay coefficient of 0.95 is used.
[0025] S3: According to the quality inspection strategy, perform multi-level progressive evaluation on the metadata description information and the data content sample to obtain multi-level evaluation results. The multi-level progressive evaluation includes at least the following in sequence: structured matching inspection based on preset rules, profile comparison based on statistical features, and semantic similarity comparison based on semantic models. Step S3 provided in this embodiment of the invention includes: S3.1: Structured matching check based on preset rules, including: Load preset structured rules, wherein the structured rules include at least value range rules, existence rules, enumeration matching rules, field combination rules, and metadata declaration consistency rules; The structured rules are executed one by one to match and check the metadata description information and the data content sample, and a rule hit list is generated. Determine whether there are any error-level hits in the rule hit list; If an error level is hit, the execution of subsequent levels in the multi-level progressive evaluation is terminated, the rule check result is output, and the process proceeds to step S4. If no error level is hit, the rule check result is output, and subsequent evaluation levels are performed.
[0026] The specific implementation method is as follows: Load the pre-defined structured rules. These rules are pre-stored in the platform's rule base, organized in JSON format. Each rule includes a rule number, rule type, rule description, severity level triggered upon a match, and penalty value. The structured rules in the rule base include at least five types: value range rules, existence rules, enumeration matching rules, field combination rules, and metadata declaration consistency rules.
[0027] Value range rules define the reasonable range of values for a single field. For example, for a field labeled as a percentage in its field declaration information, its value should be within a closed interval of 0 to 100. When executing a value range rule, all values for the field in the data content sample are iterated through, and each value is checked against the defined range. If any value exceeds the range, a value range rule hit record is generated.
[0028] Existence rules define the existence and null value rate limits of fields under specific conditions. For example, when the industry classification in the metadata description is "Medical and Health" and the data type is "Inpatient Data," the primary diagnosis code field must exist and the null value rate must not exceed 5%. When executing an existence rule, the system first checks whether the industry classification and data type in the metadata description meet the rule's triggering conditions. If they do, it further checks whether the field exists in the data content sample, i.e., whether the field name is included in the field declaration information. If it exists, the system calculates the number of null values for that field in the data content sample and divides it by the total number of records in the sample to obtain the null value rate. If the field does not exist or the null value rate exceeds 5%, an existence rule hit record is generated.
[0029] Enumeration matching rules define that a field value must belong to a certain standard enumeration set. For example, the value of the gender field must belong to the enumeration set consisting of "male" and "female". When executing an enumeration matching rule, all non-empty values of the field in the data content sample are retrieved, and each value is checked one by one to see if it belongs to the enumeration set. If there is a value that does not belong to the enumeration set, an enumeration matching rule hit record is generated.
[0030] Field combination rules define logical constraints between multiple fields. For example, when the province field value is "Hubei Province," the city field value must belong to the set of prefecture-level city names under Hubei Province. When executing a field combination rule, each record in the data sample is iterated. When the province field value of a record is "Hubei Province," it checks whether the city field value is in the set of prefecture-level city names under Hubei Province. If there is a record that does not meet the requirement, a record matching the field combination rule is generated.
[0031] Metadata declaration consistency rules define consistency constraints between the structural information of metadata declarations and the actual data structure. For example, the number of fields included in the field declaration information in the metadata description should be consistent with the actual number of fields included in the data content sample. When executing a metadata declaration consistency rule, the field declaration information is parsed from the metadata description, the field list is extracted, and the number of fields is counted; the actual field name list is extracted from the data content sample, and the number of fields is counted. If the two counts are inconsistent, a metadata declaration consistency rule hit record is generated.
[0032] After executing all structured rules one by one, all rule hit records are summarized to generate a rule hit list. The rule hit list includes the rule number, rule type, hit fields, severity level, and rule description for each hit record.
[0033] Determine if there are any error-level hits in the rule hit list. Each structured rule has a predefined severity level, including error and warning levels. An error level hit indicates a hard error, meaning the data resource has a serious problem that cannot be corrected through subsequent evaluation. If an error-level hit exists, output the rule check result and terminate the execution of subsequent levels in the multi-level progressive evaluation, directly proceeding to step S4 for scoring fusion; if no error-level hits exist, including if the rule hit list only contains warning-level hits or the rule hit list is empty, output the rule check result and continue with the profile comparison in S3.2.
[0034] S3.2: Profile comparison based on statistical features, including: Perform statistical feature extraction on the data content sample to obtain the current data profile, which includes field-level profile indicators and dataset-level profile indicators. Based on the industry classification and data type in the metadata description information, a reference data set is selected from the historical data resources that have passed the review, and a reference profile is constructed based on the reference data set; Traverse the corresponding statistical indicators in the current data profile and the reference profile, perform relative position calculations, and obtain the percentile ranking of each statistical indicator; Based on the preset quantile anomaly threshold, determine whether there are any quantile anomaly indicators in the quantile ranking; If the quantile anomaly index exists and the number of the quantile anomaly index is greater than or equal to the preset anomaly number threshold, then the execution of subsequent levels in the multi-level progressive evaluation is terminated, the profile comparison result is output, and the process proceeds to step S4. Otherwise, output the image comparison results and continue with the subsequent evaluation level.
[0035] The specific implementation method is as follows: Statistical feature extraction is performed on the data content sample obtained in step S1 to obtain the current data profile. The data content sample contains several records and several fields, with each record representing a combination of values for each field. Statistical feature extraction does not perform semantic understanding of the data values; instead, it extracts objective data features through statistical aggregation.
[0036] Field-level profiling metrics are extracted separately for each field. When a field is numeric, the extracted field-level profiling metrics include total number of records, number of non-null values, number of null values, null value rate, number of unique values, minimum value, maximum value, mean, and standard deviation. The mean is the arithmetic mean of all non-null values for that field, and the standard deviation is the square root of the dispersion of each value relative to the mean. When a field is text, the extracted field-level profiling metrics include total number of records, number of non-null values, number of null values, null value rate, number of unique values, average length, maximum length, and number of high-frequency values. When a field is date, the extracted field-level profiling metrics include total number of records, number of non-null values, null value rate, date range, and date range span.
[0037] Dataset-level profiling metrics are extracted from the entire data sample, including the total number of fields, the total number of records, the overall null value rate, and the date span. The overall null value rate is the ratio of the number of null values in all fields to the total number of records multiplied by the total number of fields.
[0038] Based on the industry classification and data type in the metadata description information, a reference data set is selected from the approved historical data resources. Approved historical data resources are those with an evaluation result of A or B in the platform, stored in the platform's reference group library. Each resource includes its industry classification, data type tag, and a data profile generated upon approval. The selection criteria for the reference data set are: same industry classification and same data type. When the number of historical data resources in the reference data set is less than 5, the selection criteria are automatically relaxed: first, the same data type is relaxed to similar data types, then the same industry classification is relaxed to similar industry classifications, until the number of reference data sets is no less than 5. Similar data types and similar industry classifications are determined according to a preset classification hierarchy structure.
[0039] A reference profile is constructed based on a reference dataset. The reference profile includes statistics for both field-level and dataset-level profile metrics, identical to those in the current data profile. For each specific metric, the value of each historical data resource for that metric is extracted from the reference dataset, and the median of the set of values is calculated as the reference value for that metric in the reference profile. When the number of elements in the set of values is even, the median is the arithmetic mean of the two middle elements.
[0040] The algorithm iterates through the corresponding statistical indicators in the current data profile and the reference profile, calculating their relative positions to obtain the quantile ranking of each indicator. For each indicator, the indicator value of the current data profile and the indicator values of each historical data resource in the reference data set are placed in the same sequence and sorted from smallest to largest. Each element in the sequence corresponds to a ranking position. The formula for calculating the quantile ranking is: quantile ranking = rank position of the current indicator value in the sequence ÷ total number of elements in the sequence. The quantile ranking value ranges from 0 to 1, where 0 indicates that the current indicator value is lower than all historical data resources in the reference set, 1 indicates that it is higher than all historical data resources, and 0.5 indicates that it is at the middle level of the reference set.
[0041] For example, the null value rate of the hospitalization days field in the current data profile is 0.05. The reference data set contains 12 historical data resources. After sorting the null value rates from smallest to largest, the current value ranks 3rd. The percentile ranking = 3 ÷ 12 = 0.25.
[0042] Based on preset quantile anomaly thresholds, the system determines whether any quantile anomaly indicators exist in the quantile rankings. The criteria for determining the quantile anomaly threshold vary depending on the nature of the indicator: for indicators where a larger value is better, such as coding compliance rate, a quantile anomaly threshold of less than 0.25 is considered anomaly; for indicators where a smaller value is better, such as null value rate, a quantile anomaly threshold of more than 0.75 is considered anomaly; for indicators where a moderate value is good, such as the number of records, a quantile anomaly threshold of less than 0.05 or more than 0.95 is considered anomaly. These anomaly thresholds are taken from boundary quantile values (0.25 and 0.75, with extreme values of 0.05 and 0.95). This selection ensures that the screening ratio of anomaly indicators is approximately 25% of the extreme range of the indicator value distribution in the reference group, providing robust statistical identification capabilities in large samples. It effectively detects indicators with significant quality deviations while controlling false alarms caused by fluctuations in the reference group.
[0043] After traversing all indicators, the total number of indicators identified as quantile outliers is counted. If quantile outliers exist and their number is greater than or equal to a preset outlier threshold, the profile comparison result is output, and the execution of subsequent levels in the multi-level progressive evaluation is terminated, directly proceeding to step S4; otherwise, the profile comparison result is output, and the semantic similarity comparison in S3.3 continues. The outlier threshold is determined based on the total number of indicators in the data profile and the platform's tolerance for quality deviations, exemplarily set at 3. The tolerance is defined as a systemic data quality deviation, rather than an occasional fluctuation of a single indicator, when a data sample deviates from similar reference resources simultaneously in at least 3 indicators.
[0044] S3.3: Semantic similarity comparison based on semantic models, including: Convert the feature information in the current data profile into content description text; Convert the content of each field in the metadata description information into metadata description text; The content description text and the metadata description text are respectively processed by word segmentation and stop word filtering, and the word segmentation results are normalized based on a preset thesaurus mapping table. Based on a pre-built semantic model, the normalized word segmentation results are vector-encoded to obtain content semantic vectors and metadata semantic vectors. The semantic model is a word vector model pre-trained based on a preset domain corpus. Calculate the cosine similarity between the content semantic vector and the metadata semantic vector to obtain the semantic comparison result.
[0045] The specific implementation method is as follows: The feature information in the current data profile is converted into content description text. The content description text presents a summary of the statistical characteristics of the data content sample in the form of natural language statements. The content description text is constructed by reading dataset-level profile indicators from the current data profile and generating templated description statements. For example, the content description text might be: "The dataset contains 5 fields and a total of 1000 records. The overall null value rate is 0.03. The fields cover common structures in the medical field, including patient identifier, admission date, discharge date, and diagnosis code." The content of each field in the metadata description information is converted into metadata description text. The metadata description text is a concatenation of the content of each field in the metadata description information in the form of natural language statements. The metadata description text is constructed as follows: the resource name, resource description, industry classification, data type, data source, and field declaration information are read sequentially from the metadata description information obtained in step S1, and then concatenated into a coherent natural language statement.
[0046] For example, the metadata description text is: "Inpatient medical record homepage data of tertiary hospitals in Hubei Province in 2024, covering inpatient medical record homepage data of 52 tertiary hospitals in Hubei Province throughout 2024, including fields such as diagnosis, surgery, and cost, totaling approximately 1.2 million records. The industry classification is medical and health, the data type is structured tabular data, and the data source is exported from the medical record management systems of various hospitals. The field declaration information includes fields such as medical record number, admission date, discharge date, primary diagnosis code, and primary surgery code." Perform word segmentation and stop word filtering on the content description text and the metadata description text respectively. Word segmentation is completed by a Chinese word segmentation tool, which divides the input text into a word sequence according to semantic units. Stop word filtering is performed based on a preset stop word list. The stop word list includes function words and punctuation marks that have no semantic contribution, such as "de (of)", "le (past tense marker)", "deng (etc.)", "gongji (altogether)", and so on. Words included in the stop word list are matched and removed one by one from the word sequence after word segmentation to obtain a filtered effective word sequence.
[0047] Based on a preset synonym mapping table, normalization processing is performed on the word segmentation result. The synonym mapping table is stored in the synonym mapping database of the platform, and each mapping record includes two fields: a source term and a standard term. During normalization processing, the word sequence after word segmentation is traversed. For each word, it is queried whether there is a matching source term in the synonym mapping table. If the matching is successful, the word is replaced with the corresponding standard term; if the matching is not successful, the original word is retained. For example, "binganhao (medical record number)" in the word segmentation result is replaced with "zhuyuanhao (hospitalization number)" after matching through the synonym mapping table.
[0048] Based on a pre-constructed semantic model, vector encoding is performed on the normalized word segmentation result. The semantic model is a word vector model pre-trained based on preset domain corpora, which maps each word in natural language text to a fixed-dimensional dense floating-point vector. During vector encoding, the word sequence of the normalized content description text and the word sequence of the metadata description text are sent to the semantic model respectively. For each text, the arithmetic mean of the word vectors of all words in the corresponding dimension is taken to obtain a text-level semantic vector with the same dimension as the word vector. The content semantic vector is obtained after vector encoding of the content description text, and the metadata semantic vector is obtained after vector encoding of the metadata description text.
[0049] Calculate the cosine similarity between the content semantic vector and the metadata semantic vector to obtain a semantic comparison result. The cosine similarity is equal to the inner product of two vectors divided by the product of the norms of the two vectors. The value range of cosine similarity is from -1 to 1. A value closer to 1 indicates that the texts represented by the two semantic vectors are semantically closer. For example, if the cosine similarity between the content semantic vector and the metadata semantic vector is 0.85, it indicates that the metadata description and the data content are highly consistent in semantics.
[0050] Through this step, the following technical effects are obtained: The system employs a three-tiered evaluation strategy, prioritizing deterministic rules and delegating probabilistic models: the rule engine captures format and structural errors and immediately blocks severely non-compliant resources; profile comparison is triggered only after rule checks are passed; and semantic comparison is performed only after the profile determines there are no systemic anomalies. These three layers are activated sequentially only when necessary, controlling the computational overhead of the semantic model while ensuring comprehensive check coverage. Profile comparison introduces historical resources of the same type within the same industry to construct a dynamic reference group, replacing fixed thresholds with quantile rankings. This, combined with differentiated anomaly detection based on indicator properties and an automatic relaxation strategy when the reference group is insufficient, allows the anomaly identification criteria to adaptively update with platform data accumulation.
[0051] S4: Based on the quality inspection strategy, the obtained inspection results are integrated and scored to generate a quality assessment report of the target data resource.
[0052] Step S4 provided in this embodiment of the invention includes: Based on the anomalies in the multi-level evaluation results, calculate the deduction value for each level according to the preset deduction rules corresponding to each level; The comprehensive quality score is obtained by deducting the deduction values at each level from the preset full score. The quality level is determined based on the preset score range in which the comprehensive quality score falls; Based on the comprehensive quality score, the quality level, and the multi-level evaluation results, a quality assessment report containing evaluation details and improvement suggestions is generated.
[0053] The specific implementation method is as follows: Based on the anomalies in the multi-level assessment results, the deduction values for each level are calculated according to the preset deduction rules corresponding to each level. The first-level deduction value is calculated based on the severity level of each hit record in the rule hit list. A rule hit with a severity level of "error" deducts 15 points per hit, and a rule hit with a severity level of "warning" deducts 5 points per hit. The deduction values of all hit records in the rule hit list are summed to obtain the first-level deduction value. If there are no errors in S3.1, and only warning hits exist, the first-level deduction value is the number of warning hits multiplied by 5 points. If there are neither errors nor warnings, the first-level deduction value is 0. The deduction value of 15 points for the error level is three times the deduction value of 5 points for the warning level, using a deduction ladder to distinguish between structural errors and warning deviations. A 15-point deduction corresponding to the error level means that even if only two structural errors occur, resulting in a deduction of 30 points, the resource score drops from the maximum of 100 points to 70 points, falling into the B-level range, making a single structural error sufficient to attract the attention of the approver.
[0054] The second-level deduction is calculated based on the number of abnormal percentile indicators in the profile comparison results. The deduction value for a single abnormal percentile indicator is related to the degree of deviation of that indicator's percentile ranking; the more severe the deviation, the higher the deduction. The deduction range for a single abnormal percentile indicator is 1 to 10 points. The calculation method for the deduction value of a single abnormal percentile indicator is as follows: For indicators where a larger value is better, when the percentile ranking is lower than the abnormal threshold, the deduction value = 10 × (abnormal threshold - percentile ranking) ÷ abnormal threshold. For indicators where a smaller value is better, when the percentile ranking is higher than the abnormal threshold, the deduction value = 10 × (percentile ranking - abnormal threshold) ÷ (1 - abnormal threshold).
[0055] For example, if a certain indicator where a smaller value is better has a percentile ranking of 0.85 and an anomaly threshold of 0.75, then the deduction value for this indicator = 10 × (0.85 - 0.75) ÷ (1 - 0.75) = 4 points. The deduction values for all percentile abnormal indicators are summed to obtain the second-level deduction value. If there are no percentile abnormal indicators in S3.2, or if the number of abnormalities does not reach the abnormality threshold and therefore the subsequent evaluation is not terminated, then the second-level deduction value is 0.
[0056] The third-level deduction value is calculated based on the cosine similarity score in the semantic comparison results. A cosine similarity score greater than or equal to 0.8 indicates semantic consistency, and the third-level deduction value is 0. A cosine similarity score greater than or equal to 0.6 but less than 0.8 indicates slight deviation, and the third-level deduction value is 10 × (0.8 - cosine similarity score) ÷ 0.2. A cosine similarity score less than 0.6 indicates serious deviation, and the third-level deduction value is 10 + 10 × (0.6 - cosine similarity score) ÷ 0.6.
[0057] For example, when the cosine similarity score is 0.85, which is greater than 0.8, the third-level deduction value is 0. When the cosine similarity score is 0.72, it is a slight deviation, and the deduction value is 10 × (0.8 - 0.72) ÷ 0.2 = 4 points. When the cosine similarity score is 0.5, it is a serious deviation, and the deduction value is 10 + 10 × (0.6 - 0.5) ÷ 0.6 ≈ 11.67 points.
[0058] The overall quality score is calculated by subtracting deductions from the preset maximum score at each level. The preset maximum score is 100 points. If a certain level of evaluation does not generate a deduction because it is not triggered (e.g., the third level is not executed according to the quality inspection strategy), then the deduction for that level is 0. Overall Quality Score = 100 - First Level Deduction - Second Level Deduction - Third Level Deduction. When the calculated overall quality score is negative, it is corrected to 0.
[0059] The quality level is determined based on the preset score range where the overall quality score falls. The correspondence between the preset score range and the quality level is as follows: an overall quality score greater than or equal to 85 is Level A; an overall quality score greater than or equal to 70 and less than 85 is Level B; an overall quality score greater than or equal to 50 and less than 70 is Level C; and an overall quality score less than 50 is Level D. Level A indicates that the metadata description and data content are highly consistent; Level B indicates that they are basically consistent but have slight deviations; Level C indicates that there are obvious inconsistencies that require rejection and correction; and Level D indicates that there are serious inconsistencies that trigger a security alarm.
[0060] For example, after a target data resource undergoes a three-layer evaluation, the first layer deduction is 0: the rule hit list only contains 2 warning hits, but the termination condition is not met to continue execution to the next layer. The second layer deduction is 4: one percentile abnormal indicator deviates, with a percentile ranking of 0.85 deducting 4 points. The third layer deduction is 4: the cosine similarity is 0.72 deducting 4 points. The overall quality score = 100 - 0 - 4 - 4 = 92 points, which falls within the range of 85 points and above, and the quality level is A.
[0061] Based on the comprehensive quality score, quality level, and multi-level evaluation results, a quality assessment report is generated, including assessment details and improvement suggestions. The assessment report includes the following sections: A header recording the resource name, supplier information, supplier reputation level, assessment time, and comprehensive quality score and level; First-level results listing detailed information on rule hits, with all rules marked as passed if none were hit; Second-level results listing the core indicator values of the current data profile, the median of the reference profile, percentile ranking, and anomaly judgment conclusion; Third-level results, displayed only when triggered, including semantic similarity score, judgment conclusion, and keyword matching details; and Improvement suggestions summarizing all anomalies from each level, sorted by deduction value from highest to lowest, with specific correction suggestions for each. The report is presented in a standardized format for approvers to view in the approval work order interface.
[0062] The following technical effects were achieved through this step: The scoring rules at different levels are correlated with the severity and degree of deviation of anomalies, enabling the comprehensive quality score to reflect the overall quality level of data resources across three dimensions: rule compliance, statistical distribution rationality, and semantic consistency. The assessment report integrates scores, grades, details at each level, and improvement suggestions into a standardized review document. Approvers can gain a comprehensive understanding of the data resource quality and make approval decisions without having to review the original data at each level individually.
[0063] Step S3 provided in this embodiment of the invention further includes: Once the target data resource passes the quality assessment and is successfully uploaded, the current data profile is saved as the profile baseline. If the data content of the target data resource is updated, statistical feature extraction is performed on the updated data to obtain an incremental data profile. The statistical indicators in the incremental data profile are traversed, and the change range is calculated by combining them with the corresponding statistical indicators in the profile baseline. Determine whether there are any statistical indicators in the calculation results of the change range that exceed the preset alarm threshold; If it exists, a quality change alarm will be generated.
[0064] The specific implementation method is as follows: Once the target data resource passes quality assessment and is uploaded, the current data profile is saved as the profile baseline. The current data profile, extracted in step S3.2, includes field-level profile metrics and dataset-level profile metrics. The current data profile is stored in the platform's profile baseline library using the resource identifier as the primary key, serving as a reference benchmark for subsequent change detection of this data resource. The profile baseline contains information at two levels: profile metric values for each field (i.e., null value rate, number of unique values, mean, standard deviation, etc.) and profile metric values for the overall dataset (i.e., number of records, overall null value rate, etc.).
[0065] When the data content of the target data resource is updated, statistical feature extraction is performed on the updated data to obtain an incremental data profile. Data updates are triggered when the supplier submits new data files or appends records through the platform's data update interface. Data updates are detected by comparing the timestamp or version number of the data resource's data file; an update is determined when the timestamp or version number differs from the value recorded at the time of upload. Statistical feature extraction is performed on the updated data, using the same categories of indicators as in step S3.2 for the current data profile, resulting in an incremental data profile.
[0066] The statistical indicators in the incremental data profile are iterated through, and the magnitude of change is calculated by combining them with the corresponding statistical indicators in the profile baseline. The formula for calculating the magnitude of change is: Magnitude of Change = |New Value of Indicator - Baseline Value| ÷ (Baseline Value + ε). Here, the new value of the indicator is the value of that indicator in the incremental data profile, the baseline value is the value of that indicator in the profile baseline, and ε is a very small positive number. The purpose of ε is to prevent the denominator from being 0 when the baseline value is 0, which would lead to an undefined division; it is set to 0.001. The magnitude of change is expressed as a percentage, ranging from 0 to positive infinity; the larger the value, the more significant the change in the indicator relative to the baseline.
[0067] For example, if the null value rate in the baseline profile of a certain data resource that has been uploaded is 0.03, and the null value rate in the updated incremental data profile is 0.08, then the change in this indicator is approximately 1.61, which means that the null value rate of this field changes by about 161% relative to the baseline.
[0068] This function determines whether any statistical indicators in the calculated change range exceed a preset alarm threshold. The preset alarm thresholds are set according to the indicator type. For the null value rate, the alarm threshold is a relative change exceeding 50% or an absolute increase in the null value rate exceeding 10 percentage points. For the number of records, the alarm threshold is a change exceeding 30% with no corresponding new data declaration in the metadata description. For the coding compliance rate, the alarm threshold is a decrease exceeding 15 percentage points. For high-frequency values in field distribution, the alarm threshold is a newly appearing value that has not appeared in the baseline high-frequency value statistics entering the top five high-frequency values. The rationale for these alarm thresholds is as follows: a doubling or 10-percentage-point increase in the null value rate signifies a clear degradation in data availability; a fluctuation of more than 30% in the number of records indicates an unexpected expansion or reduction in data coverage; a 15-percentage-point decrease in the coding compliance rate signifies a significant deterioration in data standardization; and changes in high-frequency values indicate a structural change in the composition of core data content.
[0069] When the change exceeds the alarm threshold, a quality change alarm is generated. The quality change alarm is output to the platform's approval ticket system, triggering a manual review process. The alarm content includes the name of the triggering metric, the baseline value, the updated value, the change magnitude, and the alarm level. The alarm level is determined comprehensively based on the change magnitude and metric type; an abnormally high null value rate and a decrease in coding compliance rate trigger a severe alarm, while fluctuations in the number of records and changes in high-frequency values trigger a general alarm. When the change magnitude does not exceed the alarm threshold, the incremental data profile is updated to a new profile baseline, ensuring the baseline continuously tracks the latest status of data resources.
[0070] After step S4 generates the quality assessment report, the main assessment process concludes. However, for a data quality management system, the end of one assessment is also the starting point for the next round of evolutionary iterations. The quality assessment results of the target data resources contain information that has optimization value for the system's own rules, knowledge, and benchmarks. Furthermore, each assessment result needs to be included in the supplier's historical quality assessment records as the basis for calculating reputation levels when making subsequent upload requests.
[0071] Therefore, this step performs three closed-loop evolutionary feedback and supplier record update operations after the evaluation report is generated, including: Once a new data resource passes a quality assessment, its profile information will be incorporated into the historical data resources accordingly. When there are unidentified terms in the semantic comparison results, and the data quality is confirmed to be correct by manual verification, the unidentified terms and their corresponding standard correspondences are added to the preset synonym mapping table. When the quality assessment result of the target data resource is unsuccessful, the reason for rejection is recorded, and new rule candidates are generated based on the correlation between the rejection reason obtained by manual analysis and the statistical characteristics of the data content sample. Update the supplier's historical quality assessment records based on the quality assessment results of the target data resources.
[0072] The specific implementation method is as follows: When the quality assessment result of the target data resource is Grade A or B, meaning it has passed the quality assessment, the profile information of that data resource is included in the historical data resources. The profile information, extracted in step S3.2, includes field-level profile indicators and dataset-level profile indicators. A new record is added to the platform's reference group library, using the resource identifier as the primary key, storing the industry classification, data type label, and complete profile information of the data resource. This record is included in the filtering scope of the reference data set when performing profile comparisons on other resources in S3.2, serving as the basic data for constructing the reference profile.
[0073] Simultaneously, historical data resources in the reference group library are subject to timeliness management. Historical data resources whose entry time exceeds the preset reference validity period are removed from the active reference group or have their weight in the quantile ranking calculation reduced. The reference validity period is determined based on the platform's data resource update cycle and the stability of industry data characteristics, ensuring that the reference group continuously reflects the distribution characteristics of recent data quality levels. For example, the reference validity period is 24 months; historical data resources older than 24 months are no longer included in the construction of the reference profile.
[0074] When unidentified terms are found in the semantic comparison results, and the data quality is manually confirmed to be correct, the unidentified terms and their corresponding standard correspondences are added to a pre-defined synonym mapping table. The mechanism for discovering unidentified terms is as follows: During semantic similarity comparison in step S3.3, each word in the segmented word sequence is queried in the synonym mapping table. If the query result is non-existent and the word is an actual term with domain semantics, it is marked as an unidentified term. When the resource is finally determined to pass the evaluation, the approver will discover and manually confirm the actual meaning of these unidentified terms during the approval process. If the approver determines that the unidentified term is indeed a synonym of a standard term, such as recognizing "medical record homepage" as a synonym of "inpatient medical record homepage," the pair of terms and their correspondences are added to the synonym mapping table. After addition, when subsequent resources undergo semantic similarity comparison, encountering "medical record homepage" will be automatically normalized to "inpatient medical record homepage," improving the recall and accuracy of semantic matching.
[0075] When the quality assessment result of the target data resource is Grade C or D, i.e., it fails, the reason for rejection is recorded. The reasons for rejection are selected by the approver from a pre-set list of standard rejection reasons during the rejection process, and supplementary explanation text is added. The pre-set list of standard rejection reasons covers rejection reason categories such as first-level rule hit, second-level profile comparison anomaly, and third-level semantic comparison inconsistency. The rejection reasons are associated with the profile features of the data content samples in this assessment, with the association indexed by the resource identifier. The values of the rejection reason items and each profile indicator are also recorded.
[0076] The rejection reason is associated with the profile features of the data content sample in this assessment and stored, with the association indexed by the resource identifier. The rejection reason entry and the values of each profile indicator are also recorded. When manually analyzing the rejection reason, approvers retrieve all historical rejection records associated with that reason and their corresponding profile feature data, identifying profile indicators and value ranges significantly related to the rejection reason. Based on this association, the identification result is transformed into a new rule candidate, including rule conditions and rule conclusions. The new rule candidate is submitted to operations personnel for review. After confirmation, it is added to the rule base as the basis for subsequent first-level structured rule checks or second-level profile comparisons.
[0077] Based on the quality assessment results of the target data resources, update the supplier's historical quality assessment records. Add a new record to the platform's historical quality assessment database, using the supplier identifier obtained in step S1 as the primary key, recording the rating level and assessment timestamp for this quality assessment. This record will be queried in step S2 the next time the supplier submits a listing request and used to calculate the current credit score and determine the reputation level.
[0078] Example 2, as Figure 3 As shown, based on the same inventive concept provided in Embodiment 1, this embodiment of the invention also provides a data quality assessment system based on semantic analysis, the system comprising: The request receiving and sampling module 11 is used to receive the listing request of the target data resource and determine the metadata description information and data content sample of the target data resource; The reputation assessment and strategy determination module 12 is used to obtain the historical quality assessment records of the supplier corresponding to the target data resource, determine the supplier's reputation level based on the historical quality assessment records, and determine the quality inspection strategy based on the reputation level. The multi-level progressive evaluation module 13 is used to perform multi-level progressive evaluation on the metadata description information and the data content sample according to the quality inspection strategy to obtain multi-level evaluation results. The multi-level progressive evaluation includes at least structured matching inspection based on preset rules, portrait comparison based on statistical features, and semantic similarity comparison based on semantic models. The scoring fusion and report generation module 14 is used to perform fusion scoring on the obtained inspection results according to the quality inspection strategy, and generate a quality assessment report of the target data resource.
[0079] In one embodiment, the request receiving and sampling module 11 is further configured to receive the listing request and extract the resource identifier and supplier identifier from the listing request; Based on the resource identifier, obtain the metadata description information of the target data resource, wherein the metadata description information includes at least the resource name, resource description, industry classification, data type, data source, and field declaration information; Based on the resource identifier, data sampling is performed on the target data resource to obtain the data content sample.
[0080] In one embodiment, the reputation assessment and strategy determination module 12 is used to query the corresponding historical quality assessment record based on the supplier identifier, wherein the historical quality assessment record includes the rating level corresponding to each quality assessment. Based on the preset score change corresponding to the rating level of each quality assessment in the historical quality assessment record, the current credit score is obtained by weighting the cumulative score with a preset time decay coefficient and adding it to the preset base credit score. The supplier's credit rating is determined based on the preset credit score range in which the current credit score falls. The quality inspection strategy is determined based on the preset correspondence between the reputation level and the quality inspection strategy. In the preset correspondence, different reputation levels correspond to different quality inspection strategies. The quality inspection strategy is composed of a combination of evaluation levels, and the combination of evaluation levels is a subset selected from each level of the multi-level progressive evaluation.
[0081] In one embodiment, the multi-level progressive evaluation module 13 is further used for structured matching checks based on preset rules, including: Load preset structured rules, wherein the structured rules include at least value range rules, existence rules, enumeration matching rules, field combination rules, and metadata declaration consistency rules; The structured rules are executed one by one to match and check the metadata description information and the data content sample, and a rule hit list is generated. Determine whether there are any error-level hits in the rule hit list; If an error level is hit, the execution of subsequent levels in the multi-level progressive evaluation is terminated, the rule check result is output, and the process proceeds to step S4. If no error level is hit, the rule check result is output, and subsequent evaluation levels are performed.
[0082] Profile comparison based on statistical features includes: Perform statistical feature extraction on the data content sample to obtain the current data profile, which includes field-level profile indicators and dataset-level profile indicators. Based on the industry classification and data type in the metadata description information, a reference data set is selected from the historical data resources that have passed the review, and a reference profile is constructed based on the reference data set; Traverse the corresponding statistical indicators in the current data profile and the reference profile, perform relative position calculations, and obtain the percentile ranking of each statistical indicator; Based on the preset quantile anomaly threshold, determine whether there are any quantile anomaly indicators in the quantile ranking; If the quantile anomaly index exists and the number of the quantile anomaly index is greater than or equal to the preset anomaly number threshold, then the execution of subsequent levels in the multi-level progressive evaluation is terminated, the profile comparison result is output, and the process proceeds to step S4. Otherwise, output the image comparison results and continue with the subsequent evaluation level.
[0083] Semantic similarity comparison based on semantic models includes: Convert the feature information in the current data profile into content description text; Convert the content of each field in the metadata description information into metadata description text; The content description text and the metadata description text are respectively processed by word segmentation and stop word filtering, and the word segmentation results are normalized based on a preset thesaurus mapping table. Based on a pre-built semantic model, the normalized word segmentation results are vector-encoded to obtain content semantic vectors and metadata semantic vectors. The semantic model is a word vector model pre-trained based on a preset domain corpus. Calculate the cosine similarity between the content semantic vector and the metadata semantic vector to obtain the semantic comparison result.
[0084] Semantic similarity comparison based on semantic models also includes: Once the target data resource passes the quality assessment and is successfully uploaded, the current data profile is saved as the profile baseline. If the data content of the target data resource is updated, statistical feature extraction is performed on the updated data to obtain an incremental data profile. The statistical indicators in the incremental data profile are traversed, and the change range is calculated by combining them with the corresponding statistical indicators in the profile baseline. Determine whether there are any statistical indicators in the calculation results of the change range that exceed the preset alarm threshold; If it exists, a quality change alarm will be generated.
[0085] In one embodiment, the scoring fusion and report generation module 14 is further configured to calculate the deduction value of each level according to the preset deduction rules corresponding to each level based on the anomalies in the multi-level evaluation results. The comprehensive quality score is obtained by deducting the deduction values at each level from the preset full score. The quality level is determined based on the preset score range in which the comprehensive quality score falls; Based on the comprehensive quality score, the quality level, and the multi-level evaluation results, a quality assessment report containing evaluation details and improvement suggestions is generated.
Claims
1. A data quality assessment method based on semantic analysis, characterized in that, include: S1: Receive the listing request for the target data resource and determine the metadata description information and data content sample of the target data resource; S2: Obtain the historical quality assessment records of the supplier corresponding to the target data resource, determine the supplier's credit rating based on the historical quality assessment records, and determine the quality inspection strategy based on the credit rating; S3: According to the quality inspection strategy, perform multi-level progressive evaluation on the metadata description information and the data content sample to obtain multi-level evaluation results. The multi-level progressive evaluation includes at least the following in sequence: structured matching inspection based on preset rules, profile comparison based on statistical features, and semantic similarity comparison based on semantic models. S4: Based on the quality inspection strategy, the obtained inspection results are integrated and scored to generate a quality assessment report of the target data resource.
2. The data quality assessment method based on semantic analysis as described in claim 1, characterized in that, Step S1 includes: Receive the listing request and extract the resource identifier and supplier identifier from the listing request; Based on the resource identifier, obtain the metadata description information of the target data resource, wherein the metadata description information includes at least the resource name, resource description, industry classification, data type, data source, and field declaration information; Based on the resource identifier, data sampling is performed on the target data resource to obtain the data content sample.
3. The data quality assessment method based on semantic analysis as described in claim 1, characterized in that, Step S2 includes: Based on the supplier identifier, query the corresponding historical quality assessment records, wherein the historical quality assessment records contain the rating levels corresponding to each quality assessment; Based on the preset score change corresponding to the rating level of each quality assessment in the historical quality assessment record, the current credit score is obtained by weighting the cumulative score with a preset time decay coefficient and adding it to the preset base credit score. The supplier's credit rating is determined based on the preset credit score range in which the current credit score falls. The quality inspection strategy is determined based on the preset correspondence between the reputation level and the quality inspection strategy. In the preset correspondence, different reputation levels correspond to different quality inspection strategies. The quality inspection strategy is composed of a combination of evaluation levels, and the combination of evaluation levels is a subset selected from each level of the multi-level progressive evaluation.
4. The data quality assessment method based on semantic analysis as described in claim 1, characterized in that, Step S3, the structured matching check based on preset rules, includes: Load preset structured rules, wherein the structured rules include at least value range rules, existence rules, enumeration matching rules, field combination rules, and metadata declaration consistency rules; The structured rules are executed one by one to match and check the metadata description information and the data content sample, and a rule hit list is generated. Determine whether there are any error-level hits in the rule hit list; If an error level is hit, the execution of subsequent levels in the multi-level progressive evaluation is terminated, the rule check result is output, and the process proceeds to step S4. If no error level is hit, the rule check result is output, and subsequent evaluation levels are performed.
5. The data quality assessment method based on semantic analysis as described in claim 4, characterized in that, In step S3, the profile comparison based on statistical features includes: Perform statistical feature extraction on the data content sample to obtain the current data profile, which includes field-level profile indicators and dataset-level profile indicators. Based on the industry classification and data type in the metadata description information, a reference data set is selected from the historical data resources that have passed the review, and a reference profile is constructed based on the reference data set; Traverse the corresponding statistical indicators in the current data profile and the reference profile, perform relative position calculations, and obtain the percentile ranking of each statistical indicator; Based on the preset quantile anomaly threshold, determine whether there are any quantile anomaly indicators in the quantile ranking; If the quantile anomaly index exists and the number of the quantile anomaly index is greater than or equal to the preset anomaly number threshold, then the execution of subsequent levels in the multi-level progressive evaluation is terminated, the profile comparison result is output, and the process proceeds to step S4. Otherwise, output the image comparison results and continue with the subsequent evaluation level.
6. The data quality assessment method based on semantic analysis as described in claim 5, characterized in that, In step S3, the semantic similarity comparison based on the semantic model includes: Convert the feature information in the current data profile into content description text; Convert the content of each field in the metadata description information into metadata description text; The content description text and the metadata description text are respectively processed by word segmentation and stop word filtering, and the word segmentation results are normalized based on a preset thesaurus mapping table. Based on a pre-built semantic model, the normalized word segmentation results are vector-encoded to obtain content semantic vectors and metadata semantic vectors. The semantic model is a word vector model pre-trained based on a preset domain corpus. Calculate the cosine similarity between the content semantic vector and the metadata semantic vector to obtain the semantic comparison result.
7. The data quality assessment method based on semantic analysis as described in claim 1, characterized in that, Step S4: Based on the quality inspection strategy, the obtained inspection results are integrated and scored to generate a quality assessment report for the target data resource, including: Based on the anomalies in the multi-level evaluation results, calculate the deduction value for each level according to the preset deduction rules corresponding to each level; The comprehensive quality score is obtained by deducting the deduction values at each level from the preset full score. The quality level is determined based on the preset score range in which the comprehensive quality score falls; Based on the comprehensive quality score, the quality level, and the multi-level evaluation results, a quality assessment report containing evaluation details and improvement suggestions is generated.
8. The data quality assessment method based on semantic analysis as described in claim 5, characterized in that, Semantic similarity comparison based on semantic models also includes: Once the target data resource passes the quality assessment and is successfully uploaded, the current data profile is saved as the profile baseline. If the data content of the target data resource is updated, statistical feature extraction is performed on the updated data to obtain an incremental data profile. The statistical indicators in the incremental data profile are traversed, and the change range is calculated by combining them with the corresponding statistical indicators in the profile baseline. Determine whether there are any statistical indicators in the calculation results of the change range that exceed the preset alarm threshold; If it exists, a quality change alarm will be generated.
9. The data quality assessment method based on semantic analysis as described in claim 1, characterized in that, Also includes: Once a new data resource passes a quality assessment, its profile information will be incorporated into the historical data resources accordingly. When there are unidentified terms in the semantic comparison results, and the data quality is confirmed to be correct by manual verification, the unidentified terms and their corresponding standard correspondences are added to the preset synonym mapping table. When the quality assessment result of the target data resource is unsuccessful, the reason for rejection is recorded, and a new rule candidate is generated based on the correlation between the reason for rejection obtained by manual analysis and the statistical characteristics of the data content sample. Update the supplier's historical quality assessment records based on the quality assessment results of the target data resources.
10. A data quality assessment system based on semantic analysis, characterized in that, The system for implementing the data quality assessment method based on semantic analysis according to any one of claims 1 to 9, the system comprising: The request receiving and sampling module is used to receive the listing request of the target data resource and determine the metadata description information and data content sample of the target data resource; The reputation assessment and strategy determination module is used to obtain the historical quality assessment records of the supplier corresponding to the target data resource, determine the supplier's reputation level based on the historical quality assessment records, and determine the quality inspection strategy based on the reputation level. The multi-level progressive evaluation module is used to perform multi-level progressive evaluation on the metadata description information and the data content sample according to the quality inspection strategy to obtain multi-level evaluation results. The multi-level progressive evaluation includes at least the following in sequence: structured matching inspection based on preset rules, profile comparison based on statistical features, and semantic similarity comparison based on semantic models. The scoring fusion and report generation module is used to fuse and score the obtained inspection results according to the quality inspection strategy, and generate a quality assessment report of the target data resource.