Multi-dimensional data quality evaluation method based on machine learning and industry rule base

Through a multi-dimensional data quality evaluation method based on machine learning and industry rule base, the challenges of data quality management in the medical and health industry are solved, and the accuracy, flexibility and real-timeness of data quality evaluation are achieved.

CN120013345APending Publication Date: 2025-05-16DONGHUA MEDICAL TECH CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510102308.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The medical and health industry faces challenges in data quality management, including a surge in data volume, complex business processes, numerous data tables, complex data correlation and inconsistent standards, resulting in inconsistent data quality and affecting the reliability of data and the accuracy of analysis results.

Method used

A multi-dimensional data quality evaluation method based on machine learning and industry rule database is adopted. By obtaining medical and health data to be evaluated, storing it in a classified manner, and selecting corresponding quality control rules from the industry rule database based on business categories and usage scenarios, the data is calculated in quality scores, and the quality score of each dimension indicator is generated, and the total quality score is calculated by weighted average method. Finally, the data set is constructed and the quality control evaluation model is input to generate a quality control evaluation report.

Benefits of technology

This method ensures the relevance and adaptability of quality control rules, improves the pertinence and effectiveness of evaluation, and improves the accuracy and flexibility of data quality evaluation through automated and intelligent processing methods, and can adaptively adjust according to changes in business characteristics to ensure the real-time and dynamic evaluation of evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013345A_ABST
    Figure CN120013345A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical and health data quality management, in particular to a multi-dimensional data quality evaluation method based on machine learning and an industry rule base, and the method comprises the steps: obtaining to-be-evaluated medical and health data; constructing a knowledge base; selecting a corresponding quality control rule from the industry rule base; calculating a quality score of each dimension index; calculating a total quality score; and constructing a data set based on source basic information included in the to-be-evaluated medical and health data, the business category to which the to-be-evaluated medical and health data belongs, the selected quality control rule, the quality score of each dimension index and the total quality score, inputting the data set into the quality control evaluation model, and outputting a quality control evaluation report for the data set. The technology of combining the industry rule base, the quality control evaluation knowledge base and the quality control evaluation model is adopted, overall evaluation of the data quality is automatically carried out, the correlation and adaptability of the quality control rule are ensured, and the automation, intelligence and accuracy of evaluation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical and health data quality management, and in particular to a multi-dimensional data quality evaluation method based on machine learning and an industry rule base. Background Art

[0002] In the healthcare industry, the importance of data quality management is becoming increasingly prominent, especially in the face of challenges such as surging data volume, complex business processes, numerous data tables, complex data associations, and inconsistent standards. The industry rule library plays a core role in this process. It not only contains industry standards, regulatory requirements, and best practices, but also provides a standardized and systematic framework for data quality evaluation. By selecting and customizing quality control rules from the industry rule library, the healthcare industry can ensure the accuracy and consistency of data quality evaluation, which is crucial for cross-departmental, cross-system, and even cross-border medical data sharing and analysis.

[0003] With the rapid development of medical equipment and information technology, the explosive growth of medical data has put forward higher requirements for the circulation and sharing of data. The medical business process involves multiple departments and professionals, and each link may generate a large amount of data. The accurate and timely circulation of this data is crucial to the quality and efficiency of medical services. In addition, the number of data tables in the medical information system is huge, and each data table may contain tens of thousands of records. Managing these data tables and ensuring the relevance and integrity of the data is a huge challenge.

[0004] In addition, the complex relationships between medical data, such as patient information and medical records, drug use and side effects, etc., are crucial to the correctness and accuracy of medical data analysis and application. The standards and formats of medical data vary between different medical institutions, regions and even countries, which leads to data interoperability issues and affects data sharing and analysis. At the same time, the inconsistency of data quality, due to the lack of unified data quality standards and monitoring mechanisms, makes the quality of medical data from different sources and types uneven, affecting the reliability of the data and the accuracy of the analysis results.

[0005] In order to effectively address these challenges, this application proposes a multi-dimensional data quality evaluation method based on machine learning and industry rule base. Summary of the invention

[0006] Based on this, it is necessary to provide a multi-dimensional data quality evaluation method based on machine learning and industry rule base to address the above technical problems.

[0007] According to a first aspect of the present invention, a multi-dimensional data quality evaluation method based on machine learning and an industry rule library is provided, S1, obtaining medical and health data to be evaluated, and classifying and storing the medical and health data to be evaluated according to predefined business categories; wherein the medical and health data to be evaluated include business data and source basic information, and the data type of the business data is structured two-dimensional table data; S2, based on the medical and health data to be evaluated, according to the business category, user object and actual use scenario to which the medical and health data to be evaluated belongs, selecting corresponding quality control rules from the industry rule library, each of the quality control rules including null value rate dimension indicators, integrity Dimension indicators, continuity dimension indicators, rationality dimension indicators and correlation dimension indicators; S3. Use quality control rules to calculate the quality score of the medical and health data to be evaluated, and generate the quality score of each dimensional indicator; S4. Use the weighted average method to calculate the quality score of each dimensional indicator, and calculate and obtain the total quality score; S5. Based on the source basic information contained in the medical and health data to be evaluated, the business category of the medical and health data to be evaluated, the selected quality control rules, the quality score of each dimensional indicator and the total quality score, construct a data set, input the data set into the quality control evaluation model, and output the quality control evaluation report for the data set.

[0008] Optionally, the step S1 includes: obtaining the medical and health data to be evaluated including business data and basic source information, wherein the basic source information includes the institution information to which the business data belongs, creation time information and business occurrence time information; and classifying and storing the medical and health data to be evaluated according to predefined business categories.

[0009] Optionally, before step S2, it also includes: using industry standards, regulatory requirements and industry practical experience as quality control rules, and classifying and storing the quality control rules according to predefined business categories and usage scenarios to build an industry rule library.

[0010] Optionally, in step S2, the null value rate dimension indicator refers to the ratio of the total number of fields with empty data fields in the business data to the total data volume within the accounting period and the accounting scope. The calculation formula of the null value rate dimension indicator is as follows: ; Where: Indicates the quality score of the null value rate dimension indicator, data field i represents the i-th data field, data field n represents the n-th data field, and n represents the total number of data fields; the integrity dimension indicator refers to randomly selecting a business data from a class of business data with an associated relationship within the accounting period and accounting scope, and randomly extracting sampled data from the business data, and calculating the proportion of data in the sampled data that meets the associated conditions in the total sampled data through the associated field. The calculation formula of the integrity dimension indicator is as follows: ; Where: Indicates the quality score of the completeness dimension indicator, business data j indicates the jth business data, business data m indicates the mth business data, and m indicates the total number of business data; the continuity dimension indicator refers to the ratio of the change in the amount of uploaded data within two adjacent accounting cycles and within the accounting scope. The calculation formula of the continuity dimension indicator is as follows: ; Where: It represents the quality score of the continuity dimension indicator, business data k represents the kth business data, business data p represents the pth business data, and p represents the total number of business data; the rationality dimension indicator refers to the data indicators of hospitals at different levels required by the state within the accounting period and accounting scope. The values ​​of the corresponding data indicators are counted from the business data, and the comparison is made to see whether the values ​​meet the level requirements of the hospital. The proportion of data that meet the comparison conditions in the accounting rules accounts for the total rules. The calculation formula of the rationality dimension indicator is as follows: ; Where: It represents the quality score of the rationality dimension indicator, rule h represents the hth rule, rule q represents the qth rule, q represents the total number of rules, and the result of the rule meeting the comparison condition is 0 or 1, where 0 represents the result that the rule does not meet the comparison condition, and 1 represents the result that the rule meets the comparison condition; the correlation dimension indicator refers to selecting sample data from a group of business data with a master-slave relationship within the accounting period and accounting scope, and calculating the proportion of the data volume in the sample data that meets the master-slave relationship in the total sample data volume through the association field. The calculation formula of the correlation dimension indicator is as follows: ; Where: It represents the quality score of the correlation dimension indicator, business data g represents the g-th piece of business data, business data w represents the w-th piece of business data, and w represents the total number of business data.

[0011] Optionally, the step S3 includes: according to the target accounting cycle, using quality control rules to calculate the quality score of the medical and health data to be evaluated, and generating a quality score for each dimensional indicator; wherein the target accounting cycle includes a daily accounting cycle, a weekly accounting cycle, a monthly accounting cycle, a quarterly accounting cycle, and an annual accounting cycle.

[0012] Optionally, in step S4, the total quality score is calculated using the following formula: ; Where: Represents the overall quality score.

[0013] Optionally, before step S5, it also includes: obtaining quality control evaluation results generated by historical manual evaluation, wherein the quality control evaluation results include multiple historical data sets and historical quality control evaluation reports related to the historical data sets; using a semantic vector model to vectorize all historical data sets and historical quality control evaluation reports related to the historical data sets to obtain multiple historical data set embedding vectors and historical quality control evaluation report embedding vectors related to the historical data sets, and constructing a quality control evaluation knowledge base based on the historical data set embedding vectors and the historical quality control evaluation report embedding vectors related to the historical data sets; constructing an initial quality control evaluation model based on a random forest algorithm; splitting the historical data set embedding vectors and the historical quality control evaluation report embedding vectors related to the historical data sets into a training data set and a test data set, training and optimizing the initial quality control evaluation model according to the training data set and the test data set, and outputting the trained quality control evaluation model.

[0014] Optionally, the step S5 includes: constructing a data set based on the basic information of the source contained in the medical and health data to be evaluated, the business category to which the medical and health data to be evaluated belongs, the selected quality control rules, the quality score of each dimensional indicator and the total quality score, and using a semantic vector model to vectorize the data set to obtain a data set embedding vector; inputting the data set embedding vector into the trained quality control evaluation model, outputting an embedding vector of a historical quality control evaluation report for the data set, and obtaining a historical quality control evaluation report based on the historical quality control evaluation report embedding vector, wherein the historical quality control evaluation report includes analysis result information, scoring result information, rating result information, conclusion result information, risk warning information and action suggestion information.

[0015] According to a second aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method when executing the computer program.

[0016] According to a third aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of the method when executed by a processor.

[0017] The advantages and beneficial effects of the present invention are as follows: the multi-dimensional data quality evaluation method based on machine learning and industry rule base provided by the present invention, by constructing an industry rule base, can select matching quality control rules from the industry rule base containing industry standards, regulatory requirements and industry practical experience according to specific business categories and actual usage scenarios. This method not only ensures the relevance and adaptability of the quality control rules, but also makes the quality control evaluation report more accurately meet the needs of the required industry, thereby improving the pertinence and effectiveness of the evaluation; subsequently, a machine learning algorithm is used to learn and generate a quality control evaluation knowledge base based on the quality control evaluation results previously generated based on manual assessment, to automatically perform an overall evaluation of the data quality, and provide an evaluation report and optimization suggestions. This automated and intelligent processing method not only improves the accuracy and flexibility of the evaluation, but can also adaptively adjust according to changes in business characteristics, thereby ensuring the real-time and dynamic nature of the evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is the overall flow chart of the present invention.

[0019] Figure 2 A schematic diagram of an electronic device. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific implementation methods in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0021] Embodiment 1

[0022] Figure 1 This is a flowchart of a multi-dimensional data quality evaluation method based on machine learning and industry rule base provided in Example 1 of the present invention.

[0023] Reference Figure 1 , the method comprises the following steps.

[0024] S1. Obtain the medical and health data to be evaluated, and classify and store the medical and health data according to predefined business categories; wherein the medical and health data to be evaluated includes business data and source basic information, and the data type of the business data is two-dimensional table data in structured storage.

[0025] In this embodiment, the stored medical and health data to be evaluated is obtained from the business system, and the medical and health data to be evaluated includes business data and source basic information, wherein the source basic information includes the institution information to which the business data belongs, the creation time information and the business occurrence time information.

[0026] In this embodiment, in the medical and health industry, medical and health data collection is a complex process that involves integrating residents' health records throughout their life cycle from different IT systems, institutions or regions. Medical and health data collection is a continuous process, and it is necessary to ensure that the collected data has complete source basic information, which may include the institution information, system information, creation time information, business occurrence time information, etc. to which the data belongs, in order to ensure the traceability of the data. At the same time, in the case of a large amount of collected data, a batch collection method based on a predefined time period can be adopted, which can improve the efficiency of data collection on the one hand, and reduce the impact on the business system on the other hand.

[0027] In this embodiment, the business data involved in this application refers to data generated and used in the field of medical and health care, including but not limited to patients' medical records, diagnostic information, treatment data, etc. The data type of the business data is two-dimensional table data stored in a structured manner and does not include unstructured data. At the same time, medical and health data is classified and stored according to predefined business categories.

[0028] S2. Based on the medical and health data to be evaluated, corresponding quality control rules are selected from the industry rule library according to the business category, user objects and actual usage scenarios of the medical and health data to be evaluated. Each of the quality control rules includes a null value rate dimension indicator, a completeness dimension indicator, a continuity dimension indicator, a rationality dimension indicator and a correlation dimension indicator.

[0029] In this embodiment, before step S2, it also includes: using industry standards, regulatory requirements and industry practical experience as quality control rules, and classifying and storing the quality control rules according to predefined business categories and usage scenarios to build an industry rule library.

[0030] In this embodiment, matching quality control rules are selected from the industry rule library according to the business category and user objects to which the medical and health data belongs, combined with the actual usage scenarios of the medical and health data. Each quality control rule includes a null value rate dimension indicator, an integrity dimension indicator, a continuity dimension indicator, a rationality dimension indicator, and a correlation dimension indicator. Among them, the quality scoring rules for the null value rate dimension indicator and the rationality dimension indicator are for specific business fields, the quality scoring rules for the integrity dimension indicator are for a class of business tables, the quality scoring rules for the continuity dimension indicator are for specific business tables, and the quality scoring rules for the correlation dimension indicator are for a group of business tables.

[0031] Taking the healthcare industry as an example, after determining the business categories and users of business data, we conduct an in-depth analysis of the needs of business data in specific usage scenarios. Through surveys and interviews with different user groups such as doctors, nurses, and hospital managers, we collect their specific needs for data quality in scenarios such as query, analysis, and report generation. Subsequently, we retrieve relevant quality control rules from the industry rule library, which include industry standards, regulatory requirements, and industry practice experience (i.e., industry best practices).

[0032] In order to ensure the consistency and shareability of medical and health data, it is necessary to unify data standards and perform data conversion. In this process, the determination of quality control rules is crucial and needs to be formulated based on the business characteristics, degree of informatization and industry experience of the data source. According to the "Health Information Data Element Standardization Rules" (WS / T 303-2023), the standardization of medical and health information data elements is the key to achieving data unification and interoperability. This includes data element models, attributes, naming, definitions, classifications, and content standard writing format specifications. These standardization rules help ensure the consistency and accuracy of data across the industry, thereby improving data quality.

[0033] In addition, the establishment and management of residents' electronic health records also need to follow certain standards. According to the "Basic Content of the Home Page of Residents' Electronic Health Records (Trial)", residents' health records should contain personal health identification, personal basic health information, health service activity records, etc. These contents should be dynamically extracted and updated based on unified technical requirements. Based on the above information, data quality management in the medical and health industry needs to rely on a multi-dimensional data quality score calculation method, which includes but is not limited to the null value rate, completeness, continuity, rationality and relevance of the data.

[0034] At the same time, quality control rules can be classified according to industry experience and business characteristics. For example, a set of quality control rules for the homepage of health records, a set of quality control rules for hypertension management, etc. can be flexibly and dynamically classified and stored according to actual business categories and usage scenarios, so that quality control rules can be reused in different business categories and usage scenarios to reduce labor costs.

[0035] Furthermore, the null value rate dimension indicator refers to the ratio of the total number of empty data fields in the business data to the total data volume within the accounting period and accounting scope, expressed as a percentage. The quality score of the null value rate dimension indicator can retain two decimal places; if the business data is not uploaded, the score is 100%. The null value rate dimension indicator mainly uses the presence or absence of data field values ​​as a standard for measuring the quality of the value, so as to verify the richness of the collected business data.

[0036] Taking the health information browser application as an example, its null value rate dimension indicator is based on the data display of the health information browser system. It counts the ratio of the number of empty fields in 20 basic medical business fields, including outpatient prescriptions, outpatient diagnoses, outpatient medical records, test records, examination reports, inpatient medical record homepage, discharge records, first medical records, preoperative summaries, senior physician rounds records, admission registrations, admission records, general surgical records, inpatient medical orders, discharge records, rescue records, stage summaries, preoperative discussions and daily medical records, to the total number.

[0037] Furthermore, the calculation cycle of the void rate dimension indicator includes daily accounting cycle, weekly accounting cycle, monthly accounting cycle, quarterly accounting cycle, and annual accounting cycle. The accounting scope is the business data of the institution. The calculation formula of the void rate dimension indicator is as follows: ; Where: It represents the quality score of the null value rate dimension indicator. Data field i represents the i-th data field, data field n represents the n-th data field, and n represents the total number of data fields.

[0038] Furthermore, the integrity dimension indicator refers to randomly selecting a piece of business data from a category of business data with associated relationships within the accounting period and accounting scope, and randomly extracting sampled data from the business data, and calculating the proportion of data in the sampled data that meets the associated conditions through the associated fields to the total sampled data volume, expressed as a percentage. The quality score of the integrity dimension indicator can be retained to two decimal places; if the business data is not uploaded, the score is 0%.

[0039] Since residents' health and medical information is stored in multiple business tables in a scattered manner, the scattered business tables can be summarized through associated fields. In the process of business data collection, interface developers need to summarize the various business tables in accordance with the requirements of the provincial platform to ensure that the uploaded residents' health and medical information is complete and there is no missing business table. Therefore, a way to verify the integrity of the business data is needed.

[0040] Taking daily outpatient and inpatient services as an example, based on the two most basic outpatient and inpatient data tables of "admission registration information" and "outpatient and emergency registration records", 100 "visit serial numbers" are randomly selected within the accounting period, and statistics are conducted on whether there are relevant business data in the "outpatient and emergency (including observation) medical record records", "outpatient and emergency diagnosis details records", "outpatient and emergency prescription records", "Western medicine inpatient medical record homepage", "admission record form", "discharge summary record", "first medical course record", "laboratory test information", and "inpatient medical order details information" data tables, so as to reflect the integrity of the uploaded business data.

[0041] Furthermore, the calculation cycle of the integrity dimension indicators includes daily accounting cycle, weekly accounting cycle, monthly accounting cycle, quarterly accounting cycle, and annual accounting cycle. The accounting scope is the business data of the institution. The calculation formula of the integrity dimension indicators is as follows: ; Where: It represents the quality score of the completeness dimension indicator, business data j represents the j-th business data, business data m represents the m-th business data, and m represents the total number of business data.

[0042] Furthermore, the continuity dimension indicator refers to the proportion of the change in the amount of uploaded data within two adjacent accounting cycles and within the accounting scope, expressed as a percentage. The quality score of the continuity dimension indicator can be retained to two decimal places; if no business data is uploaded, the score is 100%.

[0043] Furthermore, the continuous change of data volume can be reflected by calculating the ratio of data volume change in two adjacent calculation cycles. For example, in the outpatient registration table, the amount of new data on the 23rd was 2345, and the amount of new data on the 24th was 2568. The ratio of data change on the 24th was 9.51%, indicating that the continuous change of data was small and the new data was relatively stable.

[0044] Furthermore, the calculation cycle of the continuity dimension indicators includes daily accounting cycle, weekly accounting cycle, monthly accounting cycle, quarterly accounting cycle, and annual accounting cycle. The accounting scope is the business data of the institution. The calculation formula of the continuity dimension indicators is as follows: ; Where: It represents the quality score of the continuity dimension indicator, business data k represents the k-th business data, business data p represents the p-th business data, and p represents the total number of business data.

[0045] Furthermore, the rationality dimension indicator refers to the data indicators of hospitals at different levels required by the state within the accounting cycle and accounting scope. The corresponding data indicators are counted from the business data to compare whether the values ​​meet the hospital's level requirements. The accounting rules are used to calculate the proportion of data that meet the comparison conditions in the total rules, expressed as a percentage. The quality score of the rationality dimension indicator can be retained to two decimal places; if the business data is not uploaded, the score is 0%.

[0046] Furthermore, the calculation of rationality dimension indicators is done in a statistical way, using the national data indicators for hospitals at different levels as a yardstick to verify the statistical results of the indicators with actual data. Taking the national data indicators for Class-A tertiary hospitals as an example, the rules of data indicators include but are not limited to: the consistency rate between admission diagnosis and discharge diagnosis ≥95%, the consistency rate before and after surgery ≥90%, the consistency rate between the main clinical diagnosis and pathology ≥50%, the positive rate of X-ray computed tomography (CT) examination ≥60% (for those without CT, this item does not account for points), the positive rate of magnetic resonance imaging (MRI) examination ≥70% (for those without MRI, this item does not account for points), the positive rate of large X-ray machine examination ≥50%, the X-ray nail film rate ≥40%, the average passing rate of clinical chemistry inter-laboratory quality assessment throughout the year (VIS≤120), the average passing rate of hematology inter-laboratory quality assessment throughout the year (modified deviation index DI≤2), the average score of immunology inter-laboratory quality assessment throughout the year is above the national average score, etc. In the process of calculating the quality score of the rationality dimension indicators, the values ​​of the above data indicators are counted based on the uploaded business data to compare whether they meet the hospital's grade requirements, so as to reflect the quality of the uploaded data.

[0047] Furthermore, the calculation cycle of the rationality dimension indicators includes daily accounting cycle, weekly accounting cycle, monthly accounting cycle, quarterly accounting cycle, and annual accounting cycle. The accounting scope is the business data of the institution. The calculation formula of the rationality dimension indicators is as follows: ; Where: It represents the quality score of the rationality dimension indicator, rule h represents the hth rule, rule q represents the qth rule, q represents the total number of rules, and the result of the rule meeting the comparison condition is 0 or 1, where 0 represents the result that the rule does not meet the comparison condition, and 1 represents the result that the rule meets the comparison condition.

[0048] Furthermore, the correlation dimension indicator refers to selecting sampled data from a group of business data with a master-slave relationship within the accounting period and accounting scope, and calculating the proportion of data in the sampled data that meets the master-slave relationship to the total sampled data through the associated fields, expressed as a percentage. The quality score of the correlation dimension indicator can be retained to two decimal places; if the business data is not uploaded, the score is 0%.

[0049] Furthermore, the association involved in the present application, namely the master-slave association, is the association requirement that appears between the master and slave data tables. For example, for the data structure of the inspection result which has an inspection report header and inspection details, in order to reduce the redundancy of data storage, it will be split into a master table and a slave table for storage, and the master table and the slave table are associated using associated fields.

[0050] Furthermore, the calculation cycle of the correlation dimension indicators includes daily calculation cycle, weekly calculation cycle, monthly calculation cycle, quarterly calculation cycle, and annual calculation cycle. The calculation scope is the business data of the institution. The calculation formula of the correlation dimension indicators is as follows: ; Where: It represents the quality score of the correlation dimension indicator, business data g represents the g-th piece of business data, business data w represents the w-th piece of business data, and w represents the total number of business data.

[0051] S3. Use quality control rules to calculate the quality score of the medical and health data to be evaluated, and generate a quality score for each dimensional indicator.

[0052] In this embodiment, step S3 includes: according to the target accounting cycle, using quality control rules to calculate the quality score of the medical and health data to be evaluated, and generating a quality score for each dimensional indicator; wherein the target accounting cycle includes a daily accounting cycle, a weekly accounting cycle, a monthly accounting cycle, a quarterly accounting cycle, and an annual accounting cycle.

[0053] Furthermore, the minimum accounting cycle is day, and the maximum cycle is year. It is divided into five levels: day, week, month, quarter, and year. In the process of calculating the quality score of the medical and health data to be evaluated, the above formula can be used for direct calculation. The difference is that the accounting cycle is different, and the business data and data volume collected are different.

[0054] S4. Use the weighted average method to calculate the quality score of each dimension indicator and obtain the total quality score.

[0055] In this embodiment, the calculation formula of the total quality score is as follows: ; Where: Represents the overall quality score.

[0056] S5. Based on the basic information of the source of the medical and health data to be evaluated, the business category to which the medical and health data to be evaluated belongs, the selected quality control rules, the quality score of each dimensional indicator and the total quality score, a data set is constructed, the data set is input into the quality control evaluation model, and a quality control evaluation report for the data set is output.

[0057] In this embodiment, before step S5, it also includes: obtaining quality control evaluation results generated by historical manual evaluation, wherein the quality control evaluation results include multiple historical data sets and historical quality control evaluation reports related to the historical data sets; using a semantic vector model to vectorize all historical data sets and historical quality control evaluation reports related to the historical data sets to obtain multiple historical data set embedding vectors and historical quality control evaluation report embedding vectors related to the historical data sets, and constructing a quality control evaluation knowledge base based on the historical data set embedding vectors and the historical quality control evaluation report embedding vectors related to the historical data sets; constructing an initial quality control evaluation model based on a random forest algorithm; splitting the historical data set embedding vectors and the historical quality control evaluation report embedding vectors related to the historical data sets into a training data set and a test data set, training and optimizing the initial quality control evaluation model according to the training data set and the test data set, and outputting the trained quality control evaluation model.

[0058] Further, the core technology of step S5 involved in this application includes: constructing a quality control evaluation knowledge base, and using the constructed quality control evaluation knowledge base to train and optimize the quality control evaluation model based on machine learning, and inputting the data set into the trained quality control evaluation model to output a quality control evaluation report. The quality control evaluation knowledge base is based on a text vectorization (TextEmbedding) model, such as a semantic vector model (Embedding Model), and the quality control rules that used to be generated by manual evaluation, the quality score of each dimension indicator and the total quality score, the basic information of the source, the quality control evaluation report, etc. are divided into blocks, words, representations, extracted and vectorized to form an embedded vector, and the processed embedded vector is stored using ElasticSearch to construct a quality control evaluation knowledge base, and the quality control evaluation model is constructed based on a random forest algorithm, and is generated after training and optimization using the quality control evaluation knowledge base, mainly a decision tree model for the logical relationship between input information and output information. The above method can make up for the one-sidedness and irrationality of evaluating data quality in conventional data quality control by relying solely on the execution results of quality control rules. Since data quality evaluation is a process that requires comprehensive analysis based on information from all aspects, especially for highly professional data such as medical and health data, simply giving the execution results of quality control rules cannot meet the requirements of data quality management work aimed at improving data quality.

[0059] Furthermore, the construction and training process of the quality control evaluation model includes the following steps.

[0060] Construct a quality control evaluation knowledge base for training, wherein the quality control evaluation results generated by historical manual evaluation are obtained, and the quality control evaluation results include input information and output information. The input information is a historical data set, and the historical data set includes historical source basic information, historical business categories, historical quality control rules, historical quality scores of indicators in various dimensions, and historical total quality scores. The output information is a historical quality control evaluation report related to the historical data set, and the historical quality control evaluation report includes parsing result information, scoring result information, rating result information, conclusion result information, risk warning information, and action suggestion information. The semantic vector model (Embedding Model) is used for block segmentation, word segmentation, representation extraction and processing, and the processed historical data set embedding vector (i.e., historical source basic information embedding vector, historical business category embedding vector, historical quality control rule embedding vector, historical quality score embedding vector of indicators in various dimensions, and historical total quality score embedding vector) and the historical quality control evaluation report embedding vector related to the historical data set (i.e., parsing result embedding vector, scoring result embedding vector, rating result embedding vector, conclusion result embedding vector, risk warning information embedding vector, and action suggestion information embedding vector) are stored using ElasticSearch to construct a quality control evaluation knowledge base.

[0061] Construct a quality control evaluation model and use the Random Forest algorithm to build an initial quality control evaluation model.

[0062] Train the quality control evaluation model, split the historical data set embedding vector and the historical quality control evaluation report embedding vector related to the historical data set into a training data set and a test data set according to a predefined ratio (such as 7:3), use the training data set and the test data set to train and optimize the initial quality control evaluation model, and output the trained quality control evaluation model, that is, through machine training, construct the semantic and logical judgment between the input information and the output information.

[0063] In this embodiment, step S5 includes: constructing a data set based on the basic information of the source contained in the medical and health data to be evaluated, the business category to which the medical and health data to be evaluated belongs, the selected quality control rules, the quality score of each dimensional indicator and the total quality score, and using a semantic vector model to vectorize the data set to obtain a data set embedding vector; inputting the data set embedding vector into the trained quality control evaluation model, outputting an embedding vector of a historical quality control evaluation report for the data set, and obtaining a historical quality control evaluation report based on the embedding vector of the historical quality control evaluation report.

[0064] Furthermore, the source basic information and business type of the business data obtained in step S1, the quality control rules selected in step S2, the quality scores of the indicators of each dimension calculated in step S3 and the total quality score calculated in step S4 are used as input information and input into the trained quality control evaluation model to output an embedding vector of the historical quality control evaluation report for the data set. At the same time, based on the embedding vector of the historical quality control evaluation report, the corresponding historical quality control evaluation report is obtained.

[0065] Based on the trained quality control evaluation model, the entire data set is evaluated as a whole to determine the level of data quality. A quality control evaluation report is automatically generated that includes a data quality score and detailed analysis of each dimension. This quality control evaluation report not only points out the potential causes of data quality problems, but also provides a series of improvement suggestions.

[0066] The quality control evaluation report mainly includes a data quality overview, providing overall data quality scores and grades, and a summary of key business indicators; action recommendations, based on data analysis, provide specific action points to improve data quality; such a process ensures the continuous improvement of data quality, supports data-based decision making, and improves business efficiency and effectiveness.

[0067] In addition, the system can continuously monitor the effects of data quality improvement measures and learn from them to optimize future analysis and evaluation processes. Through continuous quality control evaluation models, it can adapt to changes in data characteristics and improve the accuracy and relevance of evaluations.

[0068] Embodiment 2

[0069] This embodiment further provides an electronic device based on the above embodiment 1. Figure 2 , Figure 2 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0070] like Figure 2 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 to a random access memory (RAM) 303. Various programs and data required for the operation of the electronic device are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0071] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, etc., an output device 307 including, for example, a liquid crystal display (LCD), a speaker, etc., a storage device 308 including, for example, a magnetic tape, a hard disk, etc., and a communication device 309. The communication device 309 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 2 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 2 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0072] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0073] Embodiment 3

[0074] Based on the above-mentioned embodiment 1, this embodiment further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned method are implemented.

[0075] It should be noted that the computer-readable medium mentioned above in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0076] In this embodiment, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperTextTransferProtocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an adhoc peer-to-peer network), as well as any currently known or future developed network.

[0077] The computer-readable medium may be included in the device, or may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains training data, converts the training data into initial data; determines an initial rule base based on the initial data, and optimizes the parameters of the initial rule base to obtain a target rule base; calculates the rules in the target rule base according to a preset activation weight calculation formula to obtain activation weights; determines abnormal information according to the test data and the activation weights.

[0078] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0079] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0080] The units described in some embodiments of the present disclosure may be implemented by software or hardware. The units described may also be provided in a processor, for example, may be described as: a processor including a data acquisition unit, a rule determination unit, a weight calculation unit, and an anomaly determination unit. The names of these units do not, in some cases, constitute limitations on the units themselves, for example, the data acquisition unit may also be described as a "unit for acquiring training data".

[0081] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that may be used include, without limitation, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.

[0082] Obviously, it should be understood by those skilled in the art that the above-mentioned various steps of the present invention can be performed in a manner different from the present invention, and the simulation method and experimental equipment include but are not limited to the above description. The above-mentioned various steps of the present invention can be performed in an order different from that here in some cases, and the steps shown or described above can be performed separately. Therefore, the present invention is not limited to any specific combination of hardware and software.

[0083] The above contents are further detailed descriptions of the present invention in combination with specific implementation methods, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.

Claims

1. A multi-dimensional data quality evaluation method based on machine learning and industry rule base, characterized in that: include: S1. Obtain the medical and health data to be evaluated, and classify and store the medical and health data to be evaluated according to predefined business categories; wherein the medical and health data to be evaluated includes business data and source basic information, and the data type of the business data is two-dimensional table data in structured storage; S2. Based on the medical and health data to be evaluated, select corresponding quality control rules from the industry rule library according to the business category, user object and actual usage scenario of the medical and health data to be evaluated, each of which includes a null value rate dimension indicator, a completeness dimension indicator, a continuity dimension indicator, a rationality dimension indicator and a correlation dimension indicator; S3. Use quality control rules to calculate the quality score of the medical and health data to be evaluated, and generate the quality score of each dimension indicator; S4. Calculate the quality score of each dimension indicator using the weighted average method to obtain the total quality score; S5. Based on the basic information of the source of the medical and health data to be evaluated, the business category to which the medical and health data to be evaluated belongs, the selected quality control rules, the quality score of each dimensional indicator and the total quality score, a data set is constructed, the data set is input into the quality control evaluation model, and a quality control evaluation report for the data set is output.

2. The multi-dimensional data quality evaluation method based on machine learning and industry rule base according to claim 1 is characterized in that: The step S1 comprises: Acquire the medical and health data to be evaluated, including business data and basic source information, wherein the basic source information includes the institution information to which the business data belongs, the creation time information, and the business occurrence time information; The medical and health data to be evaluated are classified and stored according to predefined business categories.

3. The multi-dimensional data quality evaluation method based on machine learning and industry rule base according to claim 1 is characterized in that: Before step S2, the method further includes: Industry standards, regulatory requirements and industry practical experience are used as quality control rules, and the quality control rules are classified and stored through predefined business categories and usage scenarios to build an industry rule library.

4. The multi-dimensional data quality evaluation method based on machine learning and industry rule base according to claim 1 is characterized in that: In step S2, the null value rate dimension indicator refers to the ratio of the total number of fields with empty data fields in the business data to the total data volume within the accounting period and the accounting scope. The calculation formula of the null value rate dimension indicator is as follows: Where: Indicates the quality score of the null value rate dimension indicator, data field i indicates the i-th data field, data field n indicates the n-th data field, and n indicates the total number of data fields; The integrity dimension indicator refers to randomly selecting a piece of business data from a category of business data with an associated relationship within the accounting period and accounting scope, and randomly extracting sampled data from the business data, and calculating the proportion of the amount of data in the sampled data that meets the associated conditions to the total amount of sampled data through the associated fields. The calculation formula of the integrity dimension indicator is as follows: Where: represents the quality score of the completeness dimension indicator, business data j represents the j-th business data, business data m represents the m-th business data, and m represents the total number of business data; The continuity dimension index refers to the ratio of the change in the amount of uploaded data within two adjacent accounting cycles and within the accounting scope. The calculation formula of the continuity dimension index is as follows: Where: represents the quality score of the continuity dimension indicator, business data k represents the kth business data, business data p represents the pth business data, and p represents the total number of business data; The rationality dimension indicator refers to the data indicators of hospitals of different levels required by the state within the accounting period and accounting scope. The corresponding data indicators are counted from the business data to compare whether the values ​​meet the hospital's level requirements. The proportion of data that meet the comparison conditions in the accounting rules accounts for the total rules. The calculation formula of the rationality dimension indicator is as follows: Where: represents the quality score of the rationality dimension indicator, rule h represents the hth rule, rule q represents the qth rule, q represents the total number of rules, and the result of the rule meeting the comparison condition is 0 or 1, where 0 represents the result that the rule does not meet the comparison condition, and 1 represents the result that the rule meets the comparison condition; The correlation dimension indicator refers to selecting sampled data from a group of business data with a master-slave relationship within the accounting period and accounting scope, and calculating the proportion of the data volume in the sampled data that meets the master-slave relationship to the total sampled data volume through the associated fields. The calculation formula of the correlation dimension indicator is as follows: Where: It represents the quality score of the correlation dimension indicator, business data g represents the g-th piece of business data, business data w represents the w-th piece of business data, and w represents the total number of business data.

5. The multi-dimensional data quality evaluation method based on machine learning and industry rule base according to claim 1 is characterized in that: The step S3 comprises: According to the target accounting cycle, quality control rules are used to calculate the quality score of the medical and health data to be evaluated, and a quality score for each dimensional indicator is generated; wherein the target accounting cycle includes daily accounting cycle, weekly accounting cycle, monthly accounting cycle, quarterly accounting cycle, and annual accounting cycle.

6. The multi-dimensional data quality evaluation method based on machine learning and industry rule base according to claim 1 is characterized in that: In step S4, the calculation formula of the total quality score is as follows: Where: Represents the overall quality score.

7. The multi-dimensional data quality evaluation method based on machine learning and industry rule base according to claim 1 is characterized in that: Before step S5, the method further includes: Acquire quality control evaluation results generated by historical manual evaluation, wherein the quality control evaluation results include multiple historical data sets and historical quality control evaluation reports related to the historical data sets; A semantic vector model is used to vectorize all historical data sets and historical quality control evaluation reports related to the historical data sets, and multiple historical data set embedding vectors and historical quality control evaluation report embedding vectors related to the historical data sets are obtained. Based on the historical data set embedding vectors and historical quality control evaluation report embedding vectors related to the historical data sets, a quality control evaluation knowledge base is constructed. Construct an initial quality control evaluation model based on the random forest algorithm; The historical data set embedding vector and the historical quality control evaluation report embedding vector related to the historical data set are split into a training data set and a test data set. The initial quality control evaluation model is trained and optimized based on the training data set and the test data set, and the trained quality control evaluation model is output.

8. The multi-dimensional data quality evaluation method based on machine learning and industry rule base according to claim 7 is characterized in that: The step S5 comprises: Based on the basic information of the source of the medical and health data to be evaluated, the business category to which the medical and health data to be evaluated belongs, the selected quality control rules, the quality score of each dimension indicator and the total quality score, a data set is constructed, and the data set is vectorized using a semantic vector model to obtain a data set embedding vector; The dataset embedding vector is input into the trained quality control evaluation model, and the historical quality control evaluation report embedding vector for the dataset is output. Based on the historical quality control evaluation report embedding vector, a historical quality control evaluation report is obtained, wherein the historical quality control evaluation report includes analysis result information, score result information, rating result information, conclusion result information, risk warning information and action suggestion information.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to claim 8 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to claim 8 are implemented.

Citation Information

Cited By

  • Data quality evaluation method based on government affair field

    CN120387105A

  • Trust evaluation improvement method and platform for regional health data

    CN120407555A

  • Method and device for evaluating semantic quality of data set, and electronic equipment

    CN120996027A

  • Method and apparatus for dataset semantic quality evaluation, and electronic device

    CN120996027B

  • Subjective and objective fused data quality adaptive measurement method and system and storage medium

    CN121188521A