A data quality governance method and system, electronic equipment and storage medium

By acquiring and standardizing data from multiple data sources, this data quality governance method solves the problem of incomplete quality reports caused by a single data source, realizes multi-dimensional data quality assessment and optimization strategies, and improves the effectiveness and security of data quality governance.

CN120469901BActive Publication Date: 2025-11-28北京思普艾斯科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510554446.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-11-28
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The current technology relies on a single data source, resulting in incomplete data quality reports and affecting the effectiveness of data quality governance.

Method used

By acquiring various types of raw data from multiple data sources, such as hospitals, companies, and pharmaceutical manufacturers, and after standardizing the data, it is input into the federal data quality assessment model to evaluate its completeness, consistency, and timeliness, generate a global quality assessment report, and formulate optimization strategies.

Benefits of technology

It enables multi-dimensional assessment of data quality, improves the accuracy and effectiveness of data quality governance, and ensures the security and reliability of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469901B_ABST
    Figure CN120469901B_ABST
Patent Text Reader

Abstract

The application discloses a data quality governance method and relates to the technical field of privacy computing and data governance. In the method, a plurality of original data is acquired, wherein one kind of the original data is from a data source, and the data source comprises a hospital, a company and a medicine production enterprise; the original data is subjected to standardization treatment to obtain standardized data; the standardized data is input into a preset federal data quality evaluation model for data quality evaluation to obtain a quality evaluation result corresponding to each kind of the standardized data, wherein the quality evaluation result comprises a completeness evaluation result, a consistency evaluation result and a timeliness evaluation result; the quality evaluation results are aggregated to obtain a global quality evaluation report; and a corresponding quality optimization strategy is determined according to the global quality evaluation report, and data is governed according to the quality optimization strategy. The effect of data quality governance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of privacy computing and data governance, in particular to a data quality governance method and system, an electronic device and a storage medium. BACKGROUND

[0002] With the deep penetration of Internet of Things, 5G and artificial intelligence technology, global data volume is growing at a rate of 26% per year, and data forms have also expanded from traditional structured tables to unstructured images, voice, real-time streaming data and other diversified forms. This explosive growth and deepening of technology dependence make data quality a key factor affecting decision-making reliability. Low-quality data has caused significant risk in the financial, medical and other fields, for example, incorrect data in financial transactions can cause systemic risk, and annotation errors in medical images can affect diagnostic accuracy. Therefore, effective methods are needed to improve data quality.

[0003] In the prior art, the quality of existing data is monitored by scheduling tasks, and a quality report is generated to govern data quality based on the quality report.

[0004] In the prior art, the data source is single, and the data quality is obtained by relying on a single monitoring data, which results in an incomplete quality report and low data quality governance effect. SUMMARY

[0005] In a first aspect of the present application, a data quality governance method is provided, comprising: obtaining a plurality of original data, wherein one of the original data is from a data source, and the data source includes a hospital, a company and a pharmaceutical production enterprise; performing standardization processing on the original data to obtain standardized data; inputting the standardized data into a preset federated data quality evaluation model to evaluate the data quality, to obtain a quality evaluation result corresponding to each of the standardized data, the quality evaluation result including a completeness evaluation result, a consistency evaluation result and a timeliness evaluation result; aggregating each of the quality evaluation results to obtain a global quality evaluation report; determining a corresponding quality optimization strategy according to the global quality evaluation report, and governing the data according to the quality optimization strategy.

[0006] By adopting the technical scheme, firstly, the original data from multiple data sources such as hospitals, companies and drug production enterprises is acquired, overcoming the defect of single data source in the prior art. Secondly, the standardized data is obtained by normalizing the original data, laying a unified data foundation for subsequent quality evaluation. Thirdly, the standardized data is input into the preset federal data quality evaluation model, and the data is comprehensively evaluated from the three dimensions of completeness, consistency and timeliness, obtaining a multi-dimensional quality evaluation result, avoiding the problem that the quality report is not comprehensive due to the dependence on single monitoring data in the prior art. Then, by aggregating each quality evaluation result, a global quality evaluation report containing multiple data sources and multiple evaluation dimensions is obtained, making the identification of data quality problems more comprehensive and accurate. Finally, based on the comprehensive quality evaluation report, targeted quality optimization strategies are formulated and data governance is performed, thereby improving the effect of data quality governance.

[0007] Optionally, the normalizing the original data to obtain standardized data specifically comprises: constructing a medical health knowledge base comprising disease coding specifications, diagnosis and treatment plan specifications and drug instruction specifications; importing each of the original data into a corresponding isolated area; extracting key fields of the original data in the isolated area; performing standardization processing on the key fields based on the medical health knowledge base to obtain structured data; and performing security processing on the structured data to obtain the standardized data.

[0008] By adopting the technical scheme, firstly, a medical health knowledge base comprising disease coding specifications, diagnosis and treatment plan specifications and drug instruction specifications is constructed to provide a standard basis for data normalization processing. Secondly, by importing original data of different sources into corresponding isolated areas and extracting key fields, effective isolation and feature extraction of data are realized. Thirdly, the key fields are standardized based on the medical health knowledge base, converting unstructured original data into structured data, improving the standardization and usability of the data. Finally, the structured data is processed to obtain standardized data, which not only ensures the standardization of data processing, but also ensures the security of data processing, providing a high-quality data foundation for subsequent quality evaluation.

[0009] Optionally, the security processing on the structured data to obtain the standardized data specifically comprises: replacing sensitive information in the structured data to obtain desensitized data; performing distributed storage processing on the desensitized data to obtain privacy data; and setting a visible range and an operation permission of the privacy data to obtain the standardized data.

[0010] By adopting the technical scheme, firstly, the sensitive information in the structured data is replaced to obtain the desensitized data, and the privacy information in the original data is protected. Secondly, the privacy data is obtained by performing distributed storage processing on the desensitized data, and the security and reliability of data storage are improved. Finally, the standardized data is obtained by performing fine setting on the visible range and operation permission of the privacy data, multi-level control of data access is realized, data availability is ensured while data security is ensured, and a safe and reliable data basis is provided for data quality evaluation.

[0011] Optionally, the standardized data is input into a preset federated data quality evaluation model for data quality evaluation to obtain a quality evaluation result corresponding to each standardized data, specifically including: performing integrity evaluation on the standardized data to obtain an integrity evaluation result; performing consistency evaluation on the standardized data to obtain a consistency evaluation result; performing timeliness evaluation on the standardized data to obtain a timeliness evaluation result; and performing fusion processing on the integrity evaluation result, the consistency evaluation result and the timeliness evaluation result to obtain the quality evaluation result of each data source.

[0012] By adopting the technical scheme, firstly, the standardized data is evaluated from the three dimensions of integrity, consistency and timeliness, and the core features of data quality are comprehensively covered. Secondly, the comprehensive quality evaluation result is obtained by fusion processing on the evaluation results of the three dimensions, and the one-sidedness that may be caused by single dimension evaluation is avoided. This multi-dimensional evaluation and fusion manner can more accurately reflect the data quality status of each data source, and provides a reliable decision basis for subsequent quality optimization.

[0013] Optionally, the integrity evaluation result, the consistency evaluation result and the timeliness evaluation result are fused to obtain the quality evaluation result of each data source, specifically including: constructing a feature distribution matrix based on the integrity evaluation result, the consistency evaluation result and the timeliness evaluation result, the rows of the feature distribution matrix are each data source, the columns are feature values corresponding to the integrity evaluation result, the consistency evaluation result and the timeliness evaluation result, and the feature values are normalized values reflecting quantitative indicators of each evaluation dimension, and the evaluation dimensions include integrity dimension, consistency dimension and timeliness dimension; the feature values of each data source in the feature distribution matrix are weighted calculated based on preset weight values of each evaluation dimension to obtain a comprehensive quality score of each data source; and the quality evaluation result of each data source is obtained based on a preset score benchmark and the comprehensive score.

[0014] By adopting the technical scheme, firstly, the performance of each data source in different evaluation dimensions is quantitatively represented in the form of standardized characteristic values by constructing a characteristic distribution matrix. Secondly, the characteristic values are weighted and calculated by introducing preset weight values, which reflects the importance difference of different evaluation dimensions. Finally, the comprehensive score is classified based on the preset scoring benchmark, and an intuitive quality evaluation result is obtained. This multi-dimensional fusion method based on matrix operation not only ensures the scientificity and interpretability of the evaluation process, but also improves the accuracy and reliability of the evaluation result.

[0015] Optionally, the consistency evaluation on the standardized data is performed to obtain a consistency evaluation result, specifically including: comparing the number of standard fields of each data source in the standardized data to obtain a number consistency comparison result, the standard fields including numerical standard fields and time standard fields; comparing the numerical standard fields of each data source in the standardized data to obtain a numerical consistency comparison result; comparing the time standard fields of each data source in the standardized data to obtain a time sequence consistency comparison result; and performing weighted calculation on the number consistency comparison result, the numerical consistency comparison result and the time sequence consistency comparison result to obtain the consistency evaluation result.

[0016] By adopting the technical scheme, firstly, the standardized data is comprehensively compared from three aspects of number consistency, numerical consistency and time sequence consistency. Secondly, the consistency evaluation is specially performed on the numerical fields and the time fields respectively, which improves the pertinence of the evaluation. Finally, the comprehensive consistency evaluation result is obtained by weighted calculation on the comparison results of the three aspects, which not only considers the characteristics of different types of data, but also ensures the comprehensiveness and accuracy of the evaluation result.

[0017] Optionally, the corresponding quality optimization strategy is determined according to the global quality evaluation report, specifically including: analyzing the global quality evaluation report based on preset evaluation indexes to obtain a data governance task; refining the data governance task based on historical experience and expert rules to obtain a data governance task sequence, the historical experience being a processing mode extracted from a historical data governance success case, the expert rule being a data governance best practice specification summarized by a domain expert, and the data governance task sequence including a plurality of subtasks; determining a priority weight of each subtask based on an urgency degree and resource consumption of each subtask, and performing task arrangement and computing resource allocation on the subtasks in each data governance task sequence according to the priority weight to obtain the quality optimization strategy.

[0018] By adopting the technical scheme, firstly, the global quality evaluation report is analyzed based on the preset evaluation index, and the specific problems that need to be managed are clarified. Secondly, the management task is refined by combining historical experience and expert rules, improving the feasibility and effectiveness of the management scheme. Finally, the priority is sorted and the resources are allocated by considering the task urgency and resource consumption, ensuring the efficient execution of the management process. This systematic optimization strategy formulation method not only absorbs historical experience and expert knowledge, but also considers the actual execution conditions, which can significantly improve the efficiency and effectiveness of data governance.

[0019] In a second aspect of the present application, a data quality management system is provided, specifically comprising:

[0020] A data acquisition module is configured to acquire a plurality of raw data, wherein one of the raw data is from a data source, and the data source includes a hospital, a company, and a pharmaceutical production enterprise.

[0021] A data processing module is configured to normalize the raw data to obtain standardized data.

[0022] A federal data quality evaluation module is configured to input the standardized data into a preset federal data quality evaluation model to evaluate the data quality, and obtain a quality evaluation result corresponding to each standardized data, wherein the quality evaluation result includes an integrity evaluation result, a consistency evaluation result, and a timeliness evaluation result.

[0023] A global quality aggregation module is configured to aggregate the quality evaluation results to obtain a global quality evaluation report.

[0024] A quality optimization and management module is configured to determine a corresponding quality optimization strategy based on the global quality evaluation report, and manage the data according to the quality optimization strategy.

[0025] By adopting the above technical scheme, firstly, the raw data from multiple data sources such as hospitals, companies, and pharmaceutical production enterprises is acquired, overcoming the defect of single data source in the prior art. Secondly, the standardized data is obtained by normalizing the raw data, laying a unified data foundation for subsequent quality evaluation. Thirdly, the standardized data is input into the preset federal data quality evaluation model, and the data is comprehensively evaluated from the three dimensions of integrity, consistency, and timeliness to obtain multi-dimensional quality evaluation results, avoiding the problem that the quality report is not comprehensive in the prior art due to the dependence on single monitoring data. Then, by aggregating each quality evaluation result, a global quality evaluation report containing multiple data sources and multiple evaluation dimensions is obtained, making the identification of data quality problems more comprehensive and accurate. Finally, based on the comprehensive quality evaluation report, a targeted quality optimization strategy is formulated and data management is performed, thereby significantly improving the effect of data quality management.

[0026] In a third aspect of the present application, an electronic device is provided, comprising a processor, a memory, a user interface and a network interface, the memory is configured to store instructions, the user interface and the network interface are configured to communicate with other devices, and the processor is configured to execute the instructions stored in the memory to enable the electronic device to perform the method according to any one of the preceding aspects.

[0027] In a fourth aspect of the present application, a computer readable storage medium is provided, which stores instructions, when the instructions are executed, the method according to any one of the preceding aspects is performed. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1 is a schematic diagram of an architecture of a data quality governance method disclosed by embodiments of the present application;

[0029] Figure 2 is a schematic diagram of a flow of a data quality governance method disclosed by embodiments of the present application;

[0030] Figure 3 is Figure 2 is a schematic diagram of a sub-step of step S102;

[0031] Figure 4 is Figure 3 is a schematic diagram of a sub-step of step S205;

[0032] Figure 5 is Figure 2 is a schematic diagram of a sub-step of step S103;

[0033] Figure 6 is Figure 5 is a schematic diagram of a sub-step of step S402;

[0034] Figure 7 is Figure 5 is a schematic diagram of a sub-step of step S404;

[0035] Figure 8 is Figure 2 is a schematic diagram of a sub-step of step S105;

[0036] Figure 9 is a schematic diagram of a module of a data quality governance system provided by embodiments of the present application;

[0037] Figure 10 is a schematic diagram of a structure of an electronic device disclosed by embodiments of the present application.

[0038] Reference signs: 21, data acquisition module; 22, data processing module; 23, federal data quality evaluation module; 24, global quality evaluation module; 25, quality optimization and management module; 901, processor; 902, communication bus; 903, user interface; 904, network interface; 905, memory. DETAILED DESCRIPTION

[0039] In order for those skilled in the art to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in combination with the drawings in the specification. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments.

[0040] In the description of the embodiments of the present application, the words such as "for example" or "for instance" are used to represent an example, illustration or description. Any embodiment or design scheme described as "for example" or "for instance" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words such as "for example" or "for instance" are intended to present the relevant concept in a specific way.

[0041] In the description of the embodiments of the present application, the term "a plurality of" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are used for description purposes only, and should not be interpreted as indicating or implying relative importance or implicitly indicating the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features. The terms "include", "contain", "have" and their variants mean "include but are not limited to", unless otherwise specifically emphasized.

[0042] Figure 1 An exemplary system architecture 10 is shown, which can apply to an embodiment of the data quality management method or the data quality management system of the present application.

[0043] As shown in Figure 1 The system architecture 10 can include a terminal device 11, a network 12, and a server 13. The network 12 is a medium for providing a communication link between the terminal device 11 and the server 12. The network 12 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0044] A user can use the terminal device 11 to interact with the server 13 through the network 12 to receive or send data, etc.

[0045] The terminal device 11 can be hardware or software. When the terminal device 11 is hardware, it can be various electronic devices with a display screen, including but not limited to a smartphone, a tablet computer, a laptop computer, a desktop computer, and the like. When the terminal device 11 is software, it can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (for example, multiple software or software modules for providing distributed services) or as a single software or software module. No specific limitation is made herein.

[0046] The server 13 can be a server providing various services, such as a background server for processing data displayed on the terminal device 11. The background server can analyze and process received data, and can feed back the processing result (for example, a recognition result) to the terminal device.

[0047] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, multiple software or software modules for providing distributed services) or as a single software or software module. No specific limitation is made herein.

[0048] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the system architecture is merely illustrative. According to the implementation needs, there can be any number of terminal devices, networks, and servers. In particular, in the case where the target data does not need to be obtained from a remote location, the above system architecture can not include a network, but only a terminal device or a server.

[0049] The embodiment discloses a data quality governance method, Figure 2 is a flowchart of data quality governance disclosed by the embodiment of the present application, as Figure 2 shown, including steps S101 to S105, which are as follows:

[0050] S101: Obtain multiple raw data, wherein one raw data is from one data source, and the data source includes a hospital, a company, and a pharmaceutical production enterprise.

[0051] The raw data represents data in the original form that has not been processed, such as electronic medical record of a hospital, sales data of a company, etc. One raw data is from one data source, and the data source refers to an entity or system that generates or stores the raw data, which can include a hospital, a company, and a pharmaceutical production enterprise. For example, the data source is a hospital, and the corresponding raw data can be patient medical record, drug sales data, production batch information, etc.

[0052] Specifically, first, the required original data types, such as patient medical records, drug sales data, production batch information, etc., need to be determined; then data connections or authorizations are established with different data sources; then various types of original data are obtained through data interfaces, file transfers, etc.; finally, the obtained data is preliminarily checked for integrity to ensure the accuracy and integrity of data acquisition.

[0053] In some embodiments, the acquisition and collection of original data can be achieved in various ways.

[0054] Optionally, by establishing API connection with each data source and obtaining access key, writing data request program to define required data fields, the system automatically calls API interface to obtain data according to preset time interval, and stores the obtained data in local database for unified management.

[0055] Optionally, by agreeing with the data source on the format and field specification of data export, the data source exports data files of specified format regularly, receives data files through secure file transfer, and finally uses data processing program to parse the files and import the data into the system.

[0056] It can be understood that other ways can also be used to achieve data acquisition and collection, such as direct database access, crawler collection, etc., which are not limited here.

[0057] S102: Standardizing the original data to obtain standardized data.

[0058] The standardized data represents the data after standardization, which has a unified numerical range and comparability. Standardization refers to the process of converting data to a unified scale range through a specific mathematical transformation method, including min-max standardization, Z-score standardization, etc.

[0059] Specifically, after collecting the original data, first, the distribution characteristics and value range of the data need to be analyzed, and a suitable standardization method is selected. Then, standardization calculation is performed on each feature dimension respectively, and the numerical value is mapped to the specified interval range. Finally, the processing result is verified to ensure that the converted data meets the expected distribution characteristics and statistical properties.

[0060] In some embodiments, the standardization of the original data can be achieved in various ways. Optionally, the min-max standardization method is used to process the data, first, the maximum and minimum values of each feature in the original data are calculated, then all data are converted to the [0, 1] interval according to the mapping formula, and finally the converted data is checked and the conversion parameters are saved for subsequent use.

[0061] Optionally, the Z-score standardization method can be used for data processing. By calculating the mean and standard deviation of the original data, the data can be transformed into a form that conforms to a standard normal distribution, and the transformation effect can be verified and relevant parameters can be recorded.

[0062] It is understandable that other standardization methods, such as decimal scaling and logarithmic transformation, can also be used to achieve data standardization, which is not limited here.

[0063] Reference Figure 3 , Figure 3 This is provided by the embodiments of this application. Figure 2 A flowchart illustrating a sub-step of step S102, including steps S201 to S205, is as follows:

[0064] S201: Construct a medical and health knowledge base, which includes disease coding standards, treatment plan standards, and drug instruction standards.

[0065] The medical and health knowledge base represents a systematically organized collection of professional knowledge in the field of medical and health care. Disease coding standards refer to a standardized system for uniformly coding and classifying various diseases, such as the International Classification of Diseases (ICD). Treatment protocol standards represent standardized diagnostic and treatment procedures for different diseases. Drug instruction manual standards are used to represent standardized information such as instructions for use, indications, dosage, and precautions for drugs.

[0066] Specifically, the first step is to collect various medical and health-related regulations and standards from authoritative medical institutions and data sources. Then, the collected content is organized, classified, and structured to establish relationships between diseases, treatment plans, and medications. Finally, the processed data is imported into a knowledge base system, where it is validated and maintained to ensure the accuracy and completeness of the knowledge base.

[0067] In some embodiments, a medical and health knowledge base can be constructed in various ways. Optionally, a relational database storage method can be used to construct the knowledge base. This involves establishing data tables for diseases, treatment plans, and drugs, designing relationships between tables, importing standardized medical data, and creating indexes to improve query efficiency. Finally, the knowledge base can be managed and maintained.

[0068] Optionally, a medical knowledge graph can be constructed using graph database technology, with diseases, treatment plans, and drugs as entity nodes, establishing relationship edges between entities to form a complete knowledge network structure, and providing intelligent knowledge retrieval and reasoning capabilities.

[0069] It is understandable that other data storage and management methods, such as document databases and hybrid storage, can also be used to build a medical and health knowledge base, and this is not limited here.

[0070] S202: Import each raw data into the corresponding isolated region.

[0071] Wherein, the isolated region refers to the independent storage and processing space set for data of different sources.

[0072] Specifically, first, different isolated regions are divided according to data types and security levels, and corresponding access permissions and security policies are configured for each region. Then, data classification rules are established to identify the characteristics and attributes of each raw data. Next, according to the classification results, the data is imported into the corresponding isolated region, and data integrity check is performed during the import process. Finally, the data import result is confirmed to ensure that the data in each isolated region meets the expected requirements.

[0073] In some embodiments, the import of raw data into isolated regions can be achieved in various ways.

[0074] Optionally, a physical isolation method is used, which establishes independent data partitions on different servers or storage devices, uses dedicated data transmission channels and encryption mechanisms to safely import various raw data into corresponding physical isolated regions, and implements strict access control and audit mechanisms.

[0075] Optionally, a logical isolation method is used, which establishes independent database instances or table spaces in the same storage system, sets different data schemas and access permissions, and realizes logical isolation storage and management of various data.

[0076] It can be understood that other isolation methods such as containerization isolation, hybrid isolation, etc. can also be used to realize the secure isolation storage of raw data, which is not limited here

[0077] S203: Extract the key fields of raw data in the isolated region.

[0078] Wherein, the key field represents the core data attribute that is important for data analysis and processing, such as patient basic information, diagnosis results, treatment plans, etc.

[0079] Specifically, first, the scope and type of key fields are determined according to business requirements and analysis goals. Then, for raw data in each isolated region, target fields are identified using field mapping rules. Next, field value extraction and verification are performed to ensure data integrity and accuracy. Finally, the extracted key fields are structured to facilitate subsequent processing and analysis.

[0080] In some embodiments, the extraction of key fields can be achieved in various ways. Optionally, a rule configuration method is used for field extraction, by pre-defining field extraction rules and mapping relationships, the original data is scanned and matched to identify and extract key field content that meets the rules, and data quality verification and exception handling are performed.

[0081] Optionally, an intelligent parsing technology is used for field extraction, by natural language processing and pattern recognition algorithms, the structural features of the original data are automatically analyzed to identify important field information, and intelligent extraction and labeling are performed.

[0082] It can be understood that other field extraction methods such as template matching, deep learning, etc. can also be used to achieve accurate extraction of key data, which is not limited here

[0083] S204: Standardize the key fields based on the medical health knowledge base to obtain structured data.

[0084] Among them, the structured data represents the data form after standardized processing, which has a fixed format and unified standard. The standardized processing is to convert data in different formats and expressions into a unified specification.

[0085] Specifically, first, the extracted key fields need to be matched with the standard specifications in the medical health knowledge base to identify the type of the field and the corresponding standard specification. Then, according to the matching result, the content of each field is standardized and converted, including term unification, code mapping, format conversion, etc. The converted data is then verified to ensure compliance with the standard specification requirements. Finally, the standardized data is organized into a structured form and the association between fields is established.

[0086] In some embodiments, the standardization of key fields can be achieved in various ways.

[0087] Optionally, a rule mapping method is used for standardization, by establishing a mapping dictionary of field values and standard specifications, the key fields are matched and converted one by one, abnormal values and non-standard expressions are processed, and finally the structured data that meets the specifications is generated.

[0088] Optionally, a natural language processing technology is used for standardization, by semantic analysis and entity recognition algorithms, the actual meaning of the field content is understood, it is automatically mapped to the standard terms and codes in the knowledge base, and intelligent structured conversion is performed.

[0089] It can be understood that other standardization methods such as machine learning classification, expert system, etc. can also be used to achieve the standardization and structuring of key fields, which is not limited here.

[0090] S205: Perform secure processing on structured data to obtain standardized data.

[0091] Standardized data refers to regulated data that has undergone security processing before being used for subsequent analysis. Security processing involves operations such as data anonymization and encryption to protect privacy. Structured data refers to data with a fixed format and unified standards.

[0092] Specifically, the first step is to identify sensitive information fields in the structured data to determine the scope of data requiring security processing. Then, based on the different types of sensitive information, appropriate security processing methods are selected for transformation. Next, the processed data undergoes integrity verification to ensure that data availability is not affected. Finally, the security-processed data is saved in a standardized format.

[0093] In some embodiments, secure processing of structured data can be achieved in a variety of ways.

[0094] Optionally, data anonymization can be used for security processing. Sensitive information can be masked or transformed using techniques such as replacement, masking, and truncation to protect privacy while maintaining the value of data analysis. For example, a specific name can be replaced with a random string.

[0095] Optionally, data encryption can be used for security processing. Sensitive data is encrypted using cryptographic algorithms to ensure data security during storage and transmission, while providing necessary decryption mechanisms to support subsequent use.

[0096] It is understandable that other security measures, such as data classification and access control, can be used to protect the privacy of structured data, and no specific measures are proposed here.

[0097] Reference Figure 4 , Figure 4 This is provided by the embodiments of this application. Figure 3 A flowchart illustrating a sub-step of step S205, including steps S301 to S303, is shown below:

[0098] S301: Replace sensitive information in structured data to obtain desensitized data.

[0099] Sensitive information refers to data that may involve privacy and security concerns, such as personal identification information and contact details. De-identified data refers to secure and usable data after the sensitive information has been replaced.

[0100] Specifically, first, the sensitive information field in the structured data needs to be identified to determine the data range and type that needs to be replaced. Then, according to different types of sensitive information, corresponding replacement rules and methods are selected for conversion. Next, the replaced data is verified to ensure that the replacement result meets the data format requirements. Finally, the replaced data is saved in a desensitized format.

[0101] In some embodiments, the replacement of sensitive information can be achieved in various ways.

[0102] Optionally, desensitization is performed by replacing characters, which replaces part or all of the content of sensitive information with specific characters (such as #), while maintaining the basic features of the data and achieving privacy protection, such as replacing the middle four digits of a mobile phone number with ***.

[0103] Optionally, desensitization is performed by using random value replacement, which replaces the original sensitive information with a random string or numerical value that meets the rules, maintains the distribution characteristics and analysis value of the data, and ensures that the sensitive information cannot be restored.

[0104] It can be understood that other replacement methods such as mapping conversion and interval replacement can also be used to achieve desensitization of sensitive information, which is not limited here.

[0105] S302: Perform distributed storage processing on the desensitized data to obtain private data.

[0106] Among them, the distributed storage processing represents a technical method of dispersing data storage in multiple independent nodes.

[0107] Specifically, first, the desensitized data needs to be divided into shards to determine the data distribution strategy and storage node allocation scheme. Then, the divided data is stored in different nodes to establish data indexing and association. Next, consistency check is performed on the distributed stored data to ensure data integrity and accessibility. Finally, a data access control mechanism is established to achieve privacy protection.

[0108] In some embodiments, distributed storage processing can be achieved in various ways.

[0109] Optionally, data sharding is used for storage, which divides the complete data set into multiple data shards according to specific rules and stores them on different physical nodes to reduce the risk of data leakage and improve data access efficiency.

[0110] Optionally, data replication is used for storage, which saves data copies on multiple nodes to achieve data disaster recovery and load balancing, while using a differentiated access permission control mechanism to protect data security.

[0111] It can be understood that other storage processing methods such as hierarchical storage, heterogeneous storage, etc. can also be used to realize privacy protection of desensitized data, which is not limited here.

[0112] S303: Set the visible range and operation permission of the privacy data to obtain standardized data.

[0113] Among them, the visible range represents the range of users or systems allowed to access the data. The operation permission represents the authorization level of reading, modifying, etc. The standardized data represents the standardized data after permission management.

[0114] Specifically, first, the data access strategy needs to be defined to determine the visible range and operation permission level of different roles. Then, set the corresponding access control rules for the privacy data, and establish the mapping relationship between user identity and permission. Then implement the permission verification mechanism to ensure that the data access meets the authorization requirements. Finally, associate the permission setting result with the data to form standardized data.

[0115] In some embodiments, the permission setting can be implemented in various ways.

[0116] Optionally, a role-based access control method is used to assign corresponding roles to different users by establishing the correspondence between roles and permissions, to realize fine-grained management of data access and operation, such as setting the doctor role to view patient medical records, and the nurse role to view only basic information.

[0117] Optionally, a multi-level authorization method is used for control, by setting a hierarchical structure of data access, to implement differentiated permission control at different levels, to ensure the safe circulation of data within the authorized range, while supporting dynamic adjustment of permissions.

[0118] It can be understood that other permission control methods such as attribute encryption, time control, etc. can also be used to realize the security management of privacy data, which is not limited here.

[0119] S103: Input the standardized data into a pre-set federated data quality evaluation model to evaluate the data quality, to obtain the quality evaluation result corresponding to each standardized data, which includes the completeness evaluation result, the consistency evaluation result and the timeliness evaluation result.

[0120] Among them, the data quality evaluation represents the quantitative analysis of the completeness, consistency and timeliness of the data. The pre-set federated data quality evaluation model is an analysis model used to evaluate the quality of standardized data. The quality evaluation result represents the specific measurement index of data quality.

[0121] Specifically, firstly, standardized data is input into the evaluation model, and data analysis is performed according to preset evaluation indicators. Then, evaluation scores are calculated for three dimensions: completeness, consistency, and timeliness. Next, the evaluation results for each dimension are summarized to generate a quality evaluation report. Finally, the data quality is graded and labeled based on the evaluation results.

[0122] In some embodiments, the evaluation model may include multiple evaluation dimensions.

[0123] Optionally, completeness assessment evaluates the completeness and validity of data by calculating indicators such as the proportion of non-null values ​​and the coverage of valid values ​​in data fields, and identifies missing data and anomalies.

[0124] Optionally, consistency assessment evaluates the standardization and uniformity of data by analyzing characteristics such as data format standardization and the rationality of value range, and detects data conflicts and inconsistencies.

[0125] Optionally, timeliness assessment can evaluate the real-time nature and time validity of data by using characteristics such as statistical data update frequency and timestamp distribution, and determine whether the data meets the timeliness requirements.

[0126] It is understandable that other evaluation dimensions, such as accuracy and uniqueness, can also be used to achieve a comprehensive assessment of data quality, and no limitation is made here.

[0127] Reference Figure 5 , Figure 5 This is provided by the embodiments of this application. Figure 2 A flowchart illustrating a sub-step of step S103, including steps S401 to S404, is as follows:

[0128] S401: Perform an integrity assessment on the standardized data to obtain the integrity assessment results.

[0129] The completeness assessment result is obtained after the data integrity assessment. The integrity assessment involves analyzing and evaluating the population and validity of data fields.

[0130] Specifically, the first step is to statistically analyze the data distribution of each field in the standardized data, and calculate the field fill rate and the proportion of valid values. Then, the degree of data loss and the level of validity are assessed against preset completeness standards. Next, field-level completeness scores are generated, and a weighted calculation is performed to obtain the overall completeness assessment result.

[0131] In some embodiments, integrity assessment can be achieved in a variety of ways.

[0132] Optionally, a missing value analysis method can be used for evaluation. By calculating the proportion of empty and invalid values ​​in the dataset, the location and degree of missing data can be identified, the data filling completeness can be evaluated, and a statistical report on the missing data can be generated.

[0133] Optionally, an effective value detection method can be used for evaluation. This method verifies whether the data values ​​meet the preset format specifications and value range, calculates the proportion of valid data, evaluates the quality and integrity of the data, and identifies abnormal data items.

[0134] It is understandable that other assessment methods, such as data distribution analysis and outlier detection, can also be used to achieve integrity assessment, and no limitation is made here.

[0135] S402: Perform a conformity assessment on the standardized data to obtain the conformity assessment results.

[0136] The consistency assessment result refers to the outcome obtained after a data consistency assessment. The consistency assessment involves verifying and analyzing the standardization and uniformity of the data.

[0137] Specifically, the first step is to check the format compliance of each field in the standardized data and verify the compliance of the data values. Then, compare the data with preset standards to assess the standardization and consistency of the field values. Next, calculate the field-level consistency scores and summarize them to obtain the overall consistency assessment result.

[0138] In some embodiments, consistency assessment can be achieved in a variety of ways.

[0139] Optionally, a format specification testing method can be used for evaluation. This involves verifying whether the data's expression format, units of measurement, coding rules, etc., conform to unified standards, calculating standardization indicators, and assessing the data's format consistency level.

[0140] Optionally, a logical relationship verification method can be used for evaluation. By analyzing the dependencies between data items and business rules, the logical rationality of the data can be checked, and the semantic consistency level of the data can be evaluated.

[0141] It is understandable that other evaluation methods, such as data conflict detection and standard compliance analysis, can also be used to achieve consistency assessment, and no limitation is made here.

[0142] Reference Figure 6 , Figure 6 This is provided by the embodiments of this application. Figure 5 A flowchart illustrating a sub-step of step S402, including steps S501 to S504, is as follows:

[0143] S501: Compare the number of standard fields of each data source in the standardized data, and obtain a number consistency comparison result, the standard fields including numerical standard fields and time standard fields.

[0144] The number consistency comparison result represents the comparison of the number of fields of each data source. The standard field refers to a numerical and time data attribute conforming to a specification.

[0145] Specifically, first, the number of numerical and time standard fields in each data source is counted respectively. Then, the number of fields of different data sources is cross-compared to calculate the difference degree of the number of fields. Then, the comparison and analysis result of the number consistency is generated, including the distribution of the number of fields and the difference statistics.

[0146] In some embodiments, the number consistency comparison can be implemented in various ways.

[0147] Optionally, the field counting method is used for comparison, the specific number of numerical and time fields in each data source is counted, the ratio and difference of the number of fields are calculated, and the number consistency level between the data sources is evaluated.

[0148] Optionally, the difference analysis method is used for comparison, the difference points of the number of fields between the data sources are identified, the reasons for the difference are analyzed, the completeness of the field coverage is evaluated, and a difference distribution report is generated.

[0149] It can be understood that other comparison methods such as structure comparison and mapping analysis can also be used to implement the consistency evaluation of the number of standard fields, which is not limited here.

[0150] S502: Compare the numerical standard fields of each data source in the standardized data, and obtain a numerical consistency comparison result.

[0151] The numerical standard field represents a standardized data attribute expressed in a numerical form. The numerical consistency comparison result represents the comparison of the numerical content of each data source.

[0152] Specifically, first, the corresponding numerical standard fields in each data source are extracted. Then, the value range and distribution characteristics of the fields are statistically analyzed. Then, the difference degree of the numerical values between different data sources is compared, and a similarity index is calculated. Finally, the comparison results of each field are summarized, and a numerical consistency analysis result is generated.

[0153] In some embodiments, the numerical consistency comparison can be implemented in various ways.

[0154] Optionally, a statistical feature comparison method is used, and statistical indexes such as mean, variance, and distribution interval of the value fields of each data source are calculated to compare the concentration trend and dispersion degree of the values and to evaluate the consistency level of the value distribution.

[0155] Optionally, an exact value comparison method is used, and specific values of the corresponding fields are directly compared to calculate the value deviation rate and matching degree, identify abnormal values and inconsistent points, and evaluate the accurate consistency of the values.

[0156] It can be understood that other comparison methods such as correlation analysis and trend comparison can also be used to implement the consistency evaluation of the numerical fields, which are not limited here.

[0157] S503: Comparing the time type standard fields of each data source in the standardized data to obtain a time sequence consistency comparison result.

[0158] The time type standard field represents a standardized data attribute expressed in a time format. The time sequence consistency comparison result represents the comparison of the time sequences of each data source.

[0159] Specifically, the corresponding time type standard fields in each data source are first extracted. Then, the continuity and periodicity of the time sequence are analyzed. Next, the difference in time distribution between different data sources is compared, and the time sequence similarity is calculated. Finally, the comparison results of each field are summarized to generate the time sequence consistency analysis result.

[0160] In some embodiments, the time sequence consistency comparison can be implemented in various ways.

[0161] Optionally, a time span comparison method is used, and the start time, time interval, sampling frequency, and other features of the time fields of each data source are analyzed to evaluate the continuity and integrity of the time coverage and identify the missing and abnormal time sequence data.

[0162] Optionally, a time sequence pattern comparison method is used, and the change trend, periodic characteristics, and mutation points of the time sequence are compared to evaluate the consistency level of the time sequence data and find the differences in the time sequence changes.

[0163] It can be understood that other comparison methods such as time stamp alignment and time window analysis can also be used to implement the consistency evaluation of the time type fields, which are not limited here.

[0164] S504: Weighted calculation of the quantity consistency comparison result, the value consistency comparison result, and the time sequence consistency comparison result to obtain a consistency evaluation result.

[0165] The consistency evaluation result represents the final score result of the data consistency, and the consistency includes the consistency indexes of the quantity, value, and time sequence dimensions.

[0166] Specifically, first, the weight coefficients of each comparison result need to be determined, reflecting the importance of different dimensions. Then, the comparison results are standardized to unify the scoring scale. Next, the weight is calculated by weighted summation to obtain the comprehensive evaluation score. Finally, a detailed report of the consistency evaluation is generated.

[0167] In some embodiments, the weighted calculation can be implemented in various ways.

[0168] Optionally, a linear weighting method is used for calculation. By setting the weight coefficients of each dimension evaluation result, the standardized scores are weighted and summed to obtain a comprehensive score reflecting the overall consistency level.

[0169] Optionally, a hierarchical weighting method is used for calculation. By establishing a multi-level weight system, weighted operation is performed at different levels to realize step-by-step aggregation of evaluation results and obtain more detailed consistency scores.

[0170] It can be understood that other calculation methods such as fuzzy comprehensive evaluation and analytic hierarchy process can also be used to realize the weighted calculation of the consistency evaluation result, which is not limited here.

[0171] S403: Time effectiveness evaluation is performed on the standardized data to obtain a time effectiveness evaluation result.

[0172] The time effectiveness evaluation result represents the result obtained after data time effectiveness evaluation. Time effectiveness evaluation means analyzing the time effectiveness and update timeliness of data.

[0173] Specifically, first, the time attributes of the standardized data are analyzed, and the generation time and update frequency of the data are counted. Then, according to the preset time effectiveness standard, the time effectiveness and update timeliness of the data are evaluated. Next, the time effectiveness score of the data item is calculated and aggregated to obtain the overall time effectiveness evaluation result.

[0174] In some embodiments, time effectiveness evaluation can be implemented in various ways.

[0175] Optionally, a time decay method is used for evaluation. By calculating the interval between the data and the current time, a time decay function is set to evaluate the time effectiveness level of the data and identify expired and to-be-updated data items.

[0176] Optionally, an update frequency analysis method is used for evaluation. By counting the update time interval, update regularity and other characteristics of the data, the timeliness of data maintenance is evaluated to determine whether the data update meets the time effectiveness requirements.

[0177] It can be understood that other evaluation methods such as life cycle analysis and time effectiveness threshold detection can also be used to realize time effectiveness evaluation, which is not limited here.

[0178] S404: The integrity assessment results, consistency assessment results, and timeliness assessment results are merged to obtain the quality assessment results of each data source.

[0179] Among them, the quality assessment result represents the comprehensive assessment result of the overall quality level of the data source, and the overall quality includes indicators in three dimensions: completeness, consistency and timeliness.

[0180] Specifically, the first step is to determine the weighting of the evaluation results for each dimension, reflecting the importance of the evaluation indicators. Then, the evaluation results from different dimensions are normalized to ensure consistent scoring standards. Next, a weighted fusion calculation is performed according to the weights to obtain the overall quality score of the data source.

[0181] In some embodiments, fusion processing can be implemented in a variety of ways.

[0182] Optionally, a linear fusion approach can be used. By setting weight coefficients for the completeness, consistency, and timeliness assessment results, the normalized scores are weighted and summed to obtain a comprehensive score that reflects the overall quality level of the data source.

[0183] Optionally, a multi-level fusion approach can be used. By establishing a hierarchical evaluation index system, scores can be fused at different levels to achieve a step-by-step aggregation of evaluation results and obtain a more comprehensive quality score.

[0184] It is understandable that other fusion methods, such as fuzzy comprehensive evaluation and evidence theory, can also be used to achieve the fusion processing of quality assessment results, which is not limited here.

[0185] Reference Figure 7 , Figure 7 This is provided by the embodiments of this application. Figure 5 A flowchart illustrating a sub-step of step S404, including steps S601 to S603, is as follows:

[0186] S601: Construct a feature distribution matrix based on the completeness assessment results, consistency assessment results, and timeliness assessment results. The columns of the feature distribution matrix are the feature values ​​corresponding to the completeness assessment results, consistency assessment results, and timeliness assessment results. The feature values ​​are normalized values ​​that reflect the quantitative indicators of each assessment dimension. The assessment dimensions include the completeness dimension, consistency dimension, and timeliness dimension.

[0187] The feature distribution matrix represents the correspondence between data sources and evaluation dimensions. Eigenvalues ​​are the normalized quantitative indicators of each evaluation dimension. Evaluation dimensions include three aspects: completeness, consistency, and timeliness.

[0188] Specifically, first, the original scores of each data source in different evaluation dimensions need to be extracted. Then, the evaluation scores are normalized to convert the indicators of different dimensions into characteristic values of a unified scale. Then, a matrix structure is constructed, with the data sources as rows, the evaluation dimensions as columns, and the characteristic values as matrix elements.

[0189] In some embodiments, the matrix construction can be implemented in various ways.

[0190] Optionally, a linear normalization method is used to construct the matrix. The original evaluation scores are normalized by maximum and minimum values to convert the evaluation results of each dimension into characteristic values in the interval [0, 1] and fill them into the corresponding matrix positions.

[0191] Optionally, a standardization method is used to construct the matrix. The mean and standard deviation of the evaluation scores are calculated to convert the original scores into standardized characteristic values, ensuring the comparability of indicators of different dimensions.

[0192] It can be understood that other construction methods such as quantile conversion and logarithmic transformation can also be used to generate the feature distribution matrix, which is not limited here.

[0193] S602: Based on the preset weight values of each evaluation dimension, the characteristic values of each data source in the feature distribution matrix are weighted and calculated.

[0194] Specifically, first, the preset weight values of the integrity, consistency, and timeliness dimensions need to be obtained. Then, the weight values are multiplied by the characteristic values of the corresponding dimensions in the feature distribution matrix. Then, the weighted characteristic values of each data source are summarized to obtain the comprehensive score of the data source.

[0195] In some embodiments, the weighted calculation can be implemented in various ways.

[0196] Optionally, a vector multiplication method is used for calculation. The preset weight values are constructed as a weight vector, and the vector multiplication operation is performed between the weight vector and the characteristic values of each row in the feature distribution matrix to obtain the weighted results of each data source.

[0197] Optionally, a matrix operation method is used for calculation. The preset weight values are constructed as a weight matrix, and the matrix multiplication operation is performed between the weight matrix and the feature distribution matrix to obtain the weighted results of all data sources in batches.

[0198] It can be understood that other calculation methods such as hierarchical weighting and dynamic weight can also be used to implement the weighted calculation of the characteristic values, which is not limited here.

[0199] S603: Based on the preset score benchmark and the comprehensive score, the quality evaluation results of each data source are obtained.

[0200] The quality evaluation result represents the final quality level determination of each data source.

[0201] Specifically, first, the preset score benchmark is obtained to determine the score interval of different quality levels. Then, the comprehensive score of each data source is compared with the score benchmark. Next, the quality level of the data source is determined according to the score interval, and the quality evaluation result is generated.

[0202] In some embodiments, the evaluation result determination can be achieved in various ways.

[0203] Optionally, the threshold division method is used for determination, and the score threshold of the quality level is set to map the comprehensive score of the data source to the corresponding quality level, such as excellent, good, qualified, and unqualified levels.

[0204] Optionally, the interval mapping method is used for determination, and the score interval is subdivided into multiple subintervals to establish the correspondence between the score and the quality level, thereby achieving fine evaluation of the data source quality.

[0205] It can be understood that other determination methods such as fuzzy evaluation and cluster analysis can also be used to determine the quality evaluation result, which is not limited here.

[0206] S104: Aggregating each quality evaluation result to obtain a global quality evaluation report.

[0207] The global quality evaluation report represents a comprehensive analysis report of the quality evaluation results of each data source.

[0208] Specifically, first, the quality evaluation results and detailed scores of each data source are collected. Then, the distribution of different quality levels is counted, and the common characteristics of quality problems are analyzed. Next, statistical charts and analysis explanations of data quality are generated to form the global quality evaluation report.

[0209] In some embodiments, the evaluation result aggregation can be achieved in various ways.

[0210] Optionally, the statistical analysis method is used for aggregation, and the number distribution and proportion of each quality level are calculated to evaluate the overall data quality level and identify the main types and influence range of quality problems.

[0211] Optionally, the multi-dimensional analysis method is used for aggregation, and the quality evaluation results are cross-analyzed from multiple dimensions such as integrity, consistency, and timeliness to find the correlation characteristics and distribution rules of quality problems.

[0212] It can be understood that other aggregation methods such as hierarchical summary and trend analysis can also be used to generate the global quality evaluation report, which is not limited here.

[0213] S105: According to the global quality evaluation report, determine the corresponding quality optimization strategy, and manage the data according to the quality optimization strategy.

[0214] Specifically, first, the quality problems and distribution characteristics in the global quality evaluation report need to be analyzed. Then, according to the type, severity and influence range of the problem, the corresponding optimization strategy is formulated. Then, according to the strategy, data governance is performed, and data quality is continuously improved.

[0215] In some embodiments, quality optimization can be achieved in various ways.

[0216] Optionally, a hierarchical governance method is used for optimization, the quality problems are prioritized, a hierarchical processing scheme is formulated, and quality improvement is gradually implemented according to the severity and influence range of the problem.

[0217] Optionally, a classification governance method is used for optimization, and special governance measures are formulated for quality problems in different dimensions such as integrity, consistency and timeliness, and data quality improvement is implemented in a targeted manner.

[0218] It can be understood that other optimization methods such as process reengineering and standardization construction can also be used to achieve continuous optimization of data quality, which is not limited here.

[0219] Reference Figure 8 , Figure 8 FIG. 7 is a sub-step flowchart of step S105 in FIG. 0, provided by an embodiment of the present application, comprising steps S701-S703, and the above steps are as follows:

[0220] S701: Based on the preset evaluation index, analyze the global quality evaluation report to obtain a data governance task.

[0221] The data governance task represents a specific work item that needs to be performed to improve data quality.

[0222] Specifically, first, the preset evaluation index system needs to be obtained. Then, based on the evaluation index, analyze the problems in the global quality evaluation report. Then, the identified problems are converted into executable governance tasks, and the task target and completion standard are specified.

[0223] In some embodiments, the governance task formulation can be achieved in various ways.

[0224] Optionally, an index decomposition method is used for formulation, which decomposes the evaluation index into specific governance requirements, identifies the data items and processing methods that need to be improved, and forms an executable governance task list.

[0225] Optionally, the problem-oriented approach is used to formulate, the problems in the quality evaluation report are classified and combed to determine the priority order and solution of problem improvement, and converted into specific governance tasks.

[0226] It can be understood that other formulation methods such as goal decomposition and process optimization can also be used to generate data governance tasks, which are not limited here.

[0227] S702: Based on historical experience and expert rules, the data governance task is refined to obtain a data governance task sequence, the historical experience is the processing mode extracted from the historical data governance success case, the expert rule is the data governance best practice specification summarized by the domain expert, and the data governance task sequence includes multiple sub-tasks.

[0228] Among them, the historical experience represents the general processing method extracted from the past successful governance cases. The expert rule refers to the governance experience and best practice summarized by the domain expert. The sub-task represents the specific work item after the governance task is decomposed.

[0229] Specifically, first, the historical governance cases and expert rule library need to be collected. Then, the governance task is decomposed and refined based on experience and rules. Then, the sub-tasks are formed into a task sequence according to the execution order and dependency relationship.

[0230] In some embodiments, task refinement can be achieved in various ways.

[0231] Optionally, the experience mode is used for refinement, the processing mode and method in the historical successful case are extracted, combined with the current scene for optimization and adjustment, and an executable sub-task sequence is formed.

[0232] Optionally, the rule mapping method is used for refinement, the best practice specification summarized by the expert is applied to decompose the governance task into specific operation steps and checkpoints, and a standardized task sequence is generated.

[0233] It can be understood that other refinement methods such as template reuse and scene analysis can also be used to realize the serialization of governance tasks, which are not limited here.

[0234] S703: Determine the priority weight of each sub-task based on the urgency and resource consumption of each sub-task, and perform task scheduling and resource allocation for each sub-task in the data governance task sequence according to the priority weight to obtain a quality optimization strategy.

[0235] Among them, the urgency represents the time urgency of the sub-task. The resource consumption refers to the computing and human resources required to execute the sub-task. Task scheduling and resource allocation represent the planning and arrangement of sub-task execution order and resource usage.

[0236] Specifically, first, the urgency and resource requirement of each subtask need to be evaluated. Then, the priority weight of the subtask is calculated. Next, the task arrangement and resource allocation are performed according to the weight, and an executable optimization strategy is formed.

[0237] In some embodiments, the strategy can be formulated in various ways.

[0238] Optionally, the formulation is performed by using a weighted scoring method, the urgency and resource consumption of the subtask are weighted and calculated to obtain a comprehensive priority score, the task is sorted and the resource is allocated based on the score, and an optimization strategy is formed.

[0239] Optionally, the formulation is performed by using a multi-objective optimization method, a task priority model is established, time constraints and resource limitations are considered comprehensively, a task execution plan is optimized, and the maximum use efficiency of resources is realized.

[0240] It can be understood that other formulation methods such as heuristic algorithms and constraint programming can also be used to realize the generation of the optimization strategy, which is not limited here.

[0241] Reference Figure 9 The application also provides a data quality governance system 20, which specifically comprises:

[0242] A data acquisition module 21 is configured to acquire various original data, wherein one of the original data is from a data source, and the data source includes a hospital, a company, and a drug production enterprise.

[0243] A data processing module 22 is configured to perform normalization processing on the original data to obtain standardized data.

[0244] A federal data quality evaluation module 23 is configured to input the standardized data into a preset federal data quality evaluation model to perform data quality evaluation, and obtain a quality evaluation result corresponding to each of the standardized data, wherein the quality evaluation result includes an integrity evaluation result, a consistency evaluation result, and a timeliness evaluation result.

[0245] A global quality aggregation module 24 is configured to aggregate the quality evaluation results to obtain a global quality evaluation report.

[0246] A quality optimization and governance module 25 is configured to determine a corresponding quality optimization strategy according to the global quality evaluation report, and to govern data according to the quality optimization strategy.

[0247] Optionally, the data processing module 22 is further configured to construct a medical health knowledge base including disease coding specifications, diagnosis and treatment plan specifications, and drug specification specifications; import each of the original data into a corresponding isolated area; extract key fields of the original data in the isolated area; perform standardization processing on the key fields based on the medical health knowledge base to obtain structured data; and perform security processing on the structured data to obtain the standardized data.

[0248] Optionally, the data processing module 22 is further configured to replace sensitive information in the structured data to obtain desensitized data; perform distributed storage processing on the desensitized data to obtain privacy data; and set a visible range and an operation permission of the privacy data to obtain the standardized data.

[0249] Optionally, the federated data quality evaluation module 23 is further configured to perform integrity evaluation on the standardized data to obtain an integrity evaluation result; perform consistency evaluation on the standardized data to obtain a consistency evaluation result; perform timeliness evaluation on the standardized data to obtain a timeliness evaluation result; and perform fusion processing on the integrity evaluation result, the consistency evaluation result, and the timeliness evaluation result to obtain the quality evaluation result of each data source.

[0250] Optionally, the federated data quality evaluation module 23 is further configured to construct a feature distribution matrix based on the integrity evaluation result, the consistency evaluation result, and the timeliness evaluation result, wherein the rows of the feature distribution matrix correspond to the data sources, the columns correspond to feature values of the integrity evaluation result, the consistency evaluation result, and the timeliness evaluation result, and the feature values are normalized values reflecting quantitative indicators of each evaluation dimension, the evaluation dimensions including an integrity dimension, a consistency dimension, and a timeliness dimension; perform weighted calculation on the feature values of the data sources in the feature distribution matrix based on preset weight values of the evaluation dimensions to obtain a comprehensive quality score of each data source; and obtain the quality evaluation result of each data source based on a preset score benchmark and the comprehensive score.

[0251] Optionally, the federated data quality evaluation module 23 is further configured to compare the number of standard fields of each data source in the standardized data to obtain a quantity consistency comparison result, the standard fields including numerical standard fields and time standard fields; compare the numerical standard fields of each data source in the standardized data to obtain a numerical consistency comparison result; compare the time standard fields of each data source in the standardized data to obtain a time sequence consistency comparison result; and perform weighted calculation on the quantity consistency comparison result, the numerical consistency comparison result, and the time sequence consistency comparison result to obtain the consistency evaluation result.

[0252] Optionally, the quality optimization and governance module 25 is further configured to analyze the global quality evaluation report based on preset evaluation indexes to obtain a data governance task, refine the data governance task based on historical experience and expert rules to obtain a data governance task sequence, the historical experience being a processing mode extracted from a historical data governance success case, the expert rules being data governance best practice specifications summarized by domain experts, and the data governance task sequence including a plurality of subtasks, determine a priority weight of each subtask based on an urgency degree and resource consumption of the subtask, and perform task scheduling and computing resource allocation on the subtasks in the data governance task sequence according to the priority weight to obtain the quality optimization strategy.

[0253] It should be noted that the apparatus provided in the above embodiments is only used as an example to illustrate the division of the above functional modules in realizing the functions thereof, and in actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above-described functions. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be described here.

[0254] The embodiment further discloses an electronic device 900, which refers to Figure 10 The electronic device can include at least one processor 901, at least one communication bus 902, a user interface 903, a network interface 904, and at least one memory 905.

[0255] The communication bus 902 is configured to realize the connection and communication between the components.

[0256] The user interface 903 can include a display screen (Display) and a camera (Camera), and the optional user interface can further include a standard wired interface and a wireless interface.

[0257] The network interface 904 can optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0258] The processor 901 can include one or more processing cores. The processor connects various parts within the server through various interfaces and lines, and performs various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Alternatively, the processor can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor can integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes operating systems, user interfaces, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor, but can be realized by a separate chip.

[0259] The memory 905 can include a random access memory (RAM) and a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory can also be at least one storage device located away from the aforementioned processor. As shown, the memory as a computer storage medium can include an operating system, a network communication module, a user interface module, and an application program of a data quality governance method.

[0260] In Figure 10In the electronic device 900 shown, the user interface 903 is mainly used to provide an interface for the user to input, and obtain data input by the user; and the processor 901 can be used to invoke an application program stored in the memory 905 and storing a data quality governance method, which, when executed by one or more processors, causes the electronic device to perform the method of one or more of the above-described embodiments.

[0261] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0262] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0263] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the units is only a logical function division. There can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different parts can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical or other forms.

[0264] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0265] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0266] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable memory. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned memory includes: a U disk, a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0267] The above-described are only exemplary embodiments of the present disclosure, and cannot limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon considering the disclosure. The present application is intended to cover any variations, uses or adaptive changes of the present disclosure that follow the general principles of the present disclosure and include common knowledge or conventional technical means in the technical field not recorded in the present disclosure. The scope and spirit of the present disclosure are defined by the claims.

Claims

1. A data quality governance method, characterized in that, The method includes: Acquire multiple types of raw data, wherein one type of raw data comes from a data source, and the data source includes hospitals, companies, and pharmaceutical manufacturers; The original data is then normalized to obtain standardized data; The standardized data is subjected to an integrity assessment to obtain the integrity assessment result; The number of standard fields in each data source in the standardized data is compared to obtain the quantity consistency comparison result. The standard fields include numerical standard fields and time standard fields. The numerical standard fields of each data source in the standardized data are compared to obtain the numerical consistency comparison results; The time-type standard fields of each data source in the standardized data are compared to obtain the time-series consistency comparison results. The consistency evaluation result is obtained by weighting the quantity consistency comparison result, the numerical consistency comparison result, and the temporal consistency comparison result; The timeliness of the standardized data is assessed to obtain the timeliness assessment results; The completeness assessment results, the consistency assessment results, and the timeliness assessment results are merged to obtain the quality assessment results of each data source; The quality assessment results are aggregated to obtain a global quality assessment report; Based on the global quality assessment report, a corresponding quality optimization strategy is determined, and the data is managed according to the quality optimization strategy.

2. The method according to claim 1, characterized in that, The process of normalizing the original data to obtain standardized data specifically includes: Construct a medical and health knowledge base, which includes disease coding standards, treatment plan standards, and drug instruction standards; Import each type of raw data into the corresponding isolation area; Extract the key fields of the original data within the isolated area; The key fields are standardized based on the medical and health knowledge base to obtain structured data; The structured data is processed for security purposes to obtain the standardized data.

3. The method according to claim 2, characterized in that, The process of performing security processing on the structured data to obtain the standardized data specifically includes: The sensitive information in the structured data is replaced to obtain desensitized data; The de-identified data is processed through distributed storage to obtain privacy data; The standardized data is obtained by setting the visibility scope and operation permissions of the privacy data.

4. The method according to claim 1, characterized in that, The process of fusing the completeness assessment results, the consistency assessment results, and the timeliness assessment results to obtain the quality assessment results from each data source specifically includes: Based on the completeness assessment results, the consistency assessment results, and the timeliness assessment results, a feature distribution matrix is ​​constructed. The rows of the feature distribution matrix are each data source, and the columns are the feature values ​​corresponding to the completeness assessment results, the consistency assessment results, and the timeliness assessment results. The feature values ​​are normalized values ​​that reflect the quantitative indicators of each assessment dimension. The assessment dimensions include the completeness dimension, the consistency dimension, and the timeliness dimension. Based on the preset weight values ​​of each evaluation dimension, the feature values ​​of each data source in the feature distribution matrix are weighted and calculated to obtain the comprehensive quality score of each data source. Based on the preset scoring benchmark and the comprehensive quality score, the quality assessment results of each data source are obtained.

5. The method according to claim 1, characterized in that, The step of determining the corresponding quality optimization strategy based on the global quality assessment report specifically includes: Based on the preset evaluation indicators, the global quality assessment report is analyzed to obtain data governance tasks; The data governance tasks are refined based on historical experience and expert rules to obtain a data governance task sequence. The historical experience refers to the processing patterns extracted from successful historical data governance cases, and the expert rules are the best practice specifications for data governance summarized by domain experts. The data governance task sequence contains multiple sub-tasks. The priority weight of each subtask is determined based on its urgency and resource consumption. The subtasks in each data governance task sequence are then orchestrated and computational resources are allocated according to the priority weight to obtain the quality optimization strategy.

6. A data quality governance system, characterized in that, For implementing the data quality governance method as described in claim 1, the data quality governance system comprises: The data acquisition module is used to acquire various types of raw data, wherein one type of raw data comes from a data source, and the data source includes hospitals, companies, and pharmaceutical manufacturers. The data processing module is used to perform normalization processing on the raw data to obtain standardized data; The federal data quality assessment module is used to input the standardized data into a preset federal data quality assessment model to conduct data quality assessment and obtain the quality assessment results corresponding to each type of standardized data. The quality assessment results include completeness assessment results, consistency assessment results, and timeliness assessment results. The global quality aggregation module is used to aggregate the various quality assessment results to obtain a global quality assessment report; The quality optimization and governance module is used to determine the corresponding quality optimization strategy based on the global quality assessment report, and to govern the data according to the quality optimization strategy.

7. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Equipment and method for carrying out data quality evaluation on data set based on context

    CN106056287A

  • High-quality data management system based on data management

    CN118897837A