Data quality treatment method and system, electronic equipment and storage medium
Through the data quality evaluation method that obtains and standardizes the processing from multiple data sources, the problem of incomplete quality reporting caused by a single data source is solved, and more accurate and efficient data quality governance is achieved.
Patent Information
- Application Number
- CN202510554446.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The data source in the prior art is single, resulting in insufficient comprehensive data quality reporting, affecting the effectiveness of data quality governance.
By obtaining a variety of raw data, standardized processing is carried out from multiple data sources such as hospitals, companies and pharmaceutical manufacturers, input the federal data quality assessment model for comprehensive evaluation, generate a global quality assessment report and formulate optimization strategies.
It realizes multi-dimensional evaluation of data quality and comprehensive problem identification, improving the effectiveness and efficiency of data quality governance.
Smart Images

Figure CN120469901A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of privacy computing and data governance technology, and specifically to a data quality governance method, system, electronic device and storage medium. Background Art
[0002] With the deep penetration of the Internet of Things, 5G, and artificial intelligence technologies, global data volumes are surging at an annual rate of 26%. Data formats are also expanding from traditional structured tables to diverse forms such as unstructured images, voice, and real-time streaming data. This explosive growth and increasing reliance on technology have made data quality a critical factor influencing the reliability of decision-making. Low-quality data has already posed significant risks in fields such as finance and healthcare. For example, erroneous data in financial transactions can lead to systemic risks, while errors in the annotation of medical images can affect diagnostic accuracy. Therefore, effective methods to improve data quality are urgently needed.
[0003] In the prior art, the quality of existing data is monitored by scheduling tasks, and quality reports are generated, and data quality is managed based on the quality reports.
[0004] The existing technology has a single data source and relies on a single monitoring data to obtain data quality information, resulting in incomplete quality reports and poor data quality governance effects. Summary of the Invention
[0005] In a first aspect of the present application, a data quality governance method is provided, comprising: obtaining a variety of original data, wherein one type of original data comes from a data source, and the data source includes a hospital, a company, and a pharmaceutical manufacturer; normalizing the original data to obtain standardized data; inputting the standardized data into a preset federal data quality assessment model to perform data quality assessment, and obtaining a quality assessment result corresponding to each type of the standardized data, the quality assessment result including a completeness assessment result, a consistency assessment result, and a timeliness assessment result; aggregating each of the quality assessment results to obtain a global quality assessment report; determining a corresponding quality optimization strategy based on the global quality assessment report, and governing the data based on the quality optimization strategy.
[0006] By adopting the above technical solution, firstly, by obtaining raw data from multiple data sources such as hospitals, companies and pharmaceutical manufacturers, the defect of a single data source in the existing technology is overcome. Secondly, by normalizing the raw data to obtain standardized data, a unified data foundation is laid for subsequent quality assessment. Thirdly, the standardized data is input into the preset federal data quality assessment model, and the data is comprehensively evaluated from the three dimensions of completeness, consistency and timeliness to obtain multi-dimensional quality assessment results, avoiding the problem of insufficient quality reports caused by relying on a single monitoring data in the existing technology. Then, by aggregating the results of each quality assessment, a global quality assessment report containing multiple data sources and multiple assessment dimensions is obtained, making the identification of data quality problems more comprehensive and accurate. Finally, based on the comprehensive quality assessment report, targeted quality optimization strategies are formulated and data governance is carried out, thereby improving the effectiveness of data quality governance.
[0007] Optionally, the normalization processing of the original data to obtain standardized data specifically includes: constructing a medical and health knowledge base, which includes disease coding standards, diagnosis and treatment plan standards and drug description standards; importing each type of original data into a corresponding isolation area; extracting key fields of the original data in the isolation area; standardizing the key fields based on the medical and health knowledge base to obtain structured data; and securely processing the structured data to obtain the standardized data.
[0008] By adopting the above technical solution, we first build a medical and health knowledge base that includes disease coding standards, treatment plan standards, and drug description standards to provide a standard basis for data standardization. Secondly, by importing raw data from different sources into corresponding isolated areas and extracting key fields, we achieve effective data isolation and feature extraction. Thirdly, based on the medical and health knowledge base, we standardize key fields and convert unstructured raw data into structured data, thereby improving the standardization and usability of the data. Finally, we securely process the structured data to obtain standardized data, which not only ensures the standardization of data processing but also ensures the security of data processing, providing a high-quality data foundation for subsequent quality assessment.
[0009] Optionally, the security processing of the structured data to obtain the standardized data specifically includes: replacing sensitive information in the structured data to obtain desensitized data; performing distributed storage processing on the desensitized data to obtain private data; and setting the visible scope and operation permissions of the private data to obtain the standardized data.
[0010] By adopting the above technical solution, we first replace sensitive information in structured data to obtain desensitized data, thus protecting the privacy information in the original data. Secondly, we distribute and process the desensitized data to obtain private data, improving the security and reliability of data storage. Finally, by fine-tuning the visibility and operational permissions of private data to obtain standardized data, we achieve multi-level control of data access, ensuring both data availability and security, and providing a secure and reliable data foundation for data quality assessment.
[0011] Optionally, the standardized data is input into a preset federal data quality assessment model for data quality assessment to obtain a quality assessment result corresponding to each type of standardized data, specifically including: performing an integrity assessment on the standardized data to obtain an integrity assessment result; performing a consistency assessment on the standardized data to obtain a consistency assessment result; performing a timeliness assessment on the standardized data to obtain a timeliness assessment result; and fusing the integrity assessment result, the consistency assessment result and the timeliness assessment result to obtain the quality assessment result of each data source.
[0012] By employing the above technical solution, standardized data is first evaluated along the three dimensions of completeness, consistency, and timeliness, comprehensively covering the core characteristics of data quality. Secondly, by integrating the evaluation results from these three dimensions, a comprehensive quality assessment is obtained, avoiding the potential bias of single-dimensional evaluation. This multi-dimensional evaluation and integration approach can more accurately reflect the data quality status of each data source, providing a reliable decision-making basis for subsequent quality optimization.
[0013] Optionally, the completeness assessment results, the consistency assessment results and the timeliness assessment results are fused to obtain the quality assessment results of each data source, specifically including: constructing a feature distribution matrix based on the completeness assessment results, the consistency assessment results and the timeliness assessment results, the rows of the feature distribution matrix are each data source, and the columns are eigenvalues corresponding to the completeness assessment results, the consistency assessment results and the timeliness assessment results, the eigenvalues are normalized numerical values reflecting the quantitative indicators of each assessment dimension, and the assessment dimensions include the completeness dimension, the consistency dimension and the timeliness dimension; based on the preset weight value of each assessment dimension, the eigenvalue of each data source in the feature distribution matrix is weightedly calculated to obtain a comprehensive quality score of each data source; based on a preset scoring benchmark and the comprehensive score, the quality assessment results of each data source are obtained.
[0014] By employing this technical solution, we first construct a feature distribution matrix, quantifying the performance of each data source across different evaluation dimensions using standardized eigenvalues. Secondly, we introduce preset weights to weight the eigenvalues, reflecting the varying importance of different evaluation dimensions. Finally, we categorize the overall scores based on a preset scoring benchmark, yielding intuitive quality assessment results. This multi-dimensional fusion approach, based on matrix operations, ensures the scientific nature and interpretability of the evaluation process while improving the accuracy and reliability of the results.
[0015] Optionally, the consistency assessment of the standardized data to obtain a consistency assessment result specifically includes: comparing the quantity of standard fields of each data source in the standardized data to obtain a quantity consistency comparison result, the standard fields include numerical standard fields and time standard fields; comparing the numerical standard fields of each data source in the standardized data to obtain a numerical consistency comparison result; comparing the time standard fields of each data source in the standardized data to obtain a time series consistency comparison result; and performing weighted calculation on the quantity consistency comparison result, the numerical consistency comparison result and the time series consistency comparison result to obtain the consistency assessment result.
[0016] By employing the above technical solution, standardized data is first comprehensively compared across three levels: quantitative consistency, numerical consistency, and temporal consistency. Secondly, specialized consistency assessments are performed on numerical and temporal fields, enhancing the specificity of the assessments. Finally, a weighted calculation is performed on the comparison results across the three levels to produce a comprehensive consistency assessment result, which not only takes into account the characteristics of different data types but also ensures the comprehensiveness and accuracy of the assessment results.
[0017] Optionally, the corresponding quality optimization strategy is determined based on the global quality assessment report, specifically including: analyzing the global quality assessment report based on preset evaluation indicators to obtain data governance tasks; refining the data governance tasks based on historical experience and expert rules to obtain a data governance task sequence, the historical experience is the processing mode extracted from historical data governance success cases, the expert rules are the data governance best practice specifications summarized by domain experts, and the data governance task sequence includes multiple subtasks; determining the priority weight of each subtask based on the urgency and resource consumption of each subtask, and performing task scheduling and computing resource allocation for each subtask in the data governance task sequence according to the priority weight to obtain the quality optimization strategy.
[0018] By employing the above technical solution, we first analyzed the global quality assessment report based on pre-set evaluation indicators, identifying specific issues requiring governance. Secondly, by combining historical experience and expert rules, we refined the governance tasks, improving the feasibility and effectiveness of the governance solution. Finally, by prioritizing and allocating resources based on task urgency and resource consumption, we ensured the efficient execution of the governance process. This systematic approach to optimizing strategy development, which draws on historical experience and expert knowledge while also taking into account actual implementation conditions, can significantly improve the efficiency and effectiveness of data governance.
[0019] In a second aspect of the present application, a data quality governance system is provided, specifically comprising: A data acquisition module is used to obtain a variety of raw data, wherein one type of raw data comes from a data source, including a hospital, a company, and a pharmaceutical manufacturer; A data processing module, used for normalizing the raw data to obtain standardized data; A federated data quality assessment module is configured to input the standardized data into a preset federated data quality assessment model to perform data quality assessment, and obtain a quality assessment result corresponding to each type of the standardized data, wherein the quality assessment result includes a completeness assessment result, a consistency assessment result, and a timeliness assessment result; A global quality aggregation module is used to aggregate the quality assessment results to obtain a global quality assessment report; The quality optimization and governance module is used to determine the corresponding quality optimization strategy based on the global quality assessment report and govern the data according to the quality optimization strategy.
[0020] By adopting the above technical solution, firstly, by obtaining raw data from multiple data sources such as hospitals, companies and pharmaceutical manufacturers, the defect of a single data source in the existing technology is overcome. Secondly, by normalizing the raw data to obtain standardized data, a unified data foundation is laid for subsequent quality assessment. Thirdly, the standardized data is input into the preset federal data quality assessment model, and the data is comprehensively evaluated from the three dimensions of completeness, consistency and timeliness to obtain multi-dimensional quality assessment results, avoiding the problem of insufficient quality reports caused by relying on a single monitoring data in the existing technology. Then, by aggregating the results of each quality assessment, a global quality assessment report containing multiple data sources and multiple assessment dimensions is obtained, making the identification of data quality problems more comprehensive and accurate. Finally, based on the comprehensive quality assessment report, targeted quality optimization strategies are formulated and data governance is carried out, thereby significantly improving the effectiveness of data quality governance.
[0021] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs any of the methods described above.
[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions. When the instructions are executed, any one of the methods described above is executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a schematic diagram of the architecture of a data quality management method disclosed in an embodiment of the present application; Figure 2 This is a flow chart of a data quality management method disclosed in an embodiment of the present application; Figure 3 yes Figure 2 A schematic flow chart of a sub-step of step S102; Figure 4 yes Figure 3 A schematic flow chart of a sub-step of step S205; Figure 5 yes Figure 2 A schematic flow chart of a sub-step of step S103; Figure 6 yes Figure 5 A schematic flow chart of a sub-step of step S402; Figure 7 yes Figure 5 A schematic flow chart of a sub-step of step S404; Figure 8 yes Figure 2 A schematic flow chart of a sub-step of step S105; Figure 9 This is a module diagram of a data quality management system provided by an embodiment of the present application; Figure 10 This is a structural diagram of an electronic device disclosed in an embodiment of the present application.
[0024] Explanation of the accompanying drawings: 21. Data acquisition module; 22. Data processing module; 23. Federal data quality assessment module; 24. Global quality assessment module; 25. Quality optimization and governance module; 901. Processor; 902. Communication bus; 903. User interface; 904. Network interface; 905. Memory. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0026] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0027] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0028] Figure 1 An exemplary system architecture 10 is shown to which an embodiment of a data quality governance method or a data quality governance system of the present application can be applied.
[0029] like Figure 1 As shown, the system architecture 10 may include a terminal device 11, a network 12, and a server 13. The network 12 is used to provide a medium for a communication link between the terminal device 11 and the server 12. The network 12 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0030] A user may use a terminal device 11 to interact with a server 13 via a network 12 to receive or send data, etc.
[0031] The terminal device 11 can be either hardware or software. If the terminal device 11 is hardware, it can be any electronic device with a display screen, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. If the terminal device 11 is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software programs or software modules (for example, multiple software programs or software modules used to provide distributed services), or as a single software program or software module. This is not specifically limited here.
[0032] The server 13 may be a server that provides various services, such as a background server that processes data displayed on the terminal device 11. The background server may analyze and process the received data, and may feed back the processing results (such as recognition results) to the terminal device.
[0033] It should be noted that a server can be either hardware or software. When a server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When a server is software, it can be implemented as multiple software programs or software modules (for example, multiple software programs or software modules used to provide distributed services), or as a single software program or software module. This is not specifically limited here.
[0034] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above description is merely illustrative. Any number of terminal devices, networks, and servers may be used as needed. In particular, if target data does not need to be acquired remotely, the above system architecture may not include a network, but may instead include only terminal devices or servers.
[0035] This embodiment discloses a data quality management method. Figure 2 This is a flow chart of a data quality management process disclosed in the embodiment of this application, such as Figure 2 As shown, including step S101 to step S105, the above steps are as follows: S101: Acquire multiple types of original data, wherein one type of original data comes from a data source, and the data source includes a hospital, a company, and a pharmaceutical manufacturer.
[0036] Raw data refers to data in its original form, unprocessed, such as a hospital's electronic medical records or a company's sales data. Raw data originates from a data source, which refers to the entity or system that generates or stores this raw data. This data source can include hospitals, companies, and pharmaceutical manufacturers. For example, if the data source is a hospital, the corresponding raw data might include patient visit records, pharmaceutical sales data, production batch information, and so on.
[0037] Specifically, you first need to determine the type of original data required, such as patient medical records, drug sales data, production batch information, etc.; then establish data connections or obtain authorization with different data sources; then obtain various types of original data through data interfaces, file transfers, etc.; finally, perform a preliminary integrity check on the obtained data to ensure the accuracy and completeness of the data acquisition.
[0038] In some embodiments, the acquisition and collection of raw data can be achieved in a variety of ways.
[0039] Optionally, by establishing API connections with various data sources and obtaining access keys, a data request program is written to define the required data fields. The system automatically calls the API interface to obtain data at preset time intervals and temporarily stores the obtained data in a local database for unified management.
[0040] Optionally, by agreeing on the data export format and field specifications with the data source, the data source will regularly export data files in the specified format, receive the data files through a secure file transfer method, and finally use a data processing program to parse the files and import the data into the system.
[0041] It is understandable that other methods may be used to obtain and collect data, such as direct database access, crawler collection, etc., which are not limited here.
[0042] S102: Normalize the original data to obtain standardized data.
[0043] Among them, standardized data refers to data that has been normalized and has a uniform numerical range and comparability. Normalization refers to the process of converting data into a uniform scale range through specific mathematical transformation methods. Common methods include minimum and maximum normalization, Z-score normalization, etc.
[0044] Specifically, after collecting raw data, we first need to analyze its distribution characteristics and value range and select an appropriate normalization method. We then perform normalization calculations on each feature dimension, mapping the values to a specified range. Finally, we verify the results to ensure that the converted data meets the expected distribution characteristics and statistical properties.
[0045] In some embodiments, the raw data can be normalized in a variety of ways. Optionally, the data can be processed using a min-max normalization method, which first calculates the maximum and minimum values of each feature in the raw data, then converts all data to the interval [0, 1] according to a mapping formula, and finally checks the converted data and saves the conversion parameters for subsequent use.
[0046] Optionally, use the Z-score normalization method for data processing, calculate the mean and standard deviation of the original data, transform the data into a form that conforms to the standard normal distribution, verify the transformation effect and record relevant parameters.
[0047] It is understandable that other normalization methods such as decimal scaling normalization, logarithmic transformation, etc. can also be used to achieve data standardization processing, which is not limited here.
[0048] Reference Figure 3 , Figure 3This embodiment of the present application provides Figure 2 A schematic flow chart of a sub-step of step S102 includes steps S201 to S205, which are as follows: S201: Build a medical and health knowledge base, which includes disease coding standards, diagnosis and treatment plan standards, and drug description standards.
[0049] Among them, the medical and health knowledge base represents a systematically organized data set of professional knowledge in the medical and health field. Disease coding standards refer to a standardized system for uniformly coding and classifying various diseases, such as the International Classification of Diseases (ICD). Treatment plan standards represent standardized diagnostic and treatment process guidelines for different diseases. Drug instructions standards are used to present standardized information such as drug instructions, indications, usage, dosage, and precautions.
[0050] Specifically, we first need to collect various healthcare-related specifications and standards from authoritative medical institutions and data sources. We then organize, categorize, and structure this collected information to establish relationships between diseases, treatment plans, and medications. Finally, we import this processed data into a knowledge base system, perform data verification and maintenance, and ensure the accuracy and completeness of the knowledge base.
[0051] In some embodiments, the construction of a medical and health knowledge base can be achieved through various methods. Optionally, the knowledge base can be constructed using a relational database storage method, by establishing data tables for diseases, treatment plans, and drugs, designing relationships between tables, importing standardized medical data, and establishing indexes to improve query efficiency, and finally implementing knowledge base management and maintenance functions.
[0052] Optionally, use graph database technology to build a medical knowledge graph, taking diseases, treatment plans, and drugs as entity nodes, establishing relationship edges between entities, forming a complete knowledge network structure, and providing intelligent knowledge retrieval and reasoning capabilities.
[0053] It is understandable that other data storage and management methods such as document databases, hybrid storage, etc. can also be used to build a medical and health knowledge base, which is not limited here.
[0054] S202: Import each type of original data into a corresponding isolation area.
[0055] Among them, the isolation area refers to the independent storage and processing space set up for data from different sources.
[0056] Specifically, you first need to divide different isolation zones based on data type and security level, and configure corresponding access rights and security policies for each zone. Next, establish data classification rules to identify the characteristics and attributes of each type of raw data. Based on the classification results, import the data into the corresponding isolation zone, performing a data integrity check during the import process. Finally, confirm the data import results to ensure that the data in each isolation zone meets the expected requirements.
[0057] In some embodiments, the import of the isolated area of original data can be achieved in multiple ways.
[0058] Optionally, physical isolation can be adopted by establishing independent data partitions on different servers or storage devices, using dedicated data transmission channels and encryption mechanisms to securely import various types of original data into corresponding physically isolated areas, and implementing strict access control and audit mechanisms.
[0059] Optionally, logical isolation can be used to establish independent database instances or tablespaces in the same storage system, set different data schemas and access permissions, and achieve logically isolated storage and management of various types of data.
[0060] It is understandable that other isolation methods such as containerized isolation and hybrid isolation can also be used to achieve secure isolated storage of original data, which is not limited here. S203: Extract key fields of the original data in the isolated area.
[0061] Among them, key fields represent core data attributes that are important for data analysis and processing, such as patient basic information, diagnosis results, treatment plans, etc.
[0062] Specifically, first determine the scope and type of key fields based on business requirements and analysis objectives. Then, use field mapping rules to identify target fields within the raw data in each isolated area. Next, extract and validate field values to ensure the integrity and accuracy of the extracted data. Finally, organize the extracted key fields into a structured format for easier processing and analysis.
[0063] In some embodiments, key field extraction can be achieved through a variety of methods. Optionally, field extraction can be performed using a rule-based configuration approach. By pre-defining field extraction rules and mapping relationships, the original data is scanned and matched to identify and extract key field content that meets the rules, while also performing data quality verification and exception handling.
[0064] Optionally, use intelligent parsing technology for field extraction. Through natural language processing and pattern recognition algorithms, it automatically analyzes the structural characteristics of the original data, identifies important field information, and performs intelligent extraction and annotation.
[0065] It is understandable that other field extraction methods such as template matching and deep learning can also be used to achieve accurate extraction of key data, which is not limited here. S204: Standardize key fields based on the medical and health knowledge base to obtain structured data.
[0066] Among them, structured data refers to the data form with a fixed format and unified standards after normalization processing. Normalization processing is to convert data in different formats and expressions into a unified standard.
[0067] Specifically, the extracted key fields must first be matched against the standards and specifications in the healthcare knowledge base to identify the field type and corresponding standards and specifications. Based on the matching results, the content of each field is then standardized, including terminology unification, code mapping, and format conversion. The converted data is then verified to ensure compliance with the standards and specifications. Finally, the standardized data is organized into a structured format, and relationships between fields are established.
[0068] In some embodiments, standardization of key fields can be achieved in a variety of ways.
[0069] Optionally, standardization can be performed using rule mapping. By establishing a mapping dictionary between field values and standard specifications, key fields are matched and converted one by one, abnormal values and non-standard expressions are processed, and finally structured data that meets the specifications is generated.
[0070] Optionally, natural language processing technology is used for standardization. Through semantic analysis and entity recognition algorithms, the actual meaning of field content is understood, automatically mapped to standard terms and codes in the knowledge base, and intelligent structured conversion is performed.
[0071] It is understandable that other standardized processing methods such as machine learning classification, expert systems, etc. can also be used to achieve the normalization and structuring of key fields, which is not limited here.
[0072] S205: Perform security processing on the structured data to obtain standardized data.
[0073] Standardized data refers to data that has undergone secure processing and is ready for subsequent analysis. Secure processing involves desensitizing and encrypting data to protect privacy. Structured data refers to data with a fixed format and uniform standards.
[0074] Specifically, the first step is to identify sensitive information fields within structured data and determine the scope of data requiring security processing. Based on the type of sensitive information, appropriate security processing methods are selected for conversion. The processed data is then integrity-checked to ensure data usability is not compromised. Finally, the securely processed data is saved in a standardized format.
[0075] In some embodiments, secure processing of structured data may be achieved in a variety of ways.
[0076] Optionally, data desensitization can be used for security processing, and sensitive information can be obscured or deformed through replacement, masking, truncation and other technologies to protect privacy information while maintaining the value of data analysis, such as replacing specific names with random strings.
[0077] Optionally, use data encryption for security processing, encrypt and convert sensitive data through cryptographic algorithms to ensure the security of data during storage and transmission, and provide necessary decryption mechanisms to support subsequent use.
[0078] It is understandable that other security processing methods such as data classification and access control can also be used to achieve privacy protection of structured data, which is not limited here.
[0079] Reference Figure 4 , Figure 4 This embodiment of the present application provides Figure 3 A schematic flow chart of a sub-step of step S205 includes steps S301 to S303, which are as follows: S301: Replace sensitive information in structured data to obtain desensitized data.
[0080] Sensitive information refers to data that may involve privacy and security, such as personal identity information, contact information, etc. Desensitized data refers to safe and usable data after sensitive information has been replaced.
[0081] Specifically, you first need to identify sensitive information fields within structured data and determine the scope and type of data to be replaced. Based on the different types of sensitive information, you then select corresponding replacement rules and methods for conversion. Next, you verify the replaced data to ensure it complies with data format requirements. Finally, save the replaced data in a desensitized format.
[0082] In some embodiments, replacement of sensitive information can be achieved in a variety of ways.
[0083] Optionally, desensitization can be performed by replacing part or all of the sensitive information with specific characters (such as numbers), thereby achieving privacy protection while maintaining the basic characteristics of the data. For example, replacing the middle four digits of a mobile phone number with ***.
[0084] Optionally, use random value replacement for desensitization. By generating a random string or number that complies with the rules to replace the original sensitive information, the distribution characteristics and analytical value of the data are maintained, while ensuring that the sensitive information cannot be restored.
[0085] It is understandable that other replacement methods such as mapping conversion, interval replacement, etc. can also be used to achieve desensitization of sensitive information, which is not limited here.
[0086] S302: Perform distributed storage processing on the desensitized data to obtain private data.
[0087] Among them, distributed storage processing refers to a technical method of distributing and storing data in multiple independent nodes.
[0088] Specifically, the desensitized data must first be sharded and partitioned, with a data distribution strategy and storage node allocation scheme determined. The partitioned data is then stored on different nodes, with data indexes and associations established. Consistency checks are then performed on the distributed data to ensure data integrity and accessibility. Finally, a data access control mechanism is established to ensure privacy protection.
[0089] In some embodiments, distributed storage processing can be implemented in a variety of ways.
[0090] Optionally, data sharding can be used for storage. By dividing the complete data set into multiple data shards according to specific rules and storing them on different physical nodes, the risk of data leakage can be reduced and data access efficiency can be improved.
[0091] Optionally, data replication is used for storage. By saving data copies on multiple nodes, data disaster recovery and load balancing can be achieved, and a differentiated access permission control mechanism can be used to protect data security.
[0092] It is understandable that other storage processing methods such as hierarchical storage, heterogeneous storage, etc. can also be used to achieve privacy protection of desensitized data, which is not limited here.
[0093] S303: Setting the visibility range and operation permissions of the private data to obtain standardized data.
[0094] The visible scope represents the range of users or systems allowed to access the data. Operation permissions represent the authorization level for operations such as reading and modifying data. Standardized data represents standardized data after permission management.
[0095] Specifically, data access policies must first be defined to determine the visibility and operational permissions for different roles. Access control rules must then be established for private data, mapping user identities to permissions. Next, permission verification mechanisms must be implemented to ensure that data access complies with authorization requirements. Finally, permission settings must be associated with the data to create standardized data.
[0096] In some embodiments, permission setting can be implemented in a variety of ways.
[0097] Optionally, a role-based access control method can be used to establish a correspondence between roles and permissions, and assign corresponding roles to different users to achieve refined management of data access and operations. For example, a doctor role can be set to view patient medical records, while a nurse role can only view basic information.
[0098] Optionally, use multi-level authorization for control. By setting a hierarchical structure for data access, differentiated permission control can be implemented at different levels to ensure the secure flow of data within the authorized scope, while supporting dynamic adjustment of permissions.
[0099] It is understandable that other permission control methods such as attribute encryption, time control, etc. can also be used to achieve secure management of privacy data, which is not limited here.
[0100] S103: Input the standardized data into a preset federal data quality assessment model to perform data quality assessment to obtain quality assessment results corresponding to each type of standardized data. The quality assessment results include completeness assessment results, consistency assessment results, and timeliness assessment results.
[0101] Data quality assessment refers to a quantitative analysis of data integrity, consistency, and timeliness. The pre-defined federated data quality assessment model is an analytical model used to assess the quality of standardized data. The quality assessment results represent specific metrics for measuring data quality.
[0102] Specifically, standardized data is first input into the evaluation model, and data analysis is performed based on pre-set evaluation metrics. Next, evaluation scores are calculated for completeness, consistency, and timeliness. The evaluation results for each dimension are summarized to generate a quality assessment report. Finally, data quality is graded and labeled based on the evaluation results.
[0103] In some embodiments, the assessment model may include multiple assessment dimensions.
[0104] Optionally, the completeness assessment evaluates the completeness and validity of the data and identifies missing data and anomalies by calculating indicators such as the proportion of non-null values in data fields and the coverage of valid values.
[0105] Optionally, consistency assessment evaluates the standardization and uniformity of data and detects data conflicts and inconsistencies by analyzing characteristics such as data format standardization and value range rationality.
[0106] Optionally, timeliness evaluation evaluates the real-time and time validity of data by using characteristics such as statistical data update frequency and timestamp distribution to determine whether the data meets timeliness requirements.
[0107] It is understandable that other evaluation dimensions such as accuracy, uniqueness, etc. can also be used to achieve a comprehensive evaluation of data quality, which is not limited here.
[0108] Reference Figure 5 , Figure 5 This embodiment of the present application provides Figure 2 A schematic flow chart of a sub-step of step S103 includes steps S401 to S404, which are as follows: S401: Perform integrity assessment on the standardized data to obtain integrity assessment results.
[0109] The integrity assessment result is the result obtained after the data integrity assessment. The integrity assessment means analyzing and evaluating the filling status and validity of the data fields.
[0110] Specifically, we first need to statistically analyze the data distribution of each field in the standardized data, calculating the field fill rate and the proportion of valid values. We then evaluate the degree of missing data and the level of validity against pre-set completeness standards. We then generate a field-level completeness score and perform a weighted calculation to obtain the overall completeness assessment result.
[0111] In some embodiments, integrity assessment can be achieved in a variety of ways.
[0112] Optionally, missing value analysis can be used for evaluation. By calculating the proportion of null values and invalid values in the data set, the location and extent of missing data can be identified, the data filling completeness can be evaluated, and a statistical report on missing data can be generated.
[0113] Optionally, use a valid value detection method to perform the evaluation, verify whether the data value meets the preset format specifications and value range, calculate the proportion of valid data, evaluate the quality and integrity of the data, and identify abnormal data items.
[0114] It is understandable that other evaluation methods such as data distribution analysis, outlier detection, etc. can also be used to implement integrity assessment, which is not limited here.
[0115] S402: Perform consistency assessment on the standardized data to obtain a consistency assessment result.
[0116] The consistency assessment result refers to the result obtained after the data consistency assessment. The consistency assessment means the inspection and analysis of the standardization and uniformity of the data.
[0117] Specifically, the standardized data is first checked for format compliance of each field and the compliance of the data values. The field values are then compared against the pre-set data standards to assess their standardization and consistency. Field-level consistency scores are then calculated and summarized to generate the overall consistency assessment results.
[0118] In some embodiments, consistency assessment can be achieved in a variety of ways.
[0119] Optionally, the format specification detection method can be used for evaluation, by verifying whether the data's expression format, measurement unit, coding rules, etc. comply with unified standards, calculating the normative indicators, and evaluating the data's format consistency level.
[0120] Optionally, use logical relationship verification to conduct the evaluation, by analyzing the dependencies and business rules between data items, checking the logical rationality of the data, and evaluating the semantic consistency level of the data.
[0121] It is understandable that other evaluation methods such as data conflict detection, standard compliance analysis, etc. can also be used to implement consistency evaluation, which is not limited here.
[0122] Reference Figure 6 , Figure 6 This embodiment of the present application provides Figure 5 A schematic flow chart of a sub-step of step S402 includes steps S501 to S504, which are as follows: S501: Compare the quantity of standard fields of each data source in the standardized data to obtain a quantity consistency comparison result. The standard fields include numerical standard fields and time standard fields.
[0123] The quantity consistency comparison results represent the comparison of the number of fields in each data source. Standard fields refer to numerical and time data attributes that comply with the standards.
[0124] Specifically, we first need to count the number of standard numeric and time fields in each data source. We then cross-compare the field counts across different data sources to calculate the degree of difference. We then generate a quantitative consistency analysis, including the distribution of field counts and difference statistics.
[0125] In some embodiments, quantitative identity comparison can be achieved in a variety of ways.
[0126] Optionally, a field counting method is used for comparison. By counting the specific number of numeric and time fields in each data source, the ratio and difference of the field numbers are calculated to evaluate the quantitative consistency level between data sources.
[0127] Optionally, use a difference analysis method to perform the comparison, identify the differences in the number of fields between data sources, analyze the causes of the differences, evaluate the completeness of field coverage, and generate a difference distribution report.
[0128] It is understandable that other comparison methods such as structure comparison, mapping analysis, etc. can also be used to achieve consistency assessment of the number of standard fields, which is not limited here.
[0129] S502: Compare the numerical standard fields of each data source in the standardized data to obtain a numerical consistency comparison result.
[0130] Numeric standard fields represent normalized data attributes expressed in numerical form. Numeric consistency comparison results represent the comparison of the numerical content of each data source.
[0131] Specifically, we first need to extract the corresponding numerical standard fields from each data source. We then perform a statistical analysis of the field's value range and distribution characteristics. We then compare the degree of difference in values between different data sources and calculate a similarity index. Finally, we summarize the comparisons for each field to generate the numerical consistency analysis results.
[0132] In some embodiments, numerical consistency comparison can be achieved in a variety of ways.
[0133] Optionally, a statistical feature comparison method is used to calculate statistical indicators such as the mean, variance, and distribution interval of the numerical fields of each data source, compare the central tendency and degree of dispersion of the numerical values, and evaluate the consistency level of the numerical distribution.
[0134] Optionally, use the exact value comparison method to directly compare the specific values of the corresponding fields, calculate the value deviation rate and matching degree, identify outliers and inconsistencies, and evaluate the accuracy and consistency of the values.
[0135] It is understandable that other comparison methods such as correlation analysis, trend comparison, etc. can also be used to implement consistency assessment of numerical fields, which is not limited here.
[0136] S503: Compare the time type standard fields of each data source in the standardized data to obtain a temporal consistency comparison result.
[0137] The time type standard field represents the normalized data attribute expressed in time format. The time series consistency comparison result represents the comparison of the time series of each data source.
[0138] Specifically, we first need to extract the corresponding standard time fields from each data source. We then analyze the continuity and periodicity of the time series. We then compare the differences in time distribution between different data sources and calculate the time series similarity. Finally, we summarize the comparisons of each field to generate the time series consistency analysis results.
[0139] In some embodiments, timing consistency comparison can be achieved in various ways.
[0140] Optionally, a time span comparison method is used to evaluate the continuity and completeness of time coverage and identify missing and anomalies in time series data by analyzing the start and end time, time interval, sampling frequency and other characteristics of the time fields of each data source.
[0141] Optionally, use the time series pattern comparison method to evaluate the consistency level of time series data and discover differences in time series changes by comparing pattern features such as change trends, periodic characteristics, and mutation points of the time series.
[0142] It is understandable that other comparison methods such as timestamp alignment, time window analysis, etc. can also be used to implement consistency evaluation of time-type fields, which is not limited here.
[0143] S504: Perform weighted calculation on the quantity consistency comparison results, the numerical consistency comparison results, and the time sequence consistency comparison results to obtain a consistency evaluation result.
[0144] Among them, the consistency assessment result represents the final scoring result of data consistency, and consistency includes consistency indicators in three dimensions: quantity, value, and time series.
[0145] Specifically, we first need to determine the weight coefficients for each comparison result to reflect the importance of different dimensions. We then standardize the results to achieve a consistent scoring scale. We then perform a weighted sum calculation based on the weights to obtain a comprehensive evaluation score. Finally, we generate a detailed consistency assessment report.
[0146] In some embodiments, weighted calculation can be implemented in a variety of ways.
[0147] Optionally, a linear weighted method is used for calculation. By setting the weight coefficient of the evaluation results of each dimension, the standardized scores are weighted and summed to obtain a comprehensive score reflecting the overall consistency level.
[0148] Optionally, a hierarchical weighted approach is used for calculation. By establishing a multi-level weight system and performing weighted operations at different levels, the evaluation results can be summarized step by step to obtain a more refined consistency score.
[0149] It is understandable that other calculation methods such as fuzzy comprehensive evaluation, hierarchical analysis, etc. can also be used to achieve weighted calculation of consistency assessment results, which is not limited here.
[0150] S403: Perform timeliness evaluation on the standardized data to obtain a timeliness evaluation result.
[0151] The timeliness evaluation result refers to the result obtained after the data timeliness evaluation. Timeliness evaluation refers to the analysis of the time validity and update timeliness of the data.
[0152] Specifically, we first need to analyze the time attributes of the standardized data, the generation time of the statistical data, and the update frequency. Then, we evaluate the time validity and update timeliness of the data based on pre-set timeliness standards. We then calculate the timeliness scores for each data item and summarize them to generate an overall timeliness evaluation result.
[0153] In some embodiments, timeliness evaluation can be achieved in a variety of ways.
[0154] Optionally, a time decay method is used for evaluation. By calculating the interval between the data and the current time, a time decay function is set to evaluate the timeliness level of the data and identify expired data items and data items to be updated.
[0155] Optionally, use update frequency analysis to evaluate the timeliness of data maintenance and determine whether data updates meet timeliness requirements based on characteristics such as the update time interval and update rules of statistical data.
[0156] It is understandable that other evaluation methods such as life cycle analysis, timeliness threshold detection, etc. can also be used to achieve timeliness evaluation, which is not limited here.
[0157] S404: Fusing the completeness assessment results, consistency assessment results, and timeliness assessment results to obtain quality assessment results for each data source.
[0158] Among them, the quality assessment result represents the comprehensive assessment result of the overall quality level of the data source. The overall quality includes indicators in three dimensions: completeness, consistency and timeliness.
[0159] Specifically, we first need to determine the weighting of the evaluation results for each dimension to reflect the importance of the evaluation indicator. We then normalize the evaluation results across different dimensions to unify the scoring criteria. We then perform a weighted fusion calculation based on the weights to obtain a comprehensive quality score for the data source.
[0160] In some embodiments, the fusion process can be implemented in a variety of ways.
[0161] Optionally, a linear fusion method is used for processing, and by setting weight coefficients for the completeness, consistency, and timeliness evaluation results, the normalized scores are weighted and summed to obtain a comprehensive score reflecting the overall quality level of the data source.
[0162] Optionally, a multi-layer fusion approach is used for processing. By establishing a hierarchical evaluation indicator system and integrating the scores at different levels, the evaluation results can be summarized step by step to obtain a more comprehensive quality score.
[0163] It is understandable that other fusion methods such as fuzzy comprehensive evaluation, evidence theory, etc. can also be used to achieve the fusion processing of quality assessment results, which is not limited here.
[0164] Reference Figure 7 , Figure 7 This embodiment of the present application provides Figure 5 A schematic flow chart of a sub-step of step S404 includes steps S601 to S603, which are as follows: S601: Construct a feature distribution matrix based on the completeness assessment results, consistency assessment results, and timeliness assessment results. The rows of the feature distribution matrix are the data sources, and the columns are the eigenvalues corresponding to the completeness assessment results, consistency assessment results, and timeliness assessment results. The eigenvalues are normalized numerical values reflecting the quantitative indicators of each assessment dimension. The assessment dimensions include the completeness dimension, the consistency dimension, and the timeliness dimension.
[0165] The feature distribution matrix represents the relationship between data sources and evaluation dimensions. Eigenvalues are normalized quantitative indicators for each evaluation dimension. Evaluation dimensions encompass integrity, consistency, and timeliness.
[0166] Specifically, we first need to extract the raw scores of each data source across different evaluation dimensions. We then normalize these scores, converting the indicators across different dimensions into eigenvalues on a unified scale. Next, we construct a matrix structure, with the data sources as rows, the evaluation dimensions as columns, and the eigenvalues as matrix elements.
[0167] In some embodiments, matrix construction can be achieved in a variety of ways.
[0168] Optionally, a linear normalization method is used to construct a matrix, and the evaluation results of each dimension are converted into eigenvalues in the interval [0, 1] by normalizing the original evaluation scores to their maximum and minimum values, and filling them into the corresponding matrix positions.
[0169] Optionally, a standardized processing method is used to construct a matrix, and the raw scores are converted into standardized eigenvalues by calculating the mean and standard deviation of the evaluation scores to ensure that the indicators of different dimensions are comparable.
[0170] It is understandable that other construction methods such as quantile transformation, logarithmic transformation, etc. can also be used to generate the feature distribution matrix, which is not limited here.
[0171] S602: Based on the preset weight value of each evaluation dimension, perform weighted calculation on the eigenvalue of each data source in the feature distribution matrix.
[0172] Specifically, we first need to obtain the preset weights for the completeness, consistency, and timeliness dimensions. We then multiply these weights by the eigenvalues of the corresponding dimensions in the feature distribution matrix. We then aggregate the weighted eigenvalues for each data source to obtain a comprehensive score for that data source.
[0173] In some embodiments, weighted calculation can be implemented in a variety of ways.
[0174] Optionally, vector multiplication is used for calculation, by constructing the preset weight value as a weight vector, and performing vector multiplication operation with the eigenvalue of each row in the feature distribution matrix to obtain the weighted result of each data source.
[0175] Optionally, matrix operations are used for calculation, by constructing the preset weight values into a weight matrix, and performing matrix multiplication operations with the feature distribution matrix to obtain the weighted results of all data sources in batches.
[0176] It is understandable that other calculation methods such as hierarchical weighting, dynamic weighting, etc. can also be used to implement weighted calculation of eigenvalues, which is not limited here.
[0177] S603: Based on the preset scoring benchmark and comprehensive scoring, the quality assessment results of each data source are obtained.
[0178] The quality assessment result represents the final quality level determination of each data source.
[0179] Specifically, we first need to obtain a preset scoring benchmark and define the score ranges for different quality levels. We then compare the comprehensive score of each data source with the scoring benchmark. We then determine the quality level of the data source based on the score ranges and generate a quality assessment result.
[0180] In some embodiments, the evaluation result determination can be achieved in various ways.
[0181] Optionally, a threshold division method is used for determination, and by setting a score threshold for the quality level, the comprehensive score of the data source is mapped to a corresponding quality level, such as excellent, good, qualified, and unqualified levels.
[0182] Optionally, an interval mapping method is used for determination. By subdividing the scoring interval into multiple sub-intervals, a corresponding relationship between the score and the quality level is established to achieve a refined assessment of the quality of the data source.
[0183] It is understandable that other determination methods such as fuzzy evaluation, cluster analysis, etc. can also be used to determine the quality assessment results, which is not limited here.
[0184] S104: Aggregate the quality assessment results to obtain a global quality assessment report.
[0185] Among them, the global quality assessment report is a comprehensive analysis report of the quality assessment results of each data source.
[0186] Specifically, we first need to collect the quality assessment results and detailed scores for each data source. We then compile statistics on the distribution of different quality levels and analyze the common characteristics of quality issues. We then generate statistical charts and analysis of data quality to form a global quality assessment report.
[0187] In some embodiments, aggregation of evaluation results can be achieved in a variety of ways.
[0188] Optionally, statistical analysis can be used for aggregation to calculate the quantity distribution and proportion of each quality level, evaluate the overall data quality level, and identify the main types and impact scope of quality issues.
[0189] Optionally, use multidimensional analysis to perform aggregation, and cross-analyze the quality assessment results from multiple dimensions such as completeness, consistency, and timeliness to discover the correlation characteristics and distribution patterns of quality problems.
[0190] It is understandable that other aggregation methods such as hierarchical aggregation, trend analysis, etc. can also be used to generate a global quality assessment report, which is not limited here.
[0191] S105: Determine the corresponding quality optimization strategy based on the global quality assessment report and manage data according to the quality optimization strategy.
[0192] Specifically, we first need to analyze the quality issues and distribution characteristics in the global quality assessment report. Then, we develop corresponding optimization strategies based on the type, severity, and scope of impact of the issues. We then implement data governance according to these strategies to continuously improve data quality.
[0193] In some embodiments, quality optimization can be achieved in a variety of ways.
[0194] Optionally, a hierarchical governance approach can be adopted for optimization, by prioritizing quality issues, formulating hierarchical treatment plans, and gradually implementing quality improvements according to the severity of the problems and the scope of their impact.
[0195] Optionally, use classified governance methods for optimization, formulate special governance measures for quality issues in different dimensions such as completeness, consistency, and timeliness, and implement targeted data quality improvements.
[0196] It is understandable that other optimization methods such as process reengineering and standardization construction can also be used to achieve continuous optimization of data quality, which is not limited here.
[0197] Reference Figure 8 , Figure 8 This is a schematic flow chart of a sub-step of step S105 in FIG. 0 provided in an embodiment of the present application, including steps S701 to S703, which are as follows: S701: Analyze the global quality assessment report based on the preset evaluation indicators to obtain data governance tasks.
[0198] Among them, data governance tasks represent specific work items that need to be performed to improve data quality.
[0199] Specifically, we first need to obtain a pre-defined evaluation indicator system. Then, we analyze the issues in the global quality assessment report based on these indicators. We then transform the identified issues into actionable governance tasks, clarifying the task objectives and completion standards.
[0200] In some embodiments, governance tasking can be achieved in a variety of ways.
[0201] Optionally, an indicator decomposition method can be used to formulate the plan, by breaking down the evaluation indicators into specific governance requirements, identifying the data items and processing methods that need to be improved, and forming an executable list of governance tasks.
[0202] Optionally, a problem-oriented approach can be used to classify and sort out the problems in the quality assessment report, determine the priorities and solutions for problem improvements, and convert them into clear governance tasks.
[0203] It is understandable that other formulation methods such as goal decomposition and process optimization can also be used to generate data governance tasks, which are not limited here.
[0204] S702: Refine data governance tasks based on historical experience and expert rules to obtain a data governance task sequence. Historical experience refers to the processing patterns extracted from historical data governance success cases, and expert rules refer to the data governance best practice specifications summarized by domain experts. The data governance task sequence contains multiple subtasks.
[0205] Historical experience represents general approaches drawn from successful governance cases. Expert rules refer to governance experience and best practices summarized by domain experts. Subtasks represent specific work items resulting from the decomposition of governance tasks.
[0206] Specifically, we first need to collect historical governance cases and expert rule bases. Then, we break down and refine governance tasks based on experience and rules. Then, we organize subtasks into task sequences based on execution order and dependencies.
[0207] In some embodiments, task refinement can be achieved in a variety of ways.
[0208] Optionally, refinement can be performed using an experience model approach, by extracting processing patterns and methods from historical successful cases, optimizing and adjusting them in combination with the current scenario, and forming an executable subtask sequence.
[0209] Optionally, use rule mapping to refine the process. By applying the best practices summarized by experts, governance tasks can be broken down into specific operational steps and checkpoints to generate a standardized task sequence.
[0210] It is understandable that other detailed methods such as template reuse and scenario analysis can also be used to achieve the serialization of governance tasks, which is not limited here.
[0211] S703: Determine the priority weight of each subtask based on the urgency and resource consumption of each subtask, and perform task scheduling and computing resource allocation for each subtask in the data governance task sequence according to the priority weight to obtain a quality optimization strategy.
[0212] Urgency represents the time urgency of a subtask. Resource consumption refers to the computational and human resources required to execute a subtask. Task scheduling and resource allocation represent the planning and arrangement of the subtask execution sequence and resource usage.
[0213] Specifically, we first need to assess the urgency and resource requirements of each subtask. We then calculate the subtask priority weights. We then schedule tasks and allocate resources based on these weights, forming an executable optimization strategy.
[0214] In some embodiments, policy formulation can be achieved in a variety of ways.
[0215] Optionally, a weighted scoring method is used to formulate the priority, by weighting the urgency and resource consumption of the subtasks to obtain a comprehensive priority score, and task sorting and resource allocation are performed based on the score to form an optimization strategy.
[0216] Optionally, a multi-objective optimization approach can be used to formulate the plan. By establishing a task priority model, taking into account time constraints and resource limitations, the task execution plan can be optimized to maximize resource utilization efficiency.
[0217] It is understandable that other formulation methods such as heuristic algorithms, constraint programming, etc. can also be used to generate optimization strategies, which are not limited here.
[0218] refer to Figure 9 , the present application also provides a data quality governance system 20, specifically including: The data acquisition module 21 is used to obtain a variety of raw data, wherein one type of raw data comes from a data source, including a hospital, a company, and a pharmaceutical manufacturer; The data processing module 22 is used to normalize the raw data to obtain standardized data; A federated data quality assessment module 23 is configured to input the standardized data into a preset federated data quality assessment model to perform data quality assessment, and obtain a quality assessment result corresponding to each type of the standardized data, wherein the quality assessment result includes a completeness assessment result, a consistency assessment result, and a timeliness assessment result; A global quality aggregation module 24 is configured to aggregate the quality assessment results to obtain a global quality assessment report; The quality optimization and management module 25 is used to determine a corresponding quality optimization strategy based on the global quality assessment report and manage data according to the quality optimization strategy.
[0219] Optionally, the data processing module 22 is also used to construct a medical health knowledge base, which includes disease coding specifications, diagnosis and treatment plan specifications, and drug description specifications; import each type of original data into a corresponding isolation area; extract key fields of the original data in the isolation area; standardize the key fields based on the medical health knowledge base to obtain structured data; and securely process the structured data to obtain the standardized data.
[0220] Optionally, the data processing module 22 is also used to replace sensitive information in the structured data to obtain desensitized data; perform distributed storage processing on the desensitized data to obtain private data; and set the visible scope and operation permissions of the private data to obtain the standardized data.
[0221] Optionally, the federated data quality assessment module 23 is also used to perform integrity assessment on the standardized data to obtain an integrity assessment result; perform consistency assessment on the standardized data to obtain a consistency assessment result; perform timeliness assessment on the standardized data to obtain a timeliness assessment result; and fuse the integrity assessment result, the consistency assessment result, and the timeliness assessment result to obtain the quality assessment result of each data source.
[0222] Optionally, the federated data quality assessment module 23 is further used to construct a feature distribution matrix based on the completeness assessment results, the consistency assessment results, and the timeliness assessment results, wherein the rows of the feature distribution matrix are each data source, and the columns are the eigenvalues corresponding to the completeness assessment results, the consistency assessment results, and the timeliness assessment results, and the eigenvalues are normalized numerical values reflecting the quantitative indicators of each assessment dimension, and the assessment dimensions include the completeness dimension, the consistency dimension, and the timeliness dimension; based on the preset weight values of each assessment dimension, the eigenvalues of the data sources in the feature distribution matrix are weightedly calculated to obtain the comprehensive quality score of the data sources; based on the preset scoring benchmark and the comprehensive score, the quality assessment results of the data sources are obtained.
[0223] Optionally, the federal data quality assessment module 23 is also used to compare the quantity of standard fields of each data source in the standardized data to obtain a quantity consistency comparison result, where the standard fields include numerical standard fields and time standard fields; compare the numerical standard fields of each data source in the standardized data to obtain a numerical consistency comparison result; compare the time standard fields of each data source in the standardized data to obtain a time series consistency comparison result; and perform weighted calculation on the quantity consistency comparison result, the numerical consistency comparison result and the time series consistency comparison result to obtain the consistency assessment result.
[0224] Optionally, the quality optimization and governance module 25 is also used to analyze the global quality assessment report based on preset evaluation indicators to obtain data governance tasks; refine the data governance tasks based on historical experience and expert rules to obtain a data governance task sequence, the historical experience is the processing mode extracted from historical data governance success cases, and the expert rules are the data governance best practice specifications summarized by domain experts, and the data governance task sequence includes multiple subtasks; determine the priority weight of each subtask based on the urgency and resource consumption of each subtask, and perform task scheduling and computing resource allocation for each subtask in the data governance task sequence according to the priority weight to obtain the quality optimization strategy.
[0225] It should be noted that the above embodiments provide devices that implement their functions using only the division of the above functional modules as examples. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0226] This embodiment also discloses an electronic device 900, referring to Figure 10The electronic device may include: at least one processor 901 , at least one communication bus 902 , a user interface 903 , a network interface 904 , and at least one memory 905 .
[0227] The communication bus 902 is used to implement connection and communication between these components.
[0228] The user interface 903 may include a display screen (Display) and a camera (Camera). Optional user interfaces may also include a standard wired interface and a wireless interface.
[0229] The network interface 904 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0230] The processor 901 may include one or more processing cores. The processor utilizes various interfaces and circuits to connect various components within the server. It executes instructions, programs, code sets, or instruction sets stored in memory, and accesses data stored in memory to perform various server functions and process data. Optionally, the processor may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display; and the modem handles wireless communications. It is understood that the modem may also be implemented as a separate chip, rather than integrated into the processor.
[0231] Memory 905 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory includes non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing each of the above-mentioned method embodiments, etc.; the data storage area may store data involved in each of the above-mentioned method embodiments, etc. The memory may also optionally be at least one storage device located remotely from the aforementioned processor. As shown in the figure, the memory, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a data quality management method.
[0232] exist Figure 10 In the electronic device 900 shown, the user interface 903 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 901 can be used to call an application program for a data quality management method stored in the memory 905. When executed by one or more processors, the electronic device executes one or more methods in the above embodiments.
[0233] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.
[0234] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0235] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic, such as the division of units, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0236] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0237] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0238] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.
[0239] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any variations, uses or adaptive changes of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the technical field that are not recorded in the present disclosure. The description and examples are to be regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A data quality management method, characterized in that: The method comprises: Acquire multiple types of original data, wherein one type of original data comes from a data source, including a hospital, a company, and a pharmaceutical manufacturer; performing normalization processing on the original data to obtain standardized data; Inputting the standardized data into a preset federated data quality assessment model to perform data quality assessment, and obtaining a quality assessment result corresponding to each type of the standardized data, wherein the quality assessment result includes a completeness assessment result, a consistency assessment result, and a timeliness assessment result; Aggregating the quality assessment results to obtain a global quality assessment report; According to the global quality assessment report, a corresponding quality optimization strategy is determined, and data is governed according to the quality optimization strategy.
2. The method according to claim 1, characterized in that The normalizing process of the original data to obtain standardized data specifically includes: Building a medical and health knowledge base, which includes disease coding standards, diagnosis and treatment plan standards, and drug description standards; Importing each type of raw data into a corresponding isolated area; Extracting key fields of the original data in the isolated area; Standardizing the key fields based on the medical and health knowledge base to obtain structured data; The structured data is securely processed to obtain the standardized data.
3. The method according to claim 2, characterized in that The securely processing the structured data to obtain the standardized data specifically includes: Replacing sensitive information in the structured data to obtain desensitized data; Performing distributed storage processing on the desensitized data to obtain private data; The visibility range and operation authority of the private data are set to obtain the standardized data.
4. The method according to claim 1, wherein The step of inputting the standardized data into a preset federated data quality assessment model to perform data quality assessment to obtain a quality assessment result corresponding to each type of the standardized data specifically includes: Performing integrity assessment on the standardized data to obtain an integrity assessment result; Performing consistency evaluation on the standardized data to obtain a consistency evaluation result; Performing a timeliness evaluation on the standardized data to obtain a timeliness evaluation result; The integrity assessment result, the consistency assessment result, and the timeliness assessment result are fused to obtain the quality assessment result of each data source.
5. The method according to claim 4, characterized in that The fusing of the completeness assessment result, the consistency assessment result, and the timeliness assessment result to obtain the quality assessment result of each data source specifically includes: Constructing a feature distribution matrix based on the completeness assessment result, the consistency assessment result, and the timeliness assessment result, wherein the rows of the feature distribution matrix are the data sources, and the columns are the eigenvalues corresponding to the completeness assessment result, the consistency assessment result, and the timeliness assessment result, wherein the eigenvalues are normalized numerical values reflecting the quantitative indicators of each assessment dimension, wherein the assessment dimensions include the completeness dimension, the consistency dimension, and the timeliness dimension; Based on the preset weight values of each evaluation dimension, weighted calculation is performed on the characteristic values of each data source in the characteristic distribution matrix to obtain a comprehensive quality score of each data source; Based on a preset scoring benchmark and the comprehensive score, the quality assessment results of the respective data sources are obtained.
6. The method according to claim 4, characterized in that The performing consistency assessment on the standardized data to obtain a consistency assessment result specifically includes: Comparing the quantity of standard fields of each data source in the standardized data to obtain a quantity consistency comparison result, wherein the standard fields include numerical standard fields and time standard fields; Comparing the numerical standard fields of each data source in the standardized data to obtain a numerical consistency comparison result; Comparing the time type standard fields of each data source in the standardized data to obtain a time series consistency comparison result; The quantity consistency comparison results, the numerical consistency comparison results and the time sequence consistency comparison results are weightedly calculated to obtain the consistency evaluation result.
7. The method according to claim 1, characterized in that Determining a corresponding quality optimization strategy based on the global quality assessment report specifically includes: Analyze the global quality assessment report based on preset evaluation indicators to obtain data governance tasks; The data governance tasks are refined based on historical experience and expert rules to obtain a data governance task sequence. The historical experience is the processing mode extracted from historical data governance success cases, and the expert rules are the data governance best practice specifications summarized by domain experts. The data governance task sequence includes multiple subtasks. The priority weight of each subtask is determined based on the urgency and resource consumption of each subtask, and task scheduling and computing resource allocation are performed on the subtasks in each data governance task sequence according to the priority weight to obtain the quality optimization strategy.
8. A data quality management system, characterized in that: include: A data acquisition module is used to obtain a variety of raw data, wherein one type of raw data comes from a data source, including a hospital, a company, and a pharmaceutical manufacturer; A data processing module, used for normalizing the raw data to obtain standardized data; A federated data quality assessment module is configured to input the standardized data into a preset federated data quality assessment model to perform data quality assessment, and obtain a quality assessment result corresponding to each type of the standardized data, wherein the quality assessment result includes a completeness assessment result, a consistency assessment result, and a timeliness assessment result; A global quality aggregation module is used to aggregate the quality assessment results to obtain a global quality assessment report; The quality optimization and governance module is used to determine the corresponding quality optimization strategy based on the global quality assessment report and govern the data according to the quality optimization strategy.
9. An electronic device, characterized in that: The electronic device comprises a processor, a memory, a user interface and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Equipment and method for carrying out data quality evaluation on data set based on context
CN106056287A
High-quality data management system based on data management
CN118897837A
Data quality management method and system, storage medium and electronic equipment
CN119127856A
Federal learning-based perception system evaluation optimization method and system, medium, product and terminal
CN119249482A
Data quality enrichment integration and evaluation system
US20080235288A1
Cited By
Smart city multi-source data fusion quality evaluation and encryption method and system
CN120849404A
A method and system for quality assessment and encryption of multi-source data fusion in smart cities
CN120849404B
Cross-technical route collaborative decision-making method and system of new energy automobile power system
CN121094343A
Clinical test digital standard data automatic conversion system
CN121255912A
Hydraulic power plant GIS equipment production operation data quality evaluation method and system
CN121327614A