Intelligent data cleaning and optimizing system based on deep learning

Through an intelligent data cleaning system based on deep learning, the problems of synonym recognition and data completion in medical data have been solved, and the efficient construction of high-quality knowledge graphs has been achieved, which has improved the adaptability of medical data and the efficiency of graph construction.

CN120849474APending Publication Date: 2025-10-28HANGZHOU LIUZHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510988833.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Medical data often contains complex technical terms, differing synonyms, unstructured text that is difficult to parse, and numerous missing and outlier values. Traditional cleaning methods cannot accurately identify synonyms or efficiently complete the data, which affects the quality and efficiency of knowledge graph construction.

Method used

An intelligent data cleaning system based on deep learning is adopted, including a knowledge graph determination module, a medical data screening module and a data optimization module. Deep learning algorithms and combination algorithms are used to perform multi-dimensional analysis to screen out high-quality medical data and construct a knowledge graph.

Benefits of technology

It significantly improves data adaptability, enhances the efficiency and quality of knowledge graph construction, and provides high-quality basic data for data analysis and assisted diagnosis in the medical field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849474A_ABST
    Figure CN120849474A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing. The invention provides an intelligent data cleaning and optimizing system based on deep learning. The system comprises a knowledge graph determining module which determines disease types and graph structure information related to a knowledge graph to be constructed this time; the medical data screening module screens out a first medical data set based on disease types, performs suitability analysis on each medical data in the first medical data set by using an analysis model based on map structure information, screens out the medical data with the suitability higher than a threshold value, and forms a second medical data set; and the data optimization module sorts the second medical data set according to the graph structure information, constructs a knowledge graph corresponding to the disease type, associates the second medical data set with the knowledge graph and stores the second medical data set and the knowledge graph in the database. According to the method, the knowledge graph construction efficiency and quality can be remarkably improved, so that high-quality basic data is provided for data analysis, auxiliary diagnosis and the like in the medical field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to an intelligent data cleaning and optimization system based on deep learning. Background Technology

[0002] In the wave of digital transformation in the healthcare field, constructing medical knowledge graphs has become a key technological means to integrate medical data, assist clinical decision-making, and promote medical research. Electronic medical data, such as electronic medical records, medical imaging reports, and laboratory test data, contain rich information such as the patterns of disease occurrence and development and the correlation of treatment plans, making them the core data source for constructing medical knowledge graphs. However, raw medical data has many problems, which seriously affect the quality and efficiency of knowledge graph construction. Therefore, data cleaning is necessary to filter out suitable medical data for knowledge graph construction.

[0003] However, medical data is filled with complex technical terms, and different hospitals and medical staff may use different terms to describe the same symptoms or examinations, such as "stroke" and "apoplexy." Traditional cleaning methods, lacking intelligent semantic understanding capabilities, cannot accurately identify these synonyms, resulting in incomplete extraction of key entity information. When faced with unstructured medical text data, such as electronic versions of doctors' handwritten medical records, manual or simple rule-based algorithms struggle to efficiently parse the entity relationships contained within. Furthermore, the prevalence of missing and outlier values ​​in medical data means that traditional methods can only fill them in by simple deletion or conventional statistical values, failing to accurately complete the data based on potential relationships between data points. This makes it difficult for the cleaned data to support the completeness of entity attributes and relationships in a knowledge graph.

[0004] Therefore, there is an urgent need for an intelligent data cleaning and optimization scheme based on deep learning to screen out high-quality data and achieve efficient knowledge graph construction. Summary of the Invention

[0005] In response, the present invention provides a deep learning-based intelligent data cleaning and optimization system, electronic device, computer storage medium, and computer program product to solve at least one of the above-mentioned technical problems.

[0006] In a first aspect, this invention provides an intelligent data cleaning and optimization system based on deep learning. The system includes a knowledge graph determination module, a medical data filtering module, and a data optimization module. The knowledge graph determination module is used to determine the disease types and graph structure information involved in the knowledge graph to be constructed. The medical data filtering module is used to filter out a first medical data set based on the disease types, and using an analysis model, performs suitability analysis on each medical data in the first medical data set based on the graph structure information, filtering out medical data with suitability higher than a threshold to form a second medical data set. The analysis model is constructed based on a combined algorithm including deep learning algorithms. The data optimization module is used to organize the second medical data set according to the graph structure information, construct a knowledge graph corresponding to the disease types, and store the second medical data set in a database after associating it with the knowledge graph.

[0007] In a third aspect, the present invention provides an electronic device for use in a system as described in any of the preceding claims, characterized in that: the electronic device comprises: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor.

[0008] In a fourth aspect, the present invention provides a computer storage medium for use in a system as described in any of the preceding claims, the computer storage medium storing a computer program executable by a processor.

[0009] In a fifth aspect, the present invention provides a computer program product for use in a system as described in any of the preceding claims, the computer program product comprising a computer program executable by a processor.

[0010] This invention performs multi-dimensional analysis of medical data to ensure that the selected medical data meets the requirements for knowledge graph construction, which can significantly improve data adaptability and enhance the efficiency and quality of knowledge graph construction, thereby providing high-quality basic data for data analysis and assisted diagnosis in the medical field. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of the structure of an intelligent data cleaning and optimization system based on deep learning disclosed in an embodiment of the present invention.

[0013] Figure 2This is a schematic diagram of the structure of the analysis model disclosed in the embodiments of the present invention.

[0014] Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention. Detailed Implementation

[0015] The following specific embodiments illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0016] Furthermore, the technical features involved in the different embodiments of this application described below can be combined with each other as long as they do not conflict with each other.

[0017] In smart city operation scenarios, various smart terminals continuously generate massive amounts of data, which are continuously aggregated into the smart city platform. Existing technologies employ access-based data acquisition and usage management methods, but this approach struggles to guarantee the security of data shared on the platform in the event of computer intrusion. To address this, this invention proposes the following technical solution: Figure 1 As shown in the figure, this invention discloses an intelligent data cleaning and optimization system based on deep learning. The system includes a knowledge graph determination module 100, a medical data filtering module 200, and a data optimization module 300. The knowledge graph determination module 100 is used to determine the disease types and graph structure information involved in the knowledge graph to be constructed.

[0018] This module receives user-inputted knowledge graph construction requests or, based on pre-defined rules, determines the disease types and graph structure information involved in the knowledge graph to be constructed. Disease types include specific disease domains such as "Alzheimer's disease" and "lung cancer." Graph structure information characterizes the overall framework of the knowledge graph, including the entity types to be included in the graph, such as diseases, clinical manifestations, treatments, and drug components; the types of relationships between entities, such as "disease-symptom" relationships and "drug-treatment" relationships; and the hierarchical architecture and data association patterns of the knowledge graph.

[0019] The medical data filtering module 200 is used to filter out a first medical data set based on the disease type, and use an analysis model to perform suitability analysis on each medical data in the first medical data set based on the graph structure information, and filter out medical data with suitability higher than a threshold to form a second medical data set; the analysis model is constructed based on a combined algorithm including deep learning algorithms.

[0020] First, this module uses disease type as the filtering criterion, extracting all medical data related to the disease from a massive amount of raw medical data resources, thus forming the first medical data set. For example, if the disease type is determined to be "hypertension," then data containing relevant content such as hypertension diagnosis records, treatment plans, and blood pressure monitoring data will be filtered out.

[0021] Next, using an analysis model based on a combinatorial algorithm (including deep learning algorithms), a suitability analysis is performed on each piece of medical data in the first medical dataset based on the aforementioned knowledge graph structure information. The suitability analysis needs to comprehensively consider multiple dimensions, including whether the entity and relationship information contained in the data matches the requirements of the knowledge graph structure, the accuracy (e.g., whether medical terminology is standardized), and the completeness (whether key information is missing). By setting a suitability threshold, medical data with a suitability score higher than the threshold are filtered out to form the second medical dataset. Compared to the first medical dataset, the data in the second medical dataset is of higher quality and is more suitable for constructing a knowledge graph.

[0022] The data optimization module 300 is used to sort out the second medical data set according to the graph structure information, construct a knowledge graph corresponding to the disease type, and store the second medical data set in the database after associating it with the knowledge graph.

[0023] Based on the aforementioned knowledge graph structure information, this module uses deep learning algorithms and related data processing techniques to systematically analyze the second medical dataset, deeply extracting entities and relationships from the data. Following pre-defined knowledge graph structure requirements, it classifies, integrates, and associates these entities and relationships to construct a knowledge graph corresponding to the target disease type. For example, in constructing a diabetes knowledge graph, entities such as various symptoms, treatments, and complications of diabetes are connected through relationships such as "cause," "treatment," and "accompanying," forming a structured and visualized knowledge graph.

[0024] Finally, the second medical dataset is linked to the constructed knowledge graph and stored in the database, achieving optimized storage of medical data. This storage method not only facilitates data management and maintenance but also supports subsequent applications based on the knowledge graph, such as medical data analysis, disease diagnosis assistance, and medical research, fully leveraging the value of medical data.

[0025] This invention performs multi-dimensional analysis of medical data to ensure that the selected medical data meets the requirements for knowledge graph construction, which can significantly improve data adaptability and enhance the efficiency and quality of knowledge graph construction, thereby providing high-quality basic data for data analysis and assisted diagnosis in the medical field.

[0026] Furthermore, the knowledge graph determination module 100 is specifically used for: receiving a disease type selection instruction input by a user, wherein the disease type is one or more specific disease domains; receiving several graph structure templates provided by the user, wherein the graph structure templates include a predefined set of entity types, a set of relation types, a hierarchical architecture of entities and relations, and an association pattern; and constructing the graph structure information based on the number of each graph structure template and the specific disease domain.

[0027] Users first input disease type selection instructions, allowing them to choose one or more specific disease domains, such as diabetes or lung cancer, to clarify the construction direction. Next, users import several graph structure templates, which include: a predefined set of entity types, including but not limited to diseases, clinical manifestations, treatments, and drug components; a predefined set of relationship types, including but not limited to "disease-symptom" relationships and "drug-treatment" interaction relationships; and the hierarchical structure and association patterns of entities and relationships.

[0028] After receiving the disease type selection instruction and graph structure template input by the user, the knowledge graph determination module 100 constructs the graph structure information based on the number of each graph structure template and the specific disease domain through combination, adaptation and other methods. For example, a single template for a single disease is directly applied, while multiple templates for multiple diseases are fused to generate a composite structure.

[0029] Furthermore, when the number of specific disease domains is greater than 1, the knowledge graph determination module 100 is specifically used to: acquire the entity type set, relation type set, and feature information of each specific disease domain in each graph structure template; based on the feature information, cluster and integrate the entity types in multiple graph structure templates to generate a unified entity type set; and analyze each relation type set, combine relation types that are related or extensible to form a new relation type set; and construct a hierarchical architecture and association pattern based on the unified entity type set and the new relation type set, combined with the feature information of each specific disease domain, to derive the graph structure information.

[0030] Due to the complexity of medical research and the comprehensive nature of clinical practice, many diseases often share common causes, pathological mechanisms, or clinical correlations. This necessitates the integration and analysis of composite knowledge graphs involving multiple disease domains. For example, diabetic patients often suffer from cardiovascular complications such as hypertension and atherosclerosis. The knowledge graph needs to integrate disease domains such as "diabetes," "hypertension," and "coronary heart disease," linking the pathological pathway of "insulin resistance → vascular endothelial damage → atherosclerosis," as well as the therapeutic correlation of "hypoglycemic drugs → cardiovascular protective effects" (e.g., the cardioprotective effect of GLP-1 receptor agonists). Similarly, cancer patients may experience immunodeficiency or autoimmune reactions (e.g., paraneoplastic syndromes in lung cancer patients). The knowledge graph needs to link entities such as "lung cancer," "immune cells," and "autoantibodies" to analyze the mechanism of "tumor antigen → immune escape → autoimmune dysregulation."

[0031] Therefore, users will import multiple graph structure templates corresponding to the specific disease domains mentioned above. This invention integrates these graph structure templates in the following way to obtain the appropriate graph structure information for the composite knowledge graph: First, the entity type set and relation type set in each of the multiple graph structure templates provided by the user are parsed out. For example, graph structure template A focuses on disease diagnosis, containing entity types such as "disease," "symptom," and "test indicators," as well as relation types such as "disease-symptom" and "disease-test indicators"; graph structure template B focuses on treatment plans, containing entity types such as "disease," "drug," and "surgery," as well as relation types such as "disease-drug" and "disease-surgery."

[0032] For each specific disease domain, the knowledge graph identification module 100 obtains corresponding feature information through medical knowledge base retrieval, historical case analysis, and other methods. For example, for "diabetes", it extracts features such as typical symptoms (polydipsia, polyphagia, and polyuria), complications (retinopathy), and diagnostic indicators (blood glucose levels); for "coronary heart disease", it extracts features such as angina symptoms and coronary artery stenosis.

[0033] Then, the knowledge graph determination module 100 uses natural language processing (NLP) techniques, such as word vector models (e.g., BERT), to calculate the semantic similarity between entity types. For example, "blood glucose level" in graph structure template A and "glycated hemoglobin" in graph structure template B are both diagnostic indicators for diabetes, and semantic analysis determines that they are semantically similar.

[0034] Entity types with semantic similarity exceeding a threshold are clustered and merged into the same entity type. For example, "blood glucose level," "glycated hemoglobin," and "glucose tolerance test" are merged into "diabetes diagnostic indicators." In this way, duplication and ambiguity of entity types in different templates are eliminated, generating a unified set of entity types that includes the commonalities and characteristics of all disease domains.

[0035] Next, the knowledge graph determination module 100 traverses the set of relation types in each graph structure template and analyzes the logical connections between relation types. For example, the "disease-symptom" relation and the "symptom-treatment" relation have a logical chain and can be further extended into a composite relation of "disease-symptom-treatment"; another example is that the "disease-drug" relation and the "drug-side effect" relation can be combined into a "disease-drug-side effect" relation.

[0036] By combining related or expandable relationship types, new relationship types can be formed. At the same time, deep learning algorithms (such as graph neural networks) can be used to mine potential relationships. For example, by discovering the concurrent relationship between "diabetes" and "cardiovascular disease" from a large amount of case data, a new "disease-disease (concurrent)" relationship type can be added, and finally a new set of relationship types can be constructed.

[0037] Finally, based on the knowledge logic and data associations of each disease domain, the hierarchical relationships between entities are determined. For example, "disease" is placed at the top level, "symptoms," "test indicators," and "treatment methods" are placed in the middle layer, and specific "drug names" and "surgical procedures" are placed at the bottom level, forming a tree-like hierarchical structure. At the same time, cross-level association paths are designed for the associations between multiple diseases (such as complication relationships).

[0038] Based on the set of relationship types, define the connection rules and weights (i.e., association patterns) between entities. For example, the "disease-symptom" relationship is a one-to-many connection, and the "drug-treatment" relationship is a many-to-many connection. The confidence level of the relationship is calculated as the weight using historical data statistics or machine learning models.

[0039] Furthermore, by combining hierarchical architecture and association patterns, a graph structure information applicable to multiple disease scenarios is generated, providing a complete framework for subsequent knowledge graph construction.

[0040] Furthermore, if Figure 2 As shown, the analysis model includes an entity matching analysis sub-model, a relationship verification sub-model, a data quality evaluation sub-model, and a fusion processor.

[0041] In this embodiment, the entity matching analysis sub-model embeds a rule-based matching algorithm to perform operations such as string matching, regular expression matching, and domain dictionary matching. Specifically: string matching uses edit distance (LevenshteinDistance) and Jaro-Winkler similarity to determine the text similarity of entity names (e.g., drug names, disease names).

[0042] Regular expression matching: Extract key information from entity attributes (such as test index values) and filter invalid or malformed data.

[0043] Domain dictionary matching: Entity knowledge base is built by combining medical terminology (e.g., ICD-10, SNOMED CT) and standardized matching is achieved through dictionary mapping (e.g., mapping "hypertension" to ICD-10 code I10).

[0044] The relationship verification sub-model is used to verify the authenticity of potential relationships between entities and to predict the completion of missing relationships (e.g., whether the "drug-contraindication" relationship exists). It is preferably built based on a graph neural network (GNN). The data quality evaluation sub-model is based on the analytic hierarchy process (AHP) to assign weights to quality indicators of different dimensions and generate a comprehensive data quality score.

[0045] Furthermore, the use of the analysis model to perform suitability analysis on each medical data in the first medical data set based on the graph structure information includes: the entity matching analysis sub-model extracts entity extraction results from the medical data, performs semantic matching between the entity extraction results and the unified entity type set in the graph structure information, and generates an entity coverage index.

[0046] First, extract entities from medical data (such as electronic medical records and test reports). For example, extract entities such as "type 2 diabetes", "polydipsia", "polyphagia" and "metformin" from "a type 2 diabetic patient has symptoms of polydipsia and polyphagia and is being treated with metformin".

[0047] Next, the extracted entities are semantically compared with the entity type set (e.g., "disease", "symptom", "drug") in the graph structure information. A deep learning model (e.g., BERT) is used to calculate the similarity between the entity text and standard terms (e.g., determining whether "excessive thirst" belongs to the "symptom" category). Here, the entity type set includes a predefined entity type set or a uniform entity type set.

[0048] By calculating the proportion of successfully matched entities to the total number of entities required by the graph, an entity coverage rate metric is generated. This metric reflects the degree to which entities in the medical data conform to the graph structure requirements. For example, if the graph requires three types of entities: "disease," "symptom," and "drug," and all of them match in the data, the entity coverage rate is 100%.

[0049] The relationship verification sub-model detects whether the relationships between entities in medical data cover the new set of relationship types, and performs probability prediction to fill in the missing relationship types, generating a relationship coverage index.

[0050] Check whether the relationships between entities in the medical data cover the predefined relationship types in the graph (e.g., "disease-symptom" or "disease-drug"). For example, check whether there is a "disease-symptom" relationship between "type 2 diabetes" and "polydipsia".

[0051] For relationships that are required by the graph but are not clearly defined in the data (such as "type 2 diabetes - polyphagia"), the entity features can be analyzed using graph neural networks (GNNs) to predict the likelihood of the missing relationship.

[0052] The percentage of detected and predicted relationships out of the total number of relationships required by the graph is calculated. For example, if the graph requires 5 relationships, and 3 relationships exist in the data with 1 relationship predicted for completion, then the relationship coverage rate is 80%.

[0053] The data quality evaluation sub-model performs multi-dimensional verification of the standardization of terminology, consistency of numerical logic, and continuity of timestamps in medical data, and generates a comprehensive data quality score.

[0054] The quality of medical data itself is examined in the following multi-dimensional ways: Terminology standardization: Compare whether the terminology in the medical data conforms to medical standards (e.g., ICD-11, SNOMED CT), for example, "myocardial infarction" should be standardized as "myocardial infarction".

[0055] Numerical logic consistency: Verify whether numerical data (such as blood glucose and blood pressure values) conform to physiological logic. For example, blood glucose values ​​should be within a reasonable range, and the diagnosis time should be earlier than the treatment time.

[0056] Timestamp continuity: Check whether the time sequence of medical records is consistent, for example, the examination report comes first, followed by the diagnosis conclusion.

[0057] Overall score: The above-mentioned indicators are normalized and then weighted. For example, terminology standardization accounts for 40%, numerical logic accounts for 30%, and time continuity accounts for 30%, generating a quality score between 0 and 1 (e.g., 0.9 indicates high quality).

[0058] The fusion processor performs a fusion calculation based on the entity coverage index, the relationship coverage index, and the comprehensive data quality score to determine the suitability of the medical data.

[0059] The fusion unit performs a weighted sum of entity coverage, relationship coverage, and overall data quality score according to preset weights (e.g., 40%, 30%, and 30%, respectively). If the suitability exceeds a preset threshold (e.g., 0.85), the medical data is deemed suitable for constructing a knowledge graph and included in the second medical data set.

[0060] This embodiment analyzes the deep collaboration of various sub-modules of the model and relies on deep learning technology to conduct a comprehensive evaluation of the entities, relationships, and quality of medical data. It accurately judges whether each piece of medical data is suitable for the construction of a knowledge graph, and realizes intelligent cleaning and optimization of medical data, which can significantly improve the adaptability and reliability of the data for knowledge graph construction.

[0061] Furthermore, the method of predicting the probability of completing missing relationship types includes: mapping entities in medical data to graph nodes, mapping the known relationships between entities to directed edges, and converting the attribute features of nodes and edges into vector representations; the graph neural network, based on the set of relationship types in the graph structure information, captures high-order relationship features between entities through multi-layer graph convolution operations, and calculates the probability distribution of the existence of unobserved relationship types; relationship types with predicted probabilities exceeding a threshold are identified as potential completeable relationships and included in the calculation of the relationship coverage index.

[0062] First, the medical data is structured by mapping entities (such as diseases and drugs) to graph nodes, converting the established relationships between entities (such as "disease-treatment") into directed edges, and converting the attributes of nodes and edges (such as disease name and treatment method) into vectors, thus constructing graph structure data that can be processed by graph neural networks.

[0063] Then, based on the set of relationship types in the graph structure information, the graph neural network uses multi-layer graph convolution operations to mine hidden high-order relationship features between entities, and then calculates the probability distribution of the existence of relationships that are not explicitly reflected in the data.

[0064] Finally, relation types with predicted probabilities exceeding a set threshold are identified as potential complementable relations and included in the relation coverage index calculation to ensure that relation prediction results are both accurate and reflect the characteristics of the data source, thereby improving the quality of knowledge graph construction.

[0065] Furthermore, when capturing high-order correlation features between entities through multi-layer graph convolution operations, the scaling factor of the convolutional layer and the attention weight correction factor are adaptively adjusted based on the credibility weight of the data source corresponding to the medical data.

[0066] This invention improves the extraction of the aforementioned high-order correlation features by adaptively adjusting the convolutional layer scaling factor and attention weight correction factor based on the credibility weight of the data source corresponding to the medical data during the process of capturing high-order correlation features in multi-layer graph convolution operations. For example, if the data comes from authoritative medical guidelines, the corresponding credibility weight is high, so the convolutional layer scaling factor is increased to enhance the propagation strength of entity features in the graph structure, while the attention weight correction factor is increased to make the graph neural network prioritize focusing on the potential correlations between entities in the data (such as the "drug-side effect" relationship); if the medical data is ordinary electronic medical records (which may have incomplete records or non-standard terminology), the credibility weight is low, so the graph neural network reduces the aforementioned convolutional layer scaling factor and attention weight correction factor to suppress the feature propagation range of low-quality data and avoid misjudgment of relationships due to noise (such as incorrectly associating unrelated symptoms with diseases).

[0067] Through this adaptive adjustment, the graph neural network can dynamically optimize the prediction logic based on the source characteristics of a single piece of medical data, ensuring the accuracy of relationship completion for data with different levels of credibility.

[0068] The convolutional layer scaling factor and attention weight correction factor can be adjusted, for example, using the following formula: ;in, For the Layer node feature matrix, For neighborhood aggregation functions (such as mean pooling). For convolution kernel weights, For the Layer node feature matrix For the The layer's bias matrix; This is the credibility attenuation coefficient. As a credibility weight, This is the scaling factor of the convolutional layer.

[0069] ;in, For nodes and Attention weights It is determined by the credibility weight The attention weight correction factor is used to enhance (high-confidence medical data) or suppress (low-confidence medical data) the corresponding attention weight. This is the scaling factor coefficient; They are nodes and eigenvectors.

[0070] like Figure 3As shown, embodiments of the present invention also provide an electronic device applied to a system as described in any of the preceding claims, the electronic device comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor.

[0071] This invention also provides a computer storage medium for use in a system as described in any of the preceding claims, the computer storage medium storing a computer program that can be executed by a processor.

[0072] This invention also provides a computer program product for use in a system as described in any of the preceding claims, the computer program product comprising a computer program executable by a processor.

[0073] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0074] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A deep learning-based intelligent data cleaning and optimization system, characterized in that: The system includes a knowledge graph determination module, a medical data filtering module, and a data optimization module. The knowledge graph determination module determines the disease types and graph structure information involved in the knowledge graph to be constructed. The medical data filtering module filters out a first set of medical data based on the disease types, uses an analysis model to perform suitability analysis on each medical data in the first set based on the graph structure information, and filters out medical data with suitability higher than a threshold to form a second set of medical data. The analysis model is constructed based on a combined algorithm including deep learning algorithms. The data optimization module organizes the second set of medical data according to the graph structure information, constructs a knowledge graph corresponding to the disease types, and stores the second set of medical data in a database after associating it with the knowledge graph.

2. The intelligent data cleaning and optimization system based on deep learning according to claim 1, characterized in that: The knowledge graph determination module is specifically used for: receiving a disease type selection instruction input by a user, wherein the disease type is one or more specific disease domains; receiving several graph structure templates provided by the user, wherein the graph structure templates include a predefined set of entity types, a set of relation types, a hierarchical architecture of entities and relations, and an association pattern; and constructing the graph structure information based on the number of each graph structure template and the specific disease domain.

3. The intelligent data cleaning and optimization system based on deep learning according to claim 2, characterized in that: When the number of specific disease domains is greater than 1, the knowledge graph determination module is specifically used to: obtain the entity type set, relation type set, and feature information of each specific disease domain in each graph structure template; and based on the feature information, cluster and integrate the entity types in multiple graph structure templates to generate a unified entity type set. Furthermore, the relationship type sets are analyzed, and the relationship types that are related or expandable are combined to form new relationship type sets; based on the unified entity type set and the new relationship type set, combined with the characteristic information of each specific disease domain, a hierarchical architecture and association pattern are constructed to obtain the graph structure information.

4. The intelligent data cleaning and optimization system based on deep learning according to claim 3, characterized in that: The analysis model includes an entity matching analysis sub-model, a relationship verification sub-model, a data quality evaluation sub-model, and a fusion processor.

5. The intelligent data cleaning and optimization system based on deep learning according to claim 4, characterized in that: Using an analytical model, suitability analysis is performed on each piece of medical data in the first medical dataset based on the graph structure information. This includes: the entity matching analysis sub-model extracts entity extraction results from the medical data, performs semantic matching between the entity extraction results and a unified entity type set in the graph structure information, and generates an entity coverage index; the relationship verification sub-model detects whether the relationships between entities in the medical data cover a new set of relationship types, and performs probability prediction for missing relationship types, generating a relationship coverage index; the data quality evaluation sub-model performs multi-dimensional verification of terminology standardization, numerical logic consistency, and timestamp continuity in the medical data, generating a comprehensive data quality score; and the fusion processor performs a fusion calculation based on the entity coverage index, the relationship coverage index, and the comprehensive data quality score to determine the suitability of the medical data.

6. The intelligent data cleaning and optimization system based on deep learning according to claim 5, characterized in that: The process of predicting the probability of completing missing relationship types includes: mapping entities in medical data to graph nodes, mapping known relationships between entities to directed edges, and converting the attribute features of nodes and edges into vector representations; using a graph neural network based on the set of relationship types in the graph structure information to capture high-order relationship features between entities through multi-layer graph convolution operations, and calculating the probability distribution of the existence of unobserved relationship types; and determining relationship types with predicted probabilities exceeding a threshold as potential complete relationships and incorporating them into the calculation of the relationship coverage index.

7. The intelligent data cleaning and optimization system based on deep learning according to claim 6, characterized in that: When capturing high-order correlation features between entities through multi-layer graph convolution operations, the scaling factor of the convolutional layer and the attention weight correction factor are adaptively adjusted based on the credibility weight of the data source corresponding to the medical data.

8. An electronic device, applied to the system as described in any one of claims 1-7, characterized in that: The electronic device includes: at least one processor, a memory, and a computer program stored in the memory and capable of running on the at least one processor.

9. A computer storage medium, applied to the system as described in any one of claims 1-7, characterized in that: The computer's storage medium stores computer programs that can be executed by a processor.

10. A computer program product, applied to the system as described in any one of claims 1-7, characterized in that: This computer program product contains a computer program that can be executed by a processor.