Data management system and method based on NLP
Through the NLP-based data management system, the lack of information extraction and retrieval in electronic medical record data management is solved, and the rapid and accurate disease information retrieval and medical information correlation are achieved, which improves the standardization and readability of the data.
Patent Information
- Application Number
- CN202510215684.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art lacks depth and accuracy in electronic medical record data management, making it difficult to extract key medical information, and the data management system lacks systematicity and efficiency, making it impossible to quickly retrieve and analyze valuable information.
Through an NLP-based data management system, unstructured electronic medical record data is obtained using the data interface, disease keywords are extracted after preprocessing, rule databases are built for keyword matching, classification results are optimized and classification results are built in combination with the decision tree model, and knowledge graphs are constructed, and the graph neural network is used to optimize modeling to evaluate the timeliness and credibility of diagnostic information.
It realizes rapid and accurate retrieval of specific disease medical records, clearly presents medical information correlation, standardizes the use of terms, and improves data accuracy and readability.
Smart Images

Figure CN120336540A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language processing and data management, and specifically to a data management system and method based on NLP. Background Art
[0002] With the rapid advancement of medical informatization and the extensive application of big data technology in the medical field, the amount of electronic medical record data has increased explosively, playing an important role in all aspects of the medical industry, covering multiple aspects such as clinical diagnosis, medical research, and medical management. In the process of continuous improvement of the medical system, fully exploring the value of electronic medical record data and improving the quality of medical services and scientific research level have become key requirements for the development of the medical field. The management and analysis of electronic medical record data face challenges in many aspects such as data processing difficulty and application effect.
[0003] In the current field of electronic medical record data management and analysis, traditional data processing and analysis methods expose many deficiencies when dealing with the increasing amount of data and complex application requirements. First of all, the processing of electronic medical record data lacks depth and accuracy. Traditional methods often only perform simple storage and basic statistical analysis on data, and fail to fully utilize natural language processing to deeply mine unstructured medical record texts, making it difficult to extract key medical information, including potential associations of diseases and effectiveness evaluation of treatment plans. Secondly, in the management and application of electronic medical record data, there is a lack of systematicness and efficiency. Traditional data management systems usually store different types of medical record data in isolation, without establishing an effective association and integration mechanism, and cannot dynamically adjust the organization and analysis methods of data according to changes and requirements in medical scenarios. When facing a large amount of complex medical record data, it is difficult to quickly retrieve and analyze valuable information. Summary of the Invention
[0004] The purpose of the present invention is to provide a data management system and method based on NLP to solve the problems raised in the prior art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A data management method based on NLP, the method includes the following steps:
[0006] Obtain unstructured electronic medical record data through a data interface and perform data preprocessing;
[0007] Extract disease keywords and analyze the data according to the visit time, construct a rule base for keyword matching, optimize the classification results using a decision tree model, and evaluate the classification quality;
[0008] Extract relevant entities from unstructured medical record data through NLP technology, extract the relationships between relevant entities, combine medical entities and relationships to construct a knowledge graph, calculate the similarity and causal relationships between entities by analyzing the structure of the graph nodes and edges, and optimize and model the knowledge graph in combination with graph neural networks;
[0009] Sort the electronic medical records and extract time features, evaluate the timeliness of diagnostic information and annotate it, and at the same time evaluate the credibility of the medical record data in the knowledge graph.
[0010] Obtain unstructured electronic medical record data through a data interface, including text data on patient basic information, diagnosis records, treatment records, and drug usage;
[0011] Preprocess the obtained unstructured electronic medical record data, including denoising, removing irrelevant information, redundant punctuation marks, spelling mistakes, duplicate content, and non-standard terms. Use a pre-trained medical term conversion model and combine it with a medical term library to convert non-standard terms into medical terms or standard expressions. At the same time, identify and remove invalid or irrelevant text parts;
[0012] Take the Unified Medical Language System as the core term library, combine the clinical term specifications within the hospital, establish term conversion rules, and use string matching algorithms and semantic understanding models to convert synonyms and abbreviations in the medical record data into standard medical terms;
[0013] Assign labels to the data for preliminary classification.
[0014] Assign labels to the data for preliminary classification. The specific steps are as follows:
[0015] Extract high-frequency words closely related to various diseases as disease keywords from medical literature, medical term libraries, and historical medical record data. According to the visit time, divide the data by year and season, and analyze the time distribution law of diseases;
[0016] Construct a disease keyword matching rule library, and use the edit distance algorithm to match the text in the medical record data with the keywords;
[0017] The matching of the text in the medical record data with the keywords through the edit distance algorithm includes:
[0018] Extract high-frequency words closely related to various diseases as disease keywords from medical authoritative literature, standard medical terminology databases, and a large amount of historical medical record data. Determine the screening weights of the keywords based on the diagnostic criteria and common symptom factors of the diseases. Store the keywords using a hash table, construct a disease keyword matching rule library, and establish a regular update mechanism. According to the latest research results in the medical field and newly entered medical record data, conduct a comprehensive review and update of the rule library once in a while, add new keywords in a timely manner, and delete keywords that are no longer applicable;
[0019] For the text content in the medical record data, split it according to sentences, phrases, or medical terms, and use the edit distance algorithm to calculate the edit distance between each text segment and the keywords in the rule library. When calculating the edit distance, adopt the dynamic programming algorithm to improve the calculation efficiency. Through the analysis and testing of medical text data and combined with medical professional knowledge, set the edit distance threshold for exact matching. When the edit distance between the text segment and the keyword is the set edit distance threshold for exact matching, it is determined as an exact match; set the edit distance threshold for fuzzy matching. When the edit distance is greater than the set edit distance threshold for exact matching and less than or equal to the set edit distance threshold for fuzzy matching, it is determined as a fuzzy match. Among them, the edit distance thresholds for exact matching and fuzzy matching are determined through the analysis of historical medical record data;
[0020] Introduce a word vector model. For the text segments and keywords that are successfully matched by the edit distance algorithm, further use the word vector model to calculate their semantic similarity. When the semantic similarity is lower than the set semantic similarity threshold, re-evaluate the matching results. For the case of fuzzy matching, give priority to referring to the matching results with higher semantic similarity, and select the first s matching results with reference semantic similarity;
[0021] Collect a large amount of accurately classified medical record data as a training set, select features related to disease classification, including symptom descriptions, examination results, and medication information, and use the Gini coefficient method to select the optimal features for the construction and training of decision trees;
[0022] The Gini coefficient method includes:
[0023] Collect accurately classified medical record data as a training set and preprocess the unstructured data in the training set;
[0024] For each feature in the training set, calculate its Gini coefficient. Given that there are n samples and k categories in the training set, and the number of samples in the i-th category in the dataset is n i , then the calculation formula for the Gini coefficient Gini(D) of the dataset is as follows:
[0025]
[0026] Among them, n represents the number of all unstructured electronic medical records used for model training, k represents the number of disease categories included in the unstructured electronic medical records, n i represents the number of samples of the i-th category in the dataset, and Gini(D) represents the Gini coefficient of the dataset D
[0027] For each feature, calculate the Gini coefficient after dividing the training subset, and perform a weighted sum according to the number of samples in the subset to obtain the Gini coefficient Gini(D,A) after dividing the training subset. The calculation formula is as follows:
[0028]
[0029] Among them, m represents the number of subsets obtained by dividing according to the feature A, |D j | represents the number of samples in the j-th subset, Gini(D j ) represents the Gini coefficient of the j-th subset, and Gini(D,A) represents the Gini coefficient after dividing the training subset based on the feature A;
[0030] Compare the Gini coefficients of all features, select the feature with the smallest Gini coefficient as the splitting feature of the current node, divide the training set according to the selected optimal feature to generate child nodes. For each child node, repeat the process of calculating the Gini coefficient, selecting the optimal feature and performing the division, recursively construct a decision tree. During the construction process, set the stopping condition. When all samples in the child node belong to the same category, there are no more features available for splitting, or the preset maximum tree depth is reached, stop splitting the node, mark the node as a leaf node, and determine the disease category it belongs to according to the category distribution of the samples in the node;
[0031] Use the constructed decision tree model to train the training set, so that the model learns the relationship between features and disease classification. During the training process, divide the training set into q subsets, and take one subset as the validation set in turn, and other subsets as the training set, perform multiple trainings and validations on the model, and adjust the minimum number of samples and the maximum tree depth parameters of the leaf nodes of the decision tree;
[0032] Randomly select a certain proportion (10%) of the data from the processed data for manual spot-checking to check whether the term conversion is correct and the classification is accurate; compare the processed data with the known standard dataset, calculate the accuracy index, and evaluate the quality of data processing. The specific steps are as follows:
[0033] Adopt the stratified random sampling method, stratify according to dimensions such as disease type, visit time, department, etc., and randomly select samples within each layer;
[0034] Manually and carefully check whether the terms in the extracted samples are accurately converted into standard medical terms and whether the classification results conform to the actual situation. For the problems found, record in detail the error types and the relevant information of the medical records (patient ID, visit time, etc.);
[0035] Improve the model update mechanism, which is triggered when the new data volume reaches 20% of the original training set or the model accuracy drops by more than 5%; the update process includes collecting new medical record data, re-cleaning and selecting features, and then retraining the decision tree model with the updated data to replace the original model to improve the classification performance.
[0036] Based on the unstructured data in the medical records, extract and mark the relevant entities in the electronic medical records through NLP technology. Among them, the relevant entities include disease entities, symptom entities, treatment plan entities, as well as time and dose attributes;
[0037] Based on the extracted entity information, use syntactic analysis, rule matching, and deep learning models for relation extraction, and construct a subject-relation-object triple;
[0038] Combined with the extracted medical entities and relations, use the entities as nodes in the constructed knowledge graph and the relations between entity nodes as edges in the knowledge graph to construct the knowledge graph of the electronic medical records;
[0039] Combined with the graph neural network GNN, optimize the modeling of the knowledge graph to capture the complex entities and relations in the medical records. The specific steps are as follows:
[0040] Initialize the features for each node, and convert the text description of the node into a vector representation through a word vector model as the initial feature;
[0041] Use the graph convolutional network GCN to perform convolutional processing on the nodes and edges in the knowledge graph. Each node aggregates information from its neighbor nodes to update its features. The calculation formula is as follows:
[0042]
[0043] Among them, h v (l+1) represents the feature of the node v at the l + 1 layer, N(v) represents the set of neighbor nodes of the node v, c vu represents the weight of the edge between the node v and its neighbor node u, c u represents the normalization coefficient, W (l) represents the weight matrix at the l layer, and σ represents the activation function;
[0044] Through the multi-layer convolutional operations of the GNN, capture and model the relationships between nodes, and establish the connections between diseases, drugs, and treatment plans through multi-hop information propagation; meanwhile, utilize time series information to capture the dynamic relationships of nodes over time.
[0045] Based on the constructed knowledge graph, apply the ARIMA time series analysis algorithm to sort the electronic medical records in chronological order, extract time-related features. For the diagnosis information, with reference to the release time of the latest medical research results and diagnosis and treatment guidelines, calculate the time difference between the medical record diagnosis time and the latest standard time, and evaluate the degree of timeliness decay of the diagnosis information in the current medical decision-making through the time series analysis model. According to the analysis results, label the timeliness of the medical record information, which is divided into three levels: "high timeliness", "medium timeliness", and "low timeliness".
[0046] In the constructed knowledge graph, conduct credibility assessment from aspects such as data source, data consistency, and compliance with medical knowledge. For the data source, give priority to trusting the medical record data released by authoritative medical institutions and well-known medical research institutions; in terms of data consistency, check whether the information in different parts of the medical record is contradictory; utilize the medical knowledge in the knowledge graph to judge whether the diagnosis and treatment plan in the medical record conform to the current medical cognition.
[0047] A data management system based on NLP, the system includes a data acquisition module, a data classification module, a knowledge graph construction module, and a graph neural network optimization module. The data acquisition module is used to obtain unstructured electronic medical record data through a data interface and perform data preprocessing; the data classification module is used to analyze the data by extracting disease keywords and according to the visit time, construct a rule base for keyword matching, optimize the classification results using a decision tree model, and evaluate the classification quality; the knowledge graph construction module is used to extract relevant entities from unstructured medical record data through NLP technology, extract the relationships between relevant entities, and combine medical entities and relationships to construct a knowledge graph; the graph neural network optimization module is used to calculate the similarity and causal relationship between entities by analyzing the structure of the graph nodes and edges, and optimize and model the knowledge graph in combination with the graph neural network; sort the electronic medical records and extract time features, evaluate the timeliness of the diagnosis information and make annotations, and at the same time conduct credibility assessment on the medical record data in the knowledge graph.
[0048] The data acquisition module includes an interface connection unit and a preprocessing unit. The interface connection unit is used to establish a data interface with an external data source to obtain unstructured electronic medical record text data including patient basic information, diagnosis records, treatment records, and medication usage. The preprocessing unit is used to preprocess the obtained unstructured electronic medical record data, extract disease keywords from medical literature, a thesaurus, and historical medical record data, analyze the disease time distribution pattern based on the visit time; construct a disease keyword matching rule base, use the edit distance algorithm combined with a word vector model for text matching to initially classify the data; collect the classified medical record data to train a decision tree model to optimize the classification result and achieve the initial classification of the data.
[0049] The data classification module includes a feature selection unit, a decision tree construction and training unit, and a classification evaluation unit. The feature selection unit is used to select disease classification-related features from the obtained unstructured electronic medical record text data and select the optimal features using the Gini coefficient method; The decision tree construction and training unit is used to recursively construct a decision tree according to the optimal features selected by the Gini coefficient and set stopping conditions; during the training process, cross-validation is performed to adjust the model parameters; The classification evaluation unit is used to extract data using the stratified random sampling method for manual spot-checking and comparison with a standard data set, and calculate the accuracy index to evaluate the data processing quality.
[0050] The knowledge graph construction module includes an entity extraction unit, a node construction unit, and an edge construction unit. The entity extraction unit is used to extract disease entities, symptom entities, treatment plan entities, and time and dosage attributes based on the unstructured data in the medical record using NLP technology and mark them. The node construction unit is used to construct the extracted entities into nodes in the graph. The edge construction unit establishes the edges in the graph according to the relationships between the entities to form the connection structure of the knowledge graph.
[0051] The graph neural network optimization module includes a graph convolution operation unit and a quality management unit. The graph convolution operation unit is used to optimize and model the knowledge graph in combination with the graph neural network GNN to capture complex entities and relationships in the medical record; The quality management unit is used to perform timeliness evaluation and contribution degree evaluation on the data in the knowledge graph.
[0052] Compared with the prior art, the beneficial effects of the present invention are:
[0053] 1. Extract disease keywords from multi-source data, analyze the disease time distribution pattern in combination with the visit time, construct a rule base and use the edit distance algorithm and word vector model for keyword matching, and at the same time use the Gini coefficient method to select the optimal features to construct and train a decision tree model, which helps doctors retrieve medical records of specific diseases more quickly and accurately;
[0054] 2. By leveraging NLP technology to extract entities and relationships from electronic medical records, constructing a knowledge graph, and optimizing the modeling through graph neural networks, the complex associations between medical information such as diseases, symptoms, and treatment plans can be presented more clearly. Additionally, the similarity and causal relationships between entities can be calculated by analyzing the structure of the nodes and edges in the graph;
[0055] 3. Using a pre-trained medical term conversion model in combination with a medical term library to denoise unstructured electronic medical record data, converting non-standard terms into standard expressions, unifying synonyms and abbreviations, making the use of terms in electronic medical record data more standardized, reducing the understanding barriers caused by inconsistent terms, and improving the accuracy and readability of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a schematic flow diagram of a data management method based on NLP according to the present invention;
[0057] Figure 2 It is a schematic structural diagram of a data management system based on NLP according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] In the embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution, a data management method based on NLP, and the method includes the following steps:
[0060] Obtain unstructured electronic medical record data through a data interface and perform data preprocessing;
[0061] Extract disease keywords and analyze the data according to the visit time, construct a rule library for keyword matching, optimize the classification results using a decision tree model, and evaluate the classification quality;
[0062] Extract relevant entities from unstructured medical record data through NLP technology, extract the relationships between relevant entities, combine medical entities and relationships to construct a knowledge graph, calculate the similarity and causal relationships between entities by analyzing the structure of the nodes and edges in the graph, and optimize the modeling of the knowledge graph in combination with graph neural networks;
[0063] Sort the electronic medical records and extract time features, evaluate the timeliness of diagnostic information and perform annotation, and at the same time evaluate the credibility of the medical record data in the knowledge graph.
[0064] Obtain unstructured electronic medical record data through a data interface, including text data on patients' basic information, diagnosis records, treatment records, and medication usage;
[0065] Preprocess the obtained unstructured electronic medical record data, including denoising, removing irrelevant information, redundant punctuation, spelling mistakes, duplicate content, and non-standard terms. Use a pre-trained medical term conversion model and combine it with a medical term library to convert non-standard terms into medical terms or standard expressions. At the same time, identify and remove invalid or irrelevant text parts;
[0066] Take the Unified Medical Language System as the core term library, combine it with the clinical term specifications within the hospital, establish term conversion rules, and use string matching algorithms and semantic understanding models to convert synonyms and abbreviations in the medical record data into standard medical terms;
[0067] Assign labels to the data for preliminary classification.
[0068] Specifically, obtain 1,000 unstructured electronic medical record data from the information system of a certain tertiary hospital. These medical records cover multiple departments such as internal medicine, surgery, and gynecology and obstetrics, and include patients' basic information, diagnosis records, treatment records, and medication usage. The content of one medical record is: "Patient's name, male, 56 years old. Coughing and expectorating in the past week, accompanied by low fever. After examination, there is a shadow in the lungs, diagnosed with pneumonia. Given amoxicillin capsules 0.5g, orally 3 times a day for one week.";
[0069] In the data preprocessing stage, remove irrelevant information in the medical records, such as doctors' personal signatures, some unnecessary repeated examination descriptions, etc.; remove redundant punctuation through regular expressions, and correct the redundant punctuation in "Coughing, expectorating, accompanied by low fever..." to "Coughing, expectorating, accompanied by low fever."; use a spelling check tool to check and correct possible spelling mistakes;
[0070] Use a pre-trained medical term conversion model and a medical term library to convert the non-standard expression "lung inflammation" of "pneumonia" into a standard term; based on the Unified Medical Language System and the clinical term specifications within the hospital, combine string matching algorithms and semantic understanding models to convert the abbreviation "amo capsules" of "amoxicillin capsules" into a standard expression.
[0071] Assign labels to the data for preliminary classification. The specific steps are as follows:
[0072] Extract high-frequency words closely related to various diseases as disease keywords from medical literature, medical term libraries, and historical medical record data. According to the visit time, divide the data by year and season, and analyze the time distribution law of diseases;
[0073] Construct a disease keyword matching rule base, and use the edit distance algorithm to match the text in the medical record data with the keywords;
[0074] The matching of the text in the medical record data with the keywords by the edit distance algorithm includes:
[0075] Extract high-frequency words closely related to various diseases as disease keywords from medical authoritative literature, standard medical terminology libraries, and a large amount of historical medical record data. Determine the screening weights of the keywords based on the diagnostic criteria and common symptom factors of the diseases. Store the keywords using a hash table, construct a disease keyword matching rule base, and establish a regular update mechanism. According to the latest research results in the medical field and newly entered medical record data, conduct a comprehensive review and update of the rule base once in a period of time, add new keywords in a timely manner, and delete keywords that are no longer applicable;
[0076] For the text content in the medical record data, split it according to sentences, phrases, or medical terms, and use the edit distance algorithm to calculate the edit distance between each text segment and the keywords in the rule base. When calculating the edit distance, adopt the dynamic programming algorithm to improve the calculation efficiency. Through the analysis and testing of medical text data, combined with medical professional knowledge, set the edit distance threshold for exact matching. When the edit distance between the text segment and the keyword is the set edit distance threshold for exact matching, it is determined as an exact match; set the edit distance threshold for fuzzy matching. When the edit distance is greater than the set edit distance threshold for exact matching and less than or equal to the set edit distance threshold for fuzzy matching, it is determined as a fuzzy match. Among them, the edit distance thresholds for exact matching and fuzzy matching are determined through the analysis of historical medical record data;
[0077] Introduce a word vector model. For the text segments and keywords successfully matched by the edit distance algorithm, further use the word vector model to calculate their semantic similarity. When the semantic similarity is lower than the set semantic similarity threshold, re-evaluate the matching results. For the case of fuzzy matching, preferentially refer to the matching results with higher semantic similarity, and select the first s matching results with reference semantic similarity;
[0078] Collect a large amount of accurately classified medical record data as a training set, select features related to disease classification, including symptom descriptions, examination results, and medication information, and use the Gini coefficient method to select the optimal features for the construction and training of decision trees;
[0079] The Gini coefficient method includes:
[0080] Collect accurately classified medical record data as a training set, and preprocess the unstructured data in the training set;
[0081] For each feature in the training set, calculate its Gini coefficient. Given that there are n samples and k classes in the training set, and the number of samples of the i-th class in the dataset is n i , the formula for calculating the Gini coefficient Gini(D) of the dataset is as follows:
[0082]
[0083] where n represents the number of all unstructured electronic medical records used for model training, k represents the number of disease categories included in the unstructured electronic medical records, n i represents the number of samples of the i-th class in the dataset, and Gini(D) represents the Gini coefficient of dataset D
[0084] For each feature, calculate the Gini coefficient after dividing the training subset, and perform a weighted sum according to the sample quantity of the subset to obtain the Gini coefficient Gini(D,A) after dividing the training subset. The calculation formula is as follows:
[0085]
[0086] where m represents the number of subsets obtained by dividing according to feature A, |D j | represents the number of samples in the j-th subset, Gini(D j ) represents the Gini coefficient of the j-th subset, and Gini(D,A) represents the Gini coefficient after dividing the training subset based on feature A;
[0087] Compare the Gini coefficients of all features, select the feature with the smallest Gini coefficient as the splitting feature of the current node, divide the training set according to the selected optimal feature to generate child nodes. For each child node, repeat the process of calculating the Gini coefficient, selecting the optimal feature, and performing the division to recursively construct a decision tree. During the construction process, set a stopping condition. When all samples in the child node belong to the same class, there are no more features available for splitting, or the preset maximum tree depth is reached, stop splitting the node, mark the node as a leaf node, and determine the disease category it belongs to according to the class distribution of the samples in the node;
[0088] Use the constructed decision tree model to train the training set, enabling the model to learn the relationship between features and disease classification. During the training process, divide the training set into q subsets, take one subset as the validation set in turn, and use the other subsets as the training set to perform multiple trainings and validations on the model, and adjust the minimum sample number and maximum tree depth parameters of the leaf nodes of the decision tree;
[0089] Randomly select a certain proportion (10%) of the processed data for manual spot-checking to check whether the term conversion is correct and the classification is accurate; compare the processed data with the known standard dataset, calculate the accuracy index, and evaluate the quality of data processing. The specific steps are as follows:
[0090] Adopt the stratified random sampling method, stratify according to dimensions such as disease type, visit time, department, etc., and randomly select samples within each layer;
[0091] Manually and carefully check whether the terms in the selected samples are accurately converted into standard medical terms and whether the classification results conform to the actual situation. For the discovered problems, record in detail the error types and the relevant information of the medical records (patient ID, visit time, etc.);
[0092] Improve the model update mechanism, which is triggered when the new data volume reaches 20% of the original training set or the model accuracy drops by more than 5%; the update process includes collecting new medical record data, re-cleaning and selecting features, and then retraining the decision tree model with the updated data to replace the original model to improve the classification performance.
[0093] Specifically, extract high-frequency words related to pneumonia from medical literature, medical term libraries, and historical medical record data, such as "cough", "expectoration", "fever", "lung shadow", etc. as disease keywords. According to the visit time, divide the data by year and season, and find that the number of pneumonia cases is relatively large in winter. Construct a disease keyword matching rule library, store the keywords in a hash table. For the above medical records, split the text according to sentences, phrases, or medical terms, calculate the edit distance between each text segment and the keywords using the edit distance algorithm. Set the exact match edit distance threshold to 0 and the fuzzy match edit distance threshold to 2. The edit distance between "cough" and the keyword "cough" is 0, which is determined to be an exact match; the edit distance between "low fever" and "fever" is 1, which is determined to be a fuzzy match. Introduce a word vector model to calculate the semantic similarity, and set the semantic similarity threshold to 0.8. After calculation, the semantic similarity between "low fever" and "fever" is 0.9, and the matching result is reliable. Collect 100 pieces of pneumonia medical record data that have been accurately classified as the training set, select features such as symptom descriptions, examination results, and medication information, and use the Gini coefficient method to select the optimal features to construct a decision tree model. In the training set, the Gini coefficient of the symptom description feature is the smallest and is selected as the splitting feature of the current node. Recursively construct the decision tree, set the maximum tree depth to 5, and stop splitting when all samples in the child node belong to the same category or there are no more features available for splitting. During the training process, divide the training set into 5 subsets, and take turns using one subset as the validation set and the other subsets as the training set to train and validate the model multiple times, and adjust the minimum sample number of the leaf nodes and the maximum tree depth parameters.
[0094] Based on the unstructured data in the medical record, relevant entities in the electronic medical record are extracted and marked through NLP technology, where the relevant entities include disease entities, symptom entities, treatment plan entities, as well as time and dosage attributes;
[0095] Based on the extracted entity information, syntactic analysis, rule matching, and deep learning models are used for relation extraction, and a subject-relation-object triple is constructed;
[0096] Combined with the extracted medical entities and relations, the entities are used as nodes in the constructed knowledge graph, and the relations between entity nodes are used as edges in the knowledge graph to construct the knowledge graph of the electronic medical record;
[0097] Combined with the graph neural network GNN, the knowledge graph is optimized and modeled to capture complex entities and relations in the medical record. The specific steps are as follows:
[0098] Initialize features for each node, and convert the text description of the node into a vector representation through a word vector model as the initial feature;
[0099] Use the graph convolutional network GCN to perform convolutional processing on the nodes and edges in the knowledge graph. Each node aggregates information from its neighbor nodes to update its features. The calculation formula is as follows:
[0100]
[0101] where, h v (l+1) represents the feature of node v in the l+1 layer, N(v) represents the set of neighbor nodes of node v, c vu represents the weight of the edge between node v and neighbor node u, c u represents the normalization coefficient, W (l) represents the weight matrix in the l layer, and σ represents the activation function;
[0102] Through the multi-layer convolutional operation of GNN, capture and model the relationships between nodes, and establish connections between diseases, drugs, and treatment plans through multi-hop information propagation; at the same time, use time series information to capture the dynamic relationships of nodes changing over time.
[0103] Specifically, based on the preprocessed medical record data above, use NLP technology to extract relevant entities: disease entity: "pneumonia"; symptom entities: "cough", "expectoration", "low fever"; treatment plan entity: "Amoxicillin Capsules 0.5g, orally 3 times a day for one week of treatment"; time attributes: "in the recent week", "for one week of treatment"; dosage attributes: "0.5g", "3 times a day";
[0104] Relational extraction is performed using syntactic analysis, rule matching, and deep learning models to construct subject-relation-object triples, specifically as follows: (pneumonia, symptom, cough), (pneumonia, symptom, expectoration), (pneumonia, symptom, low fever), (amoxicillin capsules, treatment, pneumonia). Taking these entities as nodes and relations as edges, an electronic medical record knowledge graph is constructed. With "pneumonia" as the central node, it connects symptom nodes such as "cough", "expectoration", "low fever", and the treatment plan node of "amoxicillin capsules". Combining with the graph neural network GNN, the knowledge graph is optimized and modeled:
[0105] The graph convolutional network GCN is used to perform convolutional processing on the nodes and edges in the knowledge graph. Each node aggregates information from its neighbor nodes to update its features. According to the current node "pneumonia" having neighbor nodes "cough" and "amoxicillin capsules", the updated features are calculated through the GCN formula. After multiple convolutional operations, the relationships between nodes are captured and modeled, establishing the connection between "pneumonia" and "amoxicillin capsules" through the "treatment" relationship. At the same time, time series information is used to capture the dynamic relationships of nodes changing over time, observing the changes in symptom nodes during the treatment process of pneumonia patients.
[0106] Based on the constructed knowledge graph, the ARIMA time series analysis algorithm is used to sort the electronic medical records in chronological order, extract time-related features. For diagnostic information, with reference to the release time of the latest medical research results and diagnosis and treatment guidelines, the time difference between the medical record diagnosis time and the latest standard time is calculated. Through the time series analysis model, the degree of timeliness attenuation of this diagnostic information in the current medical decision-making is evaluated. According to the analysis results, the medical record information is marked with timeliness, divided into three levels: "high timeliness", "medium timeliness", and "low timeliness";
[0107] In the constructed knowledge graph, credibility assessment is carried out from aspects such as data source, data consistency, and compliance with medical knowledge. For the data source, priority is given to trusting the medical record data released by authoritative medical institutions and well-known medical research institutions; in terms of data consistency, check whether the information in different parts of the medical record is contradictory; use the medical knowledge in the knowledge graph to judge whether the diagnosis and treatment plan in the medical record conform to the current medical cognition.
[0108] Specifically, based on the constructed knowledge graph, the ARIMA time series analysis algorithm is used to process the electronic medical records:
[0109] According to the latest pneumonia diagnosis and treatment guidelines released on January 1, 2024, the diagnosis time of the above medical record is October 1, 2023. The calculated time difference is 3 months. Based on the release time of the diagnosis and treatment guidelines, the diagnosis information within 0 - 6 months is determined as "high timeliness", 6 - 12 months as "medium timeliness", and more than 12 months as "low timeliness". According to the degree of timeliness decay evaluated by the model and combined with the set classification criteria, since the time difference of this medical record is 3 months, within the range of 0 - 6 months, the diagnosis information is determined as "high timeliness", marked as "high timeliness", and corresponding annotations are made;
[0110] From the perspective of data sources, this medical record comes from a tertiary hospital and has a relatively high credibility; by examining the information in different parts of the medical record, there are no contradictions among the symptom descriptions, diagnosis conclusions, and treatment plans, and the data consistency is good; using the medical knowledge in the knowledge graph to judge, the treatment of "pneumonia" with "amoxicillin capsules" conforms to the current medical understanding. After comprehensive evaluation, the credibility of this medical record data is high.
[0111] A data management system based on NLP, the system includes a data acquisition module, a data classification module, a knowledge graph construction module, and a graph neural network optimization module. The data acquisition module is used to obtain unstructured electronic medical record data through a data interface and perform data preprocessing; the data classification module is used to analyze the data by extracting disease keywords and according to the visit time, construct a rule base for keyword matching, optimize the classification results using a decision tree model, and evaluate the classification quality; the knowledge graph construction module is used to extract relevant entities from unstructured medical record data through NLP technology, extract the relationships between relevant entities, and combine medical entities and relationships to construct a knowledge graph; the graph neural network optimization module is used to calculate the similarity and causal relationship between entities by analyzing the structure of the graph nodes and edges, and optimize and model the knowledge graph in combination with the graph neural network; sort the electronic medical records and extract time features, evaluate the timeliness of the diagnosis information and make annotations, and at the same time evaluate the credibility of the medical record data in the knowledge graph.
[0112] The data acquisition module includes an interface connection unit and a preprocessing unit. The interface connection unit is used to establish a data interface with an external data source to obtain unstructured electronic medical record text data including patient basic information, diagnosis records, treatment records, and drug usage conditions. The preprocessing unit is used to preprocess the obtained unstructured electronic medical record data, extract disease keywords from medical literature, term libraries, and historical medical record data, analyze the disease time distribution law according to the visit time; construct a disease keyword matching rule base, perform text matching using the edit distance algorithm combined with the word vector model for preliminary data classification; collect the classified medical record data to train a decision tree model to optimize the classification results and achieve preliminary data classification.
[0113] The data classification module includes a feature selection unit, a decision tree construction and training unit, and a classification evaluation unit. The feature selection unit is used to select disease classification-related features from the obtained unstructured electronic medical record text data and select the optimal features using the Gini coefficient method. The decision tree construction and training unit is used to recursively construct a decision tree based on the optimal features selected by the Gini coefficient and set stopping conditions. During the training process, cross-validation is performed to adjust the model parameters. The classification evaluation unit is used to extract data by using the stratified random sampling method for manual spot-checking and comparison with the standard data set, and calculate the accuracy index to evaluate the data processing quality.
[0114] The knowledge graph construction module includes an entity extraction unit, a node construction unit, and an edge construction unit. The entity extraction unit is used to extract disease entities, symptom entities, treatment plan entities, and time and dosage attributes based on the unstructured data in the medical record by using NLP technology and mark them. The node construction unit is used to construct the extracted entities into nodes in the graph. The edge construction unit establishes edges in the graph according to the relationships between entities to form the connection structure of the knowledge graph.
[0115] The graph neural network optimization module includes a graph convolution operation unit and a quality management unit. The graph convolution operation unit is used to optimize and model the knowledge graph by combining the graph neural network GNN to capture complex entities and relationships in the medical record. The quality management unit is used to evaluate the timeliness and contribution of the data in the knowledge graph.
[0116] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A data management method based on NLP, characterized in that: The method includes the following steps: Obtain unstructured electronic medical record data through a data interface and perform data preprocessing; Extract disease keywords and analyze the data according to the visit time, construct a rule base for keyword matching, use a decision tree model to optimize the classification results, and evaluate the classification quality; Extract relevant entities from unstructured medical record data through NLP technology, extract the relationships between relevant entities, combine medical entities and relationships to construct a knowledge graph, calculate the similarity and causal relationship between entities by analyzing the structure of the nodes and edges of the graph, and optimize the modeling of the knowledge graph by combining graph neural networks; Sort the electronic medical records and extract time features, evaluate the timeliness of the diagnostic information and perform annotation, and at the same time evaluate the credibility of the medical record data in the knowledge graph.
2. A data management method based on NLP according to claim 1, characterized in that: Obtain unstructured electronic medical record data through a data interface, including text data of patient basic information, diagnosis records, treatment records, and drug usage; Perform preprocessing on the obtained unstructured electronic medical record data, including denoising, removing irrelevant information, redundant punctuation marks, spelling mistakes, duplicate content, and non-standard terms, use a pre-trained medical term conversion model, combine with a medical term library, convert non-standard terms into medical terms or standard expressions, and at the same time identify and remove invalid or irrelevant text parts; Take the Unified Medical Language System as the core term library, combine with the clinical term specifications within the hospital, establish term conversion rules, and use string matching algorithms and semantic understanding models to convert synonyms and abbreviations in the medical record data into standard medical terms; Assign labels to the data for preliminary classification.
3. The data management method based on NLP according to claim 2, wherein: Assign labels to the data for preliminary classification, and the specific steps are as follows: Extract high-frequency words closely related to various diseases from medical literature, medical term libraries, and historical medical record data as disease keywords, divide the data by year and season according to the visit time, and analyze the time distribution law of diseases; Construct a disease keyword matching rule base and use the edit distance algorithm to match the text in the medical record data with the keywords; The matching of the text in the medical record data with the keywords by the edit distance algorithm includes: Extract high-frequency words closely related to various diseases from medical authoritative literature, standard medical term libraries, and a large amount of historical medical record data as disease keywords, determine the screening weights of the keywords according to the diagnostic criteria and common symptom factors of the diseases, store the keywords using a hash table, construct a disease keyword matching rule base, and establish a regular update mechanism. According to the latest research results in the medical field and newly entered medical record data, conduct a comprehensive review and update of the rule base once in a period of time, add new keywords in a timely manner, and delete keywords that are no longer applicable; For the text content in medical record data, it is split according to sentences, phrases or medical terms. The edit distance between each text segment and the keywords in the rule base is calculated using the edit distance algorithm. When calculating the edit distance, the dynamic programming algorithm is adopted to improve the calculation efficiency. Through the analysis and testing of medical text data, combined with medical professional knowledge, the edit distance threshold for exact matching is set. When the edit distance between the text segment and the keyword is the set edit distance threshold for exact matching, it is determined as an exact match; the edit distance threshold for fuzzy matching is set. When the edit distance is greater than the set edit distance threshold for exact matching and less than or equal to the set edit distance threshold for fuzzy matching, it is determined as a fuzzy match. Among them, the edit distance threshold for exact matching and the edit distance threshold for fuzzy matching are determined through the analysis of historical medical record data; The word vector model is introduced. For the text segments and keywords that are successfully matched by the edit distance algorithm, the semantic similarity between them is further calculated using the word vector model. When the semantic similarity is lower than the set semantic similarity threshold, the matching result is re-evaluated. For the case of fuzzy matching, the matching results with higher semantic similarity are preferentially referred to, and the first s matching results with reference semantic similarity are selected; A large number of accurately classified medical record data are collected as the training set. Features related to disease classification are selected, including symptom descriptions, examination results, and medication information. The Gini coefficient method is used to select the optimal features for the construction and training of the decision tree; The Gini coefficient method includes: Collect the accurately classified medical record data as the training set and preprocess the unstructured data in the training set; For each feature in the training set, calculate its Gini coefficient. Given that there are n samples and k classes in the training set, and the number of samples of the i-th class in the dataset is n i , the calculation formula for the Gini coefficient Gini(D) of the dataset is as follows: Among them, n represents the number of all unstructured electronic medical records used for model training, k represents the number of disease categories included in the unstructured electronic medical records, and n i represents the number of samples in the i-th category in the dataset, and Gini(D) represents the Gini coefficient of the dataset D For each feature, calculate the Gini coefficient after dividing the training subset and perform weighted summation according to the sample size of the subset to obtain the Gini coefficient Gini(D,A) after dividing the training subset. The calculation formula is as follows: Among them, m represents the number of subsets obtained by dividing according to feature A, |D j | represents the number of samples in the j-th subset, Gini(D j ) represents the Gini coefficient of the j-th subset, and Gini(D, A) represents the Gini coefficient after dividing the training subset based on feature A; Compare the Gini coefficients of all features, select the feature with the smallest Gini coefficient as the splitting feature of the current node, divide the training set according to the selected optimal feature to generate child nodes. For each child node, repeat the process of calculating the Gini coefficient, selecting the optimal feature and performing the division, and recursively construct the decision tree. During the construction process, set the stopping condition. When all the samples in the child node belong to the same category, there are no more features available for splitting, or the preset maximum tree depth is reached, stop the splitting of the node, mark the node as a leaf node, and determine the disease category to which it belongs according to the category distribution of the samples in the node; Use the constructed decision tree model to train the training set to let the model learn the relationship between features and disease classification. During the training process, divide the training set into q subsets, take one subset as the validation set in turn, and the other subsets as the training set, perform multiple trainings and validations on the model, adjust the minimum sample number and maximum tree depth parameters of the leaf nodes of the decision tree, and evaluate the classification quality.
4. The data management method based on NLP according to claim 3, characterized in that: Based on the unstructured data in medical records, relevant entities in the electronic medical records are extracted and labeled through NLP technology, where the relevant entities include disease entities, symptom entities, treatment plan entities, as well as time and dosage attributes; Based on the extracted entity information, syntactic analysis, rule matching, and deep learning models are used for relation extraction, and a subject-relation-object triple is constructed; Combined with the extracted medical entities and relations, the entities are used as nodes in the constructed knowledge graph, and the relations between entity nodes are used as edges in the knowledge graph to construct the knowledge graph of electronic medical records; Combined with the graph neural network GNN, the knowledge graph is optimized and modeled to capture complex entities and relations in medical records. The specific steps are as follows: Initialize features for each node, and convert the text description of the node into a vector representation through a word vector model as the initial feature; Use the graph convolutional network GCN to perform convolutional processing on the nodes and edges in the knowledge graph. Each node aggregates information from its neighbor nodes to update its features. The calculation formula is as follows: Among them, h v (l+1) represents the feature of the node v at the (l + 1)-th layer, N(v) represents the set of neighbor nodes of the node v, and c vu represents the weight of the edge between the node v and the neighbor node u, and c u represents the normalization coefficient, and W (l) represents the weight matrix at the l-th layer, and σ represents the activation function; Through multiple-layer convolutional operations of the GNN, capture and model the relationships between nodes, and establish connections between diseases, drugs, and treatment plans through multi-hop information propagation; at the same time, use time series information to capture the dynamic relationships of nodes changing over time.
5. The data management method based on NLP according to claim 4, characterized in that: Based on the constructed knowledge graph, using the ARIMA time series analysis algorithm, sort the electronic medical records in chronological order, extract time-related features. For diagnostic information, with reference to the release time of the latest medical research results and treatment guidelines, calculate the time difference between the medical record diagnosis time and the latest standard time, and evaluate the timeliness decay degree of the diagnostic information in the current medical decision-making through the time series analysis model. According to the analysis results, perform timeliness annotation on the medical record information; In the constructed knowledge graph, perform credibility evaluation from aspects such as data source, data consistency, and compliance with medical knowledge. For the data source, give priority to trusting the medical record data released by authoritative medical institutions and well-known medical research institutions; in terms of data consistency, check whether the information in different parts of the medical record is contradictory; use the medical knowledge in the knowledge graph to judge whether the diagnosis and treatment plan in the medical record conform to the current medical cognition.
6. A data management system based on NLP, characterized in that: The system includes a data acquisition module, a data classification module, a knowledge graph construction module, and a graph neural network optimization module. The data acquisition module is used to obtain unstructured electronic medical record data through a data interface and perform data preprocessing. The data classification module is used to analyze the data by extracting disease keywords and according to the visit time, construct a rule base for keyword matching, optimize the classification results using a decision tree model, and evaluate the classification quality. The knowledge graph construction module is used to extract relevant entities from unstructured medical record data through NLP technology, extract the relationships between relevant entities, and combine medical entities and relationships to construct a knowledge graph. The graph neural network optimization module is used to calculate the similarity and causal relationship between entities by analyzing the structure of graph nodes and edges, and optimize the modeling of the knowledge graph by combining a graph neural network. Sort the electronic medical records and extract time features, evaluate the timeliness of diagnostic information and perform annotation, and at the same time evaluate the credibility of the medical record data in the knowledge graph.
7. A data management system based on NLP according to claim 6, characterized in that: The data acquisition module includes an interface connection unit and a preprocessing unit. The interface connection unit is used to establish a data interface with an external data source to obtain unstructured electronic medical record text data including patient basic information, diagnosis records, treatment records, and drug usage. The preprocessing unit is used to preprocess the obtained unstructured electronic medical record data, extract disease keywords from medical literature, a thesaurus, and historical medical record data, analyze the disease time distribution law according to the visit time; construct a disease keyword matching rule base, perform text matching using the edit distance algorithm combined with a word vector model, and perform preliminary classification of the data; collect the classified medical record data to train a decision tree model to optimize the classification results and achieve preliminary classification of the data.
8. A data management system based on NLP according to claim 7, characterized in that: The data classification module includes a feature selection unit, a decision tree construction and training unit, and a classification evaluation unit. The feature selection unit is used to select features related to disease classification from the obtained unstructured electronic medical record text data, and use the Gini coefficient method to select the optimal features. The decision tree construction and training unit is used to recursively construct a decision tree according to the optimal features selected by the Gini coefficient and set stop conditions; during the training process, perform cross-validation and adjust the model parameters. The classification evaluation unit is used to extract data using the stratified random sampling method for manual spot checks and comparison with a standard data set, and calculate the accuracy index to evaluate the data processing quality.
9. A data management system based on NLP according to claim 8, characterized in that: The knowledge graph construction module includes an entity extraction unit, a node construction unit, and an edge construction unit. The entity extraction unit is used to extract disease entities, symptom entities, treatment plan entities, and time and dosage attributes based on the unstructured data in the medical record using NLP technology and perform marking. The node construction unit is used to construct the extracted entities into nodes in the graph. The edge construction unit establishes edges in the graph according to the relationships between entities to form the connection structure of the knowledge graph.
10. A data management system based on NLP according to claim 9, characterized in that: The graph neural network optimization module includes a graph convolution operation unit and a quality management unit. The graph convolution operation unit is used to optimize and model the knowledge graph in combination with the graph neural network GNN to capture complex entities and relationships in the medical records. The quality management unit is used to evaluate the timeliness and contribution degree of the data in the knowledge graph.
Citation Information
Patent Citations
Control method for prompting annotation of on medical data
CN111028953A
Semi-supervised learning method for constructing medical knowledge graph from Chinese electronic medical records
CN112542223A
Knowledge graph construction method based on Chinese electronic medical records
CN113688255A
Drug recommendation method, device and system, electronic equipment and storage medium
CN114765075A
Knowledge graph-driven medical large model diagnosis method
CN118280562A
Cited By
Automatic generation method and system of industrial diagnosis report
CN120745572A
Multi-modal medical data intelligent association analysis system based on deep learning
CN121034512A
Multi-document question and answer method and device based on artificial intelligence and knowledge graph and medium
CN121388179A