System and method for extracting key information of medical document by using deep NLP model
Through the combination of the deep NLP model and the multi-task NLP model, the problem of lagging updating of the medical ontology library is solved, dynamic update of the ontology library and the construction of the medical knowledge graph are realized, and the timeliness and integrity of the knowledge is ensured.
Patent Information
- Application Number
- CN202510458700.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing technology has shortcomings in updating the pharmaceutical ontology library, and it is impossible to capture changes and new knowledge in the pharmaceutical field in a timely manner, resulting in lagging knowledge updates.
Using the deep NLP model, the original corpus and standardized entity dictionary were constructed by collecting multi-source medical documents, cleaning noise and aligning the standard term database. Use the pre-trained model BioBERT for domain adaptation adjustment, train multitasking NLP models, optimize model weights, process new documents and update the ontology library through threshold filtering, clustering analysis, and expert verification. The EWC algorithm is used to incrementally update the model to build a dynamic medical knowledge graph.
The dynamic expansion and improvement of the ontology library has been achieved, which can promptly reflect new knowledge and changes in the medical field and maintain the timeliness and integrity of the knowledge system.
Smart Images

Figure CN119988646A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and specifically to a system and method for using a deep NLP model to extract key information from medical documents. Background Art
[0002] In the field of medicine, information is growing explosively. Various medical documents, as important carriers of knowledge, cover a wide range of content from basic research to clinical applications. Documents such as drug instructions, medical literature, and clinical guidelines contain a large amount of valuable information, but most of them exist in unstructured form and are difficult to be directly and efficiently used by computers. Natural language processing (NLP) technology can convert unstructured medical texts into structured knowledge by cleaning and standardizing texts and using appropriate models for training and analysis, thereby constructing a medical knowledge graph and providing strong support for research, decision-making, and application in the medical field.
[0003] In the current medical knowledge management system, the ontology library is a core component, which defines the concepts, relationships and attributes in the medical field and provides a framework for the organization and understanding of knowledge. However, the existing technology has obvious deficiencies in updating the ontology library. With the rapid development of medical science, new knowledge such as new drug molecules, disease diagnosis methods, and treatment methods are constantly emerging, and new medical terms are also generated. However, the existing technology is often unable to capture these changes in a timely manner, resulting in a lag in the update of the ontology library. Summary of the invention
[0004] The purpose of the present invention is to provide a system and method for extracting key information from medical documents using a deep NLP model to solve the problems raised in the prior art.
[0005] To achieve the above object, the present invention provides the following technical solution: a method for extracting key information from medical documents using a deep NLP model, the method comprising the following steps: Step 1: Collect multi-source data, clean noise and align them according to the standard terminology library to build the original corpus, standardized entity dictionary and key information ontology library; Step 2: Use the pre-trained model BioBERT to adapt the annotated dataset, train a multi-task NLP model based on the task, and optimize the model weights based on high-frequency terms; Step 3: Process new documents with the optimized model, update the ontology library after threshold filtering, cluster analysis, and expert verification, and add new terms and annotations; Step 4: Mix the new and old annotated data and use the EWC algorithm to incrementally update the model; parse the new documents and build a dynamic medical knowledge graph.
[0006] In step 1, various types of unstructured medical documents are collected, including not only common drug instructions and medical literature, but also clinical guidelines, case reports, pharmaceutical company R&D reports and other data sources to ensure the comprehensiveness and richness of knowledge; Identify and clean various types of noise in the original document; noise types include but are not limited to format errors, irrelevant metadata, duplicate content, etc. Use regular expression matching, text format verification tools, and custom cleaning rules to comprehensively clean the original document and obtain the original corpus; For non-standard terms in the original corpus, a method based on rule matching and semantic similarity calculation is used to align them with the standard terminology library; the edit distance algorithm is used to calculate the similarity between non-standard terms and terms in the standard terminology library, and at the same time, the lexical and syntactic rules specific to the medical field are combined to determine the consistency of the terms and obtain a standardized medical entity dictionary; Organize a team of medical experts to define entities, relationships and attributes in detail based on standardized dictionaries and original corpora through specially developed annotation tools; obtain a structured key information ontology library; Collect public medical data sets, screen and evaluate them to ensure data quality and compatibility with this solution; integrate the screened public data sets with expert-annotated data, and form an annotated data set through operations such as data deduplication and format unification.
[0007] In step 2, BioBERT is selected as the basic model for domain adaptation adjustment for the annotated dataset. BioBERT has a certain pre-training foundation in the biomedical field, but it still needs to be further optimized for the specific medical knowledge graph construction needs of this project. By fine-tuning the annotated dataset, adjusting the parameters and feature representation of the model to make it more suitable for the language characteristics and knowledge structure of the medical field, an optimized text encoder is obtained. Use the optimized text encoder to encode the labeled dataset and convert the text data into a vector representation suitable for model processing; CRF (conditional random field), multi-head attention mechanism and question-answer decoder are used for joint training to build a multi-task NLP model. During the training process, the role and coordination of each component are clarified. CRF is used to process the label dependency in the sequence labeling task, the multi-head attention mechanism enhances the model's ability to focus on different parts of the text, and the question-answer decoder is used to extract specific knowledge information from the text. The attention distribution of the multi-task NLP model is analyzed. Based on the high-frequency terms in the standardized dictionary, the model's attention to key terms is enhanced by adjusting the attention weights. The frequency of occurrence and importance scores of high-frequency terms in the text are calculated, and the weights in the model's attention mechanism are adjusted according to the scores, so that the model focuses more on key terms when processing text, resulting in an optimized NLP model.
[0008] In step 3, the newly uploaded document is processed using the optimized NLP model, and the model analyzes and predicts the document content. In the prediction result processing phase, the model prediction results are filtered by a preset threshold, and the parts with prediction confidence higher than the threshold are screened out to obtain a candidate set of potential new information. For the context of the new information candidate set, the Sentence-BERT model is used to convert the text into semantic vectors, and then the K-means clustering algorithm is used to cluster these vectors. After the clustering operation, new term candidates with semantic grouping are obtained; Submit the clustering results to specially established experts for verification; the experts clearly display the relevant information of the clustering results, including the text content, contextual information and the category to which the cluster belongs of the new term candidate; the experts review and judge the new term candidate based on their own professional knowledge and field experience; formulate detailed verification standards and processes; for the new terms and related annotation data verified by experts, update them to the existing ontology library to realize the dynamic expansion and improvement of the ontology library; In step 4, the annotated data set and the newly annotated data are mixed, and the data playback method is used to ensure that the proportion of each type of data in the mixed data set is reasonable; by analyzing the distribution of different categories of data and using oversampling or undersampling techniques, the data is balanced to obtain a balanced data set.
[0009] For the optimized NLP model parameters, the elastic weight solidification (EWC) algorithm is used to protect important knowledge such as entities, relationships, and attributes that the model has learned in the labeled dataset, and obtain an incrementally updated NLP model.
[0010] For newly uploaded documents, the LayoutLM model is used for parsing; the LayoutLM model parses the document into structured text or table data to obtain accurate structured output; The parsed text is input into the incrementally updated model, and the model extracts the entities, relationships, and attributes in the text. A medical knowledge graph is constructed using the Neo4j graph database, with the extracted entities as nodes, relationships as edges, and attributes as node or edge features. A dynamically updated medical knowledge graph is implemented through graph structure design and data storage methods.
[0011] A system for extracting key information from medical documents using deep NLP models. The system includes a data preprocessing and annotation module, a multi-task model training and optimization module, a new term discovery and verification module, and an incremental learning and knowledge graph construction module. The data preprocessing and annotation module is used to collect multi-source data, clean noise and align it according to the standard terminology library, and build the original corpus, standardized entity dictionary and key information ontology library; the multi-task model training and optimization module is used to use the pre-trained model BioBERT to adapt and adjust the annotation data set, train the multi-task NLP model according to the task, and optimize the model weight based on high-frequency terms; the new term discovery and verification module is used to process new documents with the optimized model, update the ontology library after threshold filtering, cluster analysis, and expert verification, and supplement new terms and annotations; the incremental learning and knowledge graph construction module is used to mix new and old annotated data, and use the EWC algorithm to incrementally update the model; parse new documents and build a dynamic medical knowledge graph; The output end of the data preprocessing and annotation module is connected to the input end of the multi-task model training and optimization module; the output end of the multi-task model training and optimization module is connected to the input end of the new terminology discovery and verification module; the output end of the new terminology discovery and verification module is connected to the input end of the incremental learning and knowledge graph construction module and the data preprocessing and annotation module; the output end of the incremental learning and knowledge graph construction module is connected to the input end of the new terminology discovery and verification module.
[0012] The data preprocessing and annotation module includes a multi-source data acquisition unit, a data cleaning and standardization unit, and an ontology library construction and annotation unit; The multi-source data acquisition unit is used to collect medical documents and public data sets; the data cleaning and standardization unit is used to clean noise data, align standard terminology libraries, generate original corpora and standardized entity dictionaries; the ontology library construction and annotation unit is used for experts to annotate entities, relationships and attributes to form a structured key information ontology library and annotated data sets; The output end of the multi-source data acquisition unit is connected to the input end of the data cleaning and standardization unit; the output end of the data cleaning and standardization unit is connected to the input end of the ontology library construction and annotation unit; the output end of the ontology library construction and annotation unit is connected to the input end of the multi-task model training and optimization module.
[0013] The multi-task model training and optimization module includes a domain adaptation pre-training unit, a multi-task model joint training unit and a high-frequency term weight optimization unit; The domain adaptation pre-training unit is used to use BioBERT to perform domain adaptation fine-tuning on medical field texts and optimize the text encoder; the multi-task model joint training unit is used to jointly train CRF entity recognition, multi-head attention relationship extraction, and question-answer decoder attribute extraction; the high-frequency term weight optimization unit is used to enhance the model attention weight based on high-frequency terms in the standardized dictionary and improve the accuracy of key information extraction; The output end of the domain adaptation pre-training unit is connected to the input end of the multi-task model joint training unit; the output end of the multi-task model joint training unit is connected to the input end of the high-frequency term weight optimization unit; the output end of the high-frequency term weight optimization unit is connected to the input end of the new term discovery and verification module.
[0014] The new term discovery and verification module includes a new document information extraction unit, a semantic clustering analysis unit and an expert verification and update unit; The new document information extraction unit is used to process new documents with the optimized model and generate a potential new term candidate set by filtering through a preset threshold; the semantic clustering analysis unit is used to cluster the context of candidate terms using Sentence-BERT and K-means to generate semantic groupings; the expert verification and update unit is used for expert review of clustering results and to supplement new terms and annotation data into the key information ontology library; The output end of the new document information extraction unit is connected to the input end of the semantic clustering analysis unit; the output end of the semantic clustering analysis unit is connected to the input end of the expert verification and update unit; the output end of the expert verification and update unit is connected to the input end of the incremental learning and knowledge graph construction module and the data preprocessing and annotation module.
[0015] The incremental learning and knowledge graph construction module includes a hybrid data playback unit, an EWC incremental model update unit and a dynamic knowledge graph construction unit; The mixed data playback unit is used to mix new and old annotated data and balance the distribution of data sets to support incremental training; the EWC incremental model update unit is used to use elastic weight solidification EWC algorithm to update the model and protect historical knowledge from being forgotten; the dynamic knowledge graph construction unit is used to parse the new document structure, extract entity relationships, and update the knowledge graph with Neo4j; The output end of the hybrid data playback unit is connected to the input end of the EWC incremental model update unit; the output end of the EWC incremental model update unit is connected to the input end of the dynamic knowledge graph construction unit; the output end of the dynamic knowledge graph construction unit is connected to the input end of the new term discovery and verification module.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention collects common drug instructions and medical literature, and also incorporates multiple data sources such as clinical guidelines, case reports, and pharmaceutical company research and development reports to ensure the comprehensiveness and richness of knowledge. Compared with the prior art that may only rely on a single or a few data sources, it can cover a wider range of medical knowledge and provide a more solid data foundation for subsequent analysis and application; the present invention processes new documents, and after threshold filtering, cluster analysis and expert verification, updates new terms and related annotation data to the ontology library, thereby realizing dynamic expansion and improvement of the ontology library, which can timely reflect new knowledge and new changes in the medical field and maintain the timeliness and integrity of the knowledge system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A schematic diagram of the steps of the method for extracting key information from medical documents using a deep NLP model according to the present invention; Figure 2 Schematic diagram of the process of the system for extracting key information from medical documents using the deep NLP model of the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] Example: Figure 1-Figure 2 As shown, the present invention provides a technical solution, a method for extracting key information from medical documents using a deep NLP model, characterized in that the method comprises the following steps: Step 1: Collect multi-source data, clean noise and align them according to the standard terminology library to build the original corpus, standardized entity dictionary and key information ontology library; Step 2: Use the pre-trained model BioBERT to adapt the annotated dataset, train a multi-task NLP model based on the task, and optimize the model weights based on high-frequency terms; Step 3: Process new documents with the optimized model, update the ontology library after threshold filtering, cluster analysis, and expert verification, and add new terms and annotations; Step 4: Mix the new and old annotated data and use the EWC algorithm to incrementally update the model; parse the new documents and build a dynamic medical knowledge graph.
[0020] Step 1: Multi-source data collection and cleaning; Data sources: drug instructions (1,000 copies, including PDF and text formats); medical journal articles (5,000 articles, from public medical databases); hospital desensitization clinical reports (200 copies, including efficacy and side effect records); Cleaning operations: remove non-text content (such as tables and pictures in PDF) and extract plain text; use regular expressions to filter special symbols (such as "®" and "©") and garbled characters; split text by paragraphs and remove duplicate paragraphs (such as the "General Warning" section of drug instructions); Example data snippet: The Phase II clinical trial of Drug A (main ingredient: chemical X) showed that the remission rate for Disease B was 75%, and common adverse reactions included rash (incidence 15%) and dizziness (incidence 8%).
[0021] Terminology standardization alignment; Standard library matching: Use the "National Drug Standard Terminology Library" to align non-standard names; for example: "Drug A" → standard name "Chemical A_National Medicine Standard H2022001", "Disease B" → standard code "ICD-11:8A45.0"; Processing of unmatched terms: Mark as items to be reviewed by experts and temporarily store in a temporary library; Expert annotation and ontology library construction: Labeling rules: Entity type: drug, disease, ingredient, adverse reaction, test indicator; Relationship definitions: drug-treatment-disease, drug-induces-adverse-reaction, ingredient-belongs-to-drug; Annotation result example: {"text":"Drug A has a 75% remission rate for disease B, and adverse reactions include rash.","entity":[{"name":"Drug A","type":"Drug","standard code":"H2022001"},{"name":"Disease B","type":"Disease","standard code":"8A45.0"},{"name":"Rash","type":"Adverse reaction","standard code":"ADR_0032"}],"relationship":[{"subject":"H2022001","object":"8A45.0","type":"Treatment"},{"subject":"H2022001","object":"ADR_0032","type":"Trigger"}]}; Step 2: Model training and optimization; Model selection: BioBERT, a deep learning model pre-trained based on medical text; Input format: divide the text into blocks of 512 characters, retaining key position information such as drug name and disease name; parameter adjustment: initial learning rate: 0.00002; training rounds: 3 rounds; Multi-task joint training: Task design: Entity recognition task: mark the start and end positions of drugs, diseases, and adverse reactions; Relation extraction task: determine whether there are relations such as "treatment" and "initiation" in the sentence; Attribute extraction task: extract numerical attributes from the text (such as "relief rate 75%"); Model structure optimization: In the model attention layer, increase the weight of key entities such as drug names and disease names (for example: entity position attention weight × 1.5); Using a dynamic loss balancing strategy, the entity recognition loss weight accounts for 60%, the relationship extraction accounts for 30%, and the attribute extraction accounts for 10%; High-frequency term enhancement: The top 10% of terms in the standard term library are counted (such as "chemotherapy" and "tumor"). When training the model, if the text contains high-frequency terms, the loss function of the sample is weighted (for example, weight coefficient × 2). Step 3: New term discovery and dynamic update: Input example: A new drug instruction fragment: "Long-term use of drug C may cause elevated blood potassium (incidence 5%), and regular monitoring of electrolytes is recommended." Model prediction results: Entities: "Drug C" (drug, confidence 0.94), "increased blood potassium" (adverse reaction, confidence 0.68); Relationship: "Drug C-induced-increased blood potassium" (confidence 0.72); Threshold filtering: retain results with confidence ≥ 0.65 ("increased blood potassium" is retained); Semantic Clustering and Expert Review: Clustering process: extract the context of "elevated blood potassium" (such as "long-term use may lead to elevated blood potassium"); use the semantic encoding model to convert the context into a vector; group by vector similarity, and classify it into the "abnormal biochemical indicators" category together with "hyponatremia" and "elevated creatinine"; Expert decision: confirm "increased blood potassium" as a new adverse reaction, code it as "ADR_0157"; add the relationship: "drug C-induced-ADR_0157"; Ontology library and model iteration: Data playback: Mix the newly added 200 annotated data (including "increased blood potassium") with the original 6000 data; randomly sample at a ratio of 5:1 to avoid imbalance between new and old data; Incremental model training: Use elastic weight solidification technology to lock the model's recognition ability for old terms (such as "rash" and "dizziness"); only allow small adjustments to model parameters in the direction of new terms (the learning rate is set to 0.00001).
[0022] Step 4: Dynamic construction of knowledge graph; Structured analysis and storage: Table parsing results: {"drug":"drug C","attribute":[{"name":"recommended dose","value":"50mg / d"},{"name":"half-life","value":"12 hours"}]}; Graph update: New nodes: Drug C (ID: H2023001), adverse reaction "increased blood potassium" (ID: ADR_0157); New relationship: H2023001-trigger-ADR_0157; New attributes: recommended dose and half-life of Drug C.
[0023] A system for extracting key information from medical documents using deep NLP models. The system includes a data preprocessing and annotation module, a multi-task model training and optimization module, a new term discovery and verification module, and an incremental learning and knowledge graph construction module. The data preprocessing and annotation module is used to collect multi-source data, clean noise and align it according to the standard terminology library, and build the original corpus, standardized entity dictionary and key information ontology library; the multi-task model training and optimization module is used to use the pre-trained model BioBERT to adapt and adjust the annotation data set, train the multi-task NLP model according to the task, and optimize the model weight based on high-frequency terms; the new term discovery and verification module is used to process new documents with the optimized model, update the ontology library after threshold filtering, cluster analysis, and expert verification, and supplement new terms and annotations; the incremental learning and knowledge graph construction module is used to mix new and old annotated data, and use the EWC algorithm to incrementally update the model; parse new documents and build a dynamic medical knowledge graph; The output end of the data preprocessing and annotation module is connected to the input end of the multi-task model training and optimization module; the output end of the multi-task model training and optimization module is connected to the input end of the new terminology discovery and verification module; the output end of the new terminology discovery and verification module is connected to the input end of the incremental learning and knowledge graph construction module and the data preprocessing and annotation module; the output end of the incremental learning and knowledge graph construction module is connected to the input end of the new terminology discovery and verification module.
[0024] The data preprocessing and annotation module includes a multi-source data acquisition unit, a data cleaning and standardization unit, and an ontology library construction and annotation unit; The multi-source data acquisition unit is used to collect medical documents and public data sets; the data cleaning and standardization unit is used to clean noise data, align standard terminology libraries, generate original corpora and standardized entity dictionaries; the ontology library construction and annotation unit is used for experts to annotate entities, relationships and attributes to form a structured key information ontology library and annotated data sets; The output end of the multi-source data acquisition unit is connected to the input end of the data cleaning and standardization unit; the output end of the data cleaning and standardization unit is connected to the input end of the ontology library construction and annotation unit; the output end of the ontology library construction and annotation unit is connected to the input end of the multi-task model training and optimization module.
[0025] The multi-task model training and optimization module includes a domain adaptation pre-training unit, a multi-task model joint training unit and a high-frequency term weight optimization unit; The domain adaptation pre-training unit is used to use BioBERT to perform domain adaptation fine-tuning on medical field texts and optimize the text encoder; the multi-task model joint training unit is used to jointly train CRF entity recognition, multi-head attention relationship extraction, and question-answer decoder attribute extraction; the high-frequency term weight optimization unit is used to enhance the model attention weight based on high-frequency terms in the standardized dictionary and improve the accuracy of key information extraction; The output end of the domain adaptation pre-training unit is connected to the input end of the multi-task model joint training unit; the output end of the multi-task model joint training unit is connected to the input end of the high-frequency term weight optimization unit; the output end of the high-frequency term weight optimization unit is connected to the input end of the new term discovery and verification module.
[0026] The new term discovery and verification module includes a new document information extraction unit, a semantic clustering analysis unit and an expert verification and update unit; The new document information extraction unit is used to process new documents with the optimized model and generate a potential new term candidate set by filtering through a preset threshold; the semantic clustering analysis unit is used to cluster the context of candidate terms using Sentence-BERT and K-means to generate semantic groupings; the expert verification and update unit is used for expert review of clustering results and to supplement new terms and annotation data into the key information ontology library; The output end of the new document information extraction unit is connected to the input end of the semantic clustering analysis unit; the output end of the semantic clustering analysis unit is connected to the input end of the expert verification and update unit; the output end of the expert verification and update unit is connected to the input end of the incremental learning and knowledge graph construction module and the data preprocessing and annotation module.
[0027] The incremental learning and knowledge graph construction module includes a hybrid data playback unit, an EWC incremental model update unit and a dynamic knowledge graph construction unit; The mixed data playback unit is used to mix new and old annotated data and balance the distribution of data sets to support incremental training; the EWC incremental model update unit is used to use elastic weight solidification EWC algorithm to update the model and protect historical knowledge from being forgotten; the dynamic knowledge graph construction unit is used to parse the new document structure, extract entity relationships, and update the knowledge graph with Neo4j; The output end of the hybrid data playback unit is connected to the input end of the EWC incremental model update unit; the output end of the EWC incremental model update unit is connected to the input end of the dynamic knowledge graph construction unit; the output end of the dynamic knowledge graph construction unit is connected to the input end of the new term discovery and verification module.
[0028] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. A method for extracting key information from medical documents using a deep NLP model, characterized in that: The method comprises the following steps: Step 1: Collect multi-source data, clean noise and align them according to the standard terminology library to build the original corpus, standardized entity dictionary and key information ontology library; Step 2: Use the pre-trained model BioBERT to adapt the annotated dataset, train a multi-task NLP model based on the task, and optimize the model weights based on high-frequency terms; Step 3: Process new documents with the optimized model, update the ontology library after threshold filtering, cluster analysis, and expert verification, and add new terms and annotations; Step 4: Mix the new and old annotated data and use the EWC algorithm to incrementally update the model; parse the new documents and build a dynamic medical knowledge graph.
2. The method for extracting key information from medical documents using a deep NLP model according to claim 1, characterized in that: In step 1, unstructured medical documents, including instructions, reports, and literature, are collected; For the original document, clean the noise to get the original corpus; For non-standard terms in the original corpus, the standard term base is used for alignment to obtain a standardized entity dictionary; For the standardized dictionary and original corpus, entities, relations and attributes are defined through expert annotation to obtain a structured key information ontology library; Integrate public data sets and expert-annotated data to form an annotated data set.
3. The method for extracting key information from medical documents using a deep NLP model according to claim 2, characterized in that: In step 2, for the annotated dataset, the pre-trained model BioBERT is used for domain adaptation to obtain an optimized text encoder; Encode the annotated dataset using an optimized text encoder; For the encoded text, a multi-task NLP model is obtained by jointly training CRF, multi-head attention, and question-answer decoder; For the multi-task NLP model attention distribution, the weights are enhanced based on the high-frequency terms in the standardized dictionary to obtain an optimized NLP model.
4. The method for extracting key information from medical documents using a deep NLP model according to claim 3, characterized in that: In step 3, the optimized NLP model is used to process new documents, and the model prediction results are filtered based on the preset threshold to obtain a candidate set of potential new information; For the context of the new information candidate set, Sentence-BERT and K-means clustering are used to obtain semantically grouped new term candidates; The clustering results are verified through the expert platform to obtain the updated ontology library, including the newly added terms and annotation data.
5. The method for extracting key information from medical documents using a deep NLP model according to claim 4, characterized in that: In step 4, the labeled data set and the newly labeled data are mixed using data playback to obtain a balanced data set; For the optimized NLP model parameters, the elastic weight consolidation (EWC) algorithm is used to protect the entities, relationships, and attributes learned by the model in the labeled dataset to obtain an incrementally updated NLP model. For newly uploaded documents, use LayoutLM to parse them and obtain structured text or table data; For the parsed text, use the incrementally updated model to obtain entity, relationship or attribute results; For the extracted results, Neo4j is used to build a graph to obtain a dynamically updated medical knowledge graph.
6. A system for extracting key information from medical documents using a deep NLP model, applied to a method for extracting key information from medical documents using a deep NLP model as claimed in any one of claims 1 to 5, characterized in that: The system includes a data preprocessing and annotation module, a multi-task model training and optimization module, a new term discovery and verification module, and an incremental learning and knowledge graph construction module; The data preprocessing and annotation module is used to collect multi-source data, clean noise and align it according to the standard terminology library, and build the original corpus, standardized entity dictionary and key information ontology library; the multi-task model training and optimization module is used to use the pre-trained model BioBERT to adapt and adjust the annotation data set, train the multi-task NLP model according to the task, and optimize the model weight based on high-frequency terms; the new term discovery and verification module is used to process new documents with the optimized model, update the ontology library after threshold filtering, cluster analysis, and expert verification, and supplement new terms and annotations; the incremental learning and knowledge graph construction module is used to mix new and old annotated data, and use the EWC algorithm to incrementally update the model; parse new documents and build a dynamic medical knowledge graph; The output end of the data preprocessing and annotation module is connected to the input end of the multi-task model training and optimization module; the output end of the multi-task model training and optimization module is connected to the input end of the new terminology discovery and verification module; the output end of the new terminology discovery and verification module is connected to the input end of the incremental learning and knowledge graph construction module and the data preprocessing and annotation module; the output end of the incremental learning and knowledge graph construction module is connected to the input end of the new terminology discovery and verification module.
7. The system for extracting key information from medical documents using a deep NLP model according to claim 6, characterized in that: The data preprocessing and annotation module includes a multi-source data acquisition unit, a data cleaning and standardization unit, and an ontology library construction and annotation unit; The multi-source data acquisition unit is used to collect medical documents and public data sets; the data cleaning and standardization unit is used to clean noise data, align standard terminology libraries, generate original corpora and standardized entity dictionaries; the ontology library construction and annotation unit is used for experts to annotate entities, relationships and attributes to form a structured key information ontology library and annotated data sets; The output end of the multi-source data acquisition unit is connected to the input end of the data cleaning and standardization unit; the output end of the data cleaning and standardization unit is connected to the input end of the ontology library construction and annotation unit; the output end of the ontology library construction and annotation unit is connected to the input end of the multi-task model training and optimization module.
8. The system for extracting key information from medical documents using a deep NLP model according to claim 7, characterized in that: The multi-task model training and optimization module includes a domain adaptation pre-training unit, a multi-task model joint training unit and a high-frequency term weight optimization unit; The domain adaptation pre-training unit is used to use BioBERT to perform domain adaptation fine-tuning on medical field texts and optimize the text encoder; the multi-task model joint training unit is used to jointly train CRF entity recognition, multi-head attention relationship extraction, and question-answer decoder attribute extraction; the high-frequency term weight optimization unit is used to enhance the model attention weight based on high-frequency terms in the standardized dictionary and improve the accuracy of key information extraction; The output end of the domain adaptation pre-training unit is connected to the input end of the multi-task model joint training unit; the output end of the multi-task model joint training unit is connected to the input end of the high-frequency term weight optimization unit; the output end of the high-frequency term weight optimization unit is connected to the input end of the new term discovery and verification module.
9. The system for extracting key information from medical documents using a deep NLP model according to claim 8, characterized in that: The new term discovery and verification module includes a new document information extraction unit, a semantic clustering analysis unit and an expert verification and update unit; The new document information extraction unit is used to process new documents with the optimized model and generate a potential new term candidate set by filtering through a preset threshold; the semantic clustering analysis unit is used to cluster the context of candidate terms using Sentence-BERT and K-means to generate semantic groupings; the expert verification and update unit is used for expert review of clustering results and to supplement new terms and annotation data into the key information ontology library; The output end of the new document information extraction unit is connected to the input end of the semantic clustering analysis unit; the output end of the semantic clustering analysis unit is connected to the input end of the expert verification and update unit; the output end of the expert verification and update unit is connected to the input end of the incremental learning and knowledge graph construction module and the data preprocessing and annotation module.
10. The system for extracting key information from medical documents using a deep NLP model according to claim 9, characterized in that: The incremental learning and knowledge graph construction module includes a hybrid data playback unit, an EWC incremental model update unit and a dynamic knowledge graph construction unit; The mixed data playback unit is used to mix new and old annotated data and balance the distribution of data sets to support incremental training; the EWC incremental model update unit is used to update the model using the elastic weight solidification EWC algorithm to protect historical knowledge from being forgotten; The dynamic knowledge graph construction unit is used to parse the new document structure, extract entity relationships, and update the knowledge graph using Neo4j; The output end of the hybrid data playback unit is connected to the input end of the EWC incremental model update unit; the output end of the EWC incremental model update unit is connected to the input end of the dynamic knowledge graph construction unit; the output end of the dynamic knowledge graph construction unit is connected to the input end of the new term discovery and verification module.
Citation Information
Patent Citations
Address word segmentation and increment supplement method, system and equipment based on NLP
CN117474002A
Intelligent question answering system based on multi-source medical knowledge retrieval enhancement
CN119180338A
Model fine tuning method and system based on task priority and elastic weight solidification
CN119578859A
Entity relation mining method based on biomedical literature
US20230007965A1
Relation extraction system and method adapted to financial entities and fused with prior knowledge
US20240086650A1
Cited By
Text data intelligent labeling system and method based on large language model
CN121787363A
A text data intelligent labeling system and method based on a large language model
CN121787363B