Electronic medical record data structuring system based on medical language model

Through modular analysis and deep learning models, the problems of diversity of medical text structure and scarcity of labeled data are solved, efficient structure of electronic medical records is achieved, and accurate medical knowledge graphs are generated.

CN120260952APending Publication Date: 2025-07-04WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510316493.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing deep learning models are difficult to adapt to the diversity of medical documents when processing medical texts, complex relationships between medical entities and scarce high-quality labeled data, making it difficult to achieve efficient and accurate data analysis in electronic medical records structure.

Method used

The modular analysis module is used to separate electronic medical records as the main complaints and past history and other medical text modules. Combined with data annotation and deep learning models, semantic features are extracted through pre-trained language models, attention mechanisms and dynamic programming segmentation algorithms, set weighted loss functions to optimize model performance, and generate structured knowledge graphs.

Benefits of technology

It improves the accuracy and robustness of medical text segmentation, significantly improves the accuracy of entity recognition and the ability to identify complex relationships, and builds a complete medical knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260952A_ABST
    Figure CN120260952A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data structuring, and discloses an electronic medical record data structuring system based on a medical language model, and the system comprises a modular analysis module which is used for extracting an electronic medical record text and separating the electronic medical record text into medical text modules such as a main complaint module, a past history module and a current medical history module; the data annotation module is used for annotating the entities in each module and the semantic relationship between the entities; the deep learning model module is used for constructing a deep learning model for medical entity and relation extraction and training the model; and the structured knowledge generation module is used for obtaining the electronic medical record, analyzing the electronic medical record through a text data structured model after the electronic medical record is modularly analyzed, and generating a structured knowledge graph containing entity attributes and semantic relationships. According to the method, through modular design and specialized processing, the problems of medical text structure diversity, professional term complexity, data imbalance and the like are effectively solved, and the accuracy and efficiency of electronic medical record structuring are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data structuring, and more specifically, to an electronic medical record data structuring system based on a medical language model. Background Art

[0002] With the rapid development of medical informatization, electronic medical record systems have been widely used in medical institutions, accumulating a vast amount of clinical data. However, most of this data exists in the form of unstructured text and is difficult to directly use for computer analysis and knowledge mining. Transforming electronic medical record text into structured medical knowledge has become a key challenge in medical information processing.

[0003] In recent years, deep learning technology has made remarkable progress in the field of natural language processing. Currently, common models for entity and relationship extraction mainly include CNN, RNN, and the multi-task self-attention mechanism BiLSTM-CRF, etc. These models use a large amount of unsupervised data to construct multi-layer neural networks and can automatically learn text feature representations. However, when dealing with medical text, due to the complexity of medical terminology, these models have limitations: for example, different types of medical documents have different language characteristics, and existing models are difficult to adapt to this diversity; the relationships between medical entities are complex and diverse, and existing models have limited effects in dealing with complex relationships; relatively scarce high-quality medical text annotation data, etc. Summary of the Invention

[0004] In order to overcome the problems of the prior art, the present invention proposes an electronic medical record data structuring system based on a medical language model to solve the above problems.

[0005] The present invention provides the following technical solutions:

[0006] An electronic medical record data structuring system based on a medical language model, comprising:

[0007] A modular parsing module, configured to extract the text content in the electronic medical record and perform modular parsing, separating the text content into different medical text modules, where the medical text modules include a chief complaint module, a past history module, a personal history module, a menstrual history module, a family history module, and a current history module;

[0008] A data annotation module, configured to annotate the entities and the semantic relationships between entities in each module to obtain a plurality of annotation data to form an annotation data set;

[0009] A deep learning model module, configured to build a deep learning model for medical entity and relationship extraction, and use the annotation data set to train the model to obtain a text data structuring model;

[0010] A structured knowledge generation module, which is used to obtain electronic medical records that need to be structured, parse the electronic medical records modularly to obtain different medical text modules, and parse the different medical text modules through the text data structuring model to generate a structured knowledge graph containing entity attributes and semantic relationships.

[0011] Preferably, the specific steps of the modular parsing include:

[0012] Use a pre-trained language model to process the medical record text and extract the semantic vector representation between paragraphs;

[0013] Calculate the attention scores of keywords in the text and the surrounding text;

[0014] Apply a segmentation algorithm to calculate the boundary positions of text paragraphs and generate boundary coordinates;

[0015] Segment the original electronic medical record text according to the boundary coordinates to obtain different medical texts.

[0016] Preferably, the steps for annotating entities in each module include:

[0017] Define dedicated entity type tags for each medical text module, and select entity text fragments in the text through an annotation platform and assign corresponding entity type tags;

[0018] Among them, the dedicated entity types defined for each medical text module include:

[0019] Define entity type tags for symptoms, duration, and severity for the chief complaint module;

[0020] Define entity type tags for disease name, diagnosis time, treatment method, and drug name for the past history module;

[0021] Define entity type tags for living habits, occupational exposure, and allergy history for the personal history module;

[0022] Define entity type tags for age at menarche, cycle, menstrual volume, and menopause time for the menstrual history module;

[0023] Define entity type tags for kinship, genetic disease, and age of onset for the family history module;

[0024] Define entity type tags for symptoms, onset time, evolution process, and diagnosis and treatment process for the current history module.

[0025] Preferably, the steps for annotating the semantic relationships between entities in each module include:

[0026] Set entity relationship type tags, including time relationship, causal relationship, attribute relationship, modification relationship, and transformation relationship;

[0027] Formulate relationship annotation guidelines and list the judgment criteria for various types of relationships;

[0028] Use a relationship annotation tool, select the annotated entity pairs, and assign relationship type labels to them according to the relationship annotation guidelines;

[0029] Organize the annotation results into a triple data format of entity, relationship, entity.

[0030] Preferably, the specific steps for constructing a deep learning model for medical entity and relationship extraction include:

[0031] Design a modular neural network architecture, including dedicated entity recognition sub-networks and relationship extraction sub-networks for each medical text module;

[0032] Construct specific entity recognition layers for each medical text module, corresponding to the dedicated entity types of the chief complaint module, past history module, personal history module, menstrual history module, family history module, and current history module respectively;

[0033] Design a general relationship extraction network for identifying temporal relationships, causal relationships, attribute relationships, modification relationships, and transformation relationships in each module;

[0034] Construct a module selection mechanism to automatically activate the corresponding dedicated entity recognition sub-network according to the input text type;

[0035] Integrate all sub-networks to form an end-to-end modular medical text structuring model.

[0036] Preferably, the steps for training the model using the annotated dataset include:

[0037] Group the annotated dataset according to the medical text module type to form dedicated training subsets for the chief complaint module, past history module, personal history module, menstrual history module, family history module, and current history module;

[0038] Perform data augmentation on each dedicated training subset, including synonym replacement, entity transformation, and syntactic reconstruction;

[0039] Adopt a phased training strategy, first train the dedicated entity recognition sub-networks of each module, and then train the general relationship extraction network;

[0040] Set a weighted loss function to balance the training objectives of each medical text module and optimize the recognition accuracy of low-frequency entity types and complex relationship types;

[0041] Implement cross-validation, evaluate the model performance for each medical text module, and select the optimal model parameters based on the comprehensive F1 score;

[0042] Perform model integration to integrate the dedicated models of each medical text module into a unified text data structuring model.

[0043] Preferably, the specific steps of setting the weighted loss function include:

[0044] Count the occurrence frequencies of different entity types and relationship types in each medical text module;

[0045] Calculate the reciprocal of the frequency of each entity type and relationship type as the initial weight value;

[0046] Set the entity recognition loss function Le and the relationship extraction loss function Lr, and adopt weighted cross-entropy loss respectively;

[0047] For the entity recognition loss function Le, calculate according to the formula where n represents the number of entity types, α i is the weight coefficient of the i-th type of entity, y i is the true label of the i-th type of entity, and p i is the predicted probability of the i-th type of entity label;

[0048] For the relationship extraction loss function Lr, calculate according to the formula where m represents the number of relationship types, β j is the weight coefficient of the j-th type of relationship, z j is the true label of the j-th type of relationship, and q j is the predicted probability of the j-th type of relationship;

[0049] During the training process, adjust α i and β j according to the F1 score of each type on the validation set;

[0050] Combine the entity recognition loss Le and the relationship extraction loss Lr into the total loss function L. The formula of the total loss function L is L = λ1×Le + λ2×Lr, where λ1 and λ2 are the preset entity balance factor and relationship balance factor respectively.

[0051] The present invention provides an electronic medical record data structuring system based on a medical language model, which has the following beneficial effects:

[0052] By separating different medical text modules (such as chief complaint, past history, personal history, etc.) in the electronic medical record, it lays a foundation for subsequent refined processing. This module uses a pre-trained medical language model to extract text semantic features, combines the attention mechanism and the dynamic programming segmentation algorithm, can accurately capture the boundaries of different medical text modules, effectively solves the problems of diverse medical text structures and inconsistent formats, and improves the accuracy and robustness of text segmentation.

[0053] Through dedicated entity type tags and annotation specifications, such as entity types that define symptoms, duration, and severity for the chief complaint module, and entity types that define disease names, diagnosis time, etc. for the past medical history module. This modular annotation method can capture key information in various medical texts more accurately, significantly improving the accuracy of entity recognition. At the same time, by defining various semantic relationship types such as time relationships and causal relationships, the system can comprehensively express the complex associations between medical entities, providing a solid foundation for constructing a complete medical knowledge graph.

[0054] By constructing a modular deep learning model architecture, various medical entities and relationships can be extracted specifically. Especially by setting a weighted loss function and a dynamic weight adjustment mechanism, the problem of uneven distribution of medical entities and relationships is effectively solved, and the recognition ability for low-frequency entities and complex relationships is improved. Brief Description of the Drawings

[0055] Figure 1 It is a schematic diagram of the modules of an electronic medical record data structuring system based on a medical language model of the present invention. Detailed Embodiments

[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0057] Embodiment 1

[0058] Please refer to Figure 1 , in this embodiment, an electronic medical record data structuring system based on a medical language model includes:

[0059] A modular parsing module for extracting the text content in the electronic medical record and performing modular parsing, separating the text content into different medical text modules, and the medical text modules include a chief complaint module, a past medical history module, a personal history module, a menstrual history module, a family history module, and a current medical history module;

[0060] The specific steps of the modular parsing include:

[0061] Using a pre-trained language model to process the medical record text and extract the semantic vector representation between paragraphs;

[0062] Calculating the attention scores of keywords in the text and the surrounding text;

[0063] Applying a segmentation algorithm to calculate the boundary positions of text paragraphs and generating boundary coordinates;

[0064] The original electronic medical record text is segmented according to the boundary coordinates to obtain different medical texts.

[0065] In this embodiment, the modular parsing module is used to extract the text content in the electronic medical record and perform modular parsing. The process of separating the text content into different medical text modules is described as follows

[0066] First, the system receives a complete electronic medical record text data. These data may come from the hospital's electronic medical record system, and the format may be structured or semi-structured forms such as plain text, JSON, or XML. The system first converts these data into a unified plain text format, removes special characters and format tags, and retains the original paragraph structure and text content.

[0067] Next, a pre-trained medical domain language model is used to process the medical record text and extract the semantic vector representation between paragraphs. Specifically, in implementation, the system adopts a Chinese medical pre-trained language model based on the BERT architecture (such as ChineseBERT-Med or MC-BERT). This model has been pre-trained on a large number of Chinese medical literature and medical record data and can effectively capture the semantic features of medical texts. The system divides the medical record text into natural paragraphs, and each paragraph is used as an input unit and sent into the pre-trained language model. For each paragraph, the model outputs a 768-dimensional vector representation (taking the BERT-base model as an example). These vectors represent the semantic information of the paragraphs, and paragraphs with similar semantics are closer in the vector space.

[0068] Then, the system calculates the attention scores of the keywords in the text with the surrounding text. Specifically, in implementation, first, a medical domain keyword dictionary is constructed, which contains the characteristic vocabulary of each medical text module (chief complaint, past history, personal history, menstrual history, family history, and current history). For example, the characteristic words of the chief complaint module include chief complaint, reason for seeking medical treatment, symptoms, etc.; the characteristic words of the past history module include past history, past diseases, surgical history, etc. The system uses the attention mechanism of the pre-trained language model to calculate the attention scores of these keywords in each paragraph with the surrounding text. Specifically, for each paragraph, the attention weight matrix of the last layer of the pre-trained model is extracted, and the average attention score between the keyword token and other tokens is calculated. These attention scores reflect the semantic association strength between the keywords and the surrounding text and help to identify the module boundaries.

[0069] Next, the system applies a segmentation algorithm to calculate the boundary positions of text paragraphs and generates boundary coordinates. In this embodiment, a text segmentation algorithm based on dynamic programming is adopted. The algorithm first constructs a paragraph similarity matrix, where each element in the matrix represents the semantic similarity between two paragraphs. Then, the system uses the dynamic programming method to find the optimal segmentation points, so that the internal semantic similarity of each segmented module is high, while the semantic similarity between modules is low. Specifically, a cost function is set, which comprehensively considers paragraph cohesion and keyword features. The algorithm outputs a set of boundary coordinates, representing the starting and ending paragraph indices of each medical text module.

[0070] Finally, the system segments the original electronic medical record text according to the boundary coordinates to obtain different medical text modules. For each pair of boundary coordinates (start_i, end_i), the system extracts all the content from the start_i-th paragraph to the end_i-th paragraph in the original text to form a medical text module. Then, the system classifies these modules using a method that combines rules and machine learning, and labels them as the chief complaint module, past history module, personal history module, menstrual history module, family history module, or current history module. Through the above steps, the system can effectively separate the electronic medical record text into different medical text modules, laying a foundation for subsequent data annotation and structured processing.

[0071] The data annotation module is used to annotate the entities and the semantic relationships between entities in each module to obtain multiple annotation data to form an annotation dataset;

[0072] The steps for annotating the entities in each module include:

[0073] Define dedicated entity type labels for each medical text module, and select entity text fragments in the text through the annotation platform and assign the corresponding entity type labels;

[0074] Among them, defining dedicated entity types for each medical text module includes:

[0075] Define entity type labels for symptoms, duration, and severity for the chief complaint module;

[0076] Define entity type labels for disease name, diagnosis time, treatment method, and drug name for the past history module;

[0077] Set entity type labels for living habits, occupational exposure, and allergy history for the personal history module;

[0078] Set entity type labels for age of menarche, cycle, menstrual volume, and menopause time for the menstrual history module;

[0079] Set entity type labels for kinship, genetic diseases, and age of onset for the family history module;

[0080] Set entity type tags for symptoms, onset time, evolution process, and diagnosis and treatment process in the current medical history module.

[0081] The steps for annotating the semantic relationships between entities in each module include:

[0082] Set relationship type tags between entities, including time relationship, causal relationship, attribute relationship, modification relationship, and transformation relationship;

[0083] Formulate relationship annotation guidelines and list the determination criteria for various relationships;

[0084] Use a relationship annotation tool to select the annotated entity pairs and assign relationship type tags to them according to the relationship annotation guidelines;

[0085] Organize the annotation results into a triple data format of entity, relationship, entity.

[0086] In this embodiment, the data annotation module is used to annotate the entities and the semantic relationships between entities in each medical text module, and the process of obtaining multiple annotation data to form an annotation dataset is as follows.

[0087] First, the system annotates the entities in each module. For each medical text module, the system defines dedicated entity type tags, and these tags reflect the common medical information types in each module. Specifically, in implementation, for the chief complaint module, the system defines three types of entity type tags: symptoms, duration, and severity. For example, in the text where the patient's chief complaint is persistent pain in the right lower abdomen for 3 days and the pain level is moderate, persistent pain in the right lower abdomen is annotated as a symptom, 3 days is annotated as the duration, and moderate is annotated as the severity. The annotator completes the annotation by selecting the corresponding text fragment on the annotation platform and selecting the corresponding entity type from the preset drop-down menu. And similarly annotate the past medical history module, personal history module, menstrual history module, family history module, and current medical history module.

[0088] Next, the system annotates the semantic relationships between entities in each module. First, the system sets relationship type tags between entities, including time relationship, causal relationship, attribute relationship, modification relationship, and transformation relationship. The time relationship represents the chronological order of events, the causal relationship represents the causal connection between events, the attribute relationship represents the attributes or characteristics of entities, the modification relationship represents the limitation or description of entities, and the transformation relationship represents the change or transformation of entity states.

[0089] The system has developed detailed relationship annotation guidelines, listing the criteria for determining various types of relationships and typical examples. For example, the criteria for determining time relationships include: (1) two events are clearly shown to have a sequential relationship in the text; (2) the use of time connectives such as "after", "subsequently", etc.; (3) the time order can be inferred from the context. The annotation guidelines also contain a large number of examples to help annotators understand and apply these criteria.

[0090] On the annotation platform, annotators use the relationship annotation tool to select the already annotated entity pairs and assign relationship type tags to them according to the relationship annotation guidelines. For example, in the text "The patient had a cough 3 days after having a fever", the relationship between "having a fever" and "having a cough" is annotated as a time relationship; in the text "Long-term smoking leads to a decline in lung function", the relationship between "long-term smoking" and "a decline in lung function" is annotated as a causal relationship.

[0091] Finally, the system organizes the annotation results into a triple data format of entity, relationship, entity. For example, for the text "The patient's lung function declined due to long-term smoking", the system generates the following triple: (long-term smoking [lifestyle habit], causes [causal relationship], decline in lung function [symptom]). This triple format is convenient for subsequent data storage, query, and analysis, and is also suitable as training data for deep learning models.

[0092] Through the above steps, the system has obtained a large amount of high-quality annotation data, forming a comprehensive annotation dataset, providing a reliable data basis for the training of subsequent deep learning models.

[0093] The deep learning model module is used to construct a deep learning model for medical entity and relationship extraction, and use the annotation dataset to train the model to obtain a text data structuring model;

[0094] The specific steps for constructing the deep learning model for medical entity and relationship extraction include:

[0095] Design a modular neural network architecture, including dedicated entity recognition sub-networks and relationship extraction sub-networks for each medical text module;

[0096] Construct specific entity recognition layers for each medical text module, corresponding to the dedicated entity types of the chief complaint module, past history module, personal history module, menstrual history module, family history module, and current history module respectively;

[0097] Design a general relationship extraction network for identifying time relationships, causal relationships, attribute relationships, modification relationships, and transformation relationships in each module;

[0098] Construct a module selection mechanism to automatically activate the corresponding dedicated entity recognition sub-network according to the input text type;

[0099] Integrate all sub-networks to form an end-to-end modular medical text structuring model.

[0100] The steps of training the model using the labeled data set include:

[0101] Group the labeled data set according to the medical text module type to form dedicated training subsets for the chief complaint module, past history module, personal history module, menstrual history module, family history module, and current history module;

[0102] Perform data augmentation on each dedicated training subset, including synonym replacement, entity transformation, and syntactic reconstruction;

[0103] Adopt a phased training strategy, first train the dedicated entity recognition sub-networks of each module, and then train the general relation extraction network;

[0104] Set a weighted loss function to balance the training objectives of each medical text module and optimize the recognition accuracy of low-frequency entity types and complex relation types;

[0105] Implement cross-validation, evaluate the model performance for each medical text module, and select the optimal model parameters based on the comprehensive F1 score;

[0106] Perform model integration, and integrate the dedicated models of each medical text module into a unified text data structuring model.

[0107] The specific steps of setting the weighted loss function include:

[0108] Count the occurrence frequencies of different entity types and relation types in each medical text module;

[0109] Calculate the reciprocal of the frequency of each entity type and relation type as the initial weight value;

[0110] Set the entity recognition loss function Le and the relation extraction loss function Lr, and adopt weighted cross-entropy loss respectively;

[0111] For the entity recognition loss function Le, calculate according to the formula where n represents the number of entity types, α i is the weight coefficient of the i-th type of entity, y i is the true label of the i-th type of entity, and p i is the predicted probability of the i-th type of entity label;

[0112] For the relation extraction loss function Lr, calculate according to the formula where m represents the number of relation types, β j is the weight coefficient of the j-th type of relation, z j is the true label of the j-th type of relation, and q jis the predicted probability of the j-th type of relationship;

[0113] During the training process, adjust α according to the F1 scores of each type on the validation set i and β j ;

[0114] Combine the entity recognition loss Le and the relationship extraction loss Lr into the total loss function L. The formula for the total loss function L is L = λ1×Le + λ2×Lr, where λ1 and λ2 are the preset entity balance factor and relationship balance factor respectively.

[0115] In this embodiment, the deep learning model module is used to construct a deep learning model for medical entity and relationship extraction, and the process of training the model using the labeled dataset to obtain the text data structuring model is as described below.

[0116] First, the system constructs a deep learning model for medical entity and relationship extraction. Specifically, the system designs a modular neural network architecture, which includes dedicated entity recognition sub-networks and relationship extraction sub-networks for each medical text module.

[0117] The system constructs a specific entity recognition layer for each medical text module. For the chief complaint module, the entity recognition layer specifically recognizes symptom, duration, and severity entities; for the past history module, the entity recognition layer specifically recognizes disease name, diagnosis time, treatment method, and drug name entities; for the personal history module, the entity recognition layer specifically recognizes living habits, occupational exposure, and allergy history entities; for the menstrual history module, the entity recognition layer specifically recognizes age of menarche, cycle, menstrual volume, and menopause time entities; for the family history module, the entity recognition layer specifically recognizes kinship, genetic disease, and age of onset entities; for the current illness history module, the entity recognition layer specifically recognizes symptoms, onset time, evolution process, and treatment process entities.

[0118] Each entity recognition layer adopts a BiLSTM-CRF (Bidirectional Long Short-Term Memory Network - Conditional Random Field) structure, where BiLSTM is responsible for capturing the context features of the text, and the CRF layer is responsible for modeling the label sequence to ensure that the output label sequence conforms to the grammar rules. Specifically, the system first uses pre-trained medical word vectors (such as MedicalWord2Vec or MedicalBERT) to convert the input text into a word vector sequence, then extracts the context features through the BiLSTM network, and finally outputs the entity label probability distribution of each word through the CRF layer.

[0119] The system also designs a general relation extraction network to identify temporal relations, causal relations, attribute relations, modification relations, and transformation relations in each module. The relation extraction network adopts a method based on the graph convolutional network (GCN), models the relations between entities as a graph structure, and captures the semantic dependency relations between entities through multiple layers of GCN. In specific implementation, the system first constructs a syntactic dependency tree, takes the entities in the sentence as the nodes of the graph, takes the syntactic dependency relations as the edges of the graph, then performs convolutional operations on the graph through GCN, and finally predicts the relation types between entity pairs through a multi-layer perceptron (MLP) classifier. Finally, the system integrates all sub-networks to form an end-to-end modular medical text structuring model.

[0120] Next, the system uses the labeled dataset to train the model. First, the system groups the labeled dataset according to the medical text module types to form dedicated training subsets for the chief complaint module, past history module, personal history module, menstrual history module, family history module, and current history module. Each dedicated training subset contains the text of the module type and its corresponding entity and relation annotations.

[0121] Then, the system performs data augmentation on each dedicated training subset, including synonym replacement, entity transformation, and syntactic reconstruction. Synonym replacement means using a medical synonym dictionary to replace some words in the text with their synonyms, such as replacing headache with cephalalgia; entity transformation means keeping the entity type unchanged and replacing the specific content of the entity, such as replacing diabetes with hypertension; syntactic reconstruction means changing the syntactic structure of the sentence while keeping the semantics unchanged, such as changing the sentence "The patient had a fever due to a cold" to "The patient had a fever, and the reason was a cold". These data augmentation techniques can expand the training dataset and improve the generalization ability of the model.

[0122] The system adopts a phased training strategy. First, it trains the dedicated entity recognition sub-networks for each module, and then trains the general relation extraction network. In the first stage, the system uses the dedicated training subsets of each module to train the corresponding entity recognition sub-networks respectively to optimize the entity recognition performance; in the second stage, the system fixes the parameters of the trained entity recognition sub-networks and uses the complete training dataset to train the general relation extraction network; in the third stage, the system jointly fine-tunes the entire model to optimize the end-to-end performance.

[0123] The system sets a weighted loss function to balance the training objectives of each medical text module and optimize the recognition accuracy of low-frequency entity types and complex relation types. In specific implementation, the system first counts the occurrence frequencies of different entity types and relation types in each medical text module, and then calculates the reciprocal of the frequency of each entity type and relation type as the initial weight value. For example, if the occurrence frequency of the allergy history entity in the personal history module is 0.05, then its initial weight value is 1 / 0.05 = 20.

[0124] The system sets the entity recognition loss function and the relation extraction loss function, and both adopt weighted cross-entropy loss. During the training process, the system dynamically adjusts the weight coefficient α according to the F1 scores of each type on the validation set. i and α i . Specifically, for entity types or relation types with F1 scores lower than the preset threshold, the system increases their weight coefficients; for entity types or relation types with F1 scores higher than the preset threshold, the system decreases their weight coefficients. This dynamic adjustment mechanism can guide the model to pay more attention to difficult-to-recognize entity types and relation types, improving the overall performance.

[0125] Finally, the system combines the entity recognition loss and the relation extraction loss into a total loss function.

[0126] The system performs cross-validation, evaluates the model performance for each medical text module, and selects the optimal model parameters based on the comprehensive F1 score. Specifically, the system uses 5-fold cross-validation, divides each dedicated training subset into 5 parts, uses 4 parts as the training set and 1 part as the validation set each time, and conducts 5 rounds of training and validation in turn. For each validation, the system calculates the precision, recall, and F1 scores of entity recognition and relation extraction, and calculates the comprehensive F1 score as the evaluation index of the model performance. Finally, the system selects the model parameters with the highest comprehensive F1 score as the optimal model.

[0127] The system conducts model integration, integrating the dedicated models of each medical text module into a unified text data structuring model. The model integration adopts a voting mechanism. For each prediction task, the system comprehensively considers the prediction results of multiple models and selects the prediction with the highest confidence as the final output. This integration method can reduce the prediction error of a single model and improve the stability and accuracy of the overall prediction.

[0128] Through the above steps, the system constructs and trains a high-performance modular medical text structuring model, which can effectively extract medical entities and their relationships from electronic medical records, providing reliable technical support for subsequent structured knowledge generation.

[0129] The structured knowledge generation module is used to obtain the electronic medical records to be structured, modularly parse the electronic medical records to obtain different medical text modules, and parse the different medical text modules through the text data structuring model to generate a structured knowledge graph containing entity attributes and semantic relationships.

[0130] In this embodiment, the process of the structured knowledge generation module for obtaining the electronic medical records to be structured and converting them into a structured knowledge graph containing entity attributes and semantic relationships is as follows.

[0131] First, the system obtains the electronic medical record text that needs to be structured. These electronic medical records may come from hospital information systems, electronic medical record repositories, or other medical data sources. The system supports the input of electronic medical records in multiple formats, including plain text, XML, JSON, etc., and uniformly converts them into a standard format that the system can process.

[0132] Next, the system calls the aforementioned modular parsing module to perform modular parsing on the electronic medical records, obtaining different medical text modules. During the parsing process, the system identifies and separates key medical information modules such as the chief complaint module, past medical history module, personal history module, menstrual history module, family history module, and current medical history module. Each module contains specific types of medical information and has a relatively independent semantic structure.

[0133] Then, the system uses the aforementioned trained text data structuring model to parse different medical text modules. Specifically, during implementation, the system first identifies the module type, and then activates the corresponding dedicated entity recognition sub-network and general relationship extraction network to process the text within the module. The entity recognition sub-network identifies various medical entities in the module, such as symptoms, diseases, drugs, etc.; the relationship extraction network identifies various semantic relationships between entities, such as temporal relationships, causal relationships, etc.

[0134] Finally, based on the identified entities and relationships, the system generates a structured knowledge graph that includes entity attributes and semantic relationships. The knowledge graph represents the medical information in the medical record in a graphical form, where nodes represent medical entities and edges represent the semantic relationships between entities. The system outputs the knowledge graph in JSON or RDF format for subsequent storage, querying, and analysis. At the same time, the system also provides a visualization interface to intuitively display the structure and content of the knowledge graph, helping medical staff quickly understand the medical record information.

[0135] In several embodiments provided by the present invention, it should be understood that the disclosed system, apparatus, and method can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the units is only one way, and in actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0136] As described above, this is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

[0137] Finally, the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An electronic medical record data structuring system based on a medical language model, characterized in that, Including: A modular parsing module, which is used to extract the text content in the electronic medical record and perform modular parsing, separating the text content into different medical text modules. The medical text modules include a chief complaint module, a past medical history module, a personal history module, a menstrual history module, a family history module, and a current medical history module; A data annotation module, which is used to annotate the entities and the semantic relationships between entities in each module, and obtain a plurality of annotation data to form an annotation data set; A deep learning model module, which is used to build a deep learning model for medical entity and relationship extraction, and train the model using the annotation data set to obtain a text data structuring model; A structured knowledge generation module, which is used to obtain the electronic medical record that needs to be structured, perform modular parsing on the electronic medical record to obtain different medical text modules, and parse the different medical text modules through the text data structuring model to generate a structured knowledge graph containing entity attributes and semantic relationships.

2. The structured system for electronic medical record data based on a medical language model according to claim 1, wherein The specific steps of the modular parsing include: Using a pre-trained language model to process the medical record text and extract the semantic vector representation between paragraphs; Calculating the attention score between the keywords in the text and the surrounding text; Applying a segmentation algorithm to calculate the boundary positions of the text paragraphs and generate boundary coordinates; Segmenting the original electronic medical record text according to the boundary coordinates to obtain different medical texts.

3. The structured system for electronic medical record data based on a medical language model according to claim 2, wherein, The steps of annotating the entities in each module include: Defining a dedicated entity type label for each medical text module, and selecting entity text fragments in the text through an annotation platform and assigning corresponding entity type labels; Among them, the dedicated entity types defined for each medical text module include: Defining entity type labels for symptoms, duration, and severity for the chief complaint module; Defining entity type labels for disease name, diagnosis time, treatment method, and drug name for the past medical history module; Setting entity type labels for living habits, occupational exposure, and allergy history for the personal history module; Setting entity type labels for age at menarche, cycle, menstrual volume, and menopause time for the menstrual history module; Setting entity type labels for kinship, genetic disease, and age of onset for the family history module; Setting entity type labels for symptoms, onset time, evolution process, and treatment course for the current medical history module.

4. The structured system for electronic medical record data based on a medical language model according to claim 3, wherein The steps of annotating the semantic relationships between entities in each module include: Setting entity relationship type labels, including time relationship, causal relationship, attribute relationship, modification relationship, and transformation relationship; Formulating a relationship annotation guide and listing the determination criteria for various relationships; Using a relationship annotation tool, selecting the annotated entity pairs, and assigning relationship type labels to them according to the relationship annotation guide; Organizing the annotation results into a triple data format of entity, relationship, entity.

5. An electronic medical record data structuring system based on a medical language model according to claim 1, characterized in that, The specific steps of building a deep learning model for medical entity and relationship extraction include: Designing a modular neural network architecture, including dedicated entity recognition sub-networks and relationship extraction sub-networks for each medical text module; Building a specific entity recognition layer for each medical text module, corresponding to the dedicated entity types of the chief complaint module, past medical history module, personal history module, menstrual history module, family history module, and current medical history module; Design a general relation extraction network to identify temporal relations, causal relations, attribute relations, modification relations, and transformation relations in each module; Construct a module selection mechanism to automatically activate the corresponding dedicated entity recognition sub-network according to the input text type; Integrate all sub-networks to form an end-to-end modular medical text structuring model.

6. The structured system for electronic medical record data based on a medical language model according to claim 5, characterized in that, The steps of training the model using the labeled data set include: Group the labeled data set according to the medical text module type to form dedicated training subsets for the chief complaint module, past history module, personal history module, menstrual history module, family history module, and current history module; Perform data augmentation on each dedicated training subset, including synonym replacement, entity transformation, and syntactic reconstruction; Adopt a phased training strategy, first train the dedicated entity recognition sub-networks of each module, and then train the general relation extraction network; Set a weighted loss function to balance the training objectives of each medical text module and optimize the recognition accuracy of low-frequency entity types and complex relation types; Implement cross-validation, evaluate the model performance for each medical text module, and select the optimal model parameters based on the comprehensive F1 score; Perform model integration to integrate the dedicated models of each medical text module into a unified text data structuring model.

7. An electronic medical record data structuring system based on a medical language model according to claim 6, characterized in that, The specific steps of setting the weighted loss function include: Count the occurrence frequencies of different entity types and relation types in each medical text module; Calculate the reciprocal of the frequency of each entity type and relation type as the initial weight value; Set the entity recognition loss function Le and the relation extraction loss function Lr, and adopt weighted cross-entropy loss respectively; For the entity recognition loss function Le, it is calculated according to the formula , where n represents the number of entity types, α i is the weight coefficient of the i-th type of entity, y i is the true label of the i-th type of entity, and p i is the predicted probability of the label of the i-th type of entity; For the relation extraction loss function $L_r$, it is calculated according to the formula , where $m$ represents the number of relation types, $\beta$ j is the weight coefficient of the $j$-th type of relation, $z$ j is the true label of the $j$-th type of relation, and $q$ j is the predicted probability of the $j$-th type of relation; Adjust α according to the F1 scores of each type on the validation set during the training process i and β j ; Combine the entity recognition loss Le and the relation extraction loss Lr into the total loss function L. The formula of the total loss function L is L = λ1×Le + λ2×Lr, where λ1 and λ2 are the preset entity balance factor and relation balance factor respectively.

Citation Information

Cited By

  • System and method for converting electric power notification text into 5G message

    CN120429350A