Training data acquisition method, medical record structured model training method, medical record structured model training device and medical record structured equipment

By restoring, translating and identifying terms in English medical record datasets, combining them with biomedical dictionaries and large language models to generate Chinese structured data, the problems of low efficiency and accuracy of Chinese electronic structured medical record datasets are solved, and efficient and accurate training data acquisition is achieved.

CN120748593APending Publication Date: 2025-10-03TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510667491.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing methods for obtaining Chinese electronic structured medical record datasets are either inefficient or inaccurate. Manual labeling is costly and inefficient, and rule migration leads to a significant decrease in accuracy.

Method used

By restoring the medical jargon and term abbreviations in the English electronic medical record dataset to their full names, translating them into Chinese, and performing medical term recognition and attribute information extraction, combined with the biomedical dictionary tree and large language model, Chinese structured data is generated.

Benefits of technology

It has achieved the construction of accurate Chinese structured data from English medical record datasets, improved the efficiency and accuracy of obtaining training data, and reduced the need for manual labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748593A_ABST
    Figure CN120748593A_ABST
Patent Text Reader

Abstract

The invention provides a training data acquisition method, a medical record structured model training method, a medical record structured model training device and medical record structured equipment, and relates to the technical field of computer artificial intelligence. The method comprises the following steps: restoring medical voices and medical term abbreviations in a to-be-processed English medical record set into corresponding full names to obtain a restored English medical record set; translating the restored English medical record set into Chinese to obtain a Chinese translated medical record set; performing medical term recognition and attribute information extraction on each restored English medical record in the restored English medical record set to obtain English structured data of each restored English medical record; and performing information correspondence on the English structured data of each restored English medical record and the Chinese translation medical records in the corresponding Chinese translation medical record set to obtain Chinese training data. According to the method, accurate Chinese structured data can be constructed according to the English electronic medical record data set for model training, and the efficiency and accuracy of obtaining Chinese training data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer artificial intelligence technology, and in particular to a method for acquiring training data, a method and device for training a medical record structuring model, and medical record structuring equipment. Background Art

[0002] The goal of a medical record structuring model is to extract a set of key-value pairs of medically relevant information from electronic medical records (EMRs) and store them in a spreadsheet or database table. This helps standardize and integrate medical data, facilitates information sharing across different medical institutions, reduces the need for patients to repeatedly provide the same information during their visits, and improves medical efficiency. Training an EMR structuring model requires a Chinese EMR dataset. Currently, publicly available structured medical record datasets are all in English; there are no readily available Chinese EMR datasets.

[0003] Currently, obtaining Chinese electronic structured medical record datasets generally uses manual annotation or rule transfer methods. Manual annotation is costly and inefficient, requiring clinical experts to spend several hours to annotate a single medical record. Rule transfer is the direct application of English annotation rules to Chinese medical records, which can lead to a significant decrease in accuracy due to language differences. It can be seen that the existing methods for obtaining Chinese electronic structured medical record datasets are either inefficient or inaccurate. Summary of the Invention

[0004] The present invention provides a method for acquiring training data, a training method and apparatus for a medical record structuring model, and medical record structuring equipment, to address the defects of the prior art methods for acquiring Chinese electronic structured medical record data sets, which are either low in efficiency or low in accuracy. The invention realizes the construction of Chinese training data for a Chinese electronic structured medical record data set based on an English electronic medical record data set, thereby improving the efficiency and accuracy of acquiring Chinese training data.

[0005] The present invention provides a method for obtaining training data, comprising the following steps: Restoring medical jargons and abbreviations of medical terms in the to-be-processed English medical record set to their corresponding full names, thereby obtaining a restored English medical record set; translating the restored English medical record set into Chinese to obtain a Chinese translated medical record set; performing medical term recognition and attribute information extraction on each restored English medical record in the restored English medical record set to obtain English structured data for each restored English medical record, wherein the English structured data includes the medical term and attribute information related to the medical term, wherein the attribute information is used to represent condition information and / or treatment information related to the medical term; The English structured data of each restored English medical record is matched with the Chinese translated medical record in the corresponding Chinese translated medical record set to obtain Chinese structured data, which is the training data.

[0006] According to a method for acquiring training data provided by the present invention, medical term recognition and attribute information extraction are performed on each restored English medical record in the restored English medical record set to obtain English structured data of each restored English medical record, including: Performing medical term recognition on each restored English medical record according to a biomedical dictionary tree to obtain medical terms for each restored English medical record, wherein the biomedical dictionary tree is constructed according to a biomedical information ontology system; Performing narrative state analysis and related attribute extraction based on the medical terms in each restored English medical record and the corresponding English medical record text, to obtain narrative state attribute information and other attribute information corresponding to the medical terms in each restored English medical record; Taking the medical terms of each restored English medical record as core information, the narrative state attribute information and the other attribute information corresponding to the medical terms are integrated to obtain English structured data of each restored English medical record.

[0007] According to a method for acquiring training data provided by the present invention, the biomedical dictionary tree is Trie dictionary tree data, each edge of the Trie dictionary tree data represents an alphabetic character, and the path from the root node to any node constitutes a medical term string. Medical term recognition is performed on each restored English medical record based on the biomedical dictionary tree to obtain the medical term for each restored English medical record. The biomedical dictionary tree is constructed based on a biomedical information ontology system, and includes: Performing a maximum forward matching scan on each character string in each restored English medical record in the Trie dictionary tree data using a maximum forward matching algorithm to determine the initial medical term in each restored English medical record; Common terms are deleted from the initial medical terms in each restored English medical record according to a preset English common word list to obtain the medical terms in each restored English medical record, and the medical terms are linked to the term codes in the biomedical information ontology system corresponding to the medical terms.

[0008] According to a method for acquiring training data provided by the present invention, narrative state analysis and related attribute extraction are performed based on the medical terms in each restored English medical record and the corresponding English medical record text, to obtain narrative state attribute information and other attribute information corresponding to the medical terms in each restored English medical record, including: Using a pre-trained narrative state analysis model to perform narrative state analysis on the medical terms in each restored English medical record and the corresponding medical terms in each restored English medical record, outputting narrative state attribute information of each medical term in each restored English medical record, wherein the narrative state attribute information includes at least one of an existential state, a temporal state, and a subject state, the narrative state analysis model being a generative large language model trained based on a clinical narrative state annotation dataset; A general large language model is used to extract other attribute information from the medical terms in each restored English medical record and the corresponding medical terms in each restored English medical record to obtain other attribute information corresponding to each medical term in each restored English medical record.

[0009] According to a method for acquiring training data provided by the present invention, the medical terms of each restored English medical record are used as core information, and the narrative state attribute information and other attribute information corresponding to the medical terms are integrated to obtain English structured data of each restored English medical record, including: Obtaining a data structure corresponding to each medical term in each restored English medical record, wherein the data structure includes term name, semantic type, narrative state, body part, modifier, value, and unit fields; Filling the description state attribute information and other attribute information corresponding to each medical term into the corresponding fields of the data structure, and leaving empty values ​​for the fields corresponding to the attributes not mentioned; The data structure of all medical terms in each restored English medical record is organized in a list form to form the English structured data.

[0010] According to a method for obtaining training data provided by the present invention, the English structured data of each restored English medical record is matched with the Chinese translated medical record in the corresponding Chinese translated medical record set to obtain the training data, including: Establishing a correspondence between the medical term, the narrative state attribute information, and each attribute value in the other attribute information in the English structured data and the original word or original phrase in the corresponding Chinese translated medical record; If any attribute value does not have a directly corresponding original word or original phrase in the Chinese translated medical record, then translate it using a preset Chinese-English bilingual word list to obtain a translated word or translated phrase, and establish a corresponding relationship between it and the translated word or translated phrase; The Chinese structured data is generated according to the corresponding relationship.

[0011] According to a method for acquiring training data provided by the present invention, before restoring the medical jargon and medical term abbreviations in the to-be-processed English medical record set to their corresponding full names to obtain the restored English medical record set, the method further comprises: Access to publicly available English electronic medical record datasets; The English electronic medical records in the English electronic medical record dataset are de-privacyed and format-normalized using a natural language processing model to obtain the English medical record dataset to be processed. The present invention also provides a method for training a medical record structured model, comprising the following steps: The training data is obtained by using the above-mentioned training data acquisition method, wherein the training data includes Chinese electronic medical record text and its corresponding structured annotations; Taking a pre-trained Chinese language model as a base model, and using the training data to train the base model through supervised fine-tuning; The model parameters of the basic model are optimized by minimizing the difference between the structured prediction results output by the model and the structured annotations, thereby obtaining a trained medical record structured model.

[0012] The present invention also provides a device for acquiring training data, comprising the following modules: A medical record restoration module is used to restore medical jargon and medical term abbreviations in the to-be-processed English medical record set to their corresponding full names, thereby obtaining a restored English medical record set; A medical record translation module, configured to translate the restored English medical record set into Chinese to obtain a Chinese translated medical record set; a structured data module, configured to perform medical term recognition and attribute information extraction on each restored English medical record in the restored English medical record set, thereby obtaining English structured data for each restored English medical record, wherein the English structured data includes the medical term and attribute information related to the medical term, wherein the attribute information is used to represent condition information and / or treatment information related to the medical term; The data acquisition module is used to match the English structured data of each restored English medical record with the corresponding Chinese translated medical record in the Chinese translated medical record set to obtain Chinese structured data, which is the training data.

[0013] The present invention also provides a training device for a medical record structured model, comprising the following modules: A training data module, configured to obtain training data using the above-mentioned training data acquisition method, wherein the training data includes Chinese electronic medical record text and its corresponding structured annotations; A module training module, configured to use a pre-trained Chinese language model as a base model and train the base model through supervised fine-tuning using the training data; The model optimization module is used to optimize the model parameters of the basic model by minimizing the difference between the structured prediction results output by the model and the structured annotations, so as to obtain a trained medical record structured model.

[0014] The present invention also provides a medical record structuring device, including a memory, a processor, and a medical record structuring model stored in the memory and runnable on the processor. The medical record structuring model is obtained by obtaining training data using the above-mentioned training data acquisition method and performing model training using the training data.

[0015] The present invention also provides a non-transitory computer-readable storage medium on which a medical record structured model is stored. The medical record structured model is obtained by obtaining training data using the above-mentioned training data acquisition method and performing model training using the training data.

[0016] The present invention provides a method for acquiring training data, a method for training a medical record structuring model, an apparatus, and a medical record structuring device, which construct Chinese training data for a Chinese electronic structured medical record dataset based on an English electronic medical record dataset, thereby improving the efficiency and accuracy of acquiring Chinese training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 It is a flowchart of the method for obtaining training data provided by the present invention.

[0019] Figure 2 4 is a flow chart of the method for obtaining English structured data provided by the present invention.

[0020] Figure 3 This is a flowchart of obtaining structured data of Chinese electronic medical records provided by the present invention.

[0021] Figure 4 This is a flow chart of the medical record structured model training provided by the present invention.

[0022] Figure 5 It is a structural diagram of the device for acquiring training data provided by the present invention.

[0023] Figure 6 It is a structural diagram of the device for acquiring training data provided by the present invention.

[0024] Figure 7 It is a schematic diagram of the physical structure of the medical record structuring device provided by the present invention. DETAILED DESCRIPTION

[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0026] The goal of a medical record structuring model is to extract a set of key-value pairs of medically relevant information from electronic medical records (EMRs) and store them in a spreadsheet or database table. This helps standardize and integrate medical data, facilitates information sharing across different medical institutions, reduces the need for patients to repeatedly provide the same information during their visits, and improves medical efficiency. Training an EMR structuring model requires a Chinese EMR dataset. Currently, publicly available structured medical record datasets are all in English; there are no readily available Chinese EMR datasets.

[0027] Currently, obtaining Chinese electronic structured medical record datasets typically involves manual annotation or rule transfer. Manual annotation is costly and inefficient, requiring clinical experts to spend hours annotating a single medical record. Rule transfer involves directly applying English annotation rules to Chinese medical records, but this can significantly reduce accuracy due to language differences. Therefore, existing methods for obtaining Chinese electronic structured medical record datasets are either inefficient or inaccurate. Understandably, the low accuracy of Chinese electronic structured medical record datasets also makes it difficult to implement medical record structuring models based on machine learning and deep learning.

[0028] In view of this, an embodiment of the present invention provides a method for obtaining training data, which includes: restoring medical jargon and medical term abbreviations in a to-be-processed English medical record set to their corresponding full names to obtain a restored English medical record set; translating the restored English medical record set into Chinese to obtain a Chinese translated medical record set; performing medical term identification and attribute information extraction on each restored English medical record in the restored English medical record set to obtain English structured data for each restored English medical record; and matching the English structured data of each restored English medical record with the corresponding Chinese translated medical record in the Chinese translated medical record set to obtain Chinese structured data, which serves as the training data. This method can construct accurate Chinese structured data for model training based on an English electronic medical record dataset, thereby improving the efficiency and accuracy of obtaining training data.

[0029] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.

[0030] Figure 1 : is a flow chart of the method for obtaining training data provided by the present invention. The method for obtaining training data can be applied to electronic devices, which can be various types of devices with information processing capabilities during implementation. For example, the electronic device can include a personal computer, a laptop, a PDA or a server, etc.; the electronic device can also be a mobile terminal, for example, the mobile terminal can include a mobile phone, a car computer, a tablet computer or a projector, etc. Figure 1 As shown, the method includes the following: Step 101: restore the medical jargon and medical term abbreviations in the to-be-processed English medical record set to their corresponding full names, thereby obtaining a restored English medical record set.

[0031] It should be noted that the English medical record set to be processed includes a large number of medical term abbreviations and jargon, which will affect the subsequent data processing. Therefore, it is first necessary to restore the medical jargon and medical term abbreviations in the English medical record set to be processed to the corresponding full names in order to facilitate subsequent accurate data processing.

[0032] The MIMIC-III public English dataset can be used as the English medical record set to be processed. There are many methods for converting medical jargon and abbreviations in the English medical record set to their corresponding full names. For example, a general large model can be used to convert the jargon and abbreviations to their full names. The present invention does not limit the method for converting medical jargon and abbreviations in the English medical record set to their corresponding full names.

[0033] Step 102: Translate the restored English medical record set into Chinese to obtain a Chinese translated medical record set.

[0034] It should be noted that there are many ways to translate the restored English medical record set into Chinese, such as using translation software or manual translation. The present invention does not limit the method of translating the restored English medical record set into Chinese. For example, a universal large model can be used to translate the restored English medical record set into Chinese.

[0035] Step 103: Perform medical term identification and attribute information extraction on each restored English medical record in the restored English medical record set to obtain English structured data of each restored English medical record, wherein the English structured data includes the medical term and attribute information related to the medical term, and the attribute information is used to represent the condition information and / or treatment information related to the medical term.

[0036] It should be noted that there are many ways to identify medical terms and extract attribute information from each restored English medical record in the restored English medical record collection, such as using software programs or models. The present invention does not limit the method for identifying medical terms and extracting attribute information from each restored English medical record in the restored English medical record collection. After obtaining the medical terms and attribute information, data structuring is performed by integrating all relevant attribute information with the medical terms as the primary information, thereby obtaining structured English data.

[0037] For example, the structuring of English electronic medical records mainly revolves around identifying medical terms and analyzing their semantics and attributes, converting natural language statements that are difficult to use directly into a database from which information can be efficiently extracted. Common modules may include named entity recognition, entity linking or attribute standardization, narrative state analysis, body parts and modifiers, values ​​and units, etc.

[0038] Step 104: Match the English structured data of each restored English medical record with the corresponding Chinese translated medical record in the Chinese translated medical record set to obtain Chinese structured data, which is the training data.

[0039] For example, the correspondence between English integrated data and Chinese translated medical records can be achieved through a universal large model, thereby obtaining Chinese integrated data.

[0040] It is understandable that there is almost no Chinese electronic medical record data available in publicly available datasets. Therefore, the goal of the present invention is to construct sample pairs of medical record-structured data from English medical record datasets and, through translation and precise correspondence, convert them into data with a format and distribution similar to Chinese electronic medical records. The present invention provides a process for constructing a large amount of training data when Chinese (and all low-resource languages) medical record data is insufficient. By restoring and translating the jargon and abbreviations in the English medical record dataset to be processed, and then performing medical terminology recognition and attribute information extraction on each restored English medical record in the restored English medical record dataset to obtain English structured data, and then performing information correspondence between the English structured data and the translated Chinese medical record, the Chinese structured data, i.e., training data, is obtained. This improves the efficiency and accuracy of obtaining training data, thereby enabling the training of a medical record structured model.

[0041] In some embodiments, before restoring the medical jargon and medical term abbreviations in the English medical record set to be processed to the corresponding full names to obtain the restored English medical record set, the method may further include: obtaining a public English electronic medical record data set; using a natural language processing model to de-privacy and format normalize the English electronic medical records in the English electronic medical record data set to obtain the English medical record set to be processed.

[0042] It's important to note that the MIMIC-III public English dataset can be used for medical record acquisition. First, a large model, such as GPT, is used to convert de-identified information, common hospital jargon, and drug and disease abbreviations contained in the MIMIC data into full names. Furthermore, line breaks in the MIMIC medical records that are not semantically necessary are removed. This step yields English medical records with clear semantics and no formatting interference.

[0043] The GPT model is then asked to translate the English data that has been cleaned in the previous step into Chinese electronic medical records for backup. Next, the structured results obtained in English are first used and then mapped back to the Chinese electronic medical records to obtain the Chinese structured data, which is the training data.

[0044] Figure 2 Schematic diagram of the method for obtaining English structured data provided by the present invention. Figure 2 As shown, the medical term recognition and attribute information extraction are performed on each restored English medical record in the restored English medical record set to obtain English structured data of each restored English medical record, which may include: Step 201: performing medical term recognition on each restored English medical record according to a biomedical dictionary tree to obtain the medical terms of each restored English medical record, wherein the biomedical dictionary tree is constructed according to a biomedical information ontology system.

[0045] It should be noted that the biomedical information ontology system (BIOS) has a very large vocabulary. To facilitate medical terminology recognition, the present invention constructs the BIOS as a dictionary tree. Medical terminology is then identified for each restored English medical record based on the dictionary tree, obtaining the medical terminology for each restored English medical record. Using a dictionary tree matching approach can improve the efficiency of medical terminology recognition.

[0046] Step 202: Perform narrative state analysis and relevant attribute extraction based on the medical terms in each restored English medical record and its corresponding English medical record text to obtain narrative state attribute information and other attribute information corresponding to the medical terms in each restored English medical record.

[0047] It's important to note that because medical terms in electronic medical records can have different contexts, they need to be categorized based on dimensions such as presence (existing, possible), time (current, past medical history), and subject (patient, family history) to better capture narrative state. This is collectively referred to as assertion analysis. Furthermore, the body location of disease states and treatment processes, modifiers of clinical information, clinical attributes, laboratory results, and the values ​​and units of drug dosages are all crucial information for medical research and form an important component of structuring. Therefore, narrative state analysis and related attribute extraction are performed based on the medical terms and corresponding English medical record text in each restored English medical record.

[0048] Among them, the method of performing narrative state analysis and relevant attribute extraction based on the medical terms of each restored English medical record and its corresponding English medical record text can be through model extraction or program extraction, etc. The present invention does not limit the method of performing narrative state analysis and relevant attribute extraction based on the medical terms of each restored English medical record and its corresponding English medical record text.

[0049] Step 203: Taking the medical terms of each restored English medical record as core information, the narrative state attribute information and the other attribute information corresponding to the medical terms are integrated to obtain English structured data of each restored English medical record.

[0050] It should be noted that after obtaining the medical terms, narrative state attribute information, and other attribute information, information integration is required. The integration method can be based on a preset data structure or a preset data model. The present invention does not limit the method for integrating the narrative state attribute information and other attribute information corresponding to each medical term, with the medical term in each restored English medical record as the core information. For example, all term attributes can be integrated to obtain an information list, etc.

[0051] It can be understood that the English structured data obtained by the above method is obtained by directly processing the original English medical records in the restored English medical record set. The extraction results are all substrings of the original text, which avoids the error information caused by the post-translation structuring of the medical records and improves the accuracy of obtaining English structured data.

[0052] Furthermore, the biomedical dictionary tree is Trie dictionary tree data, each edge of the Trie dictionary tree data represents an alphabetic character, and the path from the root node to any node constitutes a medical term string. The medical term identification is performed on each restored English medical record according to the biomedical dictionary tree to obtain the medical term of each restored English medical record. The biomedical dictionary tree is constructed according to the biomedical information ontology system, and may include: using a maximum forward matching algorithm to perform a maximum forward matching scan on each string in each restored English medical record in the Trie dictionary tree data to determine the initial medical term in each restored English medical record; deleting common terms from the initial medical terms in each restored English medical record according to a preset English common word list to obtain the medical terms in each restored English medical record, and linking them to the term codes in the biomedical information ontology system corresponding to the medical terms.

[0053] It should be noted that in the term identification and standardization part of the present invention, the biomedical information ontology system BIOS can be used to use the Trie dictionary tree to perform maximum positive matching on the restored English medical records to identify terms and link them to the corresponding term ID code of the BIOS.

[0054] For example, the biomedical information ontology system BIOS and maximum forward matching can be used to identify terms. For the very large vocabulary of BIOS, the Trie dictionary tree algorithm can be used, with edges representing letters, and the path from the root node to a node on the tree representing a string. After the entire BIOS vocabulary is built into the dictionary tree, a search is performed for each medical record to see if a string appears in the dictionary tree, thereby achieving a match. The matched entities can automatically correspond to the term encoding of the BIOS ontology, and standardization can be implemented when necessary. That is, after medical term recognition, the identified medical terms can be standardized or mapped to standard encoding.

[0055] Some terms in the BIOS ontology are too common, non-medical, or have unclear word boundaries. First, a common English word list can be screened, and non-medical terms among the identified words can be processed using a large language model such as DeepSeek, and then deleted. DeepSeek can also be used to correct the word boundaries of some incomplete terms. Subsequently, the present invention retains the identified terms whose semantic types are the desired information to be extracted for subsequent processing.

[0056] It is understandable that using the above method to extract medical terms from each restored English medical record can improve the efficiency and accuracy of term extraction, laying the foundation for subsequent data processing.

[0057] Furthermore, the narrative state analysis and related attribute extraction based on the medical terms in each restored English medical record and the corresponding English medical record text to obtain narrative state attribute information and other attribute information corresponding to the medical terms in each restored English medical record may include: using a pre-trained narrative state analysis model to perform narrative state analysis on the medical terms in each restored English medical record and the corresponding each restored English medical record, and outputting the narrative state attribute information of each medical term in each restored English medical record, the narrative state attribute information including at least one of existence state, time state and subject state, the narrative state analysis model being a generative large language model trained based on a clinical narrative state annotation dataset; using a general large language model to extract other attribute information from the medical terms in each restored English medical record and the corresponding each restored English medical record to obtain other attribute information corresponding to each medical term in each restored English medical record.

[0058] It's important to note that the analysis of English medical record narrative states can be performed by training a dedicated large language model for narrative states using clinical narrative state annotation datasets, such as the i2b2 manually annotated dataset. This model can then be applied to the English medical record and terminology judgment step. The i2b2 (Informatics for Integrating Biology and the Bedside) annotated dataset is a public dataset released by a medical informatics research project supported by the National Institutes of Health (NIH). It is primarily used for natural language processing (NLP) tasks in electronic health records (EHRs). Its core feature is manually annotated clinical text data, which can cover a variety of medical information extraction tasks.

[0059] For example, the i2b2 dataset can be used. It contains thousands of manually annotated electronic medical records, along with corresponding named entity recognition results and the narrative state of each entity. This dataset will help fine-tune a large model based on Llama-3.1-8B specifically for generative narrative state analysis. This model achieves over 97% accuracy on the test set. This model will be used to infer narrative state analysis on all extracted medical terms and the corresponding restored English medical records. The results will serve as training data for the structured large model.

[0060] The remaining attributes can be extracted using a general-purpose model like GPT. For example, for the remaining attributes: position, modifier, value, and unit, similar prompts can be used to ensure that GPT outputs consistent extraction results. The prompts are as follows (currently, all processing is still performed in English): "This is part of the electronic medical record:\n" + paragraph + "\nFor each entity marked with double curly braces in the note, output its xx attribute (format: entity: attribute) according to the note. The xx attribute should be a substring of the note and should be a noun or noun phrase (the required part of speech for attributes). If no attribute is explicitly specified in the note, just print 'None'.\nThe following are terms:" It is understandable that using the above method to extract narrative state attribute information and other attribute information can improve extraction efficiency and accuracy, and lay the foundation for generating English structured data.

[0061] Furthermore, the medical terms of each restored English medical record are used as core information, and the narrative status attribute information and other attribute information corresponding to the medical terms are integrated to obtain the English structured data of each restored English medical record. This can include: obtaining the data structure corresponding to each medical term in each restored English medical record, the data structure including the term name, semantic type, narrative status, body part, modifier, value and unit field; filling the narrative status attribute information and other attribute information corresponding to each medical term into the corresponding fields of the data structure, and retaining empty values ​​for the fields corresponding to unmentioned attributes; organizing the data structures of all medical terms in each restored English medical record in a list form to form the English structured data.

[0062] It should be noted that the data integration process can be performed on a term-by-term basis, integrating all of the aforementioned term attributes to produce structured English data. Specifically, for all outputs, all attribute information extraction results can be organized around the extracted medical term. The final structured result can be a list, with each element based on the term and including the term's name, semantic type, narrative state, body part, modifier, value, and unit. If the original text does not have a corresponding attribute, it will be left blank. (Other attributes that will be updated and become optional include: date, purpose, etc.)

[0063] Understandably, existing electronic medical record structuring technologies generally perform these tasks independently and integrate the results. However, few methods include all of the aforementioned attribute extraction. The present invention, however, utilizes the aforementioned method to integrate medical terminology and term attribute information to generate English structured data. This allows for the generation of English structured data based on the original English text, unaffected by translation errors, and lays the foundation for subsequent generation of accurate Chinese structured data.

[0064] In some embodiments, the information correspondence between the English structured data of each restored English medical record and the corresponding Chinese translated medical record in the Chinese translated medical record set to obtain the training data may include: establishing a correspondence between the medical terms, the narrative state attribute information, and each attribute value in the other attribute information in the English structured data and the original words or original phrases in the corresponding Chinese translated medical record; if any attribute value does not have a directly corresponding original word or original phrase in the Chinese translated medical record, translating it through a preset Chinese-English comparison word list to obtain a translated word or translated phrase, and establishing a correspondence between it and the translated word or translated phrase; and generating the Chinese structured data based on the correspondence.

[0065] It should be noted that the correspondence between English integrated data and Chinese translated medical records can be achieved through a general large model, thereby obtaining Chinese integrated data.

[0066] Among them, the present invention obtains Chinese medical records from the restored English medical record set, and obtains English structured results from the restored English medical record set. However, as a structured, all extracted results should be substrings of the original text, rather than generated results that may be hallucinatory. Therefore, the present invention does not directly translate the English structured results, but allows the general large model to find each attribute in the English structured results based on the Chinese medical record translation it outputs, and which word / phrase in the translated text corresponds to it. What is obtained in this way is a structured result that completely belongs to the Chinese original text. For a small number of attributes that are not extracted from the original text, the present invention constructs a Chinese-English comparison vocabulary to cover their translation. Taking into account the inconsistency of drug names when actually used in the Chinese environment, some drug names are not translated and English is retained.

[0067] As can be understood, the solution provided by this invention combines the existing biomedical information ontology (BIOS), the narrative state manually annotated data (i2b2), and the support of large-scale model knowledge such as GPT to obtain a Chinese translation of the MIMIC dataset, as well as a structured output of terms and term attributes derived entirely from the original text within the translated medical records. This then constructs a corresponding data extraction process, resulting in Chinese structured data, or training data. This improves the efficiency and accuracy of training data acquisition, thereby enhancing model training effectiveness.

[0068] The present invention also provides a training method for a medical record structured model, which may include: obtaining training data using the above-mentioned training data acquisition method, wherein the training data includes Chinese electronic medical record text and its corresponding structured annotations; using a pre-trained Chinese large language model as a base model, and using the training data to train the base model through supervised fine-tuning; optimizing the model parameters of the base model by minimizing the difference between the structured prediction results output by the model and the structured annotations, thereby obtaining a trained medical record structured model.

[0069] It should be noted that the present invention provides a training method for a large model that can automatically structure the input Chinese electronic medical records. The method uses a pre-trained Chinese large language model as the basic model, and uses the training data to train the basic model through supervised fine-tuning; the model parameters of the basic model are optimized by minimizing the difference between the structured prediction results output by the model and the structured annotations, thereby obtaining a trained medical record structured model. Among them, the pre-trained Chinese large language model refers to a general model pre-trained by massive text, such as the Qwen-2.5-7B model. Using the pre-trained Chinese large language model to train the medical record structured model can improve training efficiency and training effect.

[0070] As can be understood, the training data provided by the above process is used to train and fine-tune the Chinese large language model, enabling it to accept user-provided Chinese raw electronic medical record text and output structured terms and attributes. This achieves the function of automatically processing the structured Chinese electronic medical records in batches while requiring relatively low computing resources. This invention utilizes the supervised fine-tuning and inference technology of the large language model.

[0071] Existing large-scale model technologies generally involve pre-training, supervised fine-tuning, and reinforcement learning. The trained large model is then used for inference. Large language model training learns patterns in the training data and generates output based on the given prompts and input. It also supports long input and output contexts. However, general-purpose large language models such as GPT-4, Wenxin Yiyan, and Claude are slow, expensive, and cannot be deployed offline, posing the risk of patient information leakage.

[0072] The medical record structuring model provided by this invention can integrate the above tasks and output all structured results at once. Specifically, based on the requirements for structuring Chinese electronic medical records, this model uses a smaller model, the pre-trained Qwen-2.5 7B model, which has basic understanding of medical knowledge, to construct a dictionary of medical record-to-medical record structured data. This model then uses multiple cards for parallel training. The resulting performance is comparable to that of a general-purpose large model, enabling local offline deployment and rapid inference execution.

[0073] Specifically, the model training and reasoning part can use the Qwen-2.5-7B model as the base, and train 3 rounds on the Chinese medical records obtained above. The obtained structured model can run reasoning on a single RTX 3090. The training data is the Chinese version of the above-mentioned MIMIC-III data. Taking into account the length limit of the basic model, the present invention only retains training samples with a total length of less than 8192 characters, and there are a total of 190,000 such samples. Starting from the Qwen-2.5-7B large Chinese model, the above training samples were trained on 8 A100 GPUs for 3 rounds, with a total time of 60 hours.

[0074] The trained medical record structured model performed well on artificially generated test data that was not part of the MIMIC model, demonstrating that the proposed data processing process effectively simulates real Chinese medical records. Furthermore, it was able to output at a speed of 2500 characters per second on a single A100 and even run on a single RTX 3090, demonstrating the model's high speed and minimal computational resource requirements.

[0075] The medical record structuring model provided by this invention can output structured results efficiently and accurately in the Chinese language domain in a single pass. This large model can run on an RTX 3090. The model outputs characters extremely quickly, taking only 1-2 seconds per real Chinese medical record on average. The large model's output includes terms and all their attributes, eliminating the need for complex multi-tasking workflows.

[0076] The medical record structured model provided by the present invention can process long text electronic medical records and avoid output duplication. The input and output of the model are limited to 8192 characters, which is longer than the length of most medical records. Compared with previous deep learning methods, the model and method of the present invention can avoid the truncation of medical records to the greatest extent, maintain the complete semantic content, and thus further improve the accuracy. In some tasks that use large language models only to extract terms, the model often outputs the same word repeatedly, especially when processing long texts. By adding all term attributes, the present invention can not only fully realize structured extraction, but also assist in positioning information based on the model to avoid output duplication.

[0077] The data processing flow of the medical record structuring model provided by the present invention has a low demand for manual annotation and can be extended to all low-resource languages. In the process of generating medical record-structured data sample pairs, the present invention mainly uses a general large language model and a biomedical ontology system to assist in annotation, which greatly reduces the requirements for manual annotation in the field of deep learning. By migrating from MIMIC data to Chinese, the present invention simulates the input of real Chinese medical records, and the model still achieves excellent performance. Therefore, this method can be extended to all low-resource languages ​​to help the structuring of electronic medical records in the corresponding languages.

[0078] The following describes an exemplary application of an embodiment of the present invention in a practical application scenario.

[0079] Figure 3 This is a flowchart of obtaining structured data of Chinese electronic medical records provided by the present invention. Figure 3 As shown, the method includes the following steps 301 to 305: Step 301: Use the electronic medical record corpus from MIMIC and process the medical records to obtain processed English medical records. A portion of an original electronic medical record is shown below: “"""Unit No:___ Admission Date:___ Discharge Date:___ Date of Birth:___ Sex: F Service: MEDICINE Allergies: Sulfur / Norvasc Attending:___ Addendum: See below Chief Complaint: abdominal pain Major Surgical or Invasive Procedure: none History of Present Illness: 84 F with PMHx of Renovascular HTN c / b NSTEMI now s / p renal stents, Gout and h / o Crohn's disease who presented to the ED on ___with RLQ pain for approx 2 days. She denies any nausea / vomiting / diarrhea or constipation but has not been taking po well and felt dehydrated.""" This type of data contains excessive jargon and abbreviations, as well as numerous blank lines after line breaks that appear not for semantics but for formatting purposes (e.g., renal; ED on; denies any). Therefore, the present invention first allows GPT to restore the abbreviations in the medical records and remove the formatting line breaks. For example, GPT would restore "84 F with PMHx of Renovascular HTN c / b NSTEMI now s / p renal stents" to "An 84-year-old female with a past medical history of renovascular hypertension complicated by non-ST elevation myocardial infarction, now status post renal stents." The restored medical record is then translated into Chinese and retained for future use.

[0080] Step 302: Using the BIOS ontology and the Trie dictionary tree method, all terms in the processed English medical records are identified. However, some BIOS terms, such as "female" and "days," are not terms. The present invention uses DeepSeek to determine whether they are medical terms and deletes those that are not. It also deletes the 3,000 most frequently occurring words in English. The present invention then retains only terms belonging to the following 16 semantic categories in the BIOS ontology: Anatomical Abnormality, Cell or Molecular Dysfunction, Chemical or Drug, Clinical Attribute, Diagnostic Procedure, Disease, Syndrome or Pathologic Function, Eukaryote, Individual Behavior, Injury or Poisoning, Laboratory Procedure, Mental or Behavioral Dysfunction, Microorganism, Neoplastic Process, Physiology, Sign, Symptom, or Finding, and Therapeutic or Preventive Procedure.

[0081] Step 303: Identify the narrative status of the terms. The present invention has additionally trained a large model to handle this task. Specifically, for a medical record from i2b2, the present invention organizes its manual annotations into the format of "1. Term 1: Status\n2. Term 2: Status\n...", and marks each term with double curly brackets in the original text. In this way, the input of the large model is the medical record text annotated with brackets, and the output is the annotation of the terms and status in the above format. After training, a generative narrative status classification large model can be obtained. For the terms extracted in 302 and the English medical records processed in 301, the present invention organizes the input as described above, and the model will output the status of all terms in the format. By splitting the output of the model using rules, the status can be mapped to each term.

[0082] Step 304: GPT is responsible for extracting all attributes except the descriptive state of the term mentioned in 303. For each attribute, only terms of certain semantic types can have this attribute (for example, medication should not have the attribute of body part, and disease should not have the attribute of purpose). The present invention asks GPT whether the corresponding terms of these semantic types have the corresponding attributes mentioned in the text, and integrates the terms into the final training data. In the medical record provided in 301, the term abdominal pain has the body part abdominal, the term RLQ pain has the value and unit 2days, and so on. An example of the integrated term is as follows: { 'phrase': 'abdominal pain', 'semantic_type': 'Sign, Symptom, or Finding', 'assertion_status': 'present', 'body_location': 'Abdominal', }, Since no other attributes of the term are mentioned in the text, those attributes are left blank and only the above ones are retained for output.

[0083] Step 305: To ensure that the results of the Chinese structured data extraction appear in the original text or can be accurately translated, the present invention has GPT translate each dictionary value in the structured results in 304 into a portion of the original text in the Chinese medical record based on the comparison of the translated Chinese in 301 with the cleaned English original text. For example, "abdominal" can be translated into "abdomen," "abdominal cavity," and so on. Direct translation may not correspond to the Chinese medical record. Therefore, if the Chinese medical record reads "abdominal pain," it is translated into "abdomen." For some words that cannot be found in the original text, such as "present" for describing a state, the present invention manually constructs a Chinese-English vocabulary table so that dictionary values ​​can be translated directly through the vocabulary table lookup (for example, "present" is translated into "existence"). By translating all English into Chinese using one of the two methods described above, the structured results of the Chinese electronic medical record data are obtained.

[0084] Figure 4 This is a flow chart of the medical record structured model training provided by the present invention. Figure 4 As shown in Figure 2, the training method of the medical record structured model includes: Step 401: Take the medical records of several different electronic medical records as input and the corresponding structured results of each record as output, thereby constructing input-output pairs. A certain number of input-output pairs constitute a batch.

[0085] Step 402: For each result that needs to be output, given the actual output up to the current position, the model predicts what the next character will be and provides a probability distribution over the vocabulary.

[0086] Step 403: Update the model parameters through reverse gradient propagation by taking the sum of the cross entropy of the next character at each position in the training data and the model probability distribution as the loss function.

[0087] Step 404: After training the batch data for three rounds, training is stopped to obtain the final model. This model can be deployed for inference on a single GPU.

[0088] Among them, in the method of structuring electronic medical records, multiple information extraction tasks are generally performed independently, and the task results are finally integrated. However, for the extraction of body parts, values ​​and units, for example, traditional rule-based methods cannot achieve a high accuracy rate, and some machine learning-based algorithms require users to specify in advance which word attributes need to be extracted, and cannot directly output all results in full. For the task of term recognition, traditional machine learning methods cannot process long medical record texts. For example, BERT-type models can generally only accept 512 characters as input. For the task of narrative state analysis, the accuracy of the model of the present invention also exceeds all current machine learning and deep learning results on the I2B2 test set. That is, so far there is no other method to complete this task in the field of Chinese electronic medical records where data is sparse.

[0089] The model training method of this invention achieves the following: 1. It transforms the problem of structuring electronic medical records into a problem of generating a single output from a large language model, thereby simplifying multiple tasks into a single one, requiring only the construction of structured electronic medical record output data for training. 2. It innovatively splits the structured Chinese electronic medical record data into structured English electronic medical record data, along with the Chinese translation of the data and the corresponding translation, thereby utilizing public datasets to fill the data gap in Chinese electronic medical records. 3. It uses a universal large model to annotate the data and employs the large model algorithm to complete the task, improving overall accuracy.

[0090] It is understood that the training data acquisition method and medical record structuring model training method provided by this invention can be used to structure Chinese medical electronic medical records in various formats. Furthermore, the data generation method is suitable for Chinese electronic medical records and can be extended to other low-resource languages. Training samples of various styles can also be generated based on standardized and semi-standardized electronic medical records to accommodate different electronic medical record formats. By adding auxiliary positioning information such as term attributes, the model can avoid the dilemma of repeated output for simple tasks.

[0091] Based on the foregoing embodiments, embodiments of the present invention provide a device for acquiring training data and a device for training a medical record structured model. The modules included in the device and the units included in each module can be implemented by a processor; of course, they can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0092] The following describes the training data acquisition device and the training device for the medical record structured model provided by the present invention. The training data acquisition device described below and the training data acquisition method described above can be referenced to each other, and the training device for the medical record structured model described below and the training method for the medical record structured model described above can be referenced to each other.

[0093] Figure 5 Schematic diagram of the structure of the device for acquiring training data provided by the present invention. Figure 5 As shown, the apparatus 500 includes a medical record restoration module 501, a medical record translation module 502, a structured data module 503, and a data acquisition module 504, wherein: The medical record restoration module 501 is used to restore the medical jargon and medical term abbreviations in the to-be-processed English medical record set to their corresponding full names, thereby obtaining a restored English medical record set; A medical record translation module 502 is configured to translate the restored English medical record set into Chinese to obtain a Chinese translated medical record set; The structured data module 503 is configured to perform medical term recognition and attribute information extraction on each restored English medical record in the restored English medical record set to obtain English structured data for each restored English medical record, wherein the English structured data includes the medical term and attribute information related to the medical term, wherein the attribute information is used to represent the condition information and / or treatment information related to the medical term; The data acquisition module 504 is used to match the English structured data of each restored English medical record with the corresponding Chinese translated medical record in the Chinese translated medical record set to obtain Chinese structured data, which is the training data.

[0094] In some embodiments, the structure data module 503 includes a term extraction unit, an attribute extraction unit and an information integration unit, wherein: The term extraction unit is used to identify medical terms for each restored English medical record based on a biomedical dictionary tree to obtain the medical terms for each restored English medical record, wherein the biomedical dictionary tree is constructed based on a biomedical information ontology system; The attribute extraction unit is configured to perform narrative state analysis and relevant attribute extraction based on the medical terms in each restored English medical record and the corresponding English medical record text, to obtain at least one of narrative state attribute information and other attribute information corresponding to the medical terms in each restored English medical record; The information integration unit is used to integrate the narrative state attribute information and the other attribute information corresponding to the medical terms with the medical terms of each restored English medical record as the core information to obtain the English structured data of each restored English medical record.

[0095] In some embodiments, the biomedical dictionary tree is Trie dictionary tree data, each edge of the Trie dictionary tree data represents an alphabetic character, and the path from the root node to any node constitutes a medical term string, and the term extraction unit is specifically used to: Performing a maximum forward matching scan on each character string in each restored English medical record in the Trie dictionary tree data using a maximum forward matching algorithm to determine the initial medical term in each restored English medical record; Common terms are deleted from the initial medical terms in each restored English medical record according to a preset English common word list to obtain the medical terms in each restored English medical record, and the medical terms are linked to the term codes in the biomedical information ontology system corresponding to the medical terms.

[0096] In some embodiments, the attribute extraction unit is specifically configured to: Using a pre-trained narrative state analysis model to perform narrative state analysis on the medical terms in each restored English medical record and the corresponding medical terms in each restored English medical record, outputting narrative state attribute information of each medical term in each restored English medical record, wherein the narrative state attribute information includes at least one of an existential state, a temporal state, and a subject state, the narrative state analysis model being a generative large language model trained based on a clinical narrative state annotation dataset; A general large language model is used to extract other attribute information from the medical terms in each restored English medical record and the corresponding medical terms in each restored English medical record to obtain other attribute information corresponding to each medical term in each restored English medical record.

[0097] In some embodiments, the information integration unit is specifically configured to: Obtaining a data structure corresponding to each medical term in each restored English medical record, wherein the data structure includes term name, semantic type, narrative state, body part, modifier, value, and unit fields; Filling the description state attribute information and other attribute information corresponding to each medical term into the corresponding fields of the data structure, and leaving empty values ​​for the fields corresponding to the attributes not mentioned; The data structure of all medical terms in each restored English medical record is organized in a list form to form the English structured data.

[0098] In some embodiments, the information integration unit is specifically configured to: Establishing a correspondence between the medical term, the narrative state attribute information, and each attribute value in the other attribute information in the English structured data and the original word or original phrase in the corresponding Chinese translated medical record; If any attribute value does not have a directly corresponding original word or original phrase in the Chinese translated medical record, then translate it using a preset Chinese-English bilingual word list to obtain a translated word or translated phrase, and establish a corresponding relationship between it and the translated word or translated phrase; The Chinese structured data is generated according to the corresponding relationship.

[0099] In some embodiments, the device also includes a preprocessing module, which is used to: obtain a public English electronic medical record data set; use a natural language processing model to de-privacy and format normalize the English electronic medical records in the English electronic medical record data set to obtain the English medical record set to be processed.

[0100] In the embodiment of the present invention, Chinese training data of a Chinese electronic structured medical record dataset can be constructed based on an English electronic medical record dataset, thereby improving the efficiency and accuracy of obtaining the Chinese training data.

[0101] Figure 6 Schematic diagram of the structure of the device for acquiring training data provided by the present invention. Figure 6 As shown, the apparatus 600 includes a training data module 601, a model training module 602, and a model optimization module 603, wherein: The training data module 601 is used to obtain training data using the above-mentioned training data acquisition method, wherein the training data includes Chinese electronic medical record text and its corresponding structured annotations; A module training module 602 is configured to use a pre-trained Chinese language model as a base model and train the base model through supervised fine-tuning using the training data; The model optimization module 603 is used to optimize the model parameters of the basic model by minimizing the difference between the structured prediction results output by the model and the structured annotations, so as to obtain a trained medical record structured model.

[0102] In the embodiment of the present invention, accurate Chinese structured data can be constructed based on the English electronic medical record data set for model training, which improves the efficiency and accuracy of obtaining training data and the effect of model training.

[0103] Figure 7 This is a schematic diagram of the physical structure of the medical record structuring device provided by the present invention. Figure 7As shown, the medical record structuring device may include: a processor 710, a communications interface 720, a memory 730 and a communication bus 740, wherein the memory stores a medical record structuring model, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logic instructions in the memory 730 to execute the medical record structured model, which is obtained by obtaining training data using the above-mentioned training data acquisition method and performing model training using the training data. The training data acquisition method includes: restoring the medical jargon and medical term abbreviations in the English medical record set to be processed into corresponding full names to obtain a restored English medical record set; translating the restored English medical record set into Chinese to obtain a Chinese translated medical record set; performing medical term identification and attribute information extraction on each restored English medical record in the restored English medical record set to obtain English structured data of each restored English medical record, the English structured data including the medical terms and attribute information related to the medical terms, the attribute information being used to represent the condition information and / or treatment information related to the medical terms; and matching the English structured data of each restored English medical record with the corresponding Chinese translated medical record in the Chinese translated medical record set to obtain Chinese structured data, which is the training data.

[0104] Furthermore, the logic instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0105] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a medical record structured model stored thereon, wherein the medical record structured model is obtained by obtaining training data using the above-mentioned training data acquisition method and performing model training using the training data. The training data acquisition method includes: restoring the medical jargon and medical term abbreviations in the English medical record set to be processed into corresponding full names to obtain a restored English medical record set; translating the restored English medical record set into Chinese to obtain a Chinese translated medical record set; performing medical term identification and attribute information extraction on each restored English medical record in the restored English medical record set to obtain English structured data of each restored English medical record, wherein the English structured data includes the medical terms and attribute information related to the medical terms, wherein the attribute information is used to represent the condition information and / or treatment information related to the medical terms; matching the English structured data of each restored English medical record with the corresponding Chinese translated medical record in the Chinese translated medical record set to obtain Chinese structured data, wherein the Chinese structured data is the training data.

[0106] The computer-readable storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0107] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0108] Program code embodied on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, radio frequency (RF), etc., or any suitable combination of the foregoing.

[0109] Computer program code for performing the operations of this specification may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0111] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for obtaining training data, characterized in that: include: Restoring medical jargons and abbreviations of medical terms in the to-be-processed English medical record set to their corresponding full names, thereby obtaining a restored English medical record set; translating the restored English medical record set into Chinese to obtain a Chinese translated medical record set; performing medical term recognition and attribute information extraction on each restored English medical record in the restored English medical record set to obtain English structured data for each restored English medical record, wherein the English structured data includes the medical term and attribute information related to the medical term, wherein the attribute information is used to represent condition information and / or treatment information related to the medical term; The English structured data of each restored English medical record is matched with the Chinese translated medical record in the corresponding Chinese translated medical record set to obtain Chinese structured data, which is the training data.

2. The method for obtaining training data according to claim 1, wherein: The medical term recognition and attribute information extraction are performed on each restored English medical record in the restored English medical record set to obtain English structured data of each restored English medical record, including: Performing medical term recognition on each restored English medical record according to a biomedical dictionary tree to obtain medical terms for each restored English medical record, wherein the biomedical dictionary tree is constructed according to a biomedical information ontology system; Performing narrative state analysis and related attribute extraction based on the medical terms in each restored English medical record and the corresponding English medical record text, to obtain narrative state attribute information and other attribute information corresponding to the medical terms in each restored English medical record; Taking the medical terms of each restored English medical record as core information, the narrative state attribute information and the other attribute information corresponding to the medical terms are integrated to obtain English structured data of each restored English medical record.

3. The method for obtaining training data according to claim 2, wherein: The biomedical dictionary tree is a Trie dictionary tree data, each edge of the Trie dictionary tree data represents an alphabetic character, and the path from the root node to any node constitutes a medical term string. Medical term identification is performed on each restored English medical record based on the biomedical dictionary tree to obtain the medical term of each restored English medical record. The biomedical dictionary tree is constructed based on a biomedical information ontology system and includes: Performing a maximum forward matching scan on each character string in each restored English medical record in the Trie dictionary tree data using a maximum forward matching algorithm to determine the initial medical term in each restored English medical record; Common terms are deleted from the initial medical terms in each restored English medical record according to a preset English common word list to obtain the medical terms in each restored English medical record, and the medical terms are linked to the term codes in the biomedical information ontology system corresponding to the medical terms.

4. The method for obtaining training data according to claim 2, wherein: The narrative state analysis and related attribute extraction are performed based on the medical terms in each restored English medical record and the corresponding English medical record text, to obtain narrative state attribute information and other attribute information corresponding to the medical terms in each restored English medical record, including: Using a pre-trained narrative state analysis model to perform narrative state analysis on the medical terms in each restored English medical record and the corresponding medical terms in each restored English medical record, outputting narrative state attribute information of each medical term in each restored English medical record, wherein the narrative state attribute information includes at least one of an existential state, a temporal state, and a subject state, the narrative state analysis model being a generative large language model trained based on a clinical narrative state annotation dataset; A general large language model is used to extract other attribute information from the medical terms in each restored English medical record and the corresponding medical terms in each restored English medical record to obtain other attribute information corresponding to each medical term in each restored English medical record.

5. The method for obtaining training data according to claim 2, wherein: The medical terminology of each restored English medical record is used as core information, and the narrative state attribute information and the other attribute information corresponding to the medical terminology are integrated to obtain English structured data of each restored English medical record, including: Obtaining a data structure corresponding to each medical term in each restored English medical record, wherein the data structure includes term name, semantic type, narrative state, body part, modifier, value, and unit fields; Filling the description state attribute information and other attribute information corresponding to each medical term into the corresponding fields of the data structure, and leaving empty values ​​for the fields corresponding to the attributes not mentioned; The data structure of all medical terms in each restored English medical record is organized in a list form to form the English structured data.

6. The method for obtaining training data according to claim 2, wherein: The step of matching the English structured data of each restored English medical record with the corresponding Chinese translated medical record in the Chinese translated medical record set to obtain the training data includes: Establishing a correspondence between the medical term, the narrative state attribute information, and each attribute value in the other attribute information in the English structured data and the original word or original phrase in the corresponding Chinese translated medical record; If any attribute value does not have a directly corresponding original word or original phrase in the Chinese translated medical record, then translate it using a preset Chinese-English bilingual word list to obtain a translated word or translated phrase, and establish a corresponding relationship between it and the translated word or translated phrase; The Chinese structured data is generated according to the corresponding relationship.

7. A method for training a medical record structured model, characterized in that: include: Acquire training data using the method for acquiring training data according to any one of claims 1 to 6, wherein the training data includes Chinese electronic medical record text and its corresponding structured annotations; Taking a pre-trained Chinese language model as a base model, and using the training data to train the base model through supervised fine-tuning; The model parameters of the basic model are optimized by minimizing the difference between the structured prediction results output by the model and the structured annotations, thereby obtaining a trained medical record structured model.

8. A device for acquiring training data, characterized in that: include: A medical record restoration module is used to restore medical jargon and medical term abbreviations in the to-be-processed English medical record set to their corresponding full names, thereby obtaining a restored English medical record set; A medical record translation module, configured to translate the restored English medical record set into Chinese to obtain a Chinese translated medical record set; a structured data module, configured to perform medical term recognition and attribute information extraction on each restored English medical record in the restored English medical record set, thereby obtaining English structured data for each restored English medical record, wherein the English structured data includes the medical term and attribute information related to the medical term, wherein the attribute information is used to represent condition information and / or treatment information related to the medical term; The data acquisition module is used to match the English structured data of each restored English medical record with the corresponding Chinese translated medical record in the Chinese translated medical record set to obtain Chinese structured data, which is the training data.

9. A training device for a medical record structured model, characterized in that: include: A training data module, configured to obtain training data using the training data acquisition method according to any one of claims 1 to 6, wherein the training data includes Chinese electronic medical record text and its corresponding structured annotations; A module training module, configured to use a pre-trained Chinese language model as a base model and train the base model through supervised fine-tuning using the training data; The model optimization module is used to optimize the model parameters of the basic model by minimizing the difference between the structured prediction results output by the model and the structured annotations, so as to obtain a trained medical record structured model.

10. A medical record structuring device, characterized in that: It includes a memory, a processor, and a medical record structured model stored in the memory and running on the processor. The medical record structured model is obtained by obtaining training data using the training data acquisition method described in any one of claims 1 to 6, and performing model training using the training data.