Structured Chinese Electronic Medical Record Generation Method, System, Device and Storage Medium

Through the three-stage training strategy and semantic retrieval model combined with thinking chain technology, the accuracy and efficiency of generating structured Chinese electronic medical records in the existing technology are solved, and high accuracy and low cost automated collation and analysis of medical data are achieved.

CN119694472BActive Publication Date: 2025-07-01PEKING UNION MEDICAL COLLEGE HOSPITAL +1

Patent Information

Application Number
CN202510205334.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-01
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The prior art is difficult to accurately generate structured Chinese electronic medical records, especially when processing complex semantics and long texts. The generalization ability is insufficient and it is difficult to meet the needs of automated collation and analysis of medical data.

Method used

The base model is trained using a three-stage training strategy. The first stage is pre-trained using the Chinese corpus, the second stage is fine-tuned using the Chinese medical data set, and the third stage is fine-tuned using real structured data combined with low-rank weight decomposition technology. Through hierarchical analysis and semantic retrieval models, combined with thinking chain technology, structured Chinese electronic medical records are reconstructed.

Benefits of technology

The accuracy and efficiency of information extraction have been significantly improved, the cost of manual intervention has been reduced, and the efficiency of automated collation of medical data and the value of data analysis have been improved. The average accuracy rate of 78.4% of the baseline model has been increased to 85.3%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119694472B_ABST
    Figure CN119694472B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and storage medium for generating structured Chinese electronic medical records, which are corresponding solutions. In the solutions: Designing a series of professional knowledge enhancement training and reasoning optimization strategies can help the base model accurately achieve automated structured output of electronic medical record information. Through comparative experiments, it is shown that the solution of the present invention improves the average accuracy of the baseline model from 78.4% to 85.3%, significantly exceeding the traditional electronic medical record structuring solution based on the BERT model, and has good zero-shot generalization ability. Compared with commercial application large models, the solution of the present invention has the characteristics of autonomous and controllable iterative development, good privacy and security protection characteristics, and low-cost advantages. The above solutions can be deployed into digital medical intelligent systems to solve the batch structured extraction requirements of Chinese electronic medical records, or can be embedded into intelligent medical record assisted consultation systems to achieve automated extraction of electronic medical record information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of structured electronic medical record information generation, and in particular, to a method, system, device and storage medium for generating structured Chinese electronic medical records. Background Art

[0002] The technology for generating structured Chinese electronic medical records is an important research direction in the cross-field of medical informatization and artificial intelligence. With the development of information technology, electronic medical record systems have gradually replaced traditional paper medical records and become an indispensable part of modern medical services. Structured electronic medical records not only improve the efficiency of medical data storage, retrieval and analysis, but also provide valuable data resources for clinical decision support, medical research and public health management. Since Chinese is a language with rich semantics and flexible grammar, it increases the difficulty of applying natural language processing (NLP) technology to Chinese electronic medical records.

[0003] There are mainly three types of methods for structuring electronic medical records as follows:

[0004] (1) NLP-based statistical methods for structuring electronic medical records: Traditional methods for structuring electronic medical records are mainly based on statistical methods, including word frequency analysis, hidden Markov model (HMM), conditional random field (CRF), etc. These methods construct the syntactic and semantic relationships of the text through statistical analysis of features such as words and phrases. For example, using CRF for named entity recognition (NER) of medical entities is a classic method. However, it is difficult to handle complex semantics, the model's capabilities rely on manual feature engineering, and its generalization ability is insufficient, and it performs poorly on problems such as long texts and weak context relevance. Therefore, it is difficult to accurately generate structured Chinese electronic medical records.

[0005] (2) Methods for structuring electronic medical records based on deep learning models: With the development of deep learning, the structuring of electronic medical records has gradually shifted to neural network models, including convolutional neural network (CNN), long short-term memory network (LSTM), attention mechanism, and models based on Transformer. These methods use a large amount of labeled data in medical texts to achieve medical information extraction through supervised learning. Combining bidirectional LSTM with CRF for medical entity recognition and relationship extraction in electronic medical records can effectively improve the model's extraction ability. The BERT model (Bidirectional Encoder Representations from Transformers) based on Transformer has also been widely applied in the medical field. However, its understanding of domain knowledge is limited. Therefore, it is difficult to accurately generate structured Chinese electronic medical records. An example of the literature of this type of method is: Chinese patent application for invention with publication number CN118095285A, "A Method and System for Named Entity Recognition of Electronic Medical Records Based on Deep Learning".

[0006] (3) Electronic medical record structuring method based on large language models: In recent years, methods based on large language models have gradually been applied to the structuring of electronic medical records. Through large-scale pre-training, large language models have strong language understanding and generation capabilities. However, when applied in the medical field, there are data privacy and security issues when calling API interfaces (application programming interfaces), and due to the uniqueness of Chinese and its complex semantic expression methods, the performance of this method also needs to be improved. A literature example of such a method is: Chinese patent application for invention "An Electronic Medical Record Structuring Method Based on a Large Language Model" with publication number CN119127979A. Summary of the Invention

[0007] The purpose of the present invention is to provide a method, system, device and storage medium for generating structured Chinese electronic medical records, which can significantly improve the accuracy and efficiency of information extraction, accurately generate structured Chinese electronic medical records, reduce the cost of manual intervention, and improve the efficiency of automatic medical data collation and the value of data analysis.

[0008] The purpose of the present invention is achieved through the following technical solutions:

[0009] A method for generating structured Chinese electronic medical records includes:

[0010] Model training: Training the selected base model using a three-stage training strategy. In the first stage, pre-training is performed using a Chinese corpus. In the second stage, full-scale fine-tuning training is performed using a Chinese medical data set. In the third stage, fine-tuning training is performed using the real structured data used in the project combined with low-rank weight decomposition fine-tuning technology;

[0011] Inference output optimization: Using the trained base model to hierarchically parse the input original Chinese medical record text to obtain several structured paragraphs; encoding the original Chinese medical record text and the content in the medical record knowledge document library into a shared semantic space, and querying the most relevant medical background knowledge or medical record structuring examples from the medical record knowledge document library as context cues through a semantic retrieval model; through the chain of thought technology, guiding the trained base model to extract information from each structured paragraph respectively in combination with the context cues to reconstruct a structured Chinese electronic medical record.

[0012] A system for generating structured Chinese electronic medical records includes:

[0013] A model training unit for model training, where the model training includes: Training the selected base model using a three-stage training strategy. In the first stage, pre-training is performed using a Chinese corpus. In the second stage, full-scale fine-tuning training is performed using a Chinese medical data set. In the third stage, fine-tuning training is performed using the real structured data used in the project combined with low-rank weight decomposition fine-tuning technology;

[0014] An inference output optimization unit for inference output optimization, where the inference output optimization includes: hierarchically parsing the input original Chinese medical record text using a trained base model to obtain a number of structured paragraphs; encoding the original Chinese medical record text and the content in the medical record knowledge document library into a shared semantic space, and querying the most relevant medical background knowledge or medical record structured examples from the medical record knowledge document library as context cues through a semantic retrieval model; and through the chain of thought technology, guiding the trained base model to extract information from each structured paragraph respectively in combination with the context cues to reconstruct a structured Chinese electronic medical record.

[0015] A processing device includes: one or more processors; a memory for storing one or more programs;

[0016] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing method.

[0017] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.

[0018] It can be seen from the technical solutions provided by the present invention described above that the present invention uses a base model for structured extraction of Chinese electronic medical records. A series of designed professional knowledge enhancement trainings (i.e., a three-stage training strategy) and inference optimization strategies can help the base model accurately achieve automated structured output of electronic medical record information. Through comparative experiments, it is shown that the solution of the present invention improves the average accuracy of the baseline model from 78.4% to 85.3%, significantly exceeding the traditional BERT model-based electronic medical record structuring solution, and has good zero-shot generalization ability. Compared with commercial application large models, the solution of the present invention has the characteristics of autonomous and controllable iterative development, good privacy and security protection characteristics, and low-cost advantages. The solution of the present invention can be deployed into a digital medical intelligent system to meet the large-scale batch structured extraction requirements of Chinese electronic medical records, or can be embedded into an intelligent medical record assisted consultation system to achieve automated extraction of electronic medical record information, promote the robotic automated process of medical electronic records, and improve the work efficiency of medical staff in processing electronic medical records. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1Flowchart of a method for generating a structured Chinese electronic medical record provided by an embodiment of the present invention;

[0021] Figure 2 Schematic diagram of a method for generating a structured Chinese electronic medical record provided by an embodiment of the present invention;

[0022] Figure 3 Schematic diagram of the visualization result of the accuracy of the model provided by an embodiment of the present invention on the self-built validation set with the number of training rounds;

[0023] Figure 4 Schematic diagram of the visualization result of the precision of the model provided by an embodiment of the present invention on the self-built validation set with the number of training rounds;

[0024] Figure 5 Schematic diagram of the visualization result of the recall rate of the model provided by an embodiment of the present invention on the self-built validation set with the number of training rounds;

[0025] Figure 6 Schematic diagram of the visualization result of the F1 score of the model provided by an embodiment of the present invention on the self-built validation set with the number of training rounds;

[0026] Figure 7 Schematic diagram of the visualization result of the F1 score on the self-built validation set provided by an embodiment of the present invention with the proportion of training data;

[0027] Figure 8 Schematic diagram of a structured Chinese electronic medical record generation system provided by an embodiment of the present invention;

[0028] Figure 9 Schematic diagram of a processing device provided by an embodiment of the present invention. Detailed implementation manners

[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] First, the following explanations are given for the terms that may be used in this article:

[0031] Descriptions using terms such as "including", "comprising", "containing", "having", or other similar semantics should be construed as non-exclusive inclusion. For example, including a technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) should be construed as not only including the explicitly listed technical feature element, but also including other technical feature elements well-known in the art that are not explicitly listed.

[0032] The term "consisting of" means excluding any technical feature element not explicitly listed. If this term is used in a claim, it will make the claim a closed type, so that it does not include technical feature elements other than the explicitly listed ones, except for conventional impurities related thereto. If this term only appears in a sub-clause of a claim, then it only limits the elements explicitly listed in that sub-clause, and the elements recorded in other sub-clauses are not excluded from the overall claim.

[0033] The following provides a detailed description of a method, system, device, and storage medium for generating a structured Chinese electronic medical record provided by the present invention. Contents not described in detail in the embodiments of the present invention belong to the prior art well-known to those skilled in the art. Conditions not specified in the embodiments of the present invention are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. Instruments not specified by the manufacturer in the embodiments of the present invention are all conventional products that can be obtained through commercial purchase.

[0034] Embodiment 1

[0035] The embodiment of the present invention provides a method for generating a structured Chinese electronic medical record, as Figure 1 shown, which mainly includes the following steps:

[0036] Step 1, model training based on a three-stage training strategy.

[0037] In the embodiment of the present invention, a selected base model is trained using a three-stage training strategy. In the first stage, pre-training is performed using a Chinese corpus. In the second stage, full-scale fine-tuning training is performed using a Chinese medical data set. In the third stage, fine-tuning training is performed using the real structured data used in the project combined with low-rank weight decomposition fine-tuning technology.

[0038] In the embodiment of the present invention, in the second stage, the pre-trained model obtained in the first stage is fine-tuned, which can enhance the model's understanding of the language in the medical field. Then, in the third stage, by using high-quality real structured data, the trained base model can better complete the task of generating a structured Chinese electronic medical record.

[0039] Those skilled in the art can understand that the foundation model has become an industry term in the fields of artificial intelligence and machine learning, and generally refers to a large language model (LLM) pre-trained on a vast amount of general data. For example, the foundation model of the present invention can be selected as Qwen2-7B, where Qwen is the model name, 2 represents the second-generation model, and 7B indicates that the model has 7 billion parameters. Of course, other large language models can also be selected as the foundation model according to needs, and the specific type of the foundation model is not limited in the present invention.

[0040] Step 2: The inference output optimization stage based on retrieval-augmented generation (RAG) and chain of thought (CoT).

[0041] In the embodiment of the present invention, the trained foundation model performs hierarchical parsing on the input original Chinese medical record text to obtain several structured paragraphs, and combines the retrieval augmentation technology to achieve high-accuracy and more fine-grained electronic structured case generation. The original Chinese medical record text and the content in the medical record knowledge document library are encoded into a shared semantic space, and the most relevant medical background knowledge or medical record structured examples are queried from the medical record knowledge document library through a semantic retrieval model as context prompts. The retrieved context prompts will guide the trained foundation model to extract information from each structured paragraph through the chain of thought technology, and finally reconstruct a structured Chinese electronic medical record by combining the extracted information.

[0042] Preferably, in the inference output optimization, the original Chinese medical record text and the medical record knowledge document library are embedded into the shared space through the semantic vector model BGE, and the Retriever-Reader architecture is introduced. Retriever represents the retriever, which is used to retrieve relevant documents from the medical record knowledge document library through the displayed query matching degree, and Reader represents the reader, which is used to extract medical background knowledge from the retrieved relevant documents.

[0043] In order to more clearly demonstrate the technical solutions provided by the present invention and the resulting technical effects, the method provided by the embodiment of the present invention will be described in detail below with specific embodiments.

[0044] In this embodiment, Qwen2-7B is used as the base model. Taking the original medical record text of ectopic pregnancy-related diseases in the Chinese obstetrics and gynecology field as an example, an automated structured output solution for electronic medical record information is introduced. The goal of this solution is to automatically extract key information from the original medical record text, such as, including but not limited to: medical information such as the patient's personal basic information, past medical history, number of cesarean sections, progesterone level, HCG level, etc. Through the structured output ability of the base model, the accuracy and efficiency of information extraction are significantly improved, the cost of manual intervention is reduced, and the efficiency of automated sorting of medical data related to ectopic pregnancy and the value of data analysis are enhanced. This solution mainly includes two parts: a three-stage training strategy and an inference output optimization based on Retrieval-Augmented Generation (RAG) and Chain of Thought (CoT). Figure 2 Shows the overall process.

[0045] 1. Three-stage training strategy.

[0046] (1) The first stage: The pre-training stage.

[0047] The base model is pre-trained based on a large-scale Chinese corpus to enable it to master a wide range of Chinese language knowledge and basic language understanding capabilities. In the pre-training stage, the Chinese medical corpora ChiMed and ChiMST, and the Chinese electronic medical record entity recognition datasets (CCKS17 - CCKS20) are mainly used for warm-up learning of knowledge related to the medical record field.

[0048] As Figure 2 shown, in this stage, according to the currently input text (Text), the next language unit (NextToken) is predicted, and a loss function is constructed based on the difference between the prediction result and the actual result included in the training data (i.e., using the true next language unit as supervision information) for model training.

[0049] (2) The second stage: The external data fine-tuning stage.

[0050] Based on the pre-trained model, the model is further fine-tuned using publicly available Chinese medical datasets to enhance the model's understanding of the language in the medical field, especially common terms and diagnostic processes in obstetrics and gynecology. In this stage, the present invention screens out data related to gynecological medical records from the Chinese medical Q&A dataset CMedQA, the named entity recognition dataset Yidu-S4k of Chinese electronic medical records, and the Chinese medical Q&A dataset, and performs full-scale fine-tuning training on the pre-trained base model.

[0051] As Figure 2As shown, in this stage, the questions in the training data are input into the pre-trained model, and the prompt for answering the questions is used as an instruction to guide the pre-trained model to predict the answers to the questions. A loss function is constructed based on the difference between the prediction result and the actual result included in the training data, and full-scale fine-tuning of the model parameters is performed.

[0052] (3) The third stage: the project data fine-tuning stage.

[0053] Combined with the high-quality real obstetrics and gynecology ectopic pregnancy medical record data collected in the project, targeted and efficient parameter fine-tuning is carried out. For example, 1.2k high-quality data were collected in the experiment. In the case of small-scale training data, the present invention can adopt the low-rank weight decomposition (LoRA) fine-tuning technique to enable the model to adapt to specific data distributions and task requirements, especially for the specific disease descriptions and feature extraction tasks related to ectopic pregnancy.

[0054] As Figure 2 shown, similarly, this stage also belongs to similar questions and instructions, guiding the model after the second-stage training to predict the structured Chinese electronic medical record (structured output). A loss function is constructed based on the difference between the prediction result and the actual result included in the training data, and fine-tuning of the model parameters is performed.

[0055] Considering that the specific training (fine-tuning) processes of each stage can be implemented with reference to conventional techniques, only some exemplary descriptions are provided below. For those not described in detail, reference can be made to conventional techniques, and the present invention will not be elaborated.

[0056] In this embodiment, in the pre-training stage (i.e., the first stage), the cross-entropy loss is used as the loss function to guide the entire training process. The AdamW optimizer (Adam optimizer with weight decay) is used, with hyperparameters β1 = 0.9, hyperparameter β2 = 0.999, and weight_decay = 0.01. The linear warm-up learning rate strategy is used, where the learning rate linearly increases from 1e-5 to 5e-4 and gradually decays after 10% of the training progress. The single-card batch size is set to 64, and gradient accumulation is set to 2 to accelerate the training speed while maintaining good convergence. To reduce the video memory consumption and accelerate the training speed, FP16 (half-precision floating-point number) mixed-precision training is used. In addition, the gradient clipping value is set to 1.0 to avoid gradient explosion. Hardware configuration: Multi-GPU (graphics processing unit) distributed training (8 48GB GPUs) is used, and the distributed data parallel strategy is utilized to accelerate the training. In the second and third stages based on full-scale fine-tuning and LoRA fine-tuning, the initial learning rate is set to 5e-5, the AdamW optimizer is adopted, with hyperparameters β1 = 0.9, hyperparameter β2 = 0.999, and weight_decay = 0.01. The learning rate scheduler uses linear warm-up (5% of the total number of steps) and then gradually decays. The single-card batch size is set to 64, and gradient accumulation is set to 2. In the LoRA parameters, the rank is set to 16, and the Alpha parameter is set to 32. Further, in the internal project data fine-tuning in the third stage, data enhancement strategies such as data rewriting and data augmentation are designed to improve the training effect of this stage.

[0057] 2. Inference output optimization part.

[0058] In the inference stage, in order to further enhance the model's knowledge of the medical record professional field, the present invention constructs a medical record knowledge document library (external knowledge base), and uses retrieval-enhanced generation and chain-of-thought prompting techniques to assist in optimizing the structured output of medical records in the inference stage.

[0059] Taking the original medical record text of ectopic pregnancy-related diseases in the above-mentioned Chinese obstetrics and gynecology field as an example, to achieve the above goals, first, an external knowledge base for obstetrics and gynecology medical records (i.e., the medical record knowledge document library) is constructed. Professional knowledge documents related to obstetrics and gynecology are extracted from the Chinese Clinical Guidelines Database (CCGD), and a part of the representative labeled data is selected from the project training data to provide additional medical background knowledge for the model. By pre-collecting the external knowledge base, additional medical background knowledge and similar medical record samples to be structurally extracted can be provided for the model to assist the model in context understanding and reasoning, thereby improving the accuracy of structured information extraction.

[0060] The main technical details involved in this part include:

[0061] (1) Reconstruction of medical record descriptions based on large models: For the input original medical record text, a method of hierarchical parsing based on the Qwen large model is used to divide the medical record text into different structured paragraphs (such as personal information, chief complaint, current medical history, etc.) according to semantic granularity, and information extraction is performed separately.

[0062] (2) Vector embedding module: The FlagEmbedding model is used to embed the electronic medical record queries to be structured and external professional knowledge documents into a shared space, and the retrieval effect is improved through explicit query matching degrees, which is suitable for processing complex semantics in medical record texts.

[0063] (3) Retriever-Reader architecture: An architecture combining Retriever (retriever) and Reader (reader) is adopted. The Retriever is used to find relevant documents from the medical record knowledge document library, and the Reader extracts the required information from these documents. For example, the BERT model or the RoBERTa model (an improved BERT model) can be used as the basic Retriever-Reader architecture.

[0064] (4) RAG strategy: Retrieval-Augmented Generation (RAG) locates the top-K most relevant Chunk blocks in the medical record knowledge document library through vector retrieval and introduces the relevant Chunk blocks into the instruction prompt of the large model input; among them, the text content in the medical record knowledge document library will be pre-divided into multiple semantic unit blocks for encoding and storage in advance. Here, the Chunk block is the semantic unit block, and the top-K Chunk blocks refer to the top K semantic unit blocks with the highest query matching degree. The specific value of K here can be set according to the actual situation or experience. They can help in the structured extraction stage to improve the generation quality, thereby obtaining better results in the extraction of medical record information. In addition to retrieving and locating external professional term knowledge, the most relevant medical record structuring examples will also be retrieved from the medical record knowledge document library as thought chain prompt examples. Through the strategy of few-shot thought chain injection, the structured extraction ability for unseen electronic medical records can be further improved.

[0065] To illustrate the performance of the above scheme of the present invention, verification and evaluation were carried out on the self-built evaluation benchmark data. The evaluation index adopted the accuracy index. The accuracy of structured information extraction for numerical type indicators reached 83.7%, and the accuracy of structured information for non-numerical indicators reached 91.4%. The performance evaluation of the present invention on the test data is shown in detail in Table 1 and Table 2.

[0066] Table 1 Performance evaluation of electronic medical record information extraction for numerical type indicators

[0067]

[0068] Table 2 Performance Evaluation of Non - numerical Index Electronic Medical Record Information Extraction

[0069]

[0070] In addition, Figures 3 to 7 is the visualization result of the validation set accuracy during the training process of the present invention. Among them, Figures 3 to 6 are, in sequence: the visualization results of the accuracy (Accuracy), precision (Precision), recall (Recall), and F1 - score of the model on the self - built validation set with the number of training rounds; Figure 7 is the visualization result of the F1 - score on the self - built validation set with the proportion of training data.

[0071] In the above - mentioned solution provided by the embodiments of the present invention, when using the base model for structured extraction of Chinese electronic medical records, a series of designed professional knowledge enhancement training and inference optimization strategies can help the base model accurately achieve automated structured output of electronic medical record information. The solution of the present invention improves the average accuracy of the baseline model (i.e., the "base model" without the three - stage training of the present invention) from 78.4% to 85.3%, significantly exceeding the traditional BERT - based model electronic medical record structuring scheme. And it has good zero - shot generalization ability. Compared with large commercial application models, the solution of the present invention has the characteristics of autonomous and controllable iterative development, good privacy and security protection characteristics, and low - cost advantages. The solution of the present invention can be deployed into a digital medical intelligent system to meet the large - scale batch structured extraction requirements of Chinese electronic medical records, or can be embedded into an intelligent medical record assisted consultation system to achieve automated electronic medical record information extraction, promote the robotic automated process of medical electronic records, and improve the work efficiency of medical staff in processing electronic medical records.

[0072] Through the description of the above - mentioned implementation manners, those skilled in the art can clearly understand that the above - mentioned embodiments can be implemented by software, or can be implemented by means of software plus a necessary general - purpose hardware platform. Based on such an understanding, the technical solutions of the above - mentioned embodiments can be embodied in the form of a software product, which can be stored in a non - volatile storage medium (which can be a CD - ROM, USB flash drive, mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0073] Embodiment 2

[0074] The present invention also provides a system for generating structured Chinese electronic medical records, which is mainly used to implement the method provided in the foregoing embodiments, such as Figure 8As shown, the system mainly includes:

[0075] A model training unit for model training. The model training includes: training a selected base model using a three-stage training strategy. In the first stage, pre-training is performed using a Chinese corpus. In the second stage, full-scale fine-tuning training is performed using a Chinese medical dataset. In the third stage, fine-tuning training is performed using the real structured data used in the project combined with the low-rank weight decomposition fine-tuning technique.

[0076] An inference output optimization unit for inference output optimization. The inference output optimization includes: using the trained base model to perform hierarchical parsing on the input original Chinese medical record text to obtain several structured paragraphs; encoding the content of the original Chinese medical record text and the medical record knowledge document library into a shared semantic space, and querying the most relevant medical background knowledge or medical record structured examples from the medical record knowledge document library as context cues through a semantic retrieval model; through the chain of thought technique, guiding the trained base model to extract information from each structured paragraph respectively in combination with the context cues to reconstruct a structured Chinese electronic medical record.

[0077] Preferably, the pre-training in the first stage using a Chinese corpus includes: the Chinese corpus used includes: the Chinese medical corpus ChiMed and ChiMST, and the Chinese electronic medical record entity recognition dataset (CCKS17 - CCKS20); during the pre-training process, the base model predicts the output of the next language unit based on the currently input text, and uses the real next language unit as supervision information to construct a loss function to guide the entire pre-training process.

[0078] Preferably, the full-scale fine-tuning training in the second stage using a Chinese medical dataset includes: the Chinese medical dataset used includes: the Chinese medical Q&A dataset CMedQA, the named entity recognition dataset Yidu-S4k of Chinese electronic medical records, and the Chinese medical Q&A dataset; screening out the content data related to the case from the relevant datasets, and performing full-scale fine-tuning training on the pre-trained base model.

[0079] Preferably, in the inference output optimization, the original Chinese medical record text and the medical record knowledge document library are embedded into the shared space through a semantic vector model, and a Retriever-Reader architecture is introduced. Retriever represents the retriever, which is used to retrieve relevant documents from the medical record knowledge document library through the displayed query matching degree, and Reader represents the reader, which is used to extract medical background knowledge from the retrieved relevant documents.

[0080] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.

[0081] Embodiment III

[0082] The present invention also provides a processing device, such as Figure 9 shown, which mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.

[0083] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.

[0084] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:

[0085] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.;

[0086] The output device can be a display terminal;

[0087] The memory can be a Random Access Memory (RAM), or a non-volatile memory, such as a disk memory.

[0088] Embodiment IV

[0089] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiments when the computer program is executed by a processor.

[0090] In the embodiments of the present invention, the readable storage medium as a computer-readable storage medium can be disposed in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a Read-Only Memory (ROM), a magnetic disk, or an optical disc.

[0091] As described above, it is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art.

Claims

1. A structured Chinese electronic medical record generation method, characterized in that: include: Model training: A three-stage training strategy is used to train the selected base model. In the first stage, the Chinese corpus is used for pre-training. In the second stage, the Chinese medical dataset is used for full fine-tuning training. In the third stage, the real structured data used in the project is combined with the low-rank weight decomposition fine-tuning technology for fine-tuning training. Reasoning output optimization: Use the trained base model to perform hierarchical analysis on the input original Chinese medical record text to obtain several structured paragraphs, that is, divide the original Chinese medical record text into different structured paragraphs according to semantic granularity; encode the original Chinese medical record text and the content in the medical medical record knowledge document library into a shared semantic space, and use the semantic retrieval model to query the medical background knowledge or medical record structured samples most relevant to the original Chinese medical record text from the medical medical record knowledge document library as contextual prompts. Among them, the original Chinese medical record text and the medical medical record knowledge document library are embedded into the shared space through the semantic vector model, and the Retriever-Reader architecture is introduced. Retriever represents the retriever, which is used to retrieve relevant documents from the medical medical record knowledge document library through the displayed query matching degree. Reader represents the reader, which is used to extract information from the retrieved relevant documents; through the thinking chain technology, combined with the context prompts, the trained base model is guided to extract information from each structured paragraph separately to reconstruct the structured Chinese electronic medical record.

2. A structured Chinese electronic medical record generation method according to claim 1, characterized in that: The first stage uses the Chinese corpus for pre-training, including: The Chinese corpora used include: Chinese medical corpus and Chinese electronic medical record entity recognition dataset; During the pre-training process, the base model predicts the output of the next language unit based on the current input text, and uses the actual next language unit as supervision information to construct a loss function to guide the entire pre-training process.

3. A structured Chinese electronic medical record generation method according to claim 1, characterized in that: The second phase uses the Chinese medical dataset for full fine-tuning training, including: The Chinese medical datasets used include: Chinese medical question and answer dataset, Chinese electronic medical record named entity recognition dataset, and Chinese medical question and answer dataset; Filter case-related content data from the Chinese medical dataset and perform full-scale fine-tuning training on the pre-trained base model.

4. A structured Chinese electronic medical record generation system, characterized in that: include: A model training unit, used for model training, wherein the model training includes: training the selected base model using a three-stage training strategy, wherein the first stage uses a Chinese corpus for pre-training, the second stage uses a Chinese medical data set for full fine-tuning training, and the third stage uses the real structured data used in the project combined with a low-rank weight decomposition fine-tuning technology for fine-tuning training; The inference output optimization unit is used for inference output optimization, and the inference output optimization includes: using the trained base model to perform hierarchical analysis on the input original Chinese medical record text to obtain a number of structured paragraphs, that is, dividing the original Chinese medical record text into different structured paragraphs according to semantic granularity; encoding the original Chinese medical record text and the content in the medical medical record knowledge document library into a shared semantic space, and querying the medical background knowledge or medical record structured samples most relevant to the original Chinese medical record text from the medical medical record knowledge document library as context prompts through a semantic retrieval model, wherein the original Chinese medical record text and the medical medical record knowledge document library are embedded into the shared space through a semantic vector model, and a Retriever-Reader architecture is introduced, Retriever represents a retriever, which is used to retrieve relevant documents from the medical medical record knowledge document library through the displayed query matching degree, and Reader represents a reader, which is used to extract information from the retrieved relevant documents; through the thinking chain technology, combined with the context prompts, the trained base model is guided to extract information from each structured paragraph separately to reconstruct a structured Chinese electronic medical record.

5. A structured Chinese electronic medical record generation system according to claim 4, characterized in that: The first stage uses the Chinese corpus for pre-training, including: The Chinese corpora used include: Chinese medical paper dataset, Chinese electronic medical record data, and medical entity recognition dataset; During the pre-training process, the base model predicts the output of the next language unit based on the current input text, and uses the actual next language unit as supervision information to construct a loss function to guide the entire pre-training process.

6. A structured Chinese electronic medical record generation system according to claim 4, characterized in that: The second phase uses the Chinese medical dataset for full fine-tuning training, including: The Chinese medical datasets used include: Chinese medical question answering dataset CMedQA, Chinese electronic medical record named entity recognition dataset Yidu-S4k, and Chinese medical question answering dataset; Content data related to cases is selected from the Chinese medical question-answering dataset CMedQA, the Chinese electronic medical record named entity recognition dataset Yidu-S4k, and the Chinese medical question-answering dataset, and the pre-trained base model is fully fine-tuned.

7. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 3.

8. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Electronic medical record named entity identification method and system based on deep learning

    CN118095285A

  • Electronic medical record structuring method based on large language model

    CN119127979A

  • Medical field language model construction and electronic medical record text structuring method and system

    CN117577254A

  • Chinese electronic medical record named entity recognition method based on medical dictionary knowledge enhancement

    CN118643833A

Cited By

  • Real-time structured medical record generation method and system supporting local regulation and control

    CN122177336A

  • A method and system for generating real-time structured medical records that supports local regulation

    CN122177336B