Chinese electronic medical record data processing method and system based on large language model
By using a multi-agent collaborative processing framework based on a large language model and a medical extraction knowledge base, the complex structure and entity extraction problems in Chinese electronic medical record data processing are solved, realizing the automated transformation from unstructured medical records to a standardized information database, thus improving processing efficiency and accuracy.
Patent Information
- Application Number
- CN202511627308.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-01-30
AI Technical Summary
Existing Chinese electronic medical record data processing technologies suffer from high rule maintenance costs, poor adaptability, difficulty in handling complex structures such as tables with multiple headers and spanning multiple pages, inability to achieve unified processing of text and tabular data, insufficient accuracy in extracting complex entity and attribute information, and lack of fuzzy time resolution capabilities.
A Chinese electronic medical record data processing method based on a large language model is adopted. By constructing a multi-agent collaborative processing framework and combining it with a medical extraction knowledge base, the document structure parsing, entity attribute extraction, and standardized processing of medical terminology and time expressions are realized, thus constructing a fully automated medical record information extraction process.
It significantly improves the accuracy of extracting complex entities and attributes from Chinese electronic medical records, realizes the automated and standardized mapping of medical terms and time expressions, breaks through the bottleneck of traditional methods in handling complex structures such as multi-header and cross-page tables, supports unified and efficient processing of text and tabular data, and builds an end-to-end automated process from original medical records to a standardized information database.
Smart Images

Figure CN121439062A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electronic medical record data processing technology, and particularly relates to a method and system for processing Chinese electronic medical record data based on a large language model. Background Technology
[0002] Existing Chinese electronic medical record (EMR) data processing technologies, based on rule-based or traditional machine learning methods, suffer from drawbacks such as high rule maintenance costs and poor adaptability. While Transformer-based pre-trained models (such as BERT and GPT) have been applied to medical text processing, they are primarily focused on general domain named entity recognition, lacking sufficient ability to standardize the processing of unstructured text specific to Chinese EMRs (such as admission records and daily progress notes). Specifically, this manifests as: inefficient processing of tabular data in EMR texts, difficulty in handling complex structures such as multi-header tables and multi-page tables, and an inability to achieve unified processing of text and tabular data; insufficient accuracy in extracting complex entity and attribute information from EMRs; and a lack of standardized parsing capabilities for ambiguous timeframes in medical scenarios (such as "3 days post-surgery" and "nearly 1 week"). Summary of the Invention
[0003] Objective: To address the shortcomings and standardization difficulties in existing technologies for processing unstructured text in electronic medical records, this invention aims to provide a Chinese electronic medical record data processing method and system based on a large language model. This method solves the following technical problems: efficiently processing complex structures such as multi-header tables and multi-page tables, achieving unified input of text and tabular data; improving the accuracy of entity and attribute extraction, enabling joint extraction of complex medical entities and their attributes; and completing the standardized mapping of medical terminology to time expressions, resolving terminological ambiguity and fuzzy time resolution issues. The invention comprehensively constructs a fully automated, end-to-end medical record information extraction and processing workflow, promoting the transformation of electronic medical records from raw unstructured data into a standardized medical information database.
[0004] In a first aspect, this invention proposes a method for processing Chinese electronic medical record data based on a large language model, the method comprising:
[0005] S1: Obtain Chinese electronic medical records and obtain preprocessed Chinese electronic medical records after data preprocessing;
[0006] S2: Input the preprocessed Chinese electronic medical record into a multi-agent based on a fine-tuned large language model. The multi-agent combines the input data with its own different processing tasks to obtain the first prompt word.
[0007] S3: Based on the first prompt word, retrieve the relevant medical knowledge from the medical knowledge base, and combine it with the first prompt word to obtain the second prompt word;
[0008] S4: Each agent calls the fine-tuned large language model according to its corresponding second prompt word, performs information extraction and standardization, and outputs the extracted medical information.
[0009] The medical extraction knowledge base is a medical knowledge graph organized in the form of "(entity)-(relationship, attribute)-(entity, attribute description)" triples. It includes a semantic dataset of standardized medical ontology and terminology, a semantic dataset of document specifications, a semantic dataset of medical clinical guidelines and rules, and a dataset of authoritative medical literature and case evidence.
[0010] The specific process of using a fine-tuned large language model to process preprocessed Chinese electronic medical records includes:
[0011] S401: Retrieve relevant medical document knowledge from the medical extraction knowledge base, use the document structure parsing agent to parse the document structure of the preprocessed Chinese electronic medical records, and output the parsed electronic medical records.
[0012] S402: Retrieve the medical knowledge base to obtain the corresponding medical knowledge, use the extraction agent to extract entity attribute relationships from the parsed electronic medical records, extract feature descriptions with medical clinical significance, and obtain the data stream of "clinical statement-original entity-entity relationship" and "time expression-time associated event object".
[0013] S403: Retrieve relevant medical terminology knowledge from the medical extraction knowledge base, and use the medical terminology standardization intelligent agent to standardize, rewrite and map the data flow of "clinical statement-original entity-entity relationship", and output standardized text expression data.
[0014] S404: Retrieve the medical extraction knowledge base to obtain the corresponding medical time expression knowledge, and use the time expression standardization intelligent agent to process "time expression-time related event object" into "specific time scale or time interval-related event" time expression standard format data.
[0015] S405: Utilize the inspection output agent to perform multi-dimensional verification operations on the standardized text expression data and time expression standard format data, and output the extracted medical information.
[0016] Secondly, this invention proposes a Chinese electronic medical record data processing system based on a large language model, used to implement the Chinese electronic medical record data processing method based on a large language model described in the first aspect of this invention. The system includes:
[0017] The data acquisition module collects Chinese electronic medical records;
[0018] The data preprocessing module preprocesses the input data to obtain preprocessed Chinese electronic medical records.
[0019] The medical knowledge base includes a semantic dataset of standardized medical ontology and terminology, a semantic dataset of document specifications, a semantic dataset of medical clinical guidelines and rules, and a dataset of authoritative medical literature and case evidence. Its content is stored in a vector database after being segmented and vectorized.
[0020] The document structure parsing module retrieves relevant knowledge from the medical knowledge base, performs document structure parsing, and outputs the parsed electronic medical record.
[0021] The information extraction module retrieves relevant knowledge from the medical extraction knowledge base, extracts feature descriptions with medical clinical significance, and obtains a data stream of "clinical statement-original entity-entity relationship" and a data stream of "temporal expression-temporally related event object".
[0022] The medical terminology standardization module retrieves relevant knowledge from the medical extraction knowledge base, standardizes and rewrites the data flow of "clinical statement-original entity-entity relationship", and outputs standardized expression data.
[0023] The Time Expression Standardization Module is used to process the data stream of "Time Expression - Time-related Event Object" into a standard format of time expression data of "Specific Time Scale or Time Interval - Related Event".
[0024] The output module is checked. This module retrieves the medical extraction knowledge base, performs multi-dimensional verification operations, and outputs the extracted medical information.
[0025] The beneficial effects of this invention are:
[0026] This invention, by constructing a collaborative processing framework of "large language model + medical extraction knowledge base + multi-agent" and a multimodal unified input scheme, significantly improves the extraction accuracy of complex entities and attributes in Chinese electronic medical records compared to existing technologies; it achieves automated and standardized mapping of medical terms and time expressions, reducing manual annotation costs and significantly improving standardization efficiency; it breaks through the bottleneck of traditional methods in processing complex structures such as multi-header and multi-page tables, supporting unified and efficient processing of text and tabular data; and it constructs an end-to-end automated process from original medical records to a standardized information database, providing high-quality structured data for downstream applications such as clinical decision support and medical data sharing, effectively promoting the construction of smart healthcare informatization. Attached Figure Description
[0027] Figure 1 This is an overall framework diagram of Embodiment 1 of the present invention;
[0028] Figure 2 This is a schematic diagram of the data processing flow of Embodiment 1 of the present invention;
[0029] Figure 3 This is a flowchart of the data processing steps using a large language model in Embodiment 1 of the present invention;
[0030] Figure 4 This is a schematic diagram of the structure of Embodiment 2 of the present invention. Detailed Implementation
[0031] The terms "first," "second," "third," "fourth," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application.
[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Embodiment 1 of this invention proposes a method for processing Chinese electronic medical record data based on a large language model.
[0034] This invention employs a technical architecture of "multi-agent functional decoupling + large language model semantic understanding + medical knowledge extraction base empowerment." The large language model serves as the foundational model for the entire system. Through collaborative operations among agents and combined with knowledge retrieval from the medical knowledge extraction base, it addresses the knowledge deficiencies encountered in complex and lengthy electronic medical records, such as fragmented context, intertwined entity relationships, and lack of standardized terminology / time. This enables a fully intelligent transformation from unstructured medical records to standardized clinical knowledge, extracting effective medical information. The extracted effective medical information can be widely applied to downstream data processing, such as predicting the probability of certain diseases and generating standardized Chinese electronic medical records.
[0035] Figure 1 This is an overall framework diagram of Embodiment 1 of the present invention. Figure 1 In this invention, the overall framework includes: acquiring original Chinese electronic medical records as input, performing data preprocessing on the input, inputting the preprocessed Chinese electronic medical records into a multi-agent based on a finely tuned large language model (i.e., a finely tuned large model for medical record text extraction), combining a medical extraction knowledge base, performing data processing on the input data, obtaining structured output, and outputting the extracted medical information.
[0036] Figure 2 This is a schematic diagram of the data processing flow of Embodiment 1 of the present invention. Figure 2In this process, multiple intelligent agents perform their respective functions. They also retrieve relevant knowledge from a medical knowledge base based on the data to be processed and the data to be processed, and use a large language model to generate answers based on their prompt words and the data to be processed. The data processing flow of Embodiment 1 of this invention includes: using a document structure parsing agent to parse the preprocessed Chinese electronic medical records and outputting the parsed electronic medical records; using an extraction agent to extract entity attribute relationships from the parsed electronic medical records, extracting clinically significant feature descriptions to obtain a data stream of "clinical statement-original entity-entity relationship" and a data stream of "time expression-time-related event object"; using a medical terminology standardization agent to standardize and rewrite the "clinical statement-original entity-entity relationship" data stream, outputting standardized text expression data; using a time expression standardization agent to process the "time expression-time-related event object" into a standard format of specific time scales or time intervals; and using an output checking agent to verify, correct, and integrate the extraction results, finally outputting the extracted medical information.
[0037] The "large language model" or "large model for fine-tuning medical record text extraction" described in this embodiment of the invention is specifically a large model for fine-tuning Chinese medical record text extraction driven by multi-agent collaboration. Its core is: a basic large language model based on Transformer, which, through a two-stage fine-tuning framework of "multi-agent task-supervised fine-tuning - agent collaborative reinforcement learning", combines core knowledge strongly related to the task in the medical extraction knowledge base, adapts to the specific needs of the document structure parsing agent, extraction agent, medical terminology standardization agent, time expression standardization agent, and examination output agent, and enhances the logical consistency of multi-agent results by relying on "agent task consistency loss" and "cross-agent interaction reward".
[0038] For example, the Qwen3-7B large language model is a Transformer-based large language model that can be trained using both supervised fine-tuning and reinforcement learning.
[0039] Reference Figure 1 , 2 As shown, the Chinese electronic medical record data processing method based on a large language model includes:
[0040] S1: Obtain Chinese electronic medical records and obtain preprocessed Chinese electronic medical records after data preprocessing.
[0041] Chinese electronic medical records are commonly used in hospitals and other similar institutions. These records contain text data, tabular data, and image data. This invention primarily addresses the data processing of text and tabular data within medical records.
[0042] Furthermore, the Chinese electronic medical records undergo data preprocessing to obtain preprocessed Chinese electronic medical records. The processing steps include:
[0043] Convert structured or unstructured text data in Chinese electronic medical records into JSON arrays containing "section" and "content" fields;
[0044] Convert structured or unstructured tabular data in Chinese electronic medical records into JSON format tabular data. JSON format tabular data includes "table name", "table header" and "data".
[0045] This method matches structured or unstructured text data in Chinese electronic medical records with the rules governing Chinese text content usage to initially divide the text paragraphs into chapters. Specifically, first, a set of chapter-division keywords is established. For example, keywords such as chief complaint, present illness, and past medical history are derived from the "WS / T 500-2016 Electronic Medical Record Sharing Document Specification" series of standards. Then, through keyword matching and segmentation, the text content can be initially divided into chapters quickly and cost-effectively. Finally, the text paragraphs are converted into JSON arrays containing "chapter" and "content" fields.
[0046] For example, a JSON array containing the "Chapter" and "Content" fields is shown below: [
[0048] {
[0049] Chapter: Present Illness
[0050] Content: "The patient developed a cough and fever 3 days ago, with a maximum temperature of 38.5°C. The symptoms were relieved after taking amoxicillin on their own."
[0051] }, ...... ]
[0054] The structured or unstructured tabular data in Chinese electronic medical records is processed and converted into JSON format tabular data. The JSON format tabular data includes "table name", "table header" and "data".
[0055] Specifically, firstly, OCR technology (such as Tesseract, CRNN algorithm, or commercial OCR service) is used to identify and extract text content and row and column coordinates of each text block from the tables of Chinese electronic medical records.
[0056] Then, the contents of the header and cells are extracted. The header is usually located in the first row of the table and is mostly standardized medical terms (such as keywords such as examination items, examination results, reference values, etc.). The header row area is quickly located by keyword matching, and the data rows are excluded. For nested headers, the main header is usually composed of multiple columns combined into one column, while the child header occupies only one column. The program can obtain the nested header data by traversing the main header column and then traversing each column from top to bottom and from left to right.
[0057] For cross-page table data, it is determined whether the number of columns and the header content are consistent. If they are consistent, they can be directly merged. If they are inconsistent, the similarity S between the cross-page tables is calculated. If the similarity S reaches a preset threshold, the table fragments from multiple pages are then merged. The formula for calculating the threshold is:
[0058]
[0059] in, Indicates the current page table Table on the next page The semantic similarity is denoted by S, where S takes values in the range [0,1]. This represents the table data on the current page. This indicates the next page of table data. (.) indicates semantic vectorization. This refers to the Euclidean norm, i.e. Norm.
[0060] In this embodiment, when When the value is greater than 0.8, cross-page table data merging is performed, and the merged table data is converted into JSON format table data. The JSON format table data includes "table name", "table header" and "data".
[0061] Example: JSON formatted table data as follows:
[0062] {
[0063] "Form Name": "Laboratory Test Form"
[0064] "Header": ["Date of Examination", ["Complete Blood Test", "White Blood Cell Count (×10)"] 9 ["Routine Blood Test", "Red Blood Cell Count (×10¹² / L)"], ["Biochemistry", "Blood Glucose (mmol / L)"], ......],
[0065] "data": [
[0066] {
[0067] Inspection Date: 2025-01-01
[0068] Complete blood count - White blood cell count (×10) 9 / L): 6.5,
[0069] "Complete Blood Count - Red Blood Cell Count (×10¹² / L)": 4.8,
[0070] "Biochemistry_Blood Glucose (mmol / L)": 5.1, ......
[0072] }, ...... ]
[0075] }
[0076] S2: Input the preprocessed Chinese electronic medical record into a multi-agent system based on a fine-tuned large language model. The multi-agent system combines the input data with its own different processing tasks to obtain the first prompt word.
[0077] It should be noted that each agent pre-sets a prompt word template according to its own functional task requirements. Combined with the input data, it obtains a complete prompt word. Based on different prompt words, the large language model can be invoked to perform different tasks and achieve different functions. For the preprocessed Chinese electronic medical records, the large language model needs to perform several types of tasks, such as document structure parsing, extracting entity attribute relationships, standardizing medical terminology, standardizing time expressions, and multi-dimensional validation.
[0078] For example, the document structure parsing prompts are shown below:
[0079] {
[0080] You are a professional medical document parsing expert. Please strictly follow the following requirements when processing the input Chinese electronic medical records:
[0081] ## Task Description
[0082] **Task:**
[0083] Based on knowledge of medical document structure, medical record text is parsed into structured chapters.
[0084] **Output format requirements:**
[0085] - Output must be in plain JSON array format
[0086] - Each array item is an object containing exactly two fields:
[0087] - "Chapter": A string representing a standardized chapter name.
[0088] - "Content": String type, representing the complete text content of this chapter.
[0089] - If a chapter in the original text does not exist, please do not create that item in the array.
[0090] - Extract strictly according to the original text order, and do not add information that does not exist in the original text.
[0091] **Common Standard Chapter Names:**
[0092] ["Chief Complaint", "Present Illness", "Past Medical History", "Personal History", "Family History", "Physical Examination", "Auxiliary Examinations", "Preliminary Diagnosis", "Diagnosis", "Treatment Plan", "Medical Orders"]
[0093] ## Processing Rules
[0094] 1. **Chapter Recognition:** Accurately identify chapter boundaries using knowledge of medical document structure.
[0095] 2. **Content Merging**: Merge related content into semantically correct chapters.
[0096] 3. **Maintaining Integrity:** Maintain the integrity of the original content; do not modify or summarize it.
[0097] 4. **Logical Division:** For non-standard structures, perform reasonable division based on the logic of medical documents.
[0098] 5. **Ambiguity Handling:** When encountering content with ambiguous boundaries, refer to medical documentation conventions for handling it.
[0099] ## Domain Knowledge Reference
[0100] {knowledge_text: The structure of medical documents retrieved from the knowledge base, including definitions of different chapters, etc.}
[0101] ## Enter text:
[0102] {input_text: The text data to be processed}
[0103] Please parse the document based on knowledge of medical document structure, outputting only a JSON array without any further interpretation.
[0104] }
[0105] S3: Based on the first prompt word, retrieve the relevant medical knowledge from the medical knowledge base, and combine it with the first prompt word to obtain the second prompt word.
[0106] Because the field of electronic medical records (EMR) has its own specialized medical knowledge, norms, and industry standards, using only a general large language model to process EMRs is ineffective. To improve the processing performance of large language models on EMRs, this invention constructs a medical extraction knowledge base to empower the large language model with medical expertise.
[0107] The described medical extraction knowledge base is specifically designed for medical information extraction and standardization. It focuses primarily on medical information extraction and standardization, and is not a general encyclopedia or factual database. Instead, it is a deeply structured, strongly semantic, and tightly integrated knowledge infrastructure that incorporates medical industry standards. Its core purpose is to provide large-scale models with accurate and authoritative medical context and standardized rules, thereby enabling them to perform high-precision information extraction, normalization, and relation construction tasks from unstructured medical record texts.
[0108] The medical extraction knowledge base includes multi-source medical datasets, which include semantic datasets of standardized medical ontologies and terminology, semantic datasets of document specifications, semantic datasets of medical clinical guidelines and rules, and datasets of authoritative medical literature and case evidence. The medical extraction knowledge base includes, but is not limited to, the following:
[0109] 1. Standardized Medical Ontology and Terminology: This is the semantic cornerstone of this medical extraction knowledge base, mainly integrating the semantics of core medical terminology standards from both domestic and international sources. For example, international standards cover the meta-terminal sets of SNOMED CT (Systematic Medical Terminology), ICD-10 / 11 (International Classification of Diseases), LOINC (Laboratory Observation Identifier Nomenclature and Coding System), RxNorm (Standardized Drug Nomenclature System), and UMLS (Unified Medical Language System); domestic medical industry standards include the "WS 364-2011 Health Information Data Meta-Value Domain Code" series of standards.
[0110] 2. Document Specifications / Industry Standards: This primarily focuses on document structure standards as a core component. For example, the "WS / T 500-2016 Electronic Medical Record Sharing Document Specification" series of standards includes XML Schema definitions for various shared documents (such as inpatient medical record cover sheets, admission records, discharge records, etc.). This ensures that the medical extraction knowledge base contains a structured blueprint of clinical documents, allowing the large language model to directly understand the proper location and format of information within standardized electronic medical records during information extraction, thus achieving precise filling from free text into standardized data elements.
[0111] 3. Clinical Guidelines and Rule Base: Stores structured clinical practice guidelines, drug interaction rules, diagnostic logic pathways, dosage calculation formulas, etc.
[0112] 4. Authoritative literature and case evidence database: containing structured chains of evidence extracted from high-quality medical journals, textbooks, and desensitized typical cases.
[0113] The medical knowledge base described is a large-scale structured medical knowledge graph organized in the form of "(entity)-(relation, attribute)-(entity, attribute description)". Entity types include: diseases, symptoms, signs, drugs, surgical procedures, laboratory tests, anatomical sites, microorganisms, and medical equipment. Relation types include: disease-has-symptoms, drug-treats-disease, drug-contraindicated-disease, test-used-for-diagnosis-disease, surgery-acts on-anatomical sites, etc. Attributes refer to the key attributes attached to entities, establishing attribute associations. For example, the entity "aspirin" is associated with its standard RxNorm code, generic name, brand name, specifications, dosage unit, and pharmacological classification. Attribute descriptions are the detailed text describing the attributes.
[0114] The medical knowledge base contains a lot of content and a large amount of data. In order to make it more convenient and efficient for large language models to access or retrieve the medical knowledge base, the medical knowledge base is fragmented and vectorized before storage.
[0115] Furthermore, the content of the extracted medical knowledge base is segmented, and each segment is vectorized to construct a multi-dimensional fusion vector. Multi-dimensional fusion vector The metadata associated with each shard is stored in a vector database, and an index is created in the vector columns.
[0116] This invention employs a hybrid slicing strategy to slice the content of the medical knowledge base, as detailed below:
[0117] (1) Rule-based segmentation: For documents with a clear hierarchical structure in the medical extraction knowledge base (such as clinical guidelines and textbooks), segmentation is performed based on their inherent chapter titles (using tools such as pdfplumber and Pandoc for format parsing). For data entries in the medical extraction knowledge base (such as terminology descriptions, a single triplet, etc.), each entry is segmented according to existing delimiters (such as periods, line breaks, etc.).
[0118] It should be noted that pdfplumber is a Python-based PDF document parsing library that provides developers with a comprehensive solution from text extraction to table analysis. It is particularly good at handling machine-generated PDF files (not suitable for scanned documents). Its core functions include text extraction, table parsing, object location, and extraction of PDF metadata (such as creation date, modification date, producer, etc.).
[0119] It should be noted that Pandoc is a powerful open-source document format conversion tool. Its core functions include bidirectional conversion between multiple formats, processing document metadata such as title, author, and date, and supporting batch conversion via scripts.
[0120] (2) Semantic boundary segmentation: For continuous texts without a clear structure in the medical extraction knowledge base, a pre-trained text segmentation model (such as BERT-based segmenter) is used for segmentation.
[0121] It's important to note that the BERT-based segmenter is a text segmentation tool or model based on the BERT model. Its core strength lies in leveraging BERT's powerful contextual understanding capabilities to perform semantic segmentation of text, rather than traditional rule-based or statistical methods. This model can accurately understand sentence boundaries and paragraph structure, achieving sentence and paragraph segmentation.
[0122] (3) Limiting the size of the slide overlap window: The length of the text segmented by the rules and the segmentation model is limited. If the length of the slide exceeds the preset threshold (e.g., 512 characters), recursive slicing is used, and a slide overlap window (e.g., overlapping 50 characters) is used to avoid being cut off in the middle of the sentence or at key information.
[0123] For each segment, vectorization is performed, integrating basic semantics, weighted medical entities, and semantic-entity nonlinear interaction features to construct a multi-dimensional fused vector. The multi-dimensional fusion vector The calculation formula is:
[0124]
[0125] In the formula, Indicates the first fusion weight. This represents the basic semantic vector of the text in each slice. The dimension is 512; Indicates the second fusion weight. This represents the weighted medical entity feature vector in each slice. Indicates the third fusion weight. This represents element-wise multiplication, used to calculate the nonlinear interaction between the basic semantic vector and the weighted medical entity feature vector associated with the segmented sentences.
[0126] The BioBERT model, pre-trained in the medical field, is used to vectorize the fragmented text, capturing the literal semantics of the text and obtaining basic semantic vectors. .
[0127] It should be noted that BioBERT is a derivative of BERT (Bidirectional Encoder Representations from Transformers), specifically designed for the biomedical field. Through pre-training on large-scale biomedical literature data, it significantly improves the ability to understand technical terms and complex text. Its core goal is to address the limitations of general-purpose BERT in biomedical text processing through domain adaptation.
[0128] The specific process of extracting medical entity features from each slice of the medical knowledge base includes:
[0129] First, medical entity recognition tools are used to extract medical entities from each segment, resulting in a list of medical entities. For example, spaCy's medical-specific word segmenter is used as a tool for entity recognition in the medical field.
[0130] It should be noted that spaCy is an open-source, industrial-grade natural language processing (NLP) library that focuses on efficient, easy-to-use, and multilingual text processing tasks. It is widely used in production environments for text analysis, information extraction, and natural language understanding tasks.
[0131] Medical Entity List Specifically, it is expressed as follows:
[0132]
[0133] In the formula, This represents the first entity extracted from the current partition. This indicates the second entity extracted from the current fragment. This indicates the number of segments extracted from the current segment. One entity, The number of entities in a shard varies depending on the specific content of the shard.
[0134] Secondly, each entity is vectorized using an entity embedding model to obtain entity vectors. For example, PubMedBERT is used as an entity embedding model. Entity vectors. The dimension is 512.
[0135] It's worth noting that PubMedBERT is a pre-trained language model specifically designed for the biomedical field. By training entirely from scratch on biomedical text, it significantly improves performance in medical natural language processing (NLP) tasks. PubMedBERT's applications can cover the entire medical NLP workflow, including medical literature analysis, clinical text processing, and medical question answering.
[0136] If there are multiple entities in the partition ( For each entity The weight of each medical entity is calculated by combining its word frequency weight and semantic association weight in the segmented sentences. The formula is:
[0137]
[0138] In the formula, Representing word frequency weights, as entities Frequency of occurrence in segmented sentences; The semantic association weight is represented by the entity embedding vector. semantic vectors of segmented sentences The cosine similarity value is calculated; finally, it is normalized until the weight sum is 1. .
[0139] Ultimately, the weighted medical entity plus feature vector That is, the vectors of all entities in the partition are ordered by weights. The weighted sum is obtained. Medical entity plus feature vector. The calculation formula is as follows:
[0140]
[0141] In the formula, Representing entities Normalized weights in segmented sentences; Representing entities The embedding vector.
[0142] If there is no entity in the fragment ( ),but Let it be a vector of all zeros.
[0143] For the basic semantic vector in each slice Weighted medical entity feature vector and semantic-entity nonlinear interaction Enhanced fusion is performed to obtain a multi-dimensional fusion vector. Example, , .
[0144] Multi-dimensional fusion vector Metadata related to each shard is stored in a vector database, and an index is created on the vector columns. Each shard-related metadata includes the shard ID, source, and publication time. wait.
[0145] For example, the vector database used is Milvus. It's worth noting that Milvus is an open-source, cloud-native vector database designed for efficient storage, management, and retrieval of large-scale, high-dimensional vector data. It is suitable for scenarios requiring fast similarity searches, such as image retrieval, natural language processing, recommendation systems, and intelligent customer service.
[0146] After the content of the medical knowledge base is segmented and vectorized, it is stored in a vector database, and an index is created in the vector columns. Thus, the medical knowledge base is established. When a question is posed to the large language model, the system first retrieves relevant knowledge based on the question content. The retrieved information is then merged with the original question to form new prompt words, which the large language model then uses to provide an answer.
[0147] Furthermore, the specific process of retrieving medical extraction knowledge bases includes:
[0148] S101: Different agents have different prompts and data text to be processed according to specific functional requirements. All these texts are used as query texts for retrieving the knowledge base and are processed into fragmented vectors to obtain query fusion vectors. .
[0149] Query fusion vector Specifically, it is expressed as follows:
[0150]
[0151] In the formula, This represents the first query vector in the query text slice. This represents the second query vector in the query text slice. Query the total number of text fragments.
[0152] S102: Based on query fusion vector For each query fusion vector Perform a semantic query on the vector database. The vector database will use its index to query the database for results related to vectors. Vectors with similar semantics and their corresponding knowledge, each vector Searchable ( =3) A total of 3 items were obtained Data entries. The m data entries returned by the vector database query are reordered, i.e., a comprehensive ranking score is calculated. The ranking is based on a combination of similarity and timeliness. The overall ranking score... The calculation formula is:
[0153]
[0154]
[0155] in, Represents semantic similarity value, The data is returned from the vector database during the query. Indicates the weight of timeliness. This represents the first weighting coefficient. This represents the second weighting coefficient. ; Indicates the current search time. Indicates the publication time of the segmented content. Indicates a time base, usually =365 days, or 1 year); This represents the attenuation coefficient, for example. , The weight of content decreases to its original value every year it expires. Prioritize the latest content. Range of values Example, 8, .
[0156] Select overall score The highest and greater than the preset threshold Each data point is used as the corresponding output for this round of retrieval. For example, the preset threshold is 0.75. .
[0157] The output of the medical knowledge base retrieved in this round, i.e. the corresponding knowledge retrieved in this round, is combined with the initial prompt word (i.e., the first prompt word) and the data text to be processed to form a new prompt word (i.e., the second prompt word), and then the large language model is called to process it and generate an answer.
[0158] S4: Each agent calls the fine-tuned large language model according to its corresponding second prompt word, performs information extraction and standardization, and outputs the extracted medical information.
[0159] The large language model assigns different agents to handle their respective tasks based on their corresponding second prompt words, depending on the data processing task.
[0160] Reference Figure 2 , 3 As shown, the specific process of processing the preprocessed Chinese electronic medical records using the fine-tuned large language model includes:
[0161] S401: Retrieve relevant medical document knowledge from the medical extraction knowledge base, use the document structure parsing agent to parse the document structure of the preprocessed Chinese electronic medical records, and output the parsed electronic medical records.
[0162] The document structure parsing agent combines the semantic analysis capabilities of a large language model with the standard chapter rules of a medical knowledge base to achieve accurate division and correction of medical record chapters. The large language model performs semantic analysis on preprocessed text fragments and tabular JSON data, and, combined with the standard chapter list and chapter boundary definition rules returned by the knowledge base, identifies chapter identifiers such as "present illness history" and completes table-chapter bindings such as "laboratory examination form belongs to the auxiliary examination chapter." The large language model verifies the chapter order (e.g., "chief complaint" must precede "present illness history") and outputs a standardized JSON structure containing fields such as "medical record type," "chapter list," and "chapter-content."
[0163] S402: Retrieve the medical knowledge base to obtain the corresponding medical knowledge, use the extraction agent to extract entity attribute relationships from the parsed electronic medical records, extract feature descriptions with medical clinical significance, and obtain the data stream of "clinical statement-original entity-entity relationship" and "time expression-time associated event object".
[0164] The extraction agent employs a parallel semantic parsing and joint extraction method using a large language model to achieve multi-dimensional extraction of clinical features, medical entities, entity relationships, and time expressions. Structured chapter data is divided into different chapters and distributed to multiple parallel processing units to improve processing efficiency. The large language model uses a joint extraction method, simultaneously extracting structured attributes of clinical features such as "cough-nature (dry cough)" and "cough-severity (moderate)" from the text based on knowledge data from the knowledge base, including clinical feature keywords, entity type sets, and relationship type sets. It also identifies medical entities such as "amoxicillin (drug entity)" and "cough (symptom entity)" and their types, and determines semantic relationships between entities such as "amoxicillin-treatment-cough". For time recognition, the large language model identifies absolute, relative, and fuzzy times (such as "2025-01-01", "day 2 after admission", "early morning") in the text and marks time-related event objects (such as "two days ago" associated with the "start of coughing" event). Finally, the data is split: "clinical features - original entities - entity relationships" is encapsulated into the first data stream, and "time expressions - time-related event objects" is encapsulated into the second data stream, which are then output to the medical terminology standardized intelligent agent and the time expression standardized intelligent agent, respectively.
[0165] S403: Retrieve relevant medical terminology knowledge from the medical extraction knowledge base, and use the medical terminology standardization intelligent agent to standardize, rewrite and map the data flow of "clinical statement-original entity-entity relationship", and output standardized text expression data.
[0166] The medical terminology standardization intelligent agent utilizes the ambiguity resolution and contextual understanding capabilities of a large language model, combined with standard terminology mapping rules from a medical knowledge base, to achieve the unification and standardization of medical terminology rewriting. For example, for original entities such as "aspirin" and "myocardial infarction," it matches them with standard terms (e.g., "aspirin" and "myocardial infarction") and their corresponding codes (e.g., RxNorm, ICD-10) from the knowledge base; for ambiguity resolution, for polysemous terms such as "cillin," the large language model combines the context (e.g., "bacterial infection relieved after taking ciliate") and the applicable scenarios in the knowledge base to determine a unique standard term; for non-standard features such as "high fever," the large language model rewrites them into standardized descriptions such as "high fever, body temperature ≥39℃."
[0167] S404: Retrieve the medical extraction knowledge base to obtain the corresponding medical time expression knowledge, and use the time expression standardization intelligent agent to process "time expression-time related event object" into "specific time scale or time interval-related event" time expression standard format data.
[0168] The time expression standardization agent achieves the standardization of time expressions through the time logic reasoning capability of the large language model. Its specific processing includes: (1) Anchor point determination: The absolute values such as "admission time (e.g., "2025-01-01")" and "document creation time" extracted from the medical records are used as time anchor points. If there is no anchor point, a marker is triggered, and no time conversion processing is performed subsequently. (2) Time conversion: The large language model calculates "the second day after admission" as an absolute time (e.g., "2025-01-03") and maps "early morning" to time intervals such as "05:00-08:00". (3) Logic verification: The large language model verifies the time logic such as "the examination time is not later than the discharge time", corrects contradictory expressions, and solves problems such as inconsistent time descriptions and logical confusion.
[0169] S405: Utilize the inspection output agent to perform multi-dimensional verification operations on the standardized text expression data and time expression standard format data, and output the extracted medical information.
[0170] The output agent for the inspection utilizes the multi-dimensional rule verification capabilities of a large language model, combined with the document specifications of a medical knowledge base, to verify, correct, and integrate the results. The specific processing includes: multi-dimensional verification: the large language model verifies whether "core entities are missing," "whether the relationship between drug dosage and treatment is consistent," and "whether the time logic is reasonable." Data integration: according to the document specifications of the knowledge base, the verified "standardized terminology," "standardized time," and "chapter structure" are integrated into a final structured JSON containing "basic information," "chapter," "entity," "relationship," and "timeline," achieving integrated data output.
[0171] In this embodiment of the invention, the fine-tuning of the large language model includes a multi-agent task-supervised fine-tuning stage and an agent-cooperative reinforcement learning stage.
[0172] Multi-agent task-supervised fine-tuning stage: Based on five types of agent-specific labeled datasets (e.g., document structure parsing labeled dataset containing "chapter-table boundary labels", entity extraction labeled dataset containing "entity-relationship labels", medical and time terminology standardization labeled dataset containing "raw terminology-standard terminology mapping labels") and open-source datasets (e.g., entity recognition (CMeEE, CHIP-2020, KUAKE-NER, etc.), relation extraction (CMeIE, TCM-KG, etc.), standard mapping (CHIP-CDN, Chinese medical terminology system, etc.), a "multi-task loss fusion" mechanism is constructed. This mechanism enables the model to simultaneously optimize the accuracy of each agent's specific task and ensures that there are no logical conflicts in the results of different agents through agent task consistency loss (e.g., the "aspirin" output by the extraction agent is consistent with the "aspirin (RxNorm encoding: 1191)" mapped by the terminology standardization agent). The specific formula is defined as follows:
[0173] Document structure parsing loss :Measure the overlap between the model's predicted structure and the actual structure:
[0174]
[0175] In the formula For real structural labels, Predict structural labels for the model. Indicates the number of elements in the set.
[0176] Entity Relationship Extraction Loss Fusion of cross-entropy loss (entity classification) and F1 loss (relationship prediction):
[0177]
[0178] In the formula ( For the number of entities, For entity type,
[0179] For the first Feature representation vectors of each entity calculate Belongs to type The probability is ; ( For the accuracy of relation prediction, (For predicting recall based on relationships). Hyperparameter (range of values) ).
[0180] Loss of medical terminology standardization Align the term vectors output by the model with the standard term vectors from the medical knowledge base.
[0181]
[0182] In the formula For the number of terms, The first output of the model A vector of original terms, Let i be the vector of the i-th standard term. The cosine similarity function is used. Temperature parameter (values) ).
[0183] Time-expression standardized loss : Measures the deviation between the predicted time interval and the standard time interval:
[0184]
[0185] In the formula The number of time expressions. The start and end values of the time interval predicted by the model. The standard time interval is marked with its start and end values.
[0186] Check for verification loss :
[0187]
[0188] In the formula To verify the number of samples, Indicates the first The group results were in line with expectations. This indicates that it did not meet expectations. The model represents the first The probability that a given validation sample will "meet expectations" is within a given range. .
[0189] Agent task consistency loss Used to constrain the logical consistency of results between preceding and subsequent agents:
[0190]
[0191] In the formula The number of agent result pairs (e.g., entity extraction result - terminology standardization result pair, terminology standardization result - inspection result pair). This is the hidden state vector of the preceding agent's result. This is the hidden state vector of the subsequent agent's result. This is the cosine similarity function.
[0192] General Supervisor Fine-tunes Losses :
[0193]
[0194]
[0195] in, This indicates the first fine-tuning weight coefficient. This represents the document structure parsing loss. This indicates the second fine-tuning weighting coefficient. This represents the loss from entity relationship extraction. This indicates the third fine-tuning weight coefficient. This indicates the loss of standardization in medical terminology. This indicates the fourth fine-tuning weight coefficient. Represents the time-normalized loss. This indicates the fifth fine-tuning weighting coefficient. This indicates that the check and verification loss has been detected. This indicates the fifth fine-tuning weighting coefficient. This represents the loss of task consistency for the intelligent agent.
[0196] Example, .
[0197] Agent Cooperative Reinforcement Learning Phase: Using the "multi-agent cooperative processing flow" as the environment, based on the supervised fine-tuned large language model, the Proximal Policy Optimization (PPO) algorithm is adopted, and a multi-dimensional reward function of "subtask accuracy reward + cross-agent interaction reward" is designed. The result correlation between the preceding and subsequent agents is strengthened through interaction rewards to achieve global consistency optimization.
[0198] State space: Environment state Defined as a set of "preprocessed Chinese medical record text and intermediate results of completed agents," such as chapter data output by document structure parsing agents and results from previous extraction agents. Its function is to provide the policy network with contextual information for the current task, including the original text input and the accumulated output of historical tasks.
[0199] Action space: Action Defined as a set of "processing results of the current agent", such as the entity list output by the entity extraction agent, the standard term mapping results output by the terminology standardization agent, etc.
[0200] Policy Network: With a supervised fine-tuned large language model as its core, the sequence generation capability of the large language model is naturally suited to the mapping requirement "from state to action", that is, the input state Simultaneously, it retrieves core knowledge relevant to the current task from the medical extraction knowledge base as context, and outputs actions compatible with historical results. .
[0201] Reward Function: Designing a multi-dimensional collaborative reward function for total reward. The formula for comprehensively evaluating "single-task accuracy" and "cross-task correlation" is as follows:
[0202]
[0203]
[0204] In the formula, each sub-award is defined as follows:
[0205] Draw rewards:
[0206] Terminology standardization award:
[0207] Time expression standardization reward:
[0208] Verification reward:
[0209] Cross-agent interaction rewards:
[0210] in, For entity extraction precision, For entity recall; For relation extraction precision, Extracting recall rates based on relationships; For example, to extract weighting coefficients, =0.65, prioritizing entity extraction weight; The number of standard terms for correct mapping, (To extract the number of correctly extracted entities). The loss is standardized for the time expression; Check the number of verification passes. The total number of checksums required; This refers to the results of previous tasks, such as entity extraction results. For results of subsequent tasks, such as terminology standardization results, For mutual information, For entropy, This represents the maximum mutual information calculated offline.
[0211] Example, weighting coefficients: 5.
[0212] The PPO strategy update objective uses the PPO clip objective function. Limit the policy update range to ensure stable convergence:
[0213]
[0214] In the formula, For large language model parameters; Indicates the old strategy The mathematical expectation of the corresponding state and action distribution; The probability of the action under the current policy. The action probability of the old strategy, i.e., the strategy of the previous iteration; For the dominant function, For the reward function, For a value network, the expected reward for predicting state s is given by the parameter . ; The clip coefficient is used to avoid policy mutations.
[0215] Based on the same inventive concept, this invention also proposes a Chinese electronic medical record data processing system based on a large language model, used in the aforementioned Chinese electronic medical record data processing method based on a large language model. The two embodiments share the same or similar technical features, which will not be elaborated further below.
[0216] Reference Figure 4 As shown, the system includes:
[0217] The data acquisition module collects Chinese electronic medical records;
[0218] The data preprocessing module preprocesses the input data to obtain preprocessed Chinese electronic medical records.
[0219] The medical knowledge base includes a semantic dataset of standardized medical ontology and terminology, a semantic dataset of document specifications, a semantic dataset of medical clinical guidelines and rules, and a dataset of authoritative medical literature and case evidence. Its content is stored in a vector database after being segmented and vectorized.
[0220] The document structure parsing module retrieves relevant knowledge from the medical knowledge base, performs document structure parsing, and outputs the parsed electronic medical record.
[0221] The information extraction module retrieves relevant knowledge from the medical extraction knowledge base, extracts feature descriptions with medical clinical significance, and obtains a data stream of "clinical statement-original entity-entity relationship" and a data stream of "temporal expression-temporally related event object".
[0222] The medical terminology standardization module retrieves relevant knowledge from the medical extraction knowledge base, standardizes and rewrites the data flow of "clinical statement-original entity-entity relationship", and outputs standardized expression data.
[0223] The Time Expression Standardization Module is used to process the data stream of "Time Expression - Time-related Event Object" into a standard format of time expression data of "Specific Time Scale or Time Interval - Related Event".
[0224] The output module is checked. This module retrieves the medical extraction knowledge base, performs multi-dimensional verification operations, and outputs the extracted medical information.
[0225] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include ROM, RAM, disk, or optical disk, etc.
[0226] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A Chinese electronic medical record data processing method based on a large language model, characterized in that, The application comprises: obtaining a Chinese electronic medical record, and obtaining a preprocessed Chinese electronic medical record after data preprocessing; inputting the preprocessed Chinese electronic medical record into a multi-agent based on a large language model after fine-tuning, and obtaining a first prompt word from the multi-agent according to different processing tasks of the multi-agent and input data; obtaining corresponding medical knowledge from a medical extraction knowledge base according to the first prompt word, and obtaining a second prompt word in combination with the first prompt word; each agent calls a large language model after fine-tuning according to the respective second prompt word, performs information extraction and standardization processing, and outputs extracted medical information. The medical extraction knowledge base is a medical knowledge graph organized in the form of a "(entity)-(relationship, attribute)-(entity, attribute description)" triple, which includes a semantic data set of standardized medical ontology and terms, a semantic data set of document specifications, a semantic data set of medical clinical guidelines and rules, and a data set of medical authoritative literature and case evidence.
2. The Chinese electronic medical record data processing method based on a large language model according to claim 1, characterized in that, The application uses a large language model after fine-tuning to process the preprocessed Chinese electronic medical record, and the specific process includes: retrieving corresponding medical document knowledge from the medical extraction knowledge base, using a document structure analysis agent to analyze the document structure of the preprocessed Chinese electronic medical record, and outputting the analyzed electronic medical record; retrieving corresponding medical knowledge from the medical extraction knowledge base, using an extraction agent to extract entity attributes and relationships from the analyzed electronic medical record, and extracting feature descriptions with medical clinical significance to obtain "clinical statement-original entity-entity relationship" data streams and "time expression-time related event object" data streams; retrieving corresponding medical terminology knowledge from the medical extraction knowledge base, using a medical terminology standardization agent to standardize and map the "clinical statement-original entity-entity relationship" data stream, and outputting standardized text expression data; retrieving corresponding medical time expression knowledge from the medical extraction knowledge base, using a time expression standardization agent to process the "time expression-time related event object" into "specific time scale or time interval-related event" time expression standard format data; using a check output agent to perform multi-dimensional checking operations on the standardized text expression data and time expression standard format data, and outputting extracted medical information.
3. The Chinese electronic medical record data processing method based on a large language model according to claim 1, characterized in that, The application preprocesses the Chinese electronic medical record to obtain a preprocessed Chinese electronic medical record, and the processing process includes: converting structured or unstructured text data in the Chinese electronic medical record into a JSON array containing "chapter" and "content" fields; converting structured or unstructured table data in the Chinese electronic medical record into JSON format table data, which includes "table name", "table header" and "data".
4. The Chinese electronic medical record data processing method based on a large language model according to claim 3, characterized in that, In the data preprocessing of the table data in the Chinese electronic medical record, the process includes: identifying and extracting text content and row and column coordinates of each text block from the table of the Chinese electronic medical record; The table header and cell content of the table data in the Chinese electronic medical record are extracted, wherein for a nested table header, the main table header is formed by combining multiple columns into one column, and the sub-table header is one column, the main table header column is traversed, and each column is traversed from top to bottom and from left to right to splice to obtain nested table header data; for cross-page table data, it is judged whether the number of table columns and the table header content are consistent, if consistent, the table data can be directly merged, if inconsistent, the similarity S of the cross-page table is calculated, if the similarity S reaches a preset threshold, the table fragments of multiple pages are merged; The extracted table header and cell content are converted into JSON format table data.
5. The Chinese electronic medical record data processing method based on a large language model according to claim 1, characterized in that, The content of the medical extraction knowledge base is fragmented, each fragment is vectorized, and a multi-dimensional fusion vector is constructed The multi-dimensional fusion vector and the metadata related to each fragment are stored in a vector database, and an index is established in a vector column; the calculation formula of the multi-dimensional fusion vector is: wherein, denotes a first fusion weight, denotes a base semantic vector for each shard of text, denotes a second fusion weight, denotes a weighted medical entity feature vector for each shard, denotes a third fusion weight, denotes an element-wise multiplication.
6. The Chinese electronic medical record data processing method based on a large language model according to claim 1, characterized in that, The specific process of retrieving the medical extraction knowledge base includes: Based on the pretreated Chinese electronic medical record, the large language model is questioned, that is, different prompt words are set according to the specific requirements combined with the input data, the prompt words are processed by slicing and vectorization, and a query fusion vector is obtained ; query fusion vector The vector database for sharded storage of medical extraction knowledge base is queried, and the vector database will be queried by index with the query fusion vector Vectors semantically close to the query fusion vector and corresponding medical knowledge A total of m pieces of data are queried The m data points are combined based on their similarity and timeliness, reordered, and a comprehensive ranking score is calculated. Select the comprehensive ranking score The highest and greater than the preset threshold Each data point is used as the corresponding output for this round of retrieval.
7. The Chinese electronic medical record data processing method based on a large language model according to claim 1, characterized in that, In the process of fine-tuning the large language model, the loss function used is: In the formula, denotes the first fine-tuning weight coefficient, denotes the document structure parsing loss, denotes the second fine-tuning weight coefficient, denotes the entity relation extraction loss, denotes the third fine-tuning weight coefficient, denotes the medical terminology standardization loss, denotes the fourth fine-tuning weight coefficient, denotes the time expression standardization loss, denotes the fifth fine-tuning weight coefficient, denotes the examination verification loss, denotes the fifth fine-tuning weight coefficient, denotes the agent task consistency loss.
8. A Chinese electronic medical record data processing system for implementing the Chinese electronic medical record data processing method based on a large language model according to claim 1, characterized in that, The system comprises: A data acquisition module acquires Chinese electronic medical records; A data preprocessing module pre-processes the input data to obtain pre-processed Chinese electronic medical records; A medical extraction knowledge base, comprising a semantic data set of standardized medical ontology and terminology, a semantic data set of document specification, a semantic data set of medical clinical guidelines and rules, and a data set of medical authoritative literature and case evidence base, the contents of which are stored in a vector database after being fragmented and vectorized; A document structure analysis module retrieves the medical extraction knowledge base to obtain corresponding knowledge, analyzes the document structure, and outputs the analyzed electronic medical record; An information extraction module retrieves the medical extraction knowledge base to obtain corresponding knowledge, extracts feature descriptions with medical clinical significance, and obtains "clinical statement-original entity-relation" data stream and time expression data stream; A medical terminology standardization module retrieves the medical extraction knowledge base to obtain corresponding knowledge, performs standardization rewriting and mapping, and outputs standardized expression data; A time expression standardization module is used to process the time expression data stream into "specific time scale or time interval-associated event" time table standard format data; An inspection output module retrieves the medical extraction knowledge base, performs multi-dimensional verification operation, and outputs the extracted medical information.
Citation Information
Cited By
Medical text information processing method and device, electronic equipment and storage medium
CN121809410A
A medical record entity-based surgical operation coding generation method, device and equipment
CN122337454A
Knowledge base construction method and device based on multi-modal agent, equipment and medium
CN122366598A