Full-process automatic voice-driven electronic medical record generation system
The medical record generation system, which utilizes high-precision voice acquisition, deep learning, and knowledge enhancement, addresses the shortcomings of existing electronic medical record systems in areas such as medical terminology recognition, structured processing, and privacy protection. It achieves efficient and secure medical record generation, meeting the standardization requirements of medical information systems.
Patent Information
- Application Number
- CN202511125205.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-12-02
AI Technical Summary
Existing electronic medical record generation systems suffer from insufficient speech recognition accuracy, weak structured processing capabilities, difficulty in handling multi-role dialogues, and inadequate privacy protection, resulting in low generation efficiency, poor accuracy, and security risks.
A high-precision microphone array and GMM-UBM voiceprint recognition model are used, combined with deep learning-based VAD technology for voice acquisition and preprocessing; Faster-Whisper-large-v3 and BiLSTM+CRF models are used for medical terminology recognition and logical contradiction detection; a large language model and medical knowledge base are integrated, and similar case retrieval and rule verification are performed through the FAISS knowledge base to generate structured medical records; and an end-to-end data protection mechanism is implemented.
It improves the accuracy of medical terminology recognition, reduces medical record generation time, ensures the structured quality and accuracy of medical records, reduces the risk of privacy leaks, meets legal and regulatory requirements, and enhances the system's integration efficiency and security.
Smart Images

Figure CN121054166A_ABST
Abstract
Description
Technical Field
[0003] This invention relates to the field of intelligent medical technology, specifically to a fully automated, voice-driven electronic medical record generation system. Background Technology
[0004] Currently, the generation of electronic medical records in medical institutions mainly relies on manual input by doctors, which has the following technical shortcomings:
[0005] (1) Inefficiency: Doctors need to type and record while taking medical advice, with each doctor spending an average of 2-3 hours a day writing medical records, which seriously encroaches on consultation time. This is especially true for older doctors, whose slow typing speed further exacerbates the efficiency problem.
[0006] (2) High error rate: Manual entry is prone to terminology errors (such as drug name, disease code errors) and logical contradictions (such as "penicillin allergy" after "no history of allergy"). Existing electronic medical record systems lack a real-time verification mechanism.
[0007] (3) Insufficient structure: Traditional voice input systems (such as CN107564571A) only realize simple conversion from voice to text, and cannot automatically extract medical entities and generate structured medical records that comply with the "Electronic Medical Record Application Specifications".
[0008] (4) Identity confusion: Existing solutions (such as CN113724695A) only distinguish between doctors and patients by voiceprint, which cannot effectively handle the scenario of multiple people interrupting during ward rounds, resulting in incorrect role labeling in the generated medical records.
[0009] (5) Security risks: Voice data is usually uploaded directly to the cloud for processing, which poses a risk of patient privacy leakage and does not comply with the requirements of the "Guidelines for the Security of Medical and Health Data".
[0010] To address the above issues, some improvements have been attempted in existing technologies:
[0011] (1) Patent CN118072901B uses the WGAN model to expand the medical speech dataset, but it does not solve the problem of low accuracy in medical terminology recognition.
[0012] (2) Patent CN116959453A achieves voiceprint verification and speech error correction, but lacks structured processing capabilities supported by medical knowledge base.
[0013] (3) Patent CN120319380A focuses on encrypted storage of electronic medical records, but does not optimize the accuracy and efficiency of the front-end generation process.
[0014] Therefore, there is an urgent need for an end-to-end solution that integrates high-precision medical ASR, knowledge-enhanced LLM medical record generation, and multi-level security protection. Summary of the Invention
[0015] (I) Technical Issues
[0016] Current electronic medical record (EMR) generation systems face a series of critical technical challenges in practical clinical applications, primarily in the following aspects: First, in the speech recognition stage, existing general-purpose speech recognition systems suffer from severely insufficient accuracy in recognizing medical terms, especially complex drug names (such as "atorvastatin calcium tablets") and disease codes (such as "ICD-10J18.9"), with error rates generally reaching 15-20%. This not only increases the workload for doctors in subsequent revisions but also poses a serious medical safety hazard due to errors in recognizing key terms. Second, regarding the structuring of medical records, the unstructured text output by traditional systems cannot meet the standardized requirements of modern Hospital Information Systems (HIS). Doctors must spend a significant amount of time manually extracting and organizing key fields such as chief complaint, present illness, and diagnosis. Statistics show that each medical record requires an average of 10-15 minutes for post-structuring processing, severely impacting clinical work efficiency. Furthermore, existing systems... The lack of an effective clinical knowledge support system is a significant problem. This system cannot automatically detect contradictions in diagnostic logic (such as the coexistence of a diagnosis of "no smoking history" and "lung squamous cell carcinoma"), nor can it effectively identify contraindications to medication use (such as the use of cephalosporin antibiotics in patients with penicillin allergies). This lack of knowledge support makes it difficult for the system's output medical records to meet clinical requirements. Furthermore, in handling multi-role dialogues, existing technologies struggle to accurately distinguish between different roles in doctor-patient conversations, often leading to speaker confusion in the generated medical records, severely impacting their clinical value. Regarding system integration, speech recognition, medical record generation, and the HIS system are often independent, requiring multiple manual data transfers, resulting in low overall efficiency and increasing the risk of human error. Finally, in terms of privacy protection, most existing solutions employ cloud-based speech processing, posing a serious risk of patient privacy data leakage and failing to meet the stringent requirements of regulations such as the "Guidelines for Medical and Health Data Security" and the "Personal Information Protection Law."
[0017] (II) Technical Solution
[0018] This invention provides an innovative, fully automated, voice-driven electronic medical record generation system. The core technology of this system achieves a revolutionary medical record generation experience through the collaborative work of the following key subsystems: In the medical-grade voice acquisition and preprocessing subsystem, we adopt an advanced multimodal input design and a high-precision microphone array with medical-grade microphones (signal-to-noise ratio ≥96dB). This system supports far-field sound pickup up to 3 meters away and noise suppression capability ≥30dB. It can simultaneously connect to various acquisition terminals such as fixed microphones in the examination room, portable doctor's recorders, and operating room-specific sound pickup devices. Based on the GMM-UBM voiceprint recognition model and keyword wake-up technology, it achieves an intelligent voice activation function with a false wake-up rate of <0.1%. Simultaneously, it integrates the WebRTCNS algorithm for environmental noise suppression and uses deep learning-based VAD (Voice Activity Detection) technology to achieve accurate voice segmentation with a minimum segment length of 300ms.
[0019] In terms of the speech recognition engine for enhanced medical terminology, this system has made multi-level technological innovations: using Faster-Whisper-large-v3 as the basic architecture, it achieves ultra-low latency processing under GPU acceleration; by constructing a medical-specific dictionary (covering disease names and ICD-10 codes, generic and brand names of drugs, examination and testing items, and abbreviations of medical terms), it achieves accurate recognition of professional terms; at the same time, it integrates a BiLSTM+CRF model to achieve terminology consistency checks (such as ensuring the consistency between the expressions "heart attack" and "myocardial infarction") and logical contradiction detection (such as identifying the contradiction between "no history of allergies" and the subsequent mention of drug allergies); in terms of speaker separation, it uses the PyAnnote tool to combine voiceprint features (based on doctor and patient voiceprint databases), dialogue context features (such as "doctor asks:" / "patient answers:" patterns), and professional terminology usage frequency analysis and other multi-dimensional information to achieve accurate role labeling.
[0020] In terms of knowledge-enhanced medical record generation systems, this solution achieves intelligent medical record generation through deep integration of a large language model and a medical knowledge base: Based on the deepseek-r1 basic model, it uses several anonymized electronic medical records (annotated with entity relationships) and the latest clinical guidelines for three-stage training (general corpus pre-training → medical corpus fine-tuning → reinforcement learning optimization). A knowledge retrieval module using the FAISS vector database (HNSW algorithm, 768 dimensions) is constructed, which can retrieve similar historical cases in real time and return the results. At the same time, a rule verification engine with multiple functions such as drug conflict detection (based on drug knowledge graph), diagnostic logic verification (symptom-disease matching degree check) and treatment plan rationality check is developed. The final output is a structured medical record in JSON format that includes both standard SOAP format (subjective, objective, assessment, and planning) and 30+ standard fields, while providing natural language descriptions to meet the needs of different use cases.
[0021] In terms of multimodal interaction and system integration, this system achieves seamless integration of clinical workflows: it has developed an intelligent interactive interface that supports drag-and-drop correction and automatic highlighting of key clinical information; it has achieved deep integration with the HIS system based on the HL7 / FHIR standard; it supports direct connection of data from the hospital's examination and testing system and real-time data synchronization across multiple terminals (PC, mobile, and hospital terminals); and it has achieved intelligent functions such as automatic form filling and medical order execution status tracking in the HIS system through RPA technology.
[0022] Regarding the end-to-end security and compliance system, this solution constructs an end-to-end medical data protection mechanism: it adopts an edge computing architecture to ensure that voice data is processed on a local A100 GPU server, implements a hierarchical desensitization strategy (name → ID_001) for encrypted transmission, establishes the principle of least privilege access control, is equipped with complete operation audit logs, passes the Level 3 certification of Information Security Protection 2.0, ensures compliance with GDPR / Personal Information Protection Law and other regulations, and uses blockchain technology to realize the storage of key operations (timestamp error <1s). Attached Figure Description
[0023] Figure 1 The overall architecture design of the system of this invention is demonstrated. It adopts a modular design scheme with a four-layer core architecture (voice acquisition layer, speech-to-text layer, medical record generation layer, and output interaction layer). The data flow and interaction relationship between each functional module are presented in detail. The voice acquisition layer is responsible for high-quality voice acquisition in the medical environment, the speech-to-text layer realizes speech recognition with enhanced medical terminology, the medical record generation layer completes the generation of knowledge-guided structured medical records, and the output interaction layer provides a multimodal doctor interaction interface and system integration capability.
[0024] Figure 2 The training framework of knowledge-enhanced LLM is systematically presented, including the large-scale medical corpus collection and annotation process in the data preparation stage, the three-step optimization strategy (pre-training, fine-tuning and reinforcement learning) in the model fine-tuning stage, and the vectorization processing and information retrieval mechanism of the knowledge retrieval module. At the same time, the key indicators and optimization methods for model performance evaluation are also shown.
[0025] Figure 3 From a system integration perspective, the paper describes the interface solution with the hospital's existing information systems (such as HIS, EMR, etc.), including key technical details such as data format conversion, interface protocol adaptation, and business process integration. Detailed Implementation
[0026] This invention aims to provide an innovative technical solution. To fully and clearly present the core concept, technical details, and advantages of this invention, the specific embodiments of the invention will be described in detail below with reference to the accompanying drawings. It should be clearly stated that the embodiments described in this specification are only a part of the many possible implementations of this invention, intended to help understand the operating mechanism of the invention, and not to limit the scope of protection of this invention. Any equivalent or improved embodiments derived from the concept of this invention without inventive effort should be considered to fall within the protection scope of this invention.
[0027] Example 1: Overall Structure of a Medical Voice Electronic Medical Record System (Based on) Figure 1 )
[0028] This invention provides an overall structural embodiment of an advanced medical voice electronic medical record system, aiming to revolutionize traditional medical record writing methods and achieve intelligent and efficient generation of medical records through voice interaction. The system is as follows: Figure 1 As shown, its design philosophy is to modularize and hierarchically organize the complex voice processing and medical record generation processes, ensuring the system's scalability, stability, and professionalism. The entire system architecture works closely together from bottom to top, forming a complete intelligent processing chain from voice to structured electronic medical records.
[0029] 3.1 Voice Acquisition Layer: The Cornerstone and Environmental Adaptation for High-Quality Voice Input
[0030] As the primary component of the system, the voice acquisition layer aims to accurately capture raw voice data in the medical environment and perform preliminary optimizations to cope with the complex and ever-changing clinical noise environment, thus laying a solid foundation for subsequent voice recognition.
[0031] 3.1.1 Medical Microphone Arrays: Directional Pickup and Noise Suppression
[0032] This embodiment employs a medical microphone array specifically designed for medical environments. Instead of consisting of a single microphone, this array is an integrated system comprising multiple high-performance microphone units, strategically deployed in key clinical locations such as doctor's offices, ward rounds, or surgical areas.
[0033] The microphone array utilizes advanced acoustic signal processing technologies, including but not limited to beamforming and sound source localization algorithms. By precisely analyzing the differences in time delay and phase of sound waves received by different microphones, the system can intelligently identify and lock onto the direction of the speaker's (e.g., a doctor's) voice. Once the sound source is located, the array forms a directional "earpiece," maximizing the intensity of the target speech signal while effectively suppressing ambient noise from other non-target directions, such as noise from the corridor outside the examination room, conversations of other medical staff, operating noise from medical equipment, or non-speech noise from the patient. This directional pickup capability significantly improves the signal-to-noise ratio of the target speech, ensuring that the captured doctor's speech remains clear and complete even in complex and noisy medical settings. Furthermore, some advanced microphone arrays may incorporate acoustic echo cancellation to eliminate echoes generated by speaker-generated speech (e.g., remote speech during telemedicine), further purifying the speech input. The microphone array can be flexibly configured with different geometries, such as linear arrays or circular arrays, to optimize its performance based on the acoustic characteristics of the specific application scenario. The output of this module is a preprocessed multi-channel raw speech signal, which provides rich and high-quality acoustic information for subsequent processing.
[0034] 3.1.2 WebRTC Noise Reduction Module: Deep Denoising and Real-time Speech Enhancement
[0035] Although the microphone array has undergone preliminary noise reduction, various environmental noises may still remain in the acquired speech signal. Therefore, this embodiment introduces a WebRTC noise reduction module for deeper and more refined purification of the speech data.
[0036] The WebRTC noise reduction technology has been widely validated and adopted due to its superior performance in real-time communication applications (such as web video conferencing). This module integrates a series of complex digital signal processing (DSP) algorithms, and its working principle mainly includes:
[0037] Spectral subtraction: By continuously estimating the spectral characteristics of background noise and subtracting it from the spectrum of a mixed signal containing speech and noise, the pure speech component can be effectively separated.
[0038] Adaptive Filters: The module contains adaptive filters that can learn and track the dynamic characteristics of noise in real time and automatically adjust their filtering parameters according to changes in noise to effectively cancel non-stationary noise (such as occasional coughs or keyboard clicks).
[0039] Voice Activity Detection (VAD): VAD technology can accurately distinguish between speech segments and non-speech (pure noise) segments. Based on the VAD's assessment, the noise reduction algorithm applies a more aggressive noise reduction strategy to non-speech segments, while maintaining moderate noise reduction on speech segments to avoid excessive damage to speech quality.
[0040] This WebRTC noise reduction module operates efficiently in real-time or near real-time, ensuring that deeply noise-reduced speech data can be transmitted to the next layer in a timely and high-quality manner. The result is a clean speech stream with a significantly improved signal-to-noise ratio and substantial reduction in background noise. This plays a crucial role in greatly improving the recognition accuracy of the speech-to-text module and the system's robustness in complex environments.
[0041] 3.2 Speech-to-Text Layer: Intelligent Conversion from Acoustic Signals to Intelligible Text
[0042] The speech-to-text layer is a crucial step in the system's conversion of human speech into machine-processable text information. This embodiment integrates cutting-edge speech recognition technology at this layer, supplemented by specialized optimizations tailored to the medical field, ensuring the accuracy and professionalism of the conversion.
[0043] 3.2.1 Faster-Whisper Module: A Highly Efficient and Accurate Speech Recognition Engine
[0044] This embodiment uses the Faster-Whisper model as the core speech recognition engine. The Faster-Whisper is not a simple general-purpose recognizer, but a version of the Whisper model that has undergone a series of deep optimizations and accelerations.
[0045] Technical Advantages: Faster-Whisper inherits the superior performance of the Whisper model in multilingual support and multi-task processing (such as speech recognition, language identification, and speech translation). Building upon this foundation, it significantly improves inference speed and reduces memory consumption through underlying technical improvements (such as using the CTranslate2 backend for inference optimization and model quantization techniques), making deployment in real-world medical applications possible and enabling near real-time speech-to-text conversion.
[0046] Its core operating principle is based on the advanced Transformer architecture, which can directly take the acoustic waveform processed by the speech acquisition layer as input. Through a complex encoder-decoder structure, it understands the acoustic features and transforms them into corresponding text sequences. The model has already learned from massive amounts of audio-text pairing data during the pre-training phase, giving it strong generalization ability and adaptability to different accents and speech rates. In this system, Faster-Whisper is responsible for accurately and efficiently converting the doctor-patient dialogue into text records, providing high-quality text input for the subsequent intelligent medical record generation module. Its output text not only contains accurate speech content but also usually includes precise timestamp information, which helps in fine-grained text segmentation and contextual analysis in subsequent processing.
[0047] 3.2.2 Medical Terminology Injection Mechanism: Ensuring the Accuracy of Technical Terminology Recognition
[0048] While Faster-Whisper possesses powerful general recognition capabilities, in the highly specialized medical field, relying solely on a general model may not fully meet the extreme accuracy requirements for recognizing specific terms. Therefore, this embodiment innovatively introduces a medical vocabulary injection mechanism as a domain-specific enhancement to the speech-to-text process.
[0049] Implementation Principle and Optimization Effect: The medical vocabulary injection mechanism is a domain-adaptive technique designed to improve the accuracy of speech recognition for specific medical terms by prioritizing matching or reinforcement learning. Its specific implementation methods may include:
[0050] Runtime dictionary dynamic enhancement: During the speech decoding process (converting acoustic features into text) in the Faster-Whisper model, a pre-built high-priority dictionary containing a large number of medical professional terms is dynamically loaded. This dictionary encompasses disease names (e.g., "pancreatitis," "coronary atherosclerosis"), drug names (e.g., "amoxicillin," "nitroglycerin"), surgical names (e.g., "appendectomy," "cardiovascular bypass surgery"), anatomical locations, examination items (e.g., "magnetic resonance imaging," "complete blood count"), and various medical abbreviations and colloquialisms. When the speech recognition model outputs candidate word sequences, if it encounters a segment with a similar pronunciation to a word in the dictionary, the system will tend to prioritize matching and output the precise medical term from the dictionary, thereby effectively correcting errors that may be caused by general word recognition (e.g., avoiding misidentifying "pancreatitis" as "hospital banquet").
[0051] Language Model Fusion and Fine-tuning: Deeper injection can be achieved by weighted fusion of a domain-specific language model trained on a large corpus of medical texts (e.g., an N-gram model built on medical texts or a more complex neural network language model) with the base language model of Faster-Whisper, or by directly performing incremental fine-tuning on parts of the Faster-Whisper language model. This fusion or fine-tuning allows the model to have a stronger awareness and preference for the contextual probability distribution of medical terminology when generating text, thereby improving the accuracy of recognizing these terms in complex contexts.
[0052] 3.3 Medical Record Generation Layer: A Smart Elevation from Unstructured Text to Intelligent Structured Medical Records
[0053] The medical record generation layer is the key to realizing the core value of this invention. It goes beyond simple text recognition and conversion, using deep semantic understanding and intelligent reasoning to transform unstructured doctor-patient dialogue text into electronic medical records that conform to medical standards, are logically clear, and have a complete structure. This process is a concentrated manifestation of the system's level of intelligence.
[0054] 3.3.1 deepseek-1 module: Deep semantic understanding and intelligent content generation
[0055] This embodiment uses the DeepSeek-1 model as the intelligent core of the medical record generation module. DeepSeek-1 is a large-scale language model (LLM) pre-trained on massive amounts of general and medical professional text data, possessing excellent natural language understanding (NLU) and natural language generation (NLG) capabilities. In the medical record generation process, the DeepSeek-1 model receives accurately recognized doctor-patient dialogue text output from the speech-to-text layer as its core input. Its internal processing flow is highly complex and intelligent, mainly including:
[0056] Global Semantic Analysis and Intent Recognition: Deepseek-1 first performs comprehensive semantic analysis on the entire doctor-patient dialogue text to understand the overall context of the dialogue and the expressive intentions of the doctor and patient. It can distinguish which parts of the dialogue are the patient's chief complaint, which are descriptions of the present medical history, which are inquiries about past medical history from the doctor, and which are discussions of physical examination findings or auxiliary examination results, etc.
[0057] High-precision Named Entity Recognition (NER): The model can accurately identify and classify all medically relevant entities mentioned in a conversation. For example, it can accurately identify and label various disease names (such as "chronic obstructive pulmonary disease" and "stroke"), symptoms (such as "dizziness", "nausea", and "fatigue"), signs (such as "arrhythmia" and "jaundice"), medications (such as "aspirin" and "metoprolol" and their dosages), examination items (such as "complete blood count", "liver function", and "gastroscopy"), anatomical locations (such as "left lower lobe of the lung" and "duodenum"), and surgical procedures (such as "cholecystectomy").
[0058] Complex Relation Extraction (RE): Deepseek-1 not only identifies entities but also extracts the semantic relationships between them. For example, it can identify that "dizziness" is a "symptom" of "hypertension," "amoxicillin" is a "drug" used to "treat" "bacterial infection," or "chest CT" is used to "diagnose" "lung infection." This relation extraction capability enables the model to construct an internal knowledge graph of the dialogue.
[0059] Medical Event Extraction (EE): The model can identify and structure medical events, such as "admission", "discharge", "surgery", "confirmed diagnosis", etc., and extract related parameters such as time, location, participants and event outcome. For example, it can identify "the patient was admitted on March 15, 2023, diagnosed with acute appendicitis and underwent appendectomy on the same day".
[0060] Structured Information Mapping and Filling: Deepseek-1 intelligently maps and fills unstructured and semi-structured information extracted from the dialogue into corresponding fields in predefined electronic medical record templates (such as inpatient medical record templates, outpatient medical record templates, surgical record templates, etc.). For example, the patient's description of "having felt chest tightness and shortness of breath for the past month" is mapped to the "Present Illness" field; the doctor's diagnosis of "acute myocardial infarction" is mapped to the "Preliminary Diagnosis" field.
[0061] Standardized Natural Language Generation: Based on the pre-filled structured information and context, as well as an understanding of medical expertise, deepseek-1 generates the final medical record text in fluent, standardized natural language that conforms to medical terminology conventions. This may involve the integration, summarization, logical reorganization, and standardized expression of medical terminology.
[0062] The power of the deepseek-1 model lies in its ability to perform complex medical logic reasoning and context-aware content generation, ensuring that the generated medical records are both accurate and in line with clinical standards, significantly improving the automation level and quality of medical record writing.
[0063] 3.3.2 FAISS Knowledge Base and Similar Case Retrieval: Enhancing the Professionalism and Accuracy of Medical Records
[0064] To further enhance the professional depth of medical record generation, ensure the rigor of content, and effectively avoid the "illusion" that large language models may produce (i.e., generating seemingly reasonable but actually inaccurate or illogical information), this embodiment creatively integrates the FAISS knowledge base and supports similar case retrieval. This is equivalent to equipping the deepseek-1 model with a dynamic, real-time queryable "medical encyclopedia" and "clinical experience base."
[0065] Construction and Content of the FAISS Knowledge Base: FAISS (Facebook AI Similarity Search) is a highly optimized C++ library specifically designed for large-scale, efficient similarity searches in high-dimensional vector spaces, particularly suitable for handling datasets with billions or even more vectors. This invention fully leverages this advantage of FAISS to construct a vast, continuously updated medical knowledge base. The data types stored in this knowledge base are extremely rich and authoritative.
[0066] Standardized medical knowledge includes, but is not limited to, the International Classification of Diseases (ICD-10, ICD-11), surgical procedure codes (ICD-9-CM, ICD-10-PCS), national drug catalogs, drug instructions, clinical practice guidelines, expert consensus, diagnostic criteria (such as diagnostic criteria for diabetes and hypertension), treatment guidelines, and various medical terminology dictionaries.
[0067] High-quality historical clinical medical record data: Hospital internal electronic medical records that have undergone rigorous anonymization, desensitization, and review. These historical cases, serving as real-world clinical experience and examples, provide valuable context and reference for the model.
[0068] Authoritative medical literature abstracts and clinical research findings: the latest research abstracts, clinical trial results, and expert interpretations from well-known medical literature databases.
[0069] Vectorization Techniques: Before being added to FAISS, all text content in the knowledge base is transformed into high-dimensional vector embeddings using high-performance pre-trained language models in the medical domain (e.g., BERT, BioBERT, ClinicalBERT, or deepseek-1 itself, specifically fine-tuned for medical text). These vectors accurately capture the semantic information of the text, bringing semantically similar medical concepts or medical record fragments closer together. FAISS then uses these high-dimensional vectors for efficient indexing (e.g., through inverted indexes, IVF-FLAT, HNSW algorithms) to support similarity queries on large datasets within milliseconds.
[0070] Similar Case Retrieval Process and Knowledge Enhancement Mechanism (RAG): When the deepseek-1 model generates actual medical records, makes diagnostic inferences, or makes treatment recommendations, the system will activate the similar case retrieval function when the model needs to verify certain information, supplement details, or seek reference at a specific decision point.
[0071] Query vector generation: Based on the context of the current doctor-patient dialogue, identified patient symptoms, preliminary diagnosis, or generated medical record fragments, the system generates one or more query vectors. These query vectors represent the semantic information for which the knowledge needs to be acquired.
[0072] FAISS High-Speed Retrieval: These query vectors are sent to the FAISS knowledge base. Based on a preset similarity metric (most commonly cosine similarity), FAISS efficiently identifies the Top-K medical knowledge entries or historical cases that are semantically closest to the query vectors throughout the entire knowledge base.
[0073] Knowledge feedback and RAG (Retrieval-Augmented Generation): Retrieved relevant knowledge (e.g., typical diagnostic criteria for a specific disease, indications and contraindications for a certain treatment plan, how doctors describe specific symptoms or make differential diagnoses in similar cases) is extracted and dynamically added to the input of the deepseek-1 model in a structured or unstructured form as additional "chain of evidence" or "reference context".
[0074] This mechanism, which deeply integrates external knowledge with LLM generation capabilities, known as "Retrieval-Augmented Generation" (RAG), significantly improves the performance of LLM in the medical field.
[0075] Fact Checking and Correction: When generating medical record content, LLM can refer to authoritative knowledge retrieved in real time to verify the accuracy and consistency of the generated diagnoses, treatment plans, or medical descriptions. If inconsistencies are found, LLM can self-correct and adjust based on retrieved external "evidence," thereby significantly reducing the risk of the model generating "illusions" or inaccurate information.
[0076] Information Supplementation and Improvement: The retrieved knowledge can be used to supplement key information that may have been missed during the LLM generation process. For example, if the medical record only mentions "chest pain," similar cases retrieved may suggest "myocardial ischemia needs to be ruled out," thereby guiding the LLM to generate more comprehensive examination recommendations.
[0077] Through the aforementioned mechanism of deep integration with the FAISS knowledge base, this embodiment ensures that the deepseek-1 model can "draw on the strengths of various sources" when generating medical records, always making decisions and expressing information based on authoritative, accurate, and up-to-date medical knowledge, thereby outputting higher quality, more professional, and more reliable electronic medical record content, significantly improving the clinical value and accuracy of medical records.
[0078] 3.4 Output Interaction Layer: Portal for System Integration and Data Flow
[0079] The output interaction layer serves as the portal for data exchange and functional integration between the system and the external world. It is responsible for outputting the generated electronic medical record content in a standardized format and seamlessly integrating it with the hospital's existing information systems, thereby achieving interconnectivity and efficient flow of medical data.
[0080] 3.4.1 JSON Output: Standardized Data Encapsulation and Interoperability
[0081] In this embodiment, the electronic medical record content, which contains rich medical information and is generated by the medical record generation layer, is standardized and encapsulated through the JSON output module.
[0082] Technical Details and Data Structure: JSON (JavaScript Object Notation) is a lightweight data-interchange format. Due to its simplicity, readability, and ease of machine parsing, it has become the preferred format for data transmission between modern web applications and information systems. The generated electronic medical record data, including but not limited to patient basic information (name, gender, age, ID), consultation information (consultation date, department, doctor), and the main content of the medical record (chief complaint, present illness, past medical history, personal history, family history, physical examination results, auxiliary examination results and reports, preliminary diagnosis, final diagnosis, treatment plan, medical orders, and treatment opinions), will all be strictly organized into objects or arrays conforming to the JSON specification. This structured JSON output ensures high data interoperability, allowing it to be easily received, stored, analyzed, and further processed by any external system that supports JSON parsing, greatly facilitating the cross-system flow and utilization of data.
[0083] 3.4.2 HIS Integration: Seamless Integration into the Hospital's Core Information System
[0084] One of the key advantages of this embodiment is that it provides an interface for interfacing with the hospital's existing HIS (Hospital Information System), enabling close integration with the hospital's core business processes.
[0085] Technical Details and Integration Model: The goal of HIS integration is to enable voice-generated medical record data to seamlessly and in real-time enter into the hospital's official medical record system, avoiding data silos and duplicate entries. This integration typically follows healthcare IT industry standards and best practices.
[0086] Typical docking process:
[0087] Patient identity synchronization: When a doctor selects a patient for treatment in the HIS, the HIS will send the patient's basic identity information and treatment context to the voice electronic medical record core system through an interface.
[0088] Voice medical record upload: After the doctor completes the voice input and generates the medical record through the core system, the confirmed medical record data (JSON format or HL7 message format) will be pushed to the corresponding medical record module in the HIS system in real time through the HIS interface.
[0089] Data mapping and verification: After receiving the data, the HIS system performs internal data mapping and verification to ensure that the data format meets its internal requirements and is accurately populated into the patient's electronic medical record.
[0090] Medical record updates and archiving: The HIS system saves the received medical record data and completes the updates, version control, and final archiving of the medical records.
[0091] This deep integration eliminates the cumbersome steps and high error rate of traditional manual medical record entry, enabling real-time updates and seamless data transfer. It not only significantly improves medical record writing efficiency and reduces operating costs, but also ensures the real-time nature, integrity, and consistency of medical record data, thereby strengthening medical quality control and patient safety management.
[0092] Example 2: Knowledge-Enhanced Large-Scale Language Model Training Method (Based on...) Figure 2 )
[0093] This invention provides a unique implementation of a knowledge-enhanced large language model (LLM) training method, aiming to transform a general-purpose LLM into an intelligent assistant with superior medical expertise and reliability. For example... Figure 2 As shown, this training method deeply integrates high-quality domain data processing, multi-stage model training, and dynamic integration of external knowledge bases, fundamentally improving the performance of LLM in medical text understanding, generation, and reasoning.
[0094] 4.1 Medical Record Data Preprocessing and Labeling: The Cornerstone of Professional Knowledge
[0095] High-quality training data is a prerequisite for training an excellent LLM. This embodiment invests a great deal of effort in this stage to ensure that the medical record data input into the model is not only large in scale, but also of high quality and with a clear structure.
[0096] 4.1.1 Medical Record Data Collection and Standardization: Building a Massive Database of Original Resources
[0097] This embodiment begins with a large-scale collection of medical record data. This data comes from real clinical practice and covers electronic medical record texts from various departments, different disease types, and different patient groups.
[0098] Data Sources and Types: Medical record data may include:
[0099] Structured medical records: such as admission records, discharge summaries, surgical records, and examination reports (such as the conclusions of blood routine, biochemistry, and imaging reports).
[0100] Semi-structured medical records: such as outpatient medical records, ward round records, and nursing records, which typically contain some free text and some structured fields.
[0101] Unstructured text: Doctor's handwritten notes (recognized by OCR), detailed descriptions of patients' complaints, etc.
[0102] Data Scale and Diversity: To ensure the generalization ability and professional depth of LLM, the collected medical record data should reach the level of billions or even hundreds of billions of words, and should cover patients of different ages, genders, regions, and disease spectrums as much as possible, in order to reduce model bias and improve the understanding of complex cases. During data collection, strict compliance with privacy regulations (such as GDPR and HIPAA) is required, and sufficient de-identification processing must be carried out to ensure the security of patient information.
[0103] 4.1.2 Cleaning and Labeling: Refined Processing to Enhance Data Value
[0104] The collected raw medical record data is not directly usable; it must undergo a rigorous cleaning and annotation process, which is a key step in transforming the raw data into learnable knowledge for the model.
[0105] Cleaning steps:
[0106] Deduplication and noise reduction: Identify and remove duplicate medical record entries, and clean up garbled text, advertising information, non-medical related content, etc.
[0107] Format unification and standardization: Unify medical record texts from different sources and in different formats into a standard format, such as unifying encoding and date formats.
[0108] Typo and grammar correction: Use language models or rule bases to perform spell checking and grammar correction on text to ensure accuracy.
[0109] De-identification: This is the most crucial step, involving the strict removal or replacement of all patient-identifying information (such as name, ID number, contact information, and precise address) to ensure data compliance with privacy regulations and protect patient privacy. Common de-identification techniques include entity replacement, hash encryption, and obfuscation.
[0110] Annotation Steps: Professional annotation is performed on the cleaned data to provide signals for supervised learning in LLM. This is typically done by annotators with medical backgrounds, supplemented by automated tools and quality verification mechanisms. Annotation types include, but are not limited to:
[0111] Named Entity Recognition (NER): Identifies and classifies medical entities in text, such as:
[0112] Diseases and diagnoses: hypertension, diabetes, acute myocardial infarction, pneumonia, stomach cancer.
[0113] Symptoms and signs: chest pain, cough, fever, palpitations, fatigue, and moist rales in both lungs.
[0114] Medications: Aspirin, penicillin, insulin, acetaminophen.
[0115] Examinations and tests: complete blood count, X-ray, CT scan, MRI, blood glucose, creatinine.
[0116] Surgery and procedures: coronary artery bypass grafting, appendectomy, gastroscopy.
[0117] Anatomical sites: heart, lungs, stomach, liver, spine.
[0118] Time and values: time of disease onset, drug dosage, and test results.
[0119] Relation Extraction (RE): Identifying semantic relationships between medical entities, for example:
[0120] Chest pain is a symptom of myocardial infarction.
[0121] Event Extraction (EE): Identifying medical events and their participants, time, location, and other information, for example:
[0122] "Hospitalization" incident: Patient XXX was "admitted" at time YYY, with the chief complaint being ZZZ.
[0123] Text classification and summary annotation: Classify medical record texts (such as diagnostic medical records, surgical medical records), or annotate key information for the generation of summaries.
[0124] These refined annotations not only provide high-quality labels for supervised learning, but also help LLM understand the complex semantics, causal relationships and clinical logic in medical texts, thereby enabling better reasoning and generation.
[0125] 4.2 Pre-training and Fine-tuning in the Medical Field: Shaping the Professional Wisdom of LLMs
[0126] LLM training is a multi-stage process designed to move individuals from general language comprehension to in-depth medical expertise.
[0127] 4.2.1 Pre-training: Laying the foundation for a common language
[0128] Technical details: Pre-training is the stage in which LLM learns general language knowledge. Mainstream pre-training tasks include:
[0129] Next Sentence Prediction (NSP): Determines whether two text segments are consecutive in the original document, which helps the model understand semantic coherence at the document level.
[0130] Causal Language Model (CLM): Similar to autoregressive models, it predicts the next word based on the preceding words. This gives the model powerful text generation capabilities.
[0131] Through pre-training, LLM learns a rich vocabulary, grammar, semantics, common sense, and general reasoning abilities, building a powerful language foundation model that provides a solid starting point for subsequent domain-specific fine-tuning.
[0132] 4.3 FAISS Knowledge Enhancement and Similar Case Retrieval: Injecting Dynamic External Intelligence
[0133] To address the potential "illusion" problem of LLM (i.e., generating seemingly reasonable but actually erroneous information) and to improve its responsiveness to the latest medical knowledge, this embodiment introduces a dynamic external knowledge enhancement mechanism.
[0134] 4.3.1 Construction and Maintenance of the FAISS Knowledge Base: An External Medical Brain
[0135] This embodiment constructs a medical knowledge base based on FAISS (Facebook AI Similarity Search). FAISS is a highly optimized C++ library for efficient similarity searching in vector spaces, especially suitable for high-dimensional vectors with billions or more elements.
[0136] Knowledge Base Content and Vectorization: This knowledge base not only contains static medical textbook knowledge, but also a constantly updated dynamic knowledge system, including:
[0137] Vectorization Method: Before being added to the knowledge base, all text content is transformed into high-dimensional vector embeddings using high-performance medical domain pre-trained language models (such as BioBERT, ClinicalBERT, or a medically fine-tuned Deepseek-1). These vectors capture the semantic information of the text, making semantically similar texts closer together in the vector space. FAISS then efficiently indexes these vectors (e.g., using inverted indexes, IVF-FLAT, HNSW) to support millisecond-level similarity queries.
[0138] 4.3.2 Similar Case Retrieval Process and Knowledge Enhancement Mechanism (RAG): "Smart Reference" for LLM
[0139] In LLM practice of generating medical records, answering medical questions, or making clinical inferences, the FAISS knowledge base plays the role of "intelligent reference".
[0140] Search process:
[0141] Contextual query generation: When the LLM is at a certain stage of generating medical records (e.g., when a diagnosis needs to be made, a treatment plan needs to be recommended, or a specific symptom needs to be described), or when a user asks a medical question, the system extracts key information from the current LLM input (such as identified doctor-patient dialogue text, part of the generated medical record content) and transforms it into a query vector.
[0142] FAISS High-Speed Retrieval: This query vector is sent to the FAISS knowledge base. Based on a preset similarity metric (most commonly cosine similarity), FAISS efficiently finds the Top-K medical knowledge entries or historical cases that are semantically closest to the query vector throughout the entire knowledge base.
[0143] Fact Checking and Correction: When generating medical record content, LLM can refer to authoritative knowledge retrieved in real time to verify the accuracy and consistency of the generated diagnoses, treatment plans, or medical descriptions. If inconsistencies are found, LLM can self-correct and adjust based on retrieved external "evidence," thereby significantly reducing the risk of the model generating "illusions" or inaccurate information.
[0144] Information Supplementation and Improvement: The retrieved knowledge can be used to supplement key information that may have been missed during the LLM generation process. For example, if the medical record only mentions "chest pain," similar cases retrieved may suggest "myocardial ischemia needs to be ruled out," thereby guiding the LLM to generate more comprehensive examination recommendations.
[0145] Enhancing professionalism and authority: By incorporating external knowledge, LLM can generate more professional medical records that align with the latest clinical practice guidelines and standards. For example, when describing the treatment of a disease, the model can directly reference and integrate the latest guideline recommendations, rather than relying solely on its internal memory.
[0146] Handling new knowledge and timeliness: Even if the internal parameters of the LLM are not updated with the latest training data, as long as the FAISS knowledge base can be updated with the latest medical research or guidelines in a timely manner, the LLM can dynamically acquire and utilize this new knowledge through the retrieval mechanism, thereby maintaining the timeliness and cutting-edge nature of its knowledge.
[0147] Enhancing traceability and interpretability: The retrieval enhancement mechanism can also provide "evidence sources" for the content generated by LLM. That is, it explicitly indicates which specific knowledge points or historical case studies the model used to generate the current medical record content, which greatly increases the credibility, transparency, and interpretability of the system.
[0148] Through the aforementioned mechanism of deep integration with the FAISS knowledge base, this embodiment ensures that the deepseek-1 model can "draw on the strengths of various sources" when generating medical records, always making decisions and expressing information based on authoritative, accurate, and up-to-date medical knowledge, thereby outputting higher quality, more professional, and more reliable electronic medical record content, significantly improving the clinical value and accuracy of medical records.
[0149] Example 3: Integration of Medical Voice Electronic Medical Record System (Based on) Figure 3 )
[0150] This invention provides an integrated embodiment of an advanced medical voice electronic medical record system. Its core lies in achieving seamless integration between the core voice electronic medical record system and the hospital's existing IT infrastructure, thereby fully integrating intelligent medical record generation capabilities into clinical workflows and improving the hospital's overall informatization level and operational efficiency. Figure 3 As shown, this integrated solution embodies the design concept of modularity and interconnectivity.
[0151] 5.1 Voice-activated Electronic Medical Record Core System: Intelligent Medical Information Hub
[0152] Technical details and functions:
[0153] Multi-source input processing: This core system can receive voice data from different front-end devices (such as microphone arrays and mobile apps), as well as possible text input. It has built-in voice processing modules (such as speech recognition and noise reduction) and natural language processing modules (such as semantic understanding, entity recognition, and relation extraction) to intelligently analyze the input data.
[0154] Intelligent Medical Record Generation Engine: The core system integrates the knowledge-enhanced LLM described in Example 2, which can intelligently generate structured and standardized electronic medical records based on the parsed doctor-patient dialogue content and a medical knowledge base. This includes the automatic filling of all medical record fields such as chief complaint, present illness, past medical history, physical examination, auxiliary examinations, diagnosis, treatment plan, and medical orders.
[0155] Medical record data management: The core system can maintain a temporary or cached medical record database to store medical record data that is being generated or awaiting review, and to perform version control.
[0156] Interface Service Layer: The core system exposes standardized API (Application Programming Interface) service interfaces, allowing other systems to submit, query, update, and delete data by calling these interfaces. These interfaces adhere to the principles of security, stability, and high performance.
[0157] User Access and Security Management: The core system has comprehensive user authentication and authorization management functions, ensuring that only authorized medical personnel can access and manipulate medical record data. Furthermore, all data transmission and storage employ encryption technology, complying with medical data security regulations (such as HIPAA and GDPR).
[0158] Log and audit functions: Record all system operations and data flow in detail for subsequent auditing and tracking.
[0159] 5.2 Front-end system integration: A convenient tool to empower medical staff
[0160] Front-end system integration is crucial to ensuring that medical staff can use the voice-based electronic medical record function efficiently and conveniently. This embodiment supports embedding core system capabilities into different terminals used daily by medical staff.
[0161] 5.2.1 HIS Terminal Integration: Seamlessly Integrating into Doctors' Workflow
[0162] The HIS (Hospital Information System) terminal access aims to allow doctors to directly utilize the RPA automatic data entry function without switching software or changing their original operating habits.
[0163] Technical details and interaction process:
[0164] Integration Modes: Two main integration modes are typically used:
[0165] Deeply embedded: Functional modules of the core voice electronic medical record system (such as voice input buttons and medical record preview panes) are directly embedded as components into the HIS client application. This may require SDK or custom development support from the HIS vendor.
[0166] External call method: HIS sends the audio stream or text information to the core system by calling the Web service API (such as RESTful API or SOAP service) exposed by the core system when the doctor needs voice input, and then receives and displays the structured medical record data returned by the core system.
[0167] Operating procedures:
[0168] The doctor selects the patients to be entered on the front end.
[0169] Click the "Voice Input" button in the HIS interface to start voice recording.
[0170] Voice data is transmitted in real time to the "Voice Electronic Medical Record Core System" for processing (voice recognition and medical record generation).
[0171] The core system instantly returns the generated structured medical record content to the front end. Doctors then choose to upload the completed structured medical record to the HIS (Hospital Information System).
[0172] The medical record information is automatically populated into the corresponding fields in the HIS (Hospital Information System). Doctors can preview, modify, and finally confirm the information in the HIS interface.
[0173] After the doctor confirms the information, the medical record data is saved to the hospital's database through the HIS's own process.
[0174] This integration approach ensures that healthcare workers can improve their work efficiency directly in a familiar environment without having to learn new and complex systems.
[0175] 5.2.2 Mobile Access: Enables medical records to be kept anytime, anywhere.
[0176] Considering the mobile needs of medical staff in scenarios such as ward rounds, emergency rooms, and consultations, this embodiment also provides mobile terminal access functionality.
[0177] Technical details and interaction process:
[0178] Mobile App Development: Develop mobile mini-programs that integrate a voice input interface, a medical record preview interface, and a client SDK for communication with the core system.
[0179] API Communication: Data interaction with the "Voice Electronic Medical Record Core System" is conducted via secure APIs (such as HTTPS-based RESTful APIs). All data transmissions are encrypted to ensure information security.
[0180] Operating procedures:
[0181] Healthcare workers log in to the app on their mobile devices, select a patient, and access the medical record interface.
[0182] Click the voice input button in the app to start dictating your medical records. The voice data is transmitted to the core system via the mobile network.
[0183] The core system processes voice messages in real time and generates medical records.
[0184] The generated medical records are displayed in real time on the mobile app interface for medical staff to review and modify.
[0185] After confirmation by medical staff, the medical record data is uploaded to the core system via the App, and then synchronized to EMR / HIS by the core system.
[0186] Mobile access greatly enhances the work flexibility of medical staff, freeing medical record recording from the constraints of fixed workstations and facilitating more efficient bedside diagnosis and treatment and information recording.
[0187] 5.3 Backend System Integration: Building an Integrated Healthcare Information Ecosystem
[0188] Backend system integration is key to achieving full interconnection and interoperability among various information systems within the hospital, ensuring that the data generated by the core voice electronic medical record system can be organically integrated into the hospital's overall information flow, forming a collaborative and information-sharing smart healthcare ecosystem.
[0189] 5.3.1 EMR System Integration: The Cornerstone of Unified Patient Health Records
[0190] EMR (Electronic Medical Record) system integration aims to seamlessly integrate structured medical record data generated by the voice electronic medical record core system into the hospital's core electronic medical record storage system, thereby building a comprehensive and unified patient health record.
[0191] Technical details and data flow:
[0192] Data push mechanism: Once the voice electronic medical record core system generates medical record data and it is confirmed by the doctor, it will proactively push the data to the EMR system through a preset, secure interface (such as ADT messages based on the HL7 standard, CDA documents, or SOAP / RESTful API). This can be done in real-time synchronization mode (pushing data as soon as the medical record is generated) or batch synchronization mode (pushing data in batches periodically) depending on the actual needs of the hospital.
[0193] Data Mapping and Transformation: The data structure generated by the core system may differ from the internal storage structure of the EMR system. Therefore, during the data push process, precise data field mapping and format conversion are performed to ensure that the data from the core system can be accurately populated into the corresponding fields and tables in the EMR system. For example, the "chiefComplaint" field in the core system is mapped to the "chief complaint" text area in the EMR, and structured diagnostic codes are mapped to the diagnostic list in the EMR.
[0194] Unified View Presentation: The ultimate goal of the integration is that, regardless of whether the medical record data was initially generated by voice, entered manually, or imported from other systems, it can be presented to medical staff in a unified, complete, and continuously updated view within the EMR system. This forms a comprehensive and longitudinal electronic health record for patients from admission to discharge, and from outpatient to inpatient care, greatly facilitating clinical decision-making and long-term health management.
[0195] 5.3.2 Medical Equipment Integration: Automated Data Capture and Intelligent Medical Record Filling
[0196] This embodiment further supports deep integration with various medical devices, aiming to achieve automated capture of device-generated data and intelligent integration into electronic medical records, thereby reducing errors from manual entry and improving data timeliness.
[0197] Technical details and data flow:
[0198] Extensive Device Data Sources: This integration solution can cover a wide range of medical device data sources, including but not limited to:
[0199] Medical imaging equipment, such as CT scanners, MRI machines, X-ray machines, and ultrasound machines. These devices typically generate imaging reports after completing the examination, which include important diagnostic conclusions (such as "Chest CT shows: nodular shadow in the upper lobe of the right lung, suggestive of infectious lesion").
[0200] Clinical laboratory equipment includes fully automated biochemical analyzers, complete blood count (CBC) analyzers, urinalysis analyzers, and microbial culture systems. These generate large amounts of structured numerical data and qualitative results.
[0201] Physiological parameter monitoring equipment includes devices such as electrocardiographs, multi-parameter monitors, ambulatory blood pressure monitors, and pulse oximeters. These devices generate real-time data on physiological parameters such as heart rate, blood pressure, body temperature, and blood oxygen saturation.
[0202] Diverse integration interfaces and protocols: To enable effective communication with different types and brands of medical devices, this system employs a variety of standardized integration interfaces and protocols:
[0203] HL7: Most clinical laboratory and physiological monitoring devices transmit test results or observations via HL7 standard messages (such as ORU^R01-Observation Report messages). This system can receive and parse these HL7 messages.
[0204] Custom API / SDK: Some specific or new devices may provide proprietary APIs or SDKs for data export. This system can be customized and adapted as needed.
[0205] Middleware or integration engine: Deploy specialized medical device integration middleware (such as MirthConnect) as a bridge for data transformation and routing. This system can integrate with these middleware to indirectly obtain device data.
[0206] Data Analysis and Intelligent Integration: After receiving data from medical devices, the core system of the voice-based electronic medical record can intelligently identify the type and format of the data (e.g., whether it is a text report, numerical result, or waveform data) and accurately extract key medical information. For example, it can extract specific diagnostic conclusions and descriptive text from imaging reports, and extract the names, values, and units of various indicators from test results, and determine their normal ranges.
[0207] Automated medical record filling: The extracted equipment data is intelligently mapped and automatically filled into the corresponding predefined fields in the patient's electronic medical record (for example, the blood routine test results from the laboratory will be automatically filled into the "Blood Routine" sub-item under "Auxiliary Examinations"; the CT diagnosis conclusion from the radiology department will be automatically filled into the "Imaging Examinations" item). This automated filling greatly improves the efficiency of medical record writing and significantly reduces the risk of errors from manual entry, while ensuring the accuracy and completeness of the medical record content.
[0208] Data consistency and traceability: Ensure consistency between device data and medical record content, and support direct tracing from medical records to the original device data source to meet the needs of clinical data auditing and quality control.
[0209] Through the aforementioned comprehensive and multi-layered system integration solution, this invention constructs a highly intelligent and interconnected smart medical information network. This network enables the automated generation of structured medical records from medical voice data, seamless collaboration with hospital HIS / EMR systems, and intelligent linkage with various medical devices. This not only significantly improves the work efficiency of medical staff and optimizes medical processes, but more importantly, by building a comprehensive, accurate, real-time, and data-sharing electronic medical record system, it lays a solid data foundation for refined hospital management, clinical decision support, and medical research, ultimately providing patients with higher quality, safer, and more convenient medical services.
Claims
1. A fully automated voice-driven electronic medical record generation system, characterized in that, include: A medical-grade speech acquisition and preprocessing subsystem is used to acquire speech data through a microphone array and perform noise suppression and speech segmentation processing. A medical terminology-enhanced speech recognition engine is used to convert the processed speech data into text and accurately recognize medical terms. A knowledge-enhanced medical record generation system for intelligently generating structured electronic medical records from the converted text; and a multimodal interaction and system integration subsystem for outputting the structured electronic medical records and seamlessly integrating them into the hospital information system.
2. The system according to claim 1, characterized in that, The medical-grade speech acquisition and preprocessing subsystem includes a medical-grade microphone array configured for far-field sound pickup and noise suppression exceeding 30dB. It is further configured to perform speech activation based on the GMM-UBM voiceprint recognition model and keyword wake-up technology. The subsystem also employs the WebRTCNS algorithm for environmental noise suppression and deep learning-based speech activity detection (VAD) technology to achieve speech segmentation with a minimum segment length of 300ms.
3. The system according to claim 1, characterized in that, The medical terminology-enhanced speech recognition engine is based on the Faster-Whisper-large-v3 architecture and is configured to enhance the recognition of professional terms through a comprehensive medical-specific dictionary. The engine also integrates a BiLSTM+CRF model for terminology consistency checks and uses PyAnnote to combine multi-dimensional features to achieve accurate separation and annotation of doctor and patient roles.
4. The system according to claim 1, characterized in that, The knowledge-enhanced medical record generation system includes a large language model based on the deepseek-r1 architecture. The model undergoes three stages of training: pre-training on a general corpus, fine-tuning on a medical corpus, and optimization through reinforcement learning. The system also includes a knowledge retrieval module based on the FAISS vector database, which is used to retrieve similar historical cases and medical knowledge in real time to enhance medical record generation.
5. The system according to claim 4, characterized in that, The knowledge-enhanced medical record generation system further includes a rule validation engine configured to perform at least one of the following functions: drug conflict detection, diagnostic logic validation, and treatment plan rationality check.
6. The system according to claim 1, characterized in that, The multimodal interaction and system integration subsystem is configured to generate JSON-formatted structured medical records that conform to SOAP format and contain multiple standard fields, and to deeply interface with the hospital information system based on standard protocols such as HL7 / FHIR.
7. The system according to claim 1, characterized in that, It also includes a full-process security and compliance subsystem, which is based on an edge computing architecture and is used to process voice data on a local server, implement a hierarchical desensitization strategy and end-to-end encrypted transmission. The subsystem also follows the principle of least privilege access control and uses blockchain technology to provide tamper-proof evidence of key operations.
Citation Information
Patent Citations
Structured electronic medical record generating method based on voice input and corresponding storage device and mobile terminal
CN107564571A
Electronic medical record generation method and device based on artificial intelligence, equipment and medium
CN113724695A
A method and system for generating outpatient electronic medical records based on speech recognition
CN118072901B
Secure storage method and system for electronic medical records
CN120319380A
Cited By
Method, device, equipment and product for automatically transcribing voice into HIS (hospital information system) of outpatient medical record
CN121687358A
Disease condition reasoning and medical record generation method and system based on big model fusion knowledge enhancement
CN121862294A