Electronic medical record automatic generation method based on voice recognition

By assigning a unique identifier to each medical conversation and combining a multi-microphone array and a speech recognition engine, real-time monitoring and verification of voice signals, and dynamic ranking of speaker priorities, the problems of incomplete identification, confusing information association, and semantic consistency in the generation of electronic medical records are solved, achieving accurate generation and management of medical records and improving the quality of medical services.

CN120690205AInactive Publication Date: 2025-09-23THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT)
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510870896.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing electronic medical record generation technology based on speech recognition has problems such as incomplete speech input identification, confusing information association, insufficient semantic consistency detection, and insufficient multi-speaker processing capabilities, which lead to inaccurate medical record generation and information omissions, affecting the scientificity and efficiency of medical decision-making.

Method used

By assigning a unique voice collection identifier to each medical conversation, combining a multi-microphone array and a voice recognition engine, the spectral distribution and interference of the voice signal are monitored and verified in real time. Multi-layer distributed semantic analysis and rapid segmentation mechanisms are used to dynamically prioritize speakers and ensure that critical information is processed first.

Benefits of technology

It achieves comprehensive and accurate identification and voice conversion of medical conversations, ensures the accuracy and completeness of medical records, reduces errors, improves the intelligence level of automatic generation of electronic medical records, optimizes diagnosis and treatment processes, and reduces the risk of medical errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690205A_ABST
    Figure CN120690205A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electronic medical record generation, and discloses an electronic medical record automatic generation method based on voice recognition, which comprises the following steps: S1, initializing voice input, distributing a unique voice acquisition identifier for a medical session, and completing identifier generation, input, storage, association and identity verification; s2, voice information intelligent recognition: converting voice into a text by using a voice recognition engine, and ensuring semantic consistency through a context verification unit and a semantic analysis unit; and S3, performing multi-speaker processing based on intelligent interference detection and resolution, positioning an interference time period and an interference source through an interference detection unit, and realizing time period distribution and priority ranking of multi-speaker voices by using a voice segmentation protocol and a linear weighting model. And finally, extracting related information from the text generated by voice conversion, and filling the related information into a medical record template of a hospital. The method improves the efficiency and accuracy of electronic medical record generation, solves the problems of multi-speaker interference and semantic logic, and is suitable for medical informatization scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic medical record generation, and in particular to a method for automatically generating electronic medical records based on speech recognition. Background Art

[0002] Against the backdrop of the rapid development of medical informatization, the efficient generation and management of electronic medical records (EMRs), as the core carrier of medical information, is crucial for improving the quality of medical services and optimizing the diagnosis and treatment process. Traditional EMR generation relies primarily on manual data entry by medical staff, a method with numerous drawbacks. Firstly, manual data entry consumes considerable time and effort, forcing medical staff to distract themselves from recording information during consultations. This can lead to distractions, hindering communication with patients and their ability to assess their condition, thus reducing diagnosis and treatment efficiency. Secondly, manual recording inevitably results in incomplete and non-standardized information due to time constraints. This can lead to misunderstandings or misinterpretation of medical record information, compromising subsequent diagnosis, treatment, and archiving, querying, and statistical analysis of medical records. Furthermore, in multi-department collaboration or referral scenarios, non-standardized record keeping can lead to poor information transfer and increase the risk of medical errors.

[0003] With the advancement of speech recognition technology, its application in electronic medical record generation has become an important direction for addressing the aforementioned issues. However, existing speech recognition-based electronic medical record generation technologies still face a number of challenges. During the speech input process, effective identification and management of medical conversations is a key issue. Traditional identification methods may lack completeness and accuracy, failing to fully record key information such as patient identity, department visited, medical record type, and conversation time. Furthermore, confusion can easily arise when linking information with the voice capture device, compromising the accuracy and traceability of subsequent medical record generation.

[0004] Another prominent issue in speech recognition is the inadequate ability to detect semantic consistency and correct logical inconsistencies. The medical field involves a vast array of specialized terminology and complex logical relationships. The conversion of speech to text can be subject to errors due to factors such as accent, speaking speed, and ambient noise, leading to logical inconsistencies or semantic deviations in the text. Failure to promptly detect and correct these issues will result in inaccurate electronic medical records, a failure to accurately reflect the diagnosis and treatment process, and a consequent impact on the scientific nature of medical decision-making.

[0005] In multi-speaker scenarios, such as doctor-patient interactions and multi-medical consultations, speech signals are susceptible to interference, leading to recognition conflicts. Existing technologies lack effective multi-speaker processing mechanisms, making it difficult to accurately distinguish speech signals from different speakers and generate text based on speaker priority. This can lead to the omission of critical information or delayed processing, compromising the integrity and validity of medical records. For example, in emergency consultations, if expert opinions are not prioritized and extracted promptly, diagnosis and treatment planning may be delayed.

[0006] In addition, the existing systems have relatively simple technical means in semantic analysis and interference detection, and are unable to achieve multi-level and distributed analysis of text logic. It is also difficult to monitor and locate the spectral distribution and interference periods of voice signals in real time and accurately, resulting in insufficient voice processing capabilities in complex medical scenarios, limiting the application scope and effect of electronic medical record automatic generation technology. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for automatically generating electronic medical records based on speech recognition to solve the problems raised in the above background technology.

[0008] To achieve the above-mentioned object, the present invention provides the following technical solution: a method for automatically generating electronic medical records based on speech recognition, the method comprising:

[0009] S1, voice input initialization: assign a unique voice collection identifier to each medical session;

[0010] S2, intelligent voice recognition: The system is equipped with a voice recognition engine that automatically converts speech into text upon input. A context verification unit is integrated into the system's processing module to check the semantic consistency of the text in real time during the conversion process. If any logical contradiction is detected, the text is automatically corrected and marked as an anomaly.

[0011] S3, multi-speaker processing based on intelligent interference detection and elimination:

[0012] S31, embedding an interference detection unit based on acoustic features in the speech recognition engine to monitor the spectrum distribution of the speech signal and the interference occurrence period in real time;

[0013] S32: To address the issue of interference periods in speech signals, a speech segmentation protocol is introduced. The speech segmentation protocol pre-allocates speech segments through a fast segmentation mechanism and generates text based on speaker priority. The speaker priority is determined using a linear weighted model, and the priority of speakers in recognition conflicts is dynamically sorted to ensure that key information is processed first.

[0014] S4 extracts information from the text generated by speech conversion and fills it into the hospital's medical record template.

[0015] Preferably, the S1 specifically includes:

[0016] S11, identification generation and content entry: Input identification data including patient identity, treatment department, medical record type, session time and estimated duration through the electronic medical record management platform. After the voice acquisition device is started, the device information is added to the identification data to form a complete identification record;

[0017] S12, identification storage and association: using the voice collection device to store the identification data, generate a corresponding voice collection identification, and associate the generated voice collection identification with the medical conversation.

[0018] Preferably, S1 also includes identity verification: after the voice collection identifier is associated, the voice collection identifier content is read through the identity authentication module and verified and compared with the identification data stored in the electronic medical record management platform. The verification content includes the integrity, accuracy and uniqueness of the voice collection identifier; the successfully associated voice collection identifier is uploaded to the electronic medical record management platform to establish a long-term mapping relationship between the voice collection identifier and the medical session.

[0019] Preferably, the S2 specifically includes:

[0020] S21, installing a speech recognition engine at the input interface of the system, and the speech recognition engine distributes and covers the collection area through a multi-microphone array;

[0021] S22, when voice input begins, triggers the voice recognition engine to detect the voice signal in real time and automatically convert it into text data, including patient description, symptom details, diagnosis opinion, timestamp and key terms. After successful conversion, the current status of the medical session in the electronic medical record management platform is updated.

[0022] Preferably, the processing module of the system is configured with a multi-layer distributed semantic analysis unit, which collects the overall logic and segmented semantic distribution of the text in real time. When the voice is converted, the semantic analysis unit detects the semantic changes of the newly generated text and generates the actual logical data of the current text. The text logic output by the voice recognition engine is compared with the actual logic collected by the semantic analysis unit. If it is detected that the logic deviation exceeds the set threshold, the current text is automatically marked as abnormal, and the actual logical content is updated in the text data.

[0023] Preferably, the S31 specifically includes:

[0024] Acoustic perception and interference identification: The interference detection unit uses an embedded acoustic analysis chip to perceive the spectrum occupancy of voice signals in real time. The acoustic analysis chip uses a spectrum distribution analysis algorithm to detect the frequency distribution of signal transmission and mark voice communication activities with overlapping signals within the same frequency band.

[0025] Time period detection and positioning: When signal interference is detected, the specific time period where the interference occurs is located by calculating the strength, spectral characteristics, and time characteristics of the voice signal.

[0026] Preferably, the spectrum distribution analysis algorithm specifically includes:

[0027] Spectrum occupancy status calculation: By analyzing the spectrum energy distribution of the voice signal, the frequency occupancy status is determined. When the spectrum energy exceeds the set energy threshold, it indicates that the frequency is occupied.

[0028] The frequency occupancy status detection condition is as follows: the spectrum occupancy status is determined by the energy threshold. Energy above the threshold indicates occupancy, and energy below the threshold indicates idleness.

[0029] Preferably, the time period set for voice signal allocation is defined as a pre-allocated time period sequence, the time period sequence comprises a plurality of continuous time windows, in which each time window corresponds to a voice segment, and the detection condition for signal strength overlap is that the sum of the strengths of the plurality of voice signals within the time period exceeds a set conflict threshold;

[0030] Interference source localization: If signal strength overlap is detected, the interference source set is located by calculating the contribution signal strength of each voice source. Voice sources whose contribution signal strength exceeds the ratio threshold are included in the interference set.

[0031] Preferably, the S32 specifically includes:

[0032] S321, fast segmentation mechanism to implement time slot pre-allocation: After interference detection, the fast segmentation mechanism is used to dynamically adjust the time slot allocation of the voice signal, assigning a specific processing time slot to each voice source to ensure that there is no overlap between time slots;

[0033] S322, Dynamic Sorting of Speaker Priorities: Generates a multidimensional feature vector based on the strength of the speech signal, the distance between the speech source and the acquisition device, and the criticality level of the medical record content. Dynamically prioritizes the multidimensional feature vectors of the conflicting speech sources using a linear weighted model.

[0034] S323, priority-driven text generation: assigning a time sequence of processing periods according to the sorting results.

[0035] Preferably, the multidimensional feature vector is represented as a feature combination of the speech source, the feature combination including signal strength, inverse distance and content key level;

[0036] The linear weighted model calculates the comprehensive priority of the voice source and sorts it based on the comprehensive priority: the comprehensive priority is balanced by the feature weight coefficient, and the inverse of the distance indicates that the voice source closer to the acquisition device has higher priority;

[0037] The fast segmentation mechanism dynamically adjusts the time period allocation of the speech signal to be expressed as a time period start time sequence, where the time period start time is determined by the reading order of the speech source, and the reading order is allocated by the priority sorting result.

[0038] In addition, the existing systems have relatively simple technical means in semantic analysis and interference detection, and are unable to achieve multi-level and distributed analysis of text logic. It is also difficult to monitor and locate the spectral distribution and interference periods of voice signals in real time and accurately, resulting in insufficient voice processing capabilities in complex medical scenarios, limiting the application scope and effect of electronic medical record automatic generation technology.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] In terms of voice input management, comprehensive and accurate identification of medical sessions is achieved by assigning a unique voice collection identifier to each medical session and recording detailed identification data such as patient identity, treatment department, medical record type, session time, estimated duration, and device information. This identity verification mechanism ensures the integrity, accuracy, and uniqueness of identification data, establishes a long-term mapping relationship between voice collection identifiers and medical sessions, and provides reliable foundational data for subsequent medical record generation, facilitating traceability, management, and querying of medical records. This effectively addresses the issues of incomplete information and confusing associations inherent in traditional identification methods.

[0041] During speech recognition, the system's speech recognition engine, combined with a multi-microphone array covering the acquisition area, accurately converts speech into text data containing patient descriptions, symptom details, diagnostic opinions, timestamps, and key terms in real time. Through a contextual verification unit and multi-layer distributed semantic analysis units, the system detects semantic consistency and logical deviations in the text in real time. When a logical contradiction or deviation exceeding a set threshold is detected, the text is automatically corrected and marked as an anomaly, while the actual logical content is updated. This ensures the accuracy and semantic coherence of electronic medical record text, effectively reduces medical record information errors caused by speech conversion errors, and provides a reliable basis for medical diagnosis and treatment.

[0042] To address interference issues in multi-speaker scenarios, an acoustic feature-based interference detection unit is embedded in the speech recognition engine. A spectrum distribution analysis algorithm is used to monitor the spectrum occupancy of speech signals and the periods of interference in real time, enabling precise location of interference sources and periods. The introduced speech segmentation protocol pre-allocates speech segments through a rapid segmentation mechanism, generates multidimensional feature vectors based on signal strength, inverse distance, and content criticality, and dynamically prioritizes speakers using a linear weighted model, ensuring that critical information is prioritized. This multi-speaker processing mechanism effectively distinguishes speech signals from different speakers, avoiding recognition conflicts caused by signal interference. This ensures that in complex scenarios such as doctor-patient communication and multi-medical consultations, important information such as key diagnostic opinions and emergency treatment recommendations are prioritized for conversion and recording. This improves the integrity and effectiveness of medical records, and provides strong support for timely and accurate medical decision-making, especially in emergency medical scenarios.

[0043] Furthermore, by pre-assigning time slots and dynamically adjusting them, this method ensures that speech signal processing time slots do not overlap, improving the system's processing efficiency and stability. Overall, this method comprehensively enhances the intelligent level of automatic electronic medical record generation, reduces the workload of medical staff, optimizes the diagnosis and treatment process, and reduces the risk of medical errors. It has important practical significance for promoting the development of medical information technology and improving the quality of medical services. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a working principle diagram of the method for automatically generating electronic medical records based on speech recognition according to the present invention;

[0045] Figure 2 Generate and associate design drawings for voice collection identification;

[0046] Figure 3 Design diagram for interference detection and time period positioning;

[0047] Figure 4 Design diagram for dynamic speaker prioritization. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0049] See also Figure 1-Figure 4 The present invention relates to a method for automatically generating electronic medical records based on speech recognition. This method uses a collaborative mechanism of speech input initialization, intelligent speech information recognition, and multi-speaker processing to achieve automatic generation and logical verification of electronic medical records. The specific implementation steps are as follows:

[0050] S1. Voice input initialization: Assign a unique voice collection identifier to each medical session. Specifically, input identification data such as patient identity, treatment department, medical record type, session time, and expected duration through the electronic medical record management platform. After the voice collection device is started, the device information (such as device ID, collection parameters) is added to the identification data to form a complete identification record containing patient information, device information, and session time. The voice collection device is then used to store the identification data and generate a corresponding voice collection identifier, which is associated with the medical session to establish an initial mapping relationship.

[0051] S2. Intelligent recognition of voice information: The system is equipped with a voice recognition engine, which covers the collection area through a multi-microphone array to achieve all-round collection of voice signals. When the voice input begins, the voice recognition engine detects the voice signal in real time, automatically converts the patient description, symptom details, diagnostic opinions and other content into text data containing timestamps and key terms, and updates the current status of the medical conversation in the electronic medical record management platform. A context verification unit is integrated into the system processing module, which contains a multi-layer distributed semantic analysis unit to collect the overall logic and segmented semantic distribution of the text in real time. During the voice conversion process, the semantic analysis unit detects the semantic changes of the newly generated text, generates actual logical data, and compares it with the text logic output by the voice recognition engine. If the logical deviation exceeds the set threshold (such as the preset semantic conflict rule), the text is automatically marked as abnormal, the logical contradiction is corrected, and the actual logical content in the text is updated.

[0052] S3. Multi-speaker processing based on intelligent interference detection and elimination:

[0053] S31. Interference Detection: An acoustic-feature-based interference detection unit is embedded in the speech recognition engine, and the acoustic analysis chip uses it to sense the spectrum occupancy of the speech signal in real time. Using a spectrum distribution analysis algorithm, the spectral energy distribution of the speech signal is calculated. When the energy exceeds a set threshold, the corresponding frequency is determined to be occupied, and voice communication activities with overlapping signals within the same frequency band are marked. Simultaneously, by analyzing the intensity, spectral characteristics, and temporal characteristics of the speech signal, the specific time period where interference occurs is located, forming a pre-allocated time period sequence consisting of multiple continuous time windows, each corresponding to a speech segment.

[0054] S32. Interference elimination: A speech segmentation protocol is introduced, and speech segments are pre-allocated through a fast segmentation mechanism. Text is generated according to speaker priority. Speaker priority is determined using a linear weighted model. This model generates a multi-dimensional feature vector based on the strength of the speech signal, the inverse of the distance between the speech source and the acquisition device (the closer the distance, the higher the priority), and the key level of the medical record content (such as diagnostic opinions have higher priority than patient descriptions). The influence of each dimension is balanced through the feature weight coefficient, and speakers in recognition conflicts are dynamically sorted to ensure that key information is processed first. After sorting, the time sequence of the processing time periods is allocated according to priority, achieving non-overlapping time period allocation and text generation.

[0055] S4 extracts information from the text generated by speech conversion and fills it into the hospital's medical record template.

[0056] The present invention will be further described below in conjunction with Examples 1 to 5:

[0057] Example 1: In the voice input initialization step, the generation and association of the voice collection identification includes the following detailed process. The electronic medical record management platform is provided with a special input interface, which has a standardized field design so that medical staff can accurately enter relevant identification data. Medical staff enter the patient's identity information through this interface, specifically including the patient's name, unique medical record number, department, and other information that can clearly identify the patient's identity and medical treatment scenario. At the same time, it is also necessary to enter the medical record type, such as outpatient medical records, inpatient medical records and other different categories, which will help with the subsequent classification management and retrieval of medical records. In addition, the session time and expected duration are also important entry content. The session time is accurate to the specific date and time, and the expected duration provides a reference for system resource allocation and process planning.

[0058] When the voice collection device is activated, it automatically adds its own relevant information to the identification data. The voice collection device may be a specialized device such as a medical microphone array, which has a unique device ID that can be used to identify and manage different collection devices. At the same time, the device's collection parameters, such as sampling rate and pickup range, are also recorded. The addition of this device information makes the identification data more complete, accurately reflecting the device environment and technical parameters of the voice collection. By integrating patient information, treatment department, medical record type, session time, expected duration, and device information, a complete identification record is formed. This record contains the basic elements of the medical conversation and relevant information about the collection device, laying the foundation for subsequent identification storage, association, and the entire electronic medical record generation process.

[0059] The voice collection device stores the complete identification record in a local cache for subsequent processing. During this storage process, the device generates a unique voice collection identifier using a hash algorithm. The hash algorithm is unidirectional and unique, ensuring that the generated identifier is unique and preventing duplicate identifiers. The generated voice collection identifier can be a string of a specific length, such as a 32-bit string. This string serves as a unique identifier for this medical session and is used to identify and manage the session throughout the system.

[0060] After the voice collection identifier is generated, it needs to be associated with the medical session. First, the identity verification process is triggered, and the content of the voice collection identifier is read through the identity verification module. It is then verified and compared with the original identifier data stored in the electronic medical record management platform. The verification content mainly includes three aspects: the integrity of the identifier, checking whether all fields in the identification record have been entered correctly and whether there are any missing fields; the accuracy, verifying whether the entered information complies with the specifications, such as whether the time format is correct and whether the patient identity information is consistent with the information already in the system; and the uniqueness, ensuring that the generated voice collection identifier is unique in the system and does not overlap with the identifiers of other medical sessions.

[0061] If, after verification, the voice collection identifier meets the requirements of completeness, accuracy, and uniqueness, the association is determined to be successful. At this point, the successfully associated voice collection identifier is uploaded to the electronic medical record management platform. The electronic medical record management platform will establish a long-term mapping relationship between the voice collection identifier and the medical session. This mapping relationship is usually achieved through a database index. The database index can improve the efficiency of data retrieval and call, so that in the subsequent electronic medical record generation, query, and management process, the corresponding medical session and its related voice collection identifier and data can be found quickly and accurately. By establishing this long-term mapping relationship, the precise association between voice data and medical record content is ensured, so that during the entire medical session, all collected voice information can be accurately mapped to the corresponding medical record, providing a reliable foundation for subsequent voice recognition, text generation, and medical record organization and storage.

[0062] During actual operation, when medical staff enter identification data, the system provides real-time prompts and verification functions. For example, if the time format is incorrect, the system will automatically issue a prompt and request re-entry to ensure the accuracy of the identification data. At the same time, the voice collection device automatically communicates with the electronic medical record management platform for information to ensure that the device information is accurately captured and recorded. During the identity verification process, if any problems are found with the identification, such as missing fields or inaccurate information, the system will promptly provide feedback to the medical staff for correction and re-verification, ensuring that only qualified voice collection identifications are associated with medical conversations and uploaded to the platform.

[0063] From entering identification data, adding device information, generating, storing, and linking identifications to verifying identity, rigorous design and processing have been implemented at every stage. This ensures the accuracy and reliability of the initial voice input step, providing a solid foundation for the entire automatic generation of electronic medical records based on voice recognition. Each step is closely interconnected and coordinated, forming a complete process. This ensures efficient and accurate identification management and data association for medical conversations, enabling the smooth execution of subsequent steps such as intelligent voice recognition and multi-speaker processing, thus achieving the automatic generation and management of electronic medical records.

[0064] Example 2: The specific implementation of the intelligent recognition step of voice information involves multiple links such as the deployment of the voice recognition engine, the collection and conversion of voice signals, and the logical verification of the semantic analysis unit. Each link works together to achieve accurate conversion of voice to text and semantic consistency management.

[0065] At the system input interface deployment level, the distributed speech recognition engine achieves full coverage of the collection area through a multi-microphone array. The multi-microphone array is usually optimized according to the spatial structure of the treatment room. For example, microphones are evenly distributed in the four corners or ceiling of the treatment room to ensure that voice signals are collected without blind spots. Each microphone has independent signal collection and transmission functions, which can simultaneously collect voice information from patients, doctors, and other relevant personnel, and transmit the collected analog signals in real time to the speech recognition engine for digital processing and fusion. This distributed deployment method can effectively improve the collection range and accuracy of voice signals and reduce the problem of sound occlusion or attenuation caused by a single collection point.

[0066] When a patient or healthcare provider begins speaking, the speech recognition engine first triggers a real-time detection mechanism. This mechanism pre-processes the input speech signal using a voice activity detection (VAD) algorithm, identifying valid speech segments and eliminating any ambient noise. For example, background noise in the treatment room (such as the sound of air conditioning or footsteps in the hallway) is filtered out by the VAD algorithm, retaining only valid speech content, such as the patient describing their symptoms, the doctor asking about their condition, or providing a diagnosis. These pre-processed, valid speech segments are then transmitted to the engine's core conversion module, which uses a deep learning model (such as a speech recognition model based on the Transformer architecture) to convert the speech signal into text data. The converted text covers multiple dimensions: the patient's description includes the main complaint (e.g., "recurring abdominal pain for the past week") and symptom details (e.g., "pain worsens after meals, accompanied by nausea"); the doctor's diagnosis includes a preliminary diagnosis (e.g., "suspected digestive system disease") and examination recommendations (e.g., "abdominal ultrasound recommended"). The system automatically adds millisecond-accurate timestamps to each piece of text data, allowing for subsequent tracing of the chronological order of voice input. The timestamp format can be "YYYY-MM-DDHH:MM:SS.SSS." Furthermore, key terms (such as "abdominal pain," "nausea," and "ultrasound examination") are automatically identified and tagged, allowing for rapid retrieval and structured processing of medical records.

[0067] The multi-layer distributed semantic analysis unit in the system's processing module plays a key role in the speech-to-text conversion process. This unit employs a hierarchical processing architecture. The bottom-level unit analyzes the grammatical structure of sentences, such as the division of subject, predicate, and object into components, and determines whether the sentence conforms to medical language standards (e.g., whether the grammatical structure of "The patient's temperature rose to 38°C" is correct). The middle-level unit constructs a semantic network for the paragraph, identifying the logical relationships between different sentences, such as causal relationships ("Coughing after catching a cold, so it is recommended to keep warm") and parallel relationships ("The patient needs to undergo a blood and urine routine test"). The top-level unit integrates the logical context of the entire text to form a holistic semantic understanding of the entire medical conversation. For example, it determines whether the diagnosis matches the symptoms described by the patient and whether the treatment recommendation complies with the diagnosis and treatment guidelines of the disease.

[0068] The real-time monitoring mechanism of the semantic analysis unit runs through the whole process of speech conversion. When a new text is generated, the semantic analysis unit first extracts the semantic features of the text (such as keywords, semantic vectors), and then conducts context correlation analysis with the previously generated text content to generate the actual logical data of the current text. For example, if a doctor records "the patient has a history of penicillin allergy" in the previous text, and "it is recommended to use penicillin antibiotics" appears in the subsequent text, the semantic analysis unit detects a logical contradiction between the front and back texts through preset medical taboo rules (such as penicillin is prohibited for patients with penicillin allergy), and at this time, it will determine that the logical deviation exceeds the preset threshold (this threshold is set based on the logical rules in the medical knowledge graph). Once an anomaly is detected, the system will automatically perform the following operations: First, mark the current text as abnormal, usually adding a special identifier (such as "[Logical anomaly]") at the end of the text; then, based on the context information and the medical knowledge base, automatically correct the logical contradiction, for example, correct "penicillin antibiotics" to "cephalosporin antibiotics"; at the same time, record the content differences before and after the correction in the text data to form a correction log for medical staff to refer to during medical record review. The corrected text will re-enter the semantic analysis process to ensure its semantic consistency with the context.

[0069] To improve the accuracy and professionalism of semantic analysis, the system pre-loads an ontology library in the medical field (such as the SNOMED CT terminology set, ICD-10 disease classification codes). The ontology library contains rich medical concepts, terms and their interrelationships, such as the symptom associations between "pneumonia" and "cough" "fever", and the treatment associations between "hypertension" and "antihypertensive drugs". When the speech recognition engine outputs text, the semantic analysis unit will match the terms in the text with the ontology library to verify the correctness and standardization of the terms. For example, if a typo like "blood pressure" (should be "blood pressure") appears in the text, the system can automatically identify and prompt for correction through the term matching function of the ontology library. In addition, the ontology library is also used to build a rule engine for semantic analysis, such as logical rules based on disease diagnosis criteria (such as "systolic blood pressure ≥ 140 mmHg and diastolic blood pressure ≥ 90 mmHg can be diagnosed as hypertension") to ensure that the medical record content conforms to medical professional norms.

[0070] At the system interaction level, the speech recognition engine and the electronic medical record management platform synchronize data through a real-time interface. Whenever the speech conversion successfully generates a piece of text data, the engine will immediately transmit the data to the platform, and the platform will update the current status of the medical conversation (such as "speech collection in progress", "text generation", "semantic verification in progress", etc.). Medical staff can view the speech conversion progress and text content in real time through the platform interface. If abnormally marked text is found, the review process can be manually triggered to confirm or adjust the content automatically corrected by the system. This human-computer collaborative mechanism not only ensures the efficiency of medical record generation, but also provides medical staff with the ability to control and intervene in medical record content, ensuring the accuracy and reliability of medical records.

[0071] The entire intelligent voice recognition process is closely aligned with the practical needs of medical conversations. Through distributed data acquisition using a multi-microphone array, voice conversion using a deep learning model, logical verification using a multi-layered semantic analysis unit, and the knowledge support of a medical ontology library, it achieves full automation from voice input to structured text generation. The seamless integration of data and control flows between these links forms an efficient and accurate voice recognition and semantic management system, providing key technical support for the automated generation of electronic medical records while meeting the strict medical requirements for accuracy, standardization, and traceability of medical record content.

[0072] Example 3: The specific workflow of the interference detection unit in multi-speaker processing covers acoustic perception, spectrum analysis, interference period location and signal conflict determination, etc., and achieves accurate identification of speech signal interference through the collaboration of hardware and algorithms.

[0073] The interference detection unit utilizes an acoustic analysis chip embedded in a speech recognition engine as its core hardware. This chip is capable of real-time voice signal acquisition and frequency domain analysis. During the acoustic perception and interference identification phase, the chip converts the analog voice signal transmitted by the microphone array into a digital signal via an analog-to-digital converter (ADC). The sampling frequency is typically no less than 44.1kHz to ensure the preservation of high-frequency speech details. After the digital signal enters the spectrum analysis module, it is first framed. Each frame is set to 25ms in length and has a frame shift of 10ms. The short-time Fourier transform (STFT) is then used to convert the time domain signal into a frequency domain representation, generating the spectral energy distribution data for each frame.

[0074] The spectrum distribution analysis algorithm realizes interference detection through the following steps: First, the spectrum occupancy state is calculated and the spectrum energy threshold is defined as E th , when the spectrum energy E in a certain frequency band satisfies E≥E thWhen , the frequency band is determined to be occupied. The spectrum energy E here is the integrated value of the power spectrum density of all frequency points in the frequency band, in decibel milliwatts (dBm). For example, if the spectrum energy of multiple consecutive frames exceeds E in the 2000-3000Hz frequency band, th (For example, the preset value is -30dBm), it is marked that there is co-frequency voice communication activity in the frequency band and it is regarded as a potential interference source.

[0075] In the segment detection and positioning phase, the interference detection unit further analyzes the strength, spectral characteristics, and temporal characteristics of the voice signal based on the spectrum analysis results. The signal strength is measured by calculating the root mean square (RMS) value of the frame energy, using the formula:

[0076]

[0077] Among them, x i Indicates the signal amplitude of the i-th sampling point in the frame, N is the number of sampling points per frame. By setting the intensity threshold RMS th , the start and end time of the valid speech segment can be identified. When the spectrum energy of a frame exceeds E th And the signal strength exceeds RMS th When the signal is received, the timestamp of the frame (provided by the system clock synchronization) is combined to determine the specific time period when the interference occurs, for example, there is signal interference between 10.0 seconds and 15.5 seconds after the start of the medical session.

[0078] In order to achieve structured management of interference periods, the system defines the time period set for voice signal allocation as a pre-allocated time period sequence T = {t1, t2, ..., t n}, where each t i Corresponding to a continuous time window, the window length is set to 200ms, and adjacent windows overlap by 50ms to avoid signal truncation. Each time window corresponds to a speech segment, which is used to carry the speech signal of a single speaker. The detection condition for signal strength overlap is: within the same time window, the sum of the strengths of multiple speech signals exceeds the set conflict threshold S th ,Right now:

[0079]

[0080] Where m is the number of speech sources in the time window, RMS k is the root mean square value of the signal strength of the kth speech source, S th is the preset conflict threshold (in dB). If this condition is met, it is determined that there is signal strength overlap in this period and interference source location is required.

[0081] Interference source location is achieved by calculating the contribution signal strength of each speech source. kThe beamforming technology calculation based on the multi-microphone array is as follows:

[0082]

[0083] Where p is the number of microphones, w j is the weighting coefficient of the jth microphone (determined by the beamforming algorithm to enhance the signal in the direction of the target speech source), x j,k is the signal amplitude of the kth speech source collected by the jth microphone. By setting the ratio threshold C th (such as 30%), the contribution signal strength exceeds C th The speech sources are included in the interference source set K={k1,k2,…,k q}, where q is the number of interference sources. For example, if the signal strength contributions of doctors and patients are 45% and 35% respectively, both exceeding 30%, then both are considered interference sources.

[0084] At the hardware level, the acoustic analysis chip utilizes a field-programmable gate array (FPGA) architecture, leveraging its parallel computing capabilities to accelerate spectrum analysis and signal processing, ensuring real-time performance. An integrated cache unit temporarily stores spectrum energy data, signal strength data, and time segment marker information. This data is exchanged with the speech recognition engine's main control module via a bus interface. The main control module triggers the subsequent speech segmentation protocol and prioritization process based on the interference source set and time segment information output by the interference detection unit.

[0085] The entire interference detection process closely integrates acoustic and temporal characteristics, achieving precise identification and location of interference in multi-speaker scenarios through multi-level signal processing algorithms. From frequency domain conversion of speech signals and spectrum occupancy analysis to signal strength-based time segmentation and interference source screening, each link is based on quantifiable parameters and algorithmic logic to ensure the objectivity and repeatability of detection results. This mechanism not only effectively identifies co-channel interference and signal overlap issues, but also provides precise interference source data for subsequent multi-speaker processing, laying the foundation for orderly speech segmentation, priority sorting, and text generation, thereby ensuring the accurate capture and logical organization of multi-source speech information during the electronic medical record generation process.

[0086] Example 4: The speech segmentation protocol in multi-speaker processing achieves interference elimination through a fast segmentation mechanism and dynamic priority sorting. The following describes its implementation in detail in conjunction with specific application scenarios.

[0087] Multi-talker scenarios are common in medical conversations, such as doctors inquiring about a patient's condition (Doctor A, Patient B), nurses supplementing medication history (Nurse C), or multiple medical professionals discussing a patient's condition during a consultation (Doctor D, Doctor E). When the interference detection unit identifies overlapping speech sources within a certain time period (e.g., Doctor A and Patient B speaking simultaneously), the system triggers the speech segmentation protocol. This first dynamically adjusts the time slot allocation of the speech signal using a fast segmentation mechanism. This mechanism allocates independent processing time slots to each speech source based on the detection results of the interference source set, ensuring no overlap between the time slots. For example, if Doctor A and Patient B are detected speaking simultaneously at the 10th second of a conversation, the system immediately initiates fast segmentation: Doctor A's speech is allocated to the 10.0-10.8 second segment, while Patient B's speech is deferred to the 10.8-11.6 second segment, with a 0.2 second gap between the two segments to prevent signal aliasing. Time slot allocation is achieved through real-time clock synchronization, ensuring that the time boundaries of each speech segment precisely correspond to the sampling instants of the acquisition device.

[0088] Dynamic speaker prioritization is a core component of the speech segmentation protocol. The system generates a multidimensional feature vector based on speech signal strength, the distance between the speech source and the acquisition device, and the criticality of the medical record content. For example, consider a conflict scenario between Doctor A and Patient B: Doctor A's speech signal strength can be calculated from the sound wave amplitude collected by the microphone array (e.g., 70 decibels). Doctor A's distance from the acquisition device is assumed to be 1 meter (the inverse distance is 1), and the content of his speech is a diagnosis (preset to "high criticality," with a value of 1). Patient B's signal strength is 65 decibels, and the distance is 2 meters (the inverse distance is 0.5), and the content is a description of his symptoms (medium criticality, with a value of 0.5). The system combines these three features into vectors [70, 1, 1] and [65, 0.5, 0.5] and calculates the overall priority using a linear weighting model. The weighting coefficients of the linear weighting model are preset based on the medical scenario requirements, such as 0.2 for signal strength, 0.3 for inverse distance, and 0.5 for content criticality. Doctor A's overall priority is: 70 × 0.2 + 1 × 0.3 + 1 × 0.5 = 14 + 0.3 + 0.5 = 14.8; Patient B's overall priority is: 65 × 0.2 + 0.5 × 0.3 + 0.5 × 0.5 = 13 + 0.15 + 0.25 = 13.4. Based on the calculation results, Doctor A has a higher priority than Patient B, so their speech segments are allocated a higher time slot.

[0089] The priority-driven text generation process strictly adheres to the sorting results. The system processes the audio segments of each voice source in descending order of priority, prioritizing high-priority sources for inclusion in the draft electronic medical record. For example, in the scenario described above, Doctor A's diagnosis, "Suspected acute gastroenteritis; fasting and fluid replacement recommended," is prioritized for conversion to text and timestamped "10:00:10.000." It is also marked as critical information (e.g., automatically bolded). Patient B's symptom description, "Vomiting three times after waking up in the morning and watery diarrhea," is processed after Doctor A's segment ends and timestamped "10:00:10.800." To prevent truncation of lower-priority audio sources, the rapid segmentation mechanism dynamically adjusts the segment length based on the duration of the speech. If Patient B's speech lasts 1.2 seconds, the system will extend the segment to 10.8-12.0 seconds to ensure that subsequent content, such as "accompanied by abdominal cramps," is fully captured.

[0090] In complex scenarios involving multiple actors (such as doctors, nurses, and patients speaking simultaneously), the dynamic priority sorting mechanism demonstrates greater adaptability. For example, Nurse C's voice signal strength is 60 decibels, and she is 1.5 meters away from the data collection device (the inverse distance is 0.67). The content is a reminder about the patient's history of drug allergies (high criticality, value 1). Its multidimensional feature vector is [60, 0.67, 1], and the comprehensive priority is calculated as: 60 × 0.2 + 0.67 × 0.3 + 1 × 0.5 = 12 + 0.201 + 0.5 = 12.701. In this case, the priority ranking is Doctor A (14.8) > Nurse C (12.701) > Patient B (13.4). Nurse C's content has a higher priority than Patient B due to its higher criticality. Her reminder, "The patient has a history of levofloxacin allergy," will be prioritized, preventing the doctor from prescribing contraindicated medications.

[0091] The system supports customizing priority weight coefficients through the electronic medical record management platform to adapt to different diagnosis and treatment scenarios. For example, in an outpatient scenario, the doctor's content key level weight can be increased to 0.6, and the patient description weight can be set to 0.3 to highlight the priority of the doctor's diagnosis; in a consultation scenario, the weight of the inverse of the distance can be set according to the professional title level (such as the distance weight coefficient of the chief physician is 0.4, and the distance weight coefficient of the resident physician is 0.2) to ensure that the opinions of senior doctors are recorded first. Weight adjustment is achieved through the slider bar or numerical input box on the platform interface. The modification takes effect in real time without restarting the system.

[0092] The time sequence of the segment start in the fast segmentation mechanism is directly driven by the priority sorting results. Segment start times for higher-priority speech sources are assigned earlier timestamps, followed by lower-priority sources. For example, if Doctor D (priority 15), Doctor E (priority 13), and Nurse F (priority 12) speak simultaneously, Doctor D's segment starts at 5.0 seconds, Doctor E's at 6.0 seconds, and Nurse F's at 7.0 seconds. The length of each segment is determined based on the actual speech duration (e.g., 2 seconds for each), ensuring that the processing order is consistent with the decision-making hierarchy in the medical process.

[0093] During the text generation process, the system adds role tags to the text content of each voice source, such as [Doctor A] and [Patient B], to facilitate medical staff to trace the speaker. Role tags are associated with the device information in the voice collection identifier. For example, the microphone device ID used by the doctor corresponds to the "doctor" role, and the bedside microphone used by the patient corresponds to the "patient" role. If multiple voice sources with the same role (such as two doctors) speak at the same time, the system distinguishes them by the physical location of the device (such as the coordinate number of the microphone array) and reflects this in the tag (such as [Doctor-Left], [Doctor-Right]).

[0094] The entire speech segmentation and prioritization process is based entirely on real-time acoustic features and pre-set rules, enabling orderly multi-speaker speech processing without manual intervention. From interference detection to time slot allocation, priority calculation, and text generation, each step is seamlessly integrated, ensuring that electronic medical records accurately reflect the information hierarchy within medical conversations, preventing key diagnostic insights from being overshadowed by secondary information. While also ensuring the complete recording of essential information such as patient descriptions and allergy history reminders, this improves the logicality and clinical value of medical records.

[0095] Example 5: This example is developed around the overall system architecture and extended functions, combining the distributed deployment of the electronic medical record management platform, the multimodal connection of voice acquisition equipment, the medical ontology library support for semantic analysis, and the customized rules for multi-speaker priority to form a complete technical implementation solution.

[0096] At the level of the overall system architecture, the electronic medical record management platform adopts a distributed database architecture (such as MySQL cluster) to improve data processing performance through master-slave replication and read-write separation mechanisms. The platform is deployed on the hospital's private cloud server and supports multi-terminal access (such as doctor's workstations and nurse station tablets). Each terminal communicates with the platform through the HTTPS protocol to ensure data transmission security. Taking the outpatient scenario of a tertiary hospital as an example, the platform needs to process the identification data, voice collection identification and generated text data of hundreds of medical sessions at the same time. The distributed database uses the technology of sharding to split and store the data according to department and time dimensions. For example, internal medicine medical record data is stored in an independent database node, and surgical data is stored in another node to improve query efficiency. When medical staff query a patient's electronic medical record through a workstation, the platform can retrieve the associated voice collection identification and corresponding text records within a millisecond response time.

[0097] The voice collection device supports dual-mode connections of Wi-Fi and Bluetooth to adapt to different diagnosis and treatment environments. In outpatient clinics, the device is usually connected to the hospital LAN via Wi-Fi, and the identification data is synchronized to the electronic medical record management platform in real time. For example, after the doctor turns on the microphone array in the clinic, the device automatically connects to the preset Wi-Fi hotspot and uploads the patient's identity information, device ID and other identification data to the platform. The platform generates a voice collection identification and pushes it back to the device to establish a real-time connection. In mobile ward rounds scenarios (such as intensive care units), the device can be connected to the medical staff's handheld terminal (such as a PDA) via Bluetooth to temporarily store the identification data, and then synchronize it to the platform in batches after the terminal is connected to the network. This dual-mode connection mechanism ensures the flexibility of the device in fixed and mobile scenarios, avoiding collection interruptions due to network environment restrictions.

[0098] The semantic analysis unit pre-loads medical domain ontology libraries (such as the SNOMEDCT terminology set and ICD-10 disease classification codes) to provide knowledge support for the logical verification of medical record content. Taking the diagnosis and treatment of respiratory diseases as an example, when a patient describes "coughing, sputum production, and fever for three days," the speech recognition engine converts the description into text. The semantic analysis unit first matches the terms "cough" and "fever" in SNOMEDCT to confirm the standardization of the terms. Then, based on the ICD-10 coding rules, it determines the disease categories that these symptoms may be associated with (such as J20-J22 acute bronchitis and J18 unspecified pneumonia). If the doctor's diagnosis is "upper respiratory tract infection," the system verifies whether the diagnosis meets the symptom combination (upper respiratory tract infection is usually characterized by nasal congestion and runny nose, and is rarely accompanied by fever) through the disease-symptom association rules in the ontology library. If there is a logical deviation, the diagnosis is automatically marked as abnormal and additional examination suggestions are prompted.

[0099] The custom priority rule function of the multi-speaker processing module is implemented through the visual interface of the electronic medical record management platform. Medical staff can enter the "Priority Management" page in the "System Configuration" menu of the platform and set the default weight coefficient for different roles (doctors, nurses, patients, and family members). For example, in a pediatric outpatient clinic scenario, doctors need to give priority to the urgent conditions of children. The doctor's signal strength weight can be set to 0.3 (higher than the default value of 0.2) and the content critical level weight can be set to 0.6 (higher than the default value of 0.5) to ensure that the doctor's diagnosis opinion is given priority in multi-speaker conflicts. When a nurse reports the child's latest vital signs (such as a heart rate of 140 beats / minute), the system calculates its priority based on the preset weights: signal strength 65 decibels × 0.3 + distance inverse 0.8 (1.25 meters from the device) × 0.2 + content criticality 0.8 (vital signs are critical information) × 0.5 = 19.5 + 0.16 + 0.4 = 20.06. This is higher than the inquiries from the accompanying family members (signal strength 60 decibels × 0.3 + distance inverse 0.5 × 0.2 + content criticality 0.3 × 0.5 = 18 + 0.1 + 0.15 = 18.25). Therefore, the nurse's voice segment is allocated a priority time period to ensure that key vital signs data is recorded in a timely manner.

[0100] In consultation scenarios, the priorities of multiple doctors can be dynamically adjusted based on their professional titles. For example, the default distance reciprocal weight for a chief physician is 0.4 (corresponding to a reciprocal weight of 0.5 at a distance of 2 meters from the device, with a priority contribution of 0.4 × 0.5 = 0.2), while the distance reciprocal weight for a resident physician is 0.2 (with a contribution of 0.2 × 0.5 = 0.1 at the same distance). When a chief physician (at a distance of 3 meters, a signal strength of 70 decibels, and a content critical level of 1) and a resident physician (at a distance of 2 meters, a signal strength of 65 decibels, and a content critical level of 0.8) speak at the same time, the chief physician's overall priority is: 70×0.2+0.33(the inverse of 3 meters)×0.4+1×0.5=14+0.132+0.5=14.632; the resident physician's priority is: 65×0.2+0.5×0.2+0.8×0.5=13+0.1+0.4=13.5. The chief physician has a higher priority due to the weight difference corresponding to his or her professional title, and his or her treatment recommendation of "transfer to ICU monitoring" takes precedence over the resident physician's "recommendation for continued observation" and is recorded in the medical record.

[0101] The human-computer interaction interface design of the system focuses on the convenience of medical operations. In the electronic medical record editing interface, medical staff can click on the voice collection icon on the timeline, listen back to the original voice recording of the corresponding period, and verify the accuracy of the text conversion. For abnormal text automatically marked by the system (such as logical corrections), the interface is highlighted in yellow, and a drop-down menu is provided to display the content comparison before and after the correction. Medical staff can choose to accept the correction, reject the correction, or manually edit the text. For example, when the system corrects "the patient is allergic to sulfonamides" to "the patient is allergic to penicillin", the medical staff can confirm the actual content of the speech by listening back to the recording. If it is found that it is a voice recognition error, the correction can be rejected and manually corrected to the correct content.

[0102] Regarding data security, the electronic medical record management platform adheres to the "Guidelines for Health and Medical Data Security" and encrypts stored identification data, voice files, and text records. Voice capture identifiers are linked to patient identity information via hash values, rather than directly storing plaintext medical record numbers, minimizing the risk of data leakage. The platform regularly backs up data, which is stored on offline physical media. Access control is also implemented, restricting access to data to only medical staff and system administrators, who are authorized to access the data within their respective roles. Other personnel are unable to view or modify data.

[0103] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0104] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for automatically generating electronic medical records based on speech recognition, characterized in that: The following steps are involved: S1, voice input initialization: assign a unique voice collection identifier to each medical session; S2, intelligent voice recognition: The system is equipped with a voice recognition engine that automatically converts speech into text upon input. A context verification unit is integrated into the system's processing module to check the semantic consistency of the text in real time during the conversion process. If any logical contradiction is detected, the text is automatically corrected and marked as an anomaly. S3, multi-speaker processing based on intelligent interference detection and elimination: S31, embedding an interference detection unit based on acoustic features in the speech recognition engine to monitor the spectrum distribution of the speech signal and the interference occurrence period in real time; S32: To address the issue of interference periods in speech signals, a speech segmentation protocol is introduced. The speech segmentation protocol pre-allocates speech segments through a fast segmentation mechanism and generates text based on speaker priority. The speaker priority is determined using a linear weighted model, and the priority of speakers in recognition conflicts is dynamically sorted to ensure that key information is processed first. S4 extracts information from the text generated by speech conversion and fills it into the hospital's medical record template.

2. The method for automatically generating electronic medical records based on speech recognition according to claim 1, characterized in that: Said S1 specifically includes: S11, identification generation and content entry: Input identification data including patient identity, treatment department, medical record type, session time and estimated duration through the electronic medical record management platform. After the voice acquisition device is started, the device information is added to the identification data to form a complete identification record; S12, identification storage and association: using the voice collection device to store the identification data, generate a corresponding voice collection identification, and associate the generated voice collection identification with the medical conversation.

3. The method for automatically generating electronic medical records based on speech recognition according to claim 2, characterized in that: The S1 also includes identity verification: after the voice collection identifier is associated, the voice collection identifier content is read through the identity authentication module and compared with the identifier data stored in the electronic medical record management platform. The verification content includes the integrity, accuracy and uniqueness of the voice collection identifier; the successfully associated voice collection identifier is uploaded to the electronic medical record management platform to establish a long-term mapping relationship between the voice collection identifier and the medical session.

4. The method for automatically generating electronic medical records based on speech recognition according to claim 1, characterized in that: The S2 specifically includes: S21, installing a speech recognition engine at the input interface of the system, and the speech recognition engine distributes and covers the collection area through a multi-microphone array; S22, when voice input begins, triggers the voice recognition engine to detect the voice signal in real time and automatically convert it into text data, including patient description, symptom details, diagnosis opinion, timestamp and key terms. After successful conversion, the current status of the medical session in the electronic medical record management platform is updated.

5. The method for automatically generating electronic medical records based on speech recognition according to claim 4, characterized in that: The processing module of the system is configured with a multi-layer distributed semantic analysis unit, which collects the overall logic and segmented semantic distribution of the text in real time. When the voice is converted, the semantic analysis unit detects the semantic changes of the newly generated text and generates the actual logical data of the current text. The text logic output by the voice recognition engine is compared with the actual logic collected by the semantic analysis unit. If it is detected that the logic deviation exceeds the set threshold, the current text is automatically marked as abnormal, and the actual logical content is updated in the text data.

6. The method for automatically generating electronic medical records based on speech recognition according to claim 1, characterized in that: The S31 specifically includes: Acoustic perception and interference identification: The interference detection unit uses an embedded acoustic analysis chip to perceive the spectrum occupancy of voice signals in real time. The acoustic analysis chip uses a spectrum distribution analysis algorithm to detect the frequency distribution of signal transmission and mark voice communication activities with overlapping signals within the same frequency band. Time period detection and positioning: When signal interference is detected, the specific time period where the interference occurs is located by calculating the strength, spectral characteristics, and time characteristics of the voice signal.

7. The method for automatically generating electronic medical records based on speech recognition according to claim 6, characterized in that: The spectrum distribution analysis algorithm specifically includes: Spectrum occupancy status calculation: By analyzing the spectrum energy distribution of the voice signal, the frequency occupancy status is determined. When the spectrum energy exceeds the set energy threshold, it indicates that the frequency is occupied. The frequency occupancy status detection condition is as follows: the spectrum occupancy status is determined by the energy threshold. Energy above the threshold indicates occupancy, and energy below the threshold indicates idleness.

8. The method for automatically generating electronic medical records based on speech recognition according to claim 7, characterized in that: The time period set for voice signal allocation is defined as a pre-allocated time period sequence, which includes multiple continuous time windows. In the time period sequence, each time window corresponds to a voice segment, and the detection condition for signal strength overlap is that the sum of the strengths of multiple voice signals within the time period exceeds a set conflict threshold; Interference source localization: If signal strength overlap is detected, the interference source set is located by calculating the contribution signal strength of each voice source. Voice sources whose contribution signal strength exceeds the ratio threshold are included in the interference set.

9. The method for automatically generating electronic medical records based on speech recognition according to claim 8, characterized in that: The S32 specifically includes: S321, fast segmentation mechanism to implement time slot pre-allocation: After interference detection, the fast segmentation mechanism is used to dynamically adjust the time slot allocation of the voice signal, assigning a specific processing time slot to each voice source to ensure that there is no overlap between time slots; S322, Dynamic Sorting of Speaker Priorities: Generates a multidimensional feature vector based on the strength of the speech signal, the distance between the speech source and the acquisition device, and the criticality level of the medical record content. Dynamically prioritizes the multidimensional feature vectors of the conflicting speech sources using a linear weighted model. S323, priority-driven text generation: assigning a time sequence of processing periods according to the sorting results.

10. The method for automatically generating electronic medical records based on speech recognition according to claim 9, characterized in that: The multidimensional feature vector is represented as a feature combination of the speech source, the feature combination including signal strength, inverse distance and content key level; The linear weighted model calculates the comprehensive priority of the voice source and sorts it based on the comprehensive priority: the comprehensive priority is balanced by the feature weight coefficient, and the inverse of the distance indicates that the voice source closer to the acquisition device has higher priority; The fast segmentation mechanism dynamically adjusts the time period allocation of the speech signal to be expressed as a time period start time sequence, where the time period start time is determined by the reading order of the speech source, and the reading order is allocated by the priority sorting result.

Citation Information

Cited By

  • Clinical nursing patient information monitoring and recording method and system

    CN121188817A

  • Medical record automatic filling method and system based on voice input

    CN121415972A

  • A method and system for automatic medical record filling based on voice input

    CN121415972B