Oral doctor's order intelligent processing method based on ai chest plate recorder system and related device
By using an AI-powered name tag recorder system to reduce noise and separate human voices from audio data at emergency scenes, and combining voiceprint recognition and text consistency comparison, structured medical order data is generated. This solves the problems of distorted medical order information and non-standard management in emergency care, and improves the accuracy and safety of medical order processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN PEOPLES HOSPITAL
- Filing Date
- 2026-05-08
- Publication Date
- 2026-06-05
AI Technical Summary
In emergency situations, existing technologies lack intelligent methods for automatically recording and managing verbal medical orders, leading to easily distorted medical order information, perfunctory verification, and non-standard data management, making it difficult to ensure the accuracy of medical order processing and medical safety.
The AI badge recorder system uses noise reduction and voice separation processing on the audio data from the emergency scene. Combined with voiceprint identity verification and voice-text consistency comparison, it generates structured medical order data, enabling precise and intelligent management of medical orders.
It improves the accuracy of medical order processing and medical safety, reduces the risk of information transmission deviation and record errors, and adapts to the rapid treatment needs of emergency and critical care scenarios.
Smart Images

Figure CN122158042A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to a method and related device for intelligent processing of oral medical orders based on an AI badge recorder system. Background Technology
[0002] In emergency and critical care settings, when a patient's condition is sudden and urgent and time is of the essence, medical staff often use verbal instructions to quickly issue treatment and medication orders on-site in order to seize the golden time for treatment and ensure the timeliness of rescue for critically ill patients.
[0003] Currently, the existing emergency verbal medical order execution and recording management model relies heavily on the traditional method of medical staff manually memorizing and verbally repeating the orders on-site, combined with subsequent paper-based recording of the orders. There is a lack of intelligent automatic recording and management methods for verbal medical orders. Summary of the Invention
[0004] This application provides a method and related device for intelligent processing of oral medical orders based on an AI badge recorder system. By performing personalized noise reduction and voice separation processing on audio data at the emergency scene, combined with voiceprint identity verification, dual-voice text consistency comparison, and data processing logic for standardized generation of structured medical order data, it achieves precise and intelligent management of oral medical orders. This effectively solves the problems of easily distorted medical order information, perfunctory verification, and non-standard data management in emergency scenarios, and improves the accuracy of medical order processing and the level of medical safety management.
[0005] In a first aspect, embodiments of this application provide an intelligent processing method for verbal medical orders based on an AI badge recorder system, applied to a server of the AI badge recorder system. The AI badge recorder system includes a first AI badge, a second AI badge, the server, and an execution terminal. The first AI badge is worn by a doctor, and the second AI badge is worn by a nurse. The method includes: acquiring real-time emergency scene audio data collected by the first AI badge and the second AI badge; the emergency scene audio data collected by the first AI badge and the second AI badge respectively also includes a device identification identifier; performing noise reduction and enhancement processing on the emergency scene audio data to obtain a first clean speech signal after the noise reduction and enhancement processing on the first AI badge, and a second clean speech signal after the noise reduction and enhancement processing on the second AI badge; determining, based on the device identification identifier, that the first clean speech signal is a doctor's speech signal and the second clean speech signal is a nurse's speech signal; and pre-configuring a fixed dialogue sequence relationship for the medical order interaction, wherein the fixed dialogue sequence relationship for the medical order interaction... The sequence relationship includes a fixed interactive logic where the doctor speaks first and the nurse repeats later, along with a preset repeating interval. The start and end timestamps of the doctor's and nurse's voice signals are extracted, and the interval between the two voice signals is calculated. When the start and end timestamps of the doctor's voice signal are found to be earlier than those of the nurse's voice signal, and the interval is not greater than the preset repeating interval, the first clean voice signal is determined to be the doctor's spoken voice signal, and the second clean voice signal is determined to be the nurse's repeated voice signal. Voiceprint recognition matching of the doctor's spoken voice signal is performed based on a pre-stored authorized medical personnel voiceprint feature database, and identity authorization is determined based on the matching result. If authorization is successful, the doctor's spoken voice signal is transcribed to generate a medical order text. The nurse's repeated voice signal is transcribed to generate a repeating text. Consistency checks are performed on the medical order text and the repeating text. If the check passes, structured medical order data is generated based on the medical order text. The medical order data is then sent to the execution terminal for the nurse to perform the corresponding medical operations.
[0006] Secondly, embodiments of this application provide an intelligent processing device for oral medical orders based on an AI badge recorder system. The AI badge recorder system includes a first AI badge, a second AI badge, a server, and an execution terminal. The first AI badge is worn by a doctor, and the second AI badge is worn by a nurse. The server is used to execute the step instructions as described in any of the methods in the first aspect. The device includes: an acquisition unit for acquiring real-time emergency scene audio data collected by the first AI badge and the second AI badge; and a processing unit for performing noise reduction and enhancement processing and human voice separation processing on the emergency scene audio data. The system obtains the doctor's spoken voice signal and the nurse's repeated voice signal; it performs voiceprint recognition and matching on the doctor's spoken voice signal according to a pre-stored authorized medical personnel voiceprint feature database, and performs identity authorization determination based on the matching result; if authorization is successful, it transcribes the doctor's spoken voice signal to generate a medical order text; it transcribes the nurse's repeated voice signal to generate a repeated text; it performs consistency verification on the medical order text and the repeated text, and if the verification is successful, it generates structured medical order data based on the medical order text; the sending unit is used to send the medical order data to the execution terminal for the nurse to perform the corresponding medical operation.
[0007] As can be seen, in this embodiment, by acquiring real-time audio data from the emergency scene collected by the AI badges worn by doctors and nurses, noise reduction and enhancement processing and voice separation processing are performed on the audio data to accurately obtain the doctor's spoken voice signal and the nurse's repeated voice signal; voiceprint recognition and matching are performed on the doctor's spoken voice signal through a pre-stored authorized medical personnel voiceprint feature database, and identity authorization is determined based on the matching results. After authorization is passed, the two voice signals are transcribed into speech to generate corresponding medical order text and repeated text; consistency verification is performed on the two types of text, and after verification, structured medical order data is generated based on the medical order text and sent to the execution terminal to guide the nurse to complete the corresponding medical operation. This application relies on a complete implementation process that integrates intelligent audio processing, identity verification, automatic text comparison and validation, and standardized generation and distribution of medical order data. It reduces the drawbacks of manual memorization, manual verification, and post-event paper recording, lowers the risk of deviations in the transmission of medical order information and errors in recording in complex emergency situations, improves the accuracy of verbal medical order verification and the efficiency of information transmission, standardizes the flow of verbal medical orders in emergency situations, enhances the safety and full traceability of diagnosis and treatment operations, and adapts to the actual application needs of rapid treatment in emergency scenarios. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic diagram of the structure of the AI badge recorder system provided in the embodiments of this application; Figure 2 This is a structural block diagram of the AI name tag provided in an embodiment of this application; Figure 3 A schematic diagram of an AI name tag provided in an embodiment of this application; Figure 4 This is a schematic diagram of the AI badge management interface provided in an embodiment of this application; Figure 5 A schematic diagram illustrating the display of structured medical order data on the execution terminal provided in this embodiment of the application; Figure 6 A flowchart illustrating the intelligent processing method for oral medical orders based on an AI badge recorder system provided in this application embodiment; Figure 7 This is a schematic diagram of the microphone distribution for an AI badge provided in an embodiment of this application; Figure 8 A schematic diagram illustrating the complete process of intelligent processing of oral medical orders based on an AI badge recorder system provided in this application embodiment; Figure 9 A functional unit structure diagram of an intelligent processing device for oral medical orders based on an AI badge recorder system provided in this application embodiment; Figure 10 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0010] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0011] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but in some embodiments includes steps or units not listed, or in some embodiments includes other steps or units inherent to these processes, methods, products, or apparatuses.
[0012] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0013] In the embodiments of this application, "and / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone; A and B exist simultaneously; B exists alone. Among them, A and B can be singular or plural.
[0014] In this embodiment, the symbol " / " can indicate that the preceding and following objects are in an "or" relationship. Alternatively, the symbol " / " can also represent a division sign, i.e., performing a division operation. For example, A / B can mean A divided by B.
[0015] In the embodiments of this application, "at least one item" or its similar expression refers to any combination of these items, including any combination of a single item or a plurality of items. "One or more" means one or more, while "multiple" means two or more. For example, "at least one item" of a, b, or c can represent the following seven cases: a, b, c; a and b; a and c; b and c; a, b, and c. Each of a, b, and c can be an element or a set containing one or more elements.
[0016] In the embodiments of this application, "equal to" can be used with "greater than" and is applicable to technical solutions used when "greater than" is used; it can also be used with "less than" and is applicable to technical solutions used when "less than" is used. When "equal to" is used with "greater than", it is not used with "less than"; when "equal to" is used with "less than", it is not used with "greater than".
[0017] Emergency and critical care settings are characterized by their suddenness and high time sensitivity. To seize the golden window for patient treatment, medical staff commonly use verbal orders to quickly issue various diagnostic, treatment, and medication instructions, making it an indispensable form of communication in the resuscitation of critically ill patients. Currently, the overall execution and recording of emergency verbal orders still largely relies on the traditional operational model of manual memorization and verification by medical staff, on-site verbal repetition for confirmation, and subsequent paper-based recording and archiving.
[0018] Due to the unique and complex environment of emergency scenes and the limitations of traditional management models, the current process of handling verbal medical orders generally has many inherent shortcomings. Emergency scenes are characterized by numerous personnel and equipment, and significant environmental noise interference, making it highly susceptible to information discrepancies during manual verbal communication of medical orders, thus compromising the accuracy of the core content. Furthermore, the traditional model lacks standardized medical staff identity verification procedures and a digitalized medical order management chain. The entire process of order creation, verification, and execution is relatively crudely managed, and manual memorization and recording methods are prone to errors and omissions in medical order information, with verification becoming a mere formality, making it difficult to guarantee the standardization of basic medical order management.
[0019] In addition, traditional verbal medical orders rely on subjective manual verification and judgment, lack standardized compliance verification and control mechanisms, have weak capabilities for preventing and controlling medical safety risks in advance, and the real-time nature of the form of post-event recording and archiving of medical orders is insufficient, making it difficult to leave traces throughout the entire process. Overall, it is difficult to adapt to the core practical needs of modern emergency and first aid work, which require rapid transmission, accurate verification, safe management and full traceability of verbal medical orders.
[0020] In summary, existing technologies generally lack intelligent and automated processing, accurate verification, and digital closed-loop management solutions for verbal medical orders that are adapted to the complex scenarios of emergency care. They are unable to meet the actual diagnostic and treatment needs of rapid transfer, accurate verification, safety control, and full traceability of emergency verbal medical orders.
[0021] To address the problems existing in the prior art, this application provides a method and related device for intelligent processing of oral medical orders based on an AI badge recorder system. The technical solution of this application will be described in detail below with reference to specific embodiments.
[0022] Before providing a detailed explanation of the intelligent processing method for oral medical orders based on the AI badge recorder system, we will first introduce the AI badge recorder system to which this method applies.
[0023] Please see Figure 1 , Figure 1This is a schematic diagram of the AI badge recorder system provided in this application embodiment. The intelligent processing method for oral medical orders based on the AI badge recorder system relies on the overall collaborative operation of the AI badge recorder system. To realize the intelligent management and control of the entire process of emergency oral medical orders, the AI badge recorder system mainly consists of three core devices: multiple AI badges, a server, and an execution terminal. Stable two-way communication connections are established between multiple AI badges, the server, and the execution terminal. Data interaction and collaborative cooperation are completed between the devices in real time, providing hardware and data support for the entire process of oral medical order audio acquisition, intelligent data processing, identity compliance verification, accurate content verification, and medical order issuance and execution.
[0024] Among them, multiple AI name tags are differentiated and assigned according to the job functions of different staff members within the hospital. They can be worn by various on-duty medical and supporting staff, such as clinical doctors, front-line nurses, hospital medical quality control management personnel, and on-site emergency assistance staff. They can fully meet the needs of on-site audio collection and work interaction recording in various medical work scenarios, such as routine diagnosis and treatment, emergency rescue, and medical quality control verification.
[0025] Please refer to Figure 2 , Figure 3 , Figure 2 This is a structural diagram of the AI name tag. Figure 3 This is a schematic diagram of the AI badge. The badge is equipped with multiple microphones arranged in a differentiated manner according to fixed directional orientations. Different microphones collect sound signals from different directions, accurately distinguishing between the target voice acquisition direction and non-target environmental noise acquisition direction. Among the multiple microphones is a main microphone, which corresponds to the target sound acquisition direction. Its installation position is precisely aligned with the mouth and nose of the medical personnel wearing the badge, specifically for acquiring the core target voice signals of medical personnel dictating medical orders and repeating verifications at close range, ensuring clear and pure effective voice acquisition. The remaining microphones are reference microphones, corresponding to multiple non-target sound acquisition directions such as lower, middle, and right sides. They are mainly used to collect ambient background noise such as the operating sounds of various equipment and environmental noise at the emergency scene. This provides a precise noise comparison data source for subsequent audio noise reduction and enhancement, and voice separation processing by the server, effectively filtering out irrelevant interference noise and improving the quality of effective voice signal acquisition.
[0026] The AI badge also integrates an audio-visual and vibration prompt module, featuring multiple prompt functions such as sound broadcasting, light flashing, and vibration reminders. It can output corresponding prompt signals based on the status of medical order collection, data transmission, identity verification results, and medical order verification progress. This allows medical staff to quickly perceive the working status of the AI badge in busy emergency situations where visibility is limited, and to promptly understand the feedback of core processes such as voice collection, data uploading, and medical order verification, ensuring that the entire medical order processing workflow is carried out synchronously and collaboratively.
[0027] Meanwhile, the AI badge itself can also integrate various other reasonably adapted component modules such as a display screen and a camera. The display screen serves only as an auxiliary structure and can be adapted to display relevant basic equipment and job-related information according to actual application needs. This application does not specifically limit the content displayed on the display screen or the specific types and quantities of added components. It should be noted that this application only provides an exemplary description of the necessary hardware structure required for the AI badge to achieve core medical order audio collection and status prompts. In actual application, the AI badge can be equipped with other adapted functional components and hardware modules according to different hospital emergency work deployments and on-site usage needs. Any reasonable expansion structure added to the core hardware infrastructure of this application falls within the protection scope of this application.
[0028] In this application, the AI badge recorder system is specifically configured with a first AI badge and a second AI badge. The first AI badge is exclusively for use by doctors responsible for issuing verbal emergency medical orders, while the second AI badge is exclusively for use by nurses responsible for receiving and implementing those orders. Each AI badge has a unique device identification identifier, and this identifier is uniquely linked to the doctor's or nurse's ID number, ensuring accurate and unambiguous matching between each AI badge device and the medical personnel wearing it.
[0029] For specific implementation details, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic diagram of the AI badge management interface. As the backend control terminal display interface, the AI badge management interface centrally displays the basic on-duty information of multiple on-duty doctors and nurses, including the name, job title, identity category identifier, employee number, and the unique AI badge device code that is uniquely bound to each medical staff member. This enables visual and unified management of all AI badge devices and the medical staff wearing them, accurate identity matching, and rapid ledger verification.
[0030] The primary function of the first AI badge is to collect raw audio data of the entire process of doctors issuing verbal emergency orders in complex emergency situations, and to continuously upload all the raw audio data collected in real time to the server. The primary function of the second AI badge is to collect raw audio data of the entire process of nurses repeating and verifying verbal orders in emergency situations, and to transmit the collected audio data to the server in real time.
[0031] As the core data processing and overall control hub of the AI badge recorder system, the server maintains regular communication connections with the first AI badge, the second AI badge, and the execution terminal. The core function of the server is to receive two channels of raw audio data uploaded by the first and second AI badges, and sequentially complete all data processing and logical judgment tasks such as audio noise reduction and enhancement, accurate separation of human voice, voiceprint identification and verification of medical staff, real-time speech transcription, consistency comparison and verification of medical order text and repeated text, and generation of standardized structured medical order data.
[0032] The execution terminal is a dedicated terminal device in the AI badge recorder system designed for nursing and medical staff to receive, view, and confirm the execution of verbal emergency medical orders. The execution terminal includes various adaptable forms such as tablet terminals for emergency and critical care departments, computer terminals for nurses at the nurse station, and portable medical operation terminals for nurses. It can be flexibly deployed according to different work layouts in the emergency and critical care scene. In practical emergency care scenarios, execution terminals are deployed long-term in core work areas such as nursing workstations and emergency nurse stations in the emergency resuscitation room. They are specifically used to receive standardized, structured medical order data issued by the server in real time after verification. Simultaneously, the core content of the emergency verbal medical order, such as the corresponding treatment items, medication names, dosages, execution frequency, applicable patient information, and operational precautions, is clearly and intuitively displayed in a standardized text interface. This allows nurses receiving the medical orders to quickly review and accurately verify all the information in the orders. At the same time, nurses can use the execution terminal to complete supporting operations such as viewing and confirming medical orders, marking execution status, and uploading operation receipts. The entire process relies on the execution terminal to achieve visualized and standardized viewing and practical management of emergency verbal medical orders, ensuring the rapid and accurate implementation of emergency medical orders.
[0033] In summary, the complete data flow process of the AI badge recorder system is as follows: The first AI badge worn by the doctor and the second AI badge worn by the nurse simultaneously collect raw audio data from the emergency scene and upload it to the server in real time. The server performs unified intelligent processing on the two audio streams, accurately separating and extracting the valid doctor's spoken voice signal and the nurse's repeated voice signal. It then sequentially completes core processing steps such as doctor's voiceprint authentication, intelligent transcription of the two voice signals, and consistency verification between the medical order text and the repeated text. After the verification result is deemed satisfactory, the server automatically generates standardized structured medical order data and simultaneously remotely distributes the compliant medical order data to the execution terminal, where it intuitively displays the specific operational content of the medical order (e.g., ...). Figure 5 (As shown).
[0034] It should be noted that the verbal recitation by nurses at the emergency scene is only for quick and short-term verification at the rescue site. It is used to immediately confirm the general meaning of the medical order and ensure rapid connection of rescue actions. Moreover, the nurse responsible for reciting the medical order at the emergency scene is often not the same person as the nurse who actually implements the medical order. Their job responsibilities are different. Therefore, the standardized and structured medical order displayed on the execution terminal must be the only official basis for execution. After the nurse reviews the complete medical order and specific operation content displayed intuitively on the execution terminal, she will accurately carry out the corresponding medical operation according to the standard information on the terminal. This completes the closed-loop data flow of the emergency verbal medical order from on-site audio collection, intelligent back-end processing, compliance and security verification to terminal issuance and execution.
[0035] In this application, only the three core hardware structures—AI badges, servers, and execution terminals—are described as examples. In addition, the AI badge recorder system of this application can be supplemented with other supporting auxiliary equipment and expansion hardware units according to the needs of actual application scenarios. This application does not impose specific limitations or unique constraints on the types, quantities, and specific forms of other additional equipment added to the system, and all such additional equipment should fall within the protection scope of this application.
[0036] The intelligent processing method for oral medical orders based on an AI badge recorder system provided in this application is applied to... Figure 1 For the server in the middle, please refer to Figure 6 , Figure 6 This is a flowchart illustrating the intelligent processing method for verbal medical orders based on an AI-powered name tag recorder system, as shown below. Figure 6 As shown, the method includes the following steps S601-S607: Step S601: Obtain the real-time audio data of the emergency scene collected by the first AI badge and the second AI badge.
[0037] In this step, the server acquires in real-time the emergency scene audio data synchronously collected by the first and second AI badges. This audio data consists of mixed audio signals recorded in real-time during the emergency rescue process. Specifically, it includes two types of mixed sound sources: valid human voice audio from medical personnel and various environmental noises inherent in the emergency scene. The first and second AI badges synchronously and directionally collect all the mixed sounds from the scene using their respective main and reference microphones. This ensures that valid voice information and corresponding environmental noise information are completely preserved, providing a complete and original audio data foundation for subsequent audio noise reduction and enhancement processing, as well as voice separation processing.
[0038] Step S602: The audio data from the emergency scene is subjected to noise reduction and enhancement processing and human voice separation processing to obtain the doctor's spoken voice signal and the nurse's repeated voice signal.
[0039] In this step, noise reduction and enhancement processing and human voice separation processing are performed sequentially.
[0040] First, noise reduction and enhancement processing is performed. The acquired audio data from the emergency scene is segmented. From the continuously recorded mixed audio, effective voice segments containing human voices are selected, and silent segments, blank segments, and invalid segments containing only environmental noise are removed. Then, the data collected by the main microphone and the reference microphone are combined to filter out the cluttered background noise at the scene, optimize the playback quality of human voices, and complete the overall audio enhancement and optimization.
[0041] The subsequent voice separation process involves classifying the optimized audio based on the unique device binding relationships of the first and second AI badges and the chronological order of medical order interactions. This process distinguishes between the doctor's spoken voice signal and the nurse's repeating voice signal. This two-stage processing effectively filters audio, optimizes sound quality, and accurately separates the two voice streams, laying a reliable data foundation for subsequent voiceprint recognition, speech-to-text transcription, and other operations.
[0042] The noise reduction and enhancement processing, as well as the voice separation processing, performed on the audio data from the emergency scene rely on the hardware microphone deployment structure of the AI badge in this application. Please refer to... Figure 2 , Figure 3 and Figure 7 , Figure 7 This is a schematic diagram of the microphone distribution for the AI badge. The AI badge includes at least one main microphone facing the target direction and at least one reference microphone facing a non-target direction; the emergency scene audio data collected by the first AI badge and the second AI badge respectively includes multi-channel audio data collected by the main microphone and the reference microphone.
[0043] The AI badge is equipped with a main microphone and a reference microphone. The main microphone is positioned according to the target direction, pointing towards the mouth of the medical staff wearing it. Figure 7 The upper-middle microphone is specifically designed for close-range acquisition of the target voice, including dictation and repetition by medical staff; reference microphones are positioned in other locations on the AI badge, following non-target directions. Figure 7 The left-center microphone, lower microphone, and right microphone are used to collect various environmental background noises at the emergency scene. The directional, zoned layout of the main and reference microphones accurately distinguishes between the target voice acquisition area and the non-target noise acquisition area, providing hardware support for subsequent audio noise reduction and enhancement processing to select effective speech segments and for voice separation processing to differentiate the speech signals corresponding to different devices.
[0044] In a specific implementation, the noise reduction and enhancement processing of the emergency scene audio data collected from any badge includes: acquiring the short-time energy and zero-crossing rate of the emergency scene audio data; identifying and filtering out effective speech segments containing human voices based on the short-time energy and the zero-crossing rate; and separating human voices from noise for the effective speech segments based on the signal difference between the main microphone and the reference microphone to obtain a clean speech signal.
[0045] In this embodiment, before performing noise reduction and enhancement processing on the audio data from the emergency scene, it is necessary to first identify and filter the effective speech segments. In actual emergency scenarios, the original audio collected in real time by the AI badge is a long-term continuous mixed signal with complex and mixed content. It not only contains effective human voices in medical staff conversations, but also contains a large number of silent segments, blank segments, and pure noise segments such as equipment noise and environmental noise. If the entire complete original audio is directly denoised and optimized, it will significantly increase the amount of data processing. Invalid and redundant segments will occupy processing resources and reduce audio processing efficiency. At the same time, silent and meaningless noise segments will also interfere with the accuracy of subsequent noise judgment and signal optimization. There are significant differences between human voices and environmental noise in the numerical range of short-time energy and zero-crossing rate. The human voice audio produced by normal medical staff is stable in signal strength and concentrated in energy, and the overall short-time energy is maintained in a high threshold range; while the signals of silent, blank, or pure environmental noise segments are weak, and the short-time energy is in a low threshold range for a long time. Meanwhile, the waveform of human voice changes smoothly, with a relatively moderate zero-crossing rate and a small fluctuation range, while the zero-crossing rate of chaotic environmental noise and equipment noise will show disordered high-frequency fluctuations. Therefore, by pre-setting reasonable energy thresholds and zero-crossing rate thresholds, comparing the calculation parameters of each audio frame frame by frame, and splicing and integrating audio frames that simultaneously meet the threshold range corresponding to human voice, it is possible to accurately identify and extract effective speech segments containing human voice, and eliminate silent segments, blank segments, and pure interference noise segments without effective human voice, thus completing the accurate selection of audio segments.
[0046] Therefore, this application prioritizes human voice detection through short-time energy and zero-crossing rate, accurately stripping away invalid and redundant segments, and retaining only the effective speech segments containing human voice interaction. This reduces the processing scope and simplifies the computational data, providing a precise and effective processing target for subsequent differentiated noise reduction optimization based on dual microphones.
[0047] Short-time energy is used to characterize the strength of the signal amplitude within a short audio frame, directly reflecting the loudness and signal strength of that segment; the formula for calculating audio short-time energy is: ; In the formula: The short-time energy of the nth frame of audio; The original discrete audio sampling signal is represented by m, where m is the sampling point number, n is the starting sampling point number of the current audio frame, and N is the number of windowed sampling points per frame.
[0048] Zero-crossing rate refers to the number of times a waveform crosses zero level within a single audio frame, used to reflect the frequency density and timbre characteristics of an audio signal. The formula for calculating the audio zero-crossing rate is: ; In the formula, Let be the short-time zero-crossing rate of the nth frame of audio; This is a sign function used to determine the positive or negative polarity of the sampling point; It represents the signal of the previous adjacent sampling point of the m-th sampling point.
[0049] By setting the energy threshold Eth and the zero-crossing rate threshold [Zmin, Zmax], when the audio frame simultaneously satisfies... Zmin≤ When the value is less than or equal to Zmax, the frame is determined to be a valid human voice frame, and consecutive valid frames are spliced together to form a complete valid speech segment.
[0050] Based on the selection of effective speech segments, noise reduction optimization is carried out by relying on the hardware sound recording structure characteristics of the AI badge.
[0051] Because the main microphone and the reference microphone are positioned in different directions, the proportions of human voice and noise components collected by the two microphones in the same emergency rescue scenario are naturally different. This application uses two core features, namely the similarity of the two signal waveforms and the difference in signal strength, to distinguish which is human voice and which is noise. Then, through human voice enhancement and noise suppression processing, a clean speech signal is finally obtained.
[0052] In some embodiments, separating human voice and noise based on the signal difference between the main microphone and the reference microphone to obtain a clean voice signal for the effective voice segment includes: acquiring a first effective voice segment signal collected by the main microphone, and acquiring a second effective voice segment signal collected by the reference microphone, wherein the human voice component collected by the main microphone is stronger than the noise component, and the noise component collected by the reference microphone is stronger than the human voice component; comparing the waveform characteristics and signal intensity of the first effective voice segment signal and the second effective voice segment signal, identifying the parts with consistent waveforms and significant intensity differences as human voice components, and identifying the parts with inconsistent waveforms and similar intensities as noise components; retaining and enhancing the human voice components, weakening or removing the noise components, to obtain a clean voice signal for the corresponding single badge.
[0053] In the specific implementation process, the first effective speech segment signal collected by the main microphone and the second effective speech segment signal collected by the reference microphone are simultaneously acquired at the same time sequence. Directly affected by the microphone installation orientation and pickup distance, the main microphone is directly facing the medical staff's speaking position, picking up the target's spoken voice at close range. The loss during voice propagation is extremely low, and the proportion of the voice component in the acquired signal is much higher than the noise component. The reference microphone is deployed in non-target areas such as the side and bottom of the name tag, far from the core voice source of the medical staff. The voice signal is significantly weakened after spatial propagation attenuation, mainly capturing diffuse environmental noise at the emergency scene, medical equipment operating noise, and other interference signals. The noise component in the acquired signal is generally stronger than the voice component. To accurately quantify the compositional differences between the two signals, a signal modeling formula is used to uniformly represent the two acquired signals. The specific signal modeling formula is as follows: ; In the formula, This represents the first valid speech segment signal captured by the main microphone. This represents the second valid speech segment signal acquired by the reference microphone. The target's original human voice sound wave signal was generated from the on-site oral description of medical staff. This represents the ambient noise signal simultaneously recorded by the main microphone. This represents the ambient noise signal synchronously recorded by the reference microphone. The human voice attenuation coefficient represents the degree of loss of the human voice signal when it propagates to the reference microphone. The human voice attenuation coefficient value corresponding to the reference microphone is much less than 1, which further solidifies the difference in the proportion of human voice and noise components in the two signals and provides a reliable data basis for subsequent sound source component differentiation.
[0054] Based on this, the overall waveform feature matching and real-time signal strength comparison analysis of the two effective speech segments were carried out simultaneously to accurately distinguish the human voice component and noise component in the audio. The human voice produced by the medical staff's narration is a fixed-point, same-source sound signal with strong temporal synchronization and uniform waveform change patterns. The two signals have a high degree of waveform overlap and strong correlation. Moreover, relying on the directional sound gain effect of the main microphone, the amplitude and energy difference of the human voice component in the two signals are significant, which meets the core judgment conditions of waveform consistency and significant intensity difference. Therefore, it is identified as the audio human voice component. On the other hand, various environmental noises at the emergency scene are irregular diffuse sound sources with scattered and disordered propagation directions and no fixed sound emission patterns. The noise waveforms recorded by the two microphones are chaotic and have extremely low matching degree. Moreover, the spatial propagation attenuation of environmental noise is basically consistent. The energy and intensity of the two noise signals are basically similar, which meets the judgment conditions of waveform inconsistency and similar intensity. Therefore, it is identified as the audio noise component.
[0055] To quantify the degree of consistency between two signal waveforms and improve the accuracy of component segmentation, a correlation coefficient calculation formula is introduced to assist in the discrimination. The specific correlation coefficient calculation formula is as follows: ; in, The waveform correlation coefficient represents the correlation coefficient between two valid voice segments, used to visually reflect the degree of overlap and synchronization between the two signals. The covariance represents the difference between the first and second effective speech segments, used to characterize the correlation between the overall trends of the two signals. The variance represents the signal of the first valid speech segment, used to reflect the amplitude of numerical fluctuations in the signal acquired by the main microphone. The variance of the second effective speech segment signal is used to reflect the numerical fluctuation of the signal acquired by the reference microphone. The difference between the human voice component and the noise component can be objectively and accurately determined by the magnitude of the correlation coefficient.
[0056] After accurately classifying and identifying the human voice and noise components, differentiated signal modulation and optimization processing is applied to both types of signal components to filter out noise interference to the maximum extent while fully preserving the original effective information of the human voice. Throughout the processing, the original complete human voice data is retained, and appropriate gain enhancement is applied to the human voice-specific frequency band to improve the strength of the human voice audio signal and speech intelligibility. Simultaneously, the accurately identified and located noise components are attenuated, suppressed, and differentially canceled to significantly weaken or even completely filter out various background interference noises, ensuring that the purity of the human voice is not compromised. This optimization and modulation process is achieved through a unified noise reduction calculation formula, the specific noise reduction formula being: ; in, This represents the clean audio signal of a single badge after final processing. 'k' represents the voice enhancement gain coefficient, which is greater than 1 and is used to specifically amplify and strengthen the voice component signal strength, improving voice clarity. This represents the noise suppression adjustment coefficient, with a value between 0 and 1, used to precisely attenuate and cancel out identified noise components. It still represents the original human voice sound wave signal of the target. The integrated environmental noise signal is represented by a reasonable ratio of human voice enhancement gain coefficient and noise suppression adjustment coefficient to achieve a precise balance between human voice enhancement and noise filtering. This results in a clean voice signal with low interference, high fidelity, and high definition, providing stable and high-quality audio data support for subsequent core business processing steps such as speech transcription and medical order text consistency verification.
[0057] As can be seen, in this embodiment, through the differentiated noise reduction optimization processing based on the signal differences between the main microphone and the reference microphone, the inherent directional sound-receiving hardware structure of the AI badge can accurately distinguish between the effective human voice component and the chaotic environmental noise component at the emergency scene without relying on complex intelligent algorithms. First, reliable separation of human voice and noise is achieved through dual-channel microphone signal modeling and waveform and intensity feature comparison. Then, targeted signal operation optimization through human voice enhancement and noise suppression not only completely preserves the core human voice details of the medical staff's oral medical orders without loss or distortion, but also filters out various invalid background interference noises such as abnormal equipment noises, personnel movement, and environmental noise at the emergency scene to the greatest extent. This effectively solves the problem of mixed audio at the emergency scene and the inaccurate speech transcription and recognition caused by human voices being easily masked by noise. It significantly improves the clarity and signal-to-noise ratio of the processed pure voice signal, providing high-quality and highly reliable audio foundation data for subsequent key steps such as voiceprint authentication, accurate speech transcription, and consistency verification of medical order texts. This ensures the overall accuracy of oral medical order recognition and verification and the safety of emergency treatment from the audio source.
[0058] In some embodiments, the emergency scene audio data collected by the first AI badge and the second AI badge respectively also includes device identification; the voice separation processing includes: acquiring a first clean voice signal after the first AI badge has undergone the noise reduction and enhancement processing, and a second clean voice signal after the second AI badge has undergone the noise reduction and enhancement processing; determining, based on the device identification, that the first clean voice signal is a doctor's voice signal and the second clean voice signal is a nurse's voice signal; performing medical order process verification on the first clean voice signal and the second clean voice signal in conjunction with the fixed dialogue sequence relationship of medical order interaction; when the verification passes, determining that the first clean voice signal is the doctor's dictated voice signal and the second clean voice signal is the nurse's repeating voice signal.
[0059] After the aforementioned noise reduction and enhancement processing, each AI badge, relying on its own main microphone and reference microphone directional sound collection hardware structure, independently performs noise reduction and voice filtering optimization processing on the original mixed audio from the emergency scene that it separately collects, ultimately obtaining a unique pure voice audio signal corresponding to each badge. Since the main microphone of each AI badge is only directed towards the mouth of the medical staff wearing it, the distant voices of non-wearing personnel are treated as environmental noise during the collection process and are uniformly filtered out in the noise reduction and enhancement stage. Therefore, the first pure voice signal obtained after the noise reduction processing of the first AI badge only retains the voice of the doctor wearing the badge and does not contain the voice of the nurse; similarly, the second pure voice signal obtained after the noise reduction processing of the second AI badge only retains the voice of the nurse wearing the badge and does not contain the voice of the doctor. The two pure voice signals are independent, have a single sound source, and do not mix with each other.
[0060] However, in the actual emergency rescue process, the pure human voice content collected and optimized by doctors and nurses through their respective AI badges does not entirely belong to the exclusive interaction of medical order handover. Before, during, and after the rescue operation, medical staff will also have various non-medical order-related daily communication dialogues, such as preoperative preparation communication, on-site work coordination, routine work arrangements, and handling of temporary matters. Although this part of the voice content is valid pure human voice after being filtered and optimized by a single badge, it does not belong to the core business interaction content of doctors dictating medical orders and nurses repeating and verifying medical orders. If all pure human voices are directly included in the subsequent medical order recognition and text verification stages, it is very easy to misidentify non-medical order dialogues and mismatch invalid voices, which will lead to errors in medical order transcription and failure of verification judgment, affecting the accuracy of medical order handover and the safety of emergency operations.
[0061] Therefore, in addition to obtaining two independent pure human voice audio signals, this application also needs to perform voice separation and process verification processing. Based on the unique device identification of each AI badge, the two pure voice signals are first distinguished according to job identity. Then, the compliance verification of the medical order process is carried out in combination with the inherent standard dialogue sequence relationship of medical order interaction. The compliant and valid medical order interaction voice segments are accurately screened and locked from all the communication voices of medical staff, and all non-medical order related irrelevant dialogue interference is eliminated to ensure that subsequent voice transcription and text verification work is only carried out on formal medical order interaction content.
[0062] In practice, each AI badge is pre-programmed with a unique device identification identifier during both factory configuration and on-site deployment. This identifier is uniquely and immutably linked to the medical staff's job title. The first AI badge is bound to the unique identifier for a doctor, and the second badge to the unique identifier for a nurse. These identifiers are also embedded in the emergency audio data collected by the corresponding badge, providing a unique basis for accurate identification of the voice signal's job title. After independent noise reduction and enhancement processing of the two audio streams, the system simultaneously retrieves the first clean voice signal from the first AI badge and the second clean voice signal from the second AI badge. By identifying the unique device identification identifier embedded in the two clean voice signals, the system can quickly and accurately match the job title, directly determining that the first clean voice signal corresponds to a doctor's voice and the second clean voice signal corresponds to a nurse's voice, thus completing the basic correspondence between different job titles.
[0063] Based on this, simply distinguishing between the doctor's and nurse's voices is insufficient to determine valid medical order interactions. This application strictly adheres to clinical medical order handover standards, ensuring that medical order interactions have a fixed and irreversible standard dialogue sequence. The doctor must first verbally issue the medical order, followed by the nurse repeating and verifying it. The dialogue order is fixed, and the interaction process is standardized and uniform. Based on this fixed dialogue sequence, the doctor's and nurse's voice signals, after identity matching, undergo medical order process sequence verification. This verifies whether the order of speech, dialogue intervals, and interaction correspondence between the two voices conform to the formal medical order handover logic, eliminating irrelevant communication voices with disordered sequences, excessively long intervals, or those not belonging to medical order interactions. After verifying that the medical order dialogue sequence is compliant and the interaction process matches, the first clean voice signal is accurately identified as the valid doctor's verbal voice signal, and the second clean voice signal is identified as the valid nurse's repeating voice signal. This completes the precise definition and separation of the specific human voice for the medical order, providing accurate and reliable dedicated voice data support for subsequent medical order voice transcription and text content comparison verification.
[0064] In specific implementations, in some embodiments, the step of verifying the medical order process of the first clean voice signal and the second clean voice signal by combining the fixed dialogue sequence relationship of the medical order interaction includes: pre-configuring the fixed dialogue sequence relationship of the medical order interaction, which includes a fixed interaction logic of the doctor dictating first and the nurse repeating later, and a preset repeating interval; extracting the start and end timestamps of the voice signals corresponding to the doctor's voice signal and the nurse's voice signal respectively, and calculating the interval between the two voice signals; detecting whether the doctor's voice signal and the nurse's voice signal belong to the continuous dialogue content of the same medical order interaction; when the start and end timestamps of the doctor's voice signal are earlier than the start and end timestamps of the nurse's voice signal, and the interval is not greater than the preset repeating interval, the medical order process verification is determined to be successful; when there is a reversal of the voice sequence and the interval exceeds the preset repeating interval, the medical order process verification is determined to be unsuccessful.
[0065] For example, a fixed timing sequence for medical order interaction is pre-configured with the doctor dictating first, followed by the nurse repeating, and a preset interval of 3 seconds for the repeating is set. In actual application, the start and end timestamps of the doctor's voice signal are extracted from 10:05:20 to 10:05:26, and the start and end timestamps of the nurse's voice signal are from 10:05:27 to 10:05:32. The calculated interval between the two voice segments is only 1 second, which satisfies both the doctor's voice timing being earlier and the interval not exceeding the preset threshold. This indicates a continuous dialogue within the same medical order interaction, and the medical order process is deemed to have passed verification. If the nurse's voice timestamp is earlier than the doctor's, or if the interval between the two voice segments exceeds 3 seconds, it indicates an incorrect dialogue timing or an excessively long interval, which does not belong to a standard medical order repeating scenario. In this case, the medical order process is deemed to have failed verification and will not be included in subsequent medical order checks.
[0066] Step S603: Perform voiceprint recognition and matching on the doctor's spoken voice signal according to the pre-stored authorized medical personnel voiceprint feature database, and perform identity authorization determination based on the matching result.
[0067] After confirming the compliance of the doctor's spoken voice signal and the nurse's repeated voice signal, it is necessary to further verify the authorization of the doctor who issued the medical order to ensure that only authorized medical personnel with legal medical practice authority issue medical orders, and to prevent medical safety risks caused by unauthorized personnel issuing medical orders without authorization.
[0068] First, voiceprint features are extracted from the acquired doctor's spoken voice signal. Through framing, windowing, and frequency domain transformation, multi-dimensional personalized voiceprint biometric features, such as Mel spectrum, formant parameters, fundamental frequency period, and timbre energy distribution, are extracted to uniquely characterize an individual's speaking habits. These features form the voiceprint feature vector to be compared for that real-time speech segment. The system pre-builds and stores a voiceprint feature database of authorized medical personnel, pre-entering standard voiceprint feature vectors of all doctors authorized to prescribe medical orders, serving as the benchmark for identity verification.
[0069] After voiceprint feature extraction is completed, the real-time collected voiceprint feature vector of the doctor's spoken speech is compared with the standard voiceprint feature vectors in the authorized medical staff voiceprint feature database one by one using a similarity metric. The voiceprint feature similarity is accurately calculated using the cosine similarity algorithm, and the calculation formula is as follows: ; Where: A is the voiceprint feature vector to be detected extracted from the doctor's current dictated speech; B is the standard voiceprint feature vector pre-stored in the authorized medical staff voiceprint feature database; It is the inner product of two sets of eigenvectors; These are the magnitudes of the corresponding feature vectors; The final calculation result is the matching similarity value of the two sets of voiceprint features. The value range is fixed between 0 and 1. The closer the value is to 1, the higher the overlap of voiceprint features and the stronger the consistency of the speaker.
[0070] The system pre-sets a fixed authorization similarity threshold T and compares the real-time calculated voiceprint similarity with this threshold. When the calculated voiceprint similarity is greater than or equal to the threshold T, the system determines that the current doctor's spoken voice matches the voiceprint of a registered doctor in the authorization database, and the doctor issuing the medical order is a registered and legally authorized person, thus the authorization determination is successful. When the calculated voiceprint similarity is less than the threshold T, it indicates that the real-time voiceprint differs significantly from the standard voiceprint of an authorized person, and a legally authorized doctor cannot be matched. The authorization determination fails, and the doctor is identified as an unauthorized person initiating the medical order. Through the above-mentioned voiceprint feature extraction, quantitative similarity calculation, and threshold comparison, the system can efficiently and accurately complete the identification of the doctor initiating the medical order. Relying on each person's unique biometric voiceprint characteristics, it achieves personnel identity verification, ensuring the legality, uniqueness, and safety of the medical order issuance process.
[0071] After the doctor's identity authorization verification is completed, the corresponding processing procedures are executed according to the verification results: If the identity authorization verification is successful, it means that the doctor who issued the medical order is an authorized person with the authority to issue legal medical orders, and the subsequent medical order voice transcription and text verification process continues; if the identity authorization verification fails, it means that the current voice is an invalid medical order instruction initiated by an unauthorized person, the system directly determines that the voice is an illegal instruction, automatically terminates all subsequent medical order processing procedures, does not perform voice transcription, text generation and content verification operations, and can issue illegal instruction reminders through the sound, light, vibration or voice prompts of the AI badge, so as to realize the automatic interception and security control of unauthorized instructions.
[0072] Step S604: When authorization is granted, the doctor's spoken voice signal is transcribed to generate a medical order text.
[0073] Provided that the doctor's identity authorization is verified, the doctor's dictated voice signal, which has been confirmed to be legal and valid through voiceprint recognition, is input into a preset speech-to-text model. The speech recognition algorithm then analyzes the doctor's dictated voice signal sentence by sentence, performs feature matching and text conversion, and converts the audio-format dictated medical orders into electronic data in text format, automatically generating standardized and structured medical order texts.
[0074] Following the same method steps as in step S604, step S605 is executed to transcribe the nurse's repeated speech signal into a repeat text.
[0075] Step S606: Perform consistency verification on the medical order text and the recitation text. If the verification passes, generate structured medical order data based on the medical order text.
[0076] In some embodiments, the consistency verification of the medical order text and the repeated text includes: performing semantic vector encoding on the medical order text and the repeated text respectively to obtain a medical order semantic vector and a repeated text semantic vector; calculating the similarity between the medical order semantic vector and the repeated text semantic vector to obtain a semantic similarity; extracting emergency medical core entities from the medical order text and the repeated text respectively to obtain a first entity and a second entity, wherein the emergency medical core entity includes at least one of drug name, treatment item, dosage, specification, usage, and frequency; obtaining the entity matching rate through entity comparison; and performing consistency verification based on the semantic similarity and the entity matching rate.
[0077] After generating standardized medical order texts and nurse recitation texts, a comprehensive consistency check needs to be performed on the two texts to ensure a complete match between the issuance and recitation of the medical order, thus avoiding safety risks caused by errors or omissions in critical medical information. This solution no longer relies solely on literal text matching, but adopts a two-layer combined verification mechanism of overall semantic matching and precise verification of core medical entities. This takes into account both the overall consistency of the context and the accuracy of critical emergency information, resulting in more rigorous and reliable verification results, and making it suitable for complex scenarios with high pressure, noise, and frequent colloquial expressions in emergency situations.
[0078] In the specific implementation process, the medical order text and the recitation text are first processed by semantic vector encoding. Using a general text semantic encoding model, the two natural language texts are transformed into digital semantic vectors with unified dimensions, resulting in the medical order semantic vector and the recitation semantic vector, respectively. Subsequently, conventional similarity calculation is performed on the two sets of semantic vectors. This semantic similarity comparison is a common and mature technology in the field of text semantic matching, which can directly quantify the degree of similarity in the overall meaning of the two texts. The higher the similarity value, the higher the overall semantic fit between the two dialogues.
[0079] At the same time, a special extraction of core entities for emergency medical care was carried out for the two texts. All key information was extracted from the medical order text to obtain the first entity set, and corresponding key information was extracted from the recitation text to obtain the second entity set. The core entities for emergency medical care specifically refer to rigid key content that is directly related to medical safety, including at least one type of core field such as drug name, treatment operation items, drug dosage, drug specifications, usage method, and drug frequency.
[0080] After entity extraction is complete, the two sets of core entities are matched and compared item by item to calculate the entity matching rate. The formula for calculating the matching rate is as follows: ; Where: r is the entity matching rate; n is the number of core entities that are completely matched between the two parties; and m is the total number of core entities extracted from the medical order text.
[0081] Finally, by combining the two core indicators of overall semantic similarity and entity matching rate obtained above, a comprehensive consistency verification of the medical order and the repeated text is completed.
[0082] In some embodiments, the consistency verification based on the semantic similarity and the entity matching rate includes: taking a preset semantic similarity threshold and an entity matching rate threshold; comparing the semantic similarity with the semantic similarity threshold and comparing the entity matching rate with the entity matching rate threshold; if the semantic similarity is not lower than the semantic similarity threshold and the entity matching rate is not lower than the entity matching rate threshold, then the verification is deemed to have passed; if the semantic similarity is lower than the semantic similarity threshold and / or the entity matching rate is lower than the entity matching rate threshold, then the verification is deemed to have failed.
[0083] In this embodiment, fixed semantic similarity thresholds and entity matching rate thresholds are pre-configured independently as the judgment benchmarks for overall semantic compliance and key medical information compliance, respectively. The semantic similarity and entity matching rate calculated in real time are then compared with the corresponding thresholds. The judgment rule is that if both conditions are met, the verification passes, and if either condition is not met, the verification fails, thus completing the final consistency result judgment.
[0084] The semantic similarity threshold is set to T1, and the entity matching rate threshold is set to T2. When the real-time semantic similarity is ≥ T1 and the entity matching rate r is ≥ T2, the overall content semantics are consistent and the key medical information is completely matched, and the consistency verification is deemed to have passed. If the semantic similarity is T1, or the entity matching rate r is < T2, or both indicators are lower than the corresponding thresholds, the consistency verification is deemed to have failed.
[0085] This dual verification method can accommodate non-critical differences such as verbal expression and word order adjustment, while accurately identifying mismatches, omissions and deviations in critical information such as medication, dosage and usage. It significantly improves the accuracy and reliability of verifying verbal medical orders in emergency care, and comprehensively ensures medical safety throughout the entire emergency treatment process.
[0086] In some embodiments, the structured medical order data includes at least: device identification information, medical staff identity information, medical order interaction time information, doctor's dictated text content, nurse's repeated text content, drug name, medication dosage, route of administration, diagnostic and treatment procedures, medical order sequence matching results, identity authorization determination results, and consistency verification results of dictated and repeated content.
[0087] In this embodiment, the structured medical order data includes at least the device identification information, medical staff identity information, medical order interaction time information, doctor's dictated text content, nurse's repeated text content, drug name, medication dosage, route of administration, diagnosis and treatment procedures, medical order sequence matching results, identity authorization determination results, and consistency verification results of dictated and repeated content. It covers multiple dimensions such as equipment, personnel, time, diagnosis and treatment content, and verification conclusions of each link, so as to realize the full process of emergency medical orders being traceable and verifiable.
[0088] See Figure 5 Based on actual emergency rescue scenarios, a complete example of structured medical order data is given below: Equipment identification area: D1 (doctor's name tag), D5 (nurse's name tag); Identity information: Doctor (employee ID: ID1), Nurse (employee ID: ID5); Medical order interaction time information: 2026-04-27 15:30:10—2026-04-27 15:30:18; Doctor's dictated text: The patient was given ibuprofen sustained-release capsules 0.2g, orally, twice daily, for continuous symptomatic pain relief treatment; The nurse repeats the text: Administer ibuprofen sustained-release capsules 0.2g orally, twice a day, to the patient for analgesia and symptomatic treatment; Doctor's order details: Drug name: Ibuprofen sustained-release capsules; Dosage: 0.2g; Route of administration: Oral; Treatment procedure: Emergency analgesia and symptomatic treatment; Medical order timing matching result: Validation passed; Identity authorization determination result: Voiceprint matching is qualified, authorization is approved; Consistency check results between spoken and repeated content: Semantic similarity and core entity matching rate both meet the standards, check passed.
[0089] This standardized and structured data format allows for the complete and intuitive retention of all core content and verification results of a single emergency medical order interaction. The data is well-organized and complete, effectively ensuring closed-loop management of verbal medical orders and meeting medical safety retention requirements.
[0090] In some embodiments, after generating structured medical order data based on the medical order text upon successful verification, the method further includes: acquiring medical safety benchmark data, which includes rational drug use data, patient allergy history data, and emergency medical quality control rule data; the rational drug use data includes applicable symptoms, standard dosage, administration route guidelines, and contraindications information; the patient allergy history data includes the patient's previous drug allergy records, allergy reaction types, and a list of contraindicated drugs; the emergency medical quality control rule data includes emergency treatment operation guidelines, medical order writing standards, and medical safety management clauses; and, based on the medical safety benchmark data, processing the structured medical order data... The system performs multi-dimensional medical safety compliance verification on medical order data. This verification includes checking the matching of treatment indications, the compatibility of medication dosage with patient vital signs, screening for drug allergy / conflict risks, and verifying the completeness of core fields entered in the medical order. When all verification items pass, the system executes the step "sending the medical order data to the execution terminal for nurses to perform corresponding medical operations." When at least one verification item fails, a corresponding medical risk warning is automatically generated. This risk warning is encrypted and pushed to the first AI badge and the second AI badge. Warning prompts are simultaneously presented to medical staff via voice broadcast and / or screen pop-ups.
[0091] After the consistency verification between the medical order text and the repeated text is passed and structured medical order data is generated, the system further introduces a medical safety compliance verification mechanism. Based on standardized medical benchmark data, it conducts multi-dimensional risk verification of the medical order content, identifies various medical safety hazards in advance, and improves the standardization and safety of emergency verbal medical orders.
[0092] The system retrieves pre-set medical safety benchmark data, which mainly includes rational drug use data, patient allergy history data, and emergency medical quality control rule data. Rational drug use data covers the applicable symptoms, standard dosage, administration route, and contraindications of drugs; patient allergy history data records past allergies, types of allergic reactions, and a list of prohibited drugs; and emergency medical quality control rule data includes emergency operation procedures, prescription standards, and safety management clauses, providing a unified basis for compliance verification.
[0093] Based on the aforementioned benchmark data, structured medical order data undergoes medical safety compliance verification, specifically including verification of treatment indications, medication dosage appropriateness, drug allergy risk screening, and integrity verification of core fields in the medical order. When all verification items pass, the medical order content is deemed compliant, and the structured medical order data is sent to the execution terminal for nurses to perform corresponding medical operations; if any verification fails, corresponding medical risk warning information is automatically generated.
[0094] The system encrypts the risk warning information and pushes it simultaneously to the first AI badge and the second AI badge. It then displays the warning prompts to medical staff through at least one method, such as voice broadcast or screen pop-up, so that medical staff can promptly discover potential problems with medical orders, verify and correct medical orders, realize early warning and timely intervention of medical risks, and effectively avoid adverse medical events such as irrational drug use, allergic reactions, and operational violations.
[0095] Step S607: The medical order data is sent to the execution terminal so that the nurse can perform the corresponding medical operation.
[0096] Please see Figure 8 , Figure 8 This is a schematic diagram illustrating the complete process of intelligent processing of verbal medical orders based on an AI badge recorder system, such as... Figure 8 As shown, S801, collect audio data from the emergency scene; S802, perform noise reduction and enhancement processing on the audio and select effective human voice segments; S803, perform dual-microphone signal comparison processing, separate human voice from environmental noise, and generate corresponding clean speech signals; S804, based on the device identification information, distinguish between the doctor's clean speech and the nurse's clean speech; S805, extract the timestamps corresponding to the two speech segments respectively, and calculate the interval duration of medical order interaction.
[0097] S806. Perform time-series verification of medical orders; if the time-series verification fails, proceed directly to the process end node, and the processing of this medical order is terminated; if the time-series verification passes, proceed to S807. Doctor's voiceprint feature matching and comparison. S808. Perform identity authorization determination; if the identity authorization determination fails, proceed directly to the process end node, and the processing of the medical order is terminated; if the identity authorization determination passes, proceed to S809. Transcribe the doctor's spoken voice and the nurse's repeated voice separately to generate the medical order text and the repeated text.
[0098] S810. Perform semantic vector encoding and entity matching calculation on the medical order text and the repeated text to obtain semantic similarity and entity matching rate; S811. Perform text consistency verification; if the text consistency verification fails, directly enter the process end node and the medical order processing is terminated; if the text consistency verification passes, execute S812. Generate structured medical order data based on the medical order text.
[0099] S813. Retrieve pre-stored medical safety benchmark data; S814. Conduct multi-dimensional medical safety compliance verification on structured medical order data; S815. Determine the medical safety results; If all verification items pass the verification, proceed to S816. Send the medical order data to the execution terminal, and then enter the process end node; If any verification item fails the verification, proceed to S817. Automatically generate medical risk warning information.
[0100] S818. Encrypt and push medical risk warning information to the AI name tags worn by doctors and nurses; S819. The AI name tags present risk warnings through voice broadcast and screen pop-up windows. Medical staff correct medical orders based on the warning prompts, and the process is temporarily suspended while waiting for the medical orders to be re-verified and processed.
[0101] As can be seen, in this embodiment, by acquiring real-time audio data from the emergency scene collected by the AI badges worn by doctors and nurses, noise reduction and enhancement processing and voice separation processing are performed on the audio data to accurately obtain the doctor's spoken voice signal and the nurse's repeated voice signal; voiceprint recognition and matching are performed on the doctor's spoken voice signal through a pre-stored authorized medical personnel voiceprint feature database, and identity authorization is determined based on the matching results. After authorization is passed, the two voice signals are transcribed into speech to generate corresponding medical order text and repeated text; consistency verification is performed on the two types of text, and after verification, structured medical order data is generated based on the medical order text and sent to the execution terminal to guide the nurse to complete the corresponding medical operation. This application relies on a complete implementation process that integrates intelligent audio processing, identity verification, automatic text comparison and validation, and standardized generation and distribution of medical order data. It reduces the drawbacks of manual memorization, manual verification, and post-event paper recording, lowers the risk of deviations in the transmission of medical order information and errors in recording in complex emergency situations, improves the accuracy of verbal medical order verification and the efficiency of information transmission, standardizes the flow of verbal medical orders in emergency situations, enhances the safety and full traceability of diagnosis and treatment operations, and adapts to the actual application needs of rapid treatment in emergency scenarios.
[0102] The above primarily describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the server includes the corresponding hardware structure and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0103] This application embodiment can divide the server into functional units according to the above method example. For example, each function can be divided into different functional units, or two or more functions can be integrated into one processing module. The integrated unit can be implemented in hardware or as a software program module. It should be noted that the unit division in this application embodiment is illustrative and only represents a logical functional division, while other division methods may be used in actual implementation.
[0104] In the case of using integrated units, please refer to Figure 9 , Figure 9 This is a functional unit structure diagram of an intelligent processing device for oral medical orders based on an AI badge recorder system. The intelligent processing device 9 for oral medical orders based on the AI badge recorder system includes: The acquisition unit 901 is used to acquire real-time audio data of the emergency scene collected by the first AI badge and the second AI badge; Processing unit 902 is used to perform noise reduction and enhancement processing and voice separation processing on the emergency scene audio data to obtain the doctor's spoken voice signal and the nurse's repeated voice signal; to perform voiceprint recognition and matching on the doctor's spoken voice signal according to a pre-stored authorized medical personnel voiceprint feature database, and to perform identity authorization determination based on the matching result; if authorization is successful, the doctor's spoken voice signal is transcribed into speech to generate medical order text; the nurse's repeated voice signal is transcribed into speech to generate repeated text; and consistency verification is performed on the medical order text and the repeated text. If the verification is successful, structured medical order data is generated based on the medical order text. The sending unit 903 is used to send the medical order data to the execution terminal so that the nurse can perform the corresponding medical operation.
[0105] As can be seen, in this embodiment, by acquiring real-time audio data from the emergency scene collected by the AI badges worn by doctors and nurses, noise reduction and enhancement processing and voice separation processing are performed on the audio data to accurately obtain the doctor's spoken voice signal and the nurse's repeated voice signal; voiceprint recognition and matching are performed on the doctor's spoken voice signal through a pre-stored authorized medical personnel voiceprint feature database, and identity authorization is determined based on the matching results. After authorization is passed, the two voice signals are transcribed into speech to generate corresponding medical order text and repeated text; consistency verification is performed on the two types of text, and after verification, structured medical order data is generated based on the medical order text and sent to the execution terminal to guide the nurse to complete the corresponding medical operation. This application relies on a complete implementation process that integrates intelligent audio processing, identity verification, automatic text comparison and validation, and standardized generation and distribution of medical order data. It reduces the drawbacks of manual memorization, manual verification, and post-event paper recording, lowers the risk of deviations in the transmission of medical order information and errors in recording in complex emergency situations, improves the accuracy of verbal medical order verification and the efficiency of information transmission, standardizes the flow of verbal medical orders in emergency situations, enhances the safety and full traceability of diagnosis and treatment operations, and adapts to the actual application needs of rapid treatment in emergency scenarios.
[0106] In some embodiments, the AI badge includes at least one main microphone facing a target direction and at least one reference microphone facing a non-target direction; the emergency scene audio data collected by the first AI badge and the second AI badge respectively includes multi-channel audio data collected by the main microphone and the reference microphone; the noise reduction and enhancement processing of the emergency scene audio data collected by any badge includes: acquiring the short-time energy and zero-crossing rate of the emergency scene audio data; identifying and filtering effective speech segments containing human voices based on the short-time energy and the zero-crossing rate; and separating human voices from noise based on the signal difference between the main microphone and the reference microphone to obtain a clean speech signal for the effective speech segments.
[0107] In some embodiments, in separating human voice and noise based on the signal difference between the main microphone and the reference microphone to obtain a clean voice signal for the effective voice segment, the processing unit 902 is configured to: acquire a first effective voice segment signal collected by the main microphone, and acquire a second effective voice segment signal collected by the reference microphone, wherein the human voice component collected by the main microphone is stronger than the noise component, and the noise component collected by the reference microphone is stronger than the human voice component; compare the waveform characteristics and signal intensity of the first effective voice segment signal and the second effective voice segment signal, identify the parts with consistent waveforms and significant intensity differences as human voice components, and identify the parts with inconsistent waveforms and similar intensities as noise components; retain and enhance the human voice components, weaken or remove the noise components, and obtain a clean voice signal for the corresponding single badge.
[0108] In some embodiments, the emergency scene audio data collected by the first AI badge and the second AI badge respectively also includes device identification; the voice separation processing includes: acquiring a first clean voice signal after the first AI badge has undergone the noise reduction and enhancement processing, and a second clean voice signal after the second AI badge has undergone the noise reduction and enhancement processing; determining, based on the device identification, that the first clean voice signal is a doctor's voice signal and the second clean voice signal is a nurse's voice signal; performing medical order process verification on the first clean voice signal and the second clean voice signal in conjunction with the fixed dialogue sequence relationship of medical order interaction; when the verification passes, determining that the first clean voice signal is the doctor's dictated voice signal and the second clean voice signal is the nurse's repeating voice signal.
[0109] In some embodiments, in verifying the medical order process of the first clean voice signal and the second clean voice signal in conjunction with the fixed dialogue sequence of the medical order interaction, the processing unit 902 is configured to: pre-configure the fixed dialogue sequence of the medical order interaction, which includes a fixed interaction logic of doctor's speech first and nurse's recitation followed by a preset recitation interval; extract the start and end timestamps of the voice signals corresponding to the doctor's voice signal and the nurse's voice signal respectively, and calculate the interval between the two voice signals; detect whether the doctor's voice signal and the nurse's voice signal belong to the continuous dialogue content of the same medical order interaction; determine that the medical order process verification passes when the start and end timestamps of the doctor's voice signal are earlier than the start and end timestamps of the nurse's voice signal and the interval is not greater than the preset recitation interval; determine that the medical order process verification fails when there is a reversal of the voice sequence and the interval exceeds the preset recitation interval.
[0110] In some embodiments, in performing consistency verification on the medical order text and the repeated text, the processing unit 902 is configured to: perform semantic vector encoding processing on the medical order text and the repeated text respectively to obtain a medical order semantic vector and a repeated text semantic vector; calculate the similarity between the medical order semantic vector and the repeated text semantic vector to obtain a semantic similarity; extract emergency medical core entities from the medical order text and the repeated text respectively to obtain a first entity and a second entity, wherein the emergency medical core entity includes at least one of drug name, treatment item, dosage, specification, usage, and frequency; obtain the entity matching rate through entity comparison; and perform consistency verification based on the semantic similarity and the entity matching rate.
[0111] In some embodiments, in performing consistency verification based on the semantic similarity and the entity matching rate, the processing unit 902 is configured to: obtain a preset semantic similarity threshold and an entity matching rate threshold; compare the semantic similarity with the semantic similarity threshold and compare the entity matching rate with the entity matching rate threshold; if the semantic similarity is not lower than the semantic similarity threshold and the entity matching rate is not lower than the entity matching rate threshold, then the verification is determined to pass; if the semantic similarity is lower than the semantic similarity threshold and / or the entity matching rate is lower than the entity matching rate threshold, then the verification is determined to fail.
[0112] In some embodiments, the structured medical order data includes at least: device identification information, medical staff identity information, medical order interaction time information, doctor's dictated text content, nurse's repeated text content, drug name, medication dosage, route of administration, diagnostic and treatment procedures, medical order sequence matching results, identity authorization determination results, and consistency verification results of dictated and repeated content.
[0113] In some embodiments, after generating structured medical order data based on the medical order text upon successful verification, the processing unit 902 is further configured to: acquire medical safety benchmark data, which includes rational drug use data, patient allergy history data, and emergency medical quality control rule data; the rational drug use data includes applicable drug symptoms, standard dosage, administration route guidelines, and contraindication information; the patient allergy history data includes the patient's previous drug allergy records, allergy reaction types, and a list of contraindicated drugs; the emergency medical quality control rule data includes emergency treatment operation guidelines, medical order writing standards, and medical safety management clauses; and, based on the medical safety benchmark data, process the medical order data. The structured medical order data undergoes multi-dimensional medical safety compliance verification. This verification includes checking the matching of treatment indications, the compatibility of medication dosage with patient vital signs, screening for drug allergy / conflict risks, and verifying the completeness of core fields entered in the medical order. When all verification items pass, the step "send the medical order data to the execution terminal for nurses to perform corresponding medical operations" is executed. When at least one verification item fails, corresponding medical risk warning information is automatically generated. This risk warning information is encrypted and pushed to the first AI badge and the second AI badge. Warning prompts are simultaneously presented to medical staff via voice broadcast and / or screen pop-ups.
[0114] Please see Figure 10 , Figure 10 This application provides a schematic diagram of the structure of a server, as shown in the embodiment of the present application. Figure 10 As shown, the server 10 includes a processor 1001, a memory 1003, a communication interface 1002, and a computer program 10031. The computer program 10031 is stored in the memory 1003 and configured to be executed by the processor 1001. The program includes a method for performing an intelligent processing method for oral medical orders based on an AI badge recorder system as described in the above embodiments.
[0115] This application provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implement the steps of any possible embodiment of the method.
[0116] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0117] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0118] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0119] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0120] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0121] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0122] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0123] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for intelligent processing of verbal medical orders based on an AI badge recorder system, characterized in that, A server is used in the AI badge recorder system, the AI badge recorder system including a first AI badge, a second AI badge, the server, and an execution terminal, wherein the first AI badge is worn by a doctor and the second AI badge is worn by a nurse, and the method includes: The system acquires real-time audio data from the first AI badge and the second AI badge; the audio data from the first AI badge and the second AI badge also includes device identification. The audio data from the emergency scene is subjected to noise reduction and enhancement processing to obtain a first clean speech signal after the first AI badge has undergone the noise reduction and enhancement processing, and a second clean speech signal after the second AI badge has undergone the noise reduction and enhancement processing. Based on the device identification, the first clean voice signal is determined to be a doctor's voice signal, and the second clean voice signal is determined to be a nurse's voice signal. A fixed dialogue sequence for the medical order interaction is pre-configured, including a fixed interaction logic where the doctor speaks first and the nurse repeats later, and a preset repeating interval. The start and end timestamps corresponding to the doctor's voice signal and the nurse's voice signal are extracted respectively, and the interval between the two voice signals is calculated. When it is detected that the start and end timestamps of the doctor's voice signal are earlier than the start and end timestamps of the nurse's voice signal, and the interval is not greater than the preset repeating interval, the first clean voice signal is determined to be the doctor's spoken voice signal, and the second clean voice signal is the nurse's repeated voice signal. The doctor's spoken voice signal is matched with a pre-stored authorized medical personnel voiceprint feature database, and the identity authorization is determined based on the matching result. Once authorization is granted, the doctor's spoken voice signal will be transcribed to generate a medical order text. The nurse's repeated voice signal is transcribed to generate a repeated text. A consistency check is performed on the medical order text and the repeated text. If the check passes, structured medical order data is generated based on the medical order text. The medical order data is sent to the execution terminal so that nurses can perform the corresponding medical operations.
2. The method according to claim 1, characterized in that, The AI badge includes at least one main microphone facing the target direction and at least one reference microphone facing a non-target direction; the emergency scene audio data collected by the first AI badge and the second AI badge respectively includes multi-channel audio data collected by the main microphone and the reference microphone; The noise reduction and enhancement processing of the first aid scene audio data collected from any name tag includes: Acquire the short-time energy and zero-crossing rate of the emergency scene audio data; identify and filter out effective speech segments containing human voices based on the short-time energy and the zero-crossing rate; For the effective speech segment, based on the signal difference between the main microphone and the reference microphone, the human voice and noise are separated to obtain a clean speech signal.
3. The method according to claim 2, characterized in that, The step of separating human voice from noise based on the signal difference between the main microphone and the reference microphone for the effective speech segment to obtain a clean speech signal includes: Acquire a first valid speech segment signal collected by the main microphone, and acquire a second valid speech segment signal collected by the reference microphone, wherein the human voice component collected by the main microphone is stronger than the noise component, and the noise component collected by the reference microphone is stronger than the human voice component. The waveform features and signal intensity of the first effective speech segment signal and the second effective speech segment signal are compared. The parts with consistent waveforms and significant differences in intensity are identified as human voice components, and the parts with inconsistent waveforms and similar intensities are identified as noise components. The human voice component is preserved and enhanced, while the noise component is weakened or removed to obtain a clean voice signal for the corresponding single badge.
4. The method according to claim 1, characterized in that, The consistency check between the medical order text and the repeated text includes: Semantic vector encoding is performed on the medical order text and the repeated text to obtain a medical order semantic vector and a repeated semantic vector; similarity is calculated between the medical order semantic vector and the repeated semantic vector to obtain semantic similarity. The emergency medical core entities are extracted from the medical order text and the repeated text respectively to obtain the first entity and the second entity. The emergency medical core entity includes at least one of the following: drug name, treatment item, dosage, specification, usage, and frequency. The entity matching rate is obtained by entity comparison. Consistency verification is performed based on the semantic similarity and the entity matching rate.
5. The method according to claim 4, characterized in that, The consistency verification based on the semantic similarity and the entity matching rate includes: Obtain the preset semantic similarity threshold and entity matching rate threshold; The semantic similarity is compared with the semantic similarity threshold, and the entity matching rate is compared with the entity matching rate threshold. If the semantic similarity is not lower than the semantic similarity threshold and the entity matching rate is not lower than the entity matching rate threshold, then the verification is deemed successful. If the semantic similarity is lower than the semantic similarity threshold, and / or the entity matching rate is lower than the entity matching rate threshold, then the verification is deemed to have failed.
6. The method according to claim 1, characterized in that, The structured medical order data includes at least: Equipment identification information, medical staff identity information, medical order interaction time information, doctor's dictated text content, nurse's repeated text content, drug name, medication dosage, route of administration, diagnosis and treatment procedures, medical order sequence matching results, identity authorization determination results, and consistency verification results of dictated and repeated content.
7. The method according to claim 1, characterized in that, After generating structured medical order data based on the medical order text upon successful verification, the method further includes: Acquire medical safety benchmark data, which includes rational drug use data, patient allergy history data, and emergency medical quality control rule data. The rational drug use data includes the applicable diseases of the drugs, standard dosage, administration route specifications, and contraindications. The patient allergy history data includes the patient's previous drug allergy records, allergy reaction types, and a list of contraindicated drugs. The emergency medical quality control rule data includes emergency diagnosis and treatment operation specifications, medical order standards, and medical safety management clauses. Based on the aforementioned medical safety benchmark data, the structured medical order data undergoes multi-dimensional medical safety compliance verification; the medical safety compliance verification includes verification of the matching of treatment indications, verification of the compatibility between medication dosage and patient vital signs, screening of drug allergy and conflict risks, and verification of the completeness of core fields filled in the medical order. When all verification items pass the test, the step "send the medical order data to the execution terminal for the nurse to perform the corresponding medical operation" is executed. When at least one verification item fails the test, a corresponding medical risk warning message will be automatically generated. The risk warning information is encrypted and pushed to the first AI badge and the second AI badge; Warning prompts are presented to medical staff simultaneously via voice broadcast and / or screen pop-ups.
8. A smart device for processing verbal medical orders based on an AI badge recorder system, characterized in that, The AI badge recorder system includes a first AI badge, a second AI badge, a server, and an execution terminal. The first AI badge is worn by a doctor, and the second AI badge is worn by a nurse. The server is used to execute the step instructions in the method as described in any one of claims 1-7. The device includes: The acquisition unit is used to acquire real-time audio data of the emergency scene collected by the first AI badge and the second AI badge; The processing unit is used to perform noise reduction and enhancement processing and voice separation processing on the emergency scene audio data to obtain the doctor's spoken voice signal and the nurse's repeated voice signal; to perform voiceprint recognition and matching on the doctor's spoken voice signal according to a pre-stored authorized medical personnel voiceprint feature database, and to perform identity authorization determination based on the matching result; if authorization is successful, the doctor's spoken voice signal is transcribed into speech to generate medical order text; the nurse's repeated voice signal is transcribed into speech to generate repeated text; and consistency verification is performed on the medical order text and the repeated text. If the verification is successful, structured medical order data is generated based on the medical order text. The sending unit is used to send the medical order data to the execution terminal so that the nurse can perform the corresponding medical operation.