An operating room voice control system and method
By combining microphone array sound source localization and a two-stage intelligent wake-up module, the problems of low wake-up accuracy and poor noise reduction effect of the operating room voice control system in complex acoustic environments are solved, and efficient voice control in noisy environments is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHENZHOU NO 1 PEOPLES HOSPITAL
- Filing Date
- 2026-06-18
- Publication Date
- 2026-07-31
AI Technical Summary
Existing operating room voice control systems suffer from low wake-up accuracy and poor noise reduction in complex acoustic environments. They are unable to effectively distinguish between valid speech and spatial sources of interference noise, resulting in frequent missed wake-ups and false wake-ups, which affects the smoothness of surgical operations.
A microphone array is used for three-dimensional sound source localization. Combined with adaptive noise reduction and a two-stage intelligent wake-up module, multiple voice signals are collected through the microphone array. The three-dimensional sound source localization is used to determine the voice signals in the effective area for adaptive noise reduction. The two-stage intelligent wake-up module performs weighted fusion of wake-up word confidence and control command probability to determine whether to wake up the system.
It significantly improves the wake-up accuracy and reliability of the system in complex acoustic environments, effectively filters out various complex spatial interferences in the operating room, provides a stable interactive basis for voice control in the operating room, and reduces the frequency of missed wake-up and false wake-up.
Smart Images

Figure CN122493849A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a voice control system and method for operating rooms. Background Technology
[0002] In modern operating rooms, medical staff frequently operate various medical devices such as surgical lights, operating tables, anesthesia machines, and monitors. Traditional manual operation not only interrupts the surgical process but also poses a risk of cross-infection. Voice control, as a non-contact interaction method, can effectively solve these problems and has therefore received widespread attention and application in the medical field.
[0003] However, the operating room environment is highly acoustically complex, with various interfering factors such as the clanging of surgical instruments, alarms from monitors, the operation of anesthesia machines, and conversations among medical staff. Existing operating room voice control systems mainly suffer from the following two shortcomings:
[0004] On the one hand, existing systems generally use a single wake-word confidence threshold judgment mechanism. When the threshold is set too high, it is easy to miss wake-ups due to environmental noise interference, and medical staff need to repeat the command multiple times; when the threshold is set too low, false wake-ups will occur frequently, interfering with normal surgical operations. This contradiction is particularly prominent in noisy operating room environments.
[0005] On the other hand, most existing systems employ single-microphone noise reduction or general beamforming techniques without spatial differentiation, which can only suppress noise in the frequency dimension and cannot distinguish between the spatial sources of valid speech and interfering noise. They cannot effectively suppress various interference signals from non-surgical areas, leading to a large amount of noise being mistakenly treated as valid speech, severely reducing the accuracy of subsequent command recognition. Summary of the Invention
[0006] To overcome the problems mentioned in the background art, the present invention proposes an operating room voice control system and method.
[0007] The technical solution of this invention is: an operating room voice control system, comprising:
[0008] Microphone array module for real-time acquisition of multiple audio signals in the operating room environment;
[0009] The sound source localization and preprocessing module is electrically connected to the microphone array module. It is used to perform three-dimensional localization of the sound source of multiple voice signals, determine whether the sound source is located within the preset effective surgical area based on the localization result, and perform adaptive noise reduction processing on the voice signal within the effective area and suppress the signal outside the effective area.
[0010] The dual-stage intelligent wake-up module is electrically connected to the sound source localization and preprocessing module. It is used to first identify the wake-up word in the preprocessed voice signal and output the wake-up word confidence level. When the wake-up word confidence level is lower than the preset high threshold but higher than the preset low threshold, it continues to collect and identify subsequent voice signals. It combines the wake-up word confidence level and the probability that the subsequent voice is a control command to determine whether to wake up the system.
[0011] The voice command recognition module is electrically connected to the dual-stage intelligent wake-up module and is used to fully recognize and semantically understand valid voice commands after the system is woken up.
[0012] The equipment control interface module is electrically connected to the voice command recognition module and is used to send control commands to the corresponding operating room medical equipment based on the semantic understanding results.
[0013] The feedback module is electrically connected to the voice command recognition module and the device control interface module, and is used to provide voice or visual feedback on the operation status to medical staff.
[0014] A voice control method for operating rooms includes the following steps:
[0015] S1: Real-time acquisition of multiple audio signals from the operating room environment via a microphone array;
[0016] S2: Perform three-dimensional localization of the sound source for the acquired multi-channel speech signals, calculate the spatial coordinates of the sound source, and determine whether the sound source is located within the preset effective surgical area;
[0017] S3: If the sound source is located within the effective surgical area, perform adaptive noise reduction processing on the speech signal in that direction; if the sound source is located in a non-effective area, identify it as noise and completely suppress it, then return to step S1.
[0018] S4: Perform wake-up word recognition on the noise-reduced speech signal, calculate and output the wake-up word confidence score;
[0019] S5: Based on the confidence level of the wake word, make a graded judgment to determine whether to wake up directly, reject directly, or enter a two-stage comprehensive judgment process;
[0020] S6: When entering the dual-stage comprehensive judgment process, continue to collect voice signals of the subsequent preset duration, pre-identify the control commands and calculate the probability that the voice is a control command, and weight and fuse the wake-up word confidence and the control command probability to obtain a comprehensive score, and determine whether to wake up the system based on the comprehensive score;
[0021] S7: After the system is woken up, it fully recognizes and understands the semantics of the subsequent voice commands and generates corresponding device control commands.
[0022] S8: Sends control commands to the corresponding operating room medical equipment via the device control interface to execute the corresponding operations;
[0023] S9: Provide feedback to medical staff regarding the success or failure of the operation through the feedback module.
[0024] Preferably, the microphone array module adopts a 6-8 element linear array or ring array, and is installed in the center of the operating room ceiling or on the operating light bracket, and its coverage area includes the entire surgical operation area; the preset effective surgical area is a cylindrical space area with a radius of 1.5-2.5 meters centered on the operating table.
[0025] Preferably, the three-dimensional localization of the sound source in step S2 specifically includes the following steps:
[0026] S21: Perform cross-correlation calculation on each audio signal to obtain the time difference of the audio signal arriving at different microphone array elements;
[0027] S22: Based on the geometric parameters of the microphone array and the calculated time difference, establish and solve a set of spatial coordinate equations to obtain the initial three-dimensional coordinates of the sound source;
[0028] S23: The initial three-dimensional coordinates are smoothed using the Kalman filter algorithm to eliminate localization noise and obtain the final sound source localization result.
[0029] Preferably, the adaptive noise reduction process described in step S3 specifically includes the following steps:
[0030] S31: Based on the sound source localization result, activate the adaptive beamformer in the corresponding direction and use the minimum variance distortionless response algorithm to enhance the speech signal in the target direction;
[0031] S32: Perform spectral subtraction post-denoising processing on the beamforming enhanced speech signal to remove residual stable background noise and obtain the preprocessed speech signal.
[0032] Preferably, the grading determination in step S5 specifically includes the following steps:
[0033] S51: Compare the calculated wake word confidence with the preset high and low thresholds;
[0034] S52: If the confidence level of the wake word is higher than the preset high threshold, it is directly determined to be a valid wake-up signal, the system is woken up and proceeds to step S7;
[0035] S53: If the confidence level of the wake word is lower than the preset low threshold, it is directly determined to be a non-wake signal, and the process returns to step S1.
[0036] S54: If the confidence level of the wake word is between the preset low threshold and the high threshold, it is determined to be a suspected wake-up signal, and the two-stage comprehensive judgment process in step S6 is entered.
[0037] Preferably, the two-stage comprehensive judgment process described in step S6 specifically includes the following steps:
[0038] S61: Continue to collect the voice signal for the next 1-3 seconds and perform the same preprocessing operation as in step S3;
[0039] S62: A lightweight speech recognition model is used to pre-identify control commands in the pre-processed subsequent speech signal, and the probability value of the speech being a commonly used control command in the operating room is output.
[0040] S63: Based on the preset weighting coefficients, the wake-up word confidence and the control command probability are weighted and fused to calculate the comprehensive wake-up score;
[0041] S64: Compare the overall wake-up score with the preset wake-up threshold. If the overall score is higher than the wake-up threshold, wake up the system and proceed to step S7. Otherwise, determine it as a non-wake-up signal and return to step S1.
[0042] Preferably, the weighted fusion in step S63 is calculated using the following formula:
[0043] ;
[0044] in, and For the weighting coefficients, satisfying ,and , The value range is 0.6-0.8. The value range is 0.2-0.4, and the preset wake-up threshold ranges from 0.65-0.75.
[0045] Preferably, the lightweight speech recognition model described in step S62 is trained only on commonly used control command keywords in the operating room scenario. These commonly used control command keywords include, but are not limited to, raising, lowering, opening, closing, moving left, moving right, moving up, and moving down.
[0046] Preferably, the device control interface described in step S8 supports at least one of the communication protocols RS232, RS485, Ethernet and Bluetooth, and is able to communicate with at least one of the operating room medical devices, including surgical lights, operating tables, anesthesia machines, monitors and electrosurgical units.
[0047] The beneficial effects of this invention are:
[0048] 1. Compared to existing technologies that use a single wake-word confidence threshold, this approach presents inherent contradictions in practical applications. When the threshold is set too high, environmental noise can easily cause a decrease in wake-word confidence, leading to frequent missed wake-ups. Medical staff must repeatedly use the wake-word to trigger the system, impacting the smoothness of surgical procedures. Conversely, when the threshold is set too low, numerous irrelevant sounds similar to the wake-word pronunciation are misinterpreted as valid wake-up signals, resulting in frequent false wake-ups and disrupting normal surgical procedures. This deficiency is particularly pronounced in the acoustically complex environment of an operating room. This invention employs a two-stage confidence fusion intelligent wake-up scheme. First, wake-word recognition is performed on the speech signal, and a confidence score is output. For suspected wake-up signals with confidence scores in the middle range, instead of directly waking or rejecting them, subsequent speech signals are collected. A lightweight model identifies the likelihood that the subsequent speech is an operating room control command. Then, the wake-word confidence score and the control command probability are weighted and fused to obtain a comprehensive score. Finally, the system is awakened based on the comprehensive score. This solution can effectively balance the contradiction between missed wake-up and false wake-up without sacrificing system response speed, significantly improve the wake-up accuracy and reliability of the system in complex acoustic environments, and provide a stable interactive foundation for voice control in the operating room.
[0049] 2. Compared to existing technologies that use single-microphone noise reduction or general beamforming noise reduction schemes without spatial differentiation, these schemes can only suppress noise in the frequency dimension and cannot distinguish between valid speech and spatial sources of interfering noise. They cannot effectively suppress various interferences from non-surgical areas in the operating room, such as the sounds of surgical instruments colliding, monitor alarms, anesthesia machine operation, and casual conversations among medical staff. This leads to a large number of interference signals being mistakenly treated as valid speech, severely reducing the accuracy of subsequent voice command recognition. This invention employs a regionalized adaptive noise reduction scheme based on microphone array sound source localization. It collects multiple speech signals through a microphone array, calculates the three-dimensional spatial coordinates of the sound source using a time-of-arrival algorithm, and pre-defines the effective area for medical staff to perform surgical operations. Only signals from sound sources within this area are considered potentially valid speech. The system then performs adaptive beamforming enhancement and post-noise reduction processing on the speech signals in that direction. All signals from non-effective areas are directly identified as environmental noise and completely suppressed. This solution achieves precise noise isolation and effective directional speech enhancement in the spatial dimension, significantly improving the targeting and effectiveness of noise reduction processing. It can effectively filter out various complex spatial interferences in the operating room, providing high-quality voice input for subsequent wake-up and command recognition. Attached Figure Description
[0050] Figure 1 The diagram shown is a structural schematic of the operating room voice control system of the present invention.
[0051] Figure 2 The diagram shown is a flowchart of the operating room voice control method of the present invention.
[0052] Figure 3 The flowchart shown is a flowchart of the wake-word confidence level judgment in the operating room voice control method of the present invention;
[0053] Figure 4 The flowchart shown is a two-stage integrated judgment process in the operating room voice control method of the present invention.
[0054] Figure 5 The diagram shows the arrangement of the microphone array and the effective area of the operating room in the operating room voice control method of the present invention. Detailed Implementation
[0055] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0056] Please see Figure 1 and Figure 2 This invention provides an embodiment: an operating room voice control system and method, aiming to solve the problems of low wake-up accuracy and poor noise reduction effect of existing operating room voice control systems in complex acoustic environments. The following is a comprehensive description of the technical solution of this invention.
[0057] I. System Overall Architecture
[0058] The operating room voice control system of this invention adopts a modular design, with data transmission between modules via a high-speed serial bus or Ethernet. The overall system latency is controlled within 500ms, meeting the requirements of real-time operation in the operating room. The system specifically includes the following modules:
[0059] 1. Microphone array module
[0060] The microphone array module is the front end for voice signal acquisition in the system. It adopts a 6-8 element linear array or a circular array. In this embodiment, an 8-element circular microphone array is preferred. The array elements are omnidirectional MEMS microphones with a sensitivity of -38dB±3dB, a frequency response range of 100Hz-10kHz, a sampling rate of 16kHz, and a quantization accuracy of 16 bits. The array diameter is 10cm, and the array elements are evenly distributed on the circumference with an element spacing of approximately 3.9cm. This spacing can effectively avoid spatial aliasing and ensure the positioning accuracy of voice signals below 2kHz.
[0061] The microphone array module is installed in the center of the operating room ceiling, at a height of 2.8-3.2 meters from the ground. In this embodiment, 3 meters is preferred. Its horizontal coverage area is a circular area with a radius of 5 meters, which can completely cover the entire surgical operation area. For operating rooms with smaller spaces, the microphone array can also be installed on the surgical light bracket, at a height of 1.2-1.8 meters from the operating table. In this case, a 6-element linear array can meet the coverage requirements.
[0062] The system has a pre-defined effective surgical area, which is a cylindrical space with a radius of 1.5-2.5 meters and a height of 0.5-2 meters, centered on the geometric center of the operating table. In this embodiment, a radius of 2 meters and a height of 1.5 meters are preferred. This area is the main activity range for medical staff during surgical operations, and only voice signals emitted by sound sources located within this area will be considered potentially effective signals by the system. The effective area can be set through a calibration program during the system installation and debugging phase, or it can be manually adjusted according to the needs of different surgeries.
[0063] 2. Sound source localization and preprocessing module
[0064] The sound source localization and preprocessing module is electrically connected to the microphone array module. It uses an ARM Cortex-A72 processor as the core computing unit with a main frequency of 1.8GHz and is equipped with 2GB DDR4 memory, enabling real-time processing of 8 channels of voice signals. This module includes a time delay estimation unit, a localization calculation unit, an effective area determination unit, an adaptive beamforming unit, and a post-noise reduction unit. Each unit implements its function through software programs.
[0065] 3. Dual-stage intelligent wake-up module
[0066] The dual-stage intelligent wake-up module is electrically connected to the sound source localization and preprocessing module. It uses an independent low-power MCU for real-time wake-up word detection. When a suspected wake-up signal is detected, the main processor is then woken up for further processing to reduce system power consumption. This module includes a wake-up word detection unit, a confidence level judgment unit, an instruction pre-recognition unit, and a comprehensive decision-making unit.
[0067] 4. Voice command recognition module
[0068] The voice command recognition module is electrically connected to the dual-stage intelligent wake-up module. It adopts an end-to-end voice recognition model optimized for operating room scenarios, with a model parameter size of approximately 50MB. It supports Mandarin Chinese recognition, with an accuracy rate of ≥98% in quiet environments and ≥92% in noisy operating room environments. This module can recognize more than 50 commonly used operating room control commands and convert the voice commands into corresponding equipment control commands.
[0069] 5. Equipment control interface module
[0070] The equipment control interface module is electrically connected to the voice command recognition module and integrates multiple communication interfaces such as RS232, RS485, Ethernet, and Bluetooth, supporting standard industrial communication protocols such as Modbus and TCP / IP. This module can communicate with mainstream operating room medical equipment such as surgical lights, operating tables, anesthesia machines, monitors, electrosurgical units, and infusion pumps, enabling non-contact control of these devices.
[0071] 6. Feedback Module
[0072] The feedback module is electrically connected to the voice command recognition module and the device control interface module, and includes a voice feedback unit and a visual feedback unit. The voice feedback unit uses an industrial-grade speech synthesis chip, which can synthesize clear and natural Chinese speech, and the volume can be adjusted within the range of 0-100dB. The visual feedback unit uses LED indicators and a small LCD display. The LED indicators are used to display the system's working status, and the LCD display is used to display the currently executed operation information.
[0073] II. Specific Implementation Procedures for Operating Room Voice Control Methods
[0074] The operating room voice control method of the present invention is executed sequentially according to the following steps S1 to S9, and the specific implementation of each step is as follows:
[0075] S1: Real-time acquisition of multiple audio signals from the operating room environment via a microphone array.
[0076] The microphone array module acquires eight audio signals from the operating room environment in real time at a sampling rate of 16kHz and a quantization precision of 16 bits. Each signal undergoes independent A / D conversion and is transmitted to the sound source localization and preprocessing module via an I2S bus. The system uses a circular buffer to store the most recent 10 seconds of audio data for subsequent processing.
[0077] S2: Perform three-dimensional localization of the acquired multi-channel speech signals, calculate the spatial coordinates of the sound sources, and determine whether the sound sources are located within the preset effective surgical area. The effective surgical area is as follows: Figure 5 As shown.
[0078] The three-dimensional localization of the sound source using a time-difference-of-arrival (TDOA) algorithm includes the following sub-steps:
[0079] S21: Perform cross-correlation calculations on each audio signal to obtain the time difference between the audio signals arriving at different microphone array elements. First, each audio signal is windowed and framed, with a frame length of 25ms and a frame shift of 10ms. A Hamming window is used as the window function. Then, the cross-correlation function between each frame signal and the reference array element signal is calculated, and the peak position of the cross-correlation function is found. The time corresponding to this peak position is the time difference between the signal arriving at that array element and the reference array element. In this embodiment, the reference array element is the center array element.
[0080] S22: Based on the geometric parameters of the microphone array and the calculated time differences, establish and solve a system of spatial coordinate equations to obtain the initial three-dimensional coordinates of the sound source. Establish a three-dimensional Cartesian coordinate system with the center of the operating table as the origin. The X-axis is parallel to the long side of the operating table, the Y-axis is parallel to the short side of the operating table, and the Z-axis is perpendicular to the ground and upwards. Based on the known coordinates of the 8 array elements and 7 independent time differences, establish a system of nonlinear equations. Solve this system of equations using Newton's iteration method to obtain the initial three-dimensional coordinates of the sound source. .
[0081] S23: The initial 3D coordinates are smoothed using a Kalman filter algorithm to eliminate localization noise and obtain the final sound source localization result. The state vector of the Kalman filter is... ,in , , These represent the velocity components of the sound source in the X, Y, and Z directions, respectively. By filtering the positioning results of five consecutive frames, the influence of random noise can be effectively suppressed, improving the positioning accuracy to within ±15 cm.
[0082] After obtaining the final sound source localization result, the effective area judgment unit compares the sound source coordinates with the preset effective surgical area to determine whether the sound source is located within the effective area.
[0083] S3: If the sound source is located within the effective surgical area, perform adaptive noise reduction processing on the speech signal in that direction; if the sound source is located in a non-effective area, determine it as noise and completely suppress it, then return to step S1.
[0084] The adaptive noise reduction process employs a combination of beamforming enhancement and post-noise reduction, specifically including the following steps:
[0085] S31: Based on the sound source localization result, the adaptive beamformer in the corresponding direction is activated, and the MVDR algorithm is used to enhance the speech signal in the target direction. The MVDR algorithm adjusts the array's weight vector so that the array's main lobe is aligned with the sound source direction, while simultaneously creating nulls in the interference direction. This maximizes the suppression of interference signals from other directions while maintaining the target speech signal without distortion. In this embodiment, the main lobe width of the beamformer is approximately 30 degrees, which can effectively enhance the speech signal in the target direction, with a suppression ratio of ≥20dB for interference from other directions.
[0086] S32: Perform spectral subtraction post-denoising processing on the beamforming enhanced speech signal to remove residual stationary background noise. Specifically, firstly, estimate the power spectrum of the noise in the silent segment of the speech signal, then subtract the estimated noise power spectrum from the power spectrum of the noisy speech to obtain the power spectrum estimate of the clean speech. To avoid musical noise, an improved spectral subtraction method is used to over-subtract the noise power spectrum and set a lower limit for the spectrum. Finally, the power spectrum is converted into a time-domain speech signal through inverse Fourier transform to obtain the preprocessed speech signal.
[0087] like Figure 5 As shown, if the sound source is located in an ineffective area, the system directly identifies the signal as environmental noise, does not perform any subsequent processing, and immediately returns to step S1 to continue collecting new voice signals.
[0088] S4: Perform wake-up word recognition on the noise-reduced speech signal, calculate and output the wake-up word confidence level.
[0089] like Figure 3 As shown, the preprocessed speech signal is transmitted to the wake word detection unit of the two-stage intelligent wake-up module. The wake word detection unit uses a lightweight wake word model trained based on TensorFlow Lite. In this embodiment, the wake word is "surgical assistant". The model training data includes speech samples from medical staff of different genders, ages, and accents, totaling no less than 2000 samples. Common background noise in the operating room is added for data augmentation to improve the robustness of the model.
[0090] The wake-word detection unit performs real-time sliding window detection on the input speech signal, with a window size of 1 second and a step size of 0.1 seconds. For each window of speech signal, the model outputs a confidence value between 0 and 1, which represents the probability that the speech segment contains a wake-word.
[0091] S5: Based on the confidence level of the wake word, a tiered judgment is made to determine whether to directly wake up, directly reject, or proceed to a two-stage comprehensive judgment process.
[0092] The grading determination in this step specifically includes the following sub-steps:
[0093] S51: Compare the calculated wake-up word confidence score with preset high and low thresholds. In this embodiment, the high threshold is set to 0.8 and the low threshold is set to 0.5. The threshold settings can be adjusted according to the actual noise level of the operating room. When the noise is high, the threshold can be appropriately increased to reduce false wake-ups; when the noise is low, the threshold can be appropriately decreased to reduce missed wake-ups.
[0094] S52: If the confidence level of the wake word is higher than the preset high threshold, i.e. >0.8, it is directly determined as a valid wake-up signal, the system is woken up and the process proceeds to step S7.
[0095] S53: If the confidence level of the wake word is lower than the preset low threshold, i.e. <0.5, it is directly determined as a non-wake signal, and the process returns to step S1.
[0096] S54: If the confidence level of the wake word is between the preset low threshold and the high threshold, i.e. 0.5≤confidence level≤0.8, it is determined to be a suspected wake-up signal and enters the two-stage comprehensive judgment process in step S6.
[0097] S6: When entering the dual-stage comprehensive judgment process, continue to collect subsequent voice signals of a preset duration, pre-identify control commands, and calculate the probability that the voice is a control command. The wake-up word confidence and the control command probability are weighted and fused to obtain a comprehensive score. Based on the comprehensive score, determine whether to wake up the system. Figure 4 As shown, the specific steps include the following:
[0098] S61: Continue to acquire the voice signal for the next 1-3 seconds and perform the same preprocessing operation as in step S3. In this embodiment, it is preferable to acquire the voice signal for the next 2 seconds. This duration ensures that the complete beginning of the control command is acquired without causing excessive system response delay.
[0099] S62: A lightweight speech recognition model is used to pre-identify control commands in the preprocessed speech signal, outputting the probability value that the speech is a commonly used control command in the operating room. This lightweight model is trained only on keywords of commonly used control commands in the operating room scenario, without performing complete semantic understanding, thus having the advantages of small size and high speed. The model parameter size is approximately 5MB, and the inference time is less than 100ms. Commonly used control command keywords include, but are not limited to, "adjust up", "adjust down", "open", "close", "move left", "move right", "move up", "move down", "stop", and "confirm".
[0100] S63: Based on preset weighting coefficients, the wake-up word confidence and control command probability are weighted and fused to calculate the comprehensive wake-up score. The weighted fusion is calculated using the following formula:
[0101] ;
[0102] in, and For the weighting coefficients, satisfying ,and The credibility of the wake word is given priority. In this embodiment, The value is 0.7. The value is 0.3. The value range is 0.6-0.8. The value ranges from 0.2 to 0.4 and can be adjusted according to actual usage. To increase system responsiveness, the value can be increased appropriately. The value can be increased appropriately if a more stable system is desired. The value of .
[0103] S64: Compare the overall wake-up score with a preset wake-up threshold. In this embodiment, the wake-up threshold is set to 0.7. If the overall score is higher than the wake-up threshold, wake up the system and proceed to step S7; otherwise, determine it as a non-wake-up signal and return to step S1.
[0104] S7: After the system wakes up, it performs complete recognition and semantic understanding of subsequent voice commands and generates corresponding device control commands.
[0105] After the system is woken up, the two-stage intelligent wake-up module transmits the voice signal to the voice command recognition module. The voice command recognition module uses a complete end-to-end voice recognition model to perform complete voice recognition and semantic understanding on subsequent input voice commands. The model can recognize complete control commands in the format of "verb + device + parameter", such as "increase the brightness of the surgical light to 80%" or "move the operating table 10 centimeters to the left".
[0106] The voice command recognition module performs semantic parsing on the recognized text commands, extracting the operation verbs, target device, and operation parameters, and converting them into a standard control command format that the device can recognize. For example, "increase the brightness of the surgical light" is converted into "Device ID: 01, Operation Code: 03, Parameter: +10".
[0107] S8: Sends control commands to the corresponding operating room medical equipment via the device control interface to execute the corresponding operations.
[0108] Based on the semantic parsing results, the device control interface module selects the corresponding communication interface and protocol, and sends control commands to the target medical device. For example, for a surgical light using an RS485 interface, control commands are sent via the RS485 bus; for a monitor using an Ethernet interface, control commands are sent via the TCP / IP protocol.
[0109] After receiving the control command, the target medical device executes the corresponding operation and returns the operation execution result to the device control interface module.
[0110] S9: Provide feedback module to medical staff with information indicating whether the operation was successful or failed.
[0111] The feedback module provides corresponding feedback to medical staff based on the execution results returned by the device. If the operation is successful, the voice feedback unit plays a synthesized voice message saying "XX operation has been performed," such as "The brightness of the surgical light has been increased," while the green LED indicator of the visual feedback unit flashes once. If the operation fails, the voice feedback unit plays a synthesized voice message saying "Operation failed, please try again," while the red LED indicator of the visual feedback unit remains lit for 3 seconds.
[0112] If no valid voice command is detected within 10 seconds after the system wakes up, it will automatically enter sleep mode and return to step S1 to continue waiting for a new wake-up signal.
[0113] III. Specific Application Examples
[0114] The following typical application scenario further illustrates the working process of this invention:
[0115] In a general surgery procedure, the surgeon needs to increase the brightness of the surgical lights to better observe the surgical area. The surgeon, positioned beside the operating table within the pre-defined effective surgical area, gives the instruction: "Surgical assistant, increase the brightness of the surgical lights."
[0116] The microphone array module acquires the voice signal and transmits the 8-channel voice data to the sound source localization and preprocessing module.
[0117] The sound source localization and preprocessing module calculates the three-dimensional coordinates of the sound source as (0.5m, 0.3m, 1.2m) using the TDOA algorithm, and determines that the coordinates are located within an effective area with a radius of 2 meters centered on the operating table.
[0118] The system activates the MVDR beamformer to enhance the speech signal in that direction and performs spectral subtraction post-denoising processing to obtain the preprocessed speech signal.
[0119] The wake word detection unit identifies the preprocessed speech signal and calculates the confidence level of the wake word "surgical assistant" to be 0.65.
[0120] Since 0.65 falls between the low threshold of 0.5 and the high threshold of 0.8, the system enters a two-stage comprehensive judgment process.
[0121] The system continues to collect the next 2 seconds of audio signal and performs the same preprocessing operation on it.
[0122] The lightweight command pre-recognition model identified subsequent speech containing control command keywords such as "adjust" and "surgical light," and calculated the probability of the control command to be 0.9.
[0123] The system calculates the overall score based on the weighted fusion formula: 0.7×0.65+0.3×0.9=0.725.
[0124] The system was woken up because 0.725 was higher than the wake-up threshold of 0.7.
[0125] The voice command recognition module recognizes and understands the semantics of the complete command "increase the brightness of the surgical light" and generates the corresponding control command.
[0126] The device control interface module sends a brightness increase command to the surgical light via the RS485 interface.
[0127] After receiving the instruction, the surgical light increases its brightness by 10% and returns a successful operation result.
[0128] The feedback module plays a synthesized voice message saying "The brightness of the surgical light has been increased," while the green LED indicator flashes once.
[0129] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. An operating room voice control system, characterized in that, include: Microphone array module for real-time acquisition of multiple audio signals in the operating room environment; The sound source localization and preprocessing module is electrically connected to the microphone array module. It is used to perform three-dimensional localization of the sound source of multiple voice signals, determine whether the sound source is located within the preset effective surgical area based on the localization result, and perform adaptive noise reduction processing on the voice signal within the effective area and suppress the signal outside the effective area. The dual-stage intelligent wake-up module is electrically connected to the sound source localization and preprocessing module. It is used to first identify the wake-up word in the preprocessed voice signal and output the wake-up word confidence level. When the wake-up word confidence level is lower than the preset high threshold but higher than the preset low threshold, it continues to collect and identify subsequent voice signals. It combines the wake-up word confidence level and the probability that the subsequent voice is a control command to determine whether to wake up the system. The voice command recognition module is electrically connected to the dual-stage intelligent wake-up module and is used to fully recognize and semantically understand valid voice commands after the system is woken up. The equipment control interface module is electrically connected to the voice command recognition module and is used to send control commands to the corresponding operating room medical equipment based on the semantic understanding results. The feedback module is electrically connected to the voice command recognition module and the device control interface module, and is used to provide voice or visual feedback on the operation status to medical staff.
2. The operating room voice control method according to claim 1, characterized in that, The microphone array module adopts a 6-8 element linear array or ring array, and is installed in the center of the operating room ceiling or on the operating light bracket. Its coverage area includes the entire surgical operation area. The effective surgical area is a cylindrical space with a radius of 1.5-2.5 meters centered on the operating table.
3. A voice control method for an operating room, applied to the system described in claims 1-2, characterized in that, Includes the following steps: S1: Real-time acquisition of multiple audio signals from the operating room environment via a microphone array; S2: Perform three-dimensional localization of the sound source for the acquired multi-channel speech signals, calculate the spatial coordinates of the sound source, and determine whether the sound source is located within the preset effective surgical area; S3: If the sound source is located within the effective surgical area, perform adaptive noise reduction processing on the speech signal in that direction; if the sound source is located in a non-effective area, identify it as noise and completely suppress it, then return to step S1. S4: Perform wake-up word recognition on the noise-reduced speech signal, calculate and output the wake-up word confidence score; S5: Based on the confidence level of the wake word, make a graded judgment to determine whether to wake up directly, reject directly, or enter a two-stage comprehensive judgment process; S6: When entering the dual-stage comprehensive judgment process, continue to collect voice signals of the subsequent preset duration, pre-identify the control commands and calculate the probability that the voice is a control command, and weight and fuse the wake-up word confidence and the control command probability to obtain a comprehensive score, and determine whether to wake up the system based on the comprehensive score; S7: After the system is woken up, it fully recognizes and understands the semantics of the subsequent voice commands and generates corresponding device control commands. S8: Sends control commands to the corresponding operating room medical equipment via the device control interface to execute the corresponding operations; S9: Provide feedback to medical staff regarding the success or failure of the operation through the feedback module.
4. The operating room voice control method according to claim 3, characterized in that, The three-dimensional localization of the sound source mentioned in step S2 specifically includes the following steps: S21: Perform cross-correlation calculation on each audio signal to obtain the time difference of the audio signal arriving at different microphone array elements; S22: Based on the geometric parameters of the microphone array and the calculated time difference, establish and solve a set of spatial coordinate equations to obtain the initial three-dimensional coordinates of the sound source; S23: The initial three-dimensional coordinates are smoothed using the Kalman filter algorithm to eliminate localization noise and obtain the final sound source localization result.
5. The operating room voice control method according to claim 4, characterized in that, The adaptive noise reduction process described in step S3 specifically includes the following steps: S31: Based on the sound source localization result, activate the adaptive beamformer in the corresponding direction and use the minimum variance distortionless response algorithm to enhance the speech signal in the target direction; S32: Perform spectral subtraction post-denoising processing on the beamforming enhanced speech signal to remove residual stable background noise and obtain the preprocessed speech signal.
6. The operating room voice control method according to claim 3, characterized in that, The grading determination described in step S5 specifically includes the following steps: S51: Compare the calculated wake word confidence with the preset high and low thresholds; S52: If the confidence level of the wake word is higher than the preset high threshold, it is directly determined to be a valid wake-up signal, the system is woken up and proceeds to step S7; S53: If the confidence level of the wake word is lower than the preset low threshold, it is directly determined to be a non-wake signal, and the process returns to step S1. S54: If the confidence level of the wake word is between the preset low threshold and the high threshold, it is determined to be a suspected wake-up signal, and the two-stage comprehensive judgment process in step S6 is entered.
7. The operating room voice control method according to claim 3, characterized in that, The two-stage comprehensive judgment process described in step S6 specifically includes the following steps: S61: Continue to collect the voice signal for the next 1-3 seconds and perform the same preprocessing operation as in step S3; S62: A lightweight speech recognition model is used to pre-identify control commands in the pre-processed subsequent speech signal, and the probability value of the speech being a commonly used control command in the operating room is output. S63: Based on the preset weighting coefficients, the wake-up word confidence and the control command probability are weighted and fused to calculate the comprehensive wake-up score; S64: Compare the overall wake-up score with the preset wake-up threshold. If the overall score is higher than the wake-up threshold, wake up the system and proceed to step S7. Otherwise, determine it as a non-wake-up signal and return to step S1.
8. The operating room voice control method according to claim 7, characterized in that, The weighted fusion described in step S63 is calculated using the following formula: ; in, and For the weighting coefficients, satisfying ,and , The value range is 0.6-0.
8. The value range is 0.2-0.4, and the preset wake-up threshold ranges from 0.65-0.
75.
9. The operating room voice control method according to claim 3, characterized in that, The lightweight speech recognition model described in step S62 is trained only on commonly used control command keywords in the operating room scenario. These commonly used control command keywords include, but are not limited to, raising, lowering, opening, closing, moving left, moving right, moving up, and moving down.
10. The operating room voice control method according to claim 3, characterized in that, The device control interface described in step S8 supports at least one of the communication protocols RS232, RS485, Ethernet and Bluetooth, and can communicate with at least one of the operating room medical devices, including surgical lights, operating tables, anesthesia machines, monitors and electrosurgical units.