Non-contact operating room intelligent voice control method supporting multi-role voiceprint recognition

By combining a dynamic environmental noise baseline model with personalized voiceprint recognition, and integrating spatial and spectral dual-domain decoupling and speech integrity detection, the problem of false triggering of voiceprint recognition caused by high-frequency equipment noise interference in the operating room was solved, thus ensuring the safety and continuity of the surgical process.

CN120977313APending Publication Date: 2025-11-18SHENZHEN YOUJIAN MEDICAL TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511410199.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In the operating room, the noise of high-frequency surgical equipment resonates with the frequency band of the doctor's voiceprint characteristics. This causes the voiceprint recognition and command parsing mechanism to be unable to accurately distinguish between valid speech and background noise under noise interference, which may lead to the accidental triggering of equipment operation and affect the safety and efficiency of surgery.

Method used

By combining a dynamic environmental noise baseline model with personalized voiceprint recognition, spatial and spectral dual-domain decoupling, and voice integrity detection, dual verification of voiceprint identity and non-contact actions is implemented to ensure the authenticity and validity of the command issuer's identity. Furthermore, a closed-loop defense mechanism is formed through permission level comparison and contextual semantic verification.

Benefits of technology

It effectively isolates the spectral interference of high-frequency surgical equipment noise and doctor's voiceprint, ensuring the safety, accuracy and continuity of the surgical process, preventing misidentification and miscontrol, and protecting patient safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977313A_ABST
    Figure CN120977313A_ABST
Patent Text Reader

Abstract

The invention discloses a non-contact operating room intelligent voice control method supporting multi-role voiceprint recognition, and relates to the technical field of voice interaction control, and the method comprises the following steps: S100, collecting full-band noise signals of all high-frequency operation equipment in an operating room in different operation states, extracting harmonic frequency, amplitude and phase features, and carrying out the recognition of the full-band noise signals; and constructing a harmonic noise characteristic database, and establishing and dynamically updating an environmental noise baseline model. According to the invention, through the dynamic environment noise baseline model and personalized voiceprint recognition, frequency spectrum isolation and accurate voice separation of high-frequency operation equipment noise and doctor voiceprint are realized. In combination with space and frequency spectrum dual-domain decoupling and voice integrity detection, and through dual verification of voiceprint identity and non-contact action, it is ensured that an instruction sender is true and effective. And a closed-loop defense mechanism is formed through permission level comparison and context semantic verification, so that error identification and error control are effectively prevented, and the safety, accuracy and continuity of the operation process are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice interaction control, in particular to a non-contact intelligent voice control method for operating room supporting multi-role voiceprint recognition. BACKGROUND

[0002] The non-contact intelligent voice control method for operating room supporting multi-role voiceprint recognition refers to the integration of multi-role voiceprint recognition and voice control technology in the operating room environment to achieve contactless intelligent operation of medical equipment, imaging systems or environmental parameters. This method can accurately identify the voiceprint features of different medical personnel involved in the operation, automatically distinguish the identity of the commander, and match the corresponding control instructions according to the identity and authority. In this way, doctors or assistants in the operating process do not need to directly touch any equipment, but can control key devices through voice alone, effectively reducing the risk of infection caused by physical contact, while improving the efficiency and immediacy of the operation, meeting the safety requirements of multi-role collaboration and authority management, and adapting to the needs of complex, sterile and dynamic operating scenarios.

[0003] The prior art has the following disadvantages: During the intelligent voice control process in the operating room, there is a technical problem of spectral resonance between the harmonic noise generated by high-frequency surgical equipment (such as ultrasonic knives and electrocoagulation devices) during operation and the frequency band of the doctor's voiceprint features. Due to the high-frequency noise signals released by these devices under certain power and working mode, the frequency, amplitude and waveform characteristics of the noise signals overlap nonlinearly with part of the frequency band of the doctor's voiceprint, causing the voiceprint recognition and instruction analysis mechanism to fail to accurately distinguish between effective voice and background noise signals under noise interference. In extreme cases, the system may misinterpret part of the waveform in the device noise as a control instruction, thus incorrectly activating the voice-controlled device operation functions in the operating room, such as adjusting the light angle and brightness of the operating lamp, rotating the field of view and adjusting the focal length of the endoscope, etc.

[0004] If such mis-triggering occurs during the execution of delicate or high-risk surgical procedures by the doctor, such as blood vessel anastomosis, nerve repair or deep organ operation, it may cause sudden changes in the field of view, device motion interference or even interruption of the doctor's operation, directly leading to instability and uncontrollability of the operation, increasing the risk of intraoperative complications, and in severe cases, endangering the patient's life safety.

[0005] The above information disclosed in the background section is only used to enhance the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0006] The application aims to provide a non-contact operating room intelligent voice control method supporting multi-role voiceprint recognition, realize spectrum isolation of high-frequency surgical equipment noise and doctor voiceprint and accurate voice separation through dynamic environmental noise baseline model and personalized voiceprint recognition, combine spatial and spectral dual-domain decoupling and voice integrity detection, and ensure that the order issuer is real and effective through voiceprint identity and non-contact action double verification, form a closed-loop defense mechanism through permission level comparison and context semantic verification, effectively prevent misidentification and miscontrol, and guarantee the safety, accuracy and continuity of the operation process, so as to solve the problems in the above background art.

[0007] In order to achieve the above-mentioned purpose, the application provides the following technical scheme: a non-contact operating room intelligent voice control method supporting multi-role voiceprint recognition, comprising the following steps: S100, collecting full-band noise signals of all high-frequency surgical equipment in the operating room under different operating states, extracting harmonic frequency, amplitude and phase characteristics, constructing a harmonic noise feature database, and establishing and dynamically updating an environmental noise baseline model; S200, based on the environmental noise baseline model, performing spectrum mapping on the voiceprint feature sequence of each doctor, identifying voiceprint segments with spectral overlap or proximity, performing feature reconstruction and spectrum correction, and generating a personalized voiceprint recognition model; S300, based on the personalized voiceprint recognition model, performing spatial and spectral dual-domain decoupling on the real-time collected multi-channel audio signals, separating out doctor voice, equipment noise and background sound sources, implementing dynamic labeling of signal sources, and ensuring that the signal sources correspond to the personalized voiceprint recognition model one by one; S400, based on the dynamically labeled doctor voice signal, performing voice integrity detection to detect continuity, speech speed stability and audio waveform consistency, and eliminating unnatural voice signals; S500, based on the voice integrity detection signal, performing double verification of voiceprint identity verification and non-contact action sensing to confirm the identity of the order initiator and the consistency of his behavior; S600, based on the double-verified order, performing permission level comparison and context semantic verification, controlling the medical equipment to execute after compliance, and completing the closed-loop process from noise isolation, signal decoupling, identity and action verification, permission and semantic verification to device control.

[0008] Preferably, step S100 comprises: Collecting full-band noise signals of all high-frequency surgical equipment in the operating room under different operating states, extracting harmonic frequency, amplitude and phase characteristics, and constructing a harmonic noise feature database; Performing multi-dimensional feature clustering on the feature vectors in the noise feature database to divide the noise categories; Perform feature mapping dimension reduction on each category of feature vector, and establish index mapping of feature, device type, power and mode; Based on the clustering and mapping results, the environmental noise baseline model is formed and a dynamic updating mechanism is set.

[0009] Preferably, step S200 includes: Based on the environmental noise baseline model, the doctor's voiceprint feature sequence is performed spectrum mapping, and the overlapping or adjacent fragments are identified; For overlapping fragments, perform feature reconstruction based on variational autoencoder, and compensate phase information through minimum phase reconstruction; The reconstructed signal is subjected to a minimum mean square error filter to perform spectrum correction, and the resonance peak and fundamental harmonic are strengthened through spectral cepstrum analysis; Based on the corrected signal, the features and classification are extracted using convolutional neural network to generate a personalized voiceprint recognition model.

[0010] Preferably, the way of strengthening the resonance peak and fundamental harmonic by spectral cepstrum analysis includes: Performing fast Fourier transform on the voiceprint signal after spectrum correction to obtain a log spectrum; Performing inverse Fourier transform on the log spectrum to obtain a cepstrum signal; Apply low-pass cepstrum filter to the cepstrum signal to retain low-order and medium-order cepstrum coefficients and suppress non-identity features in high-frequency cepstrum components; Transform the filtered cepstrum coefficients back to the frequency domain to obtain a voiceprint signal with strengthened resonance peak and fundamental harmonic.

[0011] Preferably, step S300 includes: Perform spatial positioning on multi-channel audio signals by beamforming algorithm, use spatial filter to enhance target direction signal and suppress remaining direction signal, and realize spatial dimension decoupling; Perform short-time Fourier transform and Mel frequency analysis on the spatially decoupled signal to extract spectral features; Compare and classify the spectral features according to the personalized voiceprint recognition model and the environmental noise baseline model to realize spectral dimension decoupling; Assign dynamic labels to the decoupled signals to establish the correspondence between the signal source and the voiceprint recognition model.

[0012] Preferably, the specific way of beamforming algorithm includes: Calculate the time difference of each channel signal based on the spatial coordinates of the microphone array and the sound speed, and align the signal phase by delay compensation; Calculate the optimal weight coefficient by minimum variance distortion response algorithm, and weight and superimpose each channel signal to enhance the doctor's voice signal in the target direction; The weighted signals are partitioned according to spatial directions to realize the separated output of different sound sources.

[0013] Preferably, the step S400 comprises: The time interval and zero-crossing rate detection are performed on the speech signal sequence to determine the signal continuity. The dynamic time warping algorithm is used to analyze the phoneme duration and syllable interval to determine the speech rate stability. The short-time autocorrelation function and Hilbert transform are performed on the signal to extract the fundamental frequency periodicity and waveform envelope curve to determine the audio waveform consistency. The continuity, speech rate stability and audio waveform consistency detection results are fused to filter out the non-natural speech signals.

[0014] Preferably, the step S500 comprises: Based on the personalized voiceprint recognition model, the cosine similarity and Euclidean distance double evaluation mechanism are used to verify the identity of the issuer. Based on the action data collected by the infrared depth camera and millimeter wave radar, the convolutional neural network and long short-term memory network are used to identify the doctor's action features and categories. The voiceprint identity and action category are compared for consistency to evaluate the time synchronization of the voice and action. After confirming the identity and behavior consistency, the voiceprint features, action features and time stamp of the instruction are recorded.

[0015] Preferably, the step S600 comprises: The permission database is queried according to the issuer identity information to check the permission level and surgical stage permission of the operation device. The instruction content is converted into semantic units, and the operation log, device state and surgical process model are combined to determine the consistency of the instruction semantics and the current surgical context. For the instruction that passes the permission level and semantic verification, an encrypted control package containing the instruction content, identity, permission level, semantic verification result and time stamp is generated. The control package is sent to the medical device, and the device performs the instruction after decoding and verification, and records the whole process log.

[0016] In the above technical solution, the technical effects and advantages provided by the present application are: The application effectively isolates the spectral interference of high-frequency surgical device noise and the doctor's voiceprint by dynamically establishing and updating the environmental noise baseline model, combines the personalized voiceprint recognition model with spatial and spectral decoupling, realizes the accurate separation of the doctor's voice and the background noise, and through the voice integrity detection and the voiceprint and action double verification, ensures the identity of the order issuer and the consistency of the behavior, and through the permission level comparison and the context semantic verification, strictly limits the compliance of the order in the permission and the surgical scene, and finally forms a closed-loop defense mechanism from noise isolation, signal decoupling to device safety control. The method effectively prevents the misidentification and miscontrol risk caused by device noise or unauthorized voice, and ensures the continuity, accuracy and patient safety of the surgical process. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0018] Figure 1 The method flowchart of the non-contact intelligent voice control method of the operating room supporting multi-role voiceprint recognition of the present application. DETAILED DESCRIPTION

[0019] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these example implementations are provided so that this disclosure will be thorough and complete, and will fully convey the gist of each example to those skilled in the art.

[0020] The present application provides a non-contact intelligent voice control method of an operating room supporting multi-role voiceprint recognition as shown in Figure 1 The method flowchart of the non-contact intelligent voice control method of the operating room supporting multi-role voiceprint recognition of the present application, comprising the following steps: S100, collecting all high-frequency surgical device noise signals in different operating states in the operating room, extracting the harmonic frequency, amplitude and phase characteristics of each device noise signal, constructing the corresponding harmonic noise characteristic database, establishing an environmental noise baseline model based on the database, and setting the model to a dynamic update state; In view of the interference problem of high-frequency noise generated by high-frequency surgical devices in the operating room environment to the voice control system during operation, a feature extraction and modeling based on full-band noise signal is proposed, which aims to provide accurate noise feature reference and dynamic environmental adaptability for subsequent voiceprint recognition and noise isolation. Including the following steps: All high-frequency surgical equipment in the operating room is fully sorted and classified, and the common running mode and power level of different equipment in the surgical process are determined, and a noise collection plan is developed based on this. By placing high-sensitivity full-band pickup devices at different spatial positions in the operating room, covering a wide frequency band range of 20Hz to 20kHz, it is ensured that the noise signals released by various high-frequency surgical equipment under different working conditions can be completely captured. During the collection process, each device needs to run under all its rated power and mode combinations, and the original waveform data of the noise is recorded in real time, and the corresponding operating power, current voltage and other working condition parameters are collected synchronously, so as to be associated and analyzed subsequently. This step ensures that the noise characteristics of the equipment are fully covered in the real surgical scene, avoiding the omission of noise characteristics due to single equipment working mode or insufficient collection.

[0021] For the original noise signals collected, various spectral analysis methods such as Fourier transform and wavelet packet decomposition are used to finely analyze the signals in the full frequency band, and the harmonic frequency, amplitude and phase characteristics of each device under different working conditions are extracted. In order to enhance the accuracy of feature extraction, noise signals need to be denoised and normalized to eliminate the interference of environmental background noise on noise spectrum, ensuring that the extracted harmonic characteristics accurately reflect the noise characteristics of the equipment itself. After feature extraction, for the key parameters such as harmonic frequency, primary and secondary harmonic amplitude ratio, and phase difference under each power and mode, quantitative modeling and data structure storage are carried out to form a complete data set containing device type, working mode, power level and corresponding harmonic characteristic parameters. This step realizes the deep mining of multi-dimensional features from the original noise signals, providing a solid data foundation for the subsequent dynamic noise model construction.

[0022] Based on the device noise harmonic feature data set obtained by the above extraction and modeling, a multi-dimensional feature clustering and feature mapping algorithm is used to construct a baseline model of the operating room environment noise. This model aggregates and analyzes the noise characteristics under different time, different equipment combination running to form a global noise image under a specific spatial position and equipment running combination. In order to enhance the timeliness and adaptability of the model, a dynamic updating mechanism is designed. When new equipment running state or existing equipment noise characteristics drift due to wear and tear, aging and other factors are detected, the existing noise baseline characteristics are automatically adjusted and corrected through real-time collection and model feedback mechanism, ensuring that the model is consistent with the actual noise environment in the operating room. This step ensures the dynamic evolution ability of the model, effectively solving the feature mismatch problem caused by the change of equipment running state.

[0023] The "multi-dimensional feature clustering and feature mapping algorithm" refers to a technical means for similarity aggregation and spatial mapping of multi-dimensional features such as harmonic frequency, amplitude, phase, etc. extracted from different high-frequency surgical equipment in various operating states in the operating room, through machine learning methods, aiming to sort out stable and distinguishable noise patterns from massive noise features, and then establish a baseline model that can reflect the overall noise characteristics of the operating room. In this implementation, its role is to classify and summarize noise features under different equipment, different working modes, and different combinations, and form the overall noise picture of the operating room under various possible equipment operation combinations, facilitating subsequent rapid comparison and dynamic updating. Common clustering algorithms can use K-means, DBSCAN or GaussianMixtureModel (Gaussian Mixture Model), while feature mapping algorithms can use principal component analysis (PCA), t-SNE or autoencoder for dimension reduction and feature reconstruction. The specific steps include: first, based on the extracted harmonic frequency, amplitude, phase and other feature vectors, use clustering algorithms to group similar feature data points and divide them into representative noise categories. Second, for each noise group, apply feature mapping algorithms to reduce and map high-dimensional feature vectors, retain key features and eliminate redundant information, and form low-dimensional but representative feature expressions. Third, index map the reduced feature results with the type, running power and mode labels of the equipment to form a correspondence between multi-dimensional features and equipment states. Finally, based on all clustering and mapping results, aggregate and build a baseline model of the operating room environment noise, which can quickly retrieve and compare which feature category the current noise signal belongs to in real-time audio processing, dynamically adapt to the complex and changing noise background in the operating room, and achieve accurate isolation of voiceprints and noise interference.

[0024] The dynamically updated environmental noise baseline model is applied to subsequent voiceprint recognition and speech control processes as a basic reference for noise isolation and spectrum correction. In actual application, for real-time collected audio signals, through comparison and analysis with the noise baseline model, noise frequency bands in the audio can be quickly located and marked, achieving spectrum-level isolation of doctor's voice signals and high-frequency noise signals, laying a foundation for accurate extraction of voiceprint features and prevention of false activation of speech commands. In addition, the dynamic updating capability of the environmental noise baseline model ensures that in the scenarios of replacement, aging or addition of equipment in the operating room, the high accuracy of voiceprint recognition and the safety and reliability of speech control can still be maintained, thereby realizing long-term adaptation and intelligent management and control of the complex acoustic environment in the operating room.

[0025] The purpose of this step is to lay the foundation for accurate noise recognition and isolation in the intelligent voice control process of the operating room, to ensure that the voice control system can accurately distinguish between the doctor's effective voice command and the background noise when the high-frequency surgical equipment is running, and to avoid voiceprint recognition failure or voice command mis-triggering due to noise interference. The high-frequency surgical equipment in the operating room, such as ultrasonic knives, electrocoagulation instruments, etc., will release harmonic noise with different multi-dimensional characteristics of frequency, amplitude, phase, etc. under different power and working modes. These noises not only cover a wide frequency band, but also have overlapping or adjacent characteristics with the doctor's voiceprint frequency band. By collecting all the full-band noise signals of the equipment under different working conditions and extracting their harmonic frequency, amplitude and phase characteristics, a comprehensive understanding of the acoustic characteristics of each device can be formed, and a harmonic noise feature database can be constructed. This database can support the modeling of the overall noise environment in the operating room, form an environmental noise baseline model, and serve as a core reference for noise isolation and sound source discrimination. At the same time, the model is set to a dynamic update state, which can update the model parameters in real time according to the noise characteristics of the aging, power change or new equipment of the device, ensuring that the model remains synchronized and accurate with the actual environment. This not only improves the robustness and stability of voiceprint recognition, but also provides a continuous and reliable noise countermeasure for the voice control system in the operating room, which is a prerequisite for safe, stable and error-proof control.

[0026] S200, based on the environmental noise baseline model, performing spectrum mapping analysis on the voiceprint feature sequence of each doctor, identifying the voiceprint segment in the voiceprint feature that has spectral overlap or spectral proximity with the environmental noise baseline model, and performing feature reconstruction and spectral correction on the voiceprint feature according to the spectral overlap result, to generate a personalized voiceprint recognition model under noise isolation; In view of the problem of spectral overlap and proximity between high-frequency surgical equipment noise and doctor voiceprint features in the operating room environment, a spectrum mapping analysis, feature reconstruction and spectral correction of doctor voiceprint features based on an environmental noise baseline model are proposed, aiming to construct a personalized voiceprint recognition model under noise isolation. The implementation is as follows: Based on the pre-established operating room environment noise baseline model, the voiceprint feature sequence of each doctor is analyzed by full spectrum mapping. Specifically, the Short-Time Fourier Transform (STFT) is used to window the frequency spectrum of the doctor's voiceprint signal, and the energy distribution of the voiceprint signal in time and frequency dimensions is obtained. The Continuous Wavelet Transform (CWT) is used to capture the frequency response and transient characteristics of the voiceprint signal at different scales, further improving the frequency resolution. The doctor's voiceprint full spectrum obtained by the above analysis is compared with each device noise spectrum in the environment noise baseline model. The frequency bands that overlap or are adjacent to the harmonic frequency, sub-harmonic frequency and frequency offset interval of the device noise in the voiceprint spectrum are marked, and the center frequency, bandwidth, amplitude peak and corresponding phase information of the overlapping frequency band are recorded. Through this step, the specific frequency band of the doctor's voiceprint signal in the frequency spectrum space that is susceptible to noise interference and its parameter characteristics can be accurately identified.

[0027] For the voiceprint segments of the overlapping or adjacent frequency bands marked in the mapping analysis, the reconstruction of the voiceprint features is performed. First, based on the uninterfered frequency spectrum of the doctor's voiceprint, the formants, fundamental frequency (F0), harmonic structure and energy envelope curve are extracted. Then, the Variational Autoencoder (VAE) based on deep learning is used to learn and encode the complete voiceprint features, and a mapping model of the voiceprint features in the multi-dimensional space is established. For the frequency bands interfered by noise, the decoding mechanism of the autoencoder is used to input the features of the uninterfered frequency bands, reconstruct the lost or distorted frequency band signals, and restore the voiceprint information in the overlapping area. At the same time, the minimum phase reconstruction algorithm is used to compensate the phase information, ensuring the restoration degree of the reconstructed signal in the phase continuity and overall time domain waveform, and avoiding the interference of phase distortion on the subsequent identification.

[0028] Variational Autoencoder (VAE) is a deep learning-based generative model that can learn the latent distribution of a large number of sample data, encode high-dimensional complex data (such as the frequency spectrum features of voiceprint signals) into a low-dimensional probability distribution in the latent space, and reconstruct samples close to the original data from the latent space when needed. In this step, the role of VAE is to perform deep learning and probability coding on the multi-dimensional features of the doctor's complete voiceprint (such as formants, fundamental frequency, harmonic structure, energy envelope and phase information), map the voiceprint to a continuous distribution representation in the latent space, and capture the overall pattern and individual characteristics of the voiceprint. The specific steps are as follows: The complete voiceprint signal is extracted by spectral analysis to obtain a feature vector as the input of the VAE.

[0029] The voiceprint features are mapped by the encoder network of the VAE into mean and variance parameters of a latent space, forming a corresponding probability distribution.

[0030] The latent vector is sampled from the distribution based on the reparameterization trick, and the latent vector is restored to the feature reconstruction result of the voiceprint by the decoder network of the VAE.

[0031] The encoder and decoder are continuously trained and optimized, so that the reconstructed voiceprint signal is as close as possible to the original signal in terms of spectrum, amplitude, phase, etc. In this application scenario, when noise interference causes the local features of the voiceprint to be missing or distorted, the method can implement fitting reconstruction and spectral restoration of the missing or distorted voiceprint segment based on the complete distribution information in the latent space, ensuring the integrity and continuity of the overall features of the voiceprint.

[0032] Based on the reconstructed voiceprint signal, a spectral correction operation is performed. An adaptive filtering algorithm such as a least mean square (LMS) filter is used to assign different filtering weights to each frequency band in the reconstructed signal according to the frequency proximity to the noise baseline model, and to implement amplitude attenuation and frequency band purification for overlapping or adjacent frequency bands. A spectral cepstrum analysis technique is applied simultaneously to strengthen the formant features and fundamental harmonic features in the voiceprint signal that are strongly related to individual identity, while suppressing the energy of non-identity feature frequencies. Through this step, a purified voiceprint signal with clear spectral distribution and obvious separation from noise frequency bands is finally obtained, providing high-quality input for the training of subsequent personalized recognition models.

[0033] The frequency band purification of the reconstructed signal is performed by a least mean square (LMS) filter in the adaptive filtering algorithm, and the specific steps include: According to the pre-established environmental noise baseline model, the center frequency and bandwidth of each noise frequency band are extracted, and the frequency intervals that are close or overlapping with these noise frequency bands in the spectrum of the reconstructed voiceprint signal are marked.

[0034] The frequency gap between each marked frequency band and the frequency band of the noise baseline model is calculated, and the filtering weight is determined according to the gap. The closer the frequency is to the center frequency of the noise, the greater the filtering weight is, forming an adaptive frequency weight mapping table.

[0035] Based on the weight mapping table, the LMS filter is applied to each frequency band of the reconstructed voiceprint signal. During the filtering process, the filter continuously adjusts the weight coefficients to minimize the mean square error between the input signal (reconstructed voiceprint) and the reference signal (noise model signal corresponding to the frequency band), thereby dynamically attenuating the amplitude in the overlapping or adjacent noise frequency band and suppressing noise interference.

[0036] The filtered signal is subjected to frequency band restoration and phase correction to ensure that the purified voiceprint signal effectively weakens the interference components coinciding with the noise frequency band while retaining individual identity characteristics, achieving clear and purified frequency bands and improving the recognition degree and subsequent recognition accuracy of voiceprint features.

[0037] The spectral cepstrum analysis technique is a signal processing method that further converts signals from the frequency domain to the cepstrum domain, which can effectively separate excitation sources (such as fundamental frequency and harmonics) and vocal tract responses (such as formants) in signals. In speech and voiceprint processing, it is commonly used to extract individual identity-related features. In this step, the application of spectral cepstrum analysis is to further extract and strengthen identity features such as formants and fundamental frequency harmonics from the reconstructed and filtered voiceprint signal, while suppressing non-identity-related noise frequencies and language content differences in the speech signal. The specific steps are as follows: Performing a Fast Fourier Transform (FFT) on the purified voiceprint signal to obtain its log spectrum; Performing an inverse Fourier transform on the log spectrum to obtain the cepstral representation of the signal. In the cepstral domain, low-order cepstral coefficients correspond to vocal tract characteristics (such as formants), and high-order cepstral coefficients correspond to excitation source characteristics (such as fundamental frequency and harmonic structure).

[0038] According to the experience window function or low-pass cepstral filtering, low-order and partial middle-order cepstral coefficients closely related to individual identity are retained, and components reflecting vocal tract structure and fundamental frequency characteristics are enhanced, while noise and language content differences contained in high-frequency cepstral components are suppressed, reducing non-identity feature interference.

[0039] The filtered and enhanced cepstral coefficients are transformed back to the frequency domain, reconstructing a voiceprint signal with enhanced formants and fundamental frequency harmonics and removed non-identity frequencies, which serves as the feature input for subsequent personalized voiceprint recognition model training, thereby improving the specificity and stability of the recognition model for doctor identity.

[0040] Finally, based on the completed feature reconstruction and spectrum correction of the voiceprint signal, a personalized voiceprint recognition model for each doctor is generated. Specifically, a convolutional neural network (CNN) is constructed to extract features and classify the corrected voiceprint features. The network input is the Mel-frequency cepstral coefficients (MFCC), fundamental frequency sequence, formant frequency and its change trajectory of the voiceprint, and the output is a classification label uniquely corresponding to the doctor's identity. During training, the cosine similarity loss function and the triplet loss function are combined to strengthen the model's ability to distinguish different doctor voiceprint features and the ability to aggregate the same doctor voiceprint. After the model is trained, it has the ability to accurately identify the corresponding doctor's identity and voice command under high-frequency device noise interference, ensuring the high reliability and safety of the voice control system in the complex acoustic environment of the operating room.

[0041] Aiming at the voiceprint signal after completing feature reconstruction and spectrum correction, a convolutional neural network (CNN) is used to extract deep features and classify voiceprint features, aiming to build a personalized voiceprint recognition model with strong robustness and high recognition. The specific method is as follows: First, the voiceprint signal after spectrum correction and band purification is converted into a Mel-frequency cepstral coefficient (MFCC) matrix, combined with the fundamental frequency (F0) sequence, and the formant frequency and its dynamic change trajectory reflecting the vocal tract shape, to form a multi-dimensional feature input matrix. This matrix not only retains the static spectral features of voiceprint, but also reflects the dynamic characteristics of sound production mechanism and individual anatomy, providing rich speech biometric features for the model to capture identity differences. Second, a multi-layer convolutional neural network is constructed. The first few convolutional layers extract local spectral changes and temporal dynamic features in voiceprint features through local perception and weight sharing mechanism. The pooling layer is used for dimension reduction and key feature preservation. The subsequent fully connected layer is responsible for mapping the features extracted by the multi-layer convolution to a high-dimensional identity discrimination vector. During training, a double loss function structure is designed, and cosine similarity loss function and triplet loss function are introduced. The former is used to measure the angle distance between the model output vector and the target doctor voiceprint vector in the vector space, to ensure that the multiple utterances of the same doctor are as close as possible in the vector space. The latter controls the distance between positive samples, negative samples and anchor samples, forcing the model to narrow the vector distance of the same doctor's voiceprint and widen the vector distance between different doctors, thereby improving the discrimination ability and convergence effect of the model. Finally, the model output is a classification label uniquely corresponding to the doctor's identity, realizing the accurate recognition of individual doctors. Through the comprehensive analysis of voiceprint data by deep learning, the CNN model can maintain high recognition accuracy and identity specificity when facing complex acoustic environment, noise interference and voice differences in the operating room, meeting the high safety requirements of identity verification and instruction response of voice control in the operating room.

[0042] The step is to solve the problem of overlapping and adjacent interference between the noise spectrum of the high-frequency surgical equipment in the operating room and the characteristic frequency band of the doctor's voiceprint during operation, and to ensure the accuracy and stability of voiceprint recognition in noisy environments. The harmonic noise released by the high-frequency equipment in the operating room at different powers and modes has a nonlinear overlap or proximity with the resonance peak, fundamental frequency and harmonic frequency band in the doctor's voiceprint, directly affecting the extraction and recognition of voiceprint features, and causing the recognition rate of the voiceprint model to decrease or even misjudge when disturbed by noise. By performing spectral mapping analysis on the doctor's voiceprint feature sequence based on the environmental noise baseline model, the specific frequency band of the voiceprint feature that overlaps or approaches the noise spectrum can be identified, and the voiceprint segment disturbed by noise can be accurately located. Then, through feature reconstruction and spectral correction, the easily disturbed segments are reconstructed, enhanced and purified, effectively restoring the integrity and continuity of the voiceprint signal and weakening or eliminating the erosion of noise components on the voiceprint discrimination characteristics. Finally, the generated personalized voiceprint recognition model not only accurately depicts the voiceprint characteristics of the doctor as an individual, but also has the ability to isolate and resist interference from the specific operating room noise environment, ensuring that identity recognition and voice command analysis can still be completed stably and efficiently in a dynamic environment with multiple noises and multiple people, improving the safety and reliability of the surgical voice control system.

[0043] S300, based on the personalized voiceprint recognition model, simultaneously decouples the spatial dimension and the spectral dimension of the multi-channel audio signals collected in real time in the operating room, extracts and separates independent doctor voice signals, device noise signals and background sound source signals, and dynamically marks each signal source to ensure one-to-one correspondence between the signal source and the personalized voiceprint recognition model; In view of the problems of complex audio source signals, various sound sources and mixed interference of device noise and doctor voice in the operating room environment, a dual-domain decoupling based on a personalized voiceprint recognition model is proposed to simultaneously decouple the spatial dimension and the spectral dimension of the multi-channel audio signals collected in real time, so as to accurately separate the doctor voice signals, device noise signals and background sound source signals, and dynamically mark each signal source to ensure one-to-one correspondence between the signal source and the personalized voiceprint recognition model. Specifically, the following steps are included: By deploying an array multi-channel pickup device in the operating room, audio signals from all directions are collected in real time in a uniformly distributed manner. For multi-channel audio signals, a beamforming algorithm is applied to accurately locate the sound source direction, and a spatial filter is used to enhance and isolate sound signals from different directions, realizing preliminary decoupling in the spatial dimension. By recording and tracking the spatial position information of each sound source, the physical positions of doctors, devices, backgrounds and other sound sources can be dynamically captured, forming sound source markers in the spatial dimension. This step provides a physical space basis for further separation in the spectral layer.

[0044] Beamforming algorithm is a kind of spatial signal processing technology based on multi-channel pickup array, the core purpose of which is to form a spatial filtering effect on the sound of a specific direction by weighting and superimposing the signals collected by multiple microphones in the time domain or frequency domain, so as to realize signal enhancement of the target sound source and suppression of noise in non-target direction. In the present application, the role of the beamforming algorithm is to realize the orientation determination and signal separation of the doctor's voice, equipment noise and background sound source through spatial filtering, laying a spatial foundation for subsequent fine decoupling in the frequency dimension. Commonly used beamforming algorithms include Delay-and-Sum, Minimum Variance Distortionless Response (MVDR), and adaptive beamforming such as Capon algorithm. The specific steps are as follows: Deploy the array microphone according to the preset layout in the operating room to form a geometrically symmetrical pickup network to collect audio signals in real time at different spatial positions; According to the spatial coordinates of the microphone array and the speed of sound, the time difference of the same sound source signal received by each microphone is calculated, and the phase of each channel signal is aligned through delay compensation to form a "beam pointing" to a specified direction; According to the MVDR or Capon algorithm, the optimal weight coefficient is calculated, and each channel signal is weighted and superimposed to enhance the signal pointing to the doctor's sound direction, while suppressing the signal in the non-target direction and reducing spatial interference; The signals after beamforming are partitioned in the spatial domain to form separate outputs of sound sources in different directions, completing the preliminary decoupling in the spatial dimension and providing orientation information and pure sound source input for subsequent personalized voiceprint recognition and spectral decoupling.

[0045] On the basis of completing the decoupling in the spatial dimension, spectral analysis is performed on each separated signal, and short-time Fourier transform and Mel frequency analysis method are used to extract the spectral features of the signal. Based on the pre-trained personalized voiceprint recognition model, the spectral features of all audio channels are compared with the feature templates in the doctor's personalized voiceprint model one by one. Through feature similarity calculation, the channel or spectral segment that matches the feature of the personalized voiceprint model is identified, and the spectral interval of the doctor's voice is accurately determined. At the same time, combined with the baseline model of the operating room environment noise, the feature components corresponding to the equipment noise in the spectrum are identified, and the remaining signals that are not classified are summarized as background sound source signals, forming complete decoupling in the frequency dimension.

[0046] The mel-frequency analysis method is a frequency spectrum feature extraction method designed based on the auditory perception mechanism of the human ear. By mapping the frequency features of the audio signal to the mel scale that conforms to the human ear perception, the extracted features can better reflect the essential properties of the human voice and individual voiceprint differences. In this step, the mel-frequency analysis method is used to analyze the frequency spectrum of the voiceprint signal in detail, especially to highlight the capture of the low-frequency formant and fundamental harmonic information in the doctor's voiceprint that is closely related to the identity features, and to improve the sensitivity and recognition of the personalized voiceprint recognition model to the frequency spectrum features. The specific steps include: first, the short-time Fourier transform (STFT) is performed on the collected time-domain audio signal, and the audio signal is divided into continuous short-time windows. The Fourier transform is performed on the signal in each window to obtain the frequency spectrum of the corresponding time segment, forming a time-frequency matrix of the signal. Second, the frequency spectrum energy of each time-frequency segment is filtered by a mel filter bank. The filter bank divides the frequency range according to the mel scale, with a dense low-frequency range and a sparse high-frequency range, simulating the sensitivity of the human ear to different frequencies. Third, the energy distribution after filtering is taken as the logarithm, converted into a log energy spectrum, and then subjected to discrete cosine transform (DCT) to convert the mel energy spectrum into a set of mel frequency cepstral coefficients (MFCC). This set of coefficients highly concentrates the identity features in the audio signal. Finally, the MFCC sequence is input as a frequency spectrum feature together with the fundamental frequency, formant, and other features to the personalized voiceprint recognition model for further learning and matching, achieving accurate modeling and recognition of the doctor's voiceprint in the frequency spectrum.

[0047] On the basis of spatial and spectral decoupling, a dynamic label is assigned to each type of separated signal. The dynamic label includes not only the spatial position information of the sound source, the spectral feature label, but also the identity label matching the personalized voiceprint model, the signal strength, the time stamp, and the continuity features of the signal. This dynamic labeling mechanism allows each signal to be accurately traced back to its source attribute and identity attribution in the subsequent processing and instruction analysis process, ensuring that the doctor's voice commands and the corresponding identity are one-to-one, avoiding source confusion or incorrect classification.

[0048] Based on the dynamic labeled signal, the mapping relationship between the signal source and the personalized voiceprint recognition model is maintained in real time. When the personnel in the operating room changes or the position of the sound source changes, the dynamic updating mechanism of spatial positioning and spectral features is combined to adjust the corresponding mapping of the sound source label and the voiceprint model in real time, ensuring continuous, stable, and accurate recognition of the doctor's voice signal and isolation of device noise and background sound sources.

[0049] The role of this step is to solve the problem of mixed interlacing of doctor's voice, device noise and background sound in the operating room under the environment of multiple sound sources and strong noise interference, to ensure that the effective voice signal of the doctor can be accurately separated and recognized in a complex acoustic environment, and each signal source after separation is uniquely bound with a personalized voiceprint recognition model, to ensure the accurate perception of the voice control system to identity and instructions. There are many sources of sound signals in the operating room, including the doctor's voice, high-frequency noise of various surgical devices, and background noise such as personnel conversation and instrument prompt sound. These signals are interlaced in spatial position, and the frequency components are complex and partially overlapped in frequency spectrum. If they are not separated, directly applying voiceprint recognition and speech analysis will easily cause recognition errors, mis-triggering of control instructions, and even confusion of permissions, affecting the safety and efficiency of the operation. Through this step, based on the prior knowledge of the personalized voiceprint recognition model, first, the spatial orientation and directivity filtering of the multi-channel audio signal is performed using beamforming algorithm in the spatial dimension, to preliminarily distinguish the physical position of the sound source; then, the signal features are extracted and recognized in detail in the frequency spectrum dimension combined with short-time Fourier transform and Mel frequency analysis technology, to realize accurate decoupling in the frequency dimension and completely distinguish the doctor's voice, device noise and background sound source. Finally, each separated signal source is dynamically marked, and the mapping relationship of the spatial position, spectral features, time stamp and personalized voiceprint model of the sound source is recorded, to ensure that the subsequent voice recognition and control system can clearly, stably and dynamically track and recognize the effective voice signal of the doctor, improve the response accuracy of the instructions and the identity security, and realize the stable and reliable operation of the voice control in the high-noise operating environment.

[0050] S400, based on the doctor's voice signal sequence after dynamic marking, performing voice integrity detection, the detection content including the continuity, speech speed stability and audio waveform consistency of the voice signal, eliminating non-natural voice signals that do not meet the integrity condition, to avoid noise signals being misjudged as effective voice instructions; In view of the problem that the doctor's voice signal in the operating room may be affected by device interference, non-voice signal mixing or signal fragment loss under the environment of multiple sound sources and strong noise background, voice integrity detection based on the doctor's voice signal sequence after dynamic marking is proposed. Through comprehensive detection of the continuity, speech speed stability and audio waveform consistency of the voice signal, non-natural voice signals are accurately screened out, to ensure that the signal source for subsequent voiceprint recognition and speech instruction analysis is real, stable and complete. Specifically, the following steps are included: For the dynamic marked doctor speech signal sequence, signal continuity detection is performed. Specifically, through time series analysis of the audio signal, the time interval between adjacent speech segments is calculated, and the zero crossing rate detection and short-time energy change curve are used to evaluate the continuity of the signal and the speech breakpoint condition. If there are abnormal transient interruptions, energy drops or silent intervals in the signal, and these interruptions do not conform to the speech pause rules of natural human speech (usually natural pauses are between 150 milliseconds and 250 milliseconds), it is determined that the signal has non-natural interruptions. Through continuity detection, sudden signal interruptions caused by noise impact, electromagnetic interference or device operation can be eliminated, ensuring that the detection signal is continuous and complete in time.

[0051] Based on the time series data of the signal, the speech speed stability is detected. By performing dynamic time warping (DTW) analysis on the phoneme duration, syllable interval and rhythm rhythm of the doctor's speech signal, the stability of the pronunciation rate in the entire speech signal sequence is evaluated. Human natural speech has a relatively consistent pronunciation rhythm and rate in a specific scenario. If there is an abnormal acceleration, slowing down or irregular jump in the signal sequence (for example, the syllable interval changes dramatically within a short time more than twice the standard deviation), it indicates that the signal may be disturbed by noise fragments or non-natural signals. Through speech speed stability detection, signals disguised as speech due to device mechanical sound or background continuous noise can be excluded, ensuring that the pronunciation rhythm of the detection signal is consistent with the normal speech behavior of the doctor.

[0052] The audio waveform consistency detection is performed on the dynamic marked signal sequence. Specifically, the short-time autocorrelation function and waveform envelope analysis method of the signal are used to compare the waveform form, energy distribution and frequency component of the speech signal globally and locally. Normal human voice has stable fundamental frequency amplitude, harmonic sequence and unique envelope curve. If the signal waveform is detected to have non-periodic jump, fundamental frequency continuity interruption or harmonic distortion, such as abnormal peak in a certain frequency band or abnormal energy concentration exceeding the upper limit frequency of human voice (for example, greater than 8kHz), it is determined that the signal has pseudo-signal characteristics that do not conform to natural speech. Through audio waveform consistency detection, device noise, sudden impact sound or abnormal audio in the background environment can be effectively prevented from being mistaken for doctor speech.

[0053] Short-time autocorrelation function and waveform envelope analysis are two key techniques for evaluating the stability and continuity of speech signals in time and frequency structure. In this step, short-time autocorrelation function is used to measure the periodicity and fundamental frequency stability of the signal within a short time window, by calculating the similarity of the audio signal with itself at different time delays, to identify whether there is a stable fundamental frequency and harmonic structure in the speech signal; while the waveform envelope analysis method extracts the amplitude envelope of the audio signal, depicting the energy curve of the signal over time, so as to reflect whether the overall and local energy distribution of the speech signal is smooth and consistent with the characteristics of human voice. The specific steps include: The speech signal to be detected is divided into continuous short-time windows, such as 20 milliseconds per frame, and autocorrelation function operation is performed on the signal in each frame to evaluate the similarity peak of the signal at different delays, so as to determine whether the fundamental frequency exists and changes continuously and smoothly; if the autocorrelation peak of some frames is lost or unstable, it means that there is a possibility of non-periodicity or non-natural speech in that frame.

[0054] The amplitude envelope of each frame of signal is extracted using Hilbert transform to form a waveform envelope curve, and then the smoothness, peak-to-valley interval and energy fluctuation amplitude of the global waveform envelope and local segment envelope curve are compared to detect whether there is a sudden change, excessive flatness or abnormal jitter of non-natural energy distribution.

[0055] The detection results of short-time autocorrelation and waveform envelope are fused, and when the conditions of periodic stability of fundamental frequency, smoothness of waveform envelope and natural energy fluctuation are met at the same time, the signal is determined to be natural speech, otherwise it is marked as suspected noise or abnormal sound source. This method realizes the global and local consistency detection of speech signals from the aspects of periodicity, energy dynamics and frequency components, ensuring the naturalness and integrity of the identified signal.

[0056] Signals that meet the natural speech standards through continuity detection, speech rate stability detection and audio waveform consistency detection are determined to be complete and effective doctor's speech signals, and are retained for subsequent voiceprint recognition and speech command analysis. Signals that do not pass the integrity detection are automatically marked as invalid signals or noise segments and do not participate in subsequent recognition and command response.

[0057] The role of this step is to ensure that the doctor's voice signal obtained by dynamic marking in the operating room environment has authenticity, stability and integrity before being used for subsequent voiceprint recognition and instruction analysis, avoiding misidentification or instruction mistriggering due to factors such as device noise, background sound interference or signal fragment loss. The acoustic environment in the operating room is complex, not only does it have multiple sound source voices of doctors and assistants, but also non-voice signals generated by various medical devices in operation, such as continuous high-frequency noise emitted by ultrasonic knives and electrocoagulation instruments, or occasional metal collision sounds and instrument alarm sounds. If these signals are not subjected to integrity detection after being dynamically marked as suspected voice, they are likely to be misjudged as valid doctor's instructions. Through this step, the dynamically marked voice signal sequence is comprehensively detected in terms of continuity, speech rate stability and audio waveform consistency, which can check whether there are unnatural interruptions, sudden changes in speech rate or abnormal waveforms in the signal that do not conform to the characteristics of natural human voice. Only those signals that are continuous in the time axis, have stable rhythm and have high consistency with human voice characteristics will be determined as complete and reliable doctor's voice. This not only effectively eliminates pseudo-voice caused by electromagnetic interference, signal mutation or device noise mixing, but also ensures that the subsequent voiceprint recognition model is based only on real and complete doctor's voice for identity judgment and instruction analysis, greatly improving the safety and accuracy of the voice control system, reducing the risk of mistriggering medical devices, and ensuring the continuity and safety of the operation process.

[0058] S500, for the valid voice signal detected by the voice integrity, implement voiceprint identity verification and non-contact action sensing dual verification, non-contact action sensing based on gesture or body movement data synchronized with voice, through dual verification to confirm the identity of the initiator and the consistency of the behavior of the instruction; For the valid doctor's voice signal detected by the voice integrity, a dual verification based on voiceprint identity verification and non-contact action sensing is proposed to ensure that the doctor's voice instruction is confirmed in both identity and behavior dimensions, thereby avoiding misidentification and misexecution in the high safety requirement environment of the operating room. The following steps are implemented: For the doctor's voice signal that has passed the voice integrity detection, perform voiceprint identity verification. The specific operation is: use the previously trained personalized voiceprint recognition model for each doctor to extract the Mel frequency cepstrum coefficient, fundamental frequency, formant frequency and dynamic change of the voice signal, and compare them one by one with the feature templates stored in the voiceprint model. A dual evaluation mechanism of cosine similarity and Euclidean distance is adopted to calculate the similarity score between the current voice signal and the target doctor's voiceprint features. If the similarity score is higher than the preset identity verification threshold, it is preliminarily confirmed that the instruction is issued by the target doctor. Through this step, the voiceprint level identity of the issuer is confirmed, and the interference of others misfiring or similar voice in the background is eliminated.

[0059] The cosine similarity and Euclidean distance dual evaluation mechanism is a composite judgment method for measuring the similarity of two feature vectors in a multi-dimensional space, aiming to more comprehensively and accurately evaluate the matching degree between the current speech signal and the target doctor's voiceprint features through different distance and angle measurement standards. In the present application, the mechanism plays a role in avoiding the limitations of a single indicator on similarity evaluation through dual discrimination, ensuring the accuracy and robustness of voiceprint authentication. The specific steps are as follows: The multi-dimensional acoustic features such as mel-frequency cepstral coefficients, fundamental frequency, and formant frequency of the current doctor's speech signal passing through the integrity detection are extracted to form a feature vector. The voiceprint feature vector stored in the personalized voiceprint recognition model of the target doctor is called as a reference. Secondly, the cosine similarity between the current feature vector and the reference vector is calculated, and the closer the value is to 1, the more consistent the direction of the two vectors is, i.e. the higher the similarity.

[0060] Based on the same two feature vectors, the Euclidean distance between them is further calculated to measure the straight-line distance between the two vectors in a multi-dimensional space, and the smaller the distance, the closer the actual values of the two voiceprint features.

[0061] The results of cosine similarity and Euclidean distance are normalized and combined with the set weight factor to calculate the comprehensive similarity score. When the score exceeds the preset threshold, it is determined that the speaker of the current speech signal is consistent with the target doctor, otherwise it is determined that the identity is not consistent and the subsequent instruction response is rejected. Through this dual evaluation mechanism, the angle similarity and numerical closeness are guaranteed from two levels, effectively improving the accuracy and anti-interference ability of voiceprint identity authentication.

[0062] On the basis of completing voiceprint identity authentication, a non-contact motion sensing mechanism is started in real time. Specifically, through the infrared depth camera and millimeter wave radar equipment arranged in the operating room, the gestures, body movements and body posture dynamic data of the doctor at the time of speaking are synchronously captured. Using the motion recognition model constructed based on convolutional neural network and long short-term memory network (LSTM), the captured motion data is feature extracted to identify whether the doctor is accompanied by specific gestures or body posture instructions when speaking. For example, when the doctor issues an instruction to adjust the surgical light or endoscope rotation, the corresponding actions such as pointing, swinging or stretching should be accompanied. Through the synchronization analysis of voice and action, it is verified whether the behavior and voice instruction are logically consistent.

[0063] The action recognition model based on convolutional neural network and long short-term memory (LSTM) is a deep learning method combining spatial feature extraction and time series modeling, which is dedicated to joint recognition and analysis of spatial posture and time variation in dynamic video or action data. In this step, the role of the model is to extract high-dimensional features and recognize dynamic behaviors of the doctor's gesture or body action sequence captured synchronously by the infrared depth camera and millimeter wave radar, to determine whether the doctor is accompanied by specific action instructions when speaking, so as to realize the dual instruction verification of voice and action. The specific construction and application steps are as follows: The continuously captured doctor action data (such as skeleton joint coordinates, limb posture, joint angle change, etc.) are sorted in time sequence into multiple frames of action sequences to form three-dimensional tensor input.

[0064] The convolutional neural network performs convolution operation on each frame of action image or skeleton data to extract spatial features such as gesture shape, limb stretching amplitude, action posture, etc., forming a spatial feature vector.

[0065] The spatial features of each frame extracted by the convolutional network are input into the long short-term memory network (LSTM) in time sequence, which takes advantage of its time series dependence modeling to capture the dynamic changes and evolution patterns of the action sequence in the time dimension, such as the continuity and rhythm of finger pointing, palm waving or torso twisting, and then understand the time sequence structure and behavior intention of the action.

[0066] The output of the LSTM is matched with the pre-trained action instruction label to determine whether the current action sequence is consistent with the gesture or body posture corresponding to the specific voice instruction. If the recognized action category is consistent with the voice instruction logic and the time synchronization is within the preset threshold, the behavior consistency of the instruction is confirmed. Through this model, not only static postures can be accurately identified, but also the dynamic change process of actions can be analyzed, realizing the synchronous verification of doctor's action and voice when issuing orders, and ensuring the behavior safety and identity accuracy of voice control instructions in the operating room.

[0067] Based on the voiceprint verification result and the action perception result, the consistency comparison of dual verification is performed. Specifically, the identity confirmation result of voiceprint verification and the doctor's identity and action category identified by action perception are cross-checked, and the synchronization of the time of speaking and the time of action is evaluated. The time deviation of the two is controlled within 300 milliseconds to ensure that the sound and action indeed come from the same doctor and are completed under the same behavior intention. If the identity verification and behavior perception results are consistent and the time synchronization meets the requirements, the instruction is confirmed as a valid instruction with dual reliable guarantee of identity and behavior; otherwise, the instruction is marked as verification failed and the subsequent control response is refused.

[0068] For the valid voice command that has passed the double verification, record its voiceprint features, action features, timestamp and doctor identity tag, and store them into the operation log together with the command content as the basis for post-audit and safety traceability. At the same time, based on the cumulative verification results, dynamically optimize the verification model, continuously learn and adapt to the pronunciation habits and action patterns of individual doctors, and improve the recognition accuracy and robustness of the model in the actual operation scene.

[0069] The role of this step is to comprehensively improve the identity security and behavior consistency of voice commands in the application scenario of intelligent voice control in the operating room through the double verification mechanism of voiceprint identity verification and non-contact action sensing, to ensure that the initiator of the command and its behavior intention can be accurately confirmed before receiving and executing the voice command, thereby effectively preventing the risk of misidentification, mis-triggering or malicious interference. The operating room is a place with multiple personnel, complex noise and dynamic environment, and relying solely on voice recognition may have some uncertainty, such as background noise, imitation by others, and interference from device noise, which may lead to misjudgment, so a voice-synchronized action sensing mechanism needs to be introduced as a supplementary verification means. By capturing the gestures or body movements of the doctor in real time when speaking and analyzing them synchronously with the voice command, it can be ensured that the voice command is not only from the authorized doctor himself, but also accompanied by limb movements or body changes consistent with the command content. This multi-modal cross-verification greatly improves the rigor of command verification. Through the double verification mechanism, it can not only confirm that the initiator is the doctor himself, but also verify whether he has performed the action operation consistent with the command when speaking, such as pointing to a device, waving his hand or other specified gestures, thereby effectively preventing potential safety hazards caused by accidental triggering of voice by others or misidentification of device control, ensuring the rigor, safety and reliability of voice control during the entire operation process, and meeting the high requirements of precision and authority control for surgical operations.

[0070] S600, based on the valid command confirmed by double verification, perform permission level comparison and context semantic verification, determine the compliance of the command according to the permission level of the initiator and the operation logic of the surgical scene, control the medical device to execute the corresponding command after compliance, complete the closed-loop process from noise isolation, signal decoupling, identity and action verification, permission and semantic verification to device control; To ensure the safety and compliance of voice control commands in the operating room, the effective command confirmed based on double verification is proposed to perform permission level comparison and context semantic verification, through a multi-level review mechanism, to ensure that only the command that has passed complete verification, has the correct permission and reasonable semantics can drive the medical device to execute the operation, constituting a full-process closed loop from noise isolation, signal decoupling, identity and action verification, permission and semantic verification to device control. Specifically, the following steps are included: For valid instructions that have passed through voiceprint authentication and non-contact motion perception double verification, perform permission level comparison. Permission level comparison is based on the operation permission list and level system set in advance in the medical management platform for each doctor, assistant and other surgical participants, such as the roles of the main surgeon, first assistant, and circulating nurse, which correspond to different permission levels and types of executable devices. After receiving the instruction, immediately extract the identity information of the issuer from the identity verification result and compare it with the permission level corresponding to the identity in the permission database to check whether it has the control permission of the current operation device and whether it is allowed to execute the instruction at the current surgical stage. If the permission level is insufficient, the instruction is directly blocked from being issued to prevent unauthorized personnel from operating medical devices beyond their authority.

[0071] For instructions that pass the permission comparison, perform context semantic verification. The specific method is to convert the instruction content into structured semantic units, combine the operation log, device status information and surgical process stage of the current surgical scene, and verify whether the semantics of the instruction are consistent with the current context of the surgery. For example, if the current stage is blood vessel anastomosis, the doctor's instruction to "turn up the surgical light" needs to confirm that the light brightness can indeed be adjusted at this stage, and the adjustment behavior will not affect the sterile environment and safety of the surgical field. Conversely, if the surgical link is not suitable for adjusting the light, the instruction semantics will be determined to be non-compliant, and thus execution will be refused. During the context semantic verification process, the current state of the device (such as whether it is already running, and whether the parameter has reached the upper limit) and the surgical process model are combined to dynamically evaluate the rationality of the instruction.

[0072] Based on the instructions that pass both permission level comparison and context semantic verification, the final execution permission release is performed. In this step, an instruction control package containing the instruction content, the identity of the issuing doctor, the permission level, the semantic verification result and the timestamp is generated, and is securely encrypted and transmitted to the target medical device. After receiving the instruction control package, the medical device again confirms the integrity and legality of the instruction package according to the built-in decoding and verification mechanism, and only after confirming that it is correct, it starts the corresponding function execution, such as adjusting the brightness of the surgical light, rotating the endoscope view angle or controlling the power of the electrocoagulation instrument, etc., to ensure that every device action has a traceable and verifiable compliance basis.

[0073] For each executed control instruction, all log information in the whole process of permission comparison, semantic verification, instruction issuance and device response is recorded synchronously to form a complete instruction execution chain. This log data is uploaded in real time to the surgical data recording platform for post-surgery review, quality control and safety audit. At the same time, based on the accumulated instruction execution and verification data, the permission and semantic verification models are continuously optimized to dynamically adjust the adaptability of the permission level and the surgical process model to meet the individual control needs of different surgical types and different doctor operation habits.

[0074] The role of this step is to establish a security line based on permission level comparison and context semantic verification in the application scenario of intelligent voice control in the operating room, to ensure that even after the voice signal is isolated from noise, decoupled from the signal, and double-verified for identity and action, it still needs to pass the double review of permission and semantics to finally trigger the actual operation of the medical equipment. The operating room is a place with very high requirements for safety and standards. The operation permissions of different medical personnel are clearly distinguished. For example, the roles of the lead surgeon, assistants, and circulating nurses have different ranges of operable equipment and permissions at different stages of the operation. Without permission control, there is a risk of misactivation by low-privilege personnel or others using voice prints. At the same time, the rationality of the voice command must also match the current process, patient status, and current device status to avoid semantic conflicts or triggering inappropriate device behavior at inappropriate stages. By further performing permission level comparison on valid commands, it can quickly determine whether the issuer has the authorized qualification to operate the device. Then, combined with the context information, current device status, and surgical process node during the operation, the semantics of the command are dynamically analyzed and verified for rationality. Only when the permission is compliant and the semantics match the scenario, the command is allowed to be executed. The implementation of this step ensures that the intelligent voice control system in the operating room has a complete security chain of "prevention, verification, and traceability", effectively avoiding operation errors and medical risks caused by mismatched command permissions, inappropriate semantics, or process errors, forming a full-process, closed-loop, controllable command management mechanism from acoustic signal processing to device control, and significantly improving the reliability and safety of voice control in high-risk medical scenarios.

[0075] Through the above non-contact intelligent voice control method in the operating room supporting multi-role voiceprint recognition, high-robustness voice recognition and safe control in a complex noise environment in the operating room can be achieved, and the recognition accuracy and execution safety of voice commands are significantly improved. The method effectively isolates the spectral interference of high-frequency surgical device noise and doctor voiceprints by dynamically establishing and updating an environmental noise baseline model, and realizes accurate separation of doctor voice and background noise by combining personalized voiceprint recognition models and spatial and spectral decoupling. Through voice integrity detection and voiceprint and action double verification, it ensures that the command issuer is real and consistent in behavior, and through permission level comparison and context semantic verification, it strictly limits the compliance of the command in terms of permission and surgical scenario, and finally forms a closed-loop defense mechanism from noise isolation, signal decoupling to device safety control. This method effectively prevents the risk of misidentification and miscontrol caused by device noise or unauthorized voice, and ensures the continuity, accuracy, and patient safety of the operation process.

[0076] It is apparent that for the person skilled in the art, many modifications and changes can be made to the embodiments described without departing from the spirit and scope of the application. It is therefore understood that the above-described drawings and descriptions are illustrative in nature and should not be considered limiting the scope of the present application.

[0077] It should be noted that, in this text, relationship terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or equipment including the element.

[0078] It should be understood that in various embodiments of the present application, the size of the sequence number of each process described above does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0079] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0080] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0081] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Part or all of the units can be selected to achieve the purpose of the present embodiment according to actual needs.

[0082] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit.

[0083] The above merely describes some exemplary embodiments of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0084] The above only describes some exemplary embodiments of the present application by way of illustration, and it is needless to say that the described embodiments can be modified in various ways without departing from the spirit and scope of the present application for those skilled in the art. Therefore, the above drawings and descriptions are illustrative in nature and should not be understood as limiting the protection scope of the claims of the present application.

Claims

1. A non-contact intelligent voice control method for operating rooms supporting multi-role voiceprint recognition, characterized in that: Includes the following steps: S100 collects full-band noise signals of all high-frequency surgical equipment in the operating room under different operating conditions, extracts harmonic frequency, amplitude and phase characteristics, constructs a harmonic noise feature database, and establishes and dynamically updates the environmental noise baseline model. S200, based on an environmental noise baseline model, performs spectral mapping on the voiceprint feature sequence of each doctor, identifies voiceprint segments with overlapping or adjacent spectra, performs feature reconstruction and spectral correction, and generates a personalized voiceprint recognition model. S300, based on a personalized voiceprint recognition model, performs spatial and spectral dual-domain decoupling on real-time acquired multi-channel audio signals, separates doctor's voice, equipment noise and background sound sources, and implements dynamic marking of signal sources to ensure that signal sources correspond one-to-one with personalized voiceprint recognition models. S400, based on dynamically tagged doctor's speech signals, performs speech integrity detection, detecting continuity, speech rate stability and audio waveform consistency, and eliminating non-natural speech signals; The S500 verifies the identity of the command initiator and the consistency of their behavior by using dual verification based on voiceprint authentication and contactless motion perception for signals that pass the voice integrity detection. The S600, based on instructions that have passed dual verification, performs permission level comparison and context semantic verification. After compliance, it controls the medical device to execute, completing a closed-loop process from noise isolation, signal decoupling, identity and action verification, permission and semantic verification to device control.

2. The non-contact intelligent voice control method for operating rooms supporting multi-role voiceprint recognition according to claim 1, characterized in that, Step S100 includes: Collect full-band noise signals of all high-frequency surgical equipment in the operating room under different operating conditions, extract harmonic frequency, amplitude and phase characteristics, and construct a harmonic noise feature database; Perform multidimensional feature clustering on the feature vectors in the noise feature database to classify noise categories; Perform feature mapping dimensionality reduction on the feature vectors of each category, and establish an index mapping between features and device type, power, and mode; Based on the clustering and mapping results, an environmental noise baseline model is formed, and a dynamic update mechanism is set.

3. The non-contact intelligent voice control method for operating rooms supporting multi-role voiceprint recognition according to claim 1, characterized in that, Step S200 includes: Based on the environmental noise baseline model, the doctor's voiceprint feature sequence is spectrally mapped to identify segments with overlapping or adjacent spectra. For overlapping segments, feature reconstruction is performed based on variational autoencoder, and phase information is compensated by minimum phase reconstruction; The reconstructed signal is spectrally corrected using a minimum mean square error filter, and the resonant peaks and fundamental harmonics are enhanced through cepstral analysis. Based on the corrected signal, features are extracted and classified using a convolutional neural network to generate a personalized voiceprint recognition model.

4. The non-contact intelligent voice control method for operating rooms supporting multi-role voiceprint recognition according to claim 3, characterized in that, Methods for enhancing resonant peaks and fundamental harmonics using cepstral analysis include: Perform a Fast Fourier Transform on the spectrum-corrected voiceprint signal to obtain the logarithmic spectrum; Perform an inverse Fourier transform on the logarithmic spectrum to obtain the cepstral signal; Low-pass cepstral filtering is applied to the cepstral signal to preserve low- and mid-order cepstral coefficients and suppress non-identity features in high-frequency cepstral components. The filtered cepstral coefficients are inversely transformed back to the frequency domain to obtain the acoustic signature signals of the enhanced formant and fundamental harmonic.

5. The non-contact intelligent voice control method for operating rooms supporting multi-role voiceprint recognition according to claim 1, characterized in that, Step S300 includes: Spatial localization is performed on multi-channel audio signals by beamforming algorithm, and spatial decoupling is achieved by using spatial filter to enhance the target direction signal and suppress other direction signals. Short-time Fourier transform and Mel frequency analysis are performed on the spatially decoupled signal to extract spectral features; Based on the personalized voiceprint recognition model and the environmental noise baseline model, spectral features are compared and classified to achieve spectral dimension decoupling; Dynamic labels are assigned to the decoupled signals to establish a correspondence between the signal source and the voiceprint recognition model.

6. The non-contact intelligent voice control method for operating rooms supporting multi-role voiceprint recognition according to claim 5, characterized in that, The specific methods of beamforming algorithms include: The time difference of each channel signal is calculated based on the spatial coordinates and sound speed of the microphone array, and the signal phase is aligned by delay compensation. The minimum variance distortionless response algorithm is used to calculate the optimal weighting coefficients, and the signals of each channel are weighted and superimposed to enhance the doctor's voice signal in the target direction; The weighted signal is partitioned according to spatial direction to achieve separate output of different sound sources.

7. The non-contact intelligent voice control method for operating rooms supporting multi-role voiceprint recognition according to claim 1, characterized in that, Step S400 includes: Perform time interval and zero-crossing rate detection on the speech signal sequence to determine signal continuity; The stability of speech rate is determined by analyzing phoneme duration and syllable interval based on the dynamic time warping algorithm. Perform short-time autocorrelation function and Hilbert transform on the signal to extract fundamental frequency periodicity and waveform envelope curve, and determine the consistency of audio waveform; By fusing the results of continuity, speech rate stability, and audio waveform consistency detection, non-natural speech signals are screened out.

8. The non-contact intelligent voice control method for operating rooms supporting multi-role voiceprint recognition according to claim 1, characterized in that, Step S500 includes: Based on a personalized voiceprint recognition model, a dual evaluation mechanism of cosine similarity and Euclidean distance is used to verify the identity of the giver. Based on motion data collected by infrared depth cameras and millimeter-wave radar, the characteristics and categories of doctors' movements are identified through convolutional neural networks and long short-term memory networks. A consistency comparison is performed between voiceprint identity and action category to assess the temporal synchronization between speech and action. After confirming that the identity and behavior are consistent, the voiceprint characteristics, action characteristics and timestamp of the command are recorded.

9. The non-contact intelligent voice control method for operating rooms supporting multi-role voiceprint recognition according to claim 1, characterized in that, Step S600 includes: Based on the issuer's identity information, query the permission database to verify their permission level for operating the device and their permissions during the surgical phase; The instruction content is transformed into semantic units, and the consistency between the instruction semantics and the current surgical context is determined by combining the operation log, equipment status and surgical procedure model. For instructions that pass both permission level and semantic verification, generate an encrypted control packet containing the instruction content, identity, permission level, semantic verification result, and timestamp; The control package is sent to the medical device, which decodes and verifies it before executing the instructions and recording the entire process in a log.

Citation Information

Cited By

  • Hysteroscopic surgery robot control method and system based on multi-modal perception

    CN121465745A

  • Voice instruction dynamic recognition and separation method based on multi-person voice scene

    CN121565172A

  • Voiceprint recognition method and system for vehicle-mounted child safety seat

    CN122050367A