Doctor and nurse noise reduction intercom system oriented to robot operation environment
By building a cooperative relationship network of medical teams, eye movement multi-objective tracking and intent recognition, multi-modal cooperative intent reasoning and adaptive microphone array beamforming, the complex problems of noise interference and team collaboration in the robotic surgical environment are solved, and efficient and stable intercom communication between doctors and nurses is achieved.
Patent Information
- Application Number
- CN202510744338.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
Smart Images

Figure CN120264167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical communication technology, and more specifically, to a noise reduction intercom system for doctors and nurses in a robot surgery environment. Background Art
[0002] In a robotic surgery environment, doctors and nurses need to maintain efficient and clear intercom communication during the operation to ensure the safety of the operation. However, the existing technology has the following problems: There are many noise sources in the operating room, such as ventilators, monitors, and robot movements, and medical staff need to move during the operation, resulting in unstable intercom communication quality; large and complex operations involve the collaboration of medical staff from multiple professional fields, forming a complex network of professional division of labor and collaborative relationships, and traditional intercom systems lack the ability to understand and support this complex team collaboration model; medical staff need to keep their hands sterile during the operation and cannot easily operate traditional contact communication equipment; existing systems mainly rely on voice wake-up or gesture control, with a high rate of misoperation and easy to destroy the sterile state of the surgical environment; the communication needs between team members change dynamically with the stage of the operation, and the traditional fixed-mode communication system is not adaptable enough and cannot flexibly adjust the communication strategy according to real-time scenario requirements; in complex surgical environments, the communication matching accuracy of traditional systems is low, and they lack the ability to perceive changes in the surgical stage.
[0003] Therefore, a technology is needed to solve the communication problems in the complex environment of the operating room through innovative technical means. Summary of the invention
[0004] The present invention provides a noise reduction intercom system for doctors and nurses in a robotic surgery environment, which solves the technical problems of noise interference in the operating room, complex team collaboration, and limitations on aseptic operations in related technologies.
[0005] The present invention provides a noise reduction intercom system for doctors and nurses in a robotic surgery environment, comprising: The medical team collaboration network construction module is used to construct a professional collaboration relationship map for the surgical environment through graph structure representation and role relationship analysis; Eye movement multi-target tracking and intention recognition module, which is used to monitor the eye activities of medical staff through a high-precision long-distance eye tracking camera and identify communication intentions and targets; Multimodal collaboration intention reasoning and communication mode matching module, which is used to combine eye movement interaction features with the collaboration relationship network to infer collaboration intentions and select the most suitable communication mode; Adaptive microphone array beamforming module, used to dynamically adjust the beam direction and parameters of the microphone array according to the communication mode and participant location information; A virtual meeting space construction and dynamic optimization module, which is used to create a virtual meeting space for multi-person collaboration scenarios, provide a unified logical framework for multi-person collaborative communication, and adaptively adjust communication parameters according to changes in the surgical stage.
[0006] In a preferred embodiment, the medical staff collaboration relationship network construction module includes: Collect medical staff role and relationship characteristics, including communication frequency characteristics, communication importance characteristics, and communication dependence characteristics; Construct a collaboration relationship graph: ; Among them, represents the collaboration relationship graph, the vertex set represents each role in the medical staff team, and the edge set represents the collaboration relationship between roles; Dynamically adjust the edge attributes in the collaboration relationship graph according to the real-time identified surgical stage.
[0007] In a preferred embodiment, the eye movement multi-target tracking and intention recognition module includes: Use an infrared illumination unit and a high-frame-rate imaging unit to collect the eye movement data of medical staff; Construct a multi-level eye movement intention recognition model, including a feature extraction layer, a context fusion layer, and an intention judgment layer; Determine the communication target by analyzing the gaze direction and the positions of medical staff in the operating room, and evaluate the communication urgency according to the gaze duration, pupil dilation rate, and eye movement pattern.
[0008] In a preferred embodiment, the multi-modal collaboration intention reasoning and communication mode matching module includes: Extract multi-target eye movement interaction features, including paired gaze features, mutual gaze behavior recognition, and gaze sequence analysis; Construct a collaboration intention reasoning model to calculate the probabilities of different collaboration scenarios; Select the most suitable communication mode from the point-to-point direct connection mode, the group broadcast mode, the sequential transmission mode, and the collaborative session mode.
[0009] In a preferred embodiment, the virtual meeting space construction and dynamic optimization module includes: Create a virtual meeting space model, including space topology construction, space node configuration, and session management mechanism; Optimize each communication channel according to the positions of medical staff and the environmental noise distribution; Through the analysis of historical communication data, construct a team collaboration mode model and predict future possible communication needs; Adaptive adjustment of the communication strategy according to the characteristics of different surgical stages.
[0010] In a preferred embodiment, the collaborative intention reasoning model is based on a graph neural network and a temporal attention mechanism. The input includes an eye movement interaction feature matrix, role relationship weights, and current surgical stage information, and the output includes a set of medical staff participating in the collaboration, collaboration types, and collaboration urgency.
[0011] In a preferred embodiment, the communication mode matching is based on the following rules: If the number of participants is 2 and the collaboration type is inquiry-to-answer or instruction-to-execution, select the point-to-point direct connection mode; If the number of participants is greater than 2 and there is a clear information initiator, select the group broadcast mode; If the collaboration type is information flow and there is a sequential dependency, select the sequential transmission mode; If the number of participants is greater than 2 and the collaboration type is discussion and decision-making, select the collaborative session mode.
[0012] In a preferred embodiment, the adaptive microphone array beamforming module includes: Deploy multiple sub-microphone arrays, each sub-array containing multiple omnidirectional microphones, to form a collaborative capture network; Combine visual and acoustic information for precise positioning and tracking of medical staff; Based on the communication mode and participant location information, implement beamforming algorithms including spatial filter design, minimum variance distortionless response beamforming, multi-target beamforming, and adaptive interference suppression; According to the movement of medical staff, use Kalman filtering for position prediction to achieve dynamic tracking and adjustment of the beam.
[0013] In a preferred embodiment, it further includes: A multi-modal perception unit, including multiple microphone array sub-arrays installed on the ceiling and walls of the operating room, high-precision long-distance eye movement tracking cameras, high-definition cameras, and depth sensors; A signal processing unit, including a dedicated DSP chip for real-time signal processing, a microphone array controller, and a beamforming processor; A network transmission unit for transmitting the data collected by each unit to the central processing system to achieve fusion analysis of multi-modal data.
[0014] In a preferred embodiment, a computer-readable storage medium is used to store computer-readable instructions that can run a noise reduction intercom system for doctors and nurses in a robotic surgery environment when the computer-readable instructions are read by a computer.
[0015] The beneficial effects of the present invention are as follows: Improve communication quality and efficiency: The accuracy of speaker localization in the mobile state is improved, reducing the system response latency, enhancing the effective sound capture gain, and maintaining good communication quality in the complex surgical environment with noise.
[0016] Achieve contactless control and reduce surgical risks: Realize contactless communication control through eye tracking, maintain the sterility of the surgical environment, improve the accuracy of intention recognition, can simultaneously track the gaze behaviors of multiple medical staff, and support complex teamwork.
[0017] Optimize teamwork and reduce cognitive load: Reduce the latency of key information transmission, improve the information accuracy, reduce the cognitive load of team communication, enable medical staff to focus more on the surgery itself, improve the effective signal-to-noise ratio in the surgical environment, and relieve the auditory fatigue of medical staff.
[0018] Have adaptability and scalability: The system can automatically adapt to different surgical environments and team compositions, continuously improve the communication matching accuracy, support multiple communication modes and dynamic switching mechanisms, meet the needs of surgeries with different complexities, adopt a modular design, and can be customized and deployed according to the specific hospital environment and surgical type. Description of the Drawings
[0019] Figure 1 It is a module diagram of the noise reduction intercom system for doctors and nurses in the robot surgical environment of the present invention. Detailed Embodiments
[0020] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.
[0021] In at least one embodiment of the present invention, a noise reduction intercom system for doctors and nurses in the robot surgical environment is disclosed, as Figure 1 shown, including: A medical staff team collaboration relationship network construction module, used to construct a professional collaboration relationship map for the surgical environment through graph structure representation and role relationship analysis; Specifically, it includes the following steps: Step 1.1, collect the characteristics of medical staff roles and relationships; Collect the basic information and professional role definitions of various medical staff in the surgical environment (including the surgeon, assistant physician, anesthesiologist, circulating nurse, etc.), and determine the responsibilities and communication needs of each role in different surgical stages. For each pair of medical staff role combinations, calculate the following relationship characteristics: Communication frequency characteristic: Record the typical communication frequency between specific role pairs in different surgical stages, expressed as , where and represent role i and role j respectively, represents surgical stage k; Communication importance characteristic: Assign importance weights to each pair of role relationships according to the criticality and safety dependence of surgical operations, expressed as , where and represent role i and role j respectively, represents surgical stage k; Communication dependence characteristic: Analyze which role relationships are of mandatory dependence type (such as the anesthesiologist reporting vital signs to the surgeon), and which are of auxiliary type (such as information sharing among circulating nurses), expressed as , where and represent role i and role j respectively, represents surgical stage k.
[0022] Step 1.2, Construction of the collaboration relationship graph; Based on the collected role relationship characteristics, construct the collaboration relationship graph of the surgical team , where: Vertex set represents each role in the medical staff team; Edge set represents the collaboration relationship between roles. Each edge contains an attribute set , and and represent the communication frequency, importance, and dependence characteristics respectively; The graph structure supports dynamically adjusting the attribute values of edges according to the surgical stage to adapt to the changes in the collaboration relationship in different stages.
[0023] Step 1.3, Division of the surgical process stages and dynamic adjustment of relationship weights; Divide the surgical process into multiple standard stages such as the preparation stage, anesthesia stage, operation stage, and suture stage, and define clear start and end marks for each stage.
[0024] Construct the surgical stage transition function , the function receives surgical activity data as input and outputs the current surgical stage .
[0025] According to the real-time recognized surgical stage , dynamically adjust the edge attributes in the collaboration relationship graph. The specific adjustment formula is: ; where represents the set of edge attributes between role i and role j at time t; represents role i in the medical team (such as the surgeon, anesthesiologist, etc.); represents role j in the medical team (such as the assistant physician, circulating nurse, etc.); represents the surgical stage at time t (such as the preparation period, anesthesia period, operation period, etc.); , , respectively represent the communication frequency characteristics, communication importance weights, and communication dependence characteristics of role i and role j in the surgical stage .
[0026] The output of this step is a collaboration relationship graph of the medical team that dynamically adjusts with the surgical stage. This graph records the communication frequency, importance, and dependence characteristics between different medical roles, providing a key basis for subsequent communication mode selection and priority determination.
[0027] The eye movement multi-object tracking and intention recognition module is used to monitor the eye activities of medical staff through a high-precision long-distance eye movement tracking camera and identify communication intentions and targets; Specifically, it includes the following steps: Step 2.1, eye movement acquisition and preprocessing; Install a high-precision long-distance eye movement tracking camera on the ceiling of the operating room to collect the eye movement data of medical staff. The eye movement data acquisition system includes: Infrared illumination unit: provides stable non-interfering infrared illumination to ensure high-quality eye images are obtained without affecting the normal vision of medical staff; High-frame-rate imaging unit: collects eye images at a rate of 200 frames per second to ensure capture of minute eye movements; Data preprocessing unit: filters noise, extracts pupil edges, and calculates fixation point coordinates for the collected eye images, and outputs standardized eye movement feature data, including fixation coordinates , fixation duration, pupil diameter change, etc.
[0028] The preprocessed eye movement data is represented as a sequence , where is the three-dimensional coordinate of the fixation point, is the pupil diameter, and is the fixation state feature.
[0029] Step 2.2, constructing an eye movement intention recognition model; Construct a multi-level eye movement intention recognition model to distinguish normal fixation and communication intention fixation: Feature extraction layer: Extract key features from the eye movement sequence data, including fixation stability ; pupil response ; saccade pattern features ; blink pattern ; Context fusion layer: Combine surgical context information , including the current surgical stage, the identities and positions of medical staff, etc.; Intention judgment layer: Based on the extracted features and context information, calculate the communication intention probability: ; Among them, represents the communication intention probability, with a value range of [0, 1]. The larger the value, the higher the possibility that the medical staff has a communication intention; is the communication intention probability calculation function, which is obtained through supervised learning methods; represents the fixation stability, which is calculated by the variance of the fixation point coordinates and reflects the stability of the medical staff's line of sight; represents the pupil response, that is, the pupil diameter change rate, which reflects the physiological response of the medical staff to visual stimuli; represents the saccade pattern features, including saccade speed and direction distribution information, which reflects the characteristics of the medical staff's line of sight transfer; represents the blink pattern, including blink frequency and duration, which reflects the attention state of the medical staff; represents the surgical context information, including environmental factors such as the current surgical stage, the identities and positions of medical staff, etc.
[0030] When the calculated exceeds the preset threshold, the system determines that the medical staff has a communication intention.
[0031] Step 2.3, determining the communication target and urgency; For the identified communication intention, further determine the communication target and urgency: Determining the communication target: By analyzing the fixation direction and the positions of medical staff in the operating room , determine the fixation target person; Evaluating the urgency: According to the fixation duration , calculate the communication urgency index by combining the pupil dilation rate and eye movement pattern with the current surgical stage: ; Among them, represents the communication urgency index, which is used to quantify the urgency of the communication needs of medical staff; represents the fixation stability, which is calculated by the variance of the fixation point coordinates. The smaller the value, the more stable the fixation, which may imply a stronger communication intention; represents the pupil response, that is, the change rate of pupil diameter, which reflects the physiological response of medical staff to visual stimuli. Usually, the pupil dilation is more obvious in case of emergency; represents the saccade pattern characteristics, including saccade speed and direction distribution information. In case of emergency, it may be manifested as faster and more frequent saccades; represents the blink pattern, including blink frequency and duration. Characteristic changes will occur in the blink pattern under tense or emergency states; represents the surgical context factor, which reflects the urgency of the current surgical stage. For example, high-risk stages such as critical suture and hemostasis have higher context urgency values; , , , , respectively represent the weight coefficients of fixation stability, pupil response, saccade pattern, blink pattern and surgical context factor, which are optimized and determined according to experimental data, and reflect the contribution degree of each index to the urgency evaluation.
[0032] Urgency grading: Map the calculated urgency index to four-level communication priorities: level one (routine information exchange), level two (confirmation of important information), level three (transmission of warning information), and level four (emergency intervention).
[0033] The output of this step includes the communication intention judgment result, the communication target personnel and the communication urgency level, and these information will be used for subsequent communication mode selection and microphone array beamforming control.
[0034] The multi-modal collaborative intention inference and communication mode matching module is used to combine the eye movement interaction features with the collaborative relationship network to infer the collaborative intention and select the most suitable communication mode; Specifically, it includes the following steps: Step 3.1, extraction of multi-target eye movement interaction features; Analyze the eye movement data of multiple medical staff in the operating room and extract the eye movement interaction features between them: Paired fixation feature: Calculate the behavioral features of medical staff A's fixation on medical staff B, including fixation frequency, fixation duration and fixation intensity, which is expressed as ; Mutual gaze behavior recognition: Detect the mutual gaze behavior among medical staff, calculate the duration and frequency of mutual gaze, expressed as ; Gaze sequence analysis: Analyze the gaze order pattern within a period of time, such as "A gazes at B and then B gazes at C", and extract possible information flow paths, expressed as 。
[0035] The eye movement interaction feature matrix is expressed as: , recording the eye movement interaction relationships among all medical staff in the operating room. Among them, represents the behavioral characteristics of medical staff A gazing at medical staff B, including parameters such as gaze frequency, duration, and intensity; represents the behavioral characteristics of medical staff B gazing at medical staff C; represents the behavioral characteristics of medical staff A gazing at medical staff C; Each element in the matrix is a vector, containing multi-dimensional features such as gaze frequency (number of gazes per minute), gaze duration (average duration in seconds for a single gaze), and gaze intensity (a value between 0 and 1 calculated comprehensively from pupil dilation and gaze stability); This matrix is updated dynamically, and the system will continuously update the values in the matrix according to real-time eye movement tracking data to reflect the changing interaction relationships among medical staff during the operation.
[0036] Step 3.2, Collaboration intention inference model; Combining the eye movement interaction feature matrix and the data of the medical team collaboration relationship network, construct a collaboration intention inference model, and calculate the probabilities of different collaboration scenarios: ; Among them, represents the probability of the collaboration scenario among medical staff A, B, and C. This value ranges from 0 to 1, and the larger the value, the more likely this collaboration scenario is to occur; represents the eye movement interaction characteristics of medical staff A gazing at medical staff B, including gaze frequency, gaze duration, and gaze intensity; represents the eye movement interaction characteristics of medical staff B gazing at medical staff C, including the same parameter dimensions as ; represents the eye movement interaction characteristics of medical staff A gazing at medical staff C, including the same parameter dimensions as ; represents the role relationship weight between medical staff A and B obtained from the collaboration relationship graph, reflecting the professional collaboration tightness between the two in the operation team, such as the collaboration frequency and importance of different role combinations like the surgeon and the assistant, the anesthesiologist and the nurse, etc.; Represents the role relationship weight between medical staff B and C obtained from the collaboration relationship graph, with the same definition as ; Represents the role relationship weight between medical staff A and C obtained from the collaboration relationship graph, with the same definition as ; Represents the current surgical stage, the impact factor of different surgical stages on the collaboration mode, and each stage has specific collaboration requirement characteristics. Represents the collaboration intention inference function, which calculates the probability value of a specific collaboration scenario by fusing all the above parameters.
[0037] The collaboration intention inference model realizes the accurate inference of complex collaboration intentions of multiple people by fusing the graph neural network and the temporal attention mechanism.
[0038] The model outputs include: The set of medical staff participating in the collaboration , where , , respectively represent the th, th, rd medical staff identified as currently needing to conduct communication collaboration; Represents the number of medical staff, and the system infers that these medical staff need to participate in the current communication collaboration based on eye movement interaction characteristics and collaboration relationships; Collaboration type , such as inquiry to answer, instruction to execution, information sharing, etc.; Collaboration urgency ; Step 3.3, communication mode selection and configuration; According to the collaboration intention inference result, select the most suitable one from the following four communication modes: Point-to-point direct connection mode: suitable for simple communication between two people, establishing a direct communication link between the sender and the receiver; Group broadcast mode: suitable for scenarios where one person transmits information to multiple people, establishing a one-to-many communication link; Sequential transmission mode: suitable for scenarios where information needs to be transmitted along a specific path, such as the information flow of "A→B→C"; Collaborative session mode: suitable for complex collaboration scenarios where multiple people need to communicate simultaneously, establishing a multi-point interconnected communication network.
[0039] The communication mode selection algorithm is based on the following rules: If the number of participants = 2 and the collaboration type is inquiry to answer or instruction to execution, select the point-to-point direct connection mode; If the number of participants > 2 and there is a clear information initiator, select the group broadcast mode; If the collaboration type is information flow and there is a sequential dependency relationship, select the sequential transfer mode; If the number of participants > 2 and the collaboration type is discussion and decision-making, select the collaborative session mode.
[0040] For each communication mode, the system will also configure corresponding communication parameters according to the collaboration urgency such as signal priority, bandwidth allocation, and noise reduction intensity.
[0041] Step 3.4, initialize the virtual meeting space; For the collaborative session mode, the system will initialize a virtual meeting space, and the specific steps include: Determine the set of session participants , where: represents the set composed of all participants in the virtual meeting space; , , respectively represent each medical staff participating in the session; represents the number of medical staff; Allocate virtual session identifiers ; Allocate a communication channel and priority for each participant: ; Among them, represents the communication channel configuration set allocated to participant ; is the communication channel identifier, which is used to uniquely identify the communication channel allocated to this medical staff to ensure the correct routing of information transmission; is the communication priority, which determines the information transmission order in the case of limited network resources. Medical staff with high priority (such as the surgeon) will be given priority in information transmission; is the allocated bandwidth, which is the network resources dynamically allocated according to the medical staff role and current communication needs, and affects the quality and clarity of voice transmission.
[0042] The output of this step includes the selected communication mode type, the set of communication participants, the communication parameter configuration, and (if applicable) the virtual meeting space initialization data, and this information will be used for subsequent microphone array beamforming and audio signal processing.
[0043] The adaptive microphone array beamforming module is used to dynamically adjust the beam direction and parameters of the microphone array according to the communication mode and participant location information; Specifically, it includes the following steps: Step 4.1, Multimodal Sensing Microphone Array System Structure; Construct a multimodal sensing microphone array system that combines vision, depth perception, and acoustic information. Its physical structure includes: Microphone Array Unit: Install multiple microphone sub-arrays on the ceiling and walls of the operating room. Each sub-array contains 8 omnidirectional microphones, forming a uniform circular array, and a cooperative capture network is formed between the sub-arrays; Vision Sensing Unit: Configure multiple high-definition cameras and depth sensors to track the positions and movements of medical staff and achieve the auxiliary function of personnel positioning; Signal Processing Unit: Includes a dedicated DSP chip for real-time signal processing, a microphone array controller, and a beamforming processor.
[0044] The entire system forms a unified audio capture and processing platform through network connection to achieve the fusion analysis of multimodal data.
[0045] Step 4.2, Multimodal Target Localization and Tracking; Precisely locate and track medical staff by combining vision and acoustic information: Vision Localization: Obtain the three-dimensional position coordinates of medical staff through a depth camera , achieving centimeter-level positioning accuracy; Sound Source Localization: Calculate the direction of the sound source based on the voice signals received by the microphone array. The specific method is: ; Among them, represents the time delay between microphones i and j, that is, the time difference between the sound wave arriving at the two microphones; , respectively represent the coordinate positions of microphones i and j in three-dimensional space; represents the propagation speed of sound waves in the air; is the unit vector of the sound source direction, represents the horizontal angle of the sound source, represents the vertical angle of the sound source, jointly determining the direction of the sound source in three-dimensional space; Information Fusion: Use the Kalman filtering algorithm to fuse the vision and acoustic localization results to obtain a more accurate position of the medical staff: ; Among them, is the target state vector at the current time , including the position and speed information of the medical staff; is the target state vector at the previous time ; is the state transition matrix, describing how the system state evolves from one moment to the next moment; is the process noise, representing the uncertainty in state prediction, and is usually assumed to be Gaussian white noise with zero mean; is the current time of the observation vector, that is, the position measurement of the medical staff obtained through visual and acoustic sensors; is the observation matrix, which maps the state space to the observation space; is the observation noise, representing the uncertainty in sensor measurement, and is usually also assumed to be Gaussian white noise with zero mean.
[0046] Step 4.3, adaptive beamforming algorithm; Based on the communication mode and participant position information, adaptive beamforming is implemented: Spatial filter design: Design a spatial filter for a specific direction, and control the beam directivity by controlling the weight coefficients of each unit of the microphone array: ; Among them, is the array output signal, representing the final audio signal obtained after beamforming processing; is the input signal of the m-th microphone, representing the original audio data collected by the m-th microphone at time t; is the complex conjugate weight coefficient, which is used to control the amplitude and phase of each microphone signal, and determines the directivity and gain characteristics of the beam; is the total number of microphones in the microphone array, representing the number of microphones participating in beamforming; represents the process of weighted summation of all microphone signals, which is the core operation of beamforming; Minimum Variance Distortionless Response (MVDR) beamforming: Design an optimal weight vector for a single communication target: ; Among them, is the weight vector of the minimum variance distortionless response beamforming, which is used to control the gain and phase of each unit of the microphone array; is the noise covariance matrix, which characterizes the statistical characteristics of the ambient noise and is used to suppress the interference signals from non-target directions; is the inverse matrix of the noise covariance matrix, which is used to calculate the optimal weight; is the steering vector of the target direction, which describes the phase relationship of the sound wave arriving at each unit of the microphone array from the target direction ; is the conjugate transpose of the steering vector, which is used to calculate the inner product; $\theta$ is the horizontal angle (azimuth angle) of the target sound source, representing the direction in the horizontal plane; $\varphi$ is the vertical angle (elevation angle) of the target sound source, representing the inclination angle relative to the horizontal plane; Multi - target beamforming: For multi - target scenarios such as collaborative sessions, a constrained optimization method is used to design the beamformer: ; where, $\mathbf{w}$ is the beamforming weight vector, representing the complex weight coefficients of each microphone; $\mathbf{w}^H$ is the conjugate transpose of the weight vector $\mathbf{R}$ is the input signal covariance matrix, describing the statistical characteristics of the signals received by the microphone array; $\mathbf{a}(\theta_d^k)$ is the steering vector of the $k$-th target direction, representing the phase relationship of the sound wave propagating from this direction to each microphone; $g_k$ is the gain constraint of the $k$-th target, controlling the response intensity of the beam in this direction; $K$ is the total number of targets, representing the number of sound sources that need to be simultaneously concerned; $\min_{\mathbf{w}}$ means to find the optimal weight vector $\text{s.t.}$ Adaptive interference suppression: While forming a gain in the main beam direction, form a null in the direction of the interference source: ; where, $\mathbf{w}_{LCMV}$ is the weight vector of the linearly constrained minimum variance beamforming, used to control the gain and phase of each unit of the microphone array; $\mathbf{R}$ is the covariance matrix of the input signal, describing the statistical characteristics of the signals received by the microphone array; $\mathbf{R}^{-1}$ is the inverse matrix of the covariance matrix, used to calculate the optimal weight; $\mathbf{A}$ is the steering matrix containing target and interference directions, where each column represents the steering vector of a specific direction; $\mathbf{A}^H$ is the conjugate transpose of the steering matrix, used for matrix operations; $\mathbf{I}$ is the matrix inversion operation, used to satisfy the constraint conditions; $\mathbf{g}$ is the gain constraint vector, specifying the desired response in each concerned direction. Usually, it is set to 1 for the target direction and 0 for the interference direction; and $\theta_d^k$ $\varphi_d^k$ represent the horizontal angle (azimuth angle) and vertical angle (elevation angle) of the $k$-th concerned direction respectively; is the total number of attention directions, including the target direction and the interference directions that need to be suppressed.
[0047] Step 4.4, beamforming strategy adapted to communication mode; According to the communication mode determined in Step 3, implement different beamforming strategies: Point-to-point direct connection mode: Form a two-way main beam from the speaker to the receiver, and at the same time form deep suppression in other directions; Group broadcast mode: Form a fan-shaped beam from the speaker to multiple receivers, covering all receiver positions; Sequential transmission mode: According to the information flow path, sequentially activate the beam links between corresponding nodes; Collaborative session mode: Construct a multi-point interconnected beam network, allowing direct communication between any two participants.
[0048] The beam configuration parameters are dynamically adjusted according to the communication urgency. The higher the urgency, the greater the beam gain and the deeper the lateral suppression, ensuring the clear transmission of critical information.
[0049] Step 4.5, dynamic tracking and adjustment; For the movement of medical staff, achieve dynamic tracking and adjustment of the beam: Position prediction: Based on the Kalman filter-based target tracking algorithm, predict the next moment position of the medical staff: ; where is the predicted position value at moment, based on the information at is the estimated position value at is the state transition matrix, describing the target motion mode, used to map the current state to the predicted state at the next moment; this formula represents the prediction step of the Kalman filter, predicting the next position of the medical staff through the current state and the motion model; Beam real-time update: According to the predicted position, adjust the beam direction in advance to reduce the tracking delay: ; where is the beamforming weight vector at the next moment to control the gain and phase of each unit of the microphone array; is the beam weight calculation function, mapping the predicted position to the optimal beamforming weight; is the predicted position of the medical staff based on the information at the current moment in The position coordinates at a moment; this formula indicates that the system calculates and adjusts the beamforming weights in advance according to the predicted positions of medical staff to achieve predictive tracking of the beam.
[0050] Fast response mechanism: For sudden position changes, a fast beam switching algorithm is adopted to reduce the beam adjustment time and ensure stable voice pickup quality when medical staff move quickly.
[0051] The output of this step is a beam configuration dynamically adjusted according to the communication mode and participant positions, including the weight coefficients of each microphone, the main beam direction, and the gain parameters. These configurations directly control the signal processing of the microphone array to achieve precise voice capture and noise suppression.
[0052] Virtual meeting space construction and dynamic optimization module, used to create a virtual meeting space for multi-person collaboration scenarios, provide a unified logical framework for multi-person collaborative communication, and adaptively adjust communication parameters according to changes in the surgical stage; Specifically, it includes the following steps: Step 5.1, virtual meeting space construction; Based on the collaborative session requirements, create a virtual meeting space (VirtualMeetingSpace, VMS) model to provide a unified logical framework for multi-person collaborative communication: Spatial topology construction: Establish a communication topology structure according to the participant set and collaboration relationship to determine the signal transmission path; Spatial node configuration: Assign a set of communication parameters to each participant node : ; Among them, is the signal gain parameter, which controls the amplification multiple of the node audio signal. The larger the value, the louder the volume; is the signal priority, which determines the processing order in resource competition. High-priority signals will obtain the priority transmission right; is the noise reduction intensity, which controls the intensity level of the noise suppression algorithm applied to this node. The larger the value, the stronger the noise reduction effect; is the allocated bandwidth, which represents the communication channel capacity allocated to this node and affects the quality and delay of audio transmission; Session management mechanism: Establish session control policies, including: Access control: Determine which personnel can join or leave the current session; Communication permission control: Based on the medical staff role and urgency, allocate different speaking priorities; Session lifecycle management: Dynamically adjust session resources according to the collaboration duration and status changes.
[0053] Step 5.2, Audio Channel Optimization and Noise Reduction Processing; Optimize each communication channel according to the positions of medical staff and the distribution of environmental noise: Noise Characteristic Analysis: Analyze the noise characteristics using the environmental audio captured by the microphone array, and construct a noise model matrix , where is the noise model matrix, which contains the characteristic information of various types of noise in the environment; , , respectively represent the spectral characteristics of the , , th type of noise; is the total number of noise types recognized by the system; is the frequency variable, representing the frequency points in the spectral analysis; this noise model matrix provides the basic data for targeted processing in subsequent noise reduction algorithms, enabling the system to selectively suppress different types of noise; Adaptive Noise Suppression Processing: Dynamically adjust the parameters of the noise suppression algorithm according to the noise characteristics and communication urgency: For low-priority communications: Adopt a strong noise suppression strategy to ensure speech clarity; For high-priority communications: Adopt a conservative noise suppression strategy to prioritize the real-time and integrity of information; Spatial Selective Noise Reduction: Apply directional noise reduction processing to each communication link in the virtual meeting: ; where is the processed speech signal, representing the output speech after noise reduction processing; is the adaptive filter, which is a function that changes over time and depends on the following parameters: is the time variable, indicating that the filter will be dynamically adjusted over time; is the urgency parameter, which determines the filtering intensity. The higher the urgency, the more conservative the filtering to retain more original information; is the noise model matrix, providing environmental noise characteristic information to guide the filtering process; is the original speech signal, that is, the unprocessed voice input of medical staff; represents the convolution operation, representing the process of the filter processing the original signal; Spectrum Enhancement Processing: Implement selective spectrum enhancement for medical terms and key instructions to improve the recognizability of important information: ; where is the spectrum of the output signal after spectrum enhancement processing, indicating at frequency The signal energy distribution at is the spectrum of the original input signal, representing the signal energy distribution at frequency ; is the spectrum gain function, which dynamically adjusts the gain according to frequency and the set of medical term keywords ; is the set of medical term keywords, including medical professional vocabulary and key instructions that need to be particularly enhanced; is the frequency variable, representing different frequency components in the sound signal; this formula realizes selective enhancement of the frequency intervals containing medical terms and key instructions, improving the recognizability of these important information in a noisy environment.
[0054] Step 5.3, dynamic session optimization and prediction; Based on the surgical progress and historical collaboration patterns, dynamic prediction and pre-adjustment of communication requirements are realized: Collaboration pattern learning: By analyzing historical communication data, a team collaboration pattern model is constructed: ; Among them, is the collaboration pattern model, representing the communication and collaboration pattern of the medical staff team during the operation, including information such as the communication frequency, priority relationship, and typical interaction pattern among team members; is the machine learning function, which is an algorithm for extracting and learning collaboration patterns from historical data, and can be a method based on deep learning or statistical models; is the historical communication record, which contains the communication data between medical staff during past operations, such as communication time, initiator, recipient, content type, urgency, etc.; is the surgical type and stage information, which describes the professional classification of the operation (such as cardiac surgery, neurosurgery, etc.) and different stages of the operation (such as anesthesia period, operation period, suture period, etc.); Communication requirement prediction: Based on the collaboration pattern model and the current surgical state, predict future possible communication requirements: ; Among them, is the predicted future communication requirement at time , representing the communication activities that the system predicts may occur at a certain future time point; is the prediction function, which is used to infer future communication requirements according to the collaboration pattern model and the current state; is the collaboration pattern model, which contains the data model of the team's historical collaboration pattern and communication rules; is the current surgical state, including real-time information such as the current surgical stage, the positions of medical staff, and the equipment status; is the current time point, serving as a time reference benchmark; is the predicted time window, representing the time length predicted from the current moment into the future; Predictive resource allocation: According to the prediction results, adjust the system resource allocation in advance: Pre-adjust the beam direction to pre-form an enhanced beam for the positions of medical staff where communication may occur soon; Pre-allocate processing resources to ensure that the system responds in a timely manner when the peak communication demand arrives; Dynamically adjust the noise reduction parameters to reduce the noise suppression intensity during the expected critical information transfer stage.
[0055] Step 5.4, Communication strategy adaptation for surgical stage perception; According to the characteristics of different surgical stages, adaptively adjust the communication strategy: Surgical stage identification: Based on the equipment status in the operating room, the activity characteristics of medical staff, and the changes in collaboration relationships, identify the current surgical stage: ; Among them, is the identified surgical stage; is the surgical stage identification function, an algorithm used to comprehensively analyze various input features and map them to specific surgical stages; is the set of equipment status parameters; is the set of medical staff activity characteristics, describing the behavioral characteristic data such as the position distribution, action patterns, and operation frequencies of the members of the surgical team; is the collaboration relationship change parameter, representing the dynamic characteristics of the collaboration network such as the interaction patterns, communication frequencies, and instruction flows among the members of the medical team; Stage adaptation strategy configuration: Configure corresponding communication strategies for different surgical stages: Preparation stage: Adopt loose access control and low priority differences; Critical operation stage: Adopt strict priority management to ensure that the instructions of the surgeon are preferentially transmitted; Emergency state: Activate the emergency communication mode to minimize latency and improve the reliability of information transmission.
[0056] The output of this step is a dynamically optimized virtual meeting space model, which provides a unified logical framework and resource allocation mechanism for multi-person collaborative communication, and can adaptively adjust communication parameters according to the surgical progress and collaboration requirements to ensure efficient intercom communication in a complex surgical environment.
[0057] Application example of this embodiment: This embodiment has been successfully applied to the robot-assisted cardiac valve repair surgery in the cardiothoracic surgery department of a certain tertiary hospital. The following is an example description of this application scenario.
[0058] The cardiac valve repair surgery is performed by a surgical team consisting of a surgeon, two assistant physicians, an anesthesiologist, an instrument nurse, and a circulating nurse, a total of 6 people. The cardiopulmonary bypass machine, monitor, robot system, and multiple auxiliary devices are operating simultaneously in the operating room. The average environmental noise is 75 dBSPL, and the peak can reach 88 dBSPL. The surgical process is divided into a preparation period (30 minutes), an anesthesia period (25 minutes), a thoracotomy period (40 minutes), a robot operation period (120 minutes), a suture period (35 minutes), and a finishing period (20 minutes). Team members need to maintain precise coordination throughout the surgical process, especially during the robot operation period, where the requirements for communication quality and timeliness are extremely high.
[0059] Implementation process example: Example of constructing a collaborative relationship network for the medical team: The system first collected the communication data of 32 similar surgeries completed by this cardiothoracic team in the past 6 months and identified the following key role relationship characteristics: Communication frequency characteristics between the surgeon and the anesthesiologist = 0.43 times / minute; Communication importance weight = 0.92; Communication dependence characteristics = 0.89 (forced dependence type).
[0060] Communication frequency characteristics between the surgeon and the first assistant = 0.68 times / minute; Communication importance weight = 0.87; Communication dependence characteristics = 0.78 (forced dependence type).
[0061] The system constructed a complete collaborative relationship graph where the vertex set contains 6 medical staff roles, and the edge set contains 15 pairs of role relationships, and each pair of relationships contains 3 attribute characteristics (frequency, importance, dependence). The graph automatically updates the attribute values according to the identified surgical stages. For example, during the robot operation period, the system automatically increases the communication importance weight between the surgeon and the anesthesiologist from 0.70 in the preparation period to 0.92.
[0062] Example of eye movement multi-object tracking and intention recognition: The system installed 4 high-precision long-distance eye movement tracking cameras in the operating room, covering all the activity areas of medical staff. During a specific communication event, the system captured that the surgeon continuously stared at the anesthesiologist for 2.3 seconds. At this time: Gaze stability =0.92 (coordinate variance is less than 2.1 mm); Pupil response =0.23 (pupil dilation rate is 23%); Saccade pattern characteristics =0.19 (low-frequency short-distance saccade pattern); Blinking pattern =0.12 (decreased blinking frequency).
[0063] Combined with the current critical stage of the robotic operation period, the system calculated the probability of communication intention: ; Since this value exceeded the preset threshold of 0.75, the system determined that the surgeon had a clear communication intention. Further, the system analyzed the gaze direction and the positions of the operating room personnel, identified the anesthesiologist as the communication target, and calculated the urgency index based on the eye movement characteristics , corresponding to the third-level priority (warning message transmission).
[0064] Example of multimodal collaborative intention reasoning and communication mode matching: In another communication event, the system detected a complex three-person collaboration scenario: The surgeon stared at the instrument nurse ( =0.76); The instrument nurse stared at the first assistant ( =0.65); At the same time, there was a brief mutual gaze between the first assistant and the surgeon ( =0.58).
[0065] Combined with the role relationship weights in the collaboration relationship graph ( =0.81, =0.64, =0.87) and the current robotic operation period stage information, the system inferred: .
[0066] The system recognized this as a collaboration scenario of the "tool transfer and collaborative operation" type, with an urgency of 0.75 (second-level priority). Based on this reasoning result, the system selected the "sequential transfer" communication mode, established a communication path from the surgeon → instrument nurse → first assistant, and ensured the accurate transmission of complex instructions.
[0067] Adaptive Microphone Array Beamforming Example: The system deployed 4 sub - microphone arrays in the operating room, each consisting of 8 omnidirectional microphones to form a uniform circular array, which were respectively installed at the 4 corner positions on the top of the operating room.
[0068] In the above - mentioned communication event between the surgeon and the anesthesiologist, the system obtained the position coordinates of the surgeon (2.4m, 1.8m, 1.7m) and the anesthesiologist (3.9m, 4.2m, 1.7m) according to the multi - modal target localization result. The system automatically calculated the optimal MVDR beamforming weight vector: ; This weight vector enabled Sub - array 1 to form an accurate beam pointing to the surgeon, while forming deep suppression in the directions of the ventilator and the cardiopulmonary machine. Similarly, the system configured the weight vector of Sub - array 3 to capture the anesthesiologist's voice. This adaptive beam configuration achieved high - quality voice capture in an ambient noise of 75 dBSPL.
[0069] Virtual Conference Space Construction and Dynamic Optimization Example: In a four - person collaboration scenario involving a surgeon, two assistants, and an anesthesiologist, the system automatically established a virtual conference space and assigned communication parameters to the four participants: Surgeon: ; Anesthesiologist: ; First Assistant: ; Second Assistant: ; Among them, gain represents the audio gain parameter; priority represents the communication priority parameter; noise_reduction represents the noise suppression intensity parameter; bandwidth represents the bandwidth allocation parameter. The system also based on the analysis result of the noise characteristics; For the frequency band of 63 - 250 Hz (mainly the noise of the cardiopulmonary machine): suppression intensity 0.85; For the frequency band of 500 - 1000 Hz (mainly the alarm sound of the monitor): suppression intensity 0.65; For the frequency band of 2000 - 4000 Hz (the key frequency band of speech): suppression intensity 0.40.
[0070] During the critical stage of the operation (valve repair operation period), the system predicted that there might be important communication needs between the surgeon and the anesthesiologist, adjusted the beam direction in advance and pre-allocated processing resources. When the actual communication occurred, the system response time was only 12 ms, far lower than the 50 - 100 ms of the traditional system.
[0071] Verification of technical effects: For the core technical effects of this embodiment, quantitative verification tests were carried out, and the results are as follows: Effect of improving communication quality and efficiency: The test compared the changes in key indicators before and after using this system in 32 similar operations: Comparison of the speech clarity index (Short-Time Objective Intelligibility, STOI):
[0072] Comparison of speaker localization accuracy and response latency:
[0073] Effect of improving team collaboration and cognitive load: By combining questionnaires and objective indicators, the changes in team collaboration efficiency and cognitive load were evaluated: Team collaboration efficiency index:
[0074] Results of cognitive load assessment:
[0075] The above verification results show that this embodiment significantly improves communication quality and efficiency in the actual medical scenario, optimizes the team collaboration experience, reduces the cognitive load of medical staff, and effectively solves the technical challenges faced by doctor-nurse intercom communication in the robot surgery environment.
[0076] The embodiments of the present invention have been described above, but these embodiments are not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.
Claims
1. A noise reduction intercom system for doctors and nurses in a robotic surgery environment, characterized in that It includes: A medical team collaboration relationship network construction module, which is used to construct a professional collaboration relationship map for the surgical environment through graph structure representation and role relationship analysis; An eye movement multi-target tracking and intention recognition module, which is used to monitor the eye activities of medical staff through a high-precision long-distance eye movement tracking camera to identify communication intentions and targets; A multi-modal collaboration intention reasoning and communication mode matching module, which is used to infer collaboration intentions and select the most suitable communication mode by combining eye movement interaction features and the collaboration relationship network; An adaptive microphone array beamforming module, which is used to dynamically adjust the beam direction and parameters of the microphone array according to the communication mode and participant position information; A virtual meeting space construction and dynamic optimization module, which is used to create a virtual meeting space for multi-person collaboration scenarios, provide a unified logical framework for multi-person collaboration communication, and adaptively adjust communication parameters according to changes in the surgical stage.
2. The doctor-nurse noise reduction intercom system for a robotic surgery environment according to claim 1, wherein The medical team collaboration relationship network construction module includes: Collect medical staff role and relationship features, including communication frequency features, communication importance features, and communication dependence features; Construct a collaboration relationship map: ; Among them, represents a collaborative relationship graph, and the vertex set represents each role in the medical team, and the edge set represents the collaborative relationship between roles; Dynamically adjust the edge attributes in the collaboration relationship map according to the real-time identified surgical stage.
3. The doctor-nurse noise reduction intercom system for a robotic surgery environment according to claim 1, wherein The eye movement multi-target tracking and intention recognition module includes: Use an infrared illumination unit and a high-frame-rate imaging unit to collect eye movement data of medical staff; Construct a multi-level eye movement intention recognition model, including a feature extraction layer, a context fusion layer, and an intention judgment layer; Determine communication targets by analyzing the gaze direction and the positions of medical staff in the operating room, and evaluate the communication urgency according to the gaze duration, pupil dilation rate, and eye movement pattern.
4. The doctor-nurse noise reduction intercom system for a robotic surgery environment according to claim 1, wherein, The multi-modal collaboration intention reasoning and communication mode matching module includes: Extract multi-target eye movement interaction features, including paired gaze features, mutual gaze behavior recognition, and gaze sequence analysis; Construct a collaboration intention reasoning model to calculate the probabilities of different collaboration scenarios; Select the most suitable communication mode from the point-to-point direct connection mode, the group broadcast mode, the sequential transfer mode, and the collaborative session mode.
5. The doctor-nurse noise reduction intercom system for the robot surgery environment according to claim 1, wherein The virtual meeting space construction and dynamic optimization module includes: Create a virtual meeting space model, including space topology construction, space node configuration, and session management mechanism; Optimize each communication channel according to the positions of medical staff and the environmental noise distribution; Construct a team collaboration mode model through the analysis of historical communication data and predict future possible communication needs; Adaptively adjust the communication strategy according to the characteristics of different surgical stages.
6. The doctor-nurse noise reduction intercom system for a robotic surgery environment according to claim 4, wherein The collaboration intention reasoning model is based on a graph neural network and a temporal attention mechanism. The inputs include an eye movement interaction feature matrix, role relationship weights, and current surgical stage information, and the outputs include the set of medical staff participating in the collaboration, the collaboration type, and the collaboration urgency.
7. The doctor-nurse noise reduction intercom system for the robot-assisted surgical environment according to claim 4, characterized in that, The communication mode matching is based on the following rules: If the number of participants is 2 and the collaboration type is inquiry to answer or instruction to execution, select the point-to-point direct connection mode; If the number of participants is greater than 2 and there is a clear information initiator, select the group broadcast mode; If the collaboration type is information flow and there is a sequential dependency relationship, select the sequential transfer mode; If the number of participants is greater than 2 and the collaboration type is discussion and decision-making, select the collaborative session mode.
8. The doctor-nurse noise reduction intercom system for a robotic surgery environment according to claim 1, wherein The adaptive microphone array beamforming module includes: Deploying multiple sub-microphone arrays, each sub-array containing multiple omnidirectional microphones to form a collaborative capture network; Combining visual and acoustic information for precise positioning and tracking of medical staff; Implementing beamforming algorithms including spatial filter design, minimum variance distortionless response beamforming, multi-target beamforming, and adaptive interference suppression based on communication modes and participant location information; Using Kalman filtering for position prediction according to the movement of medical staff to achieve dynamic tracking and adjustment of the beam.
9. The doctor-nurse noise reduction intercom system for the robot surgery environment according to claim 1, wherein It also includes: A multi-modal perception unit, including multiple microphone array elements installed on the ceiling and walls of the operating room, a high-precision long-distance eye movement tracking camera, a high-definition camera, and a depth sensor; A signal processing unit, including a dedicated DSP chip for real-time signal processing, a microphone array controller, and a beamforming processor; A network transmission unit for transmitting the data collected by each unit to the central processing system to achieve fusion analysis of multi-modal data.
10. A computer-readable storage medium, characterized in that, It is used to store computer-readable instructions that, when read by a computer, can run the doctor-nurse noise reduction intercom system for a robotic surgical environment as described in any one of claims 1-9.
Citation Information
Patent Citations
High-simulation multimedia ophthalmological doctor-patient communication system
CN105117589A
Fault-tolerant method for improving underwater robot networking robustness
CN118741573A