Noise reduction intercom system for doctors and nurses in robotic surgery environments

By building a communication system that integrates collaborative relationship network of medical teams, eye tracking and multimodal data, the noise interference and sterile operation restrictions in robotic surgical environments are solved, and efficient and stable intercom communication between doctors and nurses is achieved to adapt to the real-time needs of complex surgical environments.

CN120264167BActive Publication Date: 2025-08-26SHENZHEN POROS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510744338.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-08-26
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

In the robotic surgical environment, the problems of noise interference in the operating room, complex team collaboration, sterile operation restrictions and insufficient adaptability of traditional communication systems have led to unstable intercom communication quality and high misoperation rate, which cannot meet the real-time communication needs of complex surgical environments.

Method used

The medical team collaboration relationship network construction module, eye movement multi-objective tracking and intent identification module, multi-modal cooperative intention reasoning and communication pattern matching module and adaptive microphone array beamforming module are adopted. Combined with graph structure representation, eye movement tracking, multi-modal data fusion and adaptive beamforming technology, communication modes and parameters are dynamically adjusted to achieve contactless control and precise positioning.

Benefits of technology

Improve communication quality and efficiency, achieve contactless control, reduce surgical risks, optimize team collaboration, reduce cognitive load, be adaptable and scalable, adaptable to different surgical environments and team composition, and support multiple communication modes and dynamic switching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264167B_ABST
    Figure CN120264167B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of medical communication technology and discloses a noise reduction intercom system for doctors and nurses in a robotic surgery environment, comprising: a medical team collaboration relationship network construction module, which is used to construct a professional collaboration relationship map for the surgical environment through graph structure representation and role relationship analysis; an eye movement multi-target tracking and intention recognition module; a multimodal collaboration intention reasoning and communication mode matching module, which is used to combine eye movement interaction characteristics with the collaboration relationship network to infer collaboration intentions and select the most suitable communication mode; an adaptive microphone array beamforming module and a virtual conference space construction and dynamic optimization module; the present invention realizes contactless control, intelligent collaborative routing and adaptive beamforming by integrating eye movement tracking technology, social relationship network analysis and multimodal perception technology, effectively solving the technical problems of noise interference in the operating room, complex team collaboration and sterile operation restrictions, and improving the communication quality and efficiency during the operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical communication technology, and more particularly, to a noise reduction intercom system for doctors and nurses in a robotic surgery environment. Background Art

[0002] In robotic surgery, doctors and nurses need to maintain efficient and clear intercom communication during the operation to ensure surgical safety. However, existing technologies have the following problems:

[0003] There are multiple noise sources in the operating room, such as ventilators, monitors, and robot movements, and medical staff need to move during the operation, resulting in unstable intercom communication quality; large and complex operations involve the collaboration of medical staff from multiple professional fields, forming a complex network of professional division of labor and collaborative relationships. Traditional intercom systems lack the ability to understand and support this complex team collaboration model; medical staff need to keep their hands sterile during the operation and cannot easily operate traditional contact communication equipment; existing systems mainly rely on voice wake-up or gesture control, with a high error rate and easy to destroy the sterility of the operating environment; the communication needs between team members change dynamically with the stage of the operation, and the traditional fixed-mode communication system lacks adaptability and cannot flexibly adjust the communication strategy according to real-time scenario requirements; in complex surgical environments, the communication matching accuracy of traditional systems is low, and they lack the ability to perceive changes in the surgical stage.

[0004] Therefore, a technology is needed to solve the communication problems in the complex environment of the operating room through innovative technical means. Summary of the Invention

[0005] The present invention provides a noise reduction intercom system for doctors and nurses in a robotic surgery environment, which solves the technical problems of noise interference in the operating room, complex team collaboration, and sterile operation restrictions in related technologies.

[0006] The present invention provides a noise reduction intercom system for doctors and nurses in a robotic surgery environment, comprising:

[0007] A medical team collaboration network construction module is used to construct a professional collaboration relationship map for the surgical environment through graph structure representation and role relationship analysis;

[0008] Eye movement multi-target tracking and intention recognition module, which is used to monitor medical staff's eye movements through a high-precision long-distance eye tracking camera and identify communication intentions and targets;

[0009] Multimodal collaboration intention inference and communication mode matching module, which combines eye movement interaction features with the collaboration relationship network to infer collaboration intentions and select the most suitable communication mode;

[0010] Adaptive microphone array beamforming module, which is used to dynamically adjust the beam direction and parameters of the microphone array according to the communication mode and participant location information;

[0011] The virtual meeting space construction and dynamic optimization module is used to create a virtual meeting space for multi-person collaboration scenarios, provide a unified logical framework for multi-person collaborative communication, and adaptively adjust communication parameters according to changes in the surgical stage.

[0012] In a preferred embodiment, the medical team collaboration relationship network building module includes:

[0013] Collecting medical and nursing role and relationship characteristics, including communication frequency, communication importance, and communication dependency;

[0014] Build a collaborative relationship map:

[0015] ;

[0016] in, Represents the collaborative relationship graph, vertex set Represents the various roles in the medical team, edge set Indicates the collaborative relationship between roles;

[0017] According to the real-time identified surgical stages, the edge attributes in the collaborative relationship graph are dynamically adjusted.

[0018] In a preferred embodiment, the eye movement multi-target tracking and intention recognition module includes:

[0019] An infrared lighting unit and a high-frame-rate imaging unit are used to collect eye movement data of medical staff;

[0020] Construct a multi-layered eye movement intention recognition model, including feature extraction layer, context fusion layer, and intention judgment layer;

[0021] The communication target was determined by analyzing gaze direction and the position of medical staff in the operating room, and the urgency of communication was assessed based on gaze duration, pupil dilation rate, and eye movement patterns.

[0022] In a preferred embodiment, the multimodal collaboration intention reasoning and communication mode matching module includes:

[0023] Extract multi-target eye movement interaction features, including paired gaze features, mutual gaze behavior recognition, and gaze sequence analysis;

[0024] Build a collaborative intention reasoning model to calculate the probability of different collaborative scenarios;

[0025] Select the most suitable communication mode from point-to-point direct connection mode, packet broadcast mode, sequential delivery mode and collaborative session mode.

[0026] In a preferred embodiment, the virtual conference space construction and dynamic optimization module includes:

[0027] Create a virtual conference space model, including space topology construction, space node configuration, and session management mechanism;

[0028] Optimize each communication channel based on the location of medical staff and the distribution of ambient noise;

[0029] By analyzing historical communication data, we can build a team collaboration model and predict possible future communication needs;

[0030] Adaptively adjust the communication strategy according to the characteristics of different surgical stages.

[0031] In a preferred embodiment, the collaborative intention reasoning model is based on a graph neural network and a temporal attention mechanism. The input includes an eye movement interaction feature matrix, role relationship weights, and current surgical stage information. The output includes the set of medical staff participating in the collaboration, the collaboration type, and the collaboration urgency.

[0032] In a preferred embodiment, the communication pattern matching is based on the following rules:

[0033] If the number of participants is 2 and the collaboration type is inquiry to answer or command to execution, select point-to-point direct connection mode;

[0034] If the number of participants is greater than 2 and there is a clear message initiator, select the group broadcast mode;

[0035] If the collaboration type is information flow and there is a sequence dependency, select the sequential delivery mode;

[0036] If the number of participants is greater than 2 and the collaboration type is discussion and decision-making, select the collaborative conversation mode.

[0037] In a preferred embodiment, the adaptive microphone array beamforming module includes:

[0038] Deploy multiple sub-microphone arrays, each containing multiple omnidirectional microphones, to form a collaborative capture network;

[0039] Combine visual and acoustic information to accurately locate and track medical staff;

[0040] Implement beamforming algorithms including spatial filter design, minimum variance distortion-free response beamforming, multi-target beamforming, and adaptive interference suppression based on communication patterns and participant location information;

[0041] According to the movement of medical staff, Kalman filtering is used to predict their position and realize dynamic tracking and adjustment of the beam.

[0042] In a preferred embodiment, it also includes:

[0043] A multimodal perception unit, including multiple microphone arrays mounted on the operating room ceiling and walls, a high-precision long-range eye-tracking camera, a high-definition camera, and a depth sensor;

[0044] Signal processing unit, including a dedicated DSP chip for real-time signal processing, a microphone array controller, and a beamforming processor;

[0045] The network transmission unit is used to transmit the data collected by each unit to the central processing system to realize the fusion analysis of multimodal data.

[0046] In a preferred embodiment, a computer-readable storage medium is used to store computer-readable instructions, which, when read by a computer, can run a noise reduction intercom system for doctors and nurses in a robotic surgery environment.

[0047] The beneficial effects of the present invention are:

[0048] Improved communication quality and efficiency: The speaker positioning accuracy in mobile state is improved, which reduces the system response delay and increases the effective sound capture gain, thus maintaining good communication quality in noisy and complex surgical environments.

[0049] Achieve contactless control and reduce surgical risks: Eye tracking enables contactless communication control, maintains a sterile surgical environment, improves the accuracy of intent recognition, and can simultaneously track the gaze behavior of multiple medical staff to support complex team collaboration.

[0050] Optimize team collaboration and reduce cognitive load: Reduce the delay in key information transmission, improve information accuracy, and reduce the cognitive load of team communication. Medical staff can focus more on the operation itself, improve the effective signal-to-noise ratio in the surgical environment, and reduce auditory fatigue of medical staff.

[0051] Adaptability and scalability: The system can automatically adapt to different surgical environments and team compositions, with continuously improved communication matching accuracy. It supports multiple communication modes and dynamic switching mechanisms to meet the needs of surgeries of varying complexity. It adopts a modular design and can be customized based on the specific hospital environment and surgical type. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a module diagram of the noise reduction intercom system for doctors and nurses in a robotic surgery environment according to the present invention. DETAILED DESCRIPTION

[0053] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0054] At least one embodiment of the present invention discloses a noise reduction intercom system for doctors and nurses in a robotic surgery environment. Figure 1 Shown, including:

[0055] A medical team collaboration network construction module is used to construct a professional collaboration relationship map for the surgical environment through graph structure representation and role relationship analysis;

[0056] The specific steps include:

[0057] Step 1.1, collection of medical and nursing role and relationship characteristics;

[0058] Collect basic information and professional role definitions for various medical personnel in the surgical environment (including the surgeon, assistant physician, anesthesiologist, circulating nurse, etc.), and determine the responsibilities and communication needs of each role at different stages of the operation. For each pair of medical and nursing roles, calculate the following relationship characteristics:

[0059] Communication frequency characteristics: Record the typical communication frequency between specific role pairs at different stages of surgery, expressed as ,in, 、 Represent role i and role j respectively, represents the surgical stage k;

[0060] Communication importance feature: According to the criticality and safety dependency of the surgical operation, an importance weight is assigned to each pair of role relationships, which is expressed as ,in, 、 Represent role i and role j respectively, represents the surgical stage k;

[0061] Communication dependency characteristics: Analyze which role relationships are mandatory (e.g., anesthesiologists reporting vital signs to surgeons) and which are auxiliary (e.g., information sharing between circulating nurses), expressed as ,in, 、 Represent role i and role j respectively, represents the surgical stage k.

[0062] Step 1.2, collaborative relationship map construction;

[0063] Construct a collaborative relationship map of the surgical team based on the collected role relationship characteristics ,in:

[0064] Vertex Set Represents the various roles in the healthcare team;

[0065] Edge Set Represents the collaborative relationship between roles, each edge contains a set of attributes , 、 、 represent the communication frequency, importance, and dependency characteristics respectively;

[0066] The atlas structure supports the classification of surgical stages Dynamically adjust the attribute values ​​of edges to adapt to changes in collaborative relationships at different stages.

[0067] Step 1.3: Division of surgical process stages and dynamic adjustment of relationship weights;

[0068] The surgical process is divided into multiple standard stages, including preparation, anesthesia, operation, and suturing, and clear start and end signs are defined for each stage.

[0069] Constructing surgical phase transition functions , this function receives surgical activity data As input, output the current stage of surgery .

[0070] Based on real-time identification of surgical stages , dynamically adjust the edge attributes in the collaboration relationship graph. The specific adjustment formula is:

[0071] ;

[0072] in, represents the edge attribute set between role i and role j at time t; represents the role i in the medical team (such as surgeon, anesthesiologist, etc.); represents the role j in the medical team (such as assistant physician, circulating nurse, etc.); represents the surgical stage at time t (such as preparation period, anesthesia period, operation period, etc.); 、 、 Respectively represent the roles i and j in the operation stage The communication frequency characteristics, communication importance weight and communication dependency characteristics of the communication network are analyzed.

[0073] The output of this step is a collaborative relationship map of the medical team that is dynamically adjusted according to the surgical stage. This map records the communication frequency, importance, and dependency characteristics between different medical roles, providing a key basis for subsequent communication mode selection and priority determination.

[0074] Eye movement multi-target tracking and intention recognition module, which is used to monitor medical staff's eye movements through a high-precision long-distance eye tracking camera and identify communication intentions and targets;

[0075] The specific steps include:

[0076] Step 2.1, eye movement collection and preprocessing;

[0077] A high-precision, long-distance eye-tracking camera is installed on the operating room ceiling to collect eye movement data from medical staff. The eye movement data collection system includes:

[0078] Infrared lighting unit: provides stable, non-interfering infrared lighting to ensure high-quality eye images are obtained without affecting the normal vision of medical staff;

[0079] High frame rate imaging unit: captures eye images at a rate of 200 frames per second to ensure that even tiny eye movements are captured;

[0080] Data preprocessing unit: performs noise filtering, pupil edge extraction and gaze point coordinate calculation on the collected eyeball images, and outputs standardized eye movement feature data, including gaze coordinates , gaze duration, pupil diameter changes, etc.

[0081] The preprocessed eye movement data is represented as a sequence ,in, is the three-dimensional coordinate of the gaze point, is the pupil diameter, is the gaze state feature.

[0082] Step 2.2, eye movement intention recognition model construction;

[0083] Build a multi-level eye movement intention recognition model to distinguish between normal gaze and communication intent gaze:

[0084] Feature extraction layer: extracts key features from eye movement sequence data, including gaze stability ;Pupillary response ; Eye saccade pattern characteristics ; Blink mode ;

[0085] Context fusion layer: combining surgical context information , including the current stage of surgery, identity and location of medical staff;

[0086] Intent judgment layer: Calculates the probability of communication intention based on the extracted features and context information:

[0087] ;

[0088] in, Indicates the probability of communication intention, with a value range of [0, 1]. A larger value indicates a higher probability that the medical staff has the intention to communicate; The communication intention probability calculation function is obtained through supervised learning training; Indicates gaze stability, which is calculated by the variance of gaze point coordinates to reflect the stability of medical staff's sight; represents pupil response, i.e. the rate of change of pupil diameter, which reflects the physiological response of medical staff to visual stimulation; Represents the eye movement pattern characteristics, including eye movement speed and direction distribution information, reflecting the characteristics of medical staff's gaze shift; Indicates blinking patterns, including blink frequency and duration, reflecting the attention status of medical staff; Represents surgical context information, including environmental factors such as the current surgical stage, medical staff identity and location.

[0089] When the calculated When the preset threshold is exceeded, the system determines that the medical staff has the intention to communicate.

[0090] Step 2.3, communication objectives and urgency are determined;

[0091] For the identified communication intention, further determine the communication target and urgency:

[0092] Communication Targeting: By Analyzing Gaze Direction and the location of medical staff in the operating room , determine the target person;

[0093] Urgency assessment: based on gaze duration , pupil dilation rate and eye movement pattern, combined with the current surgical stage, calculate the communication urgency index:

[0094] ;

[0095] in, It represents the communication urgency index, which is used to quantify the urgency of medical staff's communication needs; Indicates gaze stability, calculated by the variance of gaze point coordinates. A smaller value indicates a more stable gaze, which may imply a stronger communication intention. It represents pupil response, that is, the rate of change of pupil diameter, which reflects the physiological response of medical staff to visual stimulation. Pupil dilation is usually more obvious in emergency situations. It represents the characteristics of eye saccade patterns, including information on eye saccade speed and direction distribution, which may manifest as faster and more frequent eye saccades in emergency situations. Indicates blinking patterns, including blink frequency and duration. Blinking patterns will change characteristically under stress or emergency conditions. It represents the surgical context factor, reflecting the urgency of the current surgical stage. For example, high-risk stages such as critical suturing and hemostasis have higher context urgency values. 、 、 、 、 They represent the weight coefficients of gaze stability, pupil response, saccade pattern, blink pattern and surgical context factor, respectively. They are optimized and determined based on experimental data, reflecting the contribution of each indicator to the urgency assessment.

[0096] Urgency rating: The calculated urgency index Mapped to four levels of communication priority: Level 1 (routine information exchange), Level 2 (important information confirmation), Level 3 (warning information transmission), and Level 4 (emergency intervention).

[0097] The output of this step includes the communication intention judgment result, the communication target person, and the communication urgency level. This information will be used for subsequent communication mode selection and microphone array beamforming control.

[0098] Multimodal collaboration intention inference and communication mode matching module, which combines eye movement interaction features with the collaboration relationship network to infer collaboration intentions and select the most suitable communication mode;

[0099] The specific steps include:

[0100] Step 3.1, multi-target eye movement interaction feature extraction;

[0101] Analyze the eye movement data of multiple medical staff in the operating room and extract the eye movement interaction features between them:

[0102] Paired gaze features: Calculate the behavioral features of medical staff A looking at medical staff B, including gaze frequency, gaze duration, and gaze intensity, expressed as ;

[0103] Mutual gaze behavior recognition: Detect mutual gaze behavior between medical staff and calculate the duration and frequency of mutual gaze, expressed as ;

[0104] Gaze sequence analysis: Analyze the gaze sequence pattern over a period of time, such as "A looks at B, then B looks at C", and extract possible information flow paths, expressed as .

[0105] The eye movement interaction feature matrix is ​​expressed as: , record the eye movement interaction between all medical staff in the operating room. Indicates the behavioral characteristics of medical staff A looking at medical staff B, including parameters such as gaze frequency, duration, and gaze intensity; Indicates the behavioral characteristics of medical staff B looking at medical staff C; Indicates the behavioral characteristics of medical staff A looking at medical staff C;

[0106] Each element in the matrix is ​​a vector containing multidimensional features such as gaze frequency (number of gazes per minute), gaze duration (the average duration of a single gaze in seconds), and gaze intensity (a value between 0 and 1 calculated by combining the degree of pupil dilation and gaze stability).

[0107] The matrix is ​​updated dynamically, and the system continuously updates the values ​​in the matrix based on real-time eye tracking data to reflect the ever-changing interactions between medical staff during the operation.

[0108] Step 3.2, collaborative intention reasoning model;

[0109] Combined with eye movement interaction feature matrix Based on the collaborative relationship network data of the medical team, we build a collaborative intention inference model and calculate the probability of different collaborative scenarios:

[0110] ;

[0111] in, Indicates the probability of the collaboration scenario between medical staff A, B, and C. The value is between 0 and 1. The larger the value, the more likely the collaboration scenario is to occur. The eye movement interaction characteristics of medical staff A looking at medical staff B include gaze frequency, gaze duration and gaze intensity; Indicates the eye movement interaction characteristics of medical staff B looking at medical staff C, including Same parameter dimensions; Indicates the eye movement interaction characteristics of medical staff A looking at medical staff C, including Same parameter dimensions; The weight of the role relationship between medical staff A and B obtained from the collaboration relationship map reflects the closeness of the professional collaboration between the two in the surgical team, such as the frequency and importance of collaboration between different role combinations such as the surgeon and assistant, anesthesiologist and nurse; Represents the role relationship weight between medical staff B and C obtained from the collaborative relationship map, and defines the same ; Represents the role relationship weight between medical staff A and C obtained from the collaborative relationship map, and defines the same ; It represents the current surgical stage. The factors affecting the collaboration mode at different surgical stages are as follows. Each stage has specific collaboration demand characteristics. It represents the collaboration intention inference function, which calculates the probability value of a specific collaboration scenario by integrating all the above parameters.

[0112] The collaborative intention reasoning model achieves accurate reasoning of complex collaborative intentions of multiple people by integrating graph neural networks and temporal attention mechanisms.

[0113] Model outputs include:

[0114] A collection of medical staff participating in the collaboration ,in 、 、 They represent the first 、 、 medical staff; represents the number of medical staff. Based on the eye movement interaction characteristics and collaboration relationships, the system infers that these medical staff need to participate in the current communication collaboration;

[0115] Collaboration Type , such as inquiry to answer, instruction to execution, information sharing, etc.;

[0116] Collaboration urgency ;

[0117] Step 3.3, communication mode selection and configuration;

[0118] Based on the collaboration intent inference results, select the most appropriate one from the following four communication modes:

[0119] Point-to-point direct connection mode: suitable for simple communication between two people, establishing a direct communication link between the sender and the receiver;

[0120] Group broadcast mode: suitable for scenarios where one person transmits information to multiple people, establishing a one-to-many communication link;

[0121] Sequential delivery mode: Applicable to scenarios where information needs to be delivered along a specific path, such as the information flow of "A→B→C";

[0122] Collaborative conversation mode: Suitable for complex collaborative scenarios where multiple people need to communicate simultaneously, establishing a multi-point interconnected communication network.

[0123] The communication mode selection algorithm is based on the following rules:

[0124] If the number of participants is 2 and the collaboration type is inquiry to answer or command to execution, select point-to-point direct connection mode;

[0125] If the number of participants is greater than 2 and there is a clear message initiator, select the group broadcast mode;

[0126] If the collaboration type is information flow and there is a sequence dependency, select the sequential delivery mode;

[0127] If the number of participants is greater than 2 and the collaboration type is discussion and decision-making, select the collaborative conversation mode.

[0128] For each communication mode, the system also Configure corresponding communication parameters, such as signal priority, bandwidth allocation, and noise reduction strength.

[0129] Step 3.4, initialization of virtual conference space;

[0130] For collaborative session mode, the system will initialize a virtual meeting space. The specific steps include:

[0131] Determine the set of session participants ,in:

[0132] Represents the set of all participants in the virtual meeting space; 、 、 Represents each medical staff participating in the conversation; Indicates the number of medical staff;

[0133] Assigning a virtual session identifier ;

[0134] Assign communication channels and priorities to each participant:

[0135] ;

[0136] in, Indicates the allocation to the participants A communication channel configuration set; Communication channel identifier, used to uniquely identify the communication channel assigned to the medical staff to ensure the correct routing of information transmission; Communication priority determines the order of information transmission when network resources are limited. High-priority medical staff (such as the surgeon) receive priority in information transmission. The allocated bandwidth, network resources dynamically allocated based on the medical staff's role and current communication needs, affects the quality and clarity of voice transmission.

[0137] The output of this step includes the selected communication mode type, the set of communication participants, the communication parameter configuration, and (if applicable) the virtual meeting space initialization data. This information will be used for subsequent microphone array beamforming and audio signal processing.

[0138] Adaptive microphone array beamforming module, which is used to dynamically adjust the beam direction and parameters of the microphone array according to the communication mode and participant location information;

[0139] The specific steps include:

[0140] Step 4.1, multimodal sensing microphone array system structure;

[0141] Build a multimodal perception microphone array system that combines vision, depth perception, and acoustic information. Its physical structure includes:

[0142] Microphone array unit: Multiple microphone sub-arrays are installed on the ceiling and walls of the operating room. Each sub-array contains 8 omnidirectional microphones, forming a uniform circular array. The sub-arrays form a collaborative capture network.

[0143] Visual perception unit: Equipped with multiple high-definition cameras and depth sensors, it is used to track the position and movements of medical staff and assist in personnel positioning.

[0144] Signal processing unit: Contains a dedicated DSP chip for real-time signal processing, a microphone array controller, and a beamforming processor.

[0145] The entire system is connected through a network to form a unified audio capture and processing platform, realizing the fusion analysis of multimodal data.

[0146] Step 4.2, multimodal target localization and tracking;

[0147] Combining visual and acoustic information for precise positioning and tracking of medical staff:

[0148] Visual positioning: Obtain the three-dimensional position coordinates of medical staff through depth cameras , achieving centimeter-level positioning accuracy;

[0149] Sound source localization: Calculate the direction of the sound source based on the voice signal received by the microphone array. The specific method is as follows:

[0150] ;

[0151] in, represents the time delay between microphones i and j, that is, the time difference between the sound waves reaching the two microphones; 、 Respectively represent the coordinate positions of microphones i and j in three-dimensional space; Indicates the speed of sound waves in air; is the unit vector of the sound source direction, represents the horizontal angle of the sound source, Represents the vertical angle of the sound source, which together determine the direction of the sound source in three-dimensional space;

[0152] Information fusion: The Kalman filter algorithm is used to fuse visual and acoustic positioning results to obtain a more accurate location of medical staff:

[0153] ;

[0154] in, For the current moment The target state vector contains the position and velocity information of the medical staff; For the previous moment The target state vector of is the state transition matrix, which describes how the system state evolves from one moment to the next; is the process noise, which represents the uncertainty in state prediction and is usually assumed to be Gaussian white noise with zero mean; For the current moment The observation vector is the position measurement of the medical staff obtained by visual and acoustic sensors; is the observation matrix, which maps the state space to the observation space; is the observation noise, which represents the uncertainty in the sensor measurement and is usually assumed to be Gaussian white noise with zero mean.

[0155] Step 4.3, adaptive beamforming algorithm;

[0156] Adaptive beamforming based on communication patterns and participant location information:

[0157] Spatial filter design: A spatial filter is designed for a specific direction, and beam directionality control is achieved by controlling the weight coefficients of each unit in the microphone array:

[0158] ;

[0159] in, is the array output signal, which represents the final audio signal obtained after beamforming processing; is the input signal of the mth microphone, representing the original audio data collected by the mth microphone at time t; is the complex conjugate weight coefficient, which is used to control the amplitude and phase of each microphone signal and determine the directivity and gain characteristics of the beam; is the total number of microphones in the microphone array, indicating the number of microphones participating in beamforming; Represents the process of weighted summation of all microphone signals, which is the core operation of beamforming;

[0160] Minimum Variance Distortionless Response (MVDR) beamforming: Design the optimal weight vector for a single communication target:

[0161] ;

[0162] in, is the weight vector of minimum variance distortionless response beamforming, which is used to control the gain and phase of each unit of the microphone array; is the noise covariance matrix, which characterizes the statistical characteristics of the ambient noise and is used to suppress interference signals from non-target directions; is the inverse matrix of the noise covariance matrix, which is used to calculate the optimal weight; is the steering vector in the target direction, describing the direction of the sound wave from the target The phase relationship of the signals arriving at each element of the microphone array; is the conjugate transpose of the steering vector, used to calculate the inner product; is the horizontal angle (azimuth) of the target sound source, indicating its direction in the horizontal plane; is the vertical angle (elevation angle) of the target sound source, indicating the inclination angle relative to the horizontal plane;

[0163] Multi-objective beamforming: For multi-objective scenarios such as collaborative conversations, a constrained optimization method is used to design the beamformer:

[0164] ;

[0165] in, is the beamforming weight vector, which represents the complex weight coefficient of each microphone; is the weight vector The conjugate transpose of is the input signal covariance matrix, which describes the statistical characteristics of the signal received by the microphone array; For the Target direction The steering vector represents the phase relationship of the sound wave propagating from this direction to each microphone; For the The gain constraint of a target controls the response strength of the beam in that direction; is the total number of targets, indicating the number of sound sources that need to be paid attention to at the same time; Represents finding the optimal weight vector Minimize the objective function; Indicates "subject to" and introduces constraints;

[0166] Adaptive interference suppression: While generating gain in the main beam direction, it also generates a null in the direction of the interference source.

[0167] ;

[0168] in, is the weight vector of the linear constrained minimum variance beamforming, which is used to control the gain and phase of each unit of the microphone array; is the covariance matrix of the input signal, which describes the statistical characteristics of the signal received by the microphone array; is the inverse matrix of the covariance matrix, which is used to calculate the optimal weight; is the steering matrix containing the target and interference directions, where each column represents a steering vector in a specific direction; It is the conjugate transpose of the orientation matrix, used for matrix operations; Inverse the matrix to satisfy the constraints; is the gain constraint vector, which specifies the expected response in each direction of interest, usually set to 1 for the target direction and 0 for the interference direction; and Respectively represent The horizontal angle (azimuth) and vertical angle (elevation) of the direction of interest; is the total number of attention directions, including the target direction and the interference direction that needs to be suppressed.

[0169] Step 4.4, beam strategy for communication mode adaptation;

[0170] Based on the communication mode determined in step 3, different beamforming strategies are implemented:

[0171] Point-to-point direct mode: forms a bidirectional main beam from the speaker to the receiver, while forming deep suppression in other directions;

[0172] Packet broadcast mode: forms a fan-shaped beam from the speaker to multiple receivers, covering all receiver locations;

[0173] Sequential transmission mode: according to the information flow path, the beam links between corresponding nodes are activated in sequence;

[0174] Collaborative session mode: Builds a multi-point interconnected beam network, allowing direct communication between any two participants.

[0175] Beam configuration parameters are dynamically adjusted according to the communication urgency. The higher the urgency, the greater the beam gain and the deeper the lateral suppression, ensuring clear transmission of critical information.

[0176] Step 4.5, dynamic tracking and adjustment;

[0177] Dynamic tracking and adjustment of the beam is achieved based on the movement of medical staff:

[0178] Position prediction: Based on the Kalman filter target tracking algorithm, the next moment position of the medical staff is predicted:

[0179] ;

[0180] in, for The predicted position value at the moment is based on Information at the moment; for The position estimate at the moment represents the best known position estimate at the moment; is the state transition matrix, which describes the target motion pattern and is used to map the current state to the predicted state at the next moment. This formula represents the prediction step of the Kalman filter, which predicts the next position of the medical staff based on the current state and the motion model.

[0181] Real-time beam update: Based on the predicted position, the beam direction is adjusted in advance to reduce tracking delay:

[0182] ;

[0183] in, For the next moment The beamforming weight vector controls the gain and phase of each unit in the microphone array; is the beam weight calculation function, which maps the predicted position to the optimal beamforming weight; Based on the current moment Information prediction for medical staff The position coordinates at the moment; this formula means that the system calculates and adjusts the beamforming weights in advance based on the predicted position of the medical staff to achieve predictive beam tracking.

[0184] Fast response mechanism: In response to sudden position changes, a fast beam switching algorithm is used to reduce beam adjustment time, ensuring stable sound pickup quality when medical staff move quickly.

[0185] The output of this step is a beam configuration that is dynamically adjusted based on the communication mode and participant location. This configuration includes the weight coefficients, main beam direction, and gain parameters for each microphone. These configurations directly control the signal processing of the microphone array, achieving accurate speech capture and noise suppression.

[0186] The virtual meeting space construction and dynamic optimization module is used to create a virtual meeting space for multi-person collaboration scenarios, provide a unified logical framework for multi-person collaborative communication, and adaptively adjust communication parameters according to changes in the surgical stage;

[0187] The specific steps include:

[0188] Step 5.1, virtual meeting space construction;

[0189] Based on collaborative session requirements, a Virtual Meeting Space (VMS) model is created to provide a unified logical framework for multi-person collaborative communication:

[0190] Spatial topology construction: Establish a communication topology based on the set of participants and collaborative relationships, and determine the signal transmission path;

[0191] Spatial node configuration: assigning a set of communication parameters to each participant node :

[0192] ;

[0193] in, The signal gain parameter controls the amplification of the node audio signal. A larger value indicates a higher volume. Signal priority determines the processing order when competing for resources. High-priority signals will be given priority transmission rights. Denoise Strength controls the intensity level of the noise suppression algorithm applied to this node. A larger value indicates a stronger noise reduction effect. The allocated bandwidth indicates the communication channel capacity allocated to the node, which affects the quality and delay of audio transmission;

[0194] Session management mechanism: Establishing session control policies, including:

[0195] Access control: determines who can join or leave the current session;

[0196] Communication authority control: assign different speaking priorities based on medical roles and urgency;

[0197] Session lifecycle management: Dynamically adjust session resources based on collaboration duration and state changes.

[0198] Step 5.2: audio channel optimization and noise reduction processing;

[0199] Optimize each communication channel based on the location of medical staff and ambient noise distribution:

[0200] Noise characteristic analysis: Use the ambient audio captured by the microphone array to analyze the noise characteristics and build a noise model matrix ,in, is the noise model matrix, which contains the characteristic information of various types of noise in the environment; 、 、 Respectively represent 、 、 Noise-like spectral characteristics; The total number of noise types identified for the system; is a frequency variable, representing the frequency point in spectrum analysis; the noise model matrix provides basic data for targeted processing of subsequent noise reduction algorithms, enabling the system to discriminately suppress different types of noise;

[0201] Adaptive noise suppression: Dynamically adjust noise suppression algorithm parameters based on noise characteristics and communication urgency:

[0202] For low-priority communications: adopt a strong noise suppression strategy to ensure voice clarity;

[0203] For high-priority communications: adopt a conservative noise suppression strategy to prioritize the timeliness and integrity of information;

[0204] Spatially Selective Noise Reduction: Applies directional noise reduction processing to each communication link in a virtual meeting:

[0205] ;

[0206] in, is the processed speech signal, which represents the output speech after noise reduction processing; is an adaptive filter, which is a time-varying function that depends on the following parameters: It is a time variable, which means that the filter will be adjusted dynamically over time; The urgency parameter determines the filtering strength. The higher the urgency, the more conservative the filtering is to retain more original information. It is the noise model matrix, which provides the environmental noise characteristics information to guide the filtering process; The original speech signal, i.e., the unprocessed speech input of medical staff; Represents the convolution operation, which represents the process of the filter processing the original signal;

[0207] Spectrum enhancement processing: Selective spectrum enhancement is implemented for medical terms and key instructions to improve the recognizability of important information:

[0208] ;

[0209] in, is the output signal spectrum after spectrum enhancement processing, indicating that the frequency Signal energy distribution at ; is the spectrum of the original input signal, which is expressed in terms of frequency Signal energy distribution at ; is the spectral gain function, according to the frequency and medical terminology keyword sets Dynamically adjust the gain size; It is a collection of medical terminology keywords, including medical professional vocabulary and key instructions that need special enhancement; is a frequency variable, representing the different frequency components in the sound signal; this formula achieves selective enhancement of the frequency range containing medical terms and key instructions, thereby improving the recognizability of these important information in a noisy environment.

[0210] Step 5.3, dynamic session optimization and prediction;

[0211] Dynamically predict and pre-adjust communication needs based on surgical progress and historical collaboration patterns:

[0212] Collaboration model learning: Build a team collaboration model by analyzing historical communication data:

[0213] ;

[0214] in, The collaboration model represents the communication and collaboration pattern of the medical team during surgery, including information such as the communication frequency, priority relationship, and typical interaction patterns among team members. A machine learning function is an algorithm used to extract and learn collaboration patterns from historical data. It can be based on deep learning or statistical models. Historical communication records include communication data between medical staff and nurses during past operations, such as communication time, initiator, receiver, content type, urgency, and other information; For the information on the type and stage of surgery, the professional classification of the surgery (e.g., cardiac surgery, neurosurgery, etc.) and the different stages of the surgery (e.g., anesthesia period, operation period, suturing period, etc.) were described;

[0215] Communication demand prediction: Based on the collaboration model and the current surgical status, predict possible future communication needs:

[0216] ;

[0217] in, For the predicted future moment The communication demand indicates that the system predicts the communication activities that may occur at a certain point in the future; is a prediction function used to infer future communication needs based on the collaboration pattern model and the current state; The collaboration model includes the data model of the team's historical collaboration patterns and communication rules; Current surgical status, including real-time information such as the current surgical stage, medical staff location, and equipment status; The current time point is used as the time reference; is the prediction time window, which indicates the length of time predicted from the current moment to the future;

[0218] Predictive resource allocation: Adjust system resource allocation in advance based on prediction results:

[0219] Pre-adjust the beam direction to pre-form an enhanced beam for the location of medical personnel where communication may be about to occur;

[0220] Pre-allocate processing resources to ensure timely system response when peak communication demands arrive;

[0221] Dynamically adjust noise reduction parameters to reduce noise suppression intensity during expected critical information delivery phases.

[0222] Step 5.4, communication strategy adaptation based on surgical stage perception;

[0223] Adaptively adjust communication strategies based on the characteristics of different surgical stages:

[0224] Surgical stage identification: Identify the current surgical stage based on the status of operating room equipment, the activity characteristics of medical staff, and changes in collaborative relationships:

[0225] ;

[0226] in, The surgical stage for identification; The surgical stage identification function is an algorithm used to map multiple input features to specific surgical stages after comprehensive analysis; is a set of device status parameters; It is a collection of medical staff activity characteristics, describing the location distribution, movement patterns, operation frequency and other behavioral characteristic data of surgical team members; is the collaborative relationship change parameter, which represents the dynamic characteristics of the collaborative network, such as the interaction pattern, communication frequency, and instruction flow among medical team members;

[0227] Phase adaptation strategy configuration: Configure corresponding communication strategies for different surgical phases:

[0228] Preparation phase: adopt loose access control and low priority differentiation;

[0229] During the critical operation phase: strict priority management is adopted to ensure that the surgeon's instructions are delivered first;

[0230] Emergency: Activate emergency communication mode to minimize delays and improve information transmission reliability.

[0231] The output of this step is a dynamically optimized virtual conference space model, which provides a unified logical framework and resource allocation mechanism for multi-person collaborative communication. It can adaptively adjust communication parameters according to surgical progress and collaboration needs, ensuring efficient intercom communication in complex surgical environments.

[0232] Application examples of this embodiment:

[0233] This implementation has been successfully applied in a robot-assisted heart valve repair surgery in the cardiothoracic surgery department of a tertiary hospital. The following is an example of this application scenario.

[0234] This heart valve repair surgery involved a six-person surgical team, including the lead surgeon, two assistant surgeons, an anesthesiologist, an instrumentation nurse, and a circulating nurse. The operating room simultaneously operated a heart-lung machine, monitors, a robotic system, and numerous auxiliary devices, resulting in an average ambient noise level of 75dBSPL, with peaks reaching 88dBSPL. The surgical process consisted of a preparation phase (30 minutes), anesthesia (25 minutes), thoracotomy (40 minutes), robotic operation (120 minutes), suturing (35 minutes), and finalization (20 minutes). Team members were required to maintain precise coordination throughout the entire procedure, especially during robotic operation, placing extremely high demands on communication quality and timeliness.

[0235] Implementation process example:

[0236] Examples of building a collaborative network for medical and nursing teams:

[0237] The system first collected communication data from 32 similar surgeries performed by the cardiac surgery team over the past six months and identified the following key role relationship characteristics:

[0238] Communication frequency characteristics between surgeons and anesthesiologists =0.43 times / minute;

[0239] Communication importance weight =0.92;

[0240] Communication dependency characteristics =0.89 (compulsive dependence).

[0241] Communication frequency characteristics between the surgeon and the first assistant =0.68 times / minute;

[0242] Communication importance weight =0.87;

[0243] Communication dependency characteristics =0.78 (compulsory dependence).

[0244] The system builds a complete collaborative relationship map , where the vertex set Contains 6 medical roles, side set The graph contains 15 pairs of role relationships, each characterized by three attributes: frequency, importance, and dependency. The graph automatically updates attribute values ​​based on the identified surgical phase. For example, during a robotic procedure, the system automatically increases the importance of communication between the surgeon and anesthesiologist from 0.70 during the preparation phase to 0.92.

[0245] Eye movement multi-target tracking and intention recognition example:

[0246] The system installed four high-precision, long-range eye-tracking cameras in the operating room, covering all areas where medical staff move. In one specific communication event, the system captured the surgeon looking at the anesthesiologist for 2.3 seconds. At this time:

[0247] Gaze stability =0.92 (coordinate variance is less than 2.1 mm);

[0248] pupil response =0.23 (pupil dilation rate 23%);

[0249] Saccadic pattern characteristics =0.19 (low-frequency, short-distance saccade pattern);

[0250] Blink Mode =0.12 (blinking frequency decreased).

[0251] Considering that the current phase is a critical one for the robot’s operation, the system calculates the probability of communication intention:

[0252] ;

[0253] Since this value exceeds the preset threshold of 0.75, the system determines that the surgeon has a clear intention to communicate. Further, the system analyzes the gaze direction and the location of the operating room personnel, identifies the anesthesiologist as the communication target, and calculates the urgency index based on the eye movement characteristics. , corresponding to the third level of priority (warning information transmission).

[0254] Multimodal collaborative intention reasoning and communication mode matching examples:

[0255] In another communication event, the system detected a complex three-person collaboration scenario:

[0256] The surgeon looks at the instrument nurse ( =0.76);

[0257] The instrument nurse looks at the first assistant ( =0.65);

[0258] At the same time, the first assistant and the surgeon briefly looked at each other ( =0.58).

[0259] Combined with the role relationship weight in the collaboration relationship graph ( =0.81, =0.64, =0.87) and the current robot operation phase information, the system infers: .

[0260] The system identified this as a "tool transfer and collaborative operation" scenario with an urgency of 0.75 (level 2 priority). Based on this reasoning, the system selected the "sequential transfer" communication mode, establishing a communication path from the surgeon to the instrument nurse to the first assistant, ensuring the precise delivery of complex instructions.

[0261] Adaptive microphone array beamforming example:

[0262] The system deploys four sub-microphone arrays in the operating room, each containing eight omnidirectional microphones to form a uniform circular array, which are installed at the four corners of the top of the operating room.

[0263] In the aforementioned communication event between the surgeon and the anesthesiologist, the system obtains the surgeon's coordinates (2.4m, 1.8m, 1.7m) and the anesthesiologist's coordinates (3.9m, 4.2m, 1.7m) based on the multimodal target positioning results. The system automatically calculates the optimal MVDR beamforming weight vector:

[0264] ;

[0265] This weight vector enables subarray 1 to form a precise beam directed toward the surgeon, while also providing deep suppression in the direction of the ventilator and heart-lung machine. Similarly, the system configures the weight vector for subarray 3 to capture the anesthesiologist's voice. This adaptive beam configuration achieves high-quality speech capture in an ambient noise level of 75dB SPL.

[0266] Examples of virtual meeting space construction and dynamic optimization:

[0267] In a four-person collaboration scenario involving a surgeon, two assistants, and an anesthesiologist, the system automatically established a virtual meeting space and assigned communication parameters to the four participants:

[0268] Surgeon:

[0269] ;

[0270] anesthetist:

[0271] ;

[0272] First Assistant:

[0273] ;

[0274] Second Assistant:

[0275] ;

[0276] Among them, gain represents the audio gain parameter; priority represents the communication priority parameter; noise_reduction represents the noise suppression strength parameter; bandwidth represents the bandwidth allocation parameter. The system also analyzes the results based on the noise characteristics.

[0277] For the 63-250Hz frequency band (mainly cardiopulmonary machine noise): suppression intensity 0.85;

[0278] For the 500-1000Hz frequency band (mainly monitor alarm sounds): suppression intensity 0.65;

[0279] For the 2000-4000 Hz frequency band (critical speech frequency band): suppression strength 0.40.

[0280] During the critical stage of the surgery (the valve repair operation), the system predicted that there might be important communication needs between the surgeon and the anesthesiologist, adjusted the beam direction in advance, and pre-allocated processing resources. When actual communication occurred, the system response time was only 12ms, far lower than the 50-100ms of traditional systems.

[0281] Technical effect verification:

[0282] A quantitative verification test was conducted on the core technical effects of this implementation method, and the results are as follows:

[0283] Communication quality and efficiency improvement effects:

[0284] The test compared the changes in key indicators before and after using this system in 32 similar surgeries:

[0285] Comparison of Speech Intelligibility Index (STOI):

[0286]

[0287] Comparison of speaker localization accuracy and response delay:

[0288]

[0289] Team collaboration and cognitive load improvement effects:

[0290] By combining questionnaires and objective indicators, changes in team collaboration efficiency and cognitive load were assessed:

[0291] Team collaboration efficiency indicators:

[0292]

[0293] Cognitive load assessment results:

[0294]

[0295] The above verification results show that this implementation significantly improves the quality and efficiency of communication in actual medical scenarios, optimizes the team collaboration experience, reduces the cognitive load of medical staff, and effectively solves the technical challenges faced by doctors and nurses in intercom communications in robotic surgery environments.

[0296] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. A noise reduction intercom system for doctors and nurses in robotic surgery environments, featuring: include: The medical team collaboration relationship network construction module is used to construct a professional collaboration relationship map for the surgical environment through graph structure representation and role relationship analysis. The medical team collaboration relationship network construction module includes: Collecting medical and nursing role and relationship characteristics, including communication frequency, communication importance, and communication dependency; Build a collaborative relationship map: ; in, Represents the collaborative relationship graph, vertex set Represents the various roles in the medical team, edge set Indicates the collaborative relationship between roles; Dynamically adjust edge attributes in the collaborative relationship graph based on real-time identified surgical stages; Eye movement multi-target tracking and intention recognition module, which is used to monitor medical staff's eye movements through a high-precision long-distance eye tracking camera and identify communication intentions and targets; The multimodal collaboration intention reasoning and communication mode matching module is used to combine eye movement interaction features with the collaboration relationship network to infer collaboration intentions and select the most suitable communication mode. The multimodal collaboration intention reasoning and communication mode matching module includes: Extract multi-target eye movement interaction features, including paired gaze features, mutual gaze behavior recognition, and gaze sequence analysis; Build a collaborative intention reasoning model to calculate the probability of different collaborative scenarios; Select the most suitable communication mode from point-to-point direct connection mode, packet broadcast mode, sequential delivery mode and collaborative session mode; Adaptive microphone array beamforming module, which is used to dynamically adjust the beam direction and parameters of the microphone array according to the communication mode and participant location information; The virtual meeting space construction and dynamic optimization module is used to create a virtual meeting space for multi-person collaboration scenarios, provide a unified logical framework for multi-person collaborative communication, and adaptively adjust communication parameters according to changes in the surgical stage.

2. The noise reduction intercom system for doctors and nurses in a robotic surgery environment according to claim 1 is characterized in that: The eye movement multi-target tracking and intention recognition module includes: An infrared lighting unit and a high-frame-rate imaging unit are used to collect eye movement data of medical staff; Construct a multi-layered eye movement intention recognition model, including feature extraction layer, context fusion layer, and intention judgment layer; The communication target was determined by analyzing gaze direction and the position of medical staff in the operating room, and the urgency of communication was assessed based on gaze duration, pupil dilation rate, and eye movement patterns.

3. The noise reduction intercom system for doctors and nurses in a robotic surgery environment according to claim 1 is characterized in that: The virtual conference space construction and dynamic optimization module includes: Create a virtual conference space model, including space topology construction, space node configuration, and session management mechanism; Optimize each communication channel based on the location of medical staff and the distribution of ambient noise; By analyzing historical communication data, we can build a team collaboration model and predict possible future communication needs; Adaptively adjust the communication strategy according to the characteristics of different surgical stages.

4. The noise reduction intercom system for doctors and nurses in a robotic surgery environment according to claim 1 is characterized in that: The collaborative intention reasoning model is based on a graph neural network and a temporal attention mechanism. Its input includes an eye movement interaction feature matrix, role relationship weights, and current surgical stage information. Its output includes the set of medical staff participating in the collaboration, the collaboration type, and the collaboration urgency.

5. The noise reduction intercom system for doctors and nurses in a robotic surgery environment according to claim 1 is characterized in that: The communication pattern matching is based on the following rules: If the number of participants is 2 and the collaboration type is inquiry to answer or command to execution, select point-to-point direct connection mode; If the number of participants is greater than 2 and there is a clear message initiator, select the group broadcast mode; If the collaboration type is information flow and there is a sequence dependency, select the sequential delivery mode; If the number of participants is greater than 2 and the collaboration type is discussion and decision-making, select the collaborative conversation mode.

6. The noise reduction intercom system for doctors and nurses in a robotic surgery environment according to claim 1 is characterized in that: The adaptive microphone array beamforming module includes: Deploy multiple sub-microphone arrays, each containing multiple omnidirectional microphones, to form a collaborative capture network; Combine visual and acoustic information to accurately locate and track medical staff; Implement beamforming algorithms including spatial filter design, minimum variance distortion-free response beamforming, multi-target beamforming, and adaptive interference suppression based on communication patterns and participant location information; According to the movement of medical staff, Kalman filtering is used to predict their position and realize dynamic tracking and adjustment of the beam.

7. The noise reduction intercom system for doctors and nurses in a robotic surgery environment according to claim 1 is characterized in that: Also includes: A multimodal perception unit, including multiple microphone arrays mounted on the operating room ceiling and walls, a high-precision long-range eye-tracking camera, a high-definition camera, and a depth sensor; Signal processing unit, including a dedicated DSP chip for real-time signal processing, a microphone array controller, and a beamforming processor; The network transmission unit is used to transmit the data collected by each unit to the central processing system to realize the fusion analysis of multimodal data.

8. A computer-readable storage medium, characterized in that It is used to store computer-readable instructions, and when the computer-readable instructions are read by a computer, it can run the doctor-nurse noise reduction intercom system for a robotic surgery environment as described in any one of claims 1-7.

Citation Information

Patent Citations

  • High-simulation multimedia ophthalmological doctor-patient communication system

    CN105117589A

  • Fault-tolerant method for improving underwater robot networking robustness

    CN118741573A