A surgical robot dialogue method and system based on audio processing technology
By converting surgical navigation data into personalized spatial audio and applying auditory significance enhancement processing, the problem of excessive visual burden on doctors and in timely perception of key information in precision surgery is solved, and the doctors "see" navigation information in auditory, improving surgical efficiency and safety.
Patent Information
- Application Number
- CN202510764792.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
In the prior art, doctors need to frequently divert their sight to view navigation information during precision surgery, resulting in low surgical efficiency and increased risk. Traditional sound prompts cannot effectively convey location-related navigation information, cannot adapt to the auditory perception differences of different doctors, and key information is easily covered up or ignored.
Convert surgical navigation data into personalized spatial audio, build an individualized auditory characteristic model by measuring the doctor's auditory spatial perception ability, adjust spatial audio parameters, build a human auditory significance calculation model, calculate the importance of navigation information in real time, and apply auditory saliency feature enhancement processing to make key information more prominent in spatial audio.
It reduces the number of sight transfer times doctors in precision surgery, improves navigation accuracy and information perception speed, reduces cognitive load and risk of misoperation, adapts to the auditory characteristics and environmental changes of different doctors, and has stable long-term use effect.
Smart Images

Figure CN120267402B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical equipment control technology, and more specifically, to a surgical robot dialogue method and system based on audio processing technology. Background Art
[0002] During precision surgeries, surgeons must simultaneously focus on the surgical area and navigation information, while also receiving and processing a wide range of audio information from multiple directions. Traditional navigation systems rely primarily on visual presentation, forcing surgeons to frequently shift their gaze, which not only impacts surgical efficiency but can also increase surgical risks. Furthermore, the complex and diverse audio information in the surgical environment poses a challenge to surgeons' limited auditory attention resources, making it difficult to accurately perceive and distinguish key information in spatial audio while maintaining focused attention on the surgery.
[0003] In the existing technology, some systems attempt to use simple sound prompts to assist navigation, but these systems generally have the following problems: First, traditional sound prompts lack spatial positioning capabilities and cannot effectively convey location-related navigation information; second, standardized audio parameters cannot adapt to the differences in auditory perception among different doctors, resulting in some doctors having difficulty accurately perceiving navigation information; in addition, in complex surgical acoustic environments, key navigation information can easily be masked by other audio information or ignored by the doctor's attention filtering mechanism. Especially in emergency situations, important information may not attract the doctor's attention in time.
[0004] Therefore, a technical solution is needed that can convert surgical navigation information into personalized spatial audio and dynamically adjust the auditory salience according to the importance of the information, so as to solve the technical problems of doctors' excessive visual burden and untimely perception of key information during precision surgery. Summary of the Invention
[0005] The present invention provides a surgical robot dialogue method and system based on audio processing technology, which solves the technical problems in the prior art of excessive visual burden on doctors during precision surgery and untimely perception of key information.
[0006] A first aspect of the present invention provides a surgical robot dialogue method based on audio processing technology, comprising the following steps:
[0007] Convert surgical navigation data into sound parameters and create a sound coding system that expresses different semantic information;
[0008] Based on the sound coding system, the doctor's auditory spatial perception ability is measured and an individualized auditory characteristic model is constructed;
[0009] Based on the doctor's auditory characteristics model, spatial audio parameters are adjusted to ensure that surgical navigation information is delivered in the optimal spatial audio format;
[0010] Based on the adjusted spatial audio parameters, a computational model of human auditory saliency is constructed to quantify the effect of different acoustic features on guiding auditory attention.
[0011] Based on the human auditory saliency calculation model, the importance of different navigation information is calculated in real time according to the surgical stage, risk level and time urgency, and a dynamic priority value is assigned to each piece of information;
[0012] Based on the dynamic priority value, the auditory saliency feature enhancement processing most suitable for the current environment is applied to high-priority information, making it more prominent in the spatial audio.
[0013] Furthermore, the step of converting the surgical navigation data into sound parameters includes:
[0014] The surgical navigation data are divided into four types according to type and function: location type, boundary type, structure type and operation prompt type;
[0015] Establishing a mapping relationship between navigation data and sound parameters, wherein position information is mapped to spatial position and pitch, boundary information is mapped to timbre and rhythm, structure information is mapped to timbre and duration, and operation prompts are mapped to specific tone sequences;
[0016] Build standardized sound coding templates for different types of navigation information to ensure that different types of information have differentiated but easily recognizable sound characteristics;
[0017] According to the importance and urgency of the navigation information, the sound coding parameters are dynamically adjusted to make important information more prominent in the sound expression.
[0018] Furthermore, the step of measuring the doctor's auditory spatial perception ability includes:
[0019] By playing standard test sounds in different directions, the doctor is asked to indicate the perceived direction of the sound and record the deviation between the actual direction and the perceived direction;
[0020] By playing sounds at different virtual distances, the accuracy of doctors' perception of changes in sound distance is measured;
[0021] Play a sound source moving in a virtual space and measure the doctor's ability to track the moving sound source and reaction time;
[0022] Based on measuring the accuracy of doctors' perception of changes in sound distance and their ability and reaction time to track moving sound sources, a mathematical model describing the characteristics of doctors' auditory spatial perception is established.
[0023] Furthermore, the step of adjusting the spatial audio parameters includes:
[0024] Selecting or synthesizing the most suitable head-related transfer function for the doctor from a head-related transfer function database according to the doctor's auditory spatial perception characteristics;
[0025] Targeted compensation for imbalances or defects in the doctor's hearing characteristics;
[0026] According to the spatial orientation perception deviation shown by the doctor in the measurement, the spatial localization parameters of the sound are adjusted to correct the perception deviation;
[0027] By adjusting the reverberation ratio and near-field effect acoustic parameters of the sound, the doctor's accurate perception of the distance to the sound source is enhanced;
[0028] During use, the doctor's response to spatial audio is continuously monitored, and parameters are dynamically adjusted to adapt to changes in the doctor's auditory adaptation during long operations.
[0029] Furthermore, the step of constructing a human auditory saliency calculation model includes:
[0030] Classify various types of information during surgery according to their importance and urgency, and assign different auditory salience processing strategies;
[0031] Generate corresponding auditory saliency parameters for information of different importance levels, including pitch, timbre, rhythm, and spatial position;
[0032] When multiple sound sources exist simultaneously, the separation between target information and background information is enhanced, and the auditory masking effect is reduced;
[0033] Designing long-term time series with clear auditory patterns enables doctors to perceive surgical progress and key stages through changes in auditory patterns;
[0034] Based on the doctor's auditory preferences and cognitive habits, personalized auditory symbols are designed to establish intuitive auditory information mapping.
[0035] Furthermore, the step of calculating the importance of different navigation information in real time includes:
[0036] Based on surgical process knowledge and risk management principles, a model for evaluating the importance of navigation information was constructed. The model considered multidimensional characteristics such as surgical stage, patient status, information type, and time urgency.
[0037] Based on surgical progress data and environmental information, the system can identify the current stage of the surgery in real time and adjust the basic priority of different types of information accordingly;
[0038] Based on the navigation information content and the current surgical status, the correlation between navigation information and safety risks is evaluated, and high-risk related information is given higher priority;
[0039] The final priority score of each navigation information is calculated based on the surgical stage, patient status, information type, time urgency, basic priority and safety risk correlation factors, and divided into four levels: emergency, important, routine and background.
[0040] Furthermore, the information transmission mechanism step is also included:
[0041] Analyze the type, complexity, and urgency of medical information and determine appropriate information delivery mechanisms;
[0042] Real-time assessment of doctors' cognitive load status to avoid information overload;
[0043] Transform complex medical information into intuitively understandable auditory patterns, reducing the cognitive cost of decoding;
[0044] When doctors need to focus on specific information, use a gradual attention guidance strategy to avoid abrupt interruptions;
[0045] Multiple auditory features are used to encode key information to improve the robustness of information transmission.
[0046] Furthermore, feedback interaction design steps are also included:
[0047] By recording doctors' interactive behaviors, we can learn their personalized auditory preferences, including volume, pitch, and information density;
[0048] Enable doctors to control the auditory information system in real time through short voice commands;
[0049] Provide touch screen or eye control alternatives for important controls to adapt to different surgical scenarios;
[0050] Dynamically adjust interaction complexity based on the doctor's familiarity with the system to reduce learning costs;
[0051] Regularly collect feedback from doctors on their experience with the system and continuously improve the interaction design.
[0052] Furthermore, it also includes experimental verification and tuning steps:
[0053] Test system performance in a simulated surgical environment to evaluate the effectiveness of auditory information delivery;
[0054] Collect usage feedback from doctors and surgical teams to understand the problems encountered in actual use of the system;
[0055] Comprehensively evaluate system performance based on objective indicators and subjective evaluation;
[0056] Based on the verification results, the system parameters and algorithms are tuned to improve system performance.
[0057] A second aspect of the present invention provides a surgical robot dialogue system based on audio processing technology, which is used to execute the above-mentioned surgical robot dialogue method based on audio processing technology, including:
[0058] The surgical navigation data sound conversion module is used to convert surgical navigation data into sound parameters and create a sound coding system that expresses different semantic information;
[0059] Doctor's auditory characteristics measurement module, used to measure doctors' auditory spatial perception ability and build an individualized auditory characteristics model;
[0060] A spatial audio customization module is used to adjust spatial audio parameters based on the doctor's auditory characteristics model to ensure that surgical navigation information is delivered in the optimal spatial audio format;
[0061] Auditory saliency analysis module, used to build a computational model of human auditory saliency and quantify the effect of different acoustic features on guiding auditory attention;
[0062] The information priority dynamic assessment module is used to calculate the importance of different navigation information in real time according to the surgical stage, risk level and time urgency, and assign a dynamic priority value to each piece of information;
[0063] The key information highlighting module is used to apply the auditory saliency feature enhancement processing that is most suitable for the current environment to high-priority information, making it more prominent in spatial audio.
[0064] The beneficial effects of the present invention are: by converting surgical navigation data into personalized spatial audio information and achieving auditory salience enhancement for key information, the technical problem of doctors needing to frequently shift their gaze to check navigation information during precision surgery is solved, and significant technical effects are achieved. Through personalized HRTF adjustment, the problem of different doctors' different 3D audio perception abilities is solved, so that doctors with weaker auditory spatial perception abilities can also accurately perceive spatial navigation information. Through auditory salience analysis and dynamic information priority evaluation, the perception speed of high-priority information is improved, and in emergency situations, it can break through the attention barrier to attract the doctor's immediate attention, reducing the risk of misoperation. The adaptive design of this method enables the system to be continuously optimized according to the doctor's auditory characteristics, workload and environmental changes, and the effect does not decay with long-term use, effectively solving the problem of reduced effect of 3D audio navigation systems in the prior art after a period of use. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is an overall flow chart of the surgical robot dialogue method based on audio processing technology of the present invention. DETAILED DESCRIPTION
[0066] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.
[0067] At least one embodiment of the present invention discloses a surgical robot dialogue method based on audio processing technology, such as Figure 1 As shown, the following steps are included:
[0068] Step 1: Convert surgical navigation data into sound parameters and create a sound coding system that expresses different semantic information;
[0069] This step converts surgical navigation data (such as target location, safety margin, key structures, etc.) into sound parameters and creates a sound coding system that expresses different semantic information. Specifically, it includes:
[0070] Step 1.1. Navigation data classification: surgical navigation data are classified into four types according to type and function: position type, boundary type, structure type, and operation prompt type;
[0071] Step 1.2, Sound Parameter Mapping: Establish a mapping relationship between navigation data and sound parameters, where position information is mapped to spatial position and pitch, boundary information is mapped to timbre and rhythm, structure information is mapped to timbre and duration, and operation prompts are mapped to specific tone sequences;
[0072] Step 1.3: Constructing semantic sound coding templates: Build standardized sound coding templates for different types of navigation information to ensure that different types of information have distinct but easily recognizable sound characteristics.
[0073] Step 1.4, dynamic adjustment of encoding parameters: Dynamically adjust the sound encoding parameters according to the importance and urgency of the navigation information to make the important information more prominent in the sound expression.
[0074] Step 2: Based on the sound coding system, measure the doctor's auditory spatial perception ability and build an individualized auditory characteristic model;
[0075] It should be understood that this step measures the doctor's perception accuracy of sound sources of different directions, distances, and movements through a rapid auditory spatial perception ability assessment program, and constructs a personalized auditory characteristic model for the doctor. Specifically, it includes:
[0076] Step 2.1. Spatial hearing accuracy measurement: Standard test tones in different directions (horizontally and vertically) are played, and the physician is asked to indicate the perceived direction of the sound. The deviation between the actual direction and the perceived direction is recorded.
[0077] In some embodiments, spatial hearing accuracy measurement can be achieved using an interactive method based on a virtual reality headset. In the virtual environment, the physician marks the perceived direction of sounds by turning their head or pointing with hand gestures. The system automatically records angular deviations and generates a spatial perception accuracy heat map. Alternatively, for situations requiring higher accuracy or where virtual reality equipment is unavailable, the system can play test tones through a surround speaker array, with the physician indicating the sound direction using a dedicated controller.
[0078] Step 2.2, distance perception assessment: by playing sounds at different virtual distances, measure the doctor's perception accuracy of changes in sound distance;
[0079] For example, in one embodiment, distance perception can be assessed using technology based on the near-field effect of the head and the adjustment of the far-field reverberation ratio. The system plays test tones with different near-field and far-field characteristics in a fixed direction. The physician then marks the perceived distance value. The system then compares the actual distance with the perceived distance to construct a personalized distance perception mapping function.
[0080] Step 2.3, dynamic sound source tracking test: Play a sound source moving in the virtual space and measure the doctor's ability and reaction time to track the moving sound source;
[0081] Alternatively, dynamic sound source tracking testing can be performed by simulating movement trajectories of varying speeds and complexity in a virtual sound field. In a preferred embodiment, the system generates sound source trajectories with various motion patterns, including linear, curved, and random jumps. The physician completes the test by continuously tracking or predicting the endpoint position. The system records tracking accuracy and response latency data.
[0082] Step 2.4: Constructing an individualized auditory characteristics model: Based on the above measurement data, a mathematical model is established to describe the doctor's auditory spatial perception characteristics. The model can be expressed as:
[0083] ;
[0084] in, represents the doctor’s personal auditory characteristic model, Indicates that the doctor is at a horizontal angle and vertical angle The spatial auditory sensitivity distribution on The first element of the collection, represents the distance perception function, which indicates the doctor's perception accuracy of sound sources at different distances. The second element of the collection, It represents the dynamic sound source tracking capability parameter, which indicates the doctor's ability to track the moving sound source and the reaction time. The third element of the collection. Indicates the horizontal angle in degrees; Indicates the vertical angle in degrees; Indicates distance, indicating the spatial distance between the sound source and the listener, in meters or centimeters, with subscript Represents an individual (personal).
[0085] In some embodiments, the personalized auditory characteristics model can be implemented using a Bayesian network structure, representing multidimensional auditory spatial perception characteristics as a probabilistic graphical model, capable of processing the correlation between perception abilities in different dimensions. For example, the system may discover that a particular doctor's directional perception is particularly accurate within the forward 60° range, but relatively weak in the rearward region. In this case, the system will perform more pronounced directional enhancement processing for the rearward information in subsequent spatial audio presentation.
[0086] Step 3: Based on the doctor's auditory characteristics model, adjust the spatial audio parameters to ensure that the surgical navigation information is delivered in the optimal spatial audio format;
[0087] Specifically include:
[0088] Step 3.1, Personalized HRTF Generation: Based on the doctor's auditory spatial perception characteristics, the most suitable head-related transfer function for the doctor is selected or synthesized from the HRTF database;
[0089] In some embodiments, personalized HRTF generation can be achieved using machine learning methods. For example, a neural network model can be used to predict the most suitable HRTF parameters based on the physician's auditory characteristics. Optionally, the system can also incorporate the geometric features of the physician's head and auricle (obtained through 3D scanning) to construct a more accurate personalized HRTF using physical acoustic simulation methods.
[0090] Step 3.2: Acoustic compensation processing: Targeted compensation is performed for imbalances or defects in the doctor's hearing characteristics, such as differences in left and right ear sensitivity, and hearing loss at specific frequencies;
[0091] For example, in one embodiment, for physicians with differing left and right ear sensitivities, the system can compensate through frequency-selective gain adjustment. Alternatively, for physicians with high-frequency hearing loss, the system can employ spectral remapping technology to shift high-frequency information to the mid-frequency region while preserving the spatial localization characteristics of sound.
[0092] Step 3.3, spatial positioning parameter adjustment: Based on the spatial orientation perception deviation displayed by the doctor during the measurement, the system adjusts the spatial positioning parameters of the sound to correct the perception deviation;
[0093] In some embodiments, spatial localization parameters can be adjusted by adjusting interaural time difference, intensity difference, and spectral cues. Optionally, the system can employ an adaptive algorithm to continuously optimize spatial localization parameters based on physician feedback during use. For example, if a physician consistently detects a systematic deviation in sound source localization in a particular direction, the algorithm will automatically adjust the spatial audio parameters in that direction until optimal perception is achieved.
[0094] Step 3.4: Enhanced distance perception: By adjusting acoustic parameters such as the reverberation ratio and proximity effect of the sound, the doctor's accurate perception of the distance to the sound source is enhanced;
[0095] For example, in a preferred embodiment, enhanced distance perception can be achieved by combining multiple acoustic cues. The system dynamically adjusts parameters such as the ratio of direct sound to reflected sound, low-frequency gain, and high-frequency attenuation to create a more realistic distance perception. Optionally, for physicians with limited distance perception, the system can incorporate additional non-auditory feedback (such as tactile feedback) as an aid.
[0096] Step 3.5, dynamic parameter adaptation: During use, the system continuously monitors the doctor's response to spatial audio and dynamically adjusts parameters to adapt to the doctor's auditory adaptation changes during long operations.
[0097] In some embodiments, dynamic parameter adaptation can be achieved by building a physician auditory fatigue model. The system monitors changes in the physician's response time to spatial audio cues and automatically adjusts parameters such as the intensity and salience of the spatial audio when a delay in response is detected, possibly due to auditory fatigue. For example, in the later stages of a long surgery, the system might increase the salience of important information or reduce the frequency of presentation of non-critical information to reduce auditory cognitive burden.
[0098] Step 4: Based on the adjusted spatial audio parameters, a human auditory saliency computational model is constructed to quantify the effect of different acoustic features on guiding auditory attention;
[0099] Specifically include:
[0100] Step 4.1: Information priority classification: Classify various types of information during surgery according to importance and urgency, and assign different auditory salience processing strategies;
[0101] In some embodiments, information prioritization can be implemented using a multi-layered hierarchical model. For example, information can be categorized into multiple levels, such as critical information (e.g., catheter contact with sensitive tissue), critical information (e.g., catheter approaching target location), and routine information (e.g., catheter movement within a safe path). Optionally, the system can dynamically adjust information priority based on the surgical stage, increasing the priority of certain types of information during critical surgical phases.
[0102] Step 4.2: Generating auditory salience parameters: Generate corresponding auditory salience parameters for information of different importance levels, including pitch, timbre, rhythm, spatial position, etc.
[0103] In an optional embodiment, saliency parameter generation can utilize a bio-inspired auditory saliency model. This model simulates the human auditory system's selective attention to sound features, utilizing features such as audio spectrum contrast and temporal abruptness to calculate a saliency map. Alternatively, the system can employ machine learning methods to optimize the saliency parameter generation algorithm by analyzing physicians' attentional responses to different combinations of sound features.
[0104] Step 4.3: Enhance auditory stream separation: When multiple sound sources are present simultaneously, enhance the separation between target information and background information to reduce auditory masking effects.
[0105] For example, in one embodiment, auditory stream separation can be achieved by adjusting the harmonic structure and temporal envelope of different sound sources. For high-priority information, the system assigns unique harmonic structures and rhythmic patterns, allowing it to form independent auditory streams. Optionally, the system can also utilize spatial separation principles to present information of different priorities at different locations in the auditory space, further enhancing the separation effect.
[0106] Step 4.4: Long-term temporal pattern design: Design a long-term temporal sequence with clear auditory patterns, enabling the surgeon to perceive the surgical progress and key stages through auditory pattern changes.
[0107] In some embodiments, long-term temporal patterns can be designed using a hierarchical rhythmic structure. For example, at the micro level, the system can use specific timbres and rhythms to represent the immediate status, while at the macro level, a musicalized melodic progression can be used to represent the progress of surgical procedures. Optionally, the system can also dynamically adjust the musical tension based on the degree of deviation between the planned surgical path and the actual operation, providing the surgeon with an intuitive sense of path compliance.
[0108] Step 4.5: Personalized auditory symbol design: Design personalized auditory symbols based on the doctor's auditory preferences and cognitive habits to establish an intuitive auditory-information mapping.
[0109] For example, in one embodiment, personalized auditory symbols can be designed based on the physician's professional background and auditory habits. For physicians with musical training, the system can employ more complex musical expressions; for physicians without musical backgrounds, more intuitive sound effects can be used. Optionally, the system can also involve physicians in the selection and optimization of auditory symbols through interactive learning, fostering a stronger connection between hearing and information.
[0110] Step 5: Based on the human auditory saliency calculation model, the importance of different navigation information is calculated in real time according to the surgical stage, risk level, and time urgency, and a dynamic priority value is assigned to each piece of information;
[0111] Specifically include:
[0112] Step 5.1: Construct an information priority assessment model: Based on surgical process knowledge and risk management principles, construct a model to assess the importance of navigation information. The model considers multiple dimensions such as surgical stage, patient status, information type, and time urgency.
[0113] In an embodiment of the present application, the information priority assessment model is implemented using a hybrid approach combining the Analytic Hierarchy Process (AHP) and Support Vector Regression (SVR). The AHP is used to establish a hierarchical structure for information priorities, categorizing priority influencing factors into surgical process categories (e.g., current surgical stage, expected next operation), risk categories (e.g., degree of deviation from the safe zone, proximity to critical tissue), patient status categories (e.g., changes in vital signs, tissue response), and time urgency categories (e.g., response time window, potential risk growth rate). Then, based on expert knowledge and historical cases, a Support Vector Regression algorithm is used to learn the nonlinear mapping relationship between each factor and the final priority.
[0114] In a spinal surgery application example, when the system detects that the surgical instrument is approaching the nerve root, the model comprehensively considers factors such as whether the current stage is decompression or fixation, the rate of change of the distance of the nerve root, and the amplitude of change of the intraoperative nerve monitoring signal, and automatically adjusts the priority of the warning information to the emergency level. It immediately attracts the doctor's attention through specific sound feature enhancement methods, allowing the doctor to adjust the operation path in time before nerve damage occurs. In the test, 92% of potential nerve root injury risk events were successfully avoided.
[0115] Step 5.2, surgical stage identification: Based on surgical progress data and environmental information, the current surgical stage is identified in real time, and the basic priority of different types of information is adjusted accordingly;
[0116] Step 5.3, Risk Level Assessment: Based on the navigation information content and the current surgical status, the relevance of the information to safety risks is assessed, with high-risk related information receiving higher priority;
[0117] Step 5.4, information priority calculation and classification: Based on the above factors, the final priority score of each navigation information is calculated and divided into four levels: emergency level, important level, routine level and background level. The priority calculation function can be expressed as:
[0118] ;
[0119] in, Display information The priority score ranges from 0 to 1. Indicates the score of relevance to the surgical stage, evaluation information The degree of relevance to the current stage of surgery, Indicates risk correlation score, assessment information The degree of association with security risks, Indicates the information practicality score, evaluates information Practical value for current surgical procedures, Indicates time urgency score, evaluation information the time urgency of the response required, 、 、 and They represent the weight coefficient of the relevance of the surgical stage, the weight coefficient of the risk relevance, the weight coefficient of the information practicality, and the weight coefficient of the time urgency, respectively. These weight coefficients are used to adjust the importance of each component in the final priority. Represents navigation information, indicating the surgical navigation data items that need to be delivered to the doctor.
[0120] Step 6: Based on the dynamic priority value, apply the auditory saliency feature enhancement processing that is most suitable for the current environment to the high-priority information, making it more prominent in the spatial audio;
[0121] Specifically include:
[0122] Step 6.1, Information transmission mechanism: Analyze the type, complexity, urgency and other characteristics of medical information to determine the appropriate information transmission mechanism; assess the physician's cognitive load in real time to avoid information overload; convert complex medical information into intuitive and understandable auditory patterns to reduce the cognitive cost of decoding; when the physician's attention needs to be focused on specific information, adopt a progressive attention guidance strategy to avoid abrupt interruptions; use multiple auditory feature encoding for key information to improve the robustness of information transmission.
[0123] In an optional embodiment, information characteristics analysis may utilize a multidimensional characteristic assessment model. This model categorizes information based on dimensions such as timeliness (urgent / non-urgent), complexity (simple / complex), spatial relevance (location-dependent / location-independent), and certainty (certain / uncertain). For example, for urgent, simple, location-dependent, and certain information such as "catheter contacting the vessel wall," the system might select a direct audio warning. For non-urgent, complex, location-dependent, but uncertain information such as "abnormal tissue density," the system might select a progressive audio prompt accompanied by a detailed explanation.
[0124] Step 6.2, Feedback Interaction Design: By recording doctors' interactive behaviors, learn their personalized auditory preferences, such as volume, pitch, and information density; enable doctors to control the auditory information system in real time through short voice commands; provide alternative methods such as touch screen or eye movement control for important controls to adapt to the needs of different surgical scenarios; dynamically adjust the interaction complexity based on the doctor's familiarity with the system to reduce learning costs; regularly collect doctors' feedback on their experience with the system and continuously improve the interaction design.
[0125] In an alternative embodiment, personalized preference learning can employ a hybrid model approach. The system first establishes an initial preference model based on explicit user feedback (e.g., manual settings adjustments). The model is then continuously updated through implicit feedback (e.g., whether the volume is lowered during use, whether information is requested to be repeated, etc.). For example, if the system detects that a physician repeatedly requests a lower volume for secondary information during a particular surgical procedure, the system will automatically learn to reduce the prominence of secondary information during that procedure.
[0126] Step 6.3, Experimental Verification and Tuning: Test system performance in a simulated surgical environment to evaluate the effectiveness of auditory information transmission; collect usage feedback from doctors and surgical teams to understand the problems encountered in actual use of the system; comprehensively evaluate system performance based on objective indicators and subjective evaluations; based on the verification results, tune system parameters and algorithms to improve system performance.
[0127] In an optional implementation, simulated scenario testing can employ a multi-level progressive verification approach. The system first undergoes basic testing in a fully controlled laboratory environment to verify various acoustic parameters and information transmission capabilities. Next, realistic ambient noise and multi-person collaboration scenarios are added to a simulated operating room to test the system's performance under interference conditions. Finally, comprehensive testing is conducted in a highly realistic surgical simulation environment to evaluate the system's overall performance under conditions close to actual surgery.
[0128] Step 6.4, System Integration and Deployment: Integrate the auditory processing module with the surgical robot system to ensure smooth information flow; support remote doctors to participate in surgical guidance and collaboration through the auditory interface; ensure that the system complies with medical device safety standards and data protection requirements; establish a system update and maintenance mechanism to ensure continuous optimization of system performance.
[0129] In an optional embodiment, the hardware and software integration utilizes a modular interface design. The system defines a standardized audio information exchange protocol, enabling compatible connections with surgical robot systems from different manufacturers. The integrated architecture comprises three layers: a hardware adaptation layer, a core processing layer, and an application interaction layer. The hardware adaptation layer is responsible for communicating with various audio input / output devices; the core processing layer includes the present invention's sound processing and auditory interaction algorithms; and the application interaction layer exchanges information with the surgical robot's control system via a standard API.
[0130] According to the embodiments of the present application, by converting surgical navigation data into personalized spatial audio information and enhancing the auditory salience of key information, the following technical effects are achieved:
[0131] First, this method enables doctors to "see" navigation information through hearing while focusing on the surgical area. Test results show that using this method reduces the number of times doctors shift their gaze during precision surgery, improves navigation accuracy, shortens surgery time, and reduces doctors' self-reported cognitive load scores.
[0132] Secondly, through personalized HRTF adjustment, this method solves the problem of different doctors' different 3D audio perception abilities, enabling doctors with weaker auditory spatial perception abilities to accurately perceive spatial navigation information, thereby improving the accuracy of spatial position perception and the universality of the navigation system.
[0133] In addition, through auditory saliency analysis and dynamic information priority assessment, this method improves the perception speed of high-priority information, can break through the attention barrier in emergency situations, attract the doctor's immediate attention, and reduce the risk of misoperation.
[0134] It can be seen that the adaptive design of this method enables the system to continuously optimize according to the doctor's auditory characteristics, workload and environmental changes, and the effect does not decrease with long-term use, which effectively solves the problem of reduced effect of 3D audio navigation systems in the existing technology after a period of use.
[0135] Real-world application examples of this implementation:
[0136] To verify the practical effect of the present invention, we conducted a series of clinical application tests in the neurosurgery department of a tertiary hospital. The following is an application example of this embodiment in deep brain tumor resection surgery:
[0137] Application scenarios:
[0138] This example uses deep brain tumor resection surgery as the application scenario. This type of surgery requires high precision, high risk, and a high degree of visual focus on the surgical area, as well as the need for real-time monitoring of multiple navigational information. In traditional surgeries, surgeons frequently need to look up at the navigation screen to determine the relative position of tools and key structures (such as the motor cortex, language areas, and vascular bundles), significantly increasing surgical time and the surgeon's cognitive burden.
[0139] This example involves five neurosurgeons with different auditory spatial perception abilities, who applied the auditory dialogue system of the present invention in 15 deep brain tumor resection surgeries and compared the results with a traditional visual navigation system.
[0140] Implementation process example:
[0141] Surgical navigation data voice:
[0142] In this example, the system fuses preoperative MRI and DTI imaging data with real-time surgical navigation data to generate four key types of navigation information:
[0143] Positional information: The relative position of the surgical instrument and the target lesion is mapped to spatial audio in different directions. Distance is mapped to pitch, with the pitch increasing as the target is approached.
[0144] Boundary information: The proximity of the device to the dangerous area (such as the motor cortex), mapped to the pulse rhythm, which accelerates when approaching the dangerous boundary;
[0145] Structural information: The density of nerve fiber bundles around the instrument is mapped to a specific timbre and duration. The higher the density, the brighter the timbre.
[0146] Operation prompts: For example, the optimal entry angle deviation prompt is mapped to a short tone sequence. The greater the deviation, the more obvious the tone change.
[0147] Physician hearing characteristics measurement:
[0148] Before implementation, the system measured the auditory characteristics of each physician. For example, a 42-year-old chief neurosurgeon demonstrated significantly lower spatial localization accuracy within a range of 120°-150° to the rear, with an average deviation of 28°. However, localization accuracy for sound sources within a ±30° range in front of the physician was very high, with an average deviation of only 5°. Furthermore, the physician's distance perception of near-field sound sources was relatively accurate (error <15%), but his perception of distance changes to far-field sound sources was less accurate (error >40%).
[0149] Based on the test results, the system built a personalized hearing characteristics model for the doctor:
[0150] ;
[0151] in, represents the doctor’s personal auditory characteristic model, Indicates that the doctor is at a horizontal angle and vertical angle The spatial auditory sensitivity distribution on The first element of the collection, represents the distance perception function, which indicates the doctor's perception accuracy of sound sources at different distances. The second element of the collection, It represents the dynamic sound source tracking capability parameter, which indicates the doctor's ability to track the moving sound source and the reaction time. The third element of the collection. Indicates the horizontal angle in degrees; Indicates the vertical angle in degrees; Indicates distance, indicating the spatial distance between the sound source and the listener, in meters or centimeters, with subscript Represents an individual (personal).
[0152] Spatial audio personalization:
[0153] For the above doctor, the system is customized as follows based on his / her hearing characteristics:
[0154] HRTF parameters were selected and adjusted to enhance directional cues within the 120°-150° range behind the patient. By increasing the binaural intensity difference and enhancing the characteristic frequency in this area, the doctor's directional perception deviation in this area was reduced from 28° to less than 10°.
[0155] Compensation processing has been performed on the distance perception of far-field sound sources, and the spectrum and reverberation characteristics of far-field sound sources have been adjusted, reducing the far-field distance perception error from 40% to within 25%;
[0156] Dynamic parameter adaptation was achieved. After the operation lasted for more than 2 hours, the system detected that the doctor's auditory reaction time was extended by about 15%, automatically enhancing the salience features of key information and maintaining the doctor's perception efficiency.
[0157] Auditory saliency analysis and information priority processing:
[0158] In an actual surgical case, when the surgical instrument approached a vital blood vessel (3 mm away), the system performed the following processing:
[0159] Through auditory saliency analysis, it is determined that this is urgent information that requires the doctor's immediate attention;
[0160] Calculate priority scores in real time (Emergency level, full score 1.0), much higher than other navigation information presented at the same time (Regular level, );
[0161] The system applied a combination of auditory salience features, including rapid pulses, high pitch, and unique timbre, while localizing the sound source 15 degrees in front of the physician (the physician's most sensitive direction).
[0162] The doctor immediately perceived the warning message and responded within 250ms, pausing the current operation and slightly adjusting the surgical path, successfully avoiding the blood vessel.
[0163] Feedback interaction and system optimization:
[0164] During use, the system learns the doctor's personalized preferences through his or her interactive behavior:
[0165] The doctor frequently used voice commands to lower the volume of background information, and the system automatically learned and lowered the baseline volume of non-urgent information during subsequent surgeries;
[0166] Through feedback data and performance analysis from 15 surgeries, the system optimized the information priority algorithm parameters for deep brain tumor surgery scenarios, especially increasing the risk assessment weight for crossing areas with dense white matter fiber bundles.
[0167] Technical effect verification:
[0168] Based on data from 15 deep brain tumor resection surgeries, the auditory dialogue system of this invention achieved significant technical results compared to traditional visual navigation systems:
[0169] Table 1: Comparison of doctors’ gaze shift and cognitive load data
[0170]
[0171] Table 2: Comparison of spatial navigation accuracy and information transmission efficiency data
[0172]
[0173] The data in Tables 1 and 2 indicate that the auditory dialogue system of the present invention achieves significant results in reducing visual acuity, lowering cognitive load, improving surgical efficiency, enhancing navigation accuracy, and accelerating information transmission. In particular, it shortens the physician's response time to key information from 1250ms to 410ms, improving surgical safety. Furthermore, through personalized auditory characteristic adjustment, the physician's position perception accuracy in a spatial audio environment is increased by 34.3%, validating the technical effectiveness of the present invention.
[0174] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.
Claims
1. A surgical robot dialogue method based on audio processing technology, characterized in that: The following steps are involved: Convert surgical navigation data into sound parameters and create a sound coding system that expresses different semantic information; Based on the sound coding system, the doctor's auditory spatial perception ability is measured and an individualized auditory characteristic model is constructed; Based on the doctor's auditory characteristics model, spatial audio parameters are adjusted to ensure that surgical navigation information is delivered in the optimal spatial audio format; Based on the adjusted spatial audio parameters, a computational model of human auditory saliency is constructed to quantify the effect of different acoustic features on guiding auditory attention. Based on the human auditory saliency calculation model, the importance of different navigation information is calculated in real time according to the surgical stage, risk level and time urgency, and a dynamic priority value is assigned to each piece of information; Based on the dynamic priority value, the auditory saliency feature enhancement processing most suitable for the current environment is applied to high-priority information, making it more prominent in the spatial audio.
2. The surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that: The step of converting the surgical navigation data into sound parameters comprises: The surgical navigation data are divided into four types according to type and function: location type, boundary type, structure type and operation prompt type; Establishing a mapping relationship between navigation data and sound parameters, wherein position information is mapped to spatial position and pitch, boundary information is mapped to timbre and rhythm, structure information is mapped to timbre and duration, and operation prompts are mapped to specific tone sequences; Build standardized sound coding templates for different types of navigation information to ensure that different types of information have differentiated but easily recognizable sound characteristics; According to the importance and urgency of the navigation information, the sound coding parameters are dynamically adjusted to make important information more prominent in the sound expression.
3. The surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that: The step of measuring the doctor's auditory spatial perception ability comprises: By playing standard test sounds in different directions, the doctor is asked to indicate the perceived direction of the sound and record the deviation between the actual direction and the perceived direction; By playing sounds at different virtual distances, the accuracy of doctors' perception of changes in sound distance is measured; Play a sound source moving in a virtual space and measure the doctor's ability to track the moving sound source and reaction time; Based on measuring the accuracy of doctors' perception of changes in sound distance and their ability and reaction time to track moving sound sources, a mathematical model describing the characteristics of doctors' auditory spatial perception is established.
4. The surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that: The step of adjusting the spatial audio parameters comprises: Selecting or synthesizing the most suitable head-related transfer function for the doctor from a head-related transfer function database according to the doctor's auditory spatial perception characteristics; Targeted compensation for imbalances or defects in the doctor's hearing characteristics; According to the spatial orientation perception deviation shown by the doctor in the measurement, the spatial localization parameters of the sound are adjusted to correct the perception deviation; By adjusting the reverberation ratio and near-field effect acoustic parameters of the sound, the doctor's accurate perception of the distance to the sound source is enhanced; During use, the doctor's response to spatial audio is continuously monitored, and parameters are dynamically adjusted to adapt to changes in the doctor's auditory adaptation during long operations.
5. The surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that: The steps of constructing a human auditory saliency computational model include: Classify various types of information during surgery according to their importance and urgency, and assign different auditory salience processing strategies; Generate corresponding auditory saliency parameters for information of different importance levels, including pitch, timbre, rhythm, and spatial position; When multiple sound sources exist simultaneously, the separation between target information and background information is enhanced, and the auditory masking effect is reduced; Designing long-term time series with clear auditory patterns enables doctors to perceive surgical progress and key stages through changes in auditory patterns; Based on the doctor's auditory preferences and cognitive habits, personalized auditory symbols are designed to establish intuitive auditory information mapping.
6. The surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that: The step of calculating the importance of different navigation information in real time includes: Based on surgical process knowledge and risk management principles, a model for evaluating the importance of navigation information was constructed. The model considered multidimensional characteristics such as surgical stage, patient status, information type, and time urgency. Based on surgical progress data and environmental information, the system can identify the current stage of the surgery in real time and adjust the basic priority of different types of information accordingly; Based on the navigation information content and the current surgical status, the correlation between navigation information and safety risks is evaluated, and high-risk related information is given higher priority; The final priority score of each navigation information is calculated based on the surgical stage, patient status, information type, time urgency, basic priority and safety risk correlation factors, and divided into four levels: emergency, important, routine and background.
7. A surgical robot dialogue system based on audio processing technology, characterized in that: A surgical robot dialogue method based on audio processing technology for executing any one of claims 1 to 6, comprising: The surgical navigation data sound conversion module is used to convert surgical navigation data into sound parameters and create a sound coding system that expresses different semantic information; Doctor's auditory characteristics measurement module, used to measure doctors' auditory spatial perception ability and build an individualized auditory characteristics model; A spatial audio customization module is used to adjust spatial audio parameters based on the doctor's auditory characteristics model to ensure that surgical navigation information is delivered in the optimal spatial audio format; Auditory saliency analysis module, used to build a computational model of human auditory saliency and quantify the effect of different acoustic features on guiding auditory attention; The information priority dynamic assessment module is used to calculate the importance of different navigation information in real time according to the surgical stage, risk level and time urgency, and assign a dynamic priority value to each piece of information; The key information highlighting module is used to apply the auditory saliency feature enhancement processing that is most suitable for the current environment to high-priority information, making it more prominent in spatial audio.
Citation Information
Patent Citations
Medical intercom system and method based on audio processing
CN118972716A
Systems and methods for techniques to process, analyze and model interactive verbal data for multiple individuals
US20230320642A1