Surgical robot dialogue method and system based on audio processing technology
By converting surgical navigation data into personalized spatial audio and adjusting auditory saliency, the problem of doctors' untimely visual burden and information perception in precision surgery is solved, and efficient and accurate navigation information transmission and risk reduction are achieved.
Patent Information
- Application Number
- CN202510764792.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Doctors need to pay attention to the surgical area and navigation information at the same time during precision surgery. Traditional navigation systems rely on visual presentation to cause excessive visual burden, and the critical information is not perceived in time. Existing audio prompts lack spatial positioning capabilities and personalization, making it difficult to accurately perceive navigation information in complex environments.
Convert surgical navigation data into personalized spatial audio, build an individualized model by measuring the doctor's auditory spatial perception ability, adjust audio parameters to enhance the auditory saliency of key information, and calculate the importance and priority of the information in real time, and dynamically adjust auditory characteristics to highlight high-priority information.
It reduces the doctor's visual transfer in precision surgery, improves the perceived accuracy and speed of navigation information, reduces the risk of misoperation, adapts to auditory differences and environmental changes of different doctors, and has stable long-term results.
Smart Images

Figure CN120267402A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical device control, and more specifically, it relates to a surgical robot dialogue method and system based on audio processing technology. Background Art
[0002] During a precision operation, the doctor needs to simultaneously focus on the surgical area and navigation information, while receiving and processing various audio information from multiple directions. Traditional navigation systems mainly rely on visual presentation, forcing the doctor to frequently shift their line of sight, which not only affects the surgical efficiency but also may increase the surgical risk. In addition, the audio information in the surgical environment is complex and diverse, and the doctor's auditory attention resources are limited, making it difficult to accurately perceive and distinguish the key information in the spatial audio while concentrating on the operation.
[0003] In the prior art, some systems attempt to use simple sound cues to assist navigation, but these systems generally have the following problems: First, traditional sound cues lack the ability of spatial localization and cannot effectively convey position-related navigation information; Second, the standardized audio parameters cannot adapt to the auditory perception differences of different doctors, resulting in some doctors having difficulty accurately perceiving the navigation information; In addition, in a complex surgical acoustic environment, the key navigation information is easily masked by other audio information or ignored by the doctor's attention filtering mechanism. Especially in an emergency, important information may not be noticed by the doctor in time.
[0004] Therefore, a technical solution that can convert surgical navigation information into personalized spatial audio and can dynamically adjust the auditory saliency according to the importance of the information is needed to solve the technical problems of the doctor having an excessive visual burden and untimely perception of key information during a precision operation. Summary of the Invention
[0005] The present invention provides a surgical robot dialogue method and system based on audio processing technology, which solves the technical problems of the doctor having an excessive visual burden and untimely perception of key information in the prior art during a precision operation.
[0006] The first aspect of the present invention provides a surgical robot dialogue method based on audio processing technology, including the following steps: Convert surgical navigation data into sound parameters and create a sound coding system expressing different semantic information; Based on the sound coding system, measure the doctor's auditory spatial perception ability and construct an individualized auditory characteristic model; Based on the doctor's auditory characteristic model, adjust the spatial audio parameters to ensure that the surgical navigation information is transmitted in the best spatial audio form; Based on the adjusted spatial audio parameters, construct a human auditory saliency calculation model to quantify the guiding effect of different acoustic features on auditory attention; Based on the human auditory saliency calculation model, calculate the importance of different navigation information in real time according to the surgical stage, risk level, and time urgency, and assign dynamic priority values to each piece of information; Based on the dynamic priority values, apply the auditory saliency feature enhancement processing that best suits the current environment to the high-priority information to make it more prominent in spatial audio.
[0007] Further, the step of converting surgical navigation data into sound parameters includes: Classify the surgical navigation data into four types: position type, boundary type, structure type, and operation prompt type according to the type and function; Establish a mapping relationship between the navigation data and the sound parameters, where the position information is mapped to the spatial position and pitch, the boundary information is mapped to the timbre and rhythm, the structure information is mapped to the timbre and duration, and the operation prompt is mapped to a specific tone sequence; For different types of navigation information, construct a standardized sound coding template to ensure that different types of information have differentiated but easily recognizable sound characteristics; Dynamically adjust the sound coding parameters according to the importance and urgency of the navigation information to make the important information more prominent in the sound expression.
[0008] Further, the step of measuring the doctor's auditory spatial perception ability includes: By playing standard test sounds from different directions, ask the doctor to indicate the perceived sound direction and record the deviation between the actual direction and the perceived direction; By playing sounds with different virtual distances, measure the doctor's perception accuracy of the sound distance change; Play a sound source moving in the virtual space and measure the doctor's ability to track the moving sound source and the reaction time; Based on the measured doctor's perception accuracy of the sound distance change and the doctor's ability to track the moving sound source and the reaction time, establish a mathematical model describing the doctor's auditory spatial perception characteristics.
[0009] Further, the step of adjusting the spatial audio parameters includes: According to the doctor's auditory spatial perception characteristics, select or synthesize the head-related transfer function that best suits the doctor from the head-related transfer function database; For the imbalance or defect in the doctor's auditory characteristics, conduct targeted compensation; According to the spatial orientation perception deviation shown by the doctor in the measurement, adjust the spatial positioning parameters of the sound to correct the perception deviation; By adjusting the reverberation ratio and near-field effect acoustic parameters of the sound, enhance the doctor's accurate perception of the sound source distance; During use, continuously monitor the doctor's response to spatial audio, and dynamically adjust parameters to adapt to the doctor's auditory adaptation changes during long surgeries.
[0010] Further, the steps of constructing the human auditory saliency calculation model include: Classify various types of information during surgery according to importance and urgency, and assign different auditory saliency processing strategies; Generate corresponding auditory saliency parameters for information at different importance levels, including pitch, timbre, rhythm, and spatial position; When multiple sound sources exist simultaneously, enhance the separation between target information and background information, and reduce the auditory masking effect; Design a long-range time series with a clear auditory pattern so that doctors can perceive the surgical process and key stages through changes in the auditory pattern; According to the doctor's auditory preferences and cognitive habits, design personalized auditory symbols and establish an intuitive auditory information mapping.
[0011] Further, the steps of calculating the importance degree of different navigation information in real time include: Based on surgical process knowledge and risk management principles, construct a model for evaluating the importance degree of navigation information, and the model considers multi-dimensional features such as surgical stage, patient status, information type, and time urgency; According to the surgical process data and environmental information, real-time identify which stage the current surgery is in, and accordingly adjust the basic priority of different types of information; According to the content of the navigation information and the current surgical state, evaluate the correlation degree between the navigation information and safety risks, and information related to high risks obtains a higher priority; Comprehensively consider factors such as surgical stage, patient status, information type, time urgency, basic priority, and safety risk correlation degree, calculate the final priority score of each piece of navigation information, and divide it into four levels: emergency level, important level, regular level, and background level.
[0012] Further, it also includes the steps of an information transmission mechanism: Analyze the type, complexity, and urgency characteristics of medical information to determine a suitable information transmission mechanism; Real-time evaluate the doctor's cognitive load status to avoid information overload; Convert complex medical information into an intuitively understandable auditory pattern to reduce the cognitive cost of decoding; When doctors need to pay attention to specific information, adopt a progressive attention guidance strategy to avoid abrupt interruption; Use multiple auditory feature encodings for key information to improve the robustness of information transmission.
[0013] Further, it also includes the steps of feedback interaction design: By recording the doctor's interaction behaviors, learning their personalized auditory preferences, including volume, pitch, and information density; Enable the doctor to perform real-time control of the auditory information system through short voice commands; Provide touchscreen or eye movement control alternatives for important controls to adapt to the needs of different surgical scenarios; Dynamically adjust the interaction complexity according to the doctor's familiarity with the system to reduce the learning cost; Regularly collect the doctor's feedback on the usage experience of the system and continuously improve the interaction design.
[0014] Furthermore, it also includes experimental verification and optimization steps: Test the system performance in a simulated surgical environment and evaluate the effectiveness of auditory information transmission; Collect usage feedback from doctors and the surgical team to understand the problems of the system in actual use; Comprehensively evaluate the system performance based on objective indicators and subjective evaluations; Based on the verification results, optimize the system parameters and algorithms to improve the system performance.
[0015] The second aspect of the present invention provides a surgical robot dialogue system based on audio processing technology for implementing the above-mentioned surgical robot dialogue method based on audio processing technology, including: A surgical navigation data sonification module for converting surgical navigation data into sound parameters and creating a sound coding system that expresses different semantic information; A doctor's auditory characteristic measurement module for measuring the doctor's auditory spatial perception ability and constructing an individualized auditory characteristic model; A spatial audio personalized customization module for adjusting spatial audio parameters based on the doctor's auditory characteristic model to ensure that surgical navigation information is transmitted in the best spatial audio form; An auditory saliency analysis module for constructing a human auditory saliency calculation model to quantify the guiding effect of different acoustic features on auditory attention; An information priority dynamic evaluation module for calculating the importance of different navigation information in real time according to the surgical stage, risk level, and time urgency, and assigning a dynamic priority value to each piece of information; A key information prominent presentation module for applying the most suitable auditory saliency features for the current environment to enhance the processing of high-priority information to make it more prominent in spatial audio.
[0016] The beneficial effects of the present invention are as follows: By converting surgical navigation data into personalized spatial audio information and enhancing the auditory salience of key information, the technical problem that doctors need to frequently shift their line of sight to view navigation information during precise surgeries is solved, achieving significant technical effects. Through personalized HRTF adjustment, the problem that different doctors have different abilities to perceive 3D audio is solved, enabling doctors with weaker auditory spatial perception abilities to accurately perceive spatial navigation information. Through auditory salience analysis and dynamic information priority assessment, the perception speed of high-priority information is improved, and in emergency situations, it can break through the attention barrier to immediately attract the doctor's attention, reducing the risk of misoperation. The adaptive design of this method enables the system to continuously optimize according to the doctor's auditory characteristics, workload, and environmental changes, and the long-term use effect does not decay, effectively solving the problem that the effect of 3D audio navigation systems in the prior art decreases after a period of use. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is the overall flowchart of the surgical robot dialogue method based on audio processing technology of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Now, the subject matter described herein will be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.
[0019] In at least one embodiment of the present invention, a surgical robot dialogue method based on audio processing technology is disclosed, as Figure 1 shown, including the following steps: Step 1: Convert surgical navigation data into sound parameters and create a sound coding system that expresses different semantic information; In this step, surgical navigation data (such as target position, safety boundary, key structure, etc.) is converted into sound parameters, and a sound coding system that expresses different semantic information is created. Specifically, it includes: Step 1.1: Navigation data classification: Classify surgical navigation data into four types: position type, boundary type, structure type, and operation prompt type according to type and function; Step 1.2: Sound parameter mapping: Establish a mapping relationship between navigation data and sound parameters, where position information is mapped to spatial position and pitch, boundary information is mapped to timbre and rhythm, structure information is mapped to timbre and duration, and operation prompts are mapped to specific tone sequences; Step 1.3, Semantic Sound Encoding Template Construction: For different types of navigation information, construct a standardized sound encoding template to ensure that different types of information have differentiated but easily recognizable sound characteristics; Step 1.4, Dynamic Adjustment of Encoding Parameters: Dynamically adjust the sound encoding parameters according to the importance and urgency of the navigation information, so that important information is more prominent in sound expression.
[0020] Step 2: Based on the sound encoding system, measure the doctor's auditory spatial perception ability and construct an individualized auditory characteristic model; It should be understood that in this step, through a rapid auditory spatial perception ability assessment program, measure the doctor's perception accuracy of sound sources in different directions, distances, and movements, and construct the individualized auditory characteristic model of this doctor. Specifically, it includes: Step 2.1, Measurement of Spatial Auditory Accuracy: By playing standard test sounds in different directions (horizontal and vertical), require the doctor to indicate the perceived sound direction, and record the deviation between the actual direction and the perceived direction; In some embodiments, the measurement of spatial auditory accuracy can be implemented by an interactive method based on a virtual reality headset. The doctor marks the perceived sound direction by turning the head or pointing with gestures in the virtual environment, and the system automatically records the angular deviation and generates a spatial perception accuracy heat map. Optionally, for higher accuracy requirements or in cases where virtual reality devices cannot be used, the system can use a surround speaker array to play test sounds, and the doctor indicates the sound direction through a dedicated controller.
[0021] Step 2.2, Evaluation of Distance Perception Ability: By playing sounds at different virtual distances, measure the doctor's perception accuracy of sound distance changes; For example, in one embodiment, the evaluation of distance perception ability can be implemented using a technology based on head near-field effect and far-field reverberation ratio adjustment. The system plays test sounds with different near-field - far-field characteristics in a fixed direction, and the doctor marks the perceived distance value. The system compares the actual distance and the perceived distance to construct an individualized distance perception mapping function.
[0022] Step 2.3, Dynamic Sound Source Tracking Test: Play a sound source moving in a virtual space, and measure the doctor's ability to track the moving sound source and the reaction time; Optionally, the dynamic sound source tracking test can be completed by simulating moving trajectories with different speeds and complexities in a virtual sound field. In a preferred embodiment, the system can generate sound source trajectories of various motion modes, including linear movement, curved movement, and random jump movement. The doctor completes the test by continuously tracking or predicting the end position, etc., and the system records the tracking accuracy and reaction delay data.
[0023] Step 2.4, Construction of individualized auditory characteristic model: Based on the above measurement data, a mathematical model describing the auditory spatial perception characteristics of the doctor is established, and this model can be expressed as: ; where, represents the doctor's personal auditory characteristic model, represents the spatial auditory sensitivity distribution of the doctor at the horizontal angle and the vertical angle , which is the first element of the set, represents the distance perception function, indicating the doctor's perception accuracy of sound sources at different distances, which is the second element of the set, represents the dynamic sound source tracking ability parameter, indicating the doctor's ability to track moving sound sources and reaction time, which is the third element of the set. represents the horizontal angle, with the unit of degree; represents the vertical angle, with the unit of degree; represents the distance, indicating the spatial distance between the sound source and the listener, with the unit of meter or centimeter, and the subscript represents personal.
[0024] In some embodiments, the individualized auditory characteristic model can be implemented using a Bayesian network structure, representing the multi-dimensional auditory spatial perception characteristics as a probabilistic graphical model, which can handle the correlations between different-dimensional perception capabilities. For example, the system may find that a specific doctor has particularly precise direction perception within a 60° range in the front, but relatively weak in the rear area. In this case, the system will perform more obvious directional enhancement processing on the rear information in subsequent spatial audio rendering.
[0025] Step 3: Based on the doctor's auditory characteristic model, adjust the spatial audio parameters to ensure that the surgical navigation information is transmitted in the best spatial audio form; Specifically including: Step 3.1, Generation of personalized HRTF: According to the doctor's auditory spatial perception characteristics, select or synthesize the head-related transfer function most suitable for the doctor from the HRTF database; In some embodiments, the generation of personalized HRTF can be implemented using machine learning methods. For example, a neural network model can be used to predict the most matching HRTF parameters from the doctor's auditory characteristic measurement results. Optionally, the system can also combine the geometric characteristics of the doctor's head and auricle (obtained through three-dimensional scanning) and use physical acoustic simulation methods to construct a more accurate personalized HRTF.
[0026] Step 3.2, Acoustic Compensation Processing: Targeted compensation is carried out for the imbalances or deficiencies in the doctor's auditory characteristics, such as the sensitivity difference between the left and right ears, the hearing loss at specific frequencies, etc. For example, in one embodiment, for a doctor with a sensitivity difference between the left and right ears, the system can achieve compensation through frequency-selective gain adjustment. Optionally, for a doctor with high-frequency hearing loss, the system can adopt a spectral remapping technique to convert high-frequency information to the mid-frequency region while maintaining the spatial localization characteristics of the sound.
[0027] Step 3.3, Spatial Localization Parameter Adjustment: According to the spatial orientation perception deviation shown by the doctor during the measurement, the system adjusts the spatial localization parameters of the sound to correct the perception deviation. In some embodiments, the spatial localization parameter adjustment can be achieved by adjusting the interaural time difference, intensity difference, and spectral cues. Optionally, the system can adopt an adaptive algorithm to continuously optimize the spatial localization parameters according to the doctor's feedback during use. For example, if it is detected that the doctor always has a systematic deviation in localizing the sound source in a certain direction, the algorithm will automatically adjust the spatial audio parameters in that direction until the best perception effect is achieved.
[0028] Step 3.4, Distance Perception Enhancement: By adjusting acoustic parameters such as the reverberation ratio and near-field effect of the sound, the doctor's accurate perception of the distance of the sound source is enhanced. For example, in a preferred embodiment, the distance perception enhancement can be achieved by combining multiple acoustic cues. The system can create a more realistic distance perception by dynamically adjusting parameters such as the ratio of direct sound to reflected sound, low-frequency gain, and high-frequency attenuation. Optionally, for doctors with relatively weak distance perception ability, the system can additionally introduce non-auditory feedback (such as tactile feedback) as an aid.
[0029] Step 3.5, Dynamic Parameter Adaptation: During use, the system continuously monitors the doctor's response effect to the spatial audio and dynamically adjusts the parameters to adapt to the auditory adaptation changes of the doctor during long-term surgeries.
[0030] In some embodiments, the dynamic parameter adaptation can be achieved by constructing a doctor's auditory fatigue model. The system monitors the change in the doctor's response time to the spatial audio cues. When a response delay possibly caused by auditory fatigue is detected, the system automatically adjusts parameters such as the intensity and salience of the spatial audio. For example, in the later stage of a long-term surgery, the system may increase the salience difference of important information or reduce the presentation frequency of non-critical information to reduce the auditory cognitive burden.
[0031] Step 4: Based on the adjusted spatial audio parameters, construct a human auditory salience calculation model to quantify the guiding effect of different acoustic features on auditory attention. Specifically, it includes: Step 4.1, Information Priority Classification: Classify various types of information during the operation according to importance and urgency, and assign different auditory salience processing strategies; In some embodiments, information priority classification can be implemented using a multi-level classification model. For example, information can be divided into multiple levels such as critical information (e.g., the catheter touches sensitive tissue), key information (e.g., the catheter approaches the target position), and routine information (e.g., the catheter moves within a safe path). Optionally, the system can also dynamically adjust the information priority according to the surgical stage, and increase the priority of certain types of information during critical surgical stages.
[0032] Step 4.2, Auditory Salience Parameter Generation: Generate corresponding auditory salience parameters for information of different importance levels, including pitch, timbre, rhythm, spatial position, etc.; In an alternative embodiment, salience parameter generation can adopt a biologically inspired auditory salience model. This model simulates the selective attention mechanism of the human auditory system to sound features, and calculates a salience map using features such as audio spectrum contrast and temporal mutability. Optionally, the system can use machine learning methods to optimize the salience parameter generation algorithm by analyzing doctors' attention responses to different combinations of sound features.
[0033] Step 4.3, Auditory Stream Separation Enhancement: When multiple sound sources coexist, enhance the separation between target information and background information, and reduce the auditory masking effect; For example, in one embodiment, auditory stream separation can be achieved by adjusting the harmonic structure and temporal envelope of different sound sources. For high-priority information, the system will assign a unique harmonic structure and rhythm pattern, making it form an independent auditory stream auditorily. Optionally, the system can also utilize the principle of spatial separation to present information of different priorities at different positions in the auditory space, further enhancing the separation effect.
[0034] Step 4.4, Long-Term Temporal Pattern Design: Design a long-term time series with a clear auditory pattern, enabling doctors to perceive the surgical process and critical stages through changes in the auditory pattern; In some embodiments, the long-term temporal pattern can adopt a hierarchical rhythm structure design. For example, the system can use specific timbres and rhythms at the micro level to represent the immediate state, and represent the progress of the surgical stage through a musicalized melody process at the macro level. Optionally, the system can also dynamically adjust the musical tension according to the deviation degree between the preset surgical path and the actual operation, providing doctors with an intuitive perception of path compliance.
[0035] Step 4.5, Personalized Auditory Symbol Design: Design personalized auditory symbols according to doctors' auditory preferences and cognitive habits, and establish an intuitive auditory-information mapping.
[0036] For example, in one embodiment, the personalized auditory symbols can be designed based on the doctor's professional background and auditory habits. For doctors with a background in music training, the system can adopt more complex musical expressions; for doctors without a music background, more intuitive sound effects can be used. Optionally, the system can also enable doctors to participate in the selection and optimization process of auditory symbols through interactive learning to establish a stronger auditory-information connection.
[0037] Step 5: Based on the human auditory saliency calculation model, calculate the importance levels of different navigation information in real time according to the surgical stage, risk level, and time urgency, and assign dynamic priority values to each piece of information. Specifically, it includes: Step 5.1: Construction of the information priority evaluation model: Based on surgical process knowledge and risk management principles, construct a model for evaluating the importance of navigation information. The model considers multi-dimensional features such as the surgical stage, patient status, information type, and time urgency. In the embodiments of the present application, the information priority evaluation model is implemented by a hybrid method combining the Analytic Hierarchy Process (AHP) and Support Vector Regression (SVR). The Analytic Hierarchy Process is used to establish the hierarchical structure of information priorities, and the priority influencing factors are divided into surgical process categories (such as the current surgical stage, expected next operation), risk categories (such as the degree of deviation from the safe area, proximity to key tissues), patient status categories (such as changes in vital signs, tissue response), and time urgency categories (such as response time window, potential risk growth rate). Then, based on expert knowledge and historical cases, the Support Vector Regression algorithm is used to learn the non-linear mapping relationship between each factor and the final priority.
[0038] In an application example of a spinal surgery, when the system detects that the surgical instrument is approaching the nerve root, the model comprehensively considers factors such as whether it is the decompression stage or the fixation stage, the rate of change of the distance of the nerve root, and the change amplitude of the intraoperative nerve monitoring signal, automatically adjusts the priority of the warning information to the emergency level, and immediately attracts the doctor's attention through a specific sound feature enhancement method, enabling the doctor to timely adjust the operation path before nerve damage occurs, and successfully avoiding 92% of potential nerve root injury risk events in the test.
[0039] Step 5.2: Identification of the surgical stage: According to the surgical process data and environmental information, identify which stage the current surgery is in real time, and accordingly adjust the basic priorities of different types of information. Step 5.3: Evaluation of the risk level: According to the content of the navigation information and the current surgical state, evaluate the degree of association between this information and safety risks, and information with a high risk correlation obtains a higher priority. Step 5.4, Information Priority Calculation and Classification: Considering the above factors, calculate the final priority score for each navigation information and classify it into four levels: emergency level, important level, regular level, and background level. The priority calculation function can be expressed as: ; where, represents the priority score of information , with a value range of 0 to 1. represents the relevance score to the surgical stage, evaluating the degree of relevance of information to the current surgical stage. represents the risk relevance score, evaluating the degree of association of information with safety risks. represents the information practicality score, evaluating the practical value of information to the current surgical operation. represents the time urgency score, evaluating the degree of time urgency for information to be responded to. , , and represent the relevance weight coefficient of the surgical stage, the risk relevance weight coefficient, the information practicality weight coefficient, and the time urgency weight coefficient respectively. These weight coefficients are used to adjust the importance of each component in the final priority. represents the navigation information, which is the surgical navigation data item to be transmitted to the doctor.
[0040] Step 6: Based on the dynamic priority value, apply the most suitable auditory saliency feature enhancement processing for the current environment to the high-priority information to make it more prominent in spatial audio; Specifically include: Step 6.1, Information Transmission Mechanism: Analyze the characteristics of medical information such as type, complexity, and urgency, determine the suitable information transmission mechanism; real-time evaluate the doctor's cognitive load status to avoid information overload; convert complex medical information into an intuitively understandable auditory pattern to reduce the decoding cognitive cost; adopt a progressive attention guidance strategy when the doctor needs to focus on specific information to avoid abrupt interruption; use multiple auditory feature encodings for key information to improve the robustness of information transmission.
[0041] In an alternative embodiment, information feature analysis may employ a multi-dimensional feature evaluation model. This model classifies information from dimensions such as timeliness (urgent / non-urgent), complexity (simple / complex), spatial relevance (location-related / location-independent), certainty (certain / uncertain), etc. For example, for information such as "the catheter touches the blood vessel wall" which is urgent, simple, location-related, and certain, the system may choose a direct voice warning; while for information such as "abnormal tissue density" which is non-urgent, complex, location-related but uncertain, the system may choose a progressive voice prompt supplemented with a detailed explanation.
[0042] Step 6.2, Feedback Interaction Design: By recording the doctor's interaction behavior, learn their personalized auditory preferences, such as volume, pitch, information density, etc.; enable the doctor to perform real-time control of the auditory information system through short voice commands; provide alternative means such as touchscreen or eye movement control for important controls to adapt to the needs of different surgical scenarios; dynamically adjust the interaction complexity according to the doctor's familiarity with the system to reduce the learning cost; regularly collect the doctor's feedback on the usage experience of the system and continuously improve the interaction design.
[0043] In an alternative embodiment, personalized preference learning may adopt a hybrid model method. The system first establishes an initial preference model based on the user's explicit feedback (such as manually adjusting settings), and then continuously updates the model through implicit feedback (such as whether the volume is lowered during use, whether the information is requested to be repeated, etc.). For example, if the system detects that the doctor repeatedly requests to reduce the volume of secondary information during a specific surgical procedure, the system will automatically learn to reduce the prominence of secondary information in such procedures.
[0044] Step 6.3, Experimental Verification and Optimization: Test the system performance in a simulated surgical environment and evaluate the effectiveness of auditory information transmission; collect usage feedback from doctors and the surgical team to understand the problems of the system in actual use; comprehensively evaluate the system performance based on objective indicators and subjective evaluations; based on the verification results, optimize the system parameters and algorithms to improve the system performance.
[0045] In an alternative embodiment, the simulated scenario test may adopt a multi-level progressive verification method. The system first conducts basic tests in a fully controlled laboratory environment to verify various acoustic parameters and information transmission functions; then adds real environmental noise and multi-person collaboration scenarios in a simulated operating room to test the performance of the system under interference conditions; finally, conducts comprehensive tests in a high-fidelity surgical simulation environment to evaluate the overall performance of the system under conditions close to actual surgery.
[0046] Step 6.4, System Integration and Deployment: Integrate the auditory processing module with the surgical robot system to ensure smooth information flow; support remote doctors to participate in surgical guidance and collaboration through the auditory interface; ensure that the system complies with medical device safety standards and data protection requirements; establish a system update and maintenance mechanism to continuously optimize the system performance.
[0047] In an alternative embodiment, the software and hardware integration adopts a modular interface design. The system defines a standardized audio information exchange protocol to support compatible connection with surgical robot systems from different manufacturers. The integration architecture includes three layers: the hardware adaptation layer, the core processing layer, and the application interaction layer. The hardware adaptation layer is responsible for communicating with various audio input / output devices; the core processing layer contains the sound processing and auditory interaction algorithms of the present invention; the application interaction layer exchanges information with the control system of the surgical robot through standard APIs.
[0048] According to the embodiments of the present application, by converting surgical navigation data into personalized spatial audio information and achieving enhanced auditory saliency for key information, the following technical effects are achieved: First, this method enables doctors to "see" navigation information audibly while focusing on the surgical area. The test results show that after using this method, the number of eye movements of doctors during precise surgeries is reduced, the navigation accuracy is improved, the surgical time is shortened, and the self-reported cognitive load score of doctors is decreased.
[0049] Second, through personalized HRTF adjustment, this method solves the problem of differences in 3D audio perception ability among different doctors, enabling doctors with relatively weak auditory spatial perception ability to accurately perceive spatial navigation information, and improving the accuracy of spatial position perception and the universality of the navigation system.
[0050] In addition, through auditory saliency analysis and dynamic information priority evaluation, this method improves the perception speed of high-priority information, can break through the attention barrier and attract the immediate attention of doctors in emergency situations, and reduces the risk of misoperation.
[0051] It can be seen that the adaptive design of this method enables the system to continuously optimize according to the auditory characteristics, workload, and environmental changes of doctors, and the long-term use effect does not decay, effectively solving the problem that the effect of 3D audio navigation systems in the prior art decreases after being used for a period of time.
[0052] Real application examples of this embodiment: To verify the actual effect of the present invention, we conducted a series of clinical application tests in the neurosurgery department of a certain tertiary hospital. The following is an application example of this embodiment in the resection of deep brain tumors: Application scenario: This example selects deep brain tumor resection surgery as the application scenario. This type of surgery has the characteristics of high operation precision requirements, large risk factors, doctors' visual attention highly concentrated on the surgical area, and the need to pay attention to multiple navigation information in real time. In traditional surgery, doctors need to frequently look up at the navigation screen to obtain the relative position information of the tool and key structures (such as the motor cortex, language area, blood vessel bundle, etc.), which significantly increases the operation time and the cognitive burden of doctors.
[0053] This example involves 5 neurosurgeons with different auditory spatial perception abilities. They respectively applied the auditory dialogue system of the present invention in 15 deep brain tumor resection surgeries and compared it with the traditional visual navigation system.
[0054] Implementation process example: Sound visualization of surgical navigation data: In this example, the system fuses preoperative MRI and DTI imaging data with real-time surgical navigation data to generate four types of key navigation information: Position information: The relative position of the surgical instrument and the target lesion is mapped to spatial audio in different directions, and the distance is mapped to the pitch. The pitch increases when approaching the target. Boundary information: The proximity of the instrument to the dangerous area (such as the motor cortex) is mapped to a pulse rhythm, and the rhythm speeds up when approaching the dangerous boundary. Structure information: The density of nerve fiber bundles around the instrument is mapped to a specific timbre and sustained sound, and the higher the density, the brighter the timbre. Operation prompt type: Such as the deviation prompt of the best entry angle, which is mapped to a short tone sequence, and the more obvious the deviation, the more obvious the tone change in the sequence.
[0055] Measurement of doctors' auditory characteristics: Before application, the system measured the auditory characteristics of each doctor. For example, a 42-year-old chief neurosurgeon showed significantly lower spatial positioning accuracy in the range of 120° - 150° behind than in other directions during the test, with an average deviation reaching 28°; while the sound source positioning accuracy in the range of ±30° in front was very high, with an average deviation of only 5°. In addition, this doctor's distance perception of near-field sound sources is relatively accurate (error < 15%), but the perception accuracy of the distance change of far-field sound sources is relatively low (error > 40%).
[0056] Based on the test results, the system constructed a personalized auditory characteristic model for this doctor: ; Among them, represents the doctor's personal auditory characteristic model, represents the doctor's horizontal angle and vertical angle The spatial auditory sensitivity distribution on is the first element of the set, which represents the distance perception function, indicating the doctor's perception accuracy of sound sources at different distances, and is the second element of the set, which represents the dynamic sound source tracking ability parameter, indicating the doctor's ability to track moving sound sources and reaction time, and is the third element of the set. represents the horizontal angle, in degrees; represents the vertical angle, in degrees; represents the distance, which is the spatial distance between the sound source and the listener, in meters or centimeters, and the subscript represents personal.
[0057] Spatial audio personalization: For the above-mentioned doctor, the system has performed the following personalization according to his auditory characteristics: Selected and adjusted the HRTF parameters to enhance the directional cues in the range of 120° - 150° behind. By increasing the interaural intensity difference and characteristic frequency enhancement in this area, the doctor's directional perception deviation in this area was reduced from 28° to within 10°; Compensated the distance perception of far-field sound sources, adjusted the spectrum and reverberation characteristics of far-field sound sources, and reduced the far-field distance perception error from 40% to within 25%; Achieved dynamic parameter adaptation. After the operation lasted for more than 2 hours, the system detected that the doctor's auditory reaction time increased by about 15%, and automatically enhanced the salient features of key information, maintaining the doctor's perception efficiency.
[0058] Auditory saliency analysis and information priority processing: In an actual surgical case, when the surgical instrument approached an important blood vessel (at a distance of 3 mm), the system performed the following processing: Through auditory saliency analysis, it was determined that this was an emergency-level information that needed to immediately attract the doctor's attention; Calculated the priority score in real time (emergency level, full score 1.0), much higher than other navigation information presented simultaneously (conventional level, ); The system applied a combination of auditory saliency features including fast pulses, high pitches, and unique timbres, and at the same time located the sound source at the 15° position in front of the doctor (the most sensitive direction area of this doctor); The doctor immediately perceived the warning information, responded within 250 ms, paused the current operation and slightly adjusted the surgical path, successfully avoiding the blood vessel.
[0059] Feedback Interaction and System Optimization: During use, the system learned the doctor's personalized preferences through their interaction behavior: This doctor often uses voice commands to lower the volume of background information, and the system automatically learned and adjusted the baseline volume of non-urgent information during subsequent surgeries; Through feedback data and performance analysis in 15 surgeries, the system optimized the information priority algorithm parameters for the deep brain tumor surgery scenario, especially increasing the risk assessment weight for areas passing through dense white matter fiber bundles.
[0060] Verification of Technical Effects: Based on the data from 15 deep brain tumor resection surgeries, the auditory dialogue system of the present invention achieved significant technical effects compared to traditional visual navigation systems: Table 1: Comparison of Doctor's Gaze Shift and Cognitive Load Data
[0061] Table 2: Comparison of Spatial Navigation Accuracy and Information Transfer Efficiency Data
[0062] From the data in Table 1 and Table 2, it can be seen that the auditory dialogue system of the present invention achieved significant effects in reducing gaze shift, lowering cognitive load, improving surgical efficiency, enhancing navigation accuracy, and accelerating information transfer. In particular, it shortened the doctor's perception reaction time for key information from 1250 ms to 410 ms, improving surgical safety; at the same time, through personalized auditory characteristic adjustment, the doctor's position perception accuracy in the spatial audio environment was increased by 34.3%, verifying the technical effects of the present invention.
[0063] The above describes the embodiments of the present invention, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.
Claims
1. A surgical robot dialogue method based on audio processing technology, characterized in that, It includes the following steps: Convert surgical navigation data into sound parameters and create a sound coding system that expresses different semantic information; Based on the sound coding system, measure the doctor's auditory spatial perception ability and construct an individualized auditory characteristic model; Based on the doctor's auditory characteristic model, adjust the spatial audio parameters to ensure that the surgical navigation information is transmitted in the best spatial audio form; Based on the adjusted spatial audio parameters, construct a human auditory saliency calculation model to quantify the guiding effect of different acoustic features on auditory attention; Based on the human auditory saliency calculation model, calculate the importance degree of different navigation information in real time according to the surgical stage, risk level and time urgency, and assign dynamic priority values to each piece of information; Based on the dynamic priority values, apply the auditory saliency features that are most suitable for the current environment to the high-priority information to make it more prominent in the spatial audio.
2. The surgical robot dialogue method based on audio processing technology according to claim 1, wherein The step of converting surgical navigation data into sound parameters includes: Classify the surgical navigation data into four types: position type, boundary type, structure type and operation prompt type according to the type and function; Establish a mapping relationship between the navigation data and the sound parameters, where the position information is mapped to the spatial position and pitch, the boundary information is mapped to the timbre and rhythm, the structure information is mapped to the timbre and duration, and the operation prompt is mapped to a specific tone sequence; For different types of navigation information, construct a standardized sound coding template to ensure that different types of information have differentiated but easily recognizable sound features; Dynamically adjust the sound coding parameters according to the importance and urgency of the navigation information to make the important information more prominent in the sound expression.
3. The surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that The step of measuring the doctor's auditory spatial perception ability includes: By playing standard test sounds in different directions, ask the doctor to indicate the perceived sound direction and record the deviation between the actual direction and the perceived direction; By playing sounds at different virtual distances, measure the doctor's perception accuracy of the sound distance change; Play a sound source moving in the virtual space and measure the doctor's ability to track the moving sound source and the reaction time; Based on measuring the doctor's perception accuracy of the sound distance change and the doctor's ability to track the moving sound source and the reaction time, establish a mathematical model that describes the doctor's auditory spatial perception characteristics.
4. The surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that The step of adjusting the spatial audio parameters includes: According to the doctor's auditory spatial perception characteristics, select or synthesize the head-related transfer function that is most suitable for the doctor from the head-related transfer function database; For the imbalance or defect in the doctor's auditory characteristics, carry out targeted compensation; According to the spatial orientation perception deviation shown by the doctor in the measurement, adjust the spatial positioning parameters of the sound to correct the perception deviation; By adjusting the reverberation ratio and near-field effect acoustic parameters of the sound, enhance the doctor's accurate perception of the sound source distance; During use, continuously monitor the doctor's reaction effect to the spatial audio and dynamically adjust the parameters to adapt to the doctor's auditory adaptation changes during long-term surgery.
5. A surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that, The step of constructing the human auditory saliency calculation model includes: Classify various types of information in the surgery according to the importance and urgency, and assign different auditory saliency processing strategies; Generate corresponding auditory saliency parameters for information with different importance levels, including pitch, timbre, rhythm, and spatial position; When multiple sound sources exist simultaneously, enhance the separation between target information and background information, and reduce the auditory masking effect; Design a long-range time series with a clear auditory pattern, enabling doctors to perceive the surgical process and key stages through changes in the auditory pattern; Design personalized auditory symbols according to doctors' auditory preferences and cognitive habits, and establish an intuitive auditory information mapping.
6. The surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that, The step of calculating the importance degree of different navigation information in real time includes: Based on surgical process knowledge and risk management principles, construct a model for evaluating the importance degree of navigation information. The model considers multi-dimensional features such as surgical stage, patient status, information type, and time urgency; According to surgical process data and environmental information, identify the current surgical stage in real time, and correspondingly adjust the basic priority of different types of information; According to the content of navigation information and the current surgical status, evaluate the correlation between navigation information and safety risks. Information related to high risks is given a higher priority; Comprehensively considering factors such as surgical stage, patient status, information type, time urgency, basic priority, and safety risk correlation, calculate the final priority score for each piece of navigation information, and classify it into four levels: emergency level, important level, routine level, and background level.
7. A method for a surgical robot to have a conversation based on audio processing technology according to claim 1, characterized in that, It also includes the information transmission mechanism steps: Analyze the type, complexity, and urgency characteristics of medical information to determine a suitable information transmission mechanism; Evaluate the doctor's cognitive load status in real time to avoid information overload; Convert complex medical information into an intuitively understandable auditory pattern to reduce the cognitive cost of decoding; When doctors need to pay attention to specific information, adopt a progressive attention guidance strategy to avoid abrupt interruptions; Encode key information with multiple auditory features to improve the robustness of information transmission.
8. A surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that, It also includes the feedback interaction design steps: By recording doctors' interaction behaviors, learn their personalized auditory preferences, including volume, pitch, and information density; Enable doctors to perform real-time control of the auditory information system through short voice commands; Provide touch screen or eye movement control alternative ways for important controls to meet the requirements of different surgical scenarios; Dynamically adjust the interaction complexity according to doctors' familiarity with the system to reduce the learning cost; Regularly collect doctors' feedback on the usage experience of the system and continuously improve the interaction design.
9. The surgical robot dialogue method based on audio processing technology according to claim 1, characterized in that, It also includes the experimental verification and optimization steps: Test the system performance in a simulated surgical environment and evaluate the effectiveness of auditory information transmission; Collect usage feedback from doctors and the surgical team to understand the problems of the system in actual use; Comprehensively evaluate the system performance based on objective indicators and subjective evaluations; Based on the verification results, optimize the system parameters and algorithms to improve the system performance.
10. A surgical robot dialogue system based on audio processing technology, characterized in that, For performing a surgical robot dialogue method based on audio processing technology as described in any one of claims 1-9, including: A surgical navigation data sonification module for converting surgical navigation data into sound parameters and creating a sound coding system that expresses different semantic information; A doctor's auditory characteristic measurement module for measuring the doctor's auditory spatial perception ability and constructing an individualized auditory characteristic model; A spatial audio personalized customization module, which is used to adjust spatial audio parameters based on the doctor's auditory characteristic model to ensure that surgical navigation information is transmitted in the best spatial audio form; An auditory saliency analysis module, which is used to build a human auditory saliency calculation model to quantify the guiding effect of different acoustic features on auditory attention; An information priority dynamic evaluation module, which is used to calculate the importance of different navigation information in real time according to the surgical stage, risk level and time urgency, and assign a dynamic priority value to each piece of information; A key information prominent presentation module, which is used to apply the auditory saliency features most suitable for the current environment to enhance the processing of high-priority information, making it more prominent in spatial audio.
Citation Information
Patent Citations
Audio augmented reality cues focused on audible information
CN117083590A
Voice information processing method and device
CN117727300A
Adaptive method of binaural hearing system
CN118413794A
Voice control intelligent instrument transfer robot system for operation assistance and control method
CN118924433A
Medical intercom system and method based on audio processing
CN118972716A