Surgical robot and audio processing method thereof

By using machine learning models in remote surgical robot systems to identify instrument types and operation types and generate simulated audio signals, the problem of missing instrument sounds is solved and the safety and accuracy of surgery are improved.

CN120616768APending Publication Date: 2025-09-12SHENZHEN JINGFENG MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510813719.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-12

Smart Images

  • Figure CN120616768A_ABST
    Figure CN120616768A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a surgical robot and an audio processing method thereof, the surgical robot comprises a main console and a patient side surgical platform which can be controlled by the main console, and the patient side surgical platform comprises a surgical instrument used for executing a surgery. According to the audio processing method, the first instrument type of the first target surgical instrument in the surgical instruments is obtained, and the analog audio signal of the first target surgical instrument is obtained based on the first instrument type, so that the surgical robot can completely feed back instrument sound in a remote surgery, and the influence of instrument sound deficiency on the operation accuracy of a doctor is avoided; and the safety of the remote operation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical devices, and in particular to a surgical robot and an audio processing method thereof. Background Art

[0002] With the innovation and development of science and technology and medical technology, surgical robotics has gradually matured. For doctors, surgical robots offer advantages such as ease of operation and high precision. For patients, surgical operations performed with surgical robots offer advantages such as minimal trauma, minimal pain, and rapid recovery, and are widely accepted by both doctors and patients.

[0003] Existing surgical robots typically consist of a main control console and a patient operating platform. The main control console controls the patient operating platform to perform surgery. With the increasing maturity of high-speed communication network technologies with low latency and high bandwidth, such as 5G and the internet, remote surgical robotics, which combines the advantages of both surgical robotics and high-speed communication networks, has emerged. Remote surgical robotics allows doctors to perform ultra-remote surgery on patients in remote locations, for example, in different countries or provinces, thereby maximizing the benefits of spatial constraints on surgical implementation and effectively alleviating the uneven distribution of medical resources.

[0004] The main control console and patient surgical platform in a remote surgical robot system are usually deployed in different countries, different provinces and cities, or different jurisdictions of the same city. Before performing remote surgery, communication between the remote main control console and the patient surgical platform needs to be established. At the same time, during remote surgery, the remote surgical robot relies on multimodal feedback to assist the doctor's operation, such as visual feedback, auditory feedback, and tactile feedback. Among them, clear and complete auditory feedback can achieve clear transmission of instructions or information between the proximal and distal ends. However, existing audio solutions may lose some instrument sounds when acquiring instrument sounds, and the lack of complete auditory feedback may reduce the accuracy of the doctor's operation and bring hidden dangers to surgical safety. Summary of the Invention

[0005] Based on this, it is necessary to provide a surgical robot and an audio processing method thereof that can fully feedback the instrument sounds during remote surgery.

[0006] In a first aspect, the present application provides a surgical robot, comprising: Main console; a patient-side surgical platform controllable by the main console; The patient-side surgical platform includes surgical instruments for performing surgery, and the patient-side surgical platform or the main console is configured as follows: obtaining a first instrument type of a first target surgical instrument among the surgical instruments; An analog audio signal of the first target surgical instrument is obtained based on the first instrument type.

[0007] Furthermore, the patient-side surgical platform or the main console is further configured to: Acquire a first operation type of the first target surgical instrument; The step of obtaining the simulated audio signal of the first target surgical instrument based on the first instrument type includes: A simulated audio signal of the first target surgical instrument is obtained based on the first instrument type and the first operation type.

[0008] Furthermore, the patient-side surgical platform further includes an endoscope for acquiring images of the surgical environment; The acquiring of a first instrument type of a first target surgical instrument among the surgical instruments comprises: The first instrument type of the first target surgical instrument in the surgical environment image is identified based on a preset machine learning model.

[0009] Furthermore, the patient-side surgical platform further includes an endoscope for acquiring images of the surgical environment; The acquiring of the first operation type of the first target surgical instrument among the surgical instruments comprises: The first operation type of the first target surgical instrument in the surgical environment image is identified based on a preset machine learning model.

[0010] Further, the first target surgical instrument includes the surgical instrument in an operating state.

[0011] Furthermore, the patient-side surgical platform further comprises an audio system for acquiring instrument sounds of the surgical instruments, and the first target surgical instruments include the surgical instruments in an operating state whose sounds have not been fully acquired; Before acquiring the first instrument type of the first target surgical instrument, the patient-side surgical platform or the main console is further configured to: determining whether all sounds of the surgical instrument in the operating state are acquired based on the instrument audio signal acquired by the audio system; If not, determine the first target surgical instrument.

[0012] Furthermore, the determining whether all sounds of the surgical instrument in the operating state are acquired based on the instrument audio signal acquired by the audio system includes: Extracting voiceprint features of the audio signal of the instrument; calculating a matching degree between the voiceprint feature and a standard voiceprint feature of the surgical instrument in the operating state; If the matching degree is greater than a preset first matching degree threshold, all sounds of the surgical instrument in the operating state are acquired; otherwise, not all sounds of the surgical instrument in the operating state are acquired.

[0013] Furthermore, the patient-side surgical platform or the main console is further configured to: obtaining a second instrument type of the surgical instrument in the operating state; The standard voiceprint feature is obtained based on the second device type.

[0014] Furthermore, obtaining the standard voiceprint feature based on the second device type includes: Based on the second device type, obtaining the standard voiceprint feature from a preset standard database; or The second device type is input into a preset machine learning model to output the standard voiceprint feature.

[0015] Furthermore, the patient-side surgical platform or the main console is further configured to: acquiring a second instrument type and a second operation type of the surgical instrument in the operating state; The standard voiceprint feature is obtained based on the second instrument type and the second operation type.

[0016] Furthermore, obtaining the standard voiceprint feature based on the second device type and the second operation type includes: Based on the second device type and the second operation type, obtaining the standard voiceprint feature from a preset standard database; or The second device type and the second operation type are input into a preset machine learning model to output the standard voiceprint feature.

[0017] Furthermore, the determining whether all sounds of the surgical instrument in the operating state are acquired based on the instrument audio signal acquired by the audio system includes: detecting a vibration signal of the surgical instrument in the operating state; Performing coherence analysis on the vibration signal and the instrument audio signal to obtain a coherence coefficient; If the coherence coefficient is greater than a preset coherence coefficient threshold, all sounds of the surgical instrument in the operating state are acquired; otherwise, not all sounds of the surgical instrument in the operating state are acquired.

[0018] Furthermore, the determining whether all sounds of the surgical instrument in the operating state are acquired based on the instrument audio signal acquired by the audio system includes: Performing short-time Fourier transform on the instrument audio signal, and obtaining each independent audio signal in the instrument audio signal based on independent component analysis; If the number of the independent audio signals is less than the number of the surgical instruments in the operating state, then the sounds of the surgical instruments in the operating state are not all acquired; If the number of the independent audio signals is equal to the number of the surgical instruments in the operating state, the surgical instrument corresponding to each of the independent audio signals is identified based on a preset machine learning model, and when the matching degree between the independent voiceprint feature of each independent audio signal and the standard independent voiceprint feature of the corresponding surgical instrument is greater than a preset second matching degree threshold, it is confirmed that the sounds of the surgical instruments in the operating state are all acquired; otherwise, the sounds of the surgical instruments in the operating state are not all acquired.

[0019] Furthermore, determining the first target surgical instrument includes: determining, among the surgical instruments corresponding to the independent audio signal, the surgical instrument whose matching degree between the independent voiceprint feature and the standard independent voiceprint feature is greater than the second matching degree threshold as a second target surgical instrument; The first target surgical instrument is obtained based on the second target surgical instrument and the surgical instrument in the operating state.

[0020] Furthermore, determining the first target surgical instrument includes: The first target surgical instrument is determined based on the voiceprint feature and the standard voiceprint feature.

[0021] Furthermore, obtaining the analog audio signal of the first target surgical instrument based on the first instrument type includes: Based on the first instrument type, obtaining a standard audio signal of the first target surgical instrument from a preset standard database as the simulated audio signal; or The first instrument type is input into a preset neural network model, and the simulated audio signal is output.

[0022] Furthermore, obtaining the simulated audio signal of the first target surgical instrument based on the first instrument type and the first operation type includes: Based on the first instrument type and the first operation type, a standard audio signal of the first target surgical instrument is obtained from a preset standard database as the simulated audio signal or The first instrument type and the first operation type are input into a preset neural network model, and the simulated audio signal is output.

[0023] Furthermore, the neural network model is a generative adversarial network or a diffusion model In a second aspect, the present application further provides an audio processing method for a surgical robot, the method being applied to a patient-side surgical platform or a main console of the surgical robot, the surgical robot comprising the main console and the patient-side surgical platform controllable by the main console, the patient-side surgical platform comprising surgical instruments for performing surgery, the method comprising: obtaining a first instrument type of a first target surgical instrument among the surgical instruments; An analog audio signal of the first target surgical instrument is obtained based on the first instrument type.

[0024] In a third aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the instructions are executed on at least one processor, the method described in the second aspect is implemented.

[0025] The audio processing method of the remote surgical robot and the remote surgical robot of the present application have the following beneficial effects: The surgical robot and audio processing method of the present application obtain the first instrument type of the first target surgical instrument in the surgical instrument and obtain the simulated audio signal of the first target surgical instrument based on the first instrument type, so that the surgical robot can fully feedback the instrument sound during the operation, avoid the impact of the lack of instrument sound on the doctor's operation accuracy, and improve the safety of remote surgery. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a schematic structural diagram of a remote surgical robot according to an embodiment of the present application; Figure 2 This is a schematic structural diagram of a surgical instrument according to one embodiment of the present application; Figure 3 This is a schematic structural diagram of a slave robot of a remote surgical robot according to an embodiment of the present application; Figure 4 This is a schematic structural diagram of an interventional surgical robot according to one embodiment of the present application; Figure 5 A network topology diagram of a remote surgical robot according to an embodiment of the present application; Figure 6 This is a module diagram of an audio system according to an embodiment of the present application; Figure 7 This is a flowchart of an audio processing method for a surgical robot according to one embodiment of the present application; Figure 8 A flowchart of an audio processing method for a surgical robot according to an embodiment of the present application for determining whether all instrument sounds have been acquired; Figure 9 A flowchart of an audio processing method for a surgical robot according to another embodiment of the present application for determining whether all instrument sounds have been acquired; Figure 10 This is a flowchart of an audio processing method for a surgical robot according to another embodiment of the present application for determining whether all instrument sounds have been acquired. DETAILED DESCRIPTION

[0027] To facilitate understanding of the present application, a more comprehensive description of the present application will be provided below with reference to the accompanying drawings. The accompanying drawings illustrate preferred embodiments of the present application. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure of the present application.

[0028] It should be noted that when an element is referred to as being "disposed on" another element, it may be directly on the other element or there may also be a centered element. When an element is considered to be "connected" to another element, it may be directly connected to the other element or there may be a centered element at the same time. When an element is considered to be "coupled" to another element, it may be directly coupled to the other element or there may be a centered element at the same time. The terms "vertical", "horizontal", "left", "right", "above", "below" and similar expressions used herein are for illustrative purposes only and are not intended to be the only embodiment. It should be understood that these spatially related terms are intended to cover different orientations of the device in use or in operation in addition to the orientations depicted in the drawings. For example, if the device is flipped in the drawings, elements or features described as being "below" or "beneath" other elements or features will be oriented "above" other elements or features. Therefore, the example term "below" can include both above and below orientations.

[0029] The terms "distal end" and "proximal end" as used herein are directional terms commonly used in the field of interventional medical devices, where "distal end" refers to the end away from the operator during surgery, and "proximal end" refers to the end close to the operator during surgery. "Coupled" as used herein can be broadly understood as any event in which two or more objects are connected in a manner that allows the absolutely coupled objects to operate together, such that there is no relative movement between the objects in at least one direction, such as a coupling of a protrusion and a groove, which can move relative to each other in the radial direction but not in the axial direction.

[0030] The term "tool" is used herein to describe a medical device that is inserted into a patient's body and used to perform a surgical or diagnostic procedure, the tool comprising an end effector, which may be a surgical tool for performing a surgical procedure, such as an electrocautery device, a clamp, a stapler, a shear, an imaging device (such as an endoscope or ultrasound probe), and the like. Some tools used in embodiments of the present application further include providing an articulated component (such as a joint assembly) for the end effector so that the position and orientation of the end effector can be manipulated and moved with one or more mechanical degrees of freedom relative to the instrument axis. Furthermore, the end effector includes functional mechanical degrees of freedom, such as opening and closing the clamp. The tool may also include stored information that can be updated by the surgical system, whereby the storage system can provide one-way or two-way communication between the tool and one or more system components. Some tools used in some embodiments may also not include providing an articulated component for the end effector.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "and / or" and "and / or" as used herein include any and all combinations of one or more of the associated listed items.

[0032] The remote surgical robot of the embodiment of the present application includes a master-slave surgical robot suitable for performing ultra-remote surgery. The master-slave surgical robot includes a remote doctor's main console and a patient-side surgical platform. The patient-side surgical platform can include different types of slave robots, including but not limited to single-port laparoscopic slave robots, multi-port laparoscopic slave robots, bronchial interventional slave robots, vascular interventional slave robots, and orthopedic slave robots. Different types of slave robots have different structural characteristics and may be suitable for the same or different types of surgeries.

[0033] For example, in Figure 1 The master-slave surgical robot shown includes a remote doctor's main console 100 and a patient-side surgical platform 200, which includes a single-port laparoscopic slave robot. The remote doctor's main console 100 can send control commands to the patient-side surgical platform 200 based on the doctor's operations to control the patient-side surgical platform 200. The remote doctor's main console 100 is also used to display images captured by the patient-side surgical platform 200. The patient-side surgical platform 200 responds to control commands sent by the remote doctor's main console 100, performs corresponding operations, and captures images of the patient's body, such as images of the surgical environment within the patient's body.

[0034] The patient-side surgical platform 200 includes a robotic arm 210 and a drive assembly disposed on the robotic arm 210. The drive assembly includes a driver 220 disposed on the robotic arm 210 and a driver 220 disposed on the driver 220. Figure 2 The surgical tool 230 is shown. The patient-side surgical platform 200 also includes a puncture device 240 mounted on the long axis 231 of the surgical tool 230. When the patient-side surgical platform 200 responds to control commands from the remote doctor's main console 100, the robotic arm 210 is used to adjust the position of the surgical tool 230, the driver 220 is used to drive the surgical tool 230 to perform the corresponding operation, and the end effector 232 of the surgical tool 230 is used to extend into the patient's body and perform surgical operations and / or obtain images of the patient's body through its distal end instrument.

[0035] Figure 3 Another patient-side surgical platform 200' is shown. This patient-side surgical platform 200' includes a multi-port laparoscope slave robot. The patient-side surgical platform 200' includes a robotic arm and multiple manipulator assemblies 230' mounted on the robotic arm. The robotic arm includes a main arm 210' and multiple adjustment arms 220' mounted on an orientation platform 215' within the main arm 210'. Different adjustment arms 220' are equipped with different manipulator assemblies 230'. The main arm 210' can adjust the position of the adjustment arms 220' and the manipulator assemblies 230', and the adjustment arms 220' can adjust the position of the manipulator assemblies 230'. The manipulator assembly 230' includes a gripper arm 240' and a medical device 250' removably mounted on the gripper arm 240'. The manipulator assembly 230' includes a parallelogram mechanism. Utilizing the parallelogram principle, the gripper arm 240' can be limited to rotational motion around a remote center of motion (RC). Multiple medical devices 250' can be inserted into the patient's body through different puncture devices 500'. It is understandable that Figure 1 The remote doctor main console 100 shown can also be used to operate Figure 3 Movement of manipulator assembly 230 ′ is shown in patient-side surgical platform 200 ′.

[0036] Figure 4An interventional surgical robot 300 is shown. The interventional surgical robot 300 is a natural cavity surgical robot, comprising a remote doctor main console and a patient-side surgical platform 320. The remote doctor main console comprises a handle 310' and an imaging cart 330 that are interconnected. The patient-side surgical platform 320 may also comprise an imaging cart 330. The patient-side surgical platform 320 is connected to a catheter instrument 340, a sensor system 350, and a control system 360 for achieving control among the catheter instrument 340, the sensor system 350, and the imaging cart 330. When the doctor performs various procedures on the patient next to the patient-side surgical platform 320, he can trigger control instructions by operating the handle 310' and transmit them to the patient-side surgical platform 320 for driving, thereby controlling the catheter instrument 340 to advance, retract, bend, and turn.

[0037] It is understood that the patient-side surgical platform 320 can generally be moved to the side of the surgical bed to engage the catheter instrument 340 and, under control commands, control the catheter instrument 340 to move vertically, horizontally, or in non-vertical and non-horizontal directions, thereby providing a better preoperative preparation angle for the operation of the catheter instrument 340. The control commands can be triggered by the doctor operating the patient-side surgical platform 320, or by the doctor directly clicking or pressing a button on the patient-side surgical platform 320. Of course, in other embodiments, the control commands can also be voice control or commands triggered by a force feedback mechanism.

[0038] like Figure 4 As shown, the patient-side surgical platform 320 may further include a base 321, a sliding base 322 that can be raised and lowered along the base 321, and two robotic arms 323 fixedly connected to the sliding base 322. The robotic arms 323 may include multiple arm segments connected at joints. The multiple arm segments provide the robotic arms 323 with multiple degrees of freedom, for example, seven degrees of freedom corresponding to the seven arm segments. A manipulator ( Figure 4 (Not shown) The manipulator of robotic arm 323 is used to engage catheter instrument 340 and, driven by the manipulator, control the distal end of catheter instrument 340 to bend and turn accordingly. The two robotic arms 323 can be identical in structure or partially identical in structure, with one robotic arm 323 engaging the inner catheter instrument 341 and the other engaging the outer catheter instrument 342. During installation, the outer catheter instrument 342 can be installed first. After the outer catheter instrument 342 is installed, the catheter of the inner catheter instrument 341 is inserted into the catheter of the outer catheter instrument 342.

[0039] The sensor system 350 has one or more subsystems for receiving information about the catheter device 340. The subsystems may include: a position sensor system; a shape sensor system for determining the position, orientation, speed, velocity, pose, and / or shape of the tip of the catheter device 340 and / or one or more segments along the catheter that may comprise the catheter device 340; and / or a visualization system for capturing images from the tip of the catheter device 340.

[0040] The imaging vehicle 330 can be provided with a display system 331 and a washing system ( Figure 4 Display system 331 is used to display images or representations of the surgical site and catheter device 340 generated by the subsystems of sensor system 350. It can also display real-time images of the surgical site and catheter device 340 captured by the visualization system. It can also use image data from imaging technologies to present images of the surgical site recorded before or during surgery. Imaging technologies include computed tomography (CT), magnetic resonance imaging (MRI), optical coherence tomography (OCT), and ultrasound.

[0041] The preoperative or intraoperative image data can be presented as a two-dimensional, three-dimensional, or four-dimensional image (e.g., time-based or rate-based information) and / or as an image from a model created based on the preoperative or intraoperative image dataset, or as a virtual navigation image. In the virtual navigation image, the actual position of the catheter device 340 is registered with the preoperative image to present a virtual image of the catheter device 340 within the surgical site to the operator from the outside.

[0042] Control system 360 includes at least one memory and at least one processor. It will be appreciated that control system 360 can be integrated into patient-side surgical platform 320 or imaging cart 330, or can be independently configured. Control system 360 can support wireless communication protocols such as IEEE 802.11, IrDA, Bluetooth, HomeRF, DECT, and wireless telemetry. Control system 360 can transmit one or more signals instructing the manipulator to move catheter device 340. Catheter device 340 can be extended to a surgical site within the body through a natural orifice or surgical incision in the patient.

[0043] Furthermore, the control system 360 may include a mechanical control system ( Figure 4 not shown) and an image processing system ( Figure 4 (not shown). The mechanical control system is used to control the movement of the catheter instrument 340 and can therefore be integrated into the patient surgical platform 320. The image processing system is used to plan the virtual navigation path and can therefore be integrated into the imaging cart 330. Of course, the various subsystems of the control system 360 are not limited to the specific ones listed above and can be reasonably configured according to actual circumstances.

[0044] Among them, the image processing system can use the above-mentioned imaging technology to image the surgical site based on the image of the surgical site recorded before or during the operation, and can also use software combined with manual input to convert the recorded image into a two-dimensional or three-dimensional synthetic image of part or the entire anatomical organ or segment. During virtual navigation, the sensor system 350 can be used to calculate the position of the catheter instrument 340 relative to the patient's anatomical structure, which can be used to generate an external tracking image and an internal virtual image of the patient's anatomical structure, so as to achieve the actual position of the catheter instrument 340 and the preoperative image. The virtual image of the catheter instrument 340 in the surgical site can be presented to the operator from the outside.

[0045] The internal catheter device 341 and the external catheter device 342 have substantially the same structural composition, and each comprises a slender and flexible inner catheter 41 and an outer catheter 42, wherein the diameter of the outer catheter 42 is slightly larger than that of the inner catheter 41, so that the inner catheter 41 can pass through the outer catheter 42 and provide a certain degree of support for the inner catheter 41, thereby enabling the inner catheter 41 to reach the target location in the patient's body, so as to facilitate operations such as tissue or cell sampling from the target location.

[0046] Certain movements of handle 310 can cause corresponding movements of catheter device 340. For example, when a doctor operates handle 310 to move the directional lever upward or downward, the movement of the directional lever can be mapped to a corresponding pitch movement of the distal end of catheter device 340. When a doctor operates handle 310 to move the directional lever left or right, the movement of the directional lever can be mapped to a corresponding yaw movement of the distal end of catheter device 340. In this embodiment, handle 310 can control the movement of the distal end of catheter device 340 within a 360-degree spatial range.

[0047] The remote surgical robot of this application also includes a server. Figure 5 As shown, remote doctor master consoles A1-A3 and patient-side surgical platforms B1-B3 of the same or different master-slave surgical robots are each connected to a server S for communication. Server S can be used to manage information sent by remote doctor master consoles A1-A3 and patient-side surgical platforms B1-B3 and can be used to relay information transmitted between remote doctor master consoles A1-A3 and patient-side surgical platforms B1-B3. Server S can also be used to achieve device interconnection, i.e., pairing, between remote doctor master consoles A1-A3 and patient-side surgical platforms B1-B3. Pairing involves connecting two devices to enable communication.

[0048] In ultra-teleoperative surgery, the remote doctor's main console and the patient-side surgical platform are deployed in different locations. For example, they can be deployed in different countries, different provinces and cities, or different jurisdictions in the same city.

[0049] The server may also be deployed at a different location than either the remote physician console or the patient-side surgical platform. For example, the server, the remote physician console, and the patient-side surgical platform may be deployed in different cities. The server may also be deployed at the same location as either the remote physician console or the patient-side surgical platform. For example, the server may be deployed independently of the remote physician console or the patient-side surgical platform located at the same location. For another example, the server may be integrated with the remote physician console or the patient-side surgical platform located at the same location. It is understood that the server may also be deployed in the cloud.

[0050] In the remote surgical robot of the present application, at least one remote doctor main console, patient-side surgical platform and server are deployed respectively.

[0051] There may be at least one remote doctor main console, at least one patient-side surgical platform, and at least one server.

[0052] When there is a single server, each remote physician main console and each patient-side surgical platform are connected to the single server. When there are multiple servers, the servers are connected via a network topology, and different remote physician main consoles and different patient-side surgical platforms can be connected to the same or different servers. The network topology between the multiple servers includes at least one of a star topology, a ring topology, a bus topology, a tree topology, a mesh topology, a virtual local area network, and a wireless topology.

[0053] The remote doctor's main console and the server, the patient's surgical platform and the server, and the servers themselves can be connected via the same or different types of high-speed communication networks. For example, these types of high-speed communication networks may include at least one of the following: broadband internet, a dedicated internet network, a 5G network, and a dedicated 5G network.

[0054] It is understood that audio transmission through the audio system is a crucial means of conveying information between the patient-side surgical platform and the remote doctor's main console during remote surgery. With the help of the audio system, the remote surgical robot can detect audio signals from the patient-side surgical platform and its vicinity, or from the remote doctor's main console and its vicinity, in real time, and reproduce the audio signals for personnel near the equipment on the other end. For example, the doctor operating the main console may communicate with the medical staff surrounding the patient to convey specific instructions or information. The audio system deployed on the patient-side surgical platform captures the instrument sounds during the surgical procedure and transmits them to the remote main console, allowing the doctor operating the main console to receive auditory feedback while remotely operating the patient-side surgical platform.

[0055] In a telesurgery robot, both the remote doctor's main console and the patient-side surgical platform will deploy at least one audio system. The audio system can be deployed independently of the remote doctor's main console or the patient-side surgical platform located at the same location, or it can be integrated into the remote doctor's main console or the patient-side surgical platform located at the same location. An audio system generally includes a sound acquisition module, a sound processing module, a sound playback module, or may also include a human-computer interaction module, such as Figure 6 shown.

[0056] The sound collection module is used to collect the sound in the audio environment where the near-end device is currently located, and transmit the collected sound to the sound processing module.

[0057] The sound processing module receives the sound from the sound acquisition module and sends it to the remote device after processing. In addition, the sound processing module also receives and processes the sound collected, processed, and emitted by the audio system deployed by the remote device through the server, and finally plays it through the sound playback module.

[0058] It is understood that the proximal device refers to the device where the audio system is deployed, that is, the device located at the same location as the audio system, and can be either the remote physician's main console or the patient-side surgical platform. The remote device refers to the device that is interconnected and communicates with the proximal device, and can be the other of the remote physician's main console and the patient-side surgical platform.

[0059] However, existing audio solutions may miss some instrument sounds when capturing them. This incomplete auditory feedback can reduce the accuracy of the surgeon's operation and pose a safety risk. For example, during remote surgery, the surgeon controls the patient's surgical platform using the main console and needs to hear real-time auditory feedback from the instruments they are operating (such as electrocoagulation scalpels and ultrasonic scalpels).

[0060] Based on this, the present application proposes a surgical robot and an audio processing method thereof. The surgical robot includes a main console and a patient-side surgical platform that can be controlled by the main console. The patient-side surgical platform includes surgical instruments for performing surgery. The audio processing method is applied to the patient-side surgical platform or the main console in the surgical robot, that is, the patient-side surgical platform or the main console in the surgical robot can execute the audio processing method described below to simulate the instrument sound of the surgical instrument. In an embodiment of the present application, the surgical robot may include a remote surgical robot.

[0061] In some embodiments, reference Figure 7 , an audio processing method provided by this application includes: S110: Obtain a first instrument type of a first target surgical instrument among the surgical instruments.

[0062] S120: Obtain an analog audio signal of a first target surgical instrument based on the first instrument type.

[0063] In step S110, the first target surgical instrument is the surgical instrument selected in the embodiment of the present application for which instrument sound simulation is required. It is understood that obtaining the first instrument type of the first target surgical instrument in the embodiment of the present application can provide a basis and conditions for subsequent simulation of the instrument sound of the first target surgical instrument, thereby improving the accuracy of the instrument sound simulation.

[0064] In some embodiments, the patient-side surgical platform further includes an endoscope for acquiring an image of the surgical environment. Step S110 may include: A first instrument type of a first target surgical instrument in a surgical environment image is identified based on a preset machine learning model.

[0065] Specifically, the patient-side surgical platform stores a pre-trained machine learning model. The pre-trained machine learning model is pre-trained and can identify the instrument type of the first target surgical instrument in the surgical environment image, that is, the first instrument type.

[0066] It is understood that when the main console executes the audio processing method of the embodiment of the present application, the surgical environment image is captured by the endoscope of the patient-side surgical platform and transmitted to the main console by the patient-side surgical platform. For example, the patient-side surgical platform transmits the surgical environment image captured by the endoscope to the main console via a remote network.

[0067] In some embodiments, in addition to obtaining the first instrument type of the first target surgical instrument for use in simulating the instrument sound of the first target surgical instrument, the first operation type of the first target surgical instrument is also obtained for use in simulating the instrument sound of the first target surgical instrument. This provides a more specific basis and conditions for subsequent simulation of the instrument sound of the first target surgical instrument, further improving the accuracy of the instrument sound simulation. It will be understood that in this embodiment, step S120 may include: obtaining a simulated audio signal of the first target surgical instrument based on the first instrument type and the first operation type.

[0068] In some embodiments, the patient-side surgical platform further includes an endoscope for acquiring an image of the surgical environment. Acquiring the first operation type of the first target surgical instrument is similar to acquiring the first instrument type, and may include: Identify a first operation type of a first target surgical instrument in a surgical environment image based on a preset machine learning model.

[0069] Specifically, the patient-side surgical platform stores a pre-trained machine learning model. The pre-trained machine learning model is pre-trained and can identify the type of operation being performed by the first target surgical instrument in the surgical environment image, that is, the first operation type.

[0070] It can be understood that the embodiment of the present application uses a pre-trained machine learning model to identify the first instrument type and the first operation type of the first target surgical instrument in the surgical environment image, which has the advantages of high recognition efficiency and high accuracy of recognition results.

[0071] In other embodiments, the first instrument type and the first operation type of the first target surgical instrument may also be determined based on a control instruction received by the patient-side surgical platform from the main console.

[0072] Specifically, the type of surgical instrument currently to be operated is determined based on the control command received by the patient-side surgical platform from the main console, thereby obtaining a first instrument type. Motion parameters (e.g., force, displacement, etc.) of each first target surgical instrument are calculated based on the control command. A preset machine learning model is then used to identify a first operation type of the first target surgical instrument based on the motion parameters. The preset machine learning model is pre-trained to output the operation type of the surgical instrument when the motion parameters of the surgical instrument are input.

[0073] Alternatively, pose data of each first target surgical instrument is estimated based on the control instruction, and a preset machine learning model is used to identify the first operation type of the first target surgical instrument based on the pose data. The preset machine learning model is pre-trained to output the operation type of the surgical instrument when the pose data of the surgical instrument is input.

[0074] Alternatively, the motion of each first target surgical instrument is estimated based on the control instruction, and a preset machine learning model is used to identify the first operation type of the first target surgical instrument based on the motion of the first target surgical instrument. The preset machine learning model is pre-trained to output the operation type of the surgical instrument when the motion of the surgical instrument is input.

[0075] In some embodiments, the first target surgical instrument comprises a surgical instrument in an operative state.

[0076] It can be understood that the embodiment of the present application simulates the instrument sound of surgical instruments in an operating state, so that the doctor can obtain complete instrument sound feedback when performing surgery by controlling the surgical instruments on the patient-side surgical platform through the main console, and there is no need to pay attention to surgical instruments that are not in an operating state during the instrument sound simulation process, which reduces the amount of calculation and improves the efficiency and real-time performance of the instrument sound simulation.

[0077] In some embodiments, the patient-side surgical platform further includes an audio system for acquiring instrument sounds of surgical instruments, and the first target surgical instruments include surgical instruments in an operating state whose sounds are not fully acquired.

[0078] Before step S110, this embodiment may further include: Based on the instrument audio signal obtained by the audio system, it is determined whether the sounds of the surgical instruments in the operating state are all obtained; if not, the first target surgical instrument is determined.

[0079] Specifically, the embodiment of the present application first determines whether the sounds of the surgical instruments in operation are fully captured based on the instrument audio signals captured by the audio system, and if the sounds of the surgical instruments in operation are not fully captured, determines the surgical instruments whose instrument sounds are not fully captured, and then executes steps S110-S120 to simulate the audio signals of these surgical instruments, i.e., the first target surgical instruments, thereby being able to fully feedback the instrument sounds during the surgery based on the instrument audio signals captured by the audio system and the simulated audio signals.

[0080] It can be understood that the embodiment of the present application simulates the instrument sound of surgical instruments in an operating state whose sounds are not fully captured. Compared with the embodiment of simulating the instrument sound of surgical instruments in an operating state, it has less simulation amount, improves the efficiency of instrument sound simulation, and the complete audio signal that is finally fed back and played is mixed with the real collected instrument audio signal, which improves the reliability of the complete audio signal, so that the doctor can obtain complete and more intuitive instrument sound feedback when performing surgery by controlling the surgical instruments on the patient-side surgical platform through the main console.

[0081] As previously mentioned, the audio system is co-located with the patient-side surgical platform. For example, the patient-side surgical platform is located in the operating room, and the audio system belongs to the patient-side surgical platform and is located in the same operating room as the patient-side surgical platform.

[0082] It should be noted that the audio system belongs to the patient-side surgical platform, that is, the patient-side surgical platform includes the audio system, which means that the audio system is controlled by the patient-side surgical platform and the audio processing method applied to the patient-side surgical platform, and the audio system can be deployed independently of the patient-side surgical platform or integrated into the patient-side surgical platform for deployment.

[0083] It is understood that the audio system included in the patient-side surgical platform can capture sounds emitted by various sound sources within the patient-side surgical platform's audio environment. The captured audio signals are represented by time-varying voltage waveforms. The audio environment of the patient-side surgical platform refers to an environment within which the patient-side surgical platform is located and contains at least one distributed sound source. Sound sources within this audio environment generally include various personnel, instruments, monitoring equipment, and the audio environment itself. Personnel include doctors, nurses, and patients, and conversations between these personnel generate audio signals. Patients' physiological activities (such as heartbeats and breathing) also generate audio signals. Instruments inserted into the patient's body for surgical or diagnostic procedures also generate audio signals. Monitoring equipment (such as monitors) monitors the patient's condition in real time and issues alarms when abnormalities occur. The audio environment itself refers to the ambient noise generated by the audio environment, which may result from reflections of sound waves from the audio environment.

[0084] Among them, the instrument audio signal acquired by the audio system of the embodiment of the present application refers to the audio signal generated by the surgical instrument acquired in real time by the audio system, that is, the audio signal generated by the surgical instrument in an operating state acquired in real time. However, considering that some instrument sounds may be missing when the audio system acquires instrument sounds, the instrument audio signal acquired by the audio system cannot be equated with the audio signal generated by the surgical instrument in an operating state. When the sound of the surgical instrument in an operating state is not fully acquired, the instrument audio signal will lack energy in some or all frequency bands compared to the audio signal generated by the surgical instrument in an operating state.

[0085] It is understood that in the embodiments of the present application: The full acquisition of the sound of the surgical instrument in operation means that the instrument audio signal acquired by the audio system has no energy loss or the energy loss value is less than the preset value in each frequency band or each key frequency band compared with the audio signal generated by the surgical instrument in operation.

[0086] The fact that the sound of a surgical instrument in operation is not fully captured includes two situations: the audio signal generated by the surgical instrument in operation is partially captured by the audio system of the patient-side surgical platform, and the fact that the audio signal is not captured at all by the audio system. The fact that the audio signal generated by the surgical instrument in operation is partially captured by the audio system of the patient-side surgical platform means that the instrument audio signal captured by the audio system is completely missing energy in some frequency bands or some key frequency bands compared to the audio signal generated by the surgical instrument in operation, or the energy missing value is greater than a preset value; the fact that the audio signal generated by the surgical instrument in operation is not fully captured by the audio system means that the audio signal generated by the surgical instrument in operation is completely missing energy in all frequency bands or all key frequency bands, or the energy missing value is greater than a preset value.

[0087] Therefore, the embodiment of the present application can determine whether the sound of the surgical instrument in operation is fully acquired based on whether the instrument audio signal acquired by the audio system has energy missing in some frequency bands or some key frequency bands or whether the energy missing value is greater than a preset value, thereby realizing the detection of the integrity of the instrument audio signal and facilitating the subsequent simulation of the missing instrument sound.

[0088] In addition, a surgical instrument in an operating state refers to a surgical instrument that the doctor is currently operating. When performing a surgical operation or diagnostic procedure, it will generate audio signals, including audio signals emitted by the surgical instrument itself and audio signals emitted by the human tissue in contact with the surgical instrument, such as the arc explosion sound emitted by the high-frequency electric knife and the bubble burst sound emitted by the human tissue in contact with the high-frequency electric knife, the audio signal obtained after demodulation of the high-frequency vibration of the ultrasonic knife when it is working, and the tissue coagulation sound emitted by the human tissue in contact with the ultrasonic knife.

[0089] It should be noted that the audio system of the present application can directly capture the sound of the surgical instrument in operation through the surgical instrument's interface to obtain the instrument's audio signal. For example, an electrical coupling interface and a wireless inductive coupling interface can be used. For example, a ring electrode can be embedded in the surgical instrument's quick-release interface to achieve non-contact audio acquisition; or, based on the principle of ultrasonic carrier waves, an ultrasonic modulated signal can be transmitted through the surgical instrument's rod.

[0090] Alternatively, the audio system of the embodiment of the present application can collect the sound of the surgical instrument in operation through an embedded audio acquisition unit. For example, a miniature microphone / microphone array can be deployed at the movable joint or end of the surgical instrument to collect the direct contact sound of the surgical instrument (such as the sound of tissue cutting); or a contact accelerometer or piezoelectric ceramic piece can be deployed at the movable joint or end of the surgical instrument to collect vibration signals as auxiliary inputs (such as the high-frequency vibration generated by the ultrasonic knife when it is working can be converted into an electrical signal). It is understandable that the microphone / microphone array can also be deployed around the endoscope and can further focus on the surgical instrument in operation and the human tissue being operated based on beamforming technology.

[0091] Because the functions or permissions of the interfaces corresponding to some surgical instruments may not be fully open, or they may not have the functions or permissions to capture and transmit audio from the surgical instruments, it may be impossible to obtain a complete audio signal when directly capturing the sound of the surgical instrument in an operating state through the surgical instrument interface. Therefore, embodiments of the present application can combine the above-mentioned two methods of obtaining audio signals through the surgical instrument interface and obtaining audio signals through the audio acquisition unit. For surgical instrument audio signals that cannot be directly obtained through the interface, they can be obtained through the audio acquisition unit, thereby improving the integrity of the instrument audio signal captured by the audio system, reducing the subsequent need for simulated instrument sound generation, improving audio processing efficiency, and improving the reliability of the complete audio signal by increasing the proportion of truly captured audio (i.e., the instrument audio signal captured by the audio system) in the final generated complete audio signal. It can be understood that when the sound of the surgical instrument in an operating state is fully captured, the complete audio signal is composed of the instrument audio signal captured by the audio system; when the sound of the surgical instrument in an operating state is not fully captured, the complete audio signal is composed of the simulated audio signal generated by the simulation and the instrument audio signal captured by the audio system.

[0092] In order to more accurately determine whether the sounds of the surgical instrument in operation are fully acquired, in some embodiments, such as Figure 8 As shown, judging whether all sounds of the surgical instrument in operation are acquired based on the instrument audio signal acquired by the audio system may include: S111. Extracting voiceprint features of the device audio signal.

[0093] S112. Calculate the matching degree between the voiceprint feature and the standard voiceprint feature of the surgical instrument in operation.

[0094] S113: If the matching degree is greater than a preset first matching degree threshold, all sounds of the surgical instrument in operation are acquired; otherwise, not all sounds of the surgical instrument in operation are acquired.

[0095] For step S111, the voiceprint features of the instrument audio signal include the spectral energy features of the instrument audio signal, or may further include the time domain and non-stationary features of the instrument audio signal to supplement the transient events and nonlinear features that cannot be captured in the frequency domain, thereby improving the accuracy and reliability of the voiceprint feature extraction and comparison results.

[0096] Specifically, in some embodiments, the voiceprint features of instrument audio signals are Mel-Frequency Cepstral Coefficients (MFCC) features. The process of extracting the voiceprint features of instrument audio signals, namely the MFCC feature extraction process, specifically includes pre-emphasis, frame splitting and windowing, time-frequency domain transformation, Mel filtering, logarithmic compression, obtaining cepstral coefficients, and taking the first n dimensions (typically the first 13-20 dimensions) of the instrument audio signal to obtain the MFCC features.

[0097] It can be understood that the process of extracting MFCC features of the instrument audio signal includes compressing the sound signal with high spectral dimensions into cepstral coefficients with low spectral dimensions, which involves discrete cosine transform.

[0098] To more completely preserve the voiceprint information of the instrument audio signal and minimize noise, in some embodiments, n is 13, meaning the first 13 dimensions are used to obtain the cepstral coefficients, resulting in MFCC features. Because the discrete cosine transform concentrates energy on low-frequency coefficients, while high-frequency coefficients are mostly detail noise, the embodiments of the present application, when performing MFCC feature extraction on the instrument audio signal, obtain the cepstral coefficients and take the first 13 dimensions, thereby preserving sufficient voiceprint information of the instrument audio signal and suppressing noise.

[0099] Regarding step S112, the standard voiceprint characteristics of the surgical instrument in operation refer to the voiceprint characteristics of the standard audio signal generated by the surgical instrument in operation, and the standard audio signal refers to the standard sound emitted by the operation currently performed by the surgical instrument in operation. It is understandable that the standard audio signal can be collected in advance in an audio environment with low background noise (for example, less than 20dB) and pre-stored in a standard database. Exemplarily, for the collection of the standard audio signal of each surgical instrument, the surgical instrument can be fixed in an anechoic chamber, and a microphone can be used to collect the sound of the surgical instrument under various typical operations at close range to obtain the standard audio signal of the surgical instrument. More specifically, for the collection of the standard audio signal of the electrocoagulation knife, the electrocoagulation knife can be fixed in an anechoic chamber, and a microphone can be used to collect the sound of the electrocoagulation knife under various typical operations (such as no-load excitation, tissue contact, coagulation, etc.) at close range.

[0100] The matching degree between the voiceprint feature and the standard voiceprint feature of the surgical instrument in operation refers to the cosine similarity between the MFCC feature and the standard voiceprint feature. Taking n as 13 as an example, the calculation process of the matching degree may include: Normalize the MFCC features by Z-score to obtain the normalized MFCC features : , in, Represents 13-dimensional MFCC features. and from the training set of the feature extraction algorithm, is the mean vector of each set of features in the training set, is the standard deviation vector of each set of features in the training set. Thus, the 13-dimensional normalized MFCC features can be obtained .

[0101] Then, calculate Cosine similarity with standard voiceprint features : , Alternatively, the matching degree between the voiceprint feature and the standard voiceprint feature of the surgical instrument in operation refers to the dynamic time warping distance between the MFCC feature and the standard voiceprint feature. In this case, the matching degree calculation process may include: Calculate the dynamic time warping distance between MFCC features and standard voiceprint features : , in, is the Euclidean distance between MFCC features and standard voiceprint features.

[0102] For step S113, according to part of the content of step S112, the matching degree between the voiceprint feature and the standard voiceprint feature of the surgical instrument in the operating state is greater than the preset first matching degree threshold. It can be that the cosine similarity between the normalized MFCC feature of the instrument audio signal and the standard voiceprint feature is greater than the preset cosine similarity threshold, or the dynamic time warping distance between the MFCC feature of the instrument audio signal and the standard voiceprint feature is greater than the preset dynamic time warping distance threshold.

[0103] Therefore, the embodiment of the present application improves the accuracy of the judgment result by extracting the voiceprint features of the instrument audio signal obtained by the audio system, and judging whether the sound of the surgical instrument in the operating state is fully obtained based on whether the matching degree between the voiceprint features and the standard voiceprint features of the surgical instrument in the operating state is greater than the preset first matching degree threshold.

[0104] In some embodiments, obtaining the standard voiceprint features of a surgical instrument in an operating state may include: A second instrument type of the surgical instrument in an operating state is obtained; and a standard voiceprint feature is obtained based on the second instrument type.

[0105] In some embodiments, obtaining a standard voiceprint feature based on the second device type may include: Based on the second instrument type, a standard voiceprint feature of the surgical instrument in an operating state is obtained from a preset standard database.

[0106] Among them, the standard voiceprint features of various types of surgical instruments are pre-stored in the standard database. The embodiment of the present application improves the efficiency of obtaining the standard voiceprint features of the surgical instruments in the operating state by directly obtaining the standard voiceprint features of the surgical instruments in the operating state from the standard database based on the second instrument type.

[0107] In other embodiments, obtaining the standard voiceprint feature based on the second device type may include: The second device type is input into a preset machine learning model to output standard voiceprint features.

[0108] Among them, the preset machine learning model is trained in advance to output the standard voiceprint features of the surgical instrument when the instrument type of the surgical instrument is input.

[0109] Thus, the embodiment of the present application can directly utilize a machine learning model to obtain the standard voiceprint features of a surgical instrument in operation based on the second instrument type, without the need to build a standard database in advance. The machine learning model can also be automatically trained and optimized. Compared to the embodiment in which a standard database is built in advance and the standard voiceprint features of a surgical instrument in operation are obtained from the standard database, this embodiment can dynamically generate more matching standard voiceprint features based on the real-time requirements of the surgical scene (such as instrument material, operating force, tissue type, etc.), which has greater flexibility. In addition, this embodiment can also combine acoustic physics models to generate standard voiceprint features that are closer to real instrument operation, without being limited by the accuracy of the recording equipment or environmental noise, thereby obtaining more accurate and reliable standard voiceprint features.

[0110] In some embodiments, obtaining the standard voiceprint features of a surgical instrument in an operating state may include: A second instrument type and a second operation type of the surgical instrument in an operating state are obtained; and a standard voiceprint feature is obtained based on the second instrument type and the second operation type.

[0111] In some embodiments, the patient-side surgical platform further includes an endoscope for acquiring an image of the surgical environment, and acquiring the second instrument type and the second operation type of the surgical instrument in an operating state may include: Based on a preset machine learning model, a second instrument type and a second operation type of a surgical instrument in an operating state in a surgical environment image are identified.

[0112] Specifically, the patient-side surgical platform stores a pre-trained machine learning model that is trained in advance to identify the instrument type and the type of operation being performed, i.e., the second instrument type and the second operation type, of the surgical instruments in operation in the surgical environment image.

[0113] It can be understood that the embodiment of the present application uses a pre-trained machine learning model to identify the second instrument type and second operation type of surgical instruments in an operating state in a surgical environment image, which has the advantages of high recognition efficiency and high accuracy of recognition results.

[0114] In other embodiments, the second instrument type and the second operation type of the surgical instrument in the operating state may also be determined based on the control instructions received by the patient-side surgical platform from the main console.

[0115] Specifically, the type of surgical instrument currently in operation is determined based on the control command received by the patient-side surgical platform from the main console to obtain a second instrument type. Motion parameters (e.g., force, displacement, etc.) of each surgical instrument in operation are calculated based on the control command. Based on the motion parameters, a preset machine learning model is used to identify the second operation type of the surgical instrument in operation. The preset machine learning model is pre-trained to output the operation type of the surgical instrument when the motion parameters of the surgical instrument are input.

[0116] Alternatively, the position data of each surgical instrument in an operating state is estimated based on the control instructions, and a preset machine learning model is used to identify the second operation type of the surgical instrument in the operating state based on the position data. The preset machine learning model is pre-trained to output the operation type of the surgical instrument when the position data of the surgical instrument is input.

[0117] Alternatively, the motion of each surgical instrument in an operating state is estimated based on the control command, and a preset machine learning model is used to identify the second operation type of the surgical instrument in an operating state based on the motion of the surgical instrument in the operating state. The preset machine learning model is pre-trained to output the operation type of the surgical instrument when the motion of the surgical instrument is input.

[0118] In some embodiments, obtaining a standard voiceprint feature based on the second device type and the second operation type may include: Based on the second instrument type and the second operation type, a standard voiceprint feature of the surgical instrument in an operating state is obtained from a preset standard database.

[0119] Among them, the standard voiceprint features of each surgical instrument are pre-stored in the standard database. The embodiment of the present application improves the efficiency of obtaining the standard voiceprint features of the surgical instrument in the operating state by directly obtaining the standard voiceprint features of the surgical instrument in the operating state from the standard database.

[0120] It can be understood that, in this embodiment, the standard voiceprint features of each surgical instrument pre-stored in the standard database include voiceprint features corresponding to standard audio signals of each surgical instrument under various typical operations.

[0121] Therefore, some embodiments of the present application obtain standard voiceprint features that match the instrument type and operation type from a standard database based on the instrument type and operation type of the surgical instrument in an operating state, so that the standard voiceprint features are more consistent with the voiceprint features of the audio signal emitted by the surgical instrument under the current operation, thereby improving the accuracy and reliability of the obtained standard voiceprint features, thereby improving the reliability of the voiceprint feature comparison results, that is, it can more accurately and reliably determine whether the sound of the surgical instrument in an operating state has been fully acquired.

[0122] In some embodiments, building a standard database may include: A standard database is established, the voiceprint features of the standard audio signals of the surgical instruments included in the patient-side surgical platform are extracted, the standard voiceprint features are obtained, and the standard audio signals and the corresponding standard voiceprint features are stored in the standard database.

[0123] As can be seen from the above, the standard audio signal refers to the standard sound emitted by the operation currently performed by the surgical instrument in the operating state. It is understandable that the standard audio signal can be collected in advance in an audio environment with low background noise (for example, less than 20dB) and pre-stored in the standard database. Exemplarily, for the collection of the standard audio signal of each surgical instrument, the surgical instrument can be fixed in an anechoic chamber, and a microphone can be used to collect the sound of the surgical instrument under various typical operations at close range to obtain the standard audio signal of the surgical instrument. More specifically, for the collection of the standard audio signal of the electrocoagulation knife, the electrocoagulation knife can be fixed in an anechoic chamber, and a microphone can be used to collect the sound of the electrocoagulation knife under various typical operations (such as no-load excitation, tissue contact, coagulation, etc.) at close range. The standard voiceprint features are obtained by feature extraction based on the standard audio signal, which includes the standard voiceprint features of the surgical instrument under various typical operations.

[0124] Furthermore, in some embodiments, the feature extraction algorithm for extracting the standard voiceprint features is the same as the feature extraction algorithm for extracting the voiceprint features of the above-mentioned instrument audio signal, so as to improve the accuracy and reliability of voiceprint feature comparison.

[0125] In other embodiments, obtaining the standard voiceprint feature based on the second device type and the second operation type may include: The second device type and the second operation type are input into a preset machine learning model to output standard voiceprint features.

[0126] Among them, the preset machine learning model is trained in advance to output the standard voiceprint features of the surgical instrument when the instrument type and operation type of the surgical instrument are input.

[0127] Thus, this embodiment of the present application can directly utilize a machine learning model to obtain standard voiceprint features of a surgical instrument in operation based on the second instrument type and the second operation type, eliminating the need to pre-build a standard database. Furthermore, the machine learning model can be automatically trained and optimized. Compared to embodiments that pre-build a standard database and obtain standard voiceprint features of a surgical instrument in operation from that database, this embodiment can dynamically generate more accurate and reliable standard voiceprint features based on the real-time requirements of the surgical scenario (e.g., instrument material, operation force, tissue type, etc.), offering greater flexibility. Furthermore, this embodiment can also incorporate acoustic physics models to generate standard voiceprint features that are closer to actual instrument operation, unconstrained by the accuracy of the recording equipment or ambient noise, thereby obtaining more accurate and reliable standard voiceprint features. Furthermore, compared to embodiments that utilize a machine learning model to obtain standard voiceprint features of a surgical instrument in operation based on the second instrument type, this embodiment's machine learning model can generate more accurate and reliable standard voiceprint features based on the instrument type and operation type. These standard voiceprint features are more consistent with the voiceprint features corresponding to the audio signals generated by the surgical instrument during the current operation.

[0128] In order to more accurately determine whether the sounds of the surgical instrument in operation are fully acquired, in other embodiments, such as Figure 9 As shown, judging whether all sounds of the surgical instrument in operation are acquired based on the instrument audio signal acquired by the audio system may further include: S111 ′, detecting a vibration signal of a surgical instrument in an operating state.

[0129] S112', performing coherence analysis on the vibration signal and the instrument audio signal to obtain a coherence coefficient.

[0130] S113': If the coherence coefficient is greater than a preset coherence coefficient threshold, all sounds of the surgical instrument in operation are acquired; otherwise, not all sounds of the surgical instrument in operation are acquired.

[0131] In step S111', specifically, an accelerometer is installed on the surgical instrument, and the accelerometer is used to detect a vibration signal of the surgical instrument in operation. It is understood that the vibration signal in the embodiment of the present application is obtained by directly measuring the mechanical vibration of the surgical instrument using the accelerometer installed on the surgical instrument, and converting the mechanical vibration of the surgical instrument into an electrical signal. This electrical signal, i.e., the vibration signal, is a voltage value that varies with time and corresponds to the intensity and frequency of the vibration of the surgical instrument.

[0132] Among them, the bandwidth of the accelerometer needs to cover the typical vibration frequencies of all surgical instruments on the patient-side surgical platform of the embodiment of the present application, and the range needs to meet the impact requirements of the surgical instruments in the operating state.

[0133] In some embodiments, the accelerometer is installed at a movable joint or a transmission component of the surgical instrument, so that the accelerometer can directly contact the vibration source, thereby improving the accuracy of vibration signal detection.

[0134] Regarding step S112', specifically, the embodiment of the present application first generates a vibration spectrum based on the vibration signal, generates a sound spectrum based on the instrument audio signal, and then calculates the coherence value of each frequency point of the vibration spectrum and the sound spectrum. : , in, is the vibration spectrum, For the sound spectrum.

[0135] Then, based on the coherence value of each frequency point , calculate the average coherence coefficient of the vibration spectrum and the sound spectrum as the coherence coefficient of the embodiment of the present application: , in, is a frequency band covering the vibration range of the surgical instruments on the patient-side surgical platform of the embodiment of the present application; Frequency band The average coherence coefficient of the vibration spectrum and the sound spectrum.

[0136] More specifically, the vibration spectrum The generation of the vibration spectrum can include: after pre-processing the vibration signal, the vibration signal is divided into short time periods (for example, 20-40ms as a frame) by framing, and then each frame signal is Fourier transformed to convert the time domain signal into a frequency domain signal to obtain the vibration spectrum. It is understandable that the vibration spectrum Displays the distribution of vibration energy at different frequencies. For example, the vibration spectrum of electrocoagulation knife There is a peak at 5kHz.

[0137] Preprocessing of the vibration signal includes removing DC offset, noise reduction, and normalization. A high-pass filter can be used to remove DC offset from the vibration signal, removing low-frequency interference in the vibration signal, such as baseline drift caused by the accelerometer itself. A band-pass filter can be used to denoise the vibration signal, retaining the key frequency band of the vibration signal. The key frequency band must cover the vibration range of the surgical instruments on the patient-side surgical platform of the embodiment of the present application, for example, the key frequency band can be 300Hz-22kHz. The amplitude of the vibration signal can be adjusted for normalization to avoid overflow in the subsequent calculation of the vibration spectrum.

[0138] Furthermore, when the vibration signal is divided into short time periods by framing, the Hamming window can be used to smooth the edge of each frame signal, that is, framing and windowing, so as to reduce the error of spectrum analysis and improve the calculated vibration spectrum. The accuracy of the coherence analysis results is improved.

[0139] Sound Spectrum The generation of the sound spectrum can include: pre-processing the instrument audio signal, dividing the instrument audio signal into short time periods (for example, 20-40ms as a frame), and then performing Fourier transform on each frame signal to convert the time domain signal into the frequency domain signal to obtain the sound spectrum. Understandably, the sound spectrum Displays the sound energy distribution of different frequencies, that is, the aforementioned spectrum energy characteristics. There is strong energy at 45kHz.

[0140] Preprocessing of the instrument audio signal includes denoising and sound pressure calibration. The denoising process performed by the present embodiment can filter out low-frequency ambient noise and high-frequency interference (such as electrosurgical noise) from the instrument audio signal. Sound pressure calibration of the instrument audio signal can convert the original voltage of the instrument audio signal into standard sound pressure units.

[0141] Furthermore, when the instrument audio signal is divided into short time periods by framing, the Hamming window can be used to smooth the edge of each frame signal, that is, framing and windowing, so as to reduce the error of spectrum analysis and improve the calculated sound spectrum. The accuracy of the coherence analysis results is improved.

[0142] For step S113', the coherence coefficient threshold is pre-set and used to compare with the coherence coefficient calculated in step S112' (ie, the average coherence coefficient ) comparison to determine whether the sounds of surgical instruments in operation are fully captured.

[0143] For example, the coherence coefficient threshold is 0.85. When , the sounds of surgical instruments in operation are all captured; when When the surgical instrument is in operation, the sound is not fully captured.

[0144] In some embodiments, the coherence coefficient threshold may include multiple sub-coherence coefficient thresholds. By comparing the sub-coherence coefficient thresholds with the coherence coefficient, more instrument sound acquisition states can be identified.

[0145] Exemplarily, the coherence coefficient threshold includes two sub-coherence coefficient thresholds, which are 0.85 and 0.5 respectively. When , the instrument audio signal is strongly correlated with the vibration signal, and the sound of the surgical instrument in operation is fully acquired; when When , the sound of the surgical instrument in operation is partially acquired; when When the instrument audio signal is decoupled from the vibration signal, it is considered that the sound of the surgical instrument in operation is not obtained at all.

[0146] Therefore, the embodiment of the present application determines whether the sound of the surgical instrument in operation is fully captured based on coherence analysis, which has higher accuracy than the aforementioned embodiment based on energy comparison of each frequency band / key frequency band. At the same time, compared with the embodiment based on voiceprint feature comparison, it has stronger resistance to environmental noise interference, and does not require complex calculations such as feature extraction and analysis, avoiding the problem of delay introduced by complex algorithms.

[0147] In order to more accurately determine whether the sounds of the surgical instrument in operation are fully acquired, in other embodiments, such as Figure 10 As shown, judging whether all sounds of the surgical instrument in operation are acquired based on the instrument audio signal acquired by the audio system may further include: S111” performs short-time Fourier transform (STFT) on the instrument audio signal and obtains each independent audio signal in the instrument audio signal based on independent component analysis (ICA).

[0148] S112 ”, if the number of independent audio signals is less than the number of surgical instruments in operation, then the sounds of the surgical instruments in operation are not all acquired.

[0149] S113", if the number of independent audio signals is equal to the number of surgical instruments in operation, the surgical instrument corresponding to each independent audio signal is identified based on the preset machine learning model, and when the matching degree between the independent voiceprint feature of each independent audio signal and the standard independent voiceprint feature of the corresponding surgical instrument is greater than the preset second matching degree threshold, it is confirmed that the sounds of the surgical instruments in operation are all acquired; otherwise, the sounds of the surgical instruments in operation are not all acquired.

[0150] Regarding step S111, it can be understood that the instrument audio signal acquired by the audio system is generally a mixed audio signal generated when multiple instruments perform operations. In the embodiment of the present application, the instrument audio signal is set to: , in, The number of sensors (such as microphones) used in the audio system to obtain the audio signal of the device. Indicates the audio signal obtained by each sensor. Specifically, the sensor The obtained audio signal is: , in, Indicates the Audio signals generated when a surgical instrument is operated; Indicates the surgical instruments to sensors The attenuation coefficient; represents the propagation delay, Surgical instruments and sensors Determined by the spatial position of Indicates ambient noise.

[0151] Specifically, the process of performing STFT on the instrument audio signal is similar to the above-mentioned process of obtaining the sound spectrum. The process is similar to that of the first step. The instrument audio signal is preprocessed and framed and windowed. Then the instrument audio signal is subjected to STFT to obtain the time-frequency matrix of the instrument audio signal: , Then, the time-frequency matrix of the instrument audio signal is solved by frequency domain ICA to solve the separation matrix , such that: , in, is the estimated spectrum of the independent audio signal, so that each independent audio signal in the instrument audio signal is obtained by inverse STFT.

[0152] It is understandable that, when the standard independent voiceprint features are also extracted based on STFT, step S111" may also be: performing a short-time Fourier transform on the instrument audio signal, and obtaining the spectrum of each independent audio signal in the instrument audio signal based on independent component analysis. That is, it is only necessary to obtain the spectrum of each independent audio signal (i.e., independent voiceprint features), and there is no need to reconstruct each independent audio signal in the instrument audio signal through inverse STFT. In this way, the amount of calculation can be reduced without affecting the subsequent comparison of the independent voiceprint features with the standard independent voiceprint features.

[0153] Furthermore, in some embodiments, the estimated spectrum of the independent audio signal is optimized by time-frequency masking, so that a purer spectrum of the independent audio signal can be extracted, thereby improving the accuracy of the judgment result. Specifically: For each sound source (surgical instrument and the human tissue it operates on) , calculate the binary time-frequency mask : , in, is an empirical threshold used to retain significant energy regions. In some embodiments, .

[0154] Then, based on the binary time-frequency mask , extract the sound source The clean spectrum after masking: , For step S112, the embodiment of the present application adopts a simple method of quantity comparison. When the number of independent audio signals is less than the number of surgical instruments in operation, it indicates that the audio signals generated by some surgical instruments in operation have not been acquired. In this way, it can be directly determined that the sounds of the surgical instruments in operation have not been fully acquired.

[0155] Regarding step S113", it can be understood that when the number of independent audio signals is equal to the number of surgical instruments in operation, it cannot be determined that the sounds of the surgical instruments in operation are fully captured, and the audio signals generated by some surgical instruments may be partially captured, resulting in a situation where the number of independent audio signals is equal to the number of surgical instruments in operation, but the sounds of the surgical instruments in operation are not fully captured. Therefore, in the embodiment of the present application, when the number of independent audio signals is equal to the number of surgical instruments in operation, the surgical instrument corresponding to each independent audio signal is identified based on a preset machine learning model, and when the degree of matching between the independent voiceprint feature of each independent audio signal and the standard independent voiceprint feature of the corresponding surgical instrument is greater than a preset second matching degree threshold, it is confirmed that the sounds of the surgical instruments in operation are fully captured; otherwise, the sounds of the surgical instruments in operation are not fully captured.

[0156] The preset machine learning model is trained in advance to output the surgical instrument corresponding to the independent audio signal when the independent audio signal is input.

[0157] In addition, the feature extraction process of extracting independent voiceprint features and the acquisition of standard independent voiceprint features are the same as the feature extraction process of extracting voiceprint features and the acquisition of standard voiceprint features described in steps S111-S113, and are not repeated here.

[0158] Unlike steps S111-S113, steps S111"-S113" perform independent audio signal estimation on the instrument audio signal and then compare the independent voiceprint features of each independent audio signal with the standard independent voiceprint features, while steps S111-S113 directly compare the voiceprint features of the instrument audio signal with the standard voiceprint features of the surgical instrument in operation. Therefore, although steps S111"-S113" are more complicated to operate than steps S111-S113, they can reduce the interference caused by the spectral overlap of the instrument audio signal, thereby obtaining a more accurate comparison result, that is, it can more accurately determine whether the sound of the surgical instrument in operation has been fully acquired.

[0159] It is understandable that if the number of independent audio signals is greater than the number of surgical instruments in operation, it may be that the short-time Fourier transform or independent component analysis process of step S111" is incorrect, or noise is mixed in. The embodiment of the present application can perform noise reduction processing on the instrument audio signal and re-execute step S111".

[0160] In order to enable the audio system to obtain clearer instrument audio signals, in some embodiments, when the audio acquisition unit based on the audio system obtains the instrument audio signal, the operating surgical instrument can be directionally picked up and / or at a fixed distance based on the positioning information of the surgical instrument in the operating state.

[0161] Among them, the positioning information of the surgical instrument in the operating state can be obtained based on the aforementioned posture data, or, in some embodiments, the audio system includes a microphone array composed of a plurality of microphones distributed in an array, and the positioning information of the surgical instrument in the operating state can be obtained through the microphone array of the audio system.

[0162] Specifically, the sound source (surgical instrument in an operating state) is located by using a microphone array, including calculating the direction and distance of the sound source relative to the microphone array.

[0163] It is understood that in the embodiments of this application, the distance of the sound source from the microphone array is comparable to the size of the microphone array, that is, it belongs to the near-field model. Taking into account the curvature of the spherical propagation of sound waves, the direction and distance of the sound source from the microphone array are directly calculated.

[0164] Specifically, the generalized cross-correlation (GCC-PHAT) is used to estimate the sound source arriving at the microphones in the microphone array. and microphone Delay difference . It is calculated by the following function: , in, For microphone The frequency domain signal, For microphone frequency domain signal.

[0165] is a weighting function used to improve robustness, such as PHAT weighting: , Get function The time corresponding to the peak As a delay difference .

[0166] Then, the direction and distance (coordinates in the reference coordinate system) of the sound source relative to the microphone array are calculated based on the estimated time delay difference.

[0167] , in, is the sound source coordinate, For microphone The coordinates of For microphone The coordinates of is the speed of sound.

[0168] It is understood that for two-dimensional positioning information, at least three microphones in the microphone array need to be used to construct at least two independent equations as above for calculation. Correspondingly, for three-dimensional positioning information, at least four microphones in the microphone array need to be used to construct at least three independent equations as above for calculation.

[0169] The embodiment of the present application uses a microphone array, which can directly calculate the coordinates of each sound source in the reference coordinate system under the near-field model, thereby obtaining the direction and distance of each surgical instrument in operation compared to the reference point (such as the microphone array).

[0170] It is understood that in embodiments of the present application, the coordinates of each microphone in the microphone array can be set using the reference point as the origin of the reference coordinate system. Furthermore, the direction and distance of the sound source relative to the microphone array can be calculated using the above algorithm, and the direction and distance of the sound source relative to the reference point can be calculated based on the relative position of the reference point and the microphone array, according to the actual setting of the reference point.

[0171] In order to improve the accuracy of positioning the surgical instrument in the operating state, in some embodiments, the number of microphones in the microphone array is greater than or equal to the number of surgical instruments on the patient-side surgical platform.

[0172] Directional sound pickup of the surgical instrument in operation according to the positioning information of the surgical instrument in operation specifically includes: Based on the positioning information of the surgical instruments in operation, the delay and weight of each channel of the microphone array are adjusted through the beamforming algorithm to form a beam pointing to each surgical instrument in operation.

[0173] The method of performing distance sound pickup on the surgical instrument in the operating state according to the positioning information of the surgical instrument in the operating state specifically includes: Based on the positioning information of the surgical instrument in operation, only the audio signals within the target distance range are retained through distance gating filtering. The target distance range includes multiple distance ranges, each distance range corresponds to a fixed-distance sound pickup operation of a sound source, and the distance range corresponding to a sound source can be [0, d], where d represents the distance between the sound source and the reference point; or, the distance range corresponding to a sound source can also be [dmin, dmax], where dmin=d-dref, dmax=d+dref, and dref is a preset distance error value. As mentioned above, in the embodiments of the present application, distance refers to the distance between the sound source and the reference point.

[0174] The embodiments of the present application can perform directional sound pickup or fixed-distance sound pickup on surgical instruments in an operating state, or perform both directional sound pickup and fixed-distance sound pickup on surgical instruments in an operating state. It is understandable that when performing both directional sound pickup and fixed-distance sound pickup on surgical instruments in an operating state, it is possible to collect and output clearer sounds from each surgical instrument in an operating state, that is, output clearer instrument audio signals.

[0175] In order to enable the audio system to obtain clearer instrument audio signals and reduce noise interference, in some embodiments, active noise cancellation (ANC) is also performed on the instrument audio signals.

[0176] Specifically, the embodiment of the present application uses fast Fourier transform to decompose the noise in the instrument audio signal into frequency domain components, identify the main noise frequency, and dynamically track noise changes based on adaptive filtering, and then delay the noise signal and invert it (180° phase difference) to form an anti-phase sound wave, which is used to destructively interfere with the noise in the instrument audio signal.

[0177] Furthermore, the embodiment of the present application can adjust the amplitude of the anti-phase sound wave so that the amplitude of the anti-phase sound wave is consistent with the amplitude of the noise, thereby maximizing the cancellation effect.

[0178] Therefore, the embodiment of the present application can accurately determine whether the sounds of the surgical instruments in operation are fully captured through the content described in step S110, providing a reliable basis for subsequent instrument sound simulation of surgical instruments whose sounds are not fully captured.

[0179] As described above, the first target surgical instrument includes surgical instruments whose sounds are not fully captured among surgical instruments in an operating state.

[0180] Specifically, embodiments of the present application can analyze the motion of an operating surgical instrument based on an endoscopically captured image of the surgical environment, such as an electric knife contacting tissue. If the motion of the surgical instrument is visually detected but the expected sound is not captured, the surgical instrument is determined to be the first target surgical instrument. Alternatively, embodiments of the present application can also determine the first target surgical instrument based on the absence of spectral energy in the instrument's audio signal.

[0181] In order to more accurately determine the first target surgical instrument and avoid subsequent over-simulation of instrument sounds, which increases the amount of calculation and affects the final audio effect, in some embodiments, determining the first target surgical instrument may include: Based on the contents described in steps S111"-S113", among the surgical instruments corresponding to the independent audio signals, the surgical instrument whose independent voiceprint feature matches the standard independent voiceprint feature greater than a second matching degree threshold is determined as the second target surgical instrument. The second target surgical instrument is the surgical instrument whose sound is fully captured among the surgical instruments in the operating state. Then, based on the second target surgical instrument and the surgical instrument in the operating state, the first target surgical instrument is obtained.

[0182] It is understandable that, since the number of independent audio signals is not necessarily equal to the number of surgical instruments in operation, and due to the influence of noise, etc., independent audio signals whose independent voiceprint features have a matching degree with the standard independent voiceprint features less than the second matching degree threshold are not necessarily generated by surgical instruments. Therefore, the embodiment of the present application cannot directly determine the first target surgical instrument based on the independent audio signals whose independent voiceprint features have a matching degree with the standard independent voiceprint features less than or equal to the second matching degree threshold. It is necessary to first determine the second target surgical instrument according to the above steps, and then remove the second target surgical instrument from the surgical instruments in operation to obtain the first target surgical instrument.

[0183] It can be understood that, on the premise that the number of independent audio signals is equal to the number of surgical instruments in operation, and each independent audio signal can correspond to a surgical instrument in operation, the embodiment of the present application can directly determine the surgical instrument corresponding to the independent audio signal, whose independent voiceprint feature has a matching degree with the standard independent voiceprint feature less than or equal to the second matching degree threshold, as the first target surgical instrument.

[0184] In order to more accurately determine the first target surgical instrument and avoid subsequent over-simulation of instrument sounds, which increases the amount of calculation and affects the final audio effect, in other embodiments, determining the first target surgical instrument may further include: A first target surgical instrument is determined based on the voiceprint features of the instrument audio signal and the standard voiceprint features of the surgical instrument in an operating state.

[0185] According to the above, the fact that the sound of the surgical instrument in operation is not fully captured includes two situations: the audio signal generated by the surgical instrument in operation is partially captured by the audio system of the patient-side surgical platform, and the fact that the audio signal is not captured at all by the audio system. Among them, the fact that the audio signal generated by the surgical instrument in operation is partially captured by the audio system of the patient-side surgical platform can be manifested as the instrument audio signal captured by the audio system is completely missing energy in some frequency bands or some key frequency bands compared to the audio signal generated by the surgical instrument in operation, or the energy missing value is greater than a preset value; the fact that the audio signal generated by the surgical instrument in operation is not fully captured by the audio system can be manifested as the audio signal generated by the surgical instrument in operation is completely missing energy in all frequency bands or all key frequency bands, or the energy missing value is greater than a preset value.

[0186] At the same time, the standard audio signals of each surgical instrument on the patient-side surgical platform have energy distribution characteristics with high discrimination in the key frequency band. The first target surgical instrument can be determined based on the energy distribution characteristics of the standard audio signals of each surgical instrument in the key frequency band and the frequency band with energy loss corresponding to the instrument audio signal.

[0187] In step S120, after determining the first target surgical instrument, embodiments of the present application can generate an analog audio signal for the first target surgical instrument on the patient-side surgical platform, or generate a first instruction via the patient-side surgical platform and send it to the main console, so that the main console generates an analog audio signal for the first target surgical instrument in response to the first instruction. The principles and steps for generating the analog audio signal for the first target surgical instrument on the patient-side surgical platform are the same as those for generating the analog audio signal for the first target surgical instrument on the main console. The simulated audio signal of the first target surgical instrument can be obtained directly by obtaining a standard audio signal template from a standard database; or, the simulated audio signal of the first target surgical instrument can be obtained by obtaining a standard audio signal template from a standard database based on the current operating status of the first target surgical instrument, which has higher accuracy; or, it can be obtained by simulating and generating the corresponding instrument sound by identifying the surgical operation currently being performed by the first target surgical instrument; or, the simulated audio signal of the first target surgical instrument can be synthesized through physical modeling. For example, for high-frequency surgical instruments (such as electric knives), a pulse modulation signal is generated based on the parameters of the electrosurgical equipment (power, frequency, etc.) to simulate the explosion sound effect. For mechanical surgical instruments (such as scissors), friction and collision sounds are synthesized based on the kinematic data (speed, acceleration, etc.) of the mechanical surgical instrument.

[0188] In order to more accurately simulate the simulated audio signal of the first target surgical instrument and improve the final audio quality, in some embodiments, generating the simulated audio signal of the first target surgical instrument may include: Based on the first instrument type, a standard audio signal of the first target surgical instrument is obtained from a preset standard database as the simulated audio signal.

[0189] As can be seen from the above, standard audio signals of various types of surgical instruments can be pre-stored in the standard database. The embodiment of the present application can obtain the standard audio signal of each first target surgical instrument from the standard database based on the first instrument type of the first target surgical instrument, and mix it to obtain the final analog audio signal.

[0190] It can be understood that directly acquiring a standard audio signal matching the instrument type from a standard database based on the instrument type as the simulated audio signal of the corresponding instrument has higher audio simulation efficiency.

[0191] In other embodiments, obtaining the simulated audio signal of the first target surgical instrument based on the first instrument type and the first operation type may include: The first instrument type is input into a preset neural network model, and a simulated audio signal is output.

[0192] The preset neural network model is trained in advance to output a simulated audio signal of the first target surgical instrument when a first instrument type of the first target surgical instrument is input.

[0193] Thus, the embodiment of the present application can directly utilize a neural network model to obtain a simulated audio signal of a first target surgical instrument based on a first instrument type, without the need to construct a standard database in advance. Furthermore, the neural network model can be automatically trained and optimized. Compared to embodiments that construct a standard database in advance and obtain the simulated audio signal of the first target surgical instrument from the standard database, this embodiment can dynamically generate a more suitable simulated audio signal based on the real-time requirements of the surgical scenario (such as instrument material, operating force, tissue type, etc.), providing greater flexibility. Furthermore, this embodiment can also combine acoustic physics models to generate simulated audio signals that are closer to actual instrument operation, without being limited by the accuracy of the recording equipment or ambient noise, thereby obtaining more accurate and reliable simulated audio signals.

[0194] In order to more accurately simulate the simulated audio signal of the first target surgical instrument, so that the simulated audio signal is more consistent with the audio signal emitted by the first target surgical instrument during the current operation, and to improve the final audio effect, in some embodiments, generating the simulated audio signal of the first target surgical instrument may include: Based on the first instrument type and the first operation type, a standard audio signal of the first target surgical instrument is acquired from a preset standard database as the simulated audio signal.

[0195] As can be seen from the above, the standard database can pre-store standard audio signals of each surgical instrument when performing various typical operations. The embodiment of the present application can obtain the standard audio signal of each first target surgical instrument from the standard database based on the first instrument type and first operation type of the first target surgical instrument, and mix it to obtain the final analog audio signal.

[0196] It can be understood that directly obtaining a standard audio signal that matches the instrument type and operation type from a standard database based on the instrument type and operation type as the analog audio signal of the corresponding instrument has higher audio simulation efficiency. Compared with the embodiment of only obtaining a standard audio signal that matches the instrument type from a standard database based on the instrument type as the analog audio signal of the corresponding instrument, this embodiment can obtain an analog audio signal that is more consistent with the audio signal emitted by the first target surgical instrument under the current operation, thereby improving the accuracy and reliability of the analog audio signal.

[0197] In other embodiments, obtaining the simulated audio signal of the first target surgical instrument based on the first instrument type and the first operation type may include: The first instrument type and the first operation type are input into a preset neural network model, and a simulated audio signal is output.

[0198] The preset neural network model is trained in advance to output a simulated audio signal of the first target surgical instrument when a first instrument type and a first operation type of the first target surgical instrument are input.

[0199] Thus, the embodiment of the present application can directly utilize a neural network model to obtain a simulated audio signal of a first target surgical instrument based on a first instrument type and a first operation type, without the need to pre-build a standard database. Furthermore, the neural network model can be automatically trained and optimized. Compared to embodiments that pre-build a standard database and obtain the simulated audio signal of the first target surgical instrument from the standard database, this embodiment can dynamically generate a more suitable simulated audio signal based on the real-time requirements of the surgical scenario (such as instrument material, operation force, tissue type, etc.), providing greater flexibility. Furthermore, this embodiment can also incorporate an acoustic physics model to generate a simulated audio signal that is closer to the actual instrument operation, without being limited by the accuracy of the recording equipment or ambient noise, thereby obtaining a more accurate and reliable simulated audio signal. Furthermore, compared to embodiments that utilize a neural network model to obtain the simulated audio signal of the first target surgical instrument based on the first instrument type, the neural network model of this embodiment can generate a more accurate and reliable simulated audio signal based on the instrument type and operation type, which is more consistent with the audio signal generated by the surgical instrument during the current operation.

[0200] Furthermore, in some embodiments, the preset neural network model is a generative adversarial network or a diffusion model. For example, inputting the first device type and the first operation type into the preset neural network model and outputting a simulated audio signal may include: The first instrument type and the first operation type are input into a preset generative adversarial network or a preset diffusion model, and a simulated audio signal is output.

[0201] It is understood that a generative adversarial network consists of two neural networks: a generator and a discriminator. The generator is used to generate realistic data (such as audio) when input with random noise, while the discriminator is used to determine whether the input data is real or generated. The generator and discriminator are optimized through adversarial training, with the objective function being to minimize the Jensen-Shannon divergence between the generated distribution and the true distribution. Specifically, a first device type and a first operation type are input into a preset generative adversarial network, and the output simulated audio signal may include: A timing input matrix is ​​constructed based on the first instrument type and the first operation type, and the timing input matrix is ​​input into a generator to output an audio waveform. The audio waveform is then input into a discriminator to output an authenticity probability. When the authenticity probability is greater than a preset probability value, the audio waveform is determined to be an analog audio signal.

[0202] In addition, the diffusion model, based on the principle of thermodynamic diffusion, gradually adds Gaussian noise to the real data in a forward process (T steps), and gradually removes the noise in a reverse process based on neural network learning. Specifically, the first device type and the first operation type are input into the preset diffusion model, and the output analog audio signal may include: In the forward diffusion process, the ambient noise of the audio environment where the surgical platform on the patient side is located is used as the diffusion substrate, and forward diffusion is performed based on adaptive step-size scheduling; then, the noisy frequency, the first instrument type, and the first operation type obtained by the forward diffusion are input into the inverse denoising network to output an analog audio signal.

[0203] The step of inputting the first device type into the preset generative adversarial network or the preset diffusion model to output the analog audio signal is the same as the aforementioned step of inputting the first device type and the first operation type into the preset generative adversarial network or the preset diffusion model to output the analog audio signal, and will not be repeated here.

[0204] Therefore, the embodiments of the present application utilize a generative adversarial network or a diffusion model to output a high-fidelity analog audio signal, improve the effect of the final audio, and more completely feedback the sound of the surgical instrument in operation.

[0205] Furthermore, in some embodiments, the audio processing method further includes: Synchronize audio and video between the audio system and the endoscope.

[0206] Specifically, the audio signal acquired by the audio system and the surgical environment image acquired by the endoscope can be synchronized by using a precise time protocol, FPGA hardware synchronization, or timestamp alignment.

[0207] Therefore, the embodiment of the present application makes it possible to use the surgical environment images captured by the endoscope to calculate the motion parameters, posture recognition or action recognition of the surgical instrument, and the standard voiceprint characteristics or simulated instrument sounds obtained thereby are more accurate and reliable.

[0208] In some embodiments, the audio processing method on the patient-side surgical platform may further include: If all sounds of the surgical instruments in operation are acquired, a second instruction is generated, and the instrument audio signal and the second instruction are sent to the main console, so that the main console plays the instrument audio signal in response to the second instruction.

[0209] Therefore, when the embodiment of the present application determines that the sounds of the surgical instruments in the operating state have been fully acquired, a second instruction is generated and the second instruction and the instrument audio signal acquired by the audio system are sent to the main console, informing the main console that the sounds of the surgical instruments in the operating state have been fully acquired and are instrument audio signals. The main console responds to the second instruction and directly plays the instrument audio signal without the need for instrument sound simulation.

[0210] Alternatively, in some other embodiments, the audio processing method at the main control console may further include: If all sounds of the surgical instruments in operation are acquired, the instrument audio signals are played directly.

[0211] To sum up, the surgical robot and audio processing method thereof proposed in the embodiments of the present application, by obtaining the first instrument type of the first target surgical instrument in the surgical instrument, and obtaining the analog audio signal of the first target surgical instrument based on the first instrument type; or, by obtaining the first instrument type and the first operation type of the first target surgical instrument in the surgical instrument, and obtaining the analog audio signal of the first target surgical instrument based on the first instrument type and the first operation type, enable the surgical robot to fully feedback the instrument sound during the operation, avoid the impact of the lack of instrument sound on the doctor's operation accuracy, and improve the safety of remote surgery.

[0212] At the same time, the surgical robot and audio processing method thereof proposed in the embodiment of the present application can also obtain instrument sounds, determine whether instrument sounds are missing, and simulate missing instrument sounds, so that the remote surgical robot can fully feedback the instrument sounds in remote surgery while mixing the real collected instrument audio signals into the complete instrument audio signals, thereby improving the reliability of the complete audio signal, so that the doctor can obtain complete and more intuitive instrument sound feedback when performing surgery by controlling the surgical instruments on the patient-side surgical platform through the main console.

[0213] An embodiment of the present application further provides a computer-readable storage medium having instructions stored therein. When the instructions are executed on at least one processor, Figure 7-10The method shown. The storage medium can be volatile memory or non-volatile memory, or can include both volatile and non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface storage, optical disk, or compact disc read-only memory (CD-ROM); magnetic surface storage can be magnetic disk storage or magnetic tape storage. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memory.

[0214] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0215] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A surgical robot, characterized in that: include: Main console; a patient-side surgical platform controllable by the main console; The patient-side surgical platform includes surgical instruments for performing surgery, and the patient-side surgical platform or the main console is configured as follows: obtaining a first instrument type of a first target surgical instrument among the surgical instruments; An analog audio signal of the first target surgical instrument is obtained based on the first instrument type.

2. The surgical robot according to claim 1, wherein: The patient-side surgical platform or the main console is further configured to: Acquire a first operation type of the first target surgical instrument; The step of obtaining the simulated audio signal of the first target surgical instrument based on the first instrument type includes: A simulated audio signal of the first target surgical instrument is obtained based on the first instrument type and the first operation type.

3. The surgical robot according to claim 1, wherein: The patient-side surgical platform further includes an endoscope for acquiring images of the surgical environment; The acquiring of a first instrument type of a first target surgical instrument among the surgical instruments comprises: The first instrument type of the first target surgical instrument in the surgical environment image is identified based on a preset machine learning model.

4. The surgical robot according to claim 2, wherein: The patient-side surgical platform further includes an endoscope for acquiring images of the surgical environment; The acquiring of the first operation type of the first target surgical instrument comprises: The first operation type of the first target surgical instrument in the surgical environment image is identified based on a preset machine learning model.

5. The surgical robot according to any one of claims 1 to 4, characterized in that: The first target surgical instrument includes the surgical instrument in an operational state.

6. The surgical robot according to any one of claims 1 to 4, characterized in that: The patient-side surgical platform further includes an audio system for acquiring instrument sounds of the surgical instruments, wherein the first target surgical instruments include the surgical instruments in an operating state whose sounds have not been fully acquired; Before acquiring the first instrument type of the first target surgical instrument, the patient-side surgical platform or the main console is further configured to: determining whether all sounds of the surgical instrument in the operating state are acquired based on the instrument audio signal acquired by the audio system; If not, determine the first target surgical instrument.

7. The surgical robot according to claim 6, wherein: The determining whether all sounds of the surgical instrument in the operating state are acquired based on the instrument audio signal acquired by the audio system includes: Extracting voiceprint features of the audio signal of the instrument; calculating a matching degree between the voiceprint feature and a standard voiceprint feature of the surgical instrument in the operating state; If the matching degree is greater than a preset first matching degree threshold, all sounds of the surgical instrument in the operating state are acquired; otherwise, not all sounds of the surgical instrument in the operating state are acquired.

8. The surgical robot according to claim 7, wherein: The patient-side surgical platform or the main console is further configured to: obtaining a second instrument type of the surgical instrument in the operating state; The standard voiceprint feature is obtained based on the second device type.

9. The surgical robot according to claim 8, wherein: The obtaining of the standard voiceprint feature based on the second device type includes: Based on the second device type, obtaining the standard voiceprint feature from a preset standard database; or The second device type is input into a preset machine learning model to output the standard voiceprint feature.

10. The surgical robot according to claim 7, wherein: The patient-side surgical platform or the main console is further configured to: acquiring a second instrument type and a second operation type of the surgical instrument in the operating state; The standard voiceprint feature is obtained based on the second instrument type and the second operation type.

11. The surgical robot according to claim 10, wherein: The obtaining of the standard voiceprint feature based on the second device type and the second operation type includes: Based on the second device type and the second operation type, obtaining the standard voiceprint feature from a preset standard database; or The second device type and the second operation type are input into a preset machine learning model to output the standard voiceprint feature.

12. The surgical robot according to claim 6, wherein: The determining whether all sounds of the surgical instrument in the operating state are acquired based on the instrument audio signal acquired by the audio system includes: detecting a vibration signal of the surgical instrument in the operating state; Performing coherence analysis on the vibration signal and the instrument audio signal to obtain a coherence coefficient; If the coherence coefficient is greater than a preset coherence coefficient threshold, all sounds of the surgical instrument in the operating state are acquired; otherwise, not all sounds of the surgical instrument in the operating state are acquired.

13. The surgical robot according to claim 6, wherein: The determining whether all sounds of the surgical instrument in the operating state are acquired based on the instrument audio signal acquired by the audio system includes: Performing short-time Fourier transform on the instrument audio signal, and obtaining each independent audio signal in the instrument audio signal based on independent component analysis; If the number of the independent audio signals is less than the number of the surgical instruments in the operating state, then the sounds of the surgical instruments in the operating state are not all acquired; If the number of the independent audio signals is equal to the number of the surgical instruments in the operating state, the surgical instrument corresponding to each of the independent audio signals is identified based on a preset machine learning model, and when the matching degree between the independent voiceprint feature of each independent audio signal and the standard independent voiceprint feature of the corresponding surgical instrument is greater than a preset second matching degree threshold, it is confirmed that the sounds of the surgical instruments in the operating state are all acquired; otherwise, the sounds of the surgical instruments in the operating state are not all acquired.

14. The surgical robot according to claim 13, wherein: Determining the first target surgical instrument includes: determining, among the surgical instruments corresponding to the independent audio signal, the surgical instrument whose matching degree between the independent voiceprint feature and the standard independent voiceprint feature is greater than the second matching degree threshold as a second target surgical instrument; The first target surgical instrument is obtained based on the second target surgical instrument and the surgical instrument in the operating state.

15. The surgical robot according to claim 7, wherein: Determining the first target surgical instrument includes: The first target surgical instrument is determined based on the voiceprint feature and the standard voiceprint feature.

16. The surgical robot according to claim 1, wherein: The step of obtaining the simulated audio signal of the first target surgical instrument based on the first instrument type includes: Based on the first instrument type, obtaining a standard audio signal of the first target surgical instrument from a preset standard database as the simulated audio signal; or The first instrument type is input into a preset neural network model, and the simulated audio signal is output.

17. The surgical robot according to claim 2, wherein: The step of obtaining the simulated audio signal of the first target surgical instrument based on the first instrument type and the first operation type includes: Based on the first instrument type and the first operation type, obtaining a standard audio signal of the first target surgical instrument from a preset standard database as the simulated audio signal; or The first instrument type and the first operation type are input into a preset neural network model, and the simulated audio signal is output.

18. The surgical robot according to claim 16 or 17, wherein: The neural network model is a generative adversarial network or a diffusion model.

19. An audio processing method for a surgical robot, characterized in that: The method is applied to a patient-side surgical platform or a main console of a surgical robot, wherein the surgical robot includes the main console and the patient-side surgical platform controllable by the main console, and the patient-side surgical platform includes surgical instruments for performing surgery. The method includes: obtaining a first instrument type of a first target surgical instrument among the surgical instruments; An analog audio signal of the first target surgical instrument is obtained based on the first instrument type.

20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed on at least one processor, the method according to claim 19 is implemented.