Remote surgery delay compensation method and system and remote surgery robot
By using predictive models to generate predictive surgical videos in remote surgery and combining them with real-world scene videos, the problem of operational lag caused by network latency was solved, enabling real-time and reliable visual feedback and improving the continuity and safety of surgery.
Patent Information
- Application Number
- CN202610137779.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-15
AI Technical Summary
Network communication delays during remote surgery can prevent doctors from receiving immediate feedback on their actions, leading to a sense of lag and stagnation, increasing cognitive load and the risk of errors, and affecting surgical safety.
Predictive surgical videos are generated using a predictive model, and the display source is dynamically selected based on network latency to provide immediate and reliable visual feedback. The display is optimized by fusing real-world video and predicted video to reduce the impact of latency.
It improves the continuity and safety of surgery, reduces the operator's fatigue and cognitive load, and ensures the smoothness and reliability of the surgical procedure.
Smart Images

Figure CN122053787A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical technology, and more specifically, to a method, system, and remote surgical robot for delay compensation in remote surgery. Background Technology
[0002] With the continuous development and implementation of remote surgery and 5G technology, the traditional surgical model of operating in the same space has evolved into a remote control and operation model or a multi-control console remote control model. That is, the remote doctor's terminal and the local patient's terminal are deployed in two different hospitals, and the remote doctor's terminal and the local patient's terminal exchange data through a network (such as: wired network and / or 5G network, wired network and / or dedicated line single network or multi-network hybrid method).
[0003] Remote surgery has greatly expanded the scope of services available to surgeons, but its core bottleneck lies in network communication latency, especially high latency and fluctuating latency (jitter). This latency prevents surgeons from receiving immediate feedback, resulting in a noticeable lag and stuttering in their operations. This not only significantly increases the cognitive load and operational fatigue of surgeons but may also introduce the risk of errors during critical surgical procedures, directly endangering patient safety.
[0004] Existing technologies mainly reduce latency by optimizing network hardware and increasing bandwidth, but they cannot eliminate the inherent latency caused by network conditions. They still cannot provide doctors with smooth, timely and reliable visual feedback, which affects the surgical process. Summary of the Invention
[0005] The purpose of this application is to provide a latency compensation method, system, and remote surgical robot for remote surgery, which uses a predictive model as a "proactive agent" for the remote surgical robot to generate predictive surgical videos. Furthermore, the remote doctor dynamically selects the optimal display source (predictive surgical video or actual scene video) by judging network latency, providing the doctor with consistently smooth, immediate, and reliable visual feedback, thereby improving surgical continuity and safety.
[0006] In a first aspect, embodiments of this application provide a delay compensation method for remote surgery. This method is applied to a remote doctor's end and includes: acquiring operation control instructions, sending the operation control instructions to a prediction model, and synchronously sending the operation control instructions to a local patient's end via a network; wherein the prediction model is used to generate a predictive surgical video based on the operation control instructions; acquiring the current network latency between the remote doctor's end and the local patient's end; and if it is determined that the current network latency is greater than or equal to a first preset latency threshold, then outputting and displaying the predictive surgical video on the remote doctor's end.
[0007] In this embodiment, after obtaining the operation control command, a predictive surgical video is generated based on the operation control command using a predictive model. When the current network latency is greater than or equal to a first preset latency threshold, the predictive surgical video is output and displayed on the remote doctor's end. This creates a sense of immediacy for the doctor even in the presence of network latency, providing consistently smooth, immediate, and reliable visual feedback, thereby improving surgical continuity and safety.
[0008] In some embodiments, the method further includes: if it is determined that the current network latency is less than a first preset latency threshold, then outputting and displaying the actual scene video transmitted back from the local patient terminal via the network on the remote doctor terminal.
[0009] In this embodiment of the application, if the current network latency is less than the first preset latency threshold, it indicates that the network status is good and the remote doctor can receive the actual scene video sent by the local patient in a timely manner. At this time, the actual scene video is directly output and displayed on the remote doctor, thereby providing the doctor with real and reliable visual feedback, enhancing the doctor's confidence, and improving the continuity and safety of the surgery.
[0010] In some embodiments, the predictive surgical video includes a non-critical area predictive video; the method further includes: if it is determined that the current network latency is less than a first preset latency threshold, then splicing the actual critical area video transmitted back from the local patient terminal via the network with the non-critical area predictive video to obtain a spliced surgical video, and outputting and displaying the spliced surgical video on the remote doctor's terminal.
[0011] In this embodiment, the local patient terminal only transmits the actual video of the key areas during surgery, reducing bandwidth consumption during image transmission, thereby lowering image transmission latency and accelerating transmission speed. This approach is suitable for scenarios with poor network conditions and reduces network requirements. Furthermore, by stitching the actual video of the key areas with the predicted video of non-key areas to obtain the stitched surgical video for output display, low-latency and reliable visual feedback is consistently provided to the doctor in the key areas. This avoids the doctor's concerns about the accuracy of the model's predictions and improves the continuity and safety of the surgery.
[0012] In some embodiments, the predictive surgical video includes a confidence level; if it is determined that the current network latency is greater than or equal to a first preset latency threshold, the predictive surgical video is output and displayed on the remote doctor's end, including: if it is determined that the current network latency is greater than or equal to the first preset latency threshold and the confidence level is less than a preset confidence threshold, the predictive surgical video is degraded to generate a fuzzy predictive video, and the fuzzy predictive video is output and displayed on the remote doctor's end.
[0013] In this embodiment, when the confidence level of the predictive surgical video is less than a preset confidence threshold, it indicates that the accuracy of the prediction result is questionable. Therefore, the predictive surgical video is degraded to guide doctors to operate cautiously and slowly, thereby improving surgical safety.
[0014] In some embodiments, if it is determined that the current network latency is greater than or equal to a first preset latency threshold, the predictive surgical video is output and displayed on the remote doctor's end, including: receiving actual scene video transmitted back from the local patient's end via the network during the display of the predictive surgical video; correcting the predictive surgical video based on the actual scene video; and outputting and displaying the corrected surgical video.
[0015] In this embodiment, even with network latency, the local patient terminal continues to send real-scene videos to the remote doctor terminal. However, the remote doctor terminal receives the real-scene videos later than it receives the predictive surgical videos generated by the predictive model. Since the real-scene videos best reflect the actual surgical situation, the predictive surgical videos are corrected upon receipt to eliminate accumulated deviations caused by model errors or environmental disturbances, ensuring consistency in the final display and improving surgical safety.
[0016] In some embodiments, a predictive surgical video is corrected based on a real-scene video, and the corrected surgical video is output and displayed, including: fusing the real-scene video and the predictive surgical video to obtain a corrected surgical video, and outputting and displaying the corrected surgical video.
[0017] This application embodiment integrates real-scene videos and predictive surgical videos, enabling a smooth transition from predictive surgical videos to real-scene videos. This reduces visual discomfort caused by drastic screen flickering or stuttering during the transition from virtual to real video, alleviates visual fatigue for doctors, and improves surgical safety.
[0018] In some embodiments, the actual scene video includes real frames to be displayed; the predictive surgical video includes predicted frames to be displayed; fusing the actual scene video and the predictive surgical video to obtain a corrected surgical video includes: calculating a motion vector field between the predicted frames to be displayed and the real frames to be displayed; calculating a deformed predicted frame of the predicted frames to be displayed and a deformed real frame of the real frames to be displayed based on the motion vector field; wherein the deformed predicted frames and the deformed real frames are temporally aligned; and obtaining the corrected surgical video by weighted fusing the deformed predicted frames and the deformed real frames.
[0019] In this embodiment, by calculating the motion vector field between the predicted frame and the real frame, the deformed predicted frame and the deformed real frame are calculated based on the motion vector field, and weighted fusion is performed based on the deformed predicted frame and the deformed real frame to obtain the corrected surgical video. This makes the displayed image smoothly and without ghosting evolve from the predictive surgical video to the actual scene video, optimizing the doctor's user experience and improving surgical safety.
[0020] In some embodiments, after obtaining the operation control command, the method further includes: using the positive kinematics model of the instrument on the local patient end to determine the theoretical projection data of the instrument on the two-dimensional image plane of the endoscope based on the operation control command; wherein the prediction model is used to generate a predictive surgical video based on the operation control command and the theoretical projection data.
[0021] In this embodiment, the theoretical projection data of the instrument on the two-dimensional image plane of the endoscope is determined based on the operation control command by using the positive kinematic model of the instrument on the local patient end. This enables the prediction model to generate predictive surgical videos based on the operation control command and the theoretical projection data. This compensates for the error of the instrument motion trajectory based purely on theoretical calculation and the high computational load of predicting the motion trajectory based purely on operation control command, thereby improving the accuracy and efficiency of the instrument motion prediction trajectory and laying the foundation for the subsequent generation of predictive surgical videos.
[0022] In some embodiments, the method further includes: using a network latency prediction model to predict the network status between the remote doctor and the local patient within a preset time period from the current moment, and obtaining the predicted network latency; if it is determined that the predicted network latency is less than a second preset latency threshold, stopping the prediction model or reducing the running frequency of the prediction model, and outputting and displaying the actual scene video transmitted from the local patient via the network on the remote doctor within a preset time period; wherein the second preset latency threshold is less than the first preset latency threshold.
[0023] In this embodiment, by deploying the prediction model on the remote doctor's end, the remote doctor can obtain predictive surgical videos "without network latency," improving the timeliness and smoothness of the video display. Furthermore, the network latency prediction model predicts the network status between the remote doctor's end and the local patient's end within a preset time period from the current moment. When the predicted network latency is determined to be less than a second preset latency threshold, the prediction model stops running or its running frequency is reduced. This reduces the resource consumption of the prediction model and the predictive surgical videos on the remote doctor's end, freeing up computing power for the remote doctor's control algorithm and improving the remote doctor's response speed.
[0024] In some embodiments, the method further includes: adjusting the prediction model online using real-scene video transmitted from a local patient terminal.
[0025] In this embodiment, the prediction model is adjusted online using real-scene videos transmitted from the local patient terminal to optimize the model, thereby improving the accuracy of predictive surgical video generation. This provides doctors with consistently smooth, immediate, and reliable visual feedback, enhancing surgical safety.
[0026] In some embodiments, the prediction model is deployed on a remote physician's end; the prediction model includes an instrument motion prediction module and a biological tissue deformation prediction module; the prediction model is used to generate predictive surgical videos based on operation control instructions, including: using the instrument motion prediction module to generate an instrument motion prediction trajectory based on operation control instructions; and using the biological tissue deformation prediction module to generate a predictive surgical video based on the instrument motion prediction trajectory and historical surgical videos transmitted back from the local patient end before the current moment.
[0027] In this embodiment, an instrument motion prediction module generates a predicted instrument motion trajectory based on operation control commands to predict instrument motion at the local patient end. A biological tissue deformation prediction module generates a predictive surgical video based on the predicted instrument motion trajectory and historical surgical videos transmitted from the local patient end up to the current moment. This predicts the deformation of biological tissues during the interaction between the instrument and the tissue, thus obtaining a predictive surgical video. In this process, the complex prediction problem is decoupled through the division of labor and cooperation among the various modules in the prediction model, improving prediction accuracy. This ensures that even with network latency, doctors receive consistently smooth, immediate, and reliable visual feedback, enhancing surgical safety.
[0028] In some embodiments, the training process of the device motion prediction module is as follows: obtaining a first training sample; wherein, the first training sample includes historical operation control commands, historical theoretical projection sequences corresponding to the historical operation control commands, and historical actual motion trajectories; using the historical operation control commands and historical theoretical projection sequences as inputs and the historical actual motion trajectories as outputs, the initial device motion prediction module is trained to generate the device motion prediction module.
[0029] In this embodiment, considering that the device interacts with biological tissue during actual operation and may undergo slight nonlinear deformations or offsets under specific poses, resulting in a subtle difference between its final pose displayed in the 2D image and the theoretical projection, the training objective of the device motion prediction module is to enable it to learn this nonlinear correction from the theoretical projection to the actual video. Specifically, by using historical operation control commands and historical theoretical projection sequences as inputs and historical actual motion trajectories as outputs, the initial device motion prediction module is trained to obtain a new module. This allows the module to effectively master this correction mapping, improving its performance and thus increasing the accuracy of device motion trajectory prediction.
[0030] In some embodiments, the training process of the biological tissue deformation prediction module is as follows: obtaining a second training sample; wherein the second training sample includes historical surgical video frames and the actual movement trajectory of the instrument extracted from the historical surgical video frames; training the initial biological tissue deformation prediction module based on the historical surgical video frames and the actual movement trajectory of the instrument to generate the biological tissue deformation prediction module.
[0031] In this embodiment, since historical surgical videos realistically record the deformation, displacement, color, texture changes, and topological changes of tissues when surgical instruments interact with various biological tissues during various surgical operations, the biological tissue deformation prediction module is trained using historical surgical video frames and the real motion trajectories of instruments extracted from historical surgical video frames. This enables the biological tissue deformation prediction module to learn extremely complex, nonlinear biological tissue mechanical behaviors and their visual representations, thereby improving the performance of the biological tissue deformation prediction module and thus improving the accuracy of predictive surgical videos.
[0032] Secondly, embodiments of this application provide a delay compensation system for remote surgery, which is deployed on a remote doctor's end. The system includes: a first acquisition module, used to acquire operation control instructions and send the operation control instructions to a local patient end via a network; a prediction module, which includes a prediction model, wherein the first acquisition module also simultaneously sends the operation control instructions to the prediction model, and the prediction model is used to generate a predictive surgical video based on the operation control instructions; a second acquisition module, used to acquire the current network latency between the remote doctor's end and the local patient end; and a display module, used to output and display the predictive surgical video on the remote doctor's end if it is determined that the current network latency is greater than or equal to a first preset latency threshold.
[0033] Thirdly, embodiments of this application provide a remote surgical robot, including a remote doctor's end, a local patient's end, and a communication module; wherein the remote doctor's end and the local patient's end communicate through the communication module; the remote doctor's end is used to execute the method steps of any embodiment of the first aspect.
[0034] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing embodiments of this application. Attached Figure Description
[0035] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 A flowchart illustrating a first method for delay compensation in remote surgery provided in this application embodiment; Figure 2 A flowchart illustrating a second method for delay compensation in remote surgery provided in this application embodiment; Figure 3 This is a schematic diagram of the electronic device structure provided in an embodiment of this application. Detailed Implementation
[0037] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.
[0038] It should be noted that all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0039] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0040] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0041] With the continuous development of medical devices, computer technology, and control technology, minimally invasive surgery has been increasingly widely used due to its advantages such as small surgical trauma, short recovery time, and less patient suffering. Minimally invasive surgical robots, with their high dexterity, high control precision, and intuitive surgical images, can avoid operational limitations, such as filtering hand tremors during operation, and are widely applicable to surgical areas such as the abdominal cavity, pelvic cavity, and thoracic cavity.
[0042] Currently, the largest category of minimally invasive surgical robots is the laparoscopic surgical robot. This type of robot typically includes a surgeon control platform (also called a surgeon's carriage, master control carriage, surgeon's console, master end, etc.) and a patient surgical platform (also called a patient carriage, surgical carriage, slave end, etc.). The surgeon's carriage is equipped with a master control arm (also called a master controller, master operator / manipulator, etc.) for the operator (e.g., the surgeon) to operate, and a stereoscopic monitor for the operator to observe the surgical scene. The surgical arm (also called a robotic arm, slave manipulator, etc.) of the patient carriage is detachably equipped with surgical instruments (also called instruments, surgical tools, slave tools, etc.) for performing actions and an endoscope for acquiring surgical images. A master-slave mapping relationship (coordinate system transformation relationship, converting the motion of the master end into the motion of the slave end) is established between the master control arm in the surgeon's carriage and the surgical instruments in the patient carriage through a control device, enabling remote motion control / teleoperation of the surgical instruments by the master control arm.
[0043] With the continuous development and implementation of remote surgery and 5G technology, the traditional surgical model of operating in the same space has evolved into a remote control and operation model or a multi-control console remote control model. That is, the remote doctor's terminal and the local patient's terminal are deployed in two different hospitals, and the remote doctor's terminal and the local patient's terminal exchange data through a network (such as: wired network and / or 5G network, wired network and / or dedicated line single network or multi-network hybrid method).
[0044] Remote surgery has greatly expanded the scope of services available to surgeons, but its core bottleneck lies in network communication latency, especially high latency and fluctuating latency (jitter). This latency prevents surgeons from receiving immediate feedback, resulting in a noticeable lag and stuttering in their operations. This not only significantly increases the cognitive load and operational fatigue of surgeons but may also introduce the risk of errors during critical surgical procedures, directly endangering patient safety.
[0045] Existing technologies mainly reduce latency by optimizing network hardware and increasing bandwidth, but they cannot eliminate the inherent latency caused by network conditions. They still cannot provide doctors with smooth, timely and reliable visual feedback, which affects the surgical process.
[0046] To address this issue, this application provides a latency compensation method for remote surgery. As a latency compensation scheme based on a data-driven prediction model and dynamic latency judgment, the prediction model acts as a "proactive agent" for the remote surgical robot, generating predictive surgical videos. Furthermore, the remote doctor dynamically selects the optimal display source (predictive surgical video or actual scene video) by judging network latency, providing the doctor with consistently smooth, immediate, and reliable visual feedback, thus improving surgical continuity and safety.
[0047] Figure 1 This is a flowchart illustrating a first method for delay compensation in remote surgery provided in an embodiment of this application. This method is applied to a remote doctor's end. Figure 1 As shown, the method includes: Step S101: Obtain operation control instructions, send the operation control instructions to the prediction model, and send the operation control instructions to the local patient terminal via network synchronization; wherein, the prediction model is used to generate predictive surgical videos based on the operation control instructions.
[0048] Step S102: Obtain the current network latency between the remote doctor's terminal and the local patient's terminal; Step S103: If it is determined that the current network latency is greater than or equal to the first preset latency threshold, the predictive surgical video is output and displayed on the remote doctor's end.
[0049] In the above implementation process, the prediction model can be deployed on a remote doctor's end or in the cloud that communicates with the remote doctor's end. The prediction model can be based on a convolutional long short-term memory network or a spatiotemporal generative adversarial network.
[0050] It is important to note that since predictive surgical videos are the display source when the actual scene video cannot be transmitted back in time from the local patient's end, when the predictive model is deployed in the cloud, it is necessary to ensure that the network latency between the remote doctor's end and the cloud is less than the network latency between the remote doctor's end and the local patient's end.
[0051] The doctor generates control commands via a remote control console. The remote control console then sends these commands to a predictive model, which generates predictive surgical videos based on the commands. Furthermore, the remote control console simultaneously transmits these commands to the local patient's device via the network.
[0052] The remote doctor's terminal monitors the network status between itself and the local patient's terminal in real time to obtain the current network latency between them. When the predictive model is deployed in the cloud, the remote doctor's terminal also needs to monitor the network status between itself and the cloud in real time.
[0053] If the current network latency is determined to be greater than or equal to the first preset latency threshold, it indicates that the current network condition is poor, and the remote doctor cannot receive the actual scene video transmitted back from the local patient in a timely manner, resulting in the remote doctor's display device being unable to display the surgical situation of the local patient in a timely manner.
[0054] To ensure the surgery can proceed smoothly, the predictive surgical video output by the predictive model is displayed. Regardless of network fluctuations between the remote doctor and the local patient, the doctor on the remote end always sees a smooth, responsive operating screen, enabling them to perform remote surgery based on this screen.
[0055] In practical implementation, network latency can be obtained in the following ways: (1) Directly read the built-in data of the programmable logic controller (PLC): Utilize the distributed clock synchronization mechanism of real-time Ethernet protocols (such as EtherCAT) to directly obtain the accurate link delay through the specific data interface of the PLC (such as Beckhoff ADS), which is suitable for real-time networks supported by the protocol.
[0056] (2) Actively send probe sequences: By deploying probes at key nodes, high-priority test data packets are periodically sent and their round-trip time is calculated, thereby accurately measuring real-time indicators such as latency, jitter and packet loss rate.
[0057] The first preset latency threshold can be determined in advance based on the historical network conditions between the remote doctor's end and the local patient's end. For example, the first preset latency threshold can be set to a value between 50 and 100 milliseconds (ms). Of course, the specific value can also be set according to actual needs or scenarios, and this application does not impose specific limitations on it.
[0058] The first preset delay threshold can also be configured manually or automatically by the remote doctor based on the real-time requirements of the surgical type. Specifically, if the surgical type has high real-time requirements, the first preset delay threshold is set to a smaller value; if the surgical type has moderate real-time requirements, the first preset delay threshold can be set to a larger value that meets the basic requirements of the surgery. For example, when performing complex surgeries with many nerves and high requirements, such as radical prostatectomy, the first preset delay threshold can be set to a smaller value, such as 30ms or 50ms; when performing simple surgeries such as partial hepatectomy, the first preset delay threshold can be set to a larger value, such as 80ms or 100ms.
[0059] In this embodiment, after obtaining the operation control command, a predictive surgical video is generated based on the operation control command using a predictive model. When the current network latency is greater than or equal to a first preset latency threshold, the predictive surgical video is output and displayed on the remote doctor's end. This creates a sense of immediacy for the doctor even in the presence of network latency, providing consistently smooth, immediate, and reliable visual feedback, thereby improving surgical continuity and safety.
[0060] In some embodiments, the method further includes: if it is determined that the current network latency is less than a first preset latency threshold, then outputting and displaying the actual scene video transmitted back from the local patient terminal via the network on the remote doctor terminal.
[0061] In the above implementation process, if the current network latency is less than the first preset latency threshold, it indicates that the network status is good and the remote doctor can receive the actual scene video sent by the local patient in a timely manner. At this time, the actual scene video is directly output and displayed on the remote doctor, thereby providing doctors with real and reliable visual feedback, enhancing doctors' confidence, and improving the continuity and safety of surgery.
[0062] In some embodiments, the predictive surgical video includes a non-critical area predictive video; the method further includes: if it is determined that the current network latency is less than a first preset latency threshold, then splicing the actual critical area video transmitted back from the local patient terminal via the network with the non-critical area predictive video to obtain a spliced surgical video, and outputting and displaying the spliced surgical video on the remote doctor's terminal.
[0063] In the above implementation process, considering that the complete video frames of the actual scene image transmitted back by the local patient terminal occupy a lot of bandwidth, in order to reduce the bandwidth occupation during the image transmission process and reduce the image transmission delay, the local patient terminal can transmit only the key areas of the actual scene image.
[0064] Specifically, the local patient end uses an endoscope to capture the original surgical video in real time. Then, a pre-trained deep learning segmentation model (such as a U-Net architecture model) is used to perform real-time semantic segmentation on each frame of the captured original surgical video, thereby dividing the original surgical video into key regions and non-key regions.
[0065] Critical areas typically refer to the focal areas that directly affect the safety and precision of the surgical procedure, such as the tip of the surgical instrument, major anatomical structures (e.g., blood vessels, nerves, tumor boundaries), and bleeding points. Non-critical areas refer to the background tissue of the surgical field, already treated areas, and instrument handles, which are not focal areas.
[0066] After receiving the actual video of the key area from the local patient, the remote doctor performs pixel-level alignment and fusion stitching with the predicted video of the non-key area to obtain the stitched surgical video, and then outputs and displays the stitched surgical video on the remote doctor's end.
[0067] Specifically, the process of pixel-level alignment and fusion of the actual video of the key area and the predicted video of the non-key area is as follows: (1) Timestamp alignment and coordinate system one: Each frame of the actual video of the key region carries a high-precision timestamp. Based on this timestamp, the prediction video of the non-key region that best matches the time is determined from the predictive surgical video, and the two are converted to the same coordinate system.
[0068] (2) Dynamic geometric alignment: In the boundary region between two adjacent frames, feature points that can represent the same anatomical structure are detected (e.g., using ORB or SIFT algorithms). Using the successfully matched feature point pairs, a homography transformation matrix is estimated using the RANSAC algorithm. This transformation matrix is then applied to perform perspective transformation on the entire non-critical region prediction video, making it geometrically aligned precisely with the actual video of the critical region.
[0069] (3) Color consistency correction: In overlapping or adjacent areas, calculate the color mean and standard deviation of the actual video in the key area and the aligned predicted video in the non-key area. For each color channel of the predicted video in the non-key area, apply a linear transformation based on statistics to map its color distribution to match the actual video in the key area.
[0070] (4) Multi-resolution smooth fusion: Create a weight mask of the same size as the complete frame to generate a fused weight map. The weight is 1 within key regions and 0 within non-key regions. Within boundary bands (e.g., 15 pixels wide), the weight transitions smoothly from 1 to 0. A Laplacian pyramid and a Gaussian pyramid are constructed for the actual video of the key regions, the corrected and aligned predicted video of the non-key regions, and the weight map, respectively. At each layer of the pyramid, the Laplacian coefficients of the actual video of the key regions and the corrected and aligned predicted video of the non-key regions are weighted and fused using the corresponding Gaussian weight map. Starting from the top layer, the fused Laplacian pyramid is reconstructed layer by layer upwards to obtain the final seamless fused image.
[0071] (5) Post-processing and output: Perform necessary post-processing on the fused image (such as selective sharpening and noise reduction), and output it as a complete frame of synthetic surgical video on the display device of the remote doctor.
[0072] In this embodiment, the local patient terminal only transmits the actual video of the key areas during surgery, reducing bandwidth consumption during image transmission, thereby lowering image transmission latency and accelerating transmission speed. This approach is suitable for scenarios with poor network conditions and reduces network requirements. Furthermore, by stitching the actual video of the key areas with the predicted video of non-key areas to obtain the stitched surgical video for output display, low-latency and reliable visual feedback is consistently provided to the doctor in the key areas. This avoids the doctor's concerns about the accuracy of the model's predictions and improves the continuity and safety of the surgery.
[0073] In some embodiments, the predictive surgical video includes a confidence level; if it is determined that the current network latency is greater than or equal to a first preset latency threshold, the predictive surgical video is output and displayed on the remote doctor's end, including: if it is determined that the current network latency is greater than or equal to the first preset latency threshold and the confidence level is less than a preset confidence threshold, the predictive surgical video is degraded to generate a fuzzy predictive video, and the fuzzy predictive video is output and displayed on the remote doctor's end.
[0074] In the above implementation process, considering that the predictive ability of the prediction model is limited and the surgical process is complex and variable, the remote doctor initiates an uncertainty assessment mechanism when deciding to display the predictive surgical video. This uncertainty assessment mechanism determines whether the predictive surgical video needs to be degraded in order to guide the doctor to operate cautiously.
[0075] Specifically, when the confidence level of a predictive surgical video is less than a preset confidence threshold, the predictive surgical video is blurred before being displayed to remind the doctor to stop or proceed with caution.
[0076] In one implementation, the lower the confidence level of the predictive surgical video, i.e., the greater the difference between the confidence level and the preset confidence threshold, the higher the degree of blurring.
[0077] After quality degradation processing, before receiving the actual scene video returned from the local patient, if the confidence level of the predictive surgical video output by the prediction model in real time is always less than the preset confidence threshold, the remote doctor will directly display the actual scene video after receiving it from the local patient. If the confidence level of the predictive surgical video output by the prediction model in real time is not less than the preset confidence threshold at any point, the predictive surgical video will be displayed. Furthermore, after receiving the actual scene video, the predictive surgical video will be corrected using the actual scene video before the corrected predictive surgical video will be displayed.
[0078] In this embodiment, when the confidence level of the predictive surgical video is less than a preset confidence threshold, it indicates that the accuracy of the prediction result is questionable. Therefore, the predictive surgical video is degraded to guide doctors to operate cautiously and slowly, thereby improving surgical safety.
[0079] In some embodiments, if it is determined that the current network latency is greater than or equal to a first preset latency threshold, the predictive surgical video is output and displayed on the remote doctor's end, including: receiving actual scene video transmitted back from the local patient's end via the network during the display of the predictive surgical video; correcting the predictive surgical video based on the actual scene video; and outputting and displaying the corrected surgical video.
[0080] In the above implementation process, even with network latency, the remote doctor's end still receives the actual scene video transmitted back from the local patient's end; however, the time of receiving the actual scene video is later than the time of receiving the predictive surgical video generated by the predictive model. Since the actual scene video reflects the actual surgical situation, if the actual scene video is received during the display of the predictive surgical video, it is used to correct the predictive surgical video.
[0081] Specifically, the predictive surgical video can be directly switched to a real-world scene video, or a linear transition can be used to switch it to a real-world scene video, thereby correcting the predictive surgical video. This eliminates the cumulative deviation caused by model errors or environmental disturbances in the predictive surgical video, ensuring the consistency of the final display and improving surgical safety.
[0082] In some embodiments, a predictive surgical video is corrected based on a real-scene video, and the corrected surgical video is output and displayed, including: fusing the real-scene video and the predictive surgical video to obtain a corrected surgical video, and outputting and displaying the corrected surgical video.
[0083] In the above implementation process, considering that direct switching may cause abrupt changes in the image, and linear gradual frame interpolation may produce ghosting, both of which will affect the doctor's visual experience, a corrected surgical video is obtained by fusing the actual scene video and the predictive surgical video.
[0084] The core purpose of fusion is to generate a series of intermediate transition frames between the predicted frame and the real frame within a continuous multi-frame display cycle, so that the picture smoothly and without ghosting evolves from the predicted state to the real state.
[0085] In some embodiments, the actual scene video includes real frames to be displayed; the predictive surgical video includes predicted frames to be displayed; fusing the actual scene video and the predictive surgical video to obtain a corrected surgical video includes: calculating a motion vector field between the predicted frames to be displayed and the real frames to be displayed; calculating a deformed predicted frame of the predicted frames to be displayed and a deformed real frame of the real frames to be displayed based on the motion vector field; wherein the deformed predicted frames and the deformed real frames are temporally aligned; and obtaining the corrected surgical video by weighted fusing the deformed predicted frames and the deformed real frames.
[0086] The specific process of integration is as follows: Step 1: Calculation of Motion Vector Field Using classic computer vision techniques such as optical flow or block matching, the real frames to be displayed are... And the prediction frame being displayed Motion estimation is performed. A dense motion vector field is calculated. Among them, each vector in the field This indicates the prediction frame being displayed. medium pixel To the actual frame to be displayed The displacement (including direction and magnitude) of the corresponding pixel.
[0087] Step 2: Temporal Interpolation and Frame Generation Assume the motion of the object is linear over extremely short time intervals. Let's assume the prediction frame being displayed... Transition to the actual frame to be displayed The entire transition process requires Frame (e.g.) (Currently in the transitional phase) frame( From 1 to ), for the transition of the first Frame, based on motion vector field and current time weight Perform the following operations: (1) Forward deformation: according to Calculate an intermediate motion vector field for the currently displayed prediction frame. Each pixel in the image is extrapolated forward to a temporary intermediate position to generate a deformed prediction frame. ).
[0088] (2) Reverse deformation: According to Calculate another intermediate motion vector field (direction and) Conversely, the actual frame to be displayed Each pixel in the image is traced back in reverse to the predicted frame being displayed. At the same temporary intermediate location, a deformed real frame is generated. .
[0089] in, This indicates an image deformation operation.
[0090] Step 3: Adaptive Weighted Fusion The two intermediate images that are already aligned on the motion trajectory are merged pixel by pixel to generate the final displayed image. Frame transition image :
[0091] in, It is a time Increasing fusion weight coefficients (e.g.) At the same time, weight It can also be fine-tuned based on the prediction uncertainty of the pixel area, biasing towards the real frame information earlier in areas with high uncertainty.
[0092] In this embodiment, by calculating the motion vector field between the predicted frame and the real frame, the deformed predicted frame and the deformed real frame are calculated based on the motion vector field, and weighted fusion is performed based on the deformed predicted frame and the deformed real frame to obtain the corrected surgical video. This makes the displayed image smoothly and without ghosting evolve from the predictive surgical video to the actual scene video, optimizing the doctor's user experience and improving surgical safety.
[0093] In some embodiments, after obtaining the operation control command, the method further includes: using the positive kinematics model of the instrument on the local patient end to determine the theoretical projection data of the instrument on the two-dimensional image plane of the endoscope based on the operation control command; wherein the prediction model is used to generate a predictive surgical video based on the operation control command and the theoretical projection data.
[0094] Theoretical projection data is a calculated trajectory of instrument movement. Therefore, determining the theoretical projection data of the instrument on the two-dimensional image plane of the endoscope based on operation control commands is also a form of prediction. Due to various errors, this theoretical projection data differs from the actual trajectory of the instrument. Therefore, the prediction model generates predictive surgical videos based on operation control commands and theoretical projection data to bridge the gap between the theoretical projection data and the actual trajectory.
[0095] In the above implementation process, for the operation control command, the theoretical position and orientation of the device end in three-dimensional space are quickly calculated using the forward kinematics model of the device on the local patient end. The theoretical position and orientation of the device end in three-dimensional space are then projected onto a two-dimensional image plane using a camera model to obtain the theoretical projection data corresponding to the operation control command.
[0096] The operation control commands and theoretical projection data are then input into the prediction model to generate predictive surgical videos.
[0097] In this embodiment, the theoretical projection data of the instrument on the two-dimensional image plane of the endoscope is determined based on the operation control command by using the positive kinematic model of the instrument on the local patient end. This enables the prediction model to generate predictive surgical videos based on the operation control command and the theoretical projection data. This compensates for the error of the instrument motion trajectory based purely on theoretical calculation and the high computational load of predicting the motion trajectory based purely on operation control command, thereby improving the accuracy and efficiency of the instrument motion prediction trajectory and laying the foundation for the subsequent generation of predictive surgical videos.
[0098] In some embodiments, the method further includes: using a network latency prediction model to predict the network status between the remote doctor and the local patient within a preset time period from the current moment, and obtaining the predicted network latency; if it is determined that the predicted network latency is less than a second preset latency threshold, stopping the prediction model or reducing the running frequency of the prediction model, and outputting and displaying the actual scene video transmitted from the local patient via the network on the remote doctor within a preset time period; wherein the second preset latency threshold is less than the first preset latency threshold.
[0099] In the above implementation process, it is considered that regardless of network conditions, the prediction model will execute the prediction action and generate predictive surgical videos as soon as it receives operation control commands. Although predictive surgical video frames can be obtained in real time, it will bring huge computational load and resource consumption on the remote doctor's end. For example, if 60 frames of images are displayed per second and the duration of a surgery is 2 hours, then 430,000 frames need to be predicted.
[0100] To address the resource consumption issue on the remote doctor's end, a network latency prediction model is used to predict the network status between the remote doctor's end and the local patient's end within a preset time period from the current moment. Based on the predicted network status, it is determined whether the prediction model should continue to operate.
[0101] Among these, network latency prediction models can be implemented using deep learning techniques. For example, Long Short-Term Memory (LSTM) networks can be used to analyze historical time-series data of network performance to predict future network conditions (such as possible future latency fluctuations), providing a basis for remote doctors to make adjustments in advance.
[0102] Specifically, if the predicted network latency is determined to be less than the second preset latency threshold, the prediction model will stop running, and within a preset time period in the future, the actual scene video transmitted from the local patient terminal via the network will be output and displayed on the remote doctor's terminal.
[0103] In one implementation, if the prediction network latency is determined to be less than a first preset latency threshold, the running frequency of the prediction model can be reduced, for example, by changing the real-time prediction of the prediction model to start the prediction model to perform prediction actions at intervals.
[0104] In this process, since the second preset delay threshold is used to determine whether the prediction model is enabled and the frequency of the prediction model's operation, for safety reasons, the second preset delay threshold is lower than the first preset delay threshold, so that the prediction model can provide predictive surgical videos to the remote doctor in a timely manner before the network conditions deteriorate.
[0105] In this embodiment, by deploying the prediction model on the remote doctor's end, the remote doctor can obtain predictive surgical videos "without network latency," improving the timeliness and smoothness of the video display. Furthermore, the network latency prediction model predicts the network status between the remote doctor's end and the local patient's end within a preset time period from the current moment. When the predicted network latency is determined to be less than a second preset latency threshold, the prediction model stops running or its running frequency is reduced. This reduces the resource consumption of the prediction model and the predictive surgical videos on the remote doctor's end, freeing up computing power for the remote doctor's control algorithm and improving the remote doctor's response speed.
[0106] In some embodiments, the method further includes: adjusting the prediction model online using real-scene video transmitted from a local patient terminal.
[0107] In the above implementation process, the remote doctor can fine-tune the prediction model online based on the differences between the actual scene video returned by the local patient and the predictive surgery video, so as to improve the performance of the prediction model.
[0108] In this embodiment, the prediction model is adjusted online using real-scene videos transmitted from the local patient terminal to optimize the model, thereby improving the accuracy of predictive surgical video generation. This provides doctors with consistently smooth, immediate, and reliable visual feedback, enhancing surgical safety.
[0109] In some embodiments, the prediction model is deployed on a remote physician's end; the prediction model includes an instrument motion prediction module and a biological tissue deformation prediction module; the prediction model is used to generate predictive surgical videos based on operation control instructions, including: using the instrument motion prediction module to generate an instrument motion prediction trajectory based on operation control instructions; and using the biological tissue deformation prediction module to generate a predictive surgical video based on the instrument motion prediction trajectory and historical surgical videos transmitted back from the local patient end before the current moment.
[0110] In the above implementation process, the prediction model adopts a step-by-step model architecture, which includes an instrument motion prediction module and a biological tissue deformation prediction module.
[0111] The device motion prediction module can be implemented using a convolutional neural network. Based on operation control commands, the device motion prediction module generates a predicted device motion trajectory, which is the future position of the device on the local patient end on a two-dimensional image (the image displayed on the remote doctor's display device).
[0112] The biological tissue deformation prediction module can be implemented using a spatiotemporal prediction neural network. This module receives the instrument motion prediction trajectory output by the instrument motion prediction module and the historical surgical video transmitted from the local patient terminal up to the current moment, inferring the complete scene changes that occur in the tissue and obtaining a predictive surgical video.
[0113] Among them, the historical surgical video transmitted before the current moment refers to the surgical video of the instruments being operated on the local patient end before the current moment during this surgery.
[0114] In this embodiment, an instrument motion prediction module generates a predicted instrument motion trajectory based on operation control commands to predict instrument motion at the local patient end. A biological tissue deformation prediction module generates a predictive surgical video based on the predicted instrument motion trajectory and historical surgical videos transmitted from the local patient end up to the current moment. This predicts the deformation of biological tissues during the interaction between the instrument and the tissue, thus obtaining a predictive surgical video. In this process, the complex prediction problem is decoupled through the division of labor and cooperation among the various modules in the prediction model, improving prediction accuracy. This ensures that even with network latency, doctors receive consistently smooth, immediate, and reliable visual feedback, enhancing surgical safety.
[0115] In some embodiments, the training process of the device motion prediction module is as follows: obtaining a first training sample; wherein, the first training sample includes historical operation control commands, historical theoretical projection sequences corresponding to the historical operation control commands, and historical actual motion trajectories; using the historical operation control commands and historical theoretical projection sequences as inputs and the historical actual motion trajectories as outputs, the initial device motion prediction module is trained to generate the device motion prediction module.
[0116] In some embodiments, the training process of the biological tissue deformation prediction module is as follows: obtaining a second training sample; wherein the second training sample includes historical surgical video frames and the actual movement trajectory of the instrument extracted from the historical surgical video frames; training the initial biological tissue deformation prediction module based on the historical surgical video frames and the actual movement trajectory of the instrument to generate the biological tissue deformation prediction module.
[0117] In the above implementation process, since the core of the prediction model is a data-driven model, and the predictive visual information includes instrument movement and tissue deformation, the training data of the prediction model comes from multiple aspects to solve the two different problems of instrument movement prediction and tissue deformation prediction respectively.
[0118] For training the instrument motion prediction module, initial sample data was collected using time-precisely aligned doctor's operation command streams, robot joint parameters, and synchronized surgical video streams. The initial sample data was preprocessed to obtain historical operation control commands, corresponding historical theoretical projection sequences, and historical actual motion trajectories.
[0119] Specifically, historical operational control commands are determined based on the doctor's instruction stream. For these historical commands, the theoretical position and orientation of the instrument's end effector in three-dimensional space are quickly calculated using the local patient-side instrument's forward kinematics model. This theoretical position and orientation are then projected onto a two-dimensional image plane using a camera model, yielding historical theoretical projection data corresponding to the historical operational control commands. Finally, the historical actual motion trajectory corresponding to the historical operational control commands is determined based on the robot's joint parameters and the synchronized surgical video stream.
[0120] Considering that instruments interact with biological tissues during actual operation and may undergo slight nonlinear deformations or shifts under specific poses, resulting in subtle differences between their final pose displayed in a 2D image and the theoretical projection, the training objective for the instrument motion prediction module is to enable it to learn this nonlinear correction from the theoretical projection to the actual video.
[0121] Specifically, the initial motion prediction module is trained by taking historical operation control commands and their corresponding historical theoretical projection sequences as inputs, and the actual historical motion trajectory (i.e., the actual pixel-level position and attitude of the device in the video) as the target output. This process allows the initial motion prediction module to effectively master this correction mapping, improving its performance and thus increasing the accuracy of the device's motion trajectory prediction.
[0122] For training the biological tissue deformation prediction module, a massive, anonymized library of existing robotic surgical videos was used as initial training samples. These video sequences realistically record the deformation, displacement, color, texture changes, and topological changes of tissues during various surgical procedures as surgical instruments interact with different biological tissues. The initial sample data was preprocessed to obtain historical surgical video frames. Furthermore, the actual movement trajectories of the instruments in these historical surgical video frames were extracted.
[0123] Because the biological tissue deformation prediction module needs to learn extremely complex, nonlinear biological tissue mechanical behavior and its visual representation, historical surgical video frames and the actual movement trajectories of instruments are used as inputs to train the module. Through training on massive amounts of data, the module learns how surrounding tissues will undergo reasonable, physically-compliant deformation when instruments move in specific ways on the local patient end. This improves the performance of the biological tissue deformation prediction module, thereby enhancing the accuracy of predictive surgical videos.
[0124] Figure 2 This is a flowchart illustrating a second method for delay compensation in remote surgery provided in an embodiment of this application. Figure 2 As can be seen, after the remote surgery begins, the doctor generates operation control instructions through the remote physician terminal and simultaneously sends these instructions to the local patient terminal and the prediction model. Upon receiving the operation control instructions, the local patient terminal performs the surgery and captures video of the actual scene, which is then sent back to the remote physician terminal. Upon receiving the operation control instructions, the prediction model generates a predictive surgical video.
[0125] To further improve surgical safety, after the predictive model generates predictive surgical videos, an uncertainty assessment is performed on the predictive surgical videos. If the confidence level of the predictive surgical video is determined to be less than a preset confidence threshold, the predictive surgical video is then processed to generate a fuzzy predictive video.
[0126] To select the optimal display source and provide doctors with consistently smooth, immediate, and reliable visual feedback, the remote doctor estimates whether the current network latency between the remote doctor and the local patient is less than a first preset latency threshold. If it is less than the first preset latency threshold, the actual scene video transmitted back from the local patient is displayed. If it is greater than or equal to the first preset latency threshold, the predictive video (predictive surgical video or fuzzy predictive video) output by the predictive model is displayed.
[0127] Even with network latency, the local patient's device is transmitting real-scene video back to the remote doctor, who then uses this video to correct the predictive video.
[0128] In addition, the prediction model is adjusted online using real-world video scenarios.
[0129] This application embodiment also provides a delay compensation system for remote surgery, which is deployed on a remote doctor's end. The system includes: a first acquisition module, used to acquire operation control instructions and send the operation control instructions to a local patient end via a network; a prediction module, which includes a prediction model, wherein the first acquisition module also simultaneously sends the operation control instructions to the prediction model, and the prediction model is used to generate a predictive surgical video based on the operation control instructions; a second acquisition module, used to acquire the current network latency between the remote doctor's end and the local patient end; and a display module, used to output and display the predictive surgical video on the remote doctor's end if it is determined that the current network latency is greater than or equal to a first preset latency threshold.
[0130] Based on the above embodiments, the display module is specifically used to: if it is determined that the current network latency is less than a first preset latency threshold, output and display the actual scene video transmitted back from the local patient terminal via the network on the remote doctor terminal.
[0131] Based on the above embodiments, the predictive surgical video includes a non-critical area predictive video; the display module is specifically used to: if it is determined that the current network latency is less than a first preset latency threshold, then the actual video of the critical area transmitted back from the local patient terminal via the network is spliced with the non-critical area predictive video to obtain the spliced surgical video, and the spliced surgical video is output and displayed on the remote doctor's terminal.
[0132] Based on the above embodiments, the predictive surgical video includes a confidence level; the display module is specifically used to: if it is determined that the current network latency is greater than or equal to a first preset latency threshold and the confidence level is less than a preset confidence threshold, then the predictive surgical video is degraded to generate a fuzzy predictive video, and the fuzzy predictive video is output and displayed on the remote doctor's end.
[0133] Based on the above embodiments, the display module is specifically used to: receive actual scene video transmitted back from the local patient terminal via the network during the process of displaying predictive surgical video; correct the predictive surgical video based on the actual scene video; and output and display the corrected surgical video.
[0134] Based on the above embodiments, the display module is specifically used to: fuse the actual scene video and the predictive surgical video to obtain the corrected surgical video, and output and display the corrected surgical video.
[0135] Based on the above embodiments, the actual scene video includes real frames to be displayed; the predictive surgical video includes predictive frames currently being displayed; the display module is specifically used to: calculate the motion vector field between the predictive frames currently being displayed and the real frames to be displayed; based on the motion vector field, calculate the deformed predictive frames of the predictive frames currently being displayed, and calculate the deformed real frames of the real frames to be displayed; wherein the deformed predictive frames and the deformed real frames are aligned in time; and by weighted fusion of the deformed predictive frames and the deformed real frames, the corrected surgical video is obtained.
[0136] Based on the above embodiments, the first acquisition module is specifically used to: determine the theoretical projection data of the instrument on the two-dimensional image plane of the endoscope based on the operation control command using the positive kinematic model of the instrument on the local patient end; wherein, the prediction model is used to generate a predictive surgical video based on the operation control command and the theoretical projection data.
[0137] Based on the above embodiments, the system further includes a network prediction module, which is used to predict the network status between the remote doctor and the local patient within a preset time period from the current moment using a network latency prediction model, and obtain the predicted network latency; if it is determined that the predicted network latency is less than a second preset latency threshold, the prediction model is stopped or the running frequency of the prediction model is reduced, and the actual scene video transmitted from the local patient via the network is output and displayed on the remote doctor within a preset time period; wherein, the second preset latency threshold is less than the first preset latency threshold.
[0138] Based on the above embodiments, the system also includes an online adjustment module, which is used to adjust the prediction model online using actual scene videos transmitted back from the local patient terminal.
[0139] Based on the above embodiments, the prediction model is deployed on the remote doctor's end; the prediction model includes an instrument motion prediction module and a biological tissue deformation prediction module; the prediction module is specifically used to: generate an instrument motion prediction trajectory based on operation control commands using the instrument motion prediction module; and generate a predictive surgical video based on the instrument motion prediction trajectory and the historical surgical video transmitted back from the local patient end before the current moment using the biological tissue deformation prediction module.
[0140] Based on the above embodiments, the system further includes a training module for training the device motion prediction module: acquiring a first training sample; wherein the first training sample includes historical operation control commands, historical theoretical projection sequences corresponding to the historical operation control commands, and historical actual motion trajectories; using the historical operation control commands and historical theoretical projection sequences as inputs and the historical actual motion trajectories as outputs, training the initial device motion prediction module to generate the device motion prediction module.
[0141] Based on the above embodiments, the training module is also used to train the biological tissue deformation prediction module: obtain a second training sample; wherein, the second training sample includes historical surgical video frames and the actual movement trajectory of the instruments extracted from the historical surgical video frames; and train the initial biological tissue deformation prediction module based on the historical surgical video frames and the actual movement trajectory of the instruments to generate the biological tissue deformation prediction module.
[0142] This application also provides a remote surgical robot, including a remote doctor terminal, a local patient terminal, and a communication module; wherein the remote doctor terminal and the local patient terminal communicate through the communication module; the remote doctor terminal is used to execute the method steps of any of the above method embodiments.
[0143] In the above implementation process, the remote doctor's end includes a main control platform, a calibration module, a prediction module (with a prediction model deployed), a latency estimation module, a display decision module, and a display module. The local patient's end is equipped with surgical instruments.
[0144] The doctor generates operation control commands through the main control platform, which are then sent to the local patient terminal via the communication module. These commands are also simultaneously sent to the prediction module. The prediction module generates predictive surgical videos using a predictive model.
[0145] The surgical instruments on the local patient end perform operations on biological tissues through operation control commands, and return actual scene videos to the remote doctor end through the communication module.
[0146] The latency estimation module on the remote doctor's end estimates the real-time latency between the remote doctor's end and the local patient's end. If the current network latency is greater than or equal to a first preset latency threshold, the display decision module decides to display the predictive surgical video, and the display module displays the predictive surgical video. After receiving the real-scene video, the correction module uses the real-scene video to correct the predictive surgical video, and the display module displays the corrected surgical video.
[0147] If the current network latency is less than the first preset latency threshold, the display decision module decides to display the actual scene video, and the display module displays the actual scene video.
[0148] The calibration module is also used to adjust the prediction model online using real-world video footage.
[0149] In summary, the beneficial effects of this application are as follows: This application constructs an intelligent closed loop centered on a data-driven predictive model on the remote doctor's end. This predictive model can instantly predict visual changes in the surgical scene based on the remote doctor's operational control commands, providing doctors with "zero-latency" operational feedback. Simultaneously, it continuously corrects the prediction using real data transmitted from the local patient's end, ensuring the eventual consistency of the "global state." This eliminates the perceived operational latency and intelligently responds to network fluctuations.
[0150] By introducing uncertainty assessment and adaptive display mechanisms, doctors are proactively alerted when model predictions are unreliable, preventing misleading information. Simultaneously, fusion technology ensures smooth transitions in the displayed images, avoiding visual abrupt changes and guaranteeing consistent and stable operation.
[0151] In addition, the predictive model can be continuously optimized using new data generated during surgery through an online learning mechanism, so that it can adapt to different individual patients and types of surgery.
[0152] Figure 3This is a schematic diagram of the electronic device structure provided in the embodiments of this application, such as... Figure 3 As shown, the electronic device includes a processor 301, a memory 302, and a bus 303; wherein the processor 301 and the memory 302 communicate with each other through the bus 303. The processor 301 is used to call program instructions in the memory 302 to execute the methods provided in the above-described method embodiments.
[0153] Processor 301 can be an integrated circuit chip with signal processing capabilities. The aforementioned processor 301 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor.
[0154] The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0155] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0156] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0157] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0158] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for delay compensation in remote surgery, characterized in that, The method is applied to a remote doctor's terminal, and the method includes: The system acquires operation control commands, sends the operation control commands to a prediction model, and synchronously sends the operation control commands to a local patient terminal via a network; wherein, the prediction model is used to generate predictive surgical videos based on the operation control commands; Obtain the current network latency between the remote doctor's terminal and the local patient's terminal; If the current network latency is determined to be greater than or equal to a first preset latency threshold, the predictive surgical video will be output and displayed on the remote doctor's end.
2. The method according to claim 1, characterized in that, The method further includes: If the current network latency is determined to be less than the first preset latency threshold, the actual scene video transmitted back from the local patient terminal via the network will be output and displayed on the remote doctor terminal.
3. The method according to claim 1, characterized in that, in, The predictive surgical video includes predictive videos of non-critical areas; the method further includes: If the current network latency is determined to be less than the first preset latency threshold, the actual video of the key area transmitted back from the local patient terminal via the network is stitched together with the predicted video of the non-key area to obtain a stitched surgical video, and the stitched surgical video is output and displayed on the remote doctor's terminal.
4. The method according to claim 1, characterized in that, in, The predictive surgical video includes a confidence level; the step of outputting and displaying the predictive surgical video on the remote doctor's end if it is determined that the current network latency is greater than or equal to a first preset latency threshold includes: If the current network latency is determined to be greater than or equal to a first preset latency threshold, and the confidence level is less than a preset confidence level threshold, then the predictive surgical video is degraded to generate a fuzzy predictive video, and the fuzzy predictive video is output and displayed on the remote doctor's end.
5. The method according to claim 1, characterized in that, The step of outputting and displaying the predictive surgical video on the remote doctor's end if the current network latency is determined to be greater than or equal to the first preset latency threshold includes: During the display of the predictive surgical video, actual scene video is received from the local patient terminal via network transmission; Based on the actual scene video, the predictive surgical video is corrected, and the corrected surgical video is output and displayed.
6. The method according to claim 5, characterized in that, The step of correcting the predictive surgical video based on the actual scene video and outputting and displaying the corrected surgical video includes: The actual scene video and the predictive surgical video are fused to obtain the corrected surgical video, and the corrected surgical video is then output and displayed.
7. The method according to claim 6, characterized in that, in, The actual scene video includes real frames to be displayed; The predictive surgical video includes the predictive frames that are being displayed; The step of fusing the actual scene video and the predictive surgical video to obtain the corrected surgical video includes: Calculate the motion vector field between the currently displayed predicted frame and the actual frame to be displayed; Based on the motion vector field, the deformed prediction frame of the currently displayed prediction frame is calculated, and the deformed real frame of the real frame to be displayed is calculated; wherein the deformed prediction frame and the deformed real frame are time-aligned. The corrected surgical video is obtained by weighted fusing the predicted deformation frames and the actual deformation frames.
8. The method according to claim 1, characterized in that, After acquiring the operation control command, the method further includes: Using the forward kinematics model of the instrument at the local patient end, the theoretical projection data of the instrument on the two-dimensional image plane of the endoscope is determined based on the operation control command; wherein, the prediction model is used to generate a predictive surgical video based on the operation control command and the theoretical projection data.
9. The method according to claim 1, characterized in that, The method further includes: The network latency prediction model is used to predict the network status between the remote doctor and the local patient within a preset time period from the current moment, thereby obtaining the predicted network latency. If the predicted network latency is determined to be less than the second preset latency threshold, the prediction model is stopped or the frequency of the prediction model is reduced, and within the preset future time period, the actual scene video transmitted from the local patient terminal via the network is output and displayed on the remote doctor terminal; wherein, the second preset latency threshold is less than the first preset latency threshold.
10. The method according to claim 1, characterized in that, The method further includes: The prediction model is adjusted online using the actual scene video transmitted back from the local patient terminal.
11. The method according to any one of claims 1-10, characterized in that, in, The prediction model is deployed on the remote doctor's end; the prediction model includes an instrument motion prediction module and a biological tissue deformation prediction module; the prediction model is used to generate predictive surgical videos based on the operation control commands, including: Using the aforementioned device motion prediction module, a device motion prediction trajectory is generated based on the operation control commands; Using the biological tissue deformation prediction module, the predictive surgical video is generated based on the instrument motion prediction trajectory and the historical surgical video transmitted back from the local patient terminal before the current moment.
12. The method according to claim 11, characterized in that, The training process of the device motion prediction module is as follows: Obtain a first training sample; wherein the first training sample includes historical operation control commands, historical theoretical projection sequences corresponding to the historical operation control commands, and historical actual motion trajectories; The historical operation control commands and the historical theoretical projection sequence are used as inputs, and the historical actual motion trajectory is used as outputs to train the initial instrument motion prediction module, thereby generating the instrument motion prediction module.
13. The method according to claim 11, characterized in that, The training process of the biological tissue deformation prediction module is as follows: Obtain a second training sample; wherein the second training sample includes historical surgical video frames and the actual movement trajectory of the instruments extracted from the historical surgical video frames; Based on the historical surgical video frames and the actual movement trajectory of the instruments, the initial biological tissue deformation prediction module is trained to generate the biological tissue deformation prediction module.
14. A delay compensation system for remote surgery, characterized in that, The system is deployed on a remote doctor's end; the system includes: The first acquisition module is used to acquire operation control instructions and send the operation control instructions to the local patient terminal via the network; The prediction module includes a prediction model. The first acquisition module also synchronously sends the operation control command to the prediction model. The prediction model is used to generate a predictive surgical video based on the operation control command. The second acquisition module is used to acquire the current network latency between the remote doctor's terminal and the local patient's terminal; The display module is used to output and display the predictive surgical video on the remote doctor's end if it is determined that the current network latency is greater than or equal to a first preset latency threshold.
15. A remote surgical robot, characterized in that, It includes a remote doctor terminal, a local patient terminal, and a communication module; wherein the remote doctor terminal and the local patient terminal communicate through the communication module; the remote doctor terminal is used to perform the method described in any one of claims 1-11.