Surgical image processing method, surgical robot system and storage medium
Through the dual-camera system and spatial interpolation processing technology, the image abnormality caused by endoscopic camera failure is solved, the continuity and clarity of images in laparoscopic surgery is ensured, and the safety and fluency of surgical operations are improved.
Patent Information
- Application Number
- CN202410031138.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-11
AI Technical Summary
During laparoscopic surgery, when the endoscopic camera fails, it makes it difficult for doctors to continue diagnosis and treatment. The prior art cannot effectively deal with camera abnormalities, affecting the continuity and safety of surgical operations.
Using a dual camera system, an image sequence from the first camera and the second camera of the endoscope is received and detected by the controller, a second camera image is inserted into the first camera image sequence using spatial interpolation processing technology, or a first camera image is inserted into the second camera image sequence to ensure the continuity and clarity of the image sequence.
When the camera is abnormal, the spatial interpolation processing technology ensures that the complete and clear image sequence is displayed in the display, reducing the risk of image abnormalities during surgery and improving the continuity and safety of surgical operations.
Smart Images

Figure CN120284170A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of medical devices, and particularly to a method for processing surgical images, a surgical robot system, and a storage medium. Background Art
[0002] Laparoscopic surgery is a surgical form that has gradually developed and been widely used in recent years. It has advantages such as small incisions, greatly reducing the patient's recovery time, discomfort experience, and postoperative side effects. Performing laparoscopic surgery, especially single-port laparoscopic surgery, with a surgical robot can optimize the surgical form through computer remote control technology.
[0003] An endoscope is an optical instrument widely used in the medical field. The endoscope can enter the patient's body through a very small opening on the patient's body or a natural body orifice, and can assist doctors in expanding the surgical field of view. However, currently, in the case where the left camera or the right camera of the endoscope fails, the endoscope will stop working, which will make it difficult for doctors to continue the diagnosis and treatment. Summary of the Invention
[0004] In some embodiments, the present disclosure provides a method for processing surgical images, including:
[0005] Receiving a first surgical image sequence from a first camera of an endoscope;
[0006] Receiving a second surgical image sequence from a second camera of the endoscope;
[0007] Detecting the first surgical image sequence;
[0008] Detecting the second surgical image sequence; and
[0009] In response to the first surgical image sequence or the second surgical image sequence satisfying an abnormal condition, using the second surgical image sequence or the first surgical image sequence to perform spatial interpolation processing on the first surgical image sequence or the second surgical image sequence.
[0010] In some embodiments, the present disclosure further provides a surgical robot system, including:
[0011] An endoscope, including a first camera and a second camera;
[0012] A controller, configured to be able to execute the method for processing surgical images according to any one of some embodiments of the present disclosure; and
[0013] A display, connected to the controller, for displaying the first surgical image sequence and / or the second surgical image sequence processed by the method for processing surgical images.
[0014] In some embodiments, the present disclosure also provides a computer-readable storage medium for storing at least one instruction, which, when executed by a computer, causes the computer to execute a method for processing surgical images according to any one of some embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for the description of the embodiments of the present disclosure. The drawings in the following description only show some embodiments of the present disclosure. For those of ordinary skill in the art, other embodiments can be obtained according to the content of the embodiments of the present disclosure and these drawings without creative efforts.
[0016] Figure 1 A flowchart showing a method for processing surgical images according to some embodiments of the present disclosure;
[0017] Figure 2 A schematic diagram showing a surgical robot according to some embodiments of the present disclosure;
[0018] Figure 3 A schematic block diagram showing the structure of an image system of a surgical robot according to some embodiments of the present disclosure;
[0019] Figure 4 A flowchart showing spatial interpolation processing of a first surgical image sequence according to some embodiments of the present disclosure;
[0020] Figure 5 A schematic diagram showing a first surgical image sequence and a second surgical image sequence according to some embodiments of the present disclosure;
[0021] Figure 6A A block diagram showing a spatial interpolation model according to some embodiments of the present disclosure;
[0022] Figure 6B A block diagram showing a surgical stage recognition model and a spatial interpolation model according to some embodiments of the present disclosure;
[0023] Figure 7 A flowchart showing a method for processing surgical images according to some embodiments of the present disclosure;
[0024] Figure 8 A flowchart showing temporal interpolation processing of a first surgical image sequence and a second surgical image sequence according to some embodiments of the present disclosure;
[0025] Figure 9 A schematic diagram showing a first surgical image sequence and a second surgical image sequence according to some other embodiments of the present disclosure;
[0026] Figure 10ABlock diagram showing a temporal interpolation model according to some embodiments of the present disclosure;
[0027] Figure 10B Block diagram showing a surgical stage recognition model and a temporal interpolation model according to some embodiments of the present disclosure. Detailed Description of Specific Embodiments
[0028] To make the technical problems solved by the present disclosure, the technical solutions adopted, and the achieved technical effects clearer, the technical solutions of the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only exemplary embodiments of the present disclosure, rather than all embodiments.
[0029] In the description of the present disclosure, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present disclosure. In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance. In the description of the present disclosure, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection or a detachable connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium; it can be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific circumstances. In the present disclosure, the end close to the operator (such as a doctor) is defined as the proximal end, the proximal part, the rear end, or the rear part, and the end close to the surgical patient is defined as the distal end, the distal part, the front end, or the front part. Those skilled in the art can understand that the embodiments of the present disclosure can be used in medical devices or surgical robots, and can also be used in other non-medical devices.
[0030] Figure 1 Flowchart showing a method 100 for processing surgical images according to some embodiments of the present disclosure. Figure 2 Schematic diagram showing a surgical robot system 200 according to some embodiments of the present disclosure. Figure 3 Schematic block diagram showing the structure of the image system of the surgical robot system 200 according to some embodiments of the present disclosure. The method 100 can be implemented or executed at least partially by hardware, software, or firmware. In some embodiments, the method 100 can be executed by a robot system. The robot system can be a surgical robot system (for example, Figure 2Various suitable robotic systems, including the illustrated surgical robotic system 200). The robotic system may also include dedicated or general-purpose robotic systems for other fields (e.g., manufacturing, machinery, etc.). In some embodiments, method 100 may be performed at least in part by a controller ( Figure 2 not shown, e.g., Figure 3 the illustrated controller 240) of the surgical robotic system 200. In some embodiments, method 100 may be implemented as computer-readable instructions. These instructions may be read and executed by a general-purpose processor or a dedicated processor (e.g., the controller of the surgical robotic system 200). For example, the controller 240 of the surgical robotic system 200 may include a processor configured to perform method 100. In some embodiments, these instructions may be stored on a computer-readable storage medium.
[0031] In some embodiments, the surgical robotic system 200 may be various suitable surgical robots, including a laparoscopic surgical robot. In some embodiments, as Figure 2 shown, the surgical robotic system 200 may include a surgical cart 210, a master cart 220, and an equipment cart 230. The controller ( Figure 2 not shown, e.g., Figure 3 the illustrated controller 240) may be disposed at any suitable position in the surgical cart 210, the master cart 220, the equipment cart 230, or the surgical robotic system 200. The surgical cart 210 may include at least one moving arm 211, and at least one moving arm 211 may be movably disposed on the surgical cart 210. In some embodiments, at least one moving arm 211 may be a positioning arm of the surgical robot, and at least one surgical instrument ( Figure 2 not shown, e.g., forceps, curved scissors, an endoscope (e.g., Figure 3 the illustrated endoscope 213), etc.) may be carried at the distal end of at least one moving arm 211. The master cart 220 may include at least one master operator 221, and at least one master operator 221 may be used to receive operations of a user on at least one master operator 221. The master cart 220 may be communicatively connected to the surgical cart 210. During surgery, the surgical cart 210 is generally located on the patient side, and the user may issue control instructions by operating at least one master operator 221 of the master cart 220 to control at least one surgical instrument carried by the surgical cart 210 to perform surgical operations on the patient. The equipment cart 230 may be communicatively connected to the surgical cart 210 and the master cart 220 respectively.
[0032] In some embodiments, as Figure 3As shown, the surgical robot system 200 may further include an endoscope 213, and the endoscope 213 may include a first camera 213a and a second camera 213b. In some embodiments, the endoscope 213 may be mounted at the end of the robotic arm 211 in the surgical trolley 210 to acquire surgical images during the surgery. The controller 240 may be respectively connected to the endoscope 213 and a display (e.g., the display 224, the displays 212, 225, 231 as shown in Figure 2 etc.) to process the surgical images received from the endoscope 213 and transmit the processed surgical images to the display for the user to view.
[0033] As Figure 1 shown, in step 110, a first sequence of surgical images from the first camera 213a of the endoscope is received. In step 120, a second sequence of surgical images from the second camera 213b of the endoscope is received. Based on step 110 and step 120, the controller 240 of the surgical robot system 200 can receive the first sequence of surgical images and the second sequence of surgical images. In some embodiments, the endoscope may be mounted at the distal end of the robotic arm 211 in the surgical robot system 200. During the surgery, the end of the endoscope can extend into the patient's body to acquire images inside the patient's body and allow the doctor to view the images inside the patient's body (e.g., the environment and lesions inside the patient's body). The first camera 213a and the second camera 213b of the endoscope may be the left camera and the right camera of the endoscope respectively, and the first sequence of surgical images and the second sequence of surgical images may be the left-eye surgical image sequence and the right-eye surgical image sequence respectively. Those skilled in the art can understand that the first sequence of surgical images may include multiple frames of first surgical images acquired by the first camera, and the second sequence of surgical images may include multiple frames of second surgical images acquired by the second camera. In some embodiments, the controller of the surgical robot system 200 may be disposed in the equipment trolley 230 of the surgical robot, and the equipment trolley 230 may be communicatively connected to the surgical trolley 210 to receive the first sequence of surgical images and the second sequence of surgical images from the endoscope.
[0034] As Figure 1 shown, in step 130, the first sequence of surgical images is detected. In step 140, the second sequence of surgical images is detected. In step 130 and step 140, the controller 240 of the surgical robot system 200 can detect the first sequence of surgical images and the second sequence of surgical images in various ways, such as detecting the pixel values of the pixel points included in the images in the first sequence of surgical images and the second sequence of surgical images, detecting the frequency domain coefficients obtained after performing fast Fourier transform processing on the images, and so on.
[0035] In some embodiments, step 130 may include: detecting whether a first surgical image in a first sequence of surgical images from a first camera meets an abnormal condition. Step 140 may include: detecting whether a second surgical image in a second sequence of surgical images from a second camera meets an abnormal condition. In some embodiments, the abnormal condition may include at least one of the following: possible abnormalities such as image freeze, image blackout, image blurring, or image occlusion.
[0036] The controller 240 of the surgical robot system 200 may determine whether a first surgical image in the first sequence of surgical images or a second surgical image in the second sequence of surgical images meets an abnormal condition in various ways. In some embodiments, detecting whether a first surgical image or a second surgical image meets an abnormal condition may include: performing a fast Fourier transform process on the first surgical image or the second surgical image to obtain the frequency domain coefficients of the first surgical image or the second surgical image; detecting whether the high-frequency coefficients in the frequency domain coefficients of the first surgical image or the second surgical image are greater than a preset threshold; and in response to detecting that the high-frequency coefficients of the first surgical image or the second surgical image are greater than the preset threshold, determining that the first surgical image or the second surgical image meets the abnormal condition. If it is detected that the high-frequency coefficients in the frequency domain coefficients of the first surgical image or the second surgical image are greater than the preset threshold, it indicates that an abnormal situation has occurred in the first surgical image or the second surgical image, such as image freeze. At this time, the doctor cannot view a clear image of the patient's body collected by the first camera or the second camera through the display. Those skilled in the art can understand that the preset threshold involved here can be set by relevant technicians based on experience.
[0037] During surgery, due to splashing of the patient's blood, body fluids, etc. and covering the camera of the endoscope, etc., the images collected by the endoscope may be blurred or occluded. In the case of image blurring, it is difficult for the user to distinguish the lesion or the patient's tissue, etc. from the seen image. In some embodiments, a fast Fourier transform process may be performed on the first surgical image or the second surgical image to obtain the frequency domain coefficients of the first surgical image or the second surgical image. If the proportion of the high-frequency components in the frequency domain coefficients is less than a preset value, it can be determined that an abnormal situation of image blurring has occurred in the first surgical image or the second surgical image.
[0038] In the case of image occlusion, for example, when the screen area blocked by blood stains reaches a preset proportion (the specific proportion can be set according to the specific situation, such as 1 / 3, etc.), the user's field of view is restricted. Performing surgical operations in this situation is extremely dangerous. In some embodiments, when it is detected that the proportion of the screen area blocked by the first surgical image or the second surgical image exceeds the preset threshold, it can be determined that the first surgical image or the second surgical image meets the abnormal condition of image occlusion.
[0039] In some embodiments, the first surgical image or the second surgical image may also be grayscale processed. If the number of pixel points whose grayscale values satisfy a preset range reaches a preset ratio, it may be determined that the first surgical image or the second surgical image meets the abnormal condition of image blackout.
[0040] Those skilled in the art can understand that the methods for detecting the first surgical image sequence and the second surgical image sequence are not limited to the several methods listed above, and may also include any suitable methods.
[0041] As Figure 1 shown, in step 150, in response to the first surgical image sequence or the second surgical image sequence meeting the abnormal condition, the second surgical image sequence is used to perform spatial interpolation processing on the first surgical image sequence, or the first surgical image sequence is used to perform spatial interpolation processing on the second surgical image sequence. Spatial interpolation processing refers to, based on the spatial relationship between the first surgical image sequence and the second surgical image sequence, using the second surgical image sequence to obtain the interpolated image in the first surgical image sequence and inserting it into the first surgical image sequence, or using the first surgical image sequence to obtain the interpolated image in the second surgical image sequence and inserting it into the second surgical image sequence.
[0042] In the case where the first surgical image sequence or the second surgical image sequence is abnormal, if no effective treatment is performed, the user will lose the view of the patient's body (for example, the view of the left camera or the right camera), making it difficult to perform surgical operations. Continuing to perform surgical operations in this case poses a greater risk. In step 150, the controller 240 of the surgical robot system 200 may perform spatial interpolation processing on the abnormal surgical image sequence. Based on this, in the case where the first surgical image sequence or the second surgical image sequence is abnormal, the user can continue to view the complete and clear first surgical image sequence and the second surgical image sequence on the display, which helps the user complete the current surgical operation being performed and is beneficial to reducing the negative impact caused by the abnormality of the surgical image.
[0043] The first camera and the second camera may be the left camera and the right camera of the endoscope respectively, and are arranged side by side at the end of the endoscope. The views of the first camera and the second camera have a specific spatial relationship and there is a certain overlap. In some embodiments, in the case where the first surgical image sequence is detected to be abnormal, in step 150, the controller 240 of the surgical robot system 200 may utilize the spatial relationship of the views of the first camera and the second camera of the endoscope to obtain the first surgical image based on the second surgical image in the second surgical image sequence, and insert the obtained first surgical image into the first surgical image sequence so that the user can observe the complete and clear first surgical image sequence on the display.
[0044] The following content takes the spatial interpolation processing of the first surgical image sequence as an example to introduce the technical details involved in spatial interpolation processing. Those skilled in the art can understand that the spatial interpolation processing of the second surgical image sequence is similar. Figure 4 FIG. shows a flowchart of spatial interpolation processing of the first surgical image sequence according to some embodiments of the present disclosure. In some embodiments, using the second surgical image sequence to perform spatial interpolation processing on the first surgical image sequence may include steps 410 to 430 as Figure 4 shown.
[0045] As Figure 4 shown, in step 410, at least one interpolation position is determined in the first surgical image sequence. The interpolation position refers to the position in the first surgical image sequence where an interpolation image needs to be inserted. Figure 5 FIG. shows a schematic diagram of the first surgical image sequence 510 and the second surgical image sequence 520 according to some embodiments of the present disclosure. As Figure 5 shown, the first surgical image sequence may include first surgical images 511 - 514, etc., and the second surgical image sequence may include second surgical images 521 - 524, etc. As Figure 5 shown, the first surgical image 511 is an image acquired by the first camera at time t0, and the second surgical image 521 is an image acquired by the second camera at time t0; the first surgical image 512 is an image acquired by the first camera at time t0 + Δt, and the second surgical image 522 is an image acquired by the second camera at time t0 + Δt; and so on for other surgical images. In some embodiments, the controller 240 of the surgical robot 200 may, in response to detecting that a first surgical image in the first surgical image sequence (e.g., Figure 5 the first surgical image 513 shown) meets an abnormal condition, use the position where the first surgical image 513 is located as the interpolation position in the first surgical image sequence. In some embodiments, the controller 240 may determine multiple interpolation positions in the first surgical image sequence, such as the positions where the first surgical images 513 and 514 are located, etc.
[0046] As Figure 4 shown, in step 420, for each interpolation position, an interpolation image is obtained based on at least one adjacent image before and adjacent to the interpolation position in the first surgical image sequence, a reference image in the second surgical image sequence with the same acquisition time as the interpolation position, and a pre-trained spatial interpolation model. In step 430, the obtained interpolation image is inserted into at least one interpolation position in the first surgical image sequence.
[0047] Since there is a specific spatial relationship and a certain overlap in the fields of view of the first camera and the second camera, the surgical images with the same acquisition time in the first surgical image sequence 510 and the second surgical image sequence 520 include regions with the same content. For example, the first surgical image 511 and the second surgical image 521 with the acquisition time of t0 include regions with the same content. Each frame image in the first surgical image sequence 510 is acquired by the first camera, and the time interval between adjacent images corresponding to the acquisition time is relatively small. Therefore, the content change between adjacent images is generally small. It can be known that for the interpolated image corresponding to the interpolation position in the first surgical image sequence, the content change is smaller compared to the adjacent images in the first surgical image sequence, and at the same time, it also includes a part of the region with the same content as a part of the region of the reference image in the second surgical image sequence. Moreover, in the endoscopic scenario, the spatial relationship of the fields of view of the first camera and the second camera remains unchanged, and there is a large overlapping region at the center of the region of interest. Therefore, the effect of spatial interpolation processing is better.
[0048] Based on this, in some embodiments, in step 420, for the interpolation position (for example, the position where the first surgical image 513 is located, hereinafter also referred to as the interpolation position t0 + 2Δt), at least one adjacent image (for example, image 512 or images 512 and 511, etc.) before and adjacent to the interpolation position t0 + 2Δt in the first surgical image sequence 510, the reference image 523 with the same acquisition time as the interpolation position in the second surgical image sequence, and a pre-trained spatial interpolation model can be used to obtain the interpolated image; in step 430, the obtained interpolated image can be inserted into the interpolation position t0 + 2Δt in the first surgical image sequence 510. In some embodiments, in step 420, for the interpolation position t0 + 3Δt, at least one adjacent image (for example, image 513 or images 513, 512, 511, etc.), the reference image 524, and a pre-trained spatial interpolation model can be used to obtain the interpolated image; in step 430, the obtained interpolated image can be inserted into the interpolation position t0 + 3Δt.
[0049] In some embodiments, the spatial interpolation model can obtain the interpolated image based on at least one adjacent image at the interpolation position in the first surgical image sequence and the reference image at the interpolation position in the second surgical image sequence. In step 420, at least one adjacent image and the reference image can be input into the pre-trained spatial interpolation model to obtain the interpolated image output by the spatial interpolation model for insertion into the interpolation position in the first surgical image sequence 510.
[0050] Figure 6A A block diagram showing a spatial interpolation model 600 according to some embodiments of the present disclosure. As Figure 6AAs shown, the spatial interpolation model 600 may include a backward prediction part 610, a spatial reference part 620, and an image generation part 630. Among them, the backward prediction part 610 may be used to obtain the extraction features of at least one adjacent image and the predicted interpolated image based on at least one adjacent image before and adjacent to the interpolation position in the first surgical image sequence. The spatial reference part 620 may be used to obtain the extraction features of the reference image based on the reference image in the second surgical image sequence with the same acquisition time as the interpolation position. The image generation part 630 may be used to obtain the interpolated image based on at least one adjacent image, the extraction features of at least one adjacent image, the reference image, and the extraction features of the reference image.
[0051] As Figure 6A shown, the backward prediction part 610 may include a first input layer 611, a first feature extraction layer 612, and a first backward prediction layer 613 connected in series in sequence.
[0052] Among them, the first input layer 611 may be used to receive at least one adjacent image (e.g., Figure 5 the images 511 and 512 shown) before and adjacent to the interpolation position (e.g., the interpolation position t0 + 2Δt in the first surgical image sequence 510 shown) in the first surgical image sequence. Figure 5
[0053] The first feature extraction layer 612 can be used to obtain the extracted features of the images 511 and 512 based on at least one frame of adjacent images 511 and 512 from the first input layer 611. The extracted features can include motion features and non-motion features. In some embodiments, the first feature extraction layer 612 can include a first feature generation sub-layer, a first attention sub-layer, and a first feature separation sub-layer (not shown in the figure) cascaded in sequence. Among them, the first feature generation sub-layer can be used to perform preliminary feature extraction on the adjacent images 511 and 512 from the first input layer 611 to obtain the rough features of the adjacent images 511 and 512, such as features related to the content included in the images 511 and 512, features related to parameters such as the brightness of the images, etc. The first attention sub-layer can be used to perform attention processing (processing based on the attention mechanism) on the rough features of the adjacent images 511 and 512 from the first feature generation sub-layer to obtain the features to be separated of the adjacent images 511 and 512. Performing attention processing on the rough features of the adjacent images 511 and 512 helps to efficiently allocate information processing resources, thereby facilitating the improvement of the performance of the spatial interpolation model. The first feature separation sub-layer can be used to perform separation processing on the features to be separated of the adjacent images 511 and 512 from the first attention sub-layer to obtain the motion features and non-motion features of the adjacent images 511 and 512. Among them, the non-motion features of the image can be used to describe the non-motion information of the image. For example, the background information in the image, such as the environment inside the patient (e.g., inside the abdominal cavity, inside the natural body cavity, etc.). The motion features of the image can be used to describe the motion information of the image, such as information about surgical instruments or excised patient tissues, etc.
[0054] The first backward prediction layer 613 can be used to obtain a predicted interpolated image based on the extracted features of the adjacent images 511 and 512 from the first feature extraction layer 612. In some embodiments, the first backward prediction layer 613 can include a motion feature prediction sub-layer and a feature fusion sub-layer. Among them, the motion feature prediction sub-layer can be used to perform backward prediction on the motion features of the adjacent images 511 and 512 to obtain predicted motion features. The predicted motion features can be used to describe the motion information of the predicted interpolated image, such as information about surgical instruments or excised patient tissues in the predicted interpolated image. The feature fusion sub-layer can be used to fuse the non-motion features of the adjacent images 511 and 512 received from the first feature extraction layer 612 and the predicted motion features received from the motion feature prediction sub-layer to obtain the predicted interpolated image.
[0055] As Figure 6A shown, the spatial reference part 620 can include a second input sub-layer 621 and a second feature extraction layer 622.
[0056] Among them, the second input layer 621 can be used to receive the acquisition time and interpolation position in the second surgical image sequence (e.g.,Figure 5 a reference image (e.g., Figure 5 the reference image 523 shown in the second surgical image sequence 520).
[0057] The second feature extraction layer 622 can be used to obtain the first extracted features of the reference image 523 based on the reference image 523 from the second input layer 621. In some embodiments, the second feature extraction layer 622 can include a second feature generation sub-layer, a second attention sub-layer, and a second feature separation sub-layer (not shown in the figure). Among them, the second feature generation sub-layer can be used to perform preliminary feature extraction on the reference image 523 from the second input layer 621 to obtain the rough features of the reference image 523. The second attention sub-layer can be used to perform attention processing (processing based on the attention mechanism) on the rough features of the reference image 523 from the second feature generation sub-layer to obtain the features to be separated of the reference image 523, so as to more efficiently allocate information processing resources. The second feature separation sub-layer can be used to separate the features to be separated of the reference image 523 from the second attention sub-layer to obtain the first extracted features of the reference image 523, such as features related to surgical instruments in the reference image 523, features related to patient tissues, etc.
[0058] As Figure 6A shown, the image generation part 630 can include an image fusion layer 631 and an output layer 632.
[0059] Among them, the image fusion layer 631 can be used to fuse the predicted interpolated image from the first backward prediction layer 613 of the backward prediction part 610 and the reference image from the second input layer 621 based on the extracted features of the adjacent images 511 and 512 from the first feature extraction layer 612 and the first extracted features of the reference image 523 of the second feature extraction layer 622 from the spatial reference part 620, to obtain a fused interpolated image. In some embodiments, the image fusion layer 631 may include a feature point matching sub-layer, an image stitching sub-layer, a pixel point weight determination sub-layer, and a fusion processing sub-layer (not shown in the figure). Among them, the feature point matching sub-layer can be used to perform feature point matching on the predicted interpolated image from the backward prediction part 610 and the reference image from the second input layer 621 (e.g., image 523) based on the extracted features of the adjacent images 511 and 512 from the first feature extraction layer 612 and the first extracted features of the reference image 523 from the spatial reference part 620, to obtain a feature point matching result. The image stitching sub-layer can be used to stitch the predicted interpolated image and the reference image 523 based on the feature point matching result from the feature point matching sub-layer, to obtain a stitched interpolated image. The pixel point weight determination sub-layer can be used to determine the weights of each pixel point included in the overlapping area of the stitched interpolated image from the image stitching sub-layer, to obtain a pixel point weight determination result. The fusion processing sub-layer can be used to perform fusion processing on the stitched interpolated image from the image stitching sub-layer based on the weight determination result from the pixel point weight determination sub-layer, to obtain a fused interpolated image. After being processed by the fusion processing sub-layer, the transition of the stitched area in the image can be made more natural, which is beneficial to improving the user experience.
[0060] The output layer 632 can be used to obtain an interpolated image based on the fused interpolated image from the image fusion layer 631, for example, perform processing such as cropping and aspect ratio adjustment on the fused interpolated image, so that the aspect ratio of the obtained interpolated image is consistent with the aspect ratio of each frame image included in the first surgical image sequence.
[0061] The present disclosure does not limit the manner in which each layer and each sub-layer of the spatial interpolation model implement their functions. Each layer and sub-layer can implement the functions they can achieve in any suitable manner. In some embodiments, the spatial interpolation model can adopt a deep neural network model, such as a recurrent neural network, a convolutional neural network, or other network models. Those skilled in the art can understand that the above is only an example, and the spatial interpolation model can also adopt any other suitable network model. In addition, the present disclosure does not limit the detailed structure of the spatial interpolation model, and the spatial interpolation model can be any suitable structure.
[0062] In some embodiments, spatially interpolating the second surgical image sequence using the first surgical image sequence may include: determining at least one interpolation position in the second surgical image sequence; for each interpolation position, obtaining an interpolated image based on at least one adjacent image before and adjacent to the interpolation position in the second surgical image sequence, a reference image in the first surgical image sequence with the same acquisition time as the interpolation position, and a pre-trained spatial interpolation model; and inserting the obtained interpolated image into at least one interpolation position in the second surgical image sequence. Spatially interpolating the second surgical image sequence using the first surgical image sequence is similar to spatially interpolating the first surgical image sequence using the second surgical image sequence. Refer to the above for details. To avoid repetition, it will not be elaborated here.
[0063] In some embodiments, the spatial interpolation model may be trained separately using the historical first surgical image sequence and the historical second surgical image sequence (e.g., the first surgical image sequence and the second surgical image sequence collected during a previous surgical procedure) as training data to obtain two spatial interpolation models corresponding to the first surgical image sequence and the second surgical image sequence, respectively. When performing spatial interpolation, the spatial interpolation model corresponding to the first surgical image sequence may be used to perform spatial interpolation on the first surgical image sequence, and the spatial interpolation model corresponding to the second surgical image sequence may be used to perform spatial interpolation on the second surgical image sequence to improve the prediction effect of the spatial interpolation model, thereby facilitating the improvement of the user experience.
[0064] Those skilled in the art can understand that the structures of the two spatial interpolation models corresponding to the first surgical image sequence and the second surgical image sequence respectively are Figure 6A similar to the structure of the spatial interpolation model 600 shown. To avoid repetition, it will not be elaborated here.
[0065] In some embodiments, the spatial interpolation model may be trained by a spatial interpolation model training method including the following steps 1 to 4, for example Figure 6A the output spatial interpolation model 600:
[0066] Step 1: Obtain a large number of surgical image sets. The surgical image sets may be surgical image sets obtained from historical surgeries performed by a surgical robot. The surgical images obtained from one surgery may form a surgical image set. One surgical image set may include a first surgical image sequence captured by a first camera of an endoscope and a second surgical image sequence captured by a second camera.
[0067] Step 2: Divide the images included in each surgical image set into data groups, where each data group includes input data and output data. In some embodiments, the simulated interpolation positions can be determined in the surgical image set, and the images at the simulated interpolation positions are determined as the output data in the data groups. Furthermore, at least one image before the simulated interpolation position can be determined as the adjacent image of the simulated interpolation position, and the image with the same acquisition time as the simulated interpolation position is determined as the reference image. The adjacent image and the reference image can be determined as the input data.
[0068] Step 3: Divide all the data groups into a training set and a test set. For example, 70% of the data groups are used as the training set, and 30% of the data groups are used as the test set.
[0069] Step 4: Train the spatial interpolation model based on the training set. In some embodiments, when training Figure 6A the spatial interpolation model 600 as shown, a first loss function can be constructed for the backward prediction part 610 of the spatial interpolation model 600, a second loss function can be constructed for the spatial reference part 620 of the spatial interpolation model 600, and a third loss function can be constructed for the image generation part 630 of the spatial interpolation model 600. In some embodiments, the backward prediction part 610 in the spatial interpolation model 600 can be trained first based on the first loss function, and then the spatial reference part 620 and the image generation part 630 can be trained respectively based on the second loss function and the third loss function, so that each part of the spatial interpolation model 600 reaches the optimal state. In some embodiments, the first loss function, the second loss function, and the third loss function can also be combined into a total loss function to jointly train the backward prediction part 610, the spatial reference part 620, and the image generation part 630 in the spatial interpolation model 600. For example, optimization algorithms such as gradient descent can be used to minimize the total loss function of the spatial interpolation model 600 to update the model parameters in the spatial interpolation model 600.
[0070] Step 5: Test the spatial interpolation model based on the test set. In some embodiments, the test set is used to evaluate and optimize the spatial interpolation model 600. The performance of the spatial interpolation model 600 is evaluated by comparing the interpolated images output by the spatial interpolation model 600 with the output data in the data groups. According to the evaluation results, the hyperparameters or the structure of the spatial interpolation model 600 can be adjusted to further improve the performance of the spatial interpolation model 600.
[0071] Those skilled in the art can understand that the above training of the spatial interpolation model is only an example, and the above training process can be adaptively adjusted according to needs or other methods can be used for training the spatial interpolation model. For example, semi-supervised training can also be adopted, etc.
[0072] In some embodiments, method 100 may further include: obtaining the surgical procedure type of the surgery; and for each interpolation position in the first surgical image sequence, obtaining an interpolated image based on at least one adjacent image before and adjacent to the interpolation position in the first surgical image sequence, a reference image in the second surgical image sequence with the same acquisition time as the interpolation position, and a pre-trained spatial interpolation model corresponding to the surgical procedure type; or for each interpolation position in the second surgical image sequence, obtaining an interpolated image based on at least one adjacent image before and adjacent to the interpolation position in the second surgical image sequence, a reference image in the first surgical image sequence with the same acquisition time as the interpolation position, and a pre-trained spatial interpolation model corresponding to the surgical procedure type.
[0073] The surgical procedure type of the surgery may include: radical prostatectomy, nephrectomy, segmentectomy, lobectomy, endometrial cancer staging, radical resection of rectal cancer, and other surgical procedures. In some embodiments, obtaining the surgical procedure type of the surgery may include receiving the surgical procedure type selected by the user. In some embodiments, the user of the surgical robot may select the surgical procedure type of the current surgery through a surgical procedure type determination device provided in the surgical robot (such as surgical robot 200). In some embodiments, the surgical procedure type determination device may be a device such as a key or button provided in the main control cart 220 of the surgical robot 200. For example, the user may select the surgical procedure type of the current surgery through a key provided on the handrest 222 of the main control cart 220.
[0074] In some embodiments, historical image data of surgeries of different surgical procedure types (for example, the first surgical image sequence and the second surgical image sequence) may be used as training data to train the spatial interpolation model corresponding to the surgical procedure type. For example, multiple sets of image data (including the first surgical image sequence and the second surgical image sequence) collected during multiple previous nephrectomies are used for model training to obtain the spatial interpolation model corresponding to nephrectomy. This spatial interpolation model can be specifically used for spatial interpolation processing of the first surgical image sequence or the second surgical image sequence collected during nephrectomy. Based on the above embodiments, using the spatial interpolation model corresponding to the surgical procedure type to obtain the interpolated image for inserting into the first surgical image sequence or the second surgical image sequence is beneficial to improving the accuracy of the interpolated image, thereby helping to improve the smoothness of the first surgical image sequence or the second surgical image sequence and enhancing the user experience.
[0075] In some embodiments, the spatial interpolation model corresponding to the surgical procedure type may further include a first surgical image sequence spatial interpolation model and a second surgical image sequence spatial interpolation model, which are respectively used to perform spatial interpolation processing on the first surgical image sequence and the second surgical image sequence. For example, when performing a nephrectomy and it is detected that the first surgical image sequence meets the abnormal conditions, the first surgical image sequence spatial interpolation model corresponding to the nephrectomy can be used to perform spatial interpolation processing on the first surgical image sequence.
[0076] The structure of the spatial interpolation model corresponding to each surgical procedure type and the functions that each layer can achieve are similar to Figure 6A the spatial interpolation model 600 shown, and for the sake of reducing repetition, it will not be elaborated here.
[0077] In some embodiments, method 100 may further include: obtaining the surgical stage where at least one interpolation position is located; and for each interpolation position in the first surgical image sequence, based on the surgical stage of the interpolation position, at least one adjacent image in the first surgical image sequence that is before and adjacent to the interpolation position, the reference image in the second surgical image sequence whose acquisition time is the same as that of the interpolation position, and the pre-trained spatial interpolation model corresponding to the surgical procedure type, obtaining an interpolated image; or for each interpolation position in the second surgical image sequence, based on the surgical stage of the interpolation position, at least one adjacent image in the second surgical image sequence that is before and adjacent to the interpolation position, the reference image in the first surgical image sequence whose acquisition time is the same as that of the interpolation position, and the pre-trained spatial interpolation model corresponding to the surgical procedure type, obtaining an interpolated image. In this embodiment, during the process of obtaining the interpolated image, the surgical stage where the interpolation position is located is considered, which is beneficial to improving the accuracy of the obtained interpolated image and enhancing the user experience.
[0078] In some embodiments, the surgical stage where the interpolation position is located can also be obtained through a pre-trained surgical stage recognition model. Based on this, in some embodiments, method 100 may further include: for each interpolation position in the first surgical image sequence, obtaining the surgical stage based on the reference image in the second surgical image sequence whose acquisition time is the same as that of the interpolation position and the pre-trained surgical stage recognition model, as the surgical stage where at least one interpolation position is located; and / or for each interpolation position in the second surgical image sequence, obtaining the surgical stage based on the reference image in the first surgical image sequence whose acquisition time is the same as that of the interpolation position and the pre-trained surgical stage recognition model, as the surgical stage where at least one interpolation position is located.
[0079] For the interpolation position in the first surgical image sequence, the surgical stage recognition model can obtain the surgical stage of the interpolation position based on the reference image of the interpolation position in the second surgical image sequence.Figure 6B A block diagram showing a surgical stage recognition model 6000 and a spatial interpolation model 600 according to some embodiments of the present disclosure. As Figure 6B shown, the surgical stage recognition model 6000 may include a third input layer 6001, a third feature extraction layer 6002, and a surgical stage determination layer 6003.
[0080] Among them, the third input layer 6001 may be used to receive a reference image (e.g., Figure 5 the reference image 523 shown) that is collected at the same time as the interpolation position (e.g., Figure 5 the interpolation position t0 + 2Δt shown).
[0081] The third feature extraction layer 6002 may be used to obtain a second extracted feature of the reference image 523 based on the reference image 523 from the third input layer 6001. In some embodiments, the third feature extraction layer 6002 may include a third feature generation sub-layer, a third attention sub-layer, and a third feature separation sub-layer (not shown in the figure). Among them, the third feature generation sub-layer may be used to perform preliminary feature extraction on the reference image 523 from the third input layer 6001 to obtain a rough feature of the reference image 523, such as a feature related to the content included in the image 523, such as a feature related to surgical instruments, the internal environment of the patient, patient tissues, etc. The third attention sub-layer may be used to perform attention processing (processing based on the attention mechanism) on the rough feature of the reference image 523 from the third feature generation sub-layer to obtain a to-be-separated feature of the reference image 523, so as to more efficiently allocate information processing resources. The third feature separation sub-layer may be used to perform separation processing on the to-be-separated feature of the reference image 523 from the third attention sub-layer to obtain a second extracted feature of the reference image 523, such as a feature related to surgical instruments in the reference image 523, a feature related to patient tissues, etc.
[0082] The surgical stage determination layer 6003 may be used to determine the surgical stage of the reference image based on the second extracted feature of the reference image 523 from the third feature determination layer 6002. For example, it may include surgical stages such as the surface fat cleaning stage, the tissue dissection stage, the blood vessel separation stage, the lesion resection stage, the tissue suture stage, etc.
[0083] In some embodiments, the surgical stage output by the surgical stage recognition model 6000 may be transmitted to the first backward prediction layer 613 in the spatial interpolation model 600. Based on this, the first backward prediction layer 613 may perform backward prediction based on the extracted features of the adjacent images 511 and 512 from the first feature extraction layer 612 and the surgical stage at the interpolation position to obtain a predicted interpolation image.
[0084] As Figure 6BAs shown, in some embodiments, the second feature extraction layer 622 in the spatial interpolation model 600 may receive, as shared features, at least a portion of the second extracted features of the reference image 523 from the third feature extraction layer 6002 of the surgical stage recognition model 6000. In some embodiments, the second feature extraction layer 622 may perform further feature extraction based on the second extracted features of the reference image 523 from the third feature extraction layer 6002 to obtain the first extracted features of the reference image 523 with more details.
[0085] In some embodiments, the same surgical stage recognition model may be used to determine the surgical stage of the first surgical image and the second surgical image. In some embodiments, the surgical stage recognition model may include two surgical stage recognition models respectively used to recognize the surgical stage of the first surgical image and the second surgical image.
[0086] In some embodiments, method 100 may further include: transmitting the spatially interpolated first surgical image sequence or the second surgical image sequence to the surgical trolley 210 and / or the master trolley 220 for display on at least one display included in the surgical robot system 200 for the user to view. In some embodiments, the 2D displays (e.g., displays 212, 225, 231, etc.) in the surgical robot system 200 may display the spatially interpolated first surgical image sequence or the spatially interpolated second surgical image sequence. The 3D display (e.g., display 224) in the surgical robot system 200 may be used to display the spatially interpolated first surgical image sequence and the second surgical image sequence.
[0087] In some embodiments, when the first surgical image sequence or the second surgical image sequence meets the abnormal condition, in addition to performing spatial interpolation processing, the controller 240 of the surgical robot 200 can also adopt other measures to reduce the impact caused by image abnormalities. In some embodiments, in response to the first surgical image sequence or the second surgical image sequence meeting the abnormal condition, the controller 240 can also send an alarm message to prompt the user of the occurrence of the abnormal situation, so that the user can perform relevant processing. In some embodiments, in response to the first surgical image sequence or the second surgical image sequence meeting the abnormal condition, the controller 240 can also send an endoscope replacement prompt message to prompt the user to replace the endoscope, so that the surgical robot system can resume normal use or eliminate the cause of the fault. In some embodiments, in response to the first surgical image sequence or the second surgical image sequence meeting the abnormal condition, the controller 240 can also send a lens cleaning prompt message to prompt the user to perform a lens cleaning operation. When the user triggers the lens cleaning operation, the moving arm 211 can retract the endoscope from the patient's body, and the moving arm 211 can automatically move to a pose convenient for lens cleaning so that the user can wipe the lens of the endoscope. In some embodiments, performing the lens cleaning operation can wipe off blood stains and the like on the endoscope lens, thereby solving abnormal situations such as blurred images or image occlusion.
[0088] In some embodiments, after the user completes the lens cleaning operation, the user can send a message indicating the completion of lens cleaning by operating a button or key set in the surgical robot 200 (for example, the user can press the "Endoscope Return" key). In some embodiments, the controller 240 can receive the message indicating the completion of lens cleaning, and in response to receiving the information indicating the completion of lens cleaning, detect the first surgical image sequence and / or the second surgical image sequence. In some embodiments, the first surgical image sequence and / or the second surgical image sequence from the endoscope can be detected when the endoscope is outside the patient's body. The technical details of detecting the first surgical image sequence and the second surgical image sequence can be referred to some of the above embodiments and will not be elaborated here. In some embodiments, if it is detected that the first surgical image sequence and the second surgical image sequence do not meet the abnormal condition, it can be indicated that the abnormal situation of the surgical image has been resolved. Thus, the controller 240 can, in response to the first surgical image sequence and the second surgical image sequence not meeting the abnormal condition, control the robotic arm 211 to move to extend the endoscope back into the patient's body, and then continue to perform the surgical operation. In some embodiments, if it is detected that the first surgical image sequence and the second surgical image sequence meet the abnormal condition, it indicates that the reason for the image abnormality is not that the endoscope lens is blocked by blood stains and the like. Thus, the controller 240 can, in response to the first surgical image sequence or the second surgical image sequence meeting the abnormal condition, send an endoscope replacement prompt message to prompt the user to replace the endoscope to eliminate the abnormal situation.
[0089] Figure 7 FIG. 700 is a flowchart showing a method for processing surgical images according to some embodiments of the present disclosure. The method 700 may be implemented or executed at least in part by hardware, software, or firmware. In some embodiments, the method 700 may be executed by a robotic system. The robotic system may be various suitable robotic systems including a surgical robotic system (e.g., Figure 2 the surgical robotic system 200 shown). The robotic system may also include dedicated or general-purpose robotic systems for other fields (e.g., manufacturing, machinery, etc.). In some embodiments, the method 700 may be executed at least in part by a controller ( Figure 2 not shown, e.g., Figure 3 the controller 240 shown) of the surgical robotic system 200. In some embodiments, the method 700 may be implemented as computer-readable instructions. These instructions may be read and executed by a general-purpose processor or a dedicated processor (e.g., the controller of the surgical robotic system 200). For example, the controller 240 of the surgical robotic system 200 may include a processor configured to execute the method 700. In some embodiments, these instructions may be stored on a computer-readable storage medium.
[0090] As Figure 7 shown, in step 710, a first sequence of surgical images is received from a first camera of an endoscope. In step 720, a second sequence of surgical images is received from a second camera of the endoscope. In step 730, the first sequence of surgical images is detected. In step 740, the second sequence of surgical images is detected. Steps 710-740 of the method 700 are similar to steps 110-140 of the method 100. For the technical details involved in steps 710-740, refer to the above.
[0091] As Figure 7 shown, in step 750, in response to the first sequence of surgical images and the second sequence of surgical images not satisfying an abnormal condition, temporal interpolation processing is performed on the first sequence of surgical images and the second sequence of surgical images. Temporal interpolation processing refers to obtaining interpolated images using the first sequence of surgical images (e.g., the variation law of the content of each image in the first sequence of surgical images over time) and inserting them into the first sequence of surgical images, or obtaining interpolated images using the second sequence of surgical images and inserting them into the second sequence of surgical images.
[0092] Common frame rates of endoscopes include 24fps (frames per second), 30fps, 60fps, etc. Those skilled in the art can understand that the higher the frame rate, the smoother the playback of video images. In step 750, the controller 240 of the surgical robot system 200 can perform temporal interpolation processing on the first surgical image sequence and the second surgical image sequence. Based on this, the frame rates of the first surgical image sequence and the second surgical image sequence can be increased, so that the images observed by the user on the display of the surgical robot are smoother, which is beneficial to improving the user experience.
[0093] Figure 8 A flowchart showing temporal interpolation processing of the first surgical image sequence and the second surgical image sequence according to some embodiments of the present disclosure. In some embodiments, performing temporal interpolation processing on the first surgical image sequence and the second surgical image sequence may include steps 810 to 830 as Figure 8 shown.
[0094] As Figure 8 shown, in step 810, at least one interpolation position is determined respectively in the first surgical image sequence and the second surgical image sequence. In some embodiments, the interpolation positions in the first surgical image sequence and the second surgical image sequence can be determined according to preset interpolation conditions. For example, the preset interpolation conditions may include: the number of frames of the image before the interpolation position in the surgical image sequence is not less than a preset number, etc.
[0095] Figure 9 A schematic diagram showing the first surgical image sequence 910 and the second surgical image sequence 920 according to some embodiments of the present disclosure. As Figure 9 shown, the first surgical image sequence may include first surgical images 911, 912, 913, etc. collected by the first camera of the endoscope, and the second surgical image sequence may include second surgical images 921, 922, 923, etc. collected by the second camera of the endoscope. As Figure 9 shown, the first surgical image 911 and the second surgical image 921 may be images collected by the first camera and the second camera at time t1 respectively; the first surgical image 912 and the second surgical image 922 are images collected by the first camera and the second camera at time t1 + Δt respectively; and so on for other surgical images.
[0096] For example, in the case where there are N frames of images before the image 911 in the first surgical image sequence and the preset number involved in the preset interpolation condition is N + 1, the interpolation position can be determined after the first surgical image 911. The present disclosure does not limit the preset number or the preset interpolation conditions. The preset number can be any suitable preset number, and the preset interpolation conditions can be any suitable interpolation conditions.
[0097] In some embodiments, the interpolation positions can also be determined according to the frame rate of the images acquired by the endoscope and the target frame rate of the played images. For example, when the frame rate of the endoscope is 30 fps and the target frame rate is 60 fps, the interpolation position can be determined as the position in the middle of two images acquired by the endoscope in the surgical image sequence. For example, it is the middle position between image 911 and image 912 shown in Figure 9 and the position corresponding to time t1 + 0.5Δt, the middle position between image 912 and image 913, and the position corresponding to time t1 + 1.5Δt, etc. When the frame rate of the endoscope is 30 fps and the target frame rate is 120 fps, three interpolation positions can be determined between image 911 and image 912, and between image 912 and image 913 as shown in Figure 9 , and the three interpolation positions can be evenly distributed.
[0098] As Figure 8 shown, in step 820, for each interpolation position, an interpolated image is obtained respectively based on at least one adjacent image before and adjacent to the interpolation position in the first surgical image sequence and the second surgical image sequence, and a pre-trained temporal interpolation model. In step 830, the obtained interpolated images are respectively inserted into at least one interpolation position in the first surgical image sequence and the second surgical image sequence.
[0099] Between adjacent first surgical images in the first surgical image sequence, there may be parts with unchanged content, such as images of the environment inside the patient's body (such as the abdominal cavity environment, etc.); there may also be parts with changing content, such as changes in the image content caused by the movement of the patient's tissues or surgical instruments, etc. Since the time interval between two adjacent frames of images is very short and the change in the content of the images is small, the movement of moving parts such as the patient's tissues can be regarded as linear motion. The change in the position of the moving part in at least one adjacent image before the interpolation position can be regarded as linear motion, and the position of the moving part in the interpolated image can be regarded as the position reached by the moving part continuing to do linear motion. Thus, the position of the moving part in the interpolated image can be determined based on the linear motion of the moving part. Based on this, in step 820, an interpolated image can be obtained based on at least one adjacent image before and adjacent to the interpolation position in the first surgical image sequence and the second surgical image sequence.
[0100] For example, in step 820, for Figure 9The interpolation position in the first surgical image sequence 910 shown (such as the position corresponding to time t1 + 0.5Δt, hereinafter also referred to as interpolation position t1 + 0.5Δt) can obtain an interpolated image 9101 based on the first surgical image 911 or at least one frame of image before image 911 and image 911. In step 830, the obtained interpolated image 9101 can be inserted into the first surgical image sequence 910 at the interpolation position t1 + 0.5Δt.
[0101] In some embodiments, a pre-trained temporal interpolation model can obtain an interpolated image based on at least one adjacent image at the interpolation position in the first surgical image sequence. In step 820, at least one adjacent image at the interpolation position in the first surgical image sequence can be input into the temporal interpolation model to obtain the interpolated image output by the temporal interpolation model.
[0102] Figure 10A A block diagram showing a temporal interpolation model 1000 according to some embodiments of the present disclosure. As Figure 10A shown, the temporal interpolation model 1000 can include an input layer 1010, a feature extraction layer 1020, and a backward prediction layer 1040. Among them, the input layer 1010 is used to receive at least one adjacent image before and adjacent to the interpolation position in the first surgical image sequence and the second surgical image sequence; the feature extraction layer 1020 is used to obtain the extracted features of at least one adjacent image based on at least one adjacent image from the input layer; and the backward prediction layer 1040 is used to obtain the interpolated image corresponding to the interpolation position based on the extracted features of at least one adjacent image from the feature extraction layer.
[0103] In some embodiments, as Figure 10A shown, the temporal interpolation model 1000 may further include an image weighting layer 1030, and the image weighting layer 1030 can be used to obtain multiple weighted images based on multiple adjacent images from the input layer and the extracted features of the adjacent images from the feature extraction layer 1020. In the process of predicting the interpolated image, the different degrees of dependence of the interpolated image on multiple adjacent images are taken into account, which is beneficial to improving the accuracy of the interpolated image prediction.
[0104] Taking at least one adjacent image including multiple adjacent images as an example, through the temporal interpolation model 1000 as Figure 10A shown, the following steps 1 to 5 can be used to obtain an interpolated image based on multiple adjacent images (such as images 911 and 912) before and adjacent to the interpolation position (for example, the interpolation position t1 + 1.5Δt in the first surgical image sequence 910) in the first surgical image sequence:
[0105] Step 1, the input layer 1010 receives the adjacent images 911 and 912 at the interpolation position t1 + 1.5Δt in the first surgical image sequence 910 of the input time interpolation model 1000.
[0106] Step 2, the feature extraction layer 1020 receives the adjacent images 911 and 912 from the input layer 1010, and performs feature extraction on the adjacent images 911 and 912 to obtain the first extracted features of the adjacent images 911 and 912. For example, the extracted features may include motion features and non - motion features.
[0107] In some embodiments, the feature extraction layer 1020 may include a feature generation sub - layer, an attention sub - layer, and a feature separation sub - layer (not shown in the figure) cascaded in sequence. In step 2, the feature generation sub - layer may be used to perform preliminary feature extraction on the adjacent images 911 and 912 received from the input layer 1010 to obtain the rough features of the adjacent images 911 and 912. The attention sub - layer may be used to perform attention processing (processing based on the attention mechanism) on the rough features of the adjacent images 911 and 912 received from the feature generation sub - layer to obtain the features to be separated of the adjacent images 911 and 912. The feature separation sub - layer may be used to perform separation processing on the features to be separated of the adjacent images 911 and 912 received from the attention sub - layer to obtain the first extracted features of the adjacent images 911 and 912, such as motion features and non - motion features. Among them, the non - motion features of the image can be used to describe the non - motion information of the image. For example, the background information in the image, such as the environment inside the patient's body (intra - abdominal cavity, natural body cavity, etc.). The motion features of the image can be used to describe the motion information of the image, such as the information of surgical instruments or the excised patient tissues, etc.
[0108] Step 3, the image weighting layer 1030 receives the extracted features of the adjacent images 911 and 912 from the feature extraction layer 1020, and the adjacent images 911 and 912 received from the input layer 1010, and performs weighting processing on the adjacent images 911 and 912 to obtain multi - frame weighted images.
[0109] In some embodiments, the image weighting layer 1030 may include a weight generation sub-layer and a weighting processing sub-layer (not shown in the figure). Among them, the weight generation sub-layer may be used to generate image weights corresponding to the adjacent images 911 and 912 according to the first extracted features of the adjacent images 911 and 912. The image weights are used to characterize the degree of dependence of the interpolated images on the adjacent images 911 and 912 respectively during the temporal interpolation process. In some embodiments, the image weights of an image may include the weights corresponding to each pixel point within the image. The weighting processing sub-layer may be used to perform weighting processing on the adjacent images 911 and 912 according to the image weights of the adjacent images 911 and 912 received from the weight generation sub-layer, and the adjacent images 911 and 912, to obtain the weighted adjacent images 911 and 912 after the weighting processing.
[0110] Step 4, the feature extraction layer 1020 receives the weighted adjacent images 911 and 912 from the image weighting layer 1030, and performs feature extraction on the weighted adjacent images 911 and 912 to obtain the weighted extraction features of the weighted adjacent images 911 and 912, for example, motion features and non-motion features.
[0111] The feature extraction layer 1020 may include a cascaded feature generation sub-layer, an attention sub-layer, and a feature separation sub-layer. In step 4, the feature generation sub-layer may be used to perform preliminary feature extraction on the weighted adjacent images 911 and 912 received from the image weighting layer 1030 to obtain the rough features of the weighted adjacent images 911 and 912. The attention sub-layer may be used to perform attention processing (processing based on the attention mechanism) on the rough features of the weighted adjacent images 911 and 912 received from the feature generation sub-layer to obtain the features to be separated of the weighted adjacent images 911 and 912. The feature separation sub-layer may be used to perform separation processing on the features to be separated of the weighted adjacent images 911 and 912 received from the attention sub-layer, to obtain the motion features and non-motion features of the weighted adjacent images 911 and 912.
[0112] Step 5, the backward prediction layer 1040 receives the weighted extraction features of the weighted adjacent images 911 and 912 from the feature extraction layer 1020, and obtains the interpolated image.
[0113] In some embodiments, the backward prediction layer 1040 may include a motion feature prediction sub-layer and a feature fusion sub-layer. Among them, the motion feature prediction sub-layer may be used to perform backward prediction on the motion features of the adjacent images 911 and 912 after weighted processing to obtain predicted motion features. The predicted motion features may be used to describe the motion information of the interpolated image, such as the information of surgical instruments or excised patient tissues in the interpolated image. The feature fusion sub-layer may be used to fuse the non-motion features of the adjacent images 911 and 912 after weighted processing received from the feature extraction layer 1020 and the predicted motion features received from the motion feature prediction sub-layer to obtain the interpolated image corresponding to the adjacent images 911 and 912.
[0114] In some embodiments, it may be set that when performing temporal interpolation processing, only one adjacent image before the interpolation position is referred to. For example, for the interpolation position t1 + 1.5Δt in the first surgical image sequence 910, the image 912 is used as the adjacent image input to the temporal interpolation model. In this case, the temporal interpolation model may only include an input layer, a feature extraction layer, and a backward prediction layer connected in series in sequence. The temporal interpolation model may obtain the interpolated image based on the adjacent image 912. The input layer may be used to receive the adjacent image 912, the feature extraction layer may be used to obtain the extracted features of the adjacent image 912 based on the adjacent image 912 from the input layer, and the backward prediction layer may be used to obtain the interpolated image based on the extracted features of the adjacent image 912 from the feature extraction layer. The technical details involved in the process of obtaining the interpolated image based on a single-frame adjacent image can refer to the above steps 1, 2, and 5. To avoid repetition, they will not be elaborated here.
[0115] The above part of the content is only introduced by taking the first surgical image sequence as an example. Those skilled in the art can understand that the temporal interpolation processing of the second surgical image sequence is similar. To avoid repetition, it will not be elaborated here.
[0116] The present disclosure does not limit the manner in which each layer and each sub-layer included in the temporal interpolation model implement their functions. Each layer and sub-layer can implement the functions that can be achieved in any suitable manner. In some embodiments, the temporal interpolation model may adopt a deep neural network model, such as a recurrent neural network, a convolutional neural network, or other network models. Those skilled in the art can understand that the above is only an example, and the temporal interpolation model can also adopt any other suitable network model. In addition, the present disclosure does not limit the detailed structure of the temporal interpolation model, and the temporal interpolation model can be any suitable structure.
[0117] In some embodiments, the historical first surgical image sequence and the historical second surgical image sequence (e.g., the first surgical image sequence and the second surgical image sequence collected during a previous surgical procedure) can be used as training data to train the temporal interpolation model respectively, so as to obtain two temporal interpolation models corresponding to the first surgical image sequence and the second surgical image sequence respectively. When performing temporal interpolation processing, the temporal interpolation model corresponding to the first surgical image sequence can be used to perform temporal interpolation processing on the first surgical image sequence, and the temporal interpolation model corresponding to the second surgical image sequence can be used to perform temporal interpolation processing on the second surgical image sequence, so as to improve the prediction effect of the temporal interpolation model, thereby facilitating the improvement of the user experience.
[0118] Those skilled in the art can understand that the structures of the two temporal interpolation models corresponding to the first surgical image sequence and the second surgical image sequence respectively are similar to the structure of the spatial interpolation model 1000 shown in Figure 10A For the sake of reducing repetition, no further description will be given.
[0119] As Figure 10A shown, the training method of the temporal interpolation model 1000 is similar to the training method of the spatial interpolation model introduced in some embodiments of the present disclosure, and will not be elaborated here. When training the temporal interpolation model, the simulated interpolation position can be determined in the surgical image set, and the image at the simulated interpolation position is determined as the output data in the data group. Furthermore, at least one frame of image before the simulated interpolation position can be determined as the adjacent image of the simulated interpolation position, and the adjacent image is used as the input data in the data group.
[0120] In some embodiments, method 700 may further include: obtaining the surgical procedure type; and for each interpolation position in the first surgical image sequence and the second surgical image sequence, obtaining an interpolated image respectively based on at least one adjacent image before and adjacent to the interpolation position in the first surgical image sequence and the second surgical image sequence, and the pre-trained temporal interpolation model corresponding to the surgical procedure type.
[0121] The surgical procedure type may include: radical prostatectomy, nephrectomy, pulmonary segmentectomy, lobectomy, endometrial cancer staging, radical resection of rectal cancer and other surgical procedures. In some embodiments, obtaining the surgical procedure type may include receiving the surgical procedure type selected by the user. In some embodiments, the user can select the surgical procedure type of the current surgery through the button provided on the handrest 222 of the main control trolley 220.
[0122] In some embodiments, historical image data of surgeries of different surgical procedure types (e.g., a first surgical image sequence and a second surgical image sequence) can be used to train a temporal interpolation model corresponding to the surgical procedure type. For example, multiple sets of image data (including a first surgical image sequence and a second surgical image sequence) collected during multiple previous lobectomies are used for model training to obtain a temporal interpolation model corresponding to lobectomy. This temporal interpolation model can be specifically used for performing temporal interpolation processing on the first surgical image sequence or the second surgical image sequence collected during lobectomy. Based on the above embodiments, using the temporal interpolation model corresponding to the surgical procedure type to obtain interpolation images for inserting into the first surgical image sequence or the second surgical image sequence is beneficial to improving the accuracy of the interpolation images, thereby helping to improve the smoothness of the first surgical image sequence or the second surgical image sequence and enhancing the user experience.
[0123] In some embodiments, the temporal interpolation model corresponding to the surgical procedure type may further include a first surgical image sequence temporal interpolation model and a second surgical image sequence temporal interpolation model, which are respectively used for performing temporal interpolation processing on the first surgical image sequence and the second surgical image sequence. For example, in the case of performing a lobectomy, the first surgical image sequence temporal interpolation model corresponding to lobectomy can be used to perform temporal interpolation processing on the first surgical image sequence, and the second surgical image sequence temporal interpolation model corresponding to lobectomy can be used to perform temporal interpolation processing on the second surgical image sequence.
[0124] The structure of the temporal interpolation model corresponding to each surgical procedure type and the functions that each layer can achieve are similar to Figure 10A the temporal interpolation model 1000 shown. To avoid repetition, it will not be elaborated further.
[0125] In some embodiments, method 700 may further include: obtaining the surgical stage at which at least one interpolation position is located; and for each interpolation position in the first surgical image sequence and the second surgical image sequence, respectively obtaining an interpolation image based on the surgical stage of the interpolation position, at least one adjacent image before and adjacent to the interpolation position in the first surgical image sequence and the second surgical image sequence, and the pre-trained temporal interpolation model corresponding to the surgical procedure type. In this embodiment, considering the surgical stage at which the interpolation position is located during the process of obtaining the interpolation image is beneficial to improving the accuracy of the obtained interpolation image and enhancing the user experience.
[0126] In some embodiments, the surgical stage at the interpolation position can also be obtained through a pre-trained surgical stage recognition model. In some embodiments, for each interpolation position, the surgical stage can be obtained based on at least one adjacent image of the interpolation position and the pre-trained surgical stage recognition model, and the obtained surgical stage can be used as the surgical stage at the interpolation position. Figure 10B A block diagram showing a surgical stage recognition model 1100 and a temporal interpolation model 1000 according to some embodiments of the present disclosure. As Figure 10B shown, the surgical stage recognition model 1100 can include a fourth input layer 1101, a fourth feature extraction layer 1102, and a surgical stage determination layer 1103.
[0127] Among them, the fourth input layer 1101 can be used to receive at least one adjacent image (e.g., Figure 9 the interpolation position t1 + 1.5Δt in the first surgical image sequence 910 shown) of the interpolation position.
[0128] The fourth feature extraction layer 1102 can be used to obtain a second extracted feature of the adjacent image 912 based on the adjacent image 912 from the fourth input layer 1101. In some embodiments, the fourth feature extraction layer 1103 can include a fourth feature generation sub-layer, a fourth attention sub-layer, and a fourth feature separation sub-layer (not shown in the figure). The functions that can be achieved by the fourth feature generation sub-layer, the fourth attention sub-layer, and the fourth feature separation sub-layer will not be elaborated here.
[0129] The surgical stage determination layer 1103 can be used to determine the surgical stage of the adjacent image based on the second extracted feature of the adjacent image 912 from the fourth feature extraction layer 1102, so as to determine the surgical stage at the interpolation position t1 + 1.5Δt.
[0130] In some embodiments, the surgical stage output by the surgical stage recognition model 1100 can be transmitted to the backward prediction layer 1040 in the temporal interpolation model 1000. Based on this, the backward prediction layer 1040 can perform backward prediction based on the extracted feature of the adjacent image 912 from the feature extraction layer 1020 and the surgical stage at the interpolation position to obtain an interpolated image.
[0131] As Figure 10BAs shown, in some embodiments, the feature extraction layer 1020 in the temporal interpolation model 1000 may receive, as shared features, at least some of the second extracted features of the adjacent images 912 from the fourth feature extraction layer 1102 of the surgical stage recognition model 1100. In some embodiments, the feature extraction layer 1020 may perform further feature extraction based on the second extracted features of the adjacent images 912 from the fourth feature extraction layer 1102 to obtain more detailed first extracted features of the adjacent images 912.
[0132] In some embodiments, during the temporal interpolation process of the first surgical image sequence and the second surgical image sequence, the first surgical image sequence and the second surgical image sequence may be continuously detected, for example, the first surgical image sequence and the second surgical image sequence may be detected at a preset period. In some embodiments, when it is detected that the first surgical image sequence or the second surgical image sequence meets the abnormal condition, method 700 may further include: in response to the first surgical image sequence or the second surgical image sequence meeting the abnormal condition, stopping the temporal interpolation process of the first surgical image sequence and the second surgical image sequence. In some embodiments, after stopping the temporal interpolation process, spatial interpolation may also be performed on the first surgical image sequence or the second surgical image sequence to reduce the impact of the abnormal image. The solution for performing spatial interpolation on the first surgical image sequence or the second surgical image sequence may refer to some embodiments of the present disclosure, and will not be elaborated here to avoid repetition.
[0133] In some embodiments, the controller 240 of the surgical robot 200 may also detect that it is unable to receive the first surgical image sequence or the second surgical image sequence. In some embodiments, the controller 240 of the surgical robot 200 may also, in response to being unable to receive the first surgical image sequence or the second surgical image sequence, perform spatial interpolation on the first surgical image sequence or the second surgical image sequence using the second surgical image sequence or the first surgical image sequence. Based on this, it is possible to avoid the sudden disappearance of the image displayed on the monitor and have a negative impact on the surgical operation being performed by the surgical robot user. The solution for performing spatial interpolation on the first surgical image sequence or the second surgical image sequence can be found in some embodiments of the present disclosure.
[0134] In some embodiments, the controller 240 of the surgical robot 200 may also, in response to being unable to receive the first surgical image sequence or the second surgical image sequence, send an alarm message to prompt the surgical robot user of the occurrence of an abnormal situation, so that the user can perform corresponding processing. In some embodiments, the controller 240 of the surgical robot 200 may also, in response to being unable to receive the first surgical image sequence or the second surgical image sequence, send an endoscope replacement prompt message to prompt the user to replace the endoscope, so that the surgical robot system can resume normal use.
[0135] Some embodiments of the present disclosure also provide a surgical robot system 200. Figure 3 The structural schematic block diagram of the surgical robot system 200 according to some embodiments of the present disclosure is shown. As Figure 3 shown, the surgical robot system 200 may include an endoscope 213, a controller 240, and a display 224. The endoscope 213 may include a first camera 213a and a second camera 213b. The first camera 213a and the second camera 213b may be the left camera and the right camera of the endoscope 213 respectively. As Figure 3 shown, the controller 240 may be communicatively connected to the endoscope, and the controller 240 may be configured to be capable of executing the processing method of the surgical image according to any one of some embodiments of the present disclosure. As Figure 3 shown, the display 224 may be communicatively connected to the controller, and the display 224 may be used to display the first surgical image sequence and / or the second surgical image sequence obtained by processing through the processing method of the surgical image (for example, method 100, 700, etc.).
[0136] Figure 3 The schematic diagram of the surgical robot system 200 according to some embodiments of the present disclosure is shown. As Figure 3 shown, the surgical robot system 200 may further include a surgical trolley 210, a main control trolley 220, and an equipment trolley 230. The surgical trolley 210 may include at least one moving arm 211, and the at least one moving arm 211 is movably arranged on the surgical trolley 210. The main control trolley 220 may be communicatively connected to the surgical trolley 210 and include at least one master operator 221 for receiving user operations. Figure 3 shown, in some embodiments, the at least one master operator 221 may include a left master operator 221a for receiving the operations of the user's left hand and a right master operator 221b for receiving the operations of the user's right hand. The mobile station 210 is usually located on the patient side and performs surgery on the patient in response to the control instructions of the main control station 220. During the surgery, the user may control at least one surgical instrument carried by the mobile station 210 to perform surgical operations by operating at least one master operator 221 in the main control station 220.
[0137] The endoscope 213 can be disposed on the operation cart 210. For example, it can be mounted at the end of the robotic arm 211. The controller 240 can be disposed on the equipment cart 230, and the equipment cart 230 can be connected to the operation cart 210 so that the controller 240 can receive the surgical images collected by the endoscope 213 from the endoscope 213. The main control cart 220 can be connected to the equipment cart 230 to receive the surgical images processed by the controller 240 from the controller 240. In some embodiments, the surgical robot system 200 may further include a display 212 disposed on the operation cart 210, displays 224 and 225 disposed on the main control cart 220, and a display 231 disposed on the equipment cart 230. At least one of these displays can be used to display the surgical images processed by the processor 240.
[0138] Those skilled in the art can understand that the surgical robot system 200 provided in this embodiment can be any suitable surgical robot system including a laparoscopic surgical robot system.
[0139] In some embodiments, the present disclosure provides a computer-readable storage medium, which can be used to store at least one instruction. When the at least one instruction is executed by a computer, it causes the computer to execute the method for processing surgical images in any one of some embodiments of the present disclosure (for example, method 100, 700, etc.).
[0140] In some embodiments, the computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above.
[0141] From the above description of the embodiments, those skilled in the art can clearly understand that the present disclosure can be implemented by means of software and necessary general hardware, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium. In some embodiments, the computer-readable storage medium may include, but is not limited to: portable computer disks, hard disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technologies, CD-ROM, digital versatile disk (DVD), HD-DVD, Blu-ray or other optical storage devices, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the required information and can be accessed by a computer, on which computer-executable instructions are stored. When the computer-executable instructions run on a machine (such as a computer device), the machine executes the control method of the present disclosure. It should be understood that the computer device may include a personal computer, a server or a network device, etc.
[0142] In some embodiments, when the first surgical image sequence (e.g., the left-eye surgical image sequence) from the first camera of the endoscope (e.g., the left camera of the endoscope) meets an abnormal condition, the controller 240 of the surgical robot system 200 can perform spatial interpolation processing on the first surgical image sequence, enabling the user to continue to view a complete and clear picture of the first surgical image sequence on the display. Based on this, it helps the user complete the surgical operation currently being performed and is beneficial to reducing the negative impact caused by abnormal surgical images.
[0143] In some embodiments, the controller 240 of the surgical robot system 200 can perform temporal interpolation processing on the first surgical image sequence and the second surgical image sequence. Based on this, the frame rate of the first surgical image sequence and the second surgical image sequence can be increased, so that the images observed by the user on the display of the surgical robot are smoother, which is beneficial to improving the user experience.
[0144] In some embodiments, the historical first surgical image sequence and the historical second surgical image sequence can be used as training data to train a spatial interpolation model and a temporal interpolation model respectively, so as to obtain two spatial interpolation models and two temporal interpolation models corresponding to the first surgical image sequence and the second surgical image sequence respectively. When performing spatial interpolation processing or temporal interpolation processing, the spatial interpolation model or the temporal interpolation model corresponding to the first surgical image sequence can be used to perform spatial interpolation processing or temporal interpolation processing on the first surgical image sequence, and the spatial interpolation model or the temporal interpolation model corresponding to the second surgical image sequence can be used to perform spatial interpolation processing or temporal interpolation processing on the second surgical image sequence, so as to improve the effect of the spatial interpolation model or the temporal interpolation model and improve the accuracy of the obtained interpolated images, thereby facilitating the improvement of the user experience.
[0145] In some embodiments, the spatial interpolation model or the temporal interpolation model corresponding to the surgical procedure type can be used to obtain the interpolated images for inserting into the first surgical image sequence or the second surgical image sequence, which is conducive to improving the accuracy of the interpolated images, thereby helping to improve the smoothness of the first surgical image sequence or the second surgical image sequence and enhancing the user experience.
[0146] In some embodiments, when performing temporal interpolation processing or spatial interpolation processing, the surgical stage where the interpolation position is located can also be input into the temporal interpolation model or the spatial interpolation model, so as to refer to the surgical stage where the interpolation position is located when predicting the interpolated images, thereby facilitating the improvement of the accuracy of the interpolated image prediction and helping to enhance the user's feeling.
[0147] Note that the above are only exemplary embodiments of the present disclosure and the applied technical principles. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present disclosure. Therefore, although the present disclosure has been described in detail through the above embodiments, the present disclosure is not limited to the above embodiments only. Without departing from the concept of the present disclosure, more other equivalent embodiments can be included, and the scope of the present disclosure is determined by the scope of the appended claims.
Claims
1. A method for processing surgical images, characterized in that, Comprising: Receiving a first sequence of surgical images from a first camera of an endoscope; Receiving a second sequence of surgical images from a second camera of the endoscope; Detecting the first sequence of surgical images; Detecting the second sequence of surgical images; And In response to the first sequence of surgical images or the second sequence of surgical images satisfying an abnormal condition, using the second sequence of surgical images or the first sequence of surgical images to perform spatial interpolation processing on the first sequence of surgical images or the second sequence of surgical images.
2. The processing method of surgical images according to claim 1, characterized in that, Further comprising: In response to the first sequence of surgical images and the second sequence of surgical images not satisfying the abnormal condition, performing temporal interpolation processing on the first sequence of surgical images and the second sequence of surgical images.
3. The processing method of the surgical image according to claim 2, wherein, Performing temporal interpolation processing on the first sequence of surgical images and the second sequence of surgical images includes: Respectively determining at least one interpolation position in the first sequence of surgical images and the second sequence of surgical images; For each interpolation position, respectively obtaining an interpolated image based on at least one adjacent image before and adjacent to the interpolation position in the first sequence of surgical images and the second sequence of surgical images and a pre-trained temporal interpolation model; and Inserting the obtained interpolated images into at least one interpolation position in the first sequence of surgical images and the second sequence of surgical images respectively.
4. The processing method of the surgical image according to claim 2 or 3, characterized in that, Further comprising: In response to the first sequence of surgical images or the second sequence of surgical images satisfying the abnormal condition, stopping performing temporal interpolation processing on the first sequence of surgical images and the second sequence of surgical images.
5. The method for processing surgical images according to claim 1, wherein Detecting the first sequence of surgical images includes: detecting whether a first surgical image in the first sequence of surgical images from the first camera satisfies an abnormal condition; and / or Detecting the second sequence of surgical images includes: detecting whether a second surgical image in the second sequence of surgical images from the second camera satisfies an abnormal condition; The abnormal condition includes at least one of the following: image freeze, image blackout, image blur, or image occlusion.
6. The processing method of the surgical image according to claim 5, wherein, Detecting whether the first surgical image or the second surgical image satisfies the abnormal condition includes: Performing fast Fourier transform processing on the first surgical image or the second surgical image to obtain the frequency domain coefficients of the first surgical image or the second surgical image; Detecting whether the high-frequency coefficients in the frequency domain coefficients of the first surgical image or the second surgical image are greater than a preset threshold; and In response to detecting that the high-frequency coefficients of the first surgical image or the second surgical image are greater than the preset threshold, determining that the first surgical image or the second surgical image satisfies the abnormal condition; or Performing fast Fourier transform processing on the first surgical image or the second surgical image to obtain the frequency domain coefficients of the first surgical image or the second surgical image; Detecting whether the proportion of the high-frequency components in the frequency domain coefficients of the first surgical image or the second surgical image is less than a preset threshold; and In response to detecting that the proportion of high-frequency components in the frequency-domain coefficients of the first surgical image or the second surgical image is less than a preset threshold, determining that the first surgical image or the second surgical image meets the abnormal condition; or Detecting whether the proportion of the area of the occlusion screen in the first surgical image or the second surgical image is greater than a preset proportion; and In response to the proportion of the area of the occlusion screen in the first surgical image or the second surgical image being greater than a preset proportion, determining that the first surgical image or the second surgical image meets the abnormal condition; or Performing grayscale processing on the first surgical image or the second surgical image; Detecting whether the proportion of pixel points whose grayscale values satisfy a preset range in the first surgical image or the second surgical image is greater than a preset value; and In response to the proportion of pixel points whose grayscale values satisfy a preset range in the first surgical image or the second surgical image being greater than a preset value, determining that the first surgical image or the second surgical image meets the abnormal condition.
7. The processing method of surgical images according to claim 1, wherein Further comprising: In response to being unable to receive the first surgical image sequence or the second surgical image sequence, using the second surgical image sequence or the first surgical image sequence to perform spatial interpolation processing on the first surgical image sequence or the second surgical image sequence.
8. The method for processing a surgical image according to claim 7, wherein Using the second surgical image sequence to perform spatial interpolation processing on the first surgical image sequence includes: Determining at least one interpolation position in the first surgical image sequence; For each interpolation position, obtaining an interpolated image based on at least one adjacent image in the first surgical image sequence that is before and adjacent to the interpolation position, a reference image in the second surgical image sequence with the same acquisition time as the interpolation position, and a pre-trained spatial interpolation model; and Inserting the obtained interpolated image into the at least one interpolation position in the first surgical image sequence; or Using the first surgical image sequence to perform spatial interpolation processing on the second surgical image sequence includes: Determining at least one interpolation position in the second surgical image sequence; For each interpolation position, obtaining an interpolated image based on at least one adjacent image in the second surgical image sequence that is before and adjacent to the interpolation position, a reference image in the first surgical image sequence with the same acquisition time as the interpolation position, and a pre-trained spatial interpolation model; and Inserting the obtained interpolated image into the at least one interpolation position in the second surgical image sequence.
9. The processing method of surgical images according to claim 7, characterized in that, Further comprising: In response to being unable to receive the first surgical image sequence or the second surgical image sequence, sending an alarm message and / or sending an endoscope replacement prompt message; and / or In response to the first surgical image sequence or the second surgical image sequence meeting the abnormal condition, sending an alarm message and / or a lens cleaning prompt message.
10. The processing method of the surgical image according to claim 9, characterized in that, Further comprising: In response to receiving the information indicating that the lens cleaning is completed, detecting the first surgical image sequence and / or the second surgical image sequence; and In response to the first surgical image sequence or the second surgical image sequence satisfying an abnormal condition, an endoscope replacement prompt message is issued.
11. The method for processing surgical images according to claim 3 or 8, characterized in that It further includes: Obtaining the surgical procedure type of the surgery; And For each interpolation position in the first surgical image sequence and the second surgical image sequence, based on at least one adjacent image before and adjacent to this interpolation position in the first surgical image sequence and the second surgical image sequence, and a pre-trained temporal interpolation model corresponding to the surgical procedure type, an interpolated image is obtained; Or For each interpolation position in the first surgical image sequence, based on at least one adjacent image before and adjacent to this interpolation position in the first surgical image sequence, a reference image in the second surgical image sequence with the same acquisition time as this interpolation position, and a pre-trained spatial interpolation model corresponding to the surgical procedure type, an interpolated image is obtained; Or For each interpolation position in the second surgical image sequence, based on at least one adjacent image before and adjacent to this interpolation position in the second surgical image sequence, a reference image in the first surgical image sequence with the same acquisition time as this interpolation position, and a pre-trained spatial interpolation model corresponding to the surgical procedure type, an interpolated image is obtained.
12. The method for processing surgical images according to claim 11, wherein It further includes: Obtaining the surgical stage where at least one interpolation position is located; And For each interpolation position in the first surgical image sequence and the second surgical image sequence, based on the surgical stage of this interpolation position, at least one adjacent image before and adjacent to this interpolation position in the first surgical image sequence and the second surgical image sequence, and a pre-trained temporal interpolation model corresponding to the surgical procedure type, an interpolated image is obtained; Or For each interpolation position in the first surgical image sequence, based on the surgical stage of this interpolation position, at least one adjacent image before and adjacent to this interpolation position in the first surgical image sequence, a reference image in the second surgical image sequence with the same acquisition time as this interpolation position, and a pre-trained spatial interpolation model corresponding to the surgical procedure type, an interpolated image is obtained; Or For each interpolation position in the second surgical image sequence, based on the surgical stage of this interpolation position, at least one adjacent image before and adjacent to this interpolation position in the second surgical image sequence, a reference image in the first surgical image sequence with the same acquisition time as this interpolation position, and a pre-trained spatial interpolation model corresponding to the surgical procedure type, an interpolated image is obtained.
13. The processing method of the surgical image according to claim 12, characterized in that, Obtaining the surgical stage where at least one interpolation position is located includes: For each interpolation position in the first surgical image sequence, based on a reference image in the second surgical image sequence with the same acquisition time as this interpolation position and a pre-trained surgical stage recognition model, the surgical stage is obtained as the surgical stage where at least one interpolation position is located; and / or For each interpolation position in the second surgical image sequence, based on the reference image in the first surgical image sequence with the same acquisition time as the interpolation position and a pre-trained surgical stage recognition model, obtain the surgical stage, which is the surgical stage where the at least one interpolation position is located.
14. The processing method of the surgical image according to claim 12, wherein The surgical stages include: Surface fat cleaning stage, tissue dissection stage, blood vessel separation stage, lesion resection stage, tissue suture stage.
15. The method for processing surgical images according to claim 1 or 2, wherein The spatial interpolation model includes: A backward prediction part, configured to obtain the extracted features and the predicted interpolation image of the at least one adjacent image based on at least one adjacent image in the first surgical image sequence before and adjacent to the interpolation position; A spatial reference part, configured to obtain the extracted features of the reference image based on the reference image in the second surgical image sequence with the same acquisition time as the interpolation position; and An image generation part, configured to obtain the interpolation image based on the at least one adjacent image, the extracted features of the at least one adjacent image, the reference image, and the extracted features of the reference image; and / or The temporal interpolation model includes: An input layer, configured to receive the first surgical image sequence and at least one adjacent image in the second surgical image sequence before and adjacent to the interpolation position; A feature extraction layer, configured to obtain the extracted features of the at least one adjacent image based on the at least one adjacent image from the input layer; and A backward prediction layer, configured to obtain the interpolation image corresponding to the interpolation position based on the extracted features of at least one adjacent image from the feature extraction layer.
16. The method for processing surgical images according to claim 15, wherein The backward prediction part includes: A first input layer, configured to receive the at least one adjacent image; A first feature extraction layer, configured to obtain the extracted features of the at least one adjacent image based on the at least one adjacent image from the first input layer; and A first backward prediction layer, configured to obtain the predicted interpolation image based on the extracted features of the at least one adjacent image from the first feature extraction layer; The spatial reference part includes: A second input layer, configured to receive the reference image; and A second feature extraction layer, configured to obtain the extracted features of the reference image based on the reference image from the second input layer; and The image generation part includes: An image fusion layer, configured to obtain a fused interpolation image based on the extracted features of the at least one adjacent image from the first feature extraction layer, the extracted features of the reference image from the second feature extraction layer, the predicted interpolation image from the first backward prediction layer, and the reference image from the second input layer; and An output layer, configured to obtain the interpolation image based on the fused interpolation image from the image fusion layer.
17. A surgical robot system, characterized in that, Includes: An endoscope, including a first camera and a second camera; A controller, configured to be capable of executing the method for processing surgical images according to any one of claims 1-16; And A display, connected to the controller, for displaying a first surgical image sequence and / or a second surgical image sequence obtained by processing the surgical images using the method for processing surgical images.
18. The surgical robot system according to claim 17, wherein, It further includes: A surgical cart, on which the endoscope is provided; An equipment cart, connected to the surgical cart, with the controller provided on the equipment cart; And A main control cart, connected to the equipment cart, including at least one main operator for receiving user operations; The display is provided on at least one of the surgical cart, the equipment cart, and the main control cart.
19. A computer-readable storage medium, characterized in that, For storing at least one instruction, which when executed by a computer causes the computer to execute the method for processing surgical images according to any one of claims 1 to 16.