Remote control method and device for surgical robot
By constructing a high-bandwidth, low-latency communication network and an image prediction model, the technological gap in remote control of surgical robots has been filled, enabling real-time control of surgical robots and precise robotic arm operation in remote environments.
Patent Information
- Application Number
- CN202510762290.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-28
AI Technical Summary
Currently, there is a lack of remote control solutions for surgical robots, making it impossible to effectively assist in surgery when highly skilled doctors cannot be present on-site.
A high-bandwidth, low-latency communication network is constructed to acquire intraoperative 3D image information through a remote information transmission channel. Based on preoperative planning and intraoperative image information, robotic arm control information is generated, and the robotic arm is controlled to perform operations in real time. Data transmission is optimized using keyframe extraction and image prediction models.
This enables real-time control of the surgical robot in a remote environment, ensuring a clear surgical field of view and precise robotic arm operation, while reducing the impact of network latency on visual feedback.
Smart Images

Figure CN120837210A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image processing technology, and more specifically, to a method and apparatus for remote control of a surgical robot. Background Technology
[0002] Current surgical robots have made great strides in assisting surgery, providing excellent support to surgeons during procedures. However, in some situations, highly skilled surgeons may be unable to reach the site, necessitating a method for remotely controlling surgical robots. Summary of the Invention
[0003] The problem addressed in this application is the current lack of remote control solutions for surgical robots.
[0004] To address the aforementioned problems, the first aspect of this application provides a method for remote control of a surgical robot, comprising:
[0005] Establish a remote information transmission channel;
[0006] Intraoperative three-dimensional image information is acquired based on the transmission channel;
[0007] Based on preoperative planning information and intraoperative 3D imaging information, robotic arm control information is generated.
[0008] Transmit and control the robotic arm to execute the robotic arm control information.
[0009] The second aspect of this application provides a manufacturing system for a method of remotely controlling a surgical robot, comprising:
[0010] The channel construction module is used to build remote information transmission channels;
[0011] Information transmission module, which is used to acquire intraoperative three-dimensional image information based on the transmission channel;
[0012] The control generation module is used to generate robotic arm control information based on preoperative planning information and intraoperative 3D image information;
[0013] A robotic arm control module is used to transmit and control the robotic arm to perform the robotic arm control information.
[0014] A third aspect of this application provides an electronic device, including: a memory and a processor; the memory being configurable to store a program, and the processor being coupled to the memory for executing the program in the memory for:
[0015] Establish a remote information transmission channel;
[0016] Intraoperative three-dimensional image information is acquired based on the transmission channel;
[0017] Based on preoperative planning information and intraoperative 3D imaging information, robotic arm control information is generated.
[0018] Transmit and control the robotic arm to execute the robotic arm control information.
[0019] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the aforementioned remote control method for a surgical robot.
[0020] In this application, intraoperative image information is transmitted through a transmission channel, thereby generating robotic arm control information at a remote location and completing robotic arm control, thus realizing remote control of the surgical robot. Attached Figure Description
[0021] Figure 1 This is a flowchart of a remote control method for a surgical robot according to an embodiment of this application;
[0022] Figure 2 This is an architecture diagram of a segmented network for a surgical robot remote control method according to an embodiment of this application;
[0023] Figure 3 This is an architectural diagram of the displacement field module of the surgical robot remote control method according to an embodiment of this application;
[0024] Figure 4 This is an architecture diagram of the class boundary module of the surgical robot remote control method according to an embodiment of this application;
[0025] Figure 5 This is an architectural diagram of a remote control device for a surgical robot according to an embodiment of this application;
[0026] Figure 6 This is an architectural diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0027] To make the above-mentioned objects, features, and advantages of this application more apparent and understandable, specific embodiments of this application will be described in detail below with reference to the accompanying drawings. Although exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0028] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application shall have the ordinary meaning as understood by one of ordinary skill in the art.
[0029] This application provides a remote control method for a surgical robot as described above, the specific scheme of which is provided by... Figures 1-2 As shown, this method can be executed by a remote control device for a surgical robot, which can be integrated into electronic devices such as computers, servers, computer clusters, and data centers. Combined with... Figure 1 As shown, the remote control method for the surgical robot includes:
[0030] S101, establish a remote information transmission channel;
[0031] In this application, a high-bandwidth, low-latency communication network is established to transmit intraoperative images, control commands, and sensor data in real time as much as possible.
[0032] Preferably, end-to-end encryption (such as TLS / SSL protocols) is implemented to prevent data leakage or tampering. An authentication mechanism is configured to ensure that only authorized users can access the system.
[0033] S102, acquire intraoperative three-dimensional image information based on the transmission channel;
[0034] In this application, three-dimensional image information of the surgical area is acquired and transmitted in real time, providing doctors with a clear surgical field of vision.
[0035] In this application, a binocular camera or a depth sensor is used to acquire images and depth information of the surgical area.
[0036] S103 generates robotic arm control information based on preoperative planning information and intraoperative 3D imaging information;
[0037] In this application, precise robotic arm control commands are generated based on preoperative planning and real-time imaging information.
[0038] S104, transmit and control the robotic arm to execute the robotic arm control information.
[0039] In this application, the generated control commands are transmitted to the robotic arm control system, and the execution process is monitored in real time.
[0040] In this application, intraoperative image information is transmitted through a transmission channel, thereby generating robotic arm control information at a remote location and completing robotic arm control, thus realizing remote control of the surgical robot.
[0041] In one specific embodiment, step S102, acquiring intraoperative three-dimensional image information based on the transmission channel, includes:
[0042] Image and video streams for acquiring intraoperative 3D imaging information;
[0043] Identify keyframes in image and video streams;
[0044] The key frame is transmitted based on the transmission channel;
[0045] Based on the keyframes, a video stream is generated.
[0046] In this application, key frames that can represent the surgical scene are extracted from the image and video stream to reduce the amount of data transmitted.
[0047] Preferably, the keyframe density is reduced in areas with less dynamic change and increased in areas with frequent action.
[0048] In this application, extracted keyframes are sent to the remote master control terminal via a transmission channel to ensure the real-time performance and integrity of the data. Transmission priorities are assigned according to the importance of the keyframes to ensure that important frames arrive first.
[0049] In this application, frame interpolation algorithms (such as optical flow interpolation, DAIN, RIFE) are used to generate transition frames between keyframes. The keyframes and interpolated / predicted frames are then combined in chronological order to form a complete video sequence.
[0050] In this application, based on keyframe extraction and video stream generation, the amount of data transmission can be significantly reduced in remote surgery while ensuring that doctors receive clear and continuous visual feedback. Through frame interpolation and frame prediction techniques, the smoothness of the video stream can be maintained even with high network latency.
[0051] In one specific implementation, identifying keyframes in an image video stream includes:
[0052] Obtain the keyframe threshold;
[0053] The image / video stream is split into sequentially set image frames;
[0054] Calculate the histogram difference between adjacent image frames;
[0055] If the histogram difference is greater than the keyframe threshold, the current image frame is considered a keyframe.
[0056] In this application, the choice of threshold depends on the degree of change in the surgical scenario: if the scenario changes significantly (e.g., the surgical instruments move frequently), a lower threshold can be set; if the scenario is relatively stable (e.g., during the observation phase), a higher threshold can be set.
[0057] Preferably, historical data is used for testing, and the keyframe extraction performance (such as coverage and redundancy) under different thresholds is statistically analyzed. The optimal threshold is found through cross-validation or grid search.
[0058] In this application, histogram calculation: calculate the color histogram (RGB or grayscale histogram) for each frame.
[0059] In this application, a distance metric method is used to calculate the difference between histograms of adjacent frames.
[0060] Preferably, the histogram differences are normalized so that their range is within [0,1] or other fixed intervals, which facilitates comparison with the threshold.
[0061] Preferably, to avoid multiple consecutive frames being marked as keyframes, a minimum interval (e.g., at least 5 frames) can be set.
[0062] In one specific implementation, after transmitting the key frame based on the transmission channel, the method further includes:
[0063] Real-time monitoring of network latency in the transmission channel;
[0064] When network latency exceeds a preset threshold, an image prediction model is obtained;
[0065] Based on the image prediction model, generate the prediction frame for the next moment corresponding to the current keyframe;
[0066] The video stream is generated based on the predicted frames.
[0067] In this application, delay measurement uses timestamp technology: a timestamp is added to each frame at the sending end, and the receiving end calculates the difference between the current time and the timestamp to obtain the delay.
[0068] In this application, a background thread or service is set up to collect and analyze latency data in real time. A sliding window algorithm is used to calculate the average latency and jitter to avoid misjudgment based on a single fluctuation.
[0069] In this application, when network latency is high, an image prediction model is used to generate a predicted frame for the next time step. The predicted frame is then combined with the real frame to generate a continuous and smooth video stream.
[0070] Preferably, if the network recovers, a smooth transition is made between the predicted and real frames upon their arrival. Interpolation techniques (such as linear interpolation or optical flow interpolation) are used to reduce the visual jump between the predicted and real frames.
[0071] The latency compensation mechanism based on an image prediction model in this application can significantly improve the quality of visual feedback in remote surgery. By monitoring network latency in real time and generating predicted frames, the continuity of the video stream can be maintained even under high latency conditions.
[0072] In one specific implementation, generating the predicted frame for the next moment corresponding to the current keyframe based on the image prediction model includes:
[0073] Input the current keyframe into the segmentation network to obtain the segmentation mask;
[0074] The current keyframe and the segmentation mask are input into the prediction network to obtain the prediction frame for the next time step.
[0075] In this application, the prediction network adopts a conditional generative adversarial network (cGAN) architecture: the generator generates prediction frames based on the input, and the discriminator is used to distinguish between generated frames and real frames to improve the generation quality.
[0076] In this application, the features of keyframes and segmentation masks are fused to ensure that the prediction network can utilize both visual and semantic information simultaneously.
[0077] In one specific implementation, the training process of the image prediction model includes:
[0078] Obtain a first sample frame, on which a real mask is marked;
[0079] The segmentation network is trained based on the first sample frame and the real mask to obtain the trained segmentation network.
[0080] Acquire a second sample frame, on which the image frame for the next moment is marked;
[0081] Keeping the parameters of the segmentation network unchanged, the image prediction model is trained based on the second sample frame to obtain the trained image prediction model.
[0082] In this application, consecutive video frame pairs are collected, with the previous frame used as the second sample frame and the next frame used as the labeled image frame.
[0083] In this application, an image prediction model is trained based on a segmentation network, enabling it to generate a prediction frame for the next time step.
[0084] In this application, the image prediction model is trained based on the second sample frame, which means that the prediction network in the image prediction model is trained.
[0085] The prediction network is a conditional generative adversarial network, and the segmentation network outputs a segmentation mask and the current keyframe. During the training of the conditional generative adversarial network, the segmentation mask is added as additional conditional information to both the generator and the discriminator.
[0086] In this application, the semantic information provided by the segmentation mask is fully utilized to improve the quality of the predicted frame.
[0087] In this application, the specific process of training the image prediction model based on the second sample frame includes:
[0088] Initialize the parameters of the prediction network (generator G and discriminator D), and load the parameters of the segmentation network;
[0089] With the parameters of generator G and segmentation network fixed, discriminator D is trained: the second sample frame is input into the segmentation network to obtain a segmentation mask; the second sample frame and the segmentation mask are input into generator G to obtain a prediction frame; the prediction frame / annotated image frame and the segmentation mask are input into discriminator D to obtain a discrimination result; the overall loss is calculated based on the discrimination result, and the parameters of discriminator D are updated according to maximizing the overall loss until convergence.
[0090] With the parameters of the discriminator D and the segmentation network fixed, the generator G is trained: the second sample frame is input into the segmentation network to obtain the segmentation mask; the second sample frame and the segmentation mask are input into the generator G to obtain the prediction frame; the prediction frame / annotated image frame and the segmentation mask are input into the discriminator D to obtain the discrimination result; the overall loss is calculated based on the discrimination result, and the parameters of the generator G are updated according to minimizing the overall loss until convergence.
[0091] The training of the discriminator D and the generator G is performed alternately until the preset conditions are met.
[0092] In one specific implementation, combined with Figure 2 As shown, the step of training the segmentation network based on the first sample frame and the real mask to obtain the trained segmentation network includes:
[0093] The first sample frame is semantically segmented to obtain a semantic mask;
[0094] Feature extraction is performed on the semantic mask to obtain the feature pyramid;
[0095] Input the feature pyramid into the displacement field module to obtain the displacement field;
[0096] Input the feature pyramid into the class boundary module to obtain the boundary map;
[0097] The displacement field and boundary map are input into the perception propagation module to obtain the prediction mask;
[0098] The overall loss is calculated based on the displacement field, the predicted mask, the real mask, and the boundary map.
[0099] The parameters of the segmentation network are iteratively divided based on the overall loss until the loss converges.
[0100] In this application, a segmentation network is used to extract a set of feature maps from the semantic mask as input, and a feature pyramid is constructed.
[0101] In this application, the real mask can be a low-resolution mask, and the segmentation network can be trained based on low-supervision and unsupervised training to obtain the trained segmentation network.
[0102] In this application, combined with Figure 3As shown, the displacement field module then applies 1×1 convolutional layers to each feature map layer for dimensionality reduction to 256, improving the efficiency of feature representation. The displacement field module integrates feature maps from different levels in a top-down manner to improve the prediction accuracy of the displacement field: the top two feature maps are upsampled, then concatenated, and processed through a 1×1 convolution. Finally, a second upsampling operation is performed to align with the concatenation of the bottom two feature maps for a second concatenation, obtaining multi-scale feature maps. Finally, a series of 1×1 convolutions are added to predict the displacement field. In the displacement field, the 2-D displacement vector learned by each pixel points to its corresponding instance centroid.
[0103] In this application, by fusing feature information at different levels, the gap between semantic information and structural details can be effectively bridged, enabling DFM to acquire rich instance-related features.
[0104] In this application, combined with Figure 4 As shown, the class boundary module generates clear class boundaries as supplementary information by calculating the semantic similarity between adjacent pixels. Specifically: First, a 1×1 convolution is used for dimensionality reduction. This is followed by bilinear upsampling, which enlarges the feature map size by 2 or 4 times as needed for subsequent feature fusion. Next, feature maps from different levels are concatenated and subjected to a single 1×1 convolution to obtain a boundary map that incorporates multi-scale information.
[0105] Preferably, in this application, when the data of the real mask is 0 (that is, the real mask is not labeled), the overall loss is calculated based on the displacement field and boundary map, the segmentation network is trained, and the trained segmentation network is obtained by utilizing the characteristic that the segmentation network can be trained in an unsupervised manner.
[0106] In one specific implementation, the step of inputting the displacement field and boundary map into the sensing propagation module to obtain the prediction mask includes:
[0107] The displacement field is masked to obtain the mask;
[0108] After multiplying the mask with the semantic mask, the result is fed into the perception propagation module along with the boundary graph to obtain the prediction mask.
[0109] In this application, the displacement field is a two-dimensional or three-dimensional vector field. Each pixel learns a 2-D displacement vector pointing to the centroid of its corresponding instance.
[0110] In this application, the corresponding loss function can be calculated and the network model can be trained using this displacement field in an unsupervised manner.
[0111] In this application, the displacement field is thresholded or otherwise processed to extract significant motion regions as masks.
[0112] In this application, a prediction mask is generated by combining a mask, a semantic mask, and a boundary graph through a perception propagation module.
[0113] In this application, the multiplication of the mask and the semantic mask combines the motion information in the displacement field with the semantic information, highlighting the semantically important motion regions.
[0114] In this application, the boundary map serves to provide contour information of the object, helping the perception propagation module to better understand the shape and structure of the object.
[0115] In this application, the perception propagation module is a neural network component used to propagate and fuse information in the feature space. It can employ U-Net or other codec architectures.
[0116] In this application, high-quality predictive masks can be generated through the feature propagation capabilities of the multi-information fusion and perception propagation modules.
[0117] In one specific implementation, the calculation of the overall loss based on the displacement field, the predicted mask, the real mask, and the boundary map includes:
[0118] Calculate displacement loss based on displacement field;
[0119] Calculate the prediction loss based on the predicted mask and the real mask;
[0120] Calculate the similarity loss based on the boundary map and the real mask;
[0121] The overall loss is calculated based on displacement loss, prediction loss, and similarity loss.
[0122] In this application, the prediction loss can be either cross-entropy loss or Dice loss.
[0123] In one implementation, before performing semantic segmentation on the image frame to obtain a semantic mask, the method further includes updating the image frame; the specific process of the update includes:
[0124] The image frame is divided into blocks to obtain independent blocks;
[0125] For each independent block, obtain the first and second neighboring blocks with different spacings;
[0126] A first feature block is generated based on the independent block and the first neighboring block;
[0127] A second feature block is generated based on the independent block and the second neighboring block.
[0128] The first and second feature blocks are compressed to obtain a compressed block.
[0129] Iterate through all the individual blocks and generate updated image frames based on the resulting compressed blocks.
[0130] In this application, the image frame is divided into blocks, that is, the image frame is divided into corresponding image blocks by using a checkerboard pattern; wherein, the image block can be at the pixel level (that is, each pixel is an image block) or other levels, and the specific division depends on the actual processing situation.
[0131] In this application, a sliding window or a fixed step size is used to divide the image into blocks of the same size.
[0132] It should be noted that if the image frame is a 3D image, then a face is selected and divided into a checkerboard pattern. Each square is a strip with a depth (the depth of the 3D image), and this strip is an image block. If the image frame is a 2D image, then the 2D image is divided into a checkerboard pattern, and each square is an image block.
[0133] Preferably, in this application, each image block has 1,001,000 pixels, thereby enabling more feature calculations between local regions while ensuring generation accuracy and reducing computational load.
[0134] In this application, an image block is selected as an independent block. The image blocks above, below, to the left, and to the right of this independent block are the first neighboring blocks; the image blocks one grid away from the top, bottom, left, and right of this independent block are the second neighboring blocks. The spacing between the first and second neighboring blocks and the independent block is different.
[0135] In this application, neighborhood information is extracted for each independent block to capture local structure.
[0136] In this application, generating the first feature block is to generate a local feature representation using an independent block and its first neighboring block. Specifically, this can be done by processing the independent block and the first neighboring block with convolutional layers and attention layers to obtain the first feature block.
[0137] In this application, the specific structure and parameters of the convolutional layer and attention layer can be obtained from the training data or determined according to the actual situation.
[0138] It should be noted that in this application, there are four first neighboring blocks and multiple first feature blocks.
[0139] In this application, the independent block and the first neighboring block are processed by convolutional layers and attention layers to obtain the first feature block. The specific process is as follows: the independent block and four neighboring blocks are concatenated together to form a multi-channel input, and the convolutional layer is used to extract features from the concatenated block; an important feature is enhanced by using a self-attention mechanism or a channel attention mechanism, the attention weight is calculated, and the output of the convolutional layer is weighted to enhance the important feature; the output of the attention layer is split into multiple feature blocks, and each feature block corresponds to the processing result of the independent block and at least one neighboring block.
[0140] In this application, a second feature block is generated to generate a broader local feature representation using the independent block and its second neighboring block. The specific generation process is the same as that of the first feature block, except that the parameters of the convolutional layer and the attention layer are different.
[0141] In this application, the generated feature blocks are compressed into a more compact representation to reduce computational cost while retaining key information. Feature compression is performed using pooling operations (such as max pooling or average pooling) or fully connected layers.
[0142] In this way, multiple first feature blocks and second feature blocks are compressed into a single compressed block, which corresponds to the size and position of the independent blocks and is used to replace them. All image blocks are replaced by the compressed blocks, resulting in an updated image frame.
[0143] In this application, each image block of the image frame is traversed to obtain the corresponding compressed block.
[0144] In this application, for image blocks / independent blocks near the edge, their first and second neighboring blocks are incomplete. In this case, the incomplete blocks are completed by copying the first and second neighboring blocks in their relative positions. For example, if the first neighboring block above an independent block does not exist, the first neighboring block below it is copied and used as the block above it.
[0145] In this application, by completing the image blocks, the processing accuracy of adjacent image blocks is greatly improved.
[0146] In this application, an adaptive adjustment module is used to capture the similarity relationship between local regions, thereby enhancing the feature representation.
[0147] This application provides a surgical robot remote control device for executing the surgical robot remote control method described above. The surgical robot remote control device will be described in detail below.
[0148] like Figure 5 As shown, the remote control device for the surgical robot includes:
[0149] Channel construction module 101, which is used to construct a remote information transmission channel;
[0150] Information transmission module 102 is used to acquire intraoperative three-dimensional image information based on the transmission channel;
[0151] The control generation module 103 is used to generate robotic arm control information based on preoperative planning information and intraoperative three-dimensional image information;
[0152] The robotic arm control module 104 is used to transmit and control the robotic arm to execute the robotic arm control information.
[0153] In one embodiment, the information transmission module 102 is further configured to:
[0154] Acquire an image and video stream of intraoperative three-dimensional imaging information; identify key frames in the image and video stream; transmit the key frames based on the transmission channel; and generate a video stream based on the key frames.
[0155] In one embodiment, the information transmission module 102 is further configured to:
[0156] Obtain the keyframe threshold; split the image video stream to obtain sequentially set image frames; calculate the histogram difference between adjacent image frames; if the histogram difference is greater than the keyframe threshold, the current image frame is regarded as a keyframe.
[0157] In one embodiment, the information transmission module 102 is further configured to:
[0158] The network latency of the transmission channel is monitored in real time; if the network latency exceeds a preset threshold, an image prediction model is obtained; based on the image prediction model, a prediction frame for the next moment corresponding to the current key frame is generated; and based on the prediction frame, the video stream is generated.
[0159] In one embodiment, the information transmission module 102 is further configured to:
[0160] The current keyframe is input into the segmentation network to obtain the segmentation mask; the current keyframe and the segmentation mask are input into the prediction network to obtain the prediction frame for the next time step.
[0161] In one embodiment, the information transmission module 102 is further configured to:
[0162] A first sample frame is acquired, on which a real mask is annotated; a segmentation network is trained based on the first sample frame and the real mask to obtain a trained segmentation network; a second sample frame is acquired, on which an image frame for the next time step is annotated; keeping the parameters of the segmentation network unchanged, an image prediction model is trained based on the second sample frame to obtain a trained image prediction model.
[0163] In one embodiment, the information transmission module 102 is further configured to:
[0164] The first sample frame is semantically segmented to obtain a semantic mask; features are extracted from the semantic mask to obtain a feature pyramid; the feature pyramid is input into the displacement field module to obtain a displacement field; the feature pyramid is input into the class boundary module to obtain a boundary map; the displacement field and the boundary map are input into the perceptual propagation module to obtain a prediction mask; the overall loss is calculated based on the displacement field, prediction mask, ground truth mask, and boundary map; the parameters of the segmentation network are iteratively calculated based on the overall loss until the loss converges.
[0165] The surgical robot remote control device provided in the above embodiments of this application corresponds to the surgical robot remote control method provided in the embodiments of this application. Therefore, the specific content in the system corresponds to the surgical robot remote control method. The specific content can be referred to the records in the surgical robot remote control method, which will not be repeated in this application.
[0166] The surgical robot remote control device provided in the above embodiments of this application and the surgical robot remote control method provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application stored therein.
[0167] The above describes the internal functions and structure of the surgical robot remote control device, such as... Figure 6 As shown, in practice, the remote control device for the surgical robot can be implemented as an electronic device, including: a memory 301 and a processor 303.
[0168] Memory 301 can be configured to store a program.
[0169] Additionally, memory 301 can also be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.
[0170] Memory 301 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Processor 303, coupled to memory 301, is used to execute programs in memory 301 for:
[0171] Establish a remote information transmission channel;
[0172] Intraoperative three-dimensional image information is acquired based on the transmission channel;
[0173] Based on preoperative planning information and intraoperative 3D imaging information, robotic arm control information is generated.
[0174] Transmit and control the robotic arm to execute the robotic arm control information.
[0175] In one implementation, the processor 303 is further configured to:
[0176] Acquire an image and video stream of intraoperative three-dimensional imaging information; identify key frames in the image and video stream; transmit the key frames based on the transmission channel; and generate a video stream based on the key frames.
[0177] In one implementation, the processor 303 is further configured to:
[0178] Obtain the keyframe threshold; split the image video stream to obtain sequentially set image frames; calculate the histogram difference between adjacent image frames; if the histogram difference is greater than the keyframe threshold, the current image frame is regarded as a keyframe.
[0179] In one implementation, the processor 303 is further configured to:
[0180] The network latency of the transmission channel is monitored in real time; if the network latency exceeds a preset threshold, an image prediction model is obtained; based on the image prediction model, a prediction frame for the next moment corresponding to the current key frame is generated; and based on the prediction frame, the video stream is generated.
[0181] In one implementation, the processor 303 is further configured to:
[0182] The current keyframe is input into the segmentation network to obtain the segmentation mask; the current keyframe and the segmentation mask are input into the prediction network to obtain the prediction frame for the next time step.
[0183] In one implementation, the processor 303 is further configured to:
[0184] A first sample frame is acquired, on which a real mask is annotated; a segmentation network is trained based on the first sample frame and the real mask to obtain a trained segmentation network; a second sample frame is acquired, on which an image frame for the next time step is annotated; keeping the parameters of the segmentation network unchanged, an image prediction model is trained based on the second sample frame to obtain a trained image prediction model.
[0185] In one implementation, the processor 303 is further configured to:
[0186] The first sample frame is semantically segmented to obtain a semantic mask; features are extracted from the semantic mask to obtain a feature pyramid; the feature pyramid is input into the displacement field module to obtain a displacement field; the feature pyramid is input into the class boundary module to obtain a boundary map; the displacement field and the boundary map are input into the perceptual propagation module to obtain a prediction mask; the overall loss is calculated based on the displacement field, prediction mask, ground truth mask, and boundary map; the parameters of the segmentation network are iteratively calculated based on the overall loss until the loss converges.
[0187] In this application, Figure 6 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 6 The components shown.
[0188] The electronic device provided in this embodiment is based on the same inventive concept as the remote control method for surgical robots provided in this application embodiment, and has the same beneficial effects as the methods adopted, run or implemented by the application stored therein.
[0189] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0190] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 a process or multiple processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0191] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0192] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0193] This application also provides a computer-readable storage medium corresponding to the surgical robot remote control method provided in the foregoing embodiments, wherein a computer program (i.e., a program product) is stored thereon. When the computer program is run by a processor, it executes the interactive image analysis assistance method for 3D aerial imaging provided in any of the foregoing embodiments.
[0194] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0195] The computer-readable storage medium provided in the above embodiments of this application and the interactive image analysis assistance method for 3D aerial imaging provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.
[0196] It should be noted that numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0197] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0198] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for remote control of a surgical robot, characterized in that, include: Establish a remote information transmission channel; Intraoperative three-dimensional image information is acquired based on the transmission channel; Based on preoperative planning information and intraoperative 3D imaging information, robotic arm control information is generated. Transmit and control the robotic arm to execute the robotic arm control information.
2. The remote control method for a surgical robot according to claim 1, characterized in that, The acquisition of intraoperative three-dimensional image information based on the transmission channel includes: Image and video streams for acquiring intraoperative 3D imaging information; Identify keyframes in image and video streams; The key frame is transmitted based on the transmission channel; Based on the keyframes, a video stream is generated.
3. The remote control method for a surgical robot according to claim 2, characterized in that, The identification of keyframes in the image video stream includes: Obtain the keyframe threshold; The image / video stream is split into sequentially set image frames; Calculate the histogram difference between adjacent image frames; If the histogram difference is greater than the keyframe threshold, the current image frame is considered a keyframe.
4. The remote control method for a surgical robot according to claim 2 or 3, characterized in that, After transmitting the key frame based on the transmission channel, the process further includes: Real-time monitoring of network latency in the transmission channel; When network latency exceeds a preset threshold, an image prediction model is obtained; Based on the image prediction model, generate the prediction frame for the next moment corresponding to the current keyframe; The video stream is generated based on the predicted frames.
5. The remote control method for a surgical robot according to claim 4, characterized in that, The step of generating the predicted frame for the next time step corresponding to the current keyframe based on the image prediction model includes: Input the current keyframe into the segmentation network to obtain the segmentation mask; The current keyframe and the segmentation mask are input into the prediction network to obtain the prediction frame for the next time step.
6. The remote control method for a surgical robot according to claim 4, characterized in that, The training process of the image prediction model includes: Obtain a first sample frame, on which a real mask is marked; The segmentation network is trained based on the first sample frame and the real mask to obtain the trained segmentation network. Acquire a second sample frame, on which the image frame for the next moment is marked; Keeping the parameters of the segmentation network unchanged, the image prediction model is trained based on the second sample frame to obtain the trained image prediction model.
7. The remote control method for a surgical robot according to claim 6, characterized in that, The step of training the segmentation network based on the first sample frame and the real mask to obtain the trained segmentation network includes: The first sample frame is semantically segmented to obtain a semantic mask; Feature extraction is performed on the semantic mask to obtain the feature pyramid; Input the feature pyramid into the displacement field module to obtain the displacement field; Input the feature pyramid into the class boundary module to obtain the boundary map; The displacement field and boundary map are input into the perception propagation module to obtain the prediction mask; The overall loss is calculated based on the displacement field, the predicted mask, the real mask, and the boundary map. The parameters of the segmentation network are iteratively divided based on the overall loss until the loss converges.
8. A remote control device for a surgical robot, characterized in that, include: The channel construction module is used to build remote information transmission channels; Information transmission module, which is used to acquire intraoperative three-dimensional image information based on the transmission channel; The control generation module is used to generate robotic arm control information based on preoperative planning information and intraoperative 3D image information; A robotic arm control module is used to transmit and control the robotic arm to perform the robotic arm control information.
9. An electronic device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor, coupled to the memory, is used to execute the program for: Establish a remote information transmission channel; Intraoperative three-dimensional image information is acquired based on the transmission channel; Based on preoperative planning information and intraoperative 3D imaging information, robotic arm control information is generated. Transmit and control the robotic arm to execute the robotic arm control information.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the remote control method for surgical robots according to any one of claims 1-7.