An AI-based intelligent split-screen remote robotic surgery communication control system and method
By using AI in remote robot surgical systems to predict and contact risk analysis of position changes from hand end effectors and non-surgical areas, the problem of delayed and insufficient prediction of motion in the prior art is solved, and the safety and accuracy of the surgery are improved.
Patent Information
- Application Number
- CN202510044815.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The existing remote robot surgical technology has delays in data transmission, which affects the doctor's real-time operation accuracy and response speed, and lacks predictions of tissue movements in non-surgical areas, increasing surgical risks.
Using the AI-based smart split-screen remote robot surgical communication control system, by obtaining the position data input by the main hand end, the real-time position data of the slave end effector and the tissue physiological movement data of the non-surgical area, a comprehensive prediction of the position changes of the hand end effector and the position shift changes of the non-surgical area are carried out, and the contact risk is analyzed, and the contact risk warning is sent to the main hand end.
Reduces the possibility of doctors accidentally contacting non-surgical areas due to communication delays, and improves the safety and accuracy of the surgery.
Smart Images

Figure CN119495413B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication control technology, and in particular to an AI-based intelligent split-screen remote robotic surgery communication control system and method. Background Art
[0002] With the rapid development of robotics technology, remote robotic surgery has gradually become one of the important surgical methods in the medical field. Remote robotic surgery enables doctors to perform precise surgical operations far away from the surgical site by transmitting the doctor's operating actions to the surgical robot in real time. This technology can not only provide high-quality medical services in remote areas or medical institutions with limited resources, but also protect the health and safety of doctors and patients when the patient's disease is contagious. However, the current remote robotic surgery technology still has many problems in practical applications. For example: the existing remote robotic surgery technology has a certain delay in the data transmission process, which may cause a time difference between the doctor's operating instructions and the actual action of the end effector of the slave hand. This delay will affect the doctor's real-time operation accuracy and reaction speed, and increase the risk of surgery.
[0003] During surgery, tissues in non-surgical areas (such as internal organs, blood vessels, etc.) may move due to physiological or external factors. Existing remote surgery systems usually only focus on the position changes of the slave end effector, without considering the prediction of tissue movement in non-surgical areas. Movement in non-surgical areas may cause accidental contact between the slave end effector and these areas, thereby increasing surgical risks. Existing technologies lack the ability to predict these movements, so doctors often need to rely on experience and intuitive judgment to avoid accidental touches. Existing remote robotic surgery technology also has a certain delay in the data transmission process, which may cause a time difference between the operation instructions issued by the doctor based on experience and intuitive judgment and the actual action of the slave end effector. This delay will further affect the doctor's real-time operation accuracy and reaction speed, increasing surgical risks. Summary of the invention
[0004] In order to overcome the defects and shortcomings of the prior art, the present invention provides an AI-based intelligent split-screen remote robotic surgery communication control system and method, which realizes a comprehensive prediction of the position changes of the slave hand end effector and the position offset changes of the non-surgical area by fusing the input position data of the master hand, the real-time position data of the slave hand end effector and the physiological movement data of the tissue in the non-surgical area.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] In a first aspect, an embodiment of the present invention provides an AI-based intelligent split-screen remote robotic surgery communication control method, comprising the following steps:
[0007] S1. Obtain the input position data of the master hand end of the Zhilian split-screen remote robot and the real-time position data of the slave hand end effector; and simultaneously obtain the physiological movement data of tissues in the non-surgical area;
[0008] S2, importing the tissue physiological movement data of the non-surgical area into the non-surgical area position deviation prediction model to predict the position deviation change of the non-surgical area;
[0009] S3, importing the input position data of the master hand end and the real-time position data of the slave hand end effector into the slave hand position change prediction model, and predicting the position change of the slave hand end effector;
[0010] S4. analyzing the contact risk between the slave hand end effector and the non-surgical area according to the position offset change prediction result of the non-surgical area and the position change prediction result of the slave hand end effector;
[0011] S5. According to the contact risk analysis results, a contact risk warning is sent to the main hand end of the intelligent split-screen remote robot.
[0012] In one implementation of the present invention, step S2 includes the following specific steps:
[0013] S21. Acquire the tissue physiological motion data of the non-surgical area through the fixed camera of the intelligent split-screen remote robot; wherein the tissue physiological motion data of the non-surgical area is acquired by: acquiring all frames of real-time video images when the end effector of the hand operates in the surgical area, and performing image segmentation on the real-time video images according to the division of the surgical area, acquiring all frames of real-time video images of the non-surgical area after the image segmentation, and using all frames of real-time video images of the non-surgical area after the image segmentation as the tissue physiological motion data of the non-surgical area;
[0014] S22, using the first frame of the non-surgical area real-time video image as a reference image for all subsequent frames of the non-surgical area real-time video image, randomly selecting multiple feature points in the reference image as tracking feature points; calculating the image gradient in the x direction and the image gradient in the y direction of each tracking feature point in the reference image within a surrounding window size of w*w; and calculating the time gradient of each tracking feature point in the t-th frame of the non-surgical area real-time video image.
[0015] S23, in the reference image of each tracking feature point, the image gradients in the x direction and the image gradients in the y direction of all pixels within the window size n around each tracking feature point are formed into a matrix to obtain the image gradient matrix of the tracking feature point ; Wherein, Ixi and Iyi are the image gradient in the x direction and the image gradient in the y direction of the i-th pixel, respectively; i is any one from 1 to n;
[0016] S24, in the t-th frame of the real-time video image of the non-surgical area, the time gradients of all pixels within the range of a window size of n around each tracking feature point are formed into a matrix to obtain the time gradient matrix of the tracking feature point in the t-th frame of the real-time video image of the non-surgical area ; Wherein, the temporal gradient is the difference between the pixel value of all pixels at the position in the real-time video image of the non-surgical area of the t-th frame and the pixel value of the pixel at the same position in the real-time video image of the non-surgical area of the t+1-th frame; Wherein, Iti represents the temporal gradient of the i-th pixel in the real-time video image of the non-surgical area of the t-th frame with each tracking feature point as the center and the surrounding window size of n;
[0017] S25, substituting the image gradient matrix of the tracking feature points and the time gradient matrix of the tracking feature points in the t-th frame non-surgical area real-time video image into the displacement vector calculation formula to calculate the displacement vector of each tracking feature point from the reference image to the t-th frame non-surgical area real-time video image; the displacement vector calculation formula is: ;
[0018] Where Vt represents the displacement vector of each tracking feature point from the reference image to the real-time video image of the non-surgical area in the tth frame, represents the transpose of the image gradient matrix about the tracked feature points, It represents the inverse matrix of the matrix in brackets; Respectively represent the displacement vector of each tracking feature point from the reference image to the real-time video image of the non-surgical area in the tth frame in the x direction and the y direction;
[0019] S26, the displacement vector of each tracking feature point from the reference image to the real-time video image of the non-surgical area of the t-th frame in the x-direction and the y-direction and the position of each tracking feature point in the reference image Substitute into the position change calculation formula to calculate the position of each tracking feature point in the real-time video image of the non-surgical area in the tth frame; the position change calculation formula is:
[0020] ;
[0021] In the formula, They respectively represent the abscissa and ordinate of the position of each tracking feature point in the real-time video image of the non-surgical area in the tth frame.
[0022] In one implementation of the present invention, step S2 further includes the following specific contents:
[0023] S27, obtaining the calculated position of each tracking feature point in all non-surgical area real-time video images; constructing a position time series of each tracking feature point according to the time sequence of the position of each tracking feature point in all non-surgical area real-time video images; based on the position time series of each tracking feature point, training a non-surgical area position offset prediction model for predicting the position of each tracking feature point in multiple frames of non-surgical area real-time video images in the future;
[0024] S28, preset the sliding step length to be L and the sliding window length to be W; use the sliding window method to convert the position of each tracking feature point in the position time series of each tracking feature point in all non-surgical area real-time video images into multiple training samples, wherein each training sample is composed of the position time series of each tracking feature point in the sliding window, use the training samples as the input of the non-surgical area position offset prediction model, use the position of each tracking feature point in the non-surgical area real-time video image with a preset sliding step length of L as the output, and train the non-surgical area position offset prediction model with the prediction accuracy as the training target; generate a non-surgical area position offset prediction model that predicts the position of each tracking feature point in multiple frames of non-surgical area real-time video images in the future according to the position of each tracking feature point in all non-surgical area real-time video images; wherein the non-surgical area position offset prediction model is a recurrent neural network model;
[0025] S29, obtaining the predicted position of each tracking feature point in the future multi-frame real-time video image of the non-surgical area; converting the position of each tracking feature point in the future multi-frame real-time video image of the non-surgical area into the world coordinate system through the image coordinate system, thereby obtaining the actual position of each tracking feature point in the future multi-frame real-time video image of the non-surgical area.
[0026] In one implementation of the present invention, step S3 includes the following specific steps:
[0027] S31, obtaining the input position data of the master hand end and the real-time position data of the slave hand end effector; wherein, obtaining the input position data of the master hand end includes: extracting the displacement of the master hand input corresponding to the time point of each frame of real-time video image as the master hand input displacement of each frame of real-time video image; taking the master hand input displacement of all frames of real-time video images as the input position data of the master hand end; the real-time position data of the slave hand end effector includes: the real-time position of the slave hand end effector in all frames of real-time video images;
[0028] S32, constructing a master-slave hand position time series according to the time sequence of the master hand input position data and the real-time position data of the slave hand end effector, and training a slave hand position change prediction model for predicting the position of the slave hand end effector in multiple frames of real-time video images in the future based on the master-slave hand position time series;
[0029] S33, preset the sliding step size to be Lc and the sliding window length to be Wc; use the sliding window method to convert the input position data of the master hand end in the master-slave hand position time series and the real-time position data of the slave hand end effector into multiple training samples, wherein each training sample is composed of the master-slave hand position data sequence in the sliding window, use the training sample as the input of the slave hand position change prediction model, use the position of the slave hand end effector in the future multi-frame real-time video image with a preset sliding step size of Lc as the output, and train the slave hand position change prediction model with the prediction accuracy as the training target; generate a slave hand position change prediction model that predicts the position of the slave hand end effector in the future multi-frame real-time video image based on the input position data of the master hand end and the real-time position data of the slave hand end effector; wherein the slave hand position change prediction model is a recurrent neural network model;
[0030] S34. Obtain the predicted position of the slave hand end effector in the future multi-frame real-time video image; convert the position of the slave hand end effector in the future multi-frame real-time video image into the world coordinate system through the image coordinate system to obtain the actual position of the slave hand end effector in the future multi-frame real-time video image.
[0031] In one implementation of the present invention, step S4 includes the following specific steps:
[0032] S41, obtaining a set of actual position coordinates of the surgical area; extracting the actual position of the slave hand end effector in the future multiple frames of real-time video images, and the actual position of each tracking feature point in the future multiple frames of real-time video images of the non-surgical area corresponding to the future multiple frames of real-time video images; and substituting them into the distance formula to calculate the distance between the slave hand end effector in each future frame of real-time video images and all the tracking feature points in the corresponding future frame of real-time video images of the non-surgical area;
[0033] S42, determining whether the actual position of the slave hand end effector in the future multiple frames of real-time video images belongs to the actual position coordinate set of the surgical area; when the actual position of the slave hand end effector belongs to the actual position coordinate set of the surgical area, substituting the distance between the slave hand end effector in each future frame of real-time video image and all the tracking feature points in the corresponding future frame of real-time video image of the non-surgical area into the contact risk coefficient calculation formula, and calculating the contact risk coefficient of the slave hand end effector and the non-surgical area in each future frame of real-time video image; the contact risk coefficient calculation formula is:
[0034] ;
[0035] Where Fr represents the contact risk coefficient between the slave hand end effector and the non-surgical area in the future r-th frame of real-time video image; min() represents the minimum value of the distance in the brackets; dq represents the distance between the slave hand end effector in the future r-th frame of real-time video image and the q-th tracking feature point in the corresponding future r-th frame of real-time video image of the non-surgical area.
[0036] In one implementation of the present invention, step S5 includes the following specific contents:
[0037] S51, when the actual position of the slave hand end effector does not belong to the actual position coordinate set of the surgical area, immediately send a contact risk warning to the master hand end of the Zhilian split-screen remote robot;
[0038] S52. When the actual position of the slave hand end effector belongs to the actual position coordinate set of the surgical area, a contact risk threshold is preset. When the contact risk coefficient between the slave hand end effector and the non-surgical area in the future rth frame of real-time video image is greater than the contact risk threshold, a contact risk warning is immediately issued to the master hand end of the intelligent split-screen remote robot.
[0039] In a second aspect, an embodiment of the present invention further provides an AI-based intelligent split-screen remote robotic surgery communication control system, comprising:
[0040] The data acquisition module is used to obtain the input position data of the master hand end of the Zhilian split-screen remote robot and the real-time position data of the slave hand end effector; at the same time, it obtains the physiological movement data of the tissue in the non-surgical area;
[0041] A non-surgical area position deviation prediction module is used to import the tissue physiological movement data of the non-surgical area into the non-surgical area position deviation prediction model to predict the position deviation change of the non-surgical area;
[0042] The slave hand position change prediction module is used to import the input position data of the master hand end and the real-time position data of the slave hand end effector into the slave hand position change prediction model to predict the position change of the slave hand end effector;
[0043] A contact risk analysis module, used for analyzing the contact risk between the slave hand end effector and the non-surgical area according to the position offset change prediction result of the non-surgical area and the position change prediction result of the slave hand end effector;
[0044] The contact risk warning module is used to send a contact risk warning to the main hand end of the Zhilian split-screen remote robot based on the contact risk analysis results;
[0045] A control module is used to control the operation of the data acquisition module, the non-surgical area position offset prediction module, the hand position change prediction module, the contact risk analysis module and the contact risk warning module.
[0046] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a processor and a memory, wherein the memory stores a computer program that can be called by the processor, and the processor executes an AI-based intelligent split-screen remote robotic surgery communication control method by calling the computer program stored in the memory.
[0047] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0048] The present invention imports the physiological movement data of tissues in the non-surgical area into the non-surgical area position offset prediction model to predict the position offset change of the non-surgical area; imports the input position data of the master hand end and the real-time position data of the slave hand end effector into the slave hand position change prediction model to predict the position change of the slave hand end effector; analyzes the contact risk between the slave hand end effector and the non-surgical area according to the prediction results of the position offset change of the non-surgical area and the prediction results of the position change of the slave hand end effector; sends a contact risk warning to the master hand end of the Zhilian split-screen remote robot according to the contact risk analysis results. It can reduce the possibility of doctors accidentally contacting the non-surgical area due to communication delays in remote surgery, thereby improving the safety of surgery. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0050] Figure 1 It is a schematic diagram of the overall process of an AI-based intelligent split-screen remote robotic surgery communication control method of the present invention;
[0051] Figure 2 This is a workflow diagram of step S5 in an AI-based intelligent split-screen remote robotic surgery communication control method of the present invention;
[0052] Figure 3 This is a structural schematic diagram of an AI-based intelligent split-screen remote robotic surgery communication control system of the present invention. DETAILED DESCRIPTION
[0053] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. The embodiments of the present invention and the technical features in the embodiments may be combined with each other unless there is a conflict.
[0054] Example 1
[0055] like Figure 1 As shown, this embodiment provides an AI-based intelligent split-screen remote robot surgery communication control method, which specifically includes the following steps:
[0056] S1. Obtain the input position data of the master hand end of the Zhilian split-screen remote robot and the real-time position data of the slave hand end effector; and simultaneously obtain the physiological movement data of tissues in the non-surgical area;
[0057] S2, importing the tissue physiological movement data of the non-surgical area into the non-surgical area position deviation prediction model to predict the position deviation change of the non-surgical area;
[0058] S3, importing the input position data of the master hand end and the real-time position data of the slave hand end effector into the slave hand position change prediction model, and predicting the position change of the slave hand end effector;
[0059] S4. analyzing the contact risk between the slave hand end effector and the non-surgical area according to the position offset change prediction result of the non-surgical area and the position change prediction result of the slave hand end effector;
[0060] S5. According to the contact risk analysis results, a contact risk warning is sent to the main hand end of the intelligent split-screen remote robot.
[0061] In this embodiment, step S2 includes the following specific steps:
[0062] S21. Obtain the tissue physiological motion data of the non-surgical area through the fixed camera of the intelligent split-screen remote robot; wherein, the method for obtaining the tissue physiological motion data of the non-surgical area is: obtaining all frames of real-time video images when the end effector of the hand operates in the surgical area, wherein the real-time video image includes a large amount of surgical environment information, specifically including the surgical area and the surrounding non-surgical area; in this embodiment, the real-time video image is segmented according to the division of the surgical area, and all frames of real-time video images of the non-surgical area after image segmentation are obtained, and all frames of real-time video images of the non-surgical area after image segmentation are used as the tissue physiological motion data of the non-surgical area; technicians in this field can separate the surgical area from the non-surgical area through image segmentation methods such as edge detection, region growth, and segmentation network (such as U-Net). The above image segmentation methods are all prior art, so they are not repeated here. This step ensures that the present application can focus on the tissue physiological motion data of the non-surgical area and eliminates the interference of the surgical area on the non-surgical area.
[0063] S22, using the first frame of the real-time video image of the non-surgical area as a reference image for all subsequent frames of the real-time video image of the non-surgical area, wherein the reference image is used to provide a benchmark for comparing the position offset changes of the non-surgical area in subsequent frames; randomly selecting multiple feature points in the reference image as tracking feature points; calculating the image gradient in the x direction and the image gradient in the y direction of each tracking feature point in the reference image within a surrounding window size of w*w; and calculating the time gradient of each tracking feature point in the t-th frame of the real-time video image of the non-surgical area.
[0064] S23, in the reference image of each tracking feature point, the image gradients in the x direction and the image gradients in the y direction of all pixels within the window size n around each tracking feature point are formed into a matrix to obtain the image gradient matrix of the tracking feature point ; Wherein, Ixi and Iyi are the image gradient in the x direction and the image gradient in the y direction of the i-th pixel respectively; i is any item from 1 to n; wherein, the image gradient can be calculated by existing technologies such as Sobel operator and Canny edge detection.
[0065] S24, in the t-th frame of the real-time video image of the non-surgical area, the time gradients of all pixels within the range of a window size of n around each tracking feature point are formed into a matrix to obtain the time gradient matrix of the tracking feature point in the t-th frame of the real-time video image of the non-surgical area ; Wherein, the temporal gradient is the difference between the pixel value of all pixels at the position in the real-time video image of the non-surgical area of the t-th frame and the pixel value of the pixel at the same position in the real-time video image of the non-surgical area of the t+1-th frame; Wherein, Iti represents the temporal gradient of the i-th pixel in the real-time video image of the non-surgical area of the t-th frame with each tracking feature point as the center and the surrounding window size of n;
[0066] S25, substituting the image gradient matrix of the tracking feature points and the time gradient matrix of the tracking feature points in the t-th frame non-surgical area real-time video image into the displacement vector calculation formula to calculate the displacement vector of each tracking feature point from the reference image to the t-th frame non-surgical area real-time video image; the displacement vector calculation formula is: ;
[0067] Where Vt represents the displacement vector of each tracking feature point from the reference image to the real-time video image of the non-surgical area in the tth frame, represents the transpose of the image gradient matrix about the tracked feature points, It represents the inverse matrix of the matrix in brackets; Respectively represent the displacement vector of each tracking feature point from the reference image to the real-time video image of the non-surgical area in the tth frame in the x direction and the y direction;
[0068] S26, the displacement vector of each tracking feature point from the reference image to the real-time video image of the non-surgical area of the t-th frame in the x-direction and the y-direction and the position of each tracking feature point in the reference image Substitute into the position change calculation formula to calculate the position of each tracking feature point in the real-time video image of the non-surgical area in the tth frame; the position change calculation formula is:
[0069] ;
[0070] In the formula, They respectively represent the abscissa and ordinate of the position of each tracking feature point in the real-time video image of the non-surgical area in the tth frame.
[0071] In this embodiment, step S2 also includes the following specific contents:
[0072] S27, obtaining the calculated position of each tracking feature point in all non-surgical area real-time video images; constructing a position time series of each tracking feature point according to the time sequence of the position of each tracking feature point in all non-surgical area real-time video images; based on the position time series of each tracking feature point, training a non-surgical area position offset prediction model for predicting the position of each tracking feature point in multiple frames of non-surgical area real-time video images in the future;
[0073] S28, preset the sliding step length to be L and the sliding window length to be W; use the sliding window method to convert the position of each tracking feature point in the position time series of each tracking feature point in all non-surgical area real-time video images into multiple training samples, wherein each training sample is composed of the position time series of each tracking feature point in the sliding window, use the training samples as the input of the non-surgical area position offset prediction model, use the position of each tracking feature point in the non-surgical area real-time video image with a preset sliding step length of L as the output, and train the non-surgical area position offset prediction model with the prediction accuracy as the training target; generate a non-surgical area position offset prediction model that predicts the position of each tracking feature point in multiple frames of non-surgical area real-time video images in the future according to the position of each tracking feature point in all non-surgical area real-time video images; wherein the non-surgical area position offset prediction model is a recurrent neural network model;
[0074] In this embodiment, the use of the sliding window method makes the training data more diverse, improves the generalization ability of the model, and can better cope with different types of physiological movements. By using the position time series within the sliding window as the input of the model, the position of each tracked feature point in the real-time video image of the non-surgical area with a preset sliding step size L is used as the output. This enables the model to learn the mapping relationship from the current multi-frame data to the future multi-frame data and realize multi-step prediction. Not only the current position information is considered, but also the position changes at multiple time points in the future, which better reflects the dynamic characteristics of tissue movement.
[0075] Combining the sliding window method with the recurrent neural network model can dynamically update and adjust the prediction results in the real-time data stream. This feature enables the system to adapt to the ever-changing tissue movement in the surgical environment and provide more timely and accurate early warning information.
[0076] S29, obtaining the predicted position of each tracking feature point in the future multi-frame real-time video image of the non-surgical area; converting the position of each tracking feature point in the future multi-frame real-time video image of the non-surgical area into the world coordinate system through the image coordinate system, to obtain the actual position of each tracking feature point in the future multi-frame real-time video image of the non-surgical area. In this embodiment, the method of converting the image coordinate system into the world coordinate system specifically includes:
[0077] S291, assuming that the predicted position coordinates of each tracking feature point in a future frame of real-time video image are (x, y);
[0078] S292, establish and solve equations ; Convert the point in the image coordinate system to a point (u, v) in the normalized coordinate system, where K is the intrinsic parameter matrix of the camera, which is composed of the lens focal length, principal point offset, and distortion parameters; technicians in this field can obtain it from the factory instructions of the camera;
[0079] S293, establish and solve equations ; The actual position (X, Y, Z) of each tracking feature point in a future frame of real-time video image in the world coordinate system is obtained, where R is the rotation matrix in the camera extrinsic parameters, and T is the translation vector in the camera extrinsic parameters; In this embodiment, those skilled in the art can calibrate the fixed camera of the Zhilian split-screen remote robot by using computer vision libraries such as OpenCV to obtain the camera intrinsic parameter K and extrinsic parameter [R|T]. In this embodiment, the accuracy of the camera intrinsic and extrinsic parameters is ensured by camera calibration, the geometric transformation realizes the conversion of the position of each tracking feature point from the image coordinate system to the normalized coordinate system, the depth information restores the three-dimensional coordinates, and finally the actual position of each tracking feature point in the world coordinate system is obtained by the inverse transformation of the extrinsic parameters; The safety and accuracy of the operation are improved.
[0080] In this embodiment, step S3 includes the following specific steps:
[0081] S31, obtaining the input position data of the master hand end and the real-time position data of the slave hand end effector; wherein, obtaining the input position data of the master hand end includes: extracting the displacement of the master hand input corresponding to the time point of each frame of real-time video image as the master hand input displacement of each frame of real-time video image; taking the master hand input displacement of all frames of real-time video images as the input position data of the master hand end; the real-time position data of the slave hand end effector includes: the real-time position of the slave hand end effector in all frames of real-time video images;
[0082] S32, constructing a master-slave hand position time series according to the time sequence of the master hand input position data and the real-time position data of the slave hand end effector, and training a slave hand position change prediction model for predicting the position of the slave hand end effector in multiple frames of real-time video images in the future based on the master-slave hand position time series;
[0083] S33, preset the sliding step size to be Lc and the sliding window length to be Wc; use the sliding window method to convert the input position data of the master hand end in the master-slave hand position time series and the real-time position data of the slave hand end effector into multiple training samples, wherein each training sample is composed of the master-slave hand position data sequence in the sliding window, use the training sample as the input of the slave hand position change prediction model, use the position of the slave hand end effector in the future multi-frame real-time video image with a preset sliding step size of Lc as the output, and train the slave hand position change prediction model with the prediction accuracy as the training target; generate a slave hand position change prediction model that predicts the position of the slave hand end effector in the future multi-frame real-time video image based on the input position data of the master hand end and the real-time position data of the slave hand end effector; wherein the slave hand position change prediction model is a recurrent neural network model;
[0084] S34, obtaining the predicted position of the slave hand end effector in the future multi-frame real-time video image; converting the position of the slave hand end effector in the future multi-frame real-time video image into the world coordinate system by using the image coordinate system to obtain the actual position of the slave hand end effector in the future multi-frame real-time video image;
[0085] In this embodiment, step S4 includes the following specific steps:
[0086] S41, obtaining a set of actual position coordinates of the surgical area; extracting the actual position of the slave hand end effector in the future multiple frames of real-time video images, and the actual position of each tracking feature point in the future multiple frames of real-time video images of the non-surgical area corresponding to the future multiple frames of real-time video images; and substituting them into the distance formula to calculate the distance between the slave hand end effector in each future frame of real-time video images and all the tracking feature points in the corresponding future frame of real-time video images of the non-surgical area;
[0087] S42, determining whether the actual position of the slave hand end effector in the future multiple frames of real-time video images belongs to the actual position coordinate set of the surgical area; when the actual position of the slave hand end effector belongs to the actual position coordinate set of the surgical area, substituting the distance between the slave hand end effector in each future frame of real-time video image and all the tracking feature points in the corresponding future frame of real-time video image of the non-surgical area into the contact risk coefficient calculation formula, and calculating the contact risk coefficient of the slave hand end effector and the non-surgical area in each future frame of real-time video image; the contact risk coefficient calculation formula is:
[0088] ;
[0089] Where Fr represents the contact risk coefficient between the slave hand end effector and the non-surgical area in the future r-th frame of real-time video image; min() represents the minimum value of the distance in the brackets; dq represents the distance between the slave hand end effector in the future r-th frame of real-time video image and the q-th tracking feature point in the corresponding future r-th frame of real-time video image of the non-surgical area.
[0090] In this embodiment, if Figure 2 As shown, step S5 includes the following specific contents:
[0091] S51, when the actual position of the slave hand end effector does not belong to the actual position coordinate set of the surgical area, immediately send a contact risk warning to the master hand end of the Zhilian split-screen remote robot;
[0092] S52. When the actual position of the slave hand end effector belongs to the actual position coordinate set of the surgical area, a contact risk threshold is preset. When the contact risk coefficient between the slave hand end effector and the non-surgical area in the future rth frame of real-time video image is greater than the contact risk threshold, a contact risk warning is immediately issued to the master hand end of the intelligent split-screen remote robot; wherein, the value of the contact risk threshold is determined as follows: obtaining the actual position of the slave hand end effector and the actual position of the tracking feature point in 5000 sets of real-time video images; substituting the actual position of the slave hand end effector and the actual position of the tracking feature point into the distance formula to calculate the shortest distance between the slave hand end effector and the tracking feature point; obtaining the contact risk judgment result of the slave hand end effector and the non-surgical area; importing the shortest distance between the slave hand end effector and the tracking feature point and the contact risk judgment result of the slave hand end effector and the non-surgical area into the fitting software, and outputting the value of the corresponding contact risk threshold that meets the highest contact risk judgment accuracy.
[0093] Example 2
[0094] like Figure 3 As shown, this embodiment provides an AI-based intelligent split-screen remote robotic surgery communication control system, including:
[0095] The data acquisition module is used to obtain the input position data of the master hand end of the Zhilian split-screen remote robot and the real-time position data of the slave hand end effector; at the same time, it obtains the physiological movement data of the tissue in the non-surgical area;
[0096] A non-surgical area position deviation prediction module is used to import the tissue physiological movement data of the non-surgical area into the non-surgical area position deviation prediction model to predict the position deviation change of the non-surgical area;
[0097] The slave hand position change prediction module is used to import the input position data of the master hand end and the real-time position data of the slave hand end effector into the slave hand position change prediction model to predict the position change of the slave hand end effector;
[0098] A contact risk analysis module, used for analyzing the contact risk between the slave hand end effector and the non-surgical area according to the position offset change prediction result of the non-surgical area and the position change prediction result of the slave hand end effector;
[0099] The contact risk warning module is used to send a contact risk warning to the main hand end of the Zhilian split-screen remote robot based on the contact risk analysis results;
[0100] The control module is used to control the operation of the data acquisition module, the non-surgical area position deviation prediction module, the hand position change prediction module, the contact risk analysis module and the contact risk warning module.
[0101] The above-mentioned parameters and steps for each unit module to realize the corresponding functions in the AI-based intelligent split-screen remote robot surgery communication control system of the present invention can refer to the parameters and steps in the embodiment of the AI-based intelligent split-screen remote robot surgery communication control method above, which will not be repeated here.
[0102] Example 3
[0103] An electronic device according to an embodiment of the present invention comprises: a processor and a memory, wherein the memory stores a computer program that can be called by the processor, and the processor executes an AI-based intelligent split-screen remote robot surgery communication control method by calling the computer program stored in the memory. It should be noted that all computer programs of an AI-based intelligent split-screen remote robot surgery communication control method are implemented in C language, wherein the data acquisition module, the non-surgical area position offset prediction module, the hand position change prediction module, the contact risk analysis module, the contact risk warning module and the control module are all controlled by a remote server.
[0104] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0105] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0106] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical disks.
[0107] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. An AI-based intelligent split-screen remote robotic surgery communication control method, characterized in that: The steps include: S1. Obtain the input position data of the master hand end of the Zhilian split-screen remote robot and the real-time position data of the slave hand end effector; and simultaneously obtain the physiological movement data of tissues in the non-surgical area; S2, importing the tissue physiological movement data of the non-surgical area into the non-surgical area position deviation prediction model to predict the position deviation change of the non-surgical area; S3, importing the input position data of the master hand end and the real-time position data of the slave hand end effector into the slave hand position change prediction model, and predicting the position change of the slave hand end effector; S4. analyzing the contact risk between the slave hand end effector and the non-surgical area according to the position offset change prediction result of the non-surgical area and the position change prediction result of the slave hand end effector; S5. Send a contact risk warning to the main hand end of the Zhilian split-screen remote robot based on the contact risk analysis results; The step S2 specifically includes: obtaining the position of each tracking feature point in all non-surgical area real-time video images, and inputting the position into the non-surgical area position offset prediction model for prediction, so as to obtain the position of each tracking feature point in the future multiple frames of non-surgical area real-time video images, and converting the position of each tracking feature point in the future multiple frames of non-surgical area real-time video images into the world coordinate system by means of the image coordinate system, so as to obtain the actual position of each tracking feature point in the future multiple frames of non-surgical area real-time video images; The step S3 specifically includes: obtaining the input position data of the master hand end and the real-time position data of the slave hand end effector, inputting the input position data of the master hand end and the real-time position data of the slave hand end effector into the slave hand position change prediction model as a recurrent neural network model for prediction, and obtaining the position of the slave hand end effector in the future multi-frame real-time video image; converting the position of the slave hand end effector in the future multi-frame real-time video image into the world coordinate system by means of the image coordinate system, and obtaining the actual position of the slave hand end effector in the future multi-frame real-time video image; The step S4 comprises the following specific steps: S41, obtaining a set of actual position coordinates of the surgical area; extracting the actual position of the slave hand end effector in the future multiple frames of real-time video images, and the actual position of each tracking feature point in the future multiple frames of real-time video images of the non-surgical area corresponding to the future multiple frames of real-time video images; and substituting them into the distance formula to calculate the distance between the slave hand end effector in each future frame of real-time video images and all the tracking feature points in the corresponding future frame of real-time video images of the non-surgical area; S42, determining whether the actual position of the slave hand end effector in the future multiple frames of real-time video images belongs to the actual position coordinate set of the surgical area; when the actual position of the slave hand end effector belongs to the actual position coordinate set of the surgical area, substituting the distance between the slave hand end effector in each future frame of real-time video image and all the tracking feature points in the corresponding future frame of real-time video image of the non-surgical area into the contact risk coefficient calculation formula, and calculating the contact risk coefficient of the slave hand end effector and the non-surgical area in each future frame of real-time video image; the contact risk coefficient calculation formula is: ; Where Fr represents the contact risk coefficient between the slave hand end effector and the non-surgical area in the future r-th frame of real-time video image; min() represents the minimum value of the distance in the brackets; dq represents the distance between the slave hand end effector in the future r-th frame of real-time video image and the q-th tracking feature point in the corresponding future r-th frame of real-time video image of the non-surgical area.
2. According to claim 1, the AI-based intelligent split-screen remote robotic surgery communication control method is characterized in that: The step S2 comprises the following specific steps: S21. Acquire the tissue physiological motion data of the non-surgical area through the fixed camera of the intelligent split-screen remote robot; wherein the tissue physiological motion data of the non-surgical area is acquired by: acquiring all frames of real-time video images when the end effector of the hand operates in the surgical area, and performing image segmentation on the real-time video images according to the division of the surgical area, acquiring all frames of real-time video images of the non-surgical area after the image segmentation, and using all frames of real-time video images of the non-surgical area after the image segmentation as the tissue physiological motion data of the non-surgical area; S22, using the first frame of the non-surgical area real-time video image as a reference image for all subsequent frames of the non-surgical area real-time video image, randomly selecting multiple feature points in the reference image as tracking feature points; calculating the image gradient in the x direction and the image gradient in the y direction of each tracking feature point in the reference image within a surrounding window size of w*w; and calculating the time gradient of each tracking feature point in the t-th frame of the non-surgical area real-time video image; S23, in the reference image of each tracking feature point, the image gradients in the x direction and the image gradients in the y direction of all pixels within the window size n around each tracking feature point are formed into a matrix to obtain the image gradient matrix of the tracking feature point ; Wherein, Ixi and Iyi are the image gradient in the x direction and the image gradient in the y direction of the i-th pixel, respectively; i is any one from 1 to n; S24, in the t-th frame of the real-time video image of the non-surgical area, the time gradients of all pixels within the range of a window size of n around each tracking feature point are formed into a matrix to obtain the time gradient matrix of the tracking feature point in the t-th frame of the real-time video image of the non-surgical area ; Wherein, the temporal gradient is the difference between the pixel value of all pixels at the position in the real-time video image of the non-surgical area of the t-th frame and the pixel value of the pixel at the same position in the real-time video image of the non-surgical area of the t+1-th frame; Wherein, Iti represents the temporal gradient of the i-th pixel in the real-time video image of the non-surgical area of the t-th frame with each tracking feature point as the center and the surrounding window size of n; S25, substituting the image gradient matrix of the tracking feature points and the time gradient matrix of the tracking feature points in the t-th frame non-surgical area real-time video image into the displacement vector calculation formula to calculate the displacement vector of each tracking feature point from the reference image to the t-th frame non-surgical area real-time video image; the displacement vector calculation formula is: ; Where Vt represents the displacement vector of each tracking feature point from the reference image to the real-time video image of the non-surgical area in the tth frame, represents the transpose of the image gradient matrix about the tracked feature points, It represents the inverse matrix of the matrix in brackets; Respectively represent the displacement vector of each tracking feature point from the reference image to the real-time video image of the non-surgical area in the tth frame in the x direction and the y direction; S26, the displacement vector of each tracking feature point from the reference image to the real-time video image of the non-surgical area of the t-th frame in the x-direction and the y-direction and the position of each tracking feature point in the reference image Substitute into the position change calculation formula to calculate the position of each tracking feature point in the real-time video image of the non-surgical area in the tth frame; the position change calculation formula is: ; In the formula, They respectively represent the abscissa and ordinate of the position of each tracking feature point in the real-time video image of the non-surgical area in the tth frame.
3. According to claim 2, the AI-based intelligent split-screen remote robotic surgery communication control method is characterized in that: The step S2 also includes the following specific contents: S27, obtaining the calculated position of each tracking feature point in all non-surgical area real-time video images; constructing a position time series of each tracking feature point according to the time sequence of the position of each tracking feature point in all non-surgical area real-time video images; Based on the position time series of each tracking feature point, a non-surgical area position offset prediction model is trained to predict the position of each tracking feature point in multiple frames of non-surgical area real-time video images in the future; S28, presetting the sliding step length to L and the sliding window length to W; A sliding window method is used to convert the position of each tracking feature point in the position time series of each tracking feature point in all non-surgical area real-time video images into multiple training samples, wherein each training sample is composed of the position time series of each tracking feature point in the sliding window, the training samples are used as the input of the non-surgical area position offset prediction model, the position of each tracking feature point in the non-surgical area real-time video image with a preset sliding step size of L is used as the output, and the prediction accuracy is used as the training target to train the non-surgical area position offset prediction model; a non-surgical area position offset prediction model is generated that predicts the position of each tracking feature point in multiple frames of non-surgical area real-time video images in the future according to the position of each tracking feature point in all non-surgical area real-time video images; wherein the non-surgical area position offset prediction model is a recurrent neural network model; S29, obtaining the predicted position of each tracking feature point in the future multi-frame real-time video image of the non-surgical area; converting the position of each tracking feature point in the future multi-frame real-time video image of the non-surgical area into the world coordinate system through the image coordinate system, thereby obtaining the actual position of each tracking feature point in the future multi-frame real-time video image of the non-surgical area.
4. According to claim 3, the AI-based intelligent split-screen remote robotic surgery communication control method is characterized in that: The step S3 comprises the following specific steps: S31, obtaining the input position data of the master hand end and the real-time position data of the slave hand end effector; wherein, obtaining the input position data of the master hand end includes: extracting the displacement of the master hand input corresponding to the time point of each frame of real-time video image as the master hand input displacement of each frame of real-time video image; taking the master hand input displacement of all frames of real-time video images as the input position data of the master hand end; the real-time position data of the slave hand end effector includes: the real-time position of the slave hand end effector in all frames of real-time video images; S32, constructing a master-slave hand position time series according to the time sequence of the master hand input position data and the real-time position data of the slave hand end effector, and training a slave hand position change prediction model for predicting the position of the slave hand end effector in multiple frames of real-time video images in the future based on the master-slave hand position time series; S33, preset the sliding step size to be Lc and the sliding window length to be Wc; use the sliding window method to convert the input position data of the master hand end in the master-slave hand position time series and the real-time position data of the slave hand end effector into multiple training samples, wherein each training sample is composed of the master-slave hand position data sequence in the sliding window, use the training sample as the input of the slave hand position change prediction model, use the position of the slave hand end effector in the future multi-frame real-time video image with a preset sliding step size of Lc as the output, and train the slave hand position change prediction model with the prediction accuracy as the training target; generate a slave hand position change prediction model that predicts the position of the slave hand end effector in the future multi-frame real-time video image based on the input position data of the master hand end and the real-time position data of the slave hand end effector; wherein the slave hand position change prediction model is a recurrent neural network model; S34. Obtain the predicted position of the slave hand end effector in the future multi-frame real-time video image; convert the position of the slave hand end effector in the future multi-frame real-time video image into the world coordinate system through the image coordinate system to obtain the actual position of the slave hand end effector in the future multi-frame real-time video image.
5. According to claim 4, the AI-based intelligent split-screen remote robotic surgery communication control method is characterized in that: The step S5 includes the following specific contents: S51, when the actual position of the slave hand end effector does not belong to the actual position coordinate set of the surgical area, immediately send a contact risk warning to the master hand end of the Zhilian split-screen remote robot; S52. When the actual position of the slave hand end effector belongs to the actual position coordinate set of the surgical area, a contact risk threshold is preset. When the contact risk coefficient between the slave hand end effector and the non-surgical area in the future rth frame of real-time video image is greater than the contact risk threshold, a contact risk warning is immediately issued to the master hand end of the intelligent split-screen remote robot.
6. An AI-based intelligent split-screen remote robotic surgery communication control system, which is implemented based on an AI-based intelligent split-screen remote robotic surgery communication control method according to any one of claims 1 to 5, characterized in that: The system comprises: The data acquisition module is used to obtain the input position data of the master hand end of the Zhilian split-screen remote robot and the real-time position data of the slave hand end effector; at the same time, it obtains the physiological movement data of the tissue in the non-surgical area; A non-surgical area position deviation prediction module is used to import the tissue physiological movement data of the non-surgical area into the non-surgical area position deviation prediction model to predict the position deviation change of the non-surgical area; The slave hand position change prediction module is used to import the input position data of the master hand end and the real-time position data of the slave hand end effector into the slave hand position change prediction model to predict the position change of the slave hand end effector; A contact risk analysis module, used for analyzing the contact risk between the slave hand end effector and the non-surgical area according to the position offset change prediction result of the non-surgical area and the position change prediction result of the slave hand end effector; The contact risk warning module is used to send a contact risk warning to the main hand end of the Zhilian split-screen remote robot based on the contact risk analysis results; A control module is used to control the operation of the data acquisition module, the non-surgical area position offset prediction module, the hand position change prediction module, the contact risk analysis module and the contact risk warning module.
7. An electronic device comprising: A processor and a memory, wherein the memory stores a computer program that can be called by the processor; characterized in that the processor executes an AI-based intelligent split-screen remote robot surgery communication control method as described in any one of claims 1 to 5 by calling the computer program stored in the memory.
Citation Information
Patent Citations
Regional risk early warning method and system for surgical robot
CN115661709A
Reactive interaction for robotic applications and other automation systems
CN116787423A