Control system and method of percutaneous surgical robot

By performing 3D reconstruction and real-time image registration on preoperative image data, a predicted target pose is generated, which solves the problem of real-time dynamic updating and precise adjustment of existing transoral surgical robot systems, and improves the safety and accuracy of the surgery.

CN121730992APending Publication Date: 2026-03-27JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing transoral surgical robot systems lack the ability to update and precisely adjust in real time in three dimensions, which limits the accuracy of surgical operations, especially when the surgical area changes and it is difficult to effectively provide real-time feedback.

Method used

By collecting preoperative medical imaging data for 3D reconstruction, combining real-time images acquired by a 3D endoscope for geometric registration, a real-time 3D point cloud of the laryngeal cavity is generated. The predicted target pose is then generated by combining natural motion laws. Inverse kinematics solutions and risk control are used to regulate the autonomous control of the robot.

Benefits of technology

High-precision 3D model updates have been achieved, ensuring that the robot's end effector can accurately respond to changes in the morphology of the laryngeal cavity, thereby improving the safety, accuracy, and flexibility of the surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121730992A_ABST
    Figure CN121730992A_ABST
Patent Text Reader

Abstract

The invention discloses a control system and method for an oral surgery robot, and relates to the technical field of medical robot control, and the method comprises the steps: collecting preoperative medical image data of a laryngeal cavity region, and generating a static laryngeal cavity three-dimensional structural body through anatomical segmentation, contour extraction and three-dimensional reconstruction; according to the static laryngeal cavity three-dimensional structural body, zero calibration is carried out on the oral surgical robot, and RCM constraints are established; under the constraint of RCM, a three-dimensional endoscope is used for collecting an operation area image in real time, a real-time laryngeal cavity local three-dimensional point cloud is generated through stereoscopic vision reconstruction, geometric registration is conducted on the real-time laryngeal cavity local three-dimensional point cloud and a static laryngeal cavity three-dimensional structural body, and the real-time tail end pose is determined; and the primary control quantity and the three types of real-time monitoring data are fused together, a comprehensive risk index is calculated, the autonomous control proportion of the robot is adjusted by using the comprehensive risk index, and a risk guidance control quantity is generated. The safety, accuracy and flexibility of the operation are improved, and finally more efficient and safer oral operation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical robot control technology, and in particular to control systems and methods for transoral surgical robots. Background Technology

[0002] With the development of minimally invasive surgery, transoral surgery, as a surgical method with great clinical application potential, has gradually become a common technique in otolaryngology, stomatology, and the treatment of some head and neck tumors. Because it does not require large incisions, avoiding the trauma of traditional open surgery, it has been widely used clinically due to its smaller incision, shorter recovery period, and fewer complications. With the advancement of robotics technology, robot-assisted transoral surgical systems have emerged, which can assist doctors in performing complex surgical procedures under high-precision control.

[0003] Existing preoperative imaging data is mostly used for preoperative planning, and its ability to collect and process real-time information during surgery remains significantly limited. Although some transoral surgical robots can utilize 3D images for assisted localization, these technologies often only provide static anatomical images, lacking dynamic, real-time updated 3D information. This limits the precision of robot operation during execution, especially when changes occur in the surgical area or when manipulating the patient's natural anatomical structures, making it difficult to effectively provide real-time feedback on the surgical situation. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a control method for transoral surgical robots to solve the problem of lack of real-time dynamic three-dimensional updates and precise adjustments in the prior art.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a control method for a transoral surgical robot, which includes acquiring preoperative medical imaging data of the laryngeal cavity region and generating a static three-dimensional structure of the laryngeal cavity through anatomical segmentation, contour extraction and three-dimensional reconstruction. Zero-point calibration of the transoral surgical robot is performed based on the static three-dimensional structure of the laryngeal cavity to establish RCM constraints. Under the RCM constraints, images of the surgical area are acquired in real time using a three-dimensional endoscope. Real-time local three-dimensional point clouds of the laryngeal cavity are generated through stereo vision reconstruction and geometrically registered with the static three-dimensional structure of the laryngeal cavity to determine the real-time end pose. The static three-dimensional structure of the laryngeal cavity is non-rigidly registered with the real-time local three-dimensional point cloud of the laryngeal cavity, and the natural motion law of key anatomical regions is extracted by combining the real-time end pose to generate the predicted end pose of the target. The preoperative planned path and predicted target end pose are solved by inverse kinematics and attitude constraint analysis to generate primary control variables. The primary control variables are then integrated with three types of real-time monitoring data to calculate a comprehensive risk index. The comprehensive risk index is used to adjust the autonomous control ratio of the robot to generate risk-guided control variables.

[0007] As a preferred embodiment of the control method for the transoral surgical robot of the present invention, the preoperative medical imaging data includes thin-slice CT images and respiratory status images.

[0008] In a preferred embodiment of the control method for the transoral surgical robot of the present invention, the step of generating a static three-dimensional structure of the laryngeal cavity through anatomical segmentation, contour extraction, and three-dimensional reconstruction includes the following specific steps. The preoperative medical imaging data was denoised and normalized to obtain standard preoperative medical imaging data. The Otsu algorithm was used to perform anatomical segmentation on the standard preoperative medical imaging data to distinguish between soft and hard tissues in the laryngeal cavity and obtain laryngeal cavity segmentation images. The outer contour of the laryngeal cavity region is extracted from the laryngeal cavity segmentation image, and the internal structure of the laryngeal cavity is segmented from the laryngeal cavity segmentation image using a region growing algorithm; The outer contour of the laryngeal cavity region and the internal structure of the laryngeal cavity are converted into a uniform three-dimensional voxel mesh, and three-dimensional point cloud data are formed by three-dimensional rasterization. Each point cloud in the 3D point cloud data is connected to its adjacent point clouds to form a triangular mesh, generating a 3D surface structure of the laryngeal cavity. The static 3D structure of the laryngeal cavity is obtained through smoothing and hole filling.

[0009] In a preferred embodiment of the control method for the transoral surgical robot described in this invention, the steps of performing zero-point calibration on the transoral surgical robot based on the static three-dimensional structure of the laryngeal cavity and establishing RCM constraints are as follows: From the static three-dimensional structure of the laryngeal cavity, a landmark structure of the laryngeal cavity region is selected as a reference point. The position of the robot's end effector is aligned with the reference point to determine the zero point of the robot's coordinate system. Based on the zero point position of the robot coordinate system, the coordinate transformation relationship between the robot coordinate system and the static three-dimensional structure of the laryngeal cavity is established through rigid transformation, and RCM constraints are set.

[0010] As a preferred embodiment of the control method for the transoral surgical robot of the present invention, the step of acquiring images of the surgical area in real time using a three-dimensional endoscope under RCM constraints, and generating a real-time local three-dimensional point cloud of the laryngeal cavity through stereoscopic vision reconstruction, includes the following specific steps. Under RCM constraints, dual cameras in a 3D endoscope are used to acquire images of the surgical area from different angles in real time to obtain dual-view surgical area images; grayscale processing is performed on the dual-view surgical area images, corner feature points are extracted and image alignment is performed to obtain aligned surgical area images; The disparity values ​​between aligned surgical area images are calculated, the depth value of each pixel is estimated using the disparity values, and the depth value of each pixel is converted into three-dimensional coordinates using triangulation to generate a real-time local three-dimensional point cloud of the laryngeal cavity.

[0011] In a preferred embodiment of the control method for the transoral surgical robot of the present invention, the specific steps for determining the real-time end-effector pose are as follows: The closest point between the real-time local 3D point cloud of the laryngeal cavity and the static 3D structure of the laryngeal cavity is matched and the distance error between the two is minimized to obtain the registered 3D point cloud; By registering the correspondence between the 3D point cloud and the static 3D structure of the throat cavity, the real-time end effector pose of the robot in the throat cavity region is calculated and determined.

[0012] As a preferred embodiment of the control method for the transoral surgical robot of the present invention, the steps of performing non-rigid registration between the static three-dimensional structure of the laryngeal cavity and the real-time local three-dimensional point cloud of the laryngeal cavity, and extracting the natural motion patterns of key anatomical regions based on the real-time end-effector pose to generate the predicted end-effector pose are as follows. Under a unified reference coordinate system, the static three-dimensional structure of the laryngeal cavity is used as the reference three-dimensional structure. The real-time local three-dimensional point cloud of the laryngeal cavity is non-rigidly registered with the reference three-dimensional structure, and the overall deformation relationship is obtained by calculating the local deformation value. By analyzing the dynamic changes of the laryngeal cavity in relation to the real-time end-effector pose and overall deformation, the movement patterns of anatomical landmarks during the operation are determined, and the relative positional changes between the real-time end-effector pose and anatomical landmarks are calculated to obtain the natural movement patterns of key anatomical areas. By summarizing multiple relative position changes and natural motion laws, the predicted target end pose is generated.

[0013] As a preferred embodiment of the control method for the transoral surgical robot of the present invention, the step of generating primary control variables by solving the preoperative planned path and the predicted target end-effector pose through inverse kinematics and attitude constraint analysis is as follows: By combining the preoperative planned path with the predicted target end-effector pose, the inverse kinematics algorithm is used to calculate the joint angles of the robot's end effector. The pose error between the current pose of the robot end effector and the predicted target end effector pose is calculated, and the change in the angle of each joint is adjusted based on the pose error. The joint angle of the robot end effector is corrected through a feedback mechanism to obtain the primary control quantity.

[0014] In a preferred embodiment of the control method for the transoral surgical robot described in this invention, the steps of integrating the primary control quantity with three types of real-time monitoring data to calculate a comprehensive risk index, and then using the comprehensive risk index to adjust the robot's autonomous control ratio to generate a risk-guided control quantity are as follows. Collect visual monitoring data, mechanical and morphological monitoring data, and safety monitoring data to form three types of real-time monitoring data; The primary control parameters are integrated with three types of real-time monitoring data to obtain a comprehensive control information set, and a comprehensive risk index is calculated based on the risk assessment weights. Based on the comprehensive risk index, the proportion of autonomous control of the robot is adjusted, and risk-guided control parameters are generated in combination with the control requirements of the surgical robot.

[0015] In a second aspect, the present invention provides a control system for a transoral surgical robot, including an image acquisition and reconstruction module for acquiring preoperative medical image data of the laryngeal cavity region and generating a static three-dimensional structure of the laryngeal cavity through anatomical segmentation, contour extraction and three-dimensional reconstruction. The calibration and constraint module is used to perform zero-point calibration on the transoral surgical robot based on the static three-dimensional structure of the laryngeal cavity and establish RCM constraints. Under the RCM constraints, the surgical area image is acquired in real time using a three-dimensional endoscope, and a real-time local three-dimensional point cloud of the laryngeal cavity is generated through stereo vision reconstruction. This point cloud is then geometrically registered with the static three-dimensional structure of the laryngeal cavity to determine the real-time end pose. The point cloud registration calculation module is used to perform non-rigid registration between the static three-dimensional structure of the laryngeal cavity and the real-time local three-dimensional point cloud of the laryngeal cavity, and to extract the natural motion law of key anatomical regions by combining the real-time end pose to generate the predicted end pose of the target. The risk control adjustment module is used to generate primary control variables by solving the preoperative planned path and the predicted target end pose through inverse kinematics and attitude constraint analysis. The primary control variables are then integrated with three types of real-time monitoring data to calculate a comprehensive risk index. The comprehensive risk index is used to adjust the autonomous control ratio of the robot and generate risk-guided control variables.

[0016] The beneficial effects of this invention are as follows: by combining preoperative thin-slice CT images and respiratory status images, the anatomical structure and dynamic changes of the laryngeal cavity can be accurately grasped, providing a high-precision three-dimensional model for the surgical robot and reducing the risk of intraoperative misoperation; by using a three-dimensional endoscope to generate a real-time local three-dimensional point cloud of the laryngeal cavity, combined with real-time feedback of RCM constraints, it is ensured that the robot's end effector can accurately respond to changes in the morphology of the laryngeal cavity, improving the safety, accuracy and flexibility of the surgery, and ultimately achieving more efficient and safer transoral surgical operations. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the control method for a transoral surgical robot.

[0019] Figure 2 This is a schematic diagram of the control system of a transoral surgical robot.

[0020] Figure 3 A flowchart for generating a static three-dimensional structure of the laryngeal cavity.

[0021] Figure 4 This is a flowchart for real-time generation of local 3D point clouds in the laryngeal cavity.

[0022] Figure 5 This is a comparison chart showing the adjustment of the end effector path error under RCM control conditions.

[0023] Figure 6 A comparative chart of strategies for responding to the risk index of the surgical procedure. Detailed Implementation

[0024] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0025] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0026] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0027] Reference Figures 1-6 As one embodiment of the present invention, this embodiment provides a control method for a transoral surgical robot, comprising the following steps: S1. Collect preoperative medical imaging data of the laryngeal cavity region, and generate a static three-dimensional structure of the laryngeal cavity through anatomical segmentation, contour extraction and three-dimensional reconstruction.

[0028] It should be noted that the patient's head and neck are fixed on a multi-slice spiral CT scanning bed, and the thin-slice CT slice thickness, pitch and scanning range are set. Continuous thin-slice CT scanning is performed on the laryngeal cavity to obtain thin-slice CT images. Subsequently, combined with a respiratory monitoring device, thin-slice CT acquisition is triggered at the end of inspiration and the end of expiration to obtain respiratory status images at different respiratory phases. The thin-slice CT images and respiratory status images are stored together as preoperative medical imaging data.

[0029] S1.1. Denoise and normalize the preoperative medical image data to obtain standard preoperative medical image data. Use the Otsu algorithm to perform anatomical segmentation on the standard preoperative medical image data to distinguish between soft and hard tissues in the laryngeal cavity and obtain laryngeal cavity segmentation images.

[0030] Furthermore, spatial domain filtering algorithms, such as mean filtering or Gaussian filtering, are used to perform convolution operations on the neighborhood of each pixel in the preoperative medical image data to reduce high-frequency noise. Subsequently, a linear transformation is performed on the grayscale value of each pixel in the denoised preoperative medical image data to map the grayscale value to a uniform grayscale range, thereby eliminating grayscale differences caused by different acquisition conditions and obtaining standard preoperative medical image data. The Otsu algorithm is used to traverse the grayscale histogram of the standard preoperative medical image data to calculate the intra-class variance corresponding to each candidate threshold, and the grayscale threshold that minimizes the intra-class variance is selected as the segmentation threshold. Each pixel in the standard preoperative medical image data is divided into soft tissue pixels and hard tissue pixels in the laryngeal cavity region according to the segmentation threshold, resulting in a laryngeal cavity segmentation image.

[0031] S1.2 Extract the outer contour of the laryngeal cavity region from the laryngeal cavity segmentation image, and use a region growing algorithm to segment the internal structure of the laryngeal cavity from the laryngeal cavity segmentation image.

[0032] Furthermore, contour extraction processing is performed on the laryngeal cavity segmentation image. A connected component tracking method or edge detection method is used to search for pixels belonging to the laryngeal cavity region boundary in the segmented image pixel by pixel. All interconnected laryngeal cavity region boundary pixels in the segmented image are sequentially connected to form the complete outer contour of the laryngeal cavity region. Seed pixels located inside the laryngeal cavity region are selected in the segmented image. A region growing algorithm is used to compare the grayscale value and label information of each candidate neighbor pixel adjacent to the seed pixel. When the grayscale value and label information of the candidate neighbor pixel meet the preset region growing similarity condition, the candidate neighbor pixel is added to the laryngeal cavity internal structure region. The neighborhood expansion and similarity determination steps are repeated, using the updated laryngeal cavity internal structure region boundary pixels as the new growth starting point, until there are no more candidate neighbor pixels in the segmented image that meet the region growing similarity condition. The laryngeal cavity internal structure is then segmented from the segmented image.

[0033] It should be noted that the region growth similarity condition is set for the region growth seed pixels by using the gray values ​​of the region growth seed pixels in the laryngeal cavity segmentation image as reference gray values.

[0034] Region growing is an image segmentation method based on pixel similarity. It selects seed pixels located within the target region and recursively expands neighboring pixels based on the similarity of features such as grayscale value, label information, or texture. Neighboring pixels that meet the region growing similarity conditions are continuously added to the current region, thereby gradually forming a complete and coherent target region structure.

[0035] S1.3. Convert the outer contour of the laryngeal cavity region and the internal structure of the laryngeal cavity into a uniform three-dimensional voxel mesh, and form three-dimensional point cloud data through three-dimensional rasterization.

[0036] Furthermore, based on the spatial resolution of preoperative medical imaging data, a uniform three-dimensional voxel grid covering the laryngeal cavity region is divided in a three-dimensional coordinate space at fixed intervals. The pixel coordinates of the outer contour of the laryngeal cavity region and the internal structure of the laryngeal cavity in each tomographic slice are mapped to the three-dimensional coordinate space. The corresponding voxel index of each pixel in the uniform three-dimensional voxel grid is calculated and the corresponding voxel is marked as an occupied voxel. Then, three-dimensional rasterization processing is performed on all occupied voxels in the uniform three-dimensional voxel grid. A three-dimensional coordinate point is generated at the geometric center of each occupied voxel and all three-dimensional coordinate points are stored in coordinate order to form three-dimensional point cloud data.

[0037] S1.4 Connect each point cloud in the three-dimensional point cloud data with the adjacent point clouds to form a triangular mesh to generate a three-dimensional surface structure of the laryngeal cavity. Obtain the static three-dimensional structure of the laryngeal cavity through smoothing and hole filling.

[0038] Furthermore, for each point cloud in the 3D point cloud data, the nearest neighboring point clouds are searched based on spatial proximity. Every three spatially close and non-collinear point clouds are combined into a triangular patch. By traversing the 3D point cloud data, a set of triangular patches covering the laryngeal cavity region is generated, thereby converting the 3D point cloud data into a 3D surface structure of the laryngeal cavity composed of multiple triangular patches. By performing a weighted average calculation on the coordinates of each vertex and the coordinates of adjacent vertices in the 3D surface structure of the laryngeal cavity, geometrical abrupt changes caused by local noise are reduced. By traversing the edges of all triangular patches in the 3D surface structure of the laryngeal cavity and identifying the edges occupied by only a single triangular patch, all edges occupied by only a single triangular patch are combined into a closed boundary loop according to the connection order. New triangular patches are generated inside the closed boundary loop to fill the surface gaps, resulting in a continuous, closed, and static 3D structure of the laryngeal cavity that can completely represent the shape of the laryngeal cavity region.

[0039] S2. Perform zero-point calibration on the transoral surgical robot based on the static three-dimensional structure of the laryngeal cavity and establish RCM constraints.

[0040] S2.1 Select a landmark structure in the laryngeal cavity region from the static three-dimensional laryngeal cavity structure as a reference point, align the position of the robot end effector with the reference point, and determine the zero point position of the robot coordinate system.

[0041] Furthermore, in the static three-dimensional structure of the laryngeal cavity, a landmark structure of the laryngeal cavity region that can be clearly identified simultaneously in preoperative images and intraoperative vision is first selected, and the spatial position of the landmark structure of the laryngeal cavity region in the static three-dimensional structure of the laryngeal cavity is determined as a reference point; then, guided by intraoperative images, the robot end effector is gradually moved along a preset entry path to the actual anatomical position corresponding to the reference point, and the working tip of the robot end effector is aligned with the position corresponding to the reference point; after the working tip of the end effector is aligned with the reference point, the position of the working tip of the end effector is set as the zero point of the robot coordinate system.

[0042] It should be noted that the entry path is based on the spatial connectivity between the oral cavity entrance, pharyngeal passage, and target area of ​​the laryngeal cavity in the static three-dimensional structure of the laryngeal cavity. A series of spatial location points that do not conflict with the soft tissue and hard tissue of the laryngeal cavity are selected in sequence according to the anatomical passage from the oral cavity entrance to the reference point. These spatial location points are then connected into a continuous trajectory line in the order of advancing along the anatomical passage.

[0043] S2.2. Based on the zero point position of the robot coordinate system, establish the coordinate transformation relationship between the robot coordinate system and the static three-dimensional structure of the laryngeal cavity through rigid transformation, and set RCM constraints.

[0044] Furthermore, a reference point corresponding to the zero point of the robot coordinate system is selected in the static laryngeal cavity three-dimensional structure as the coordinate alignment benchmark. Through rigid rotation and rigid translation operations that keep the distance and angle relationships unchanged, the static laryngeal cavity three-dimensional structure is aligned to the robot coordinate system as a whole, so that each point in the static laryngeal cavity three-dimensional structure corresponds to a unique spatial position in the robot coordinate system, thereby forming a coordinate transformation relationship between the robot coordinate system and the static laryngeal cavity three-dimensional structure. Based on the coordinate transformation relationship between the robot coordinate system and the static laryngeal cavity three-dimensional structure, a position corresponding to the patient's oral cavity entrance or laryngeal cavity entrance is selected in the static laryngeal cavity three-dimensional structure as the remote motion center position. The spatial position of the remote motion center position in the robot coordinate system is used as the rotation constraint center, which restricts the robot end effector to rotate around the remote motion center position without producing a translation that goes out of the remote motion center position, thus setting the RCM constraint.

[0045] S3. Under RCM constraints, images of the surgical area are acquired in real time using a three-dimensional endoscope. Real-time local three-dimensional point clouds of the laryngeal cavity are generated through stereo vision reconstruction and geometrically registered with the static three-dimensional structure of the laryngeal cavity to determine the real-time end pose.

[0046] S3.1 Under RCM constraints, use dual cameras in a three-dimensional endoscope to acquire images of the surgical area from different angles in real time, and obtain dual-view surgical area images.

[0047] Furthermore, under RCM constraints, the insertion position of the 3D endoscope is kept fixed, and the 3D endoscope is controlled to adjust its viewing angle around the remote motion center position so that the left camera and the right camera in the 3D endoscope are respectively aimed at the surgical area. The left camera in the 3D endoscope continuously acquires the left-view image sequence of the surgical area, and the right camera in the 3D endoscope synchronously acquires the right-view image sequence of the surgical area. The two sets of images are then paired one by one according to the acquisition time sequence to form a dual-view surgical area image.

[0048] S3.2. Perform grayscale processing on the dual-view surgical area images, extract corner feature points, and obtain the aligned surgical area image through image alignment.

[0049] Furthermore, for each frame of the left and right view images in the dual-view surgical region image, the weighted color channel grayscale value is calculated pixel by pixel, and the dual-view surgical region image is converted into a grayscale image. In the grayscale image, a corner detection algorithm is applied to calculate the grayscale gradient or structure tensor of each pixel's neighborhood, and the intensity of grayscale change in the horizontal and vertical directions of the pixel is evaluated based on the corner response function. Subsequently, thresholding and non-maximum suppression operations are performed on the corner response values ​​to retain the pixels with the most prominent response values ​​in the neighborhood, thereby obtaining corner feature points distributed at the structural abrupt change locations in the image, and feature matching is completed based on the correspondence of corner feature points. After completing the corner feature point matching, geometric correction is performed on the left and right view grayscale images to make the left and right view grayscale images correspond in the same row direction, resulting in an aligned surgical region image.

[0050] It should be noted that corner detection algorithms are a class of feature extraction methods used to identify locations in an image where local gray-level changes are significant. They analyze the gray-level gradient changes in the neighborhood of pixels in multiple directions to find pixels that have large gray-level differences in two directions simultaneously. These pixels are typically located at sharp geometric transitions or in textured local areas, and are called corners. Common corner detection algorithms include Harris corner detection, Shi-Tomasi corner detection, and FAST feature point detection.

[0051] S3.3 Calculate the disparity value between the aligned surgical area images, estimate the depth value of each pixel through the disparity value, convert the depth value of each pixel into three-dimensional coordinates through triangulation, and generate a real-time local three-dimensional point cloud of the laryngeal cavity.

[0052] Furthermore, in the aligned surgical area image, for each pixel position, corresponding pixels with the same row coordinates in the left-view grayscale image and the right-view grayscale image are found. The disparity value is obtained by calculating the difference in column coordinates of the two corresponding pixels, and a disparity value image is generated for all pixel positions. Based on the baseline distance parameter between the left and right cameras in the 3D endoscope (referring to the fixed physical distance between the left and right cameras in the 3D endoscope) and the imaging focal length parameter of the aligned surgical area image, the disparity value of each pixel position in the disparity value image is substituted into the binocular stereo vision depth estimation formula to obtain the pixel depth value corresponding to each pixel position. Finally, based on the pixel depth value of each pixel position, the row and column coordinates of the pixel position in the aligned surgical area image, and the intrinsic and extrinsic parameters of the 3D endoscope, the pixel depth value is converted into a 3D coordinate point in the surgical area coordinate system by triangulation. All 3D coordinate points are summarized according to their spatial distribution relationship to generate a real-time local 3D point cloud of the laryngeal cavity.

[0053] The expression for calculating the pixel depth value at each pixel location is: ; In the formula, Indicates the pixel row coordinates as Pixel column coordinates are The pixel depth value at a location, which is the spatial distance (in meters) from the corresponding point to the imaging plane. This indicates that the disparity value image is located at the pixel row coordinates. Pixel column coordinates are The disparity value of the location (in pixels); The physical focal length of the three-dimensional endoscopic imaging optical system (in meters, which can be obtained, for example, from the endoscopic imaging parameters); This indicates the baseline distance (in meters) between the left and right cameras in a 3D endoscope, which can be obtained, for example, through equipment calibration. This represents the physical size (in meters per pixel) of a single pixel in the aligned surgical area image on the imaging plane. The disparity value is obtained by comparing the pixel difference between the column index of the left-view grayscale image at a certain row coordinate position and the column index of the right-view grayscale image at the same row coordinate position. Since the number of pixels corresponding to the disparity value is converted into the actual spatial length on the imaging plane through the pixel physical size parameter, a unified measurement and comparable calculation between the number of pixels and the spatial length are achieved in this depth calculation expression.

[0054] S3.4 Match the nearest point between the real-time local 3D point cloud of the laryngeal cavity and the static 3D structure of the laryngeal cavity and minimize the distance error between them to obtain the registered 3D point cloud.

[0055] Furthermore, the nearest 3D point on the surface of the static 3D structure of the laryngeal cavity is searched point by point in the real-time local 3D point cloud of the laryngeal cavity to form a set of matching point pairs. Then, the rigid rotation parameters and rigid translation parameters of the real-time local 3D point cloud of the laryngeal cavity are adjusted using the set of matching point pairs to continuously reduce the distance error between the matching point pairs. When the distance error tends to stabilize, the real-time local 3D point cloud of the laryngeal cavity after transformation by rigid rotation parameters and rigid translation parameters is used as the registration 3D point cloud.

[0056] S3.5 By registering the correspondence between the three-dimensional point cloud and the static three-dimensional structure of the throat cavity, calculate and determine the real-time end effector pose of the robot end effector in the throat cavity region.

[0057] Furthermore, by utilizing the spatial correspondence between the registered 3D point cloud and the corresponding points in the static 3D laryngeal cavity structure, the rigid transformation parameters from the real-time laryngeal cavity local coordinate system to the static laryngeal cavity 3D structure coordinate system are obtained. The rigid transformation parameters are combined with the calibration transformation relationship between the 3D endoscope imaging coordinate system and the robot coordinate system to obtain the spatial position and orientation of the robot end effector's working tip in the static laryngeal cavity 3D structure coordinate system. The spatial position and orientation of the robot end effector's working tip in the static laryngeal cavity 3D structure coordinate system are used as the real-time end effector pose of the robot end effector in the laryngeal cavity region.

[0058] S4. Perform non-rigid registration between the static three-dimensional structure of the laryngeal cavity and the real-time local three-dimensional point cloud of the laryngeal cavity, and extract the natural motion law of key anatomical regions by combining the real-time end pose to generate the predicted end pose of the target.

[0059] S4.1 Under a unified reference coordinate system, the static three-dimensional structure of the laryngeal cavity is used as the reference three-dimensional structure. The real-time local three-dimensional point cloud of the laryngeal cavity is non-rigidly registered with the reference three-dimensional structure, and the overall deformation relationship is obtained by calculating the local deformation value.

[0060] Furthermore, the real-time local 3D point cloud of the laryngeal cavity is transformed into a coordinate system consistent with the static 3D structure of the laryngeal cavity through a rigid transformation relationship, and the static 3D structure of the laryngeal cavity is used as the reference 3D structure. A non-rigid registration algorithm is used to establish a spatial correspondence between the reference 3D structure and the real-time local 3D point cloud of the laryngeal cavity. By iteratively adjusting the non-rigid deformation parameters, the real-time local 3D point cloud of the laryngeal cavity gradually conforms to the surface morphology of the static 3D structure of the laryngeal cavity. In each deformation update process, the displacement vector between the corresponding point in the static 3D structure of the laryngeal cavity and the corresponding point in the real-time local 3D point cloud of the laryngeal cavity is recorded as the local deformation value. The local deformation values ​​of all corresponding points in the static 3D structure of the laryngeal cavity are combined in spatial order to obtain the overall deformation relationship covering the laryngeal cavity region.

[0061] It should be noted that the non-rigid registration algorithm is an alignment method that allows point clouds or 3D models to undergo local deformation during the registration process. By establishing a deformable deformation field in the data to be registered or using a control point mesh, the registration error and deformation smoothness are minimized simultaneously in the iterative optimization, so that the deformed point cloud can gradually fit the reference structure, thereby obtaining a spatial correspondence with elastic matching characteristics.

[0062] S4.2 By analyzing the dynamic changes of the laryngeal cavity in relation to the real-time end-effector pose and overall deformation, the motion patterns of anatomical landmarks during the operation are determined, and the relative positional changes between the real-time end-effector pose and anatomical landmarks are calculated to obtain the natural motion patterns of key anatomical areas.

[0063] Furthermore, multiple anatomical landmarks, such as the glottic margin and the anterior edge of the epiglottis, were selected within the static three-dimensional structure of the laryngeal cavity. The three-dimensional displacement trajectory of each anatomical landmark in continuous time frames was read using the overall deformation relationship. Statistical processing of the displacement magnitude, direction, and time period of the anatomical landmark's three-dimensional displacement trajectory identified the periodic motion pattern and slow drift trend of the anatomical landmark during the surgical procedure. In each time frame, the three-dimensional coordinates of the working tip of the robot's end effector in the real-time end-effector pose were subtracted from the three-dimensional coordinates of each anatomical landmark in the corresponding time frame to obtain the relative position change sequence between the real-time end-effector pose and the anatomical landmark. The amplitude, direction, and time variation patterns of the relative position change sequence were then analyzed and summarized to form the natural motion law of key anatomical regions describing the motion direction, amplitude, and period of key anatomical regions over time.

[0064] It should be noted that anatomical landmarks that significantly affect surgical safety and operational accuracy, such as the glottis, epiglottis, and laryngeal inlet, are selected. Then, the spatial motion range of these anatomical landmarks within a continuous time frame is determined using the overall deformation relationship. The local area covered by the anatomical landmarks with a large spatial motion range, frequent position changes, and close distance to the working tip of the robot's end effector is taken as the key anatomical area.

[0065] S4.3 Summarize multiple relative position changes and natural motion laws to generate the predicted target end pose.

[0066] Furthermore, the relative position changes of the robot end effector's working tip relative to different anatomical landmarks in each time frame are organized in chronological order, and the periodic change trend and displacement direction trend in the relative position changes are extracted based on the natural motion law of key anatomical regions. Then, the periodic change trend and displacement direction trend are superimposed on the real-time end effector pose to extrapolate the spatial position and attitude of the robot end effector's working tip in the next time period, forming the predicted target end effector pose.

[0067] S5. The preoperative planned path and the predicted target end pose are used to generate primary control variables through inverse kinematics solution and attitude constraint analysis.

[0068] S5.1 Combine the preoperative planned path with the predicted target end pose, and use the inverse kinematics algorithm to calculate the joint angles of the robot end effector.

[0069] Furthermore, the path target point that is closest to the predicted target end-effector pose is selected on the preoperative planning path, and the spatial position of the path target point is combined with the attitude information of the predicted target end-effector pose to generate the target end-effector pose of the robot. The target end-effector pose is used as the input for inverse kinematics solution. Under the conditions of satisfying RCM constraints, joint range of motion constraints and joint collision constraints, the angles of each joint of the robot end-effector are solved by inverse kinematics algorithm.

[0070] It should be noted that the joint motion range constraint is achieved by retrieving the joint motion parameters generated during the manufacturing or calibration phase of the robot end effector, using the lower limit and upper limit of the movable angle of each joint as the constraint boundary, and limiting the real-time joint angle of the joint to the range between the corresponding lower limit and upper limit of the movable angle. Joint collision constraints identify the spatial range of soft and hard tissues in the laryngeal cavity region within the static three-dimensional structure of the laryngeal cavity, and set a minimum safe distance threshold based on the intraoperative allowable instrument contact safety boundary. The minimum safe distance between the joint link and the static three-dimensional structure of the laryngeal cavity is not less than the minimum safe distance threshold as a constraint condition. When the inverse kinematics solution results in the minimum safe distance being less than the minimum safe distance threshold, it is considered a collision, thus prohibiting the setting of that joint angle combination. The minimum safe distance threshold is determined based on the spatial distribution characteristics of the glottic region, laryngeal inlet region, and surrounding soft tissues in the static three-dimensional structure of the laryngeal cavity, and is set as the minimum separation distance that the instrument link should maintain between the instrument and the anatomical structure.

[0071] The minimum safe distance threshold is set based on the spatial distribution characteristics of soft and hard tissues inside the laryngeal cavity and the safe boundary of instrument contact allowed during surgery. The example value ranges from 2mm to 5mm.

[0072] S5.2 Calculate the pose error between the current pose of the robot end effector and the predicted target end effector pose, and calculate the change in the angle of each joint based on the pose error. Correct the joint angle of the robot end effector through a feedback mechanism to obtain the primary control quantity.

[0073] Furthermore, the pose error is obtained by performing a difference operation between the current position vector and current pose angle vector of the robot end effector and the predicted position vector and pose angle vector of the predicted target end effector pose. The pose error vector is then converted into a joint angle change vector in joint space using the Jacobian matrix of the robot end effector. Given the pose error vector and the joint angle adjustment coefficient, the change in each joint angle is calculated, and the joint angle change vector is superimposed with the current joint angle vector to form a new joint angle command vector. By using the new joint angle command vector as feedback control output, a primary control quantity is generated to drive the robot end effector to approximate the predicted target end effector pose.

[0074] It should be noted that, given the positioning posture error vector and the joint angle adjustment scaling factor, the change in each joint angle is calculated using the following expression: ; In the formula, It is the first in the vector of joint angle change. The change in angle of each joint, in radians; Represents the joint angle adjustment ratio matrix The Middle The joint angle adjustment ratio coefficient corresponding to the first joint is a dimensionless scaling factor used to control the effect of pose error on the first joint angle change when mapped to the joint angle change. The response amplitude of each joint; Indicates the first The pose error component affects the first The linear influence coefficient of the joint angle change is obtained by taking the partial derivative of the end displacement and attitude change with respect to each joint angle, and is used to convert the pose error into the angle adjustment in the joint space. It is the index of the pose error components; It is an index of the joints; It is the first in the pose error vector There are three pose error components. The first three components are examples of position error components, with the unit being meters. The last three components are examples of attitude error components, with the unit being radians.

[0075] S6. Integrate the primary control quantity with three types of real-time monitoring data to calculate the comprehensive risk index. Use the comprehensive risk index to adjust the robot's autonomous control ratio and generate risk-guided control quantity.

[0076] S6.1 Collect visual monitoring data, mechanical and morphological monitoring data, and safety monitoring data to form three types of real-time monitoring data.

[0077] Furthermore, visual monitoring data is generated by real-time acquisition of images of the surgical area and information such as the distance between instruments and tissues through a 3D endoscope and external camera equipment. At the same time, mechanical and morphological monitoring data is generated by real-time acquisition of information such as contact force, joint torque and instrument bending shape using force sensors and deformation sensors. Simultaneously, information such as heart rate, blood pressure and emergency stop signals are collected from vital sign monitoring equipment and robot operation status monitoring device to generate safety monitoring data, thereby obtaining three types of real-time monitoring data.

[0078] S6.2 Integrate primary control quantities with three types of real-time monitoring data to obtain a comprehensive control information set, and calculate a comprehensive risk index based on risk assessment weights.

[0079] Furthermore, primary control parameters, visual monitoring data, mechanical and morphological monitoring data, and safety monitoring data are aggregated according to time sequence and data source to form a comprehensive control information set. Subsequently, multiple risk indicators reflecting the degree of surgical risk are extracted based on the comprehensive control information set, and the multiple risk indicators are weighted according to risk assessment weights to obtain a comprehensive risk index.

[0080] It should be noted that the risk assessment weights are determined by organizing and comparing past surgical records, and by setting different risk assessment weights according to the degree of correlation between the risk indicators and the risk events, based on the differences in the magnitude, speed, and frequency of changes of different risk indicators before and after the occurrence of the risk event.

[0081] S6.3. Based on the comprehensive risk index, adjust the proportion of autonomous control of the robot, and generate risk-guided control quantity in combination with the control requirements of the surgical robot.

[0082] Furthermore, based on the range of the comprehensive risk index, the corresponding robot autonomous control ratio is selected from the comparison table of comprehensive risk index and robot autonomous control ratio. The robot autonomous control ratio is used to determine the weight relationship between the primary control quantity and the manual control quantity. Combining the requirements of the surgical robot control requirements regarding the upper limit of motion speed, end force limitation and sensitive area working mode, the primary control quantity obtained under the robot autonomous control ratio constraint is subjected to amplitude scaling, speed limitation and sensitive area deceleration processing to form a set of control instructions that meet the control requirements of the surgical robot. The set of control instructions is used as the risk-guided control quantity.

[0083] This embodiment also provides a control system for the transoral surgical robot, including: an image acquisition and reconstruction module, used to acquire preoperative medical image data of the laryngeal cavity region, and generate a static three-dimensional structure of the laryngeal cavity through anatomical segmentation, contour extraction and three-dimensional reconstruction; The calibration and constraint module is used to perform zero-point calibration on the transoral surgical robot based on the static three-dimensional structure of the laryngeal cavity and establish RCM constraints. Under the RCM constraints, the surgical area image is acquired in real time using a three-dimensional endoscope, and a real-time local three-dimensional point cloud of the laryngeal cavity is generated through stereo vision reconstruction. This point cloud is then geometrically registered with the static three-dimensional structure of the laryngeal cavity to determine the real-time end pose. The point cloud registration calculation module is used to perform non-rigid registration between the static three-dimensional structure of the laryngeal cavity and the real-time local three-dimensional point cloud of the laryngeal cavity, and to extract the natural motion law of key anatomical regions by combining the real-time end pose to generate the predicted end pose of the target. The risk control adjustment module is used to generate primary control variables by solving the preoperative planned path and the predicted target end pose through inverse kinematics and attitude constraint analysis. The primary control variables are then integrated with three types of real-time monitoring data to calculate a comprehensive risk index. The comprehensive risk index is used to adjust the autonomous control ratio of the robot and generate risk-guided control variables.

[0084] In summary, this invention achieves the following: by combining preoperative thin-slice CT images and respiratory status images, it accurately grasps the anatomical structure and dynamic changes of the laryngeal cavity, providing a high-precision three-dimensional model for the surgical robot and reducing the risk of intraoperative errors; by using a three-dimensional endoscope to generate real-time local three-dimensional point clouds of the laryngeal cavity, combined with real-time feedback of RCM constraints, it ensures that the robot's end effector can accurately respond to changes in the morphology of the laryngeal cavity, improving the safety, accuracy, and flexibility of the surgery, and ultimately achieving more efficient and safer transoral surgical operations.

[0085] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A control method for a transoral surgical robot, characterized in that: include, Preoperative medical imaging data of the laryngeal cavity region were collected, and a static three-dimensional structure of the laryngeal cavity was generated through anatomical segmentation, contour extraction, and three-dimensional reconstruction. Zero-point calibration of the transoral surgical robot is performed based on the static three-dimensional structure of the laryngeal cavity to establish RCM constraints. Under the RCM constraints, images of the surgical area are acquired in real time using a three-dimensional endoscope. Real-time local three-dimensional point clouds of the laryngeal cavity are generated through stereo vision reconstruction and geometrically registered with the static three-dimensional structure of the laryngeal cavity to determine the real-time end pose. The static three-dimensional structure of the laryngeal cavity is non-rigidly registered with the real-time local three-dimensional point cloud of the laryngeal cavity, and the natural motion law of key anatomical regions is extracted by combining the real-time end pose to generate the predicted end pose of the target. The preoperative planned path and predicted target end pose are solved by inverse kinematics and attitude constraint analysis to generate primary control variables. The primary control variables are then integrated with three types of real-time monitoring data to calculate a comprehensive risk index. The comprehensive risk index is used to adjust the autonomous control ratio of the robot to generate risk-guided control variables.

2. The control method for the transoral surgical robot as described in claim 1, characterized in that: The preoperative medical imaging data includes thin-section CT images and images of respiratory status.

3. The control method for the transoral surgical robot as described in claim 2, characterized in that: The process involves anatomical segmentation, contour extraction, and 3D reconstruction to generate a static 3D structure of the laryngeal cavity. The specific steps are as follows: The preoperative medical imaging data was denoised and normalized to obtain standard preoperative medical imaging data. The Otsu algorithm was used to perform anatomical segmentation on the standard preoperative medical imaging data to distinguish between soft and hard tissues in the laryngeal cavity and obtain laryngeal cavity segmentation images. The outer contour of the laryngeal cavity region is extracted from the laryngeal cavity segmentation image, and the internal structure of the laryngeal cavity is segmented from the laryngeal cavity segmentation image using a region growing algorithm; The outer contour of the laryngeal cavity region and the internal structure of the laryngeal cavity are converted into a uniform three-dimensional voxel mesh, and three-dimensional point cloud data are formed by three-dimensional rasterization. Each point cloud in the 3D point cloud data is connected to its adjacent point clouds to form a triangular mesh, generating a 3D surface structure of the laryngeal cavity. The static 3D structure of the laryngeal cavity is obtained through smoothing and hole filling.

4. The control method for the transoral surgical robot as described in claim 3, characterized in that: The process of performing zero-point calibration on the transoral surgical robot based on the static three-dimensional structure of the laryngeal cavity and establishing RCM constraints involves the following specific steps. From the static three-dimensional structure of the laryngeal cavity, a landmark structure of the laryngeal cavity region is selected as a reference point. The position of the robot's end effector is aligned with the reference point to determine the zero point of the robot's coordinate system. Based on the zero point position of the robot coordinate system, the coordinate transformation relationship between the robot coordinate system and the static three-dimensional structure of the laryngeal cavity is established through rigid transformation, and RCM constraints are set.

5. The control method for the transoral surgical robot as described in claim 4, characterized in that: Under RCM constraints, images of the surgical area are acquired in real time using a 3D endoscope, and a real-time local 3D point cloud of the laryngeal cavity is generated through stereoscopic vision reconstruction. The specific steps are as follows. Under RCM constraints, dual cameras in a 3D endoscope are used to acquire images of the surgical area from different angles in real time to obtain dual-view surgical area images; grayscale processing is performed on the dual-view surgical area images, corner feature points are extracted and image alignment is performed to obtain aligned surgical area images; The disparity values ​​between aligned surgical area images are calculated, the depth value of each pixel is estimated using the disparity values, and the depth value of each pixel is converted into three-dimensional coordinates using triangulation to generate a real-time local three-dimensional point cloud of the laryngeal cavity.

6. The control method for the transoral surgical robot as described in claim 1, characterized in that: The specific steps for determining the real-time end-effector pose are as follows. The nearest point between the real-time local 3D point cloud of the laryngeal cavity and the static 3D structure of the laryngeal cavity is matched and the distance error between the two is minimized to obtain the registered 3D point cloud; By registering the correspondence between the 3D point cloud and the static 3D structure of the throat cavity, the real-time end effector pose of the robot in the throat cavity region is calculated and determined.

7. The control method for the transoral surgical robot as described in claim 6, characterized in that: The process involves non-rigid registration of the static 3D laryngeal cavity structure with the real-time local 3D point cloud of the laryngeal cavity, and combining this with the extraction of natural motion patterns of key anatomical regions based on the real-time distal end pose to generate the predicted distal end pose. The specific steps are as follows: Under a unified reference coordinate system, the static three-dimensional structure of the laryngeal cavity is used as the reference three-dimensional structure. The real-time local three-dimensional point cloud of the laryngeal cavity is non-rigidly registered with the reference three-dimensional structure, and the overall deformation relationship is obtained by calculating the local deformation value. By analyzing the dynamic changes of the laryngeal cavity in relation to the real-time end-effector pose and overall deformation, the movement patterns of anatomical landmarks during the operation are determined, and the relative positional changes between the real-time end-effector pose and anatomical landmarks are calculated to obtain the natural movement patterns of key anatomical areas. By summarizing multiple relative position changes and natural motion laws, the predicted target end pose is generated.

8. The control method for the transoral surgical robot as described in claim 7, characterized in that: The process of generating primary control variables by solving inverse kinematics and attitude constraint analysis based on the preoperative planned path and predicted target end pose is as follows: By combining the preoperative planned path with the predicted target end-effector pose, the inverse kinematics algorithm is used to calculate the joint angles of the robot's end effector. The pose error between the current pose of the robot end effector and the predicted target end effector pose is calculated, and the change in the angle of each joint is adjusted based on the pose error. The joint angle of the robot end effector is corrected through a feedback mechanism to obtain the primary control quantity.

9. The control method for the transoral surgical robot as described in claim 8, characterized in that: The process involves integrating primary control parameters with three types of real-time monitoring data to calculate a comprehensive risk index. This comprehensive risk index is then used to adjust the robot's autonomous control ratio, generating risk-guided control parameters. The specific steps are as follows: Collect visual monitoring data, mechanical and morphological monitoring data, and safety monitoring data to form three types of real-time monitoring data; The primary control parameters are integrated with three types of real-time monitoring data to obtain a comprehensive control information set, and a comprehensive risk index is calculated based on the risk assessment weights. Based on the comprehensive risk index, the proportion of autonomous control of the robot is adjusted, and risk-guided control parameters are generated in combination with the control requirements of the surgical robot.

10. A control system for a transoral surgical robot, based on the control method for a transoral surgical robot according to any one of claims 1 to 9, characterized in that: include, The image acquisition and reconstruction module is used to acquire preoperative medical image data of the laryngeal cavity region and generate a static three-dimensional structure of the laryngeal cavity through anatomical segmentation, contour extraction and three-dimensional reconstruction. The calibration and constraint module is used to perform zero-point calibration on the transoral surgical robot based on the static three-dimensional structure of the laryngeal cavity and establish RCM constraints. Under the RCM constraints, the surgical area image is acquired in real time using a three-dimensional endoscope, and a real-time local three-dimensional point cloud of the laryngeal cavity is generated through stereo vision reconstruction. This point cloud is then geometrically registered with the static three-dimensional structure of the laryngeal cavity to determine the real-time end pose. The point cloud registration calculation module is used to perform non-rigid registration between the static three-dimensional structure of the laryngeal cavity and the real-time local three-dimensional point cloud of the laryngeal cavity, and to extract the natural motion law of key anatomical regions by combining the real-time end pose to generate the predicted end pose of the target. The risk control adjustment module is used to generate primary control variables by solving the preoperative planned path and the predicted target end pose through inverse kinematics and attitude constraint analysis. The primary control variables are then integrated with three types of real-time monitoring data to calculate a comprehensive risk index. The comprehensive risk index is used to adjust the autonomous control ratio of the robot and generate risk-guided control variables.