Control system and method of percutaneous surgical robot

By constructing the visual feature registration of the three-dimensional model of the throat and the 3D endoscopic stereoscopic image, combined with preoperative planning, joint control instructions are generated, the problems of navigation reference drift and trajectory offset in traditional systems are solved, and high-precision control of the transoral surgical robot is realized.

CN120241258AInactive Publication Date: 2025-07-04JILIN UNIVERSITY

Patent Information

Application Number
CN202510475249.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing oral surgical robot control system is difficult to achieve precise motion control of the multi-degree of freedom robot arm in the narrow oropharyngeal cavity. Due to the patient's position changes, soft tissue deformation and respiratory movement, it leads to a decrease in navigation reference drift and the spatial matching accuracy of the end effector of the robot arm and the target anatomical structure, and there is a risk of accidental injury to healthy tissue or lesions.

Method used

The three-dimensional model of the larynx is constructed by obtaining thin-layer CT scanning data covering the neck and throat areas, combining the surgical field stereoscopic images collected by 3D endoscopy to perform semantic-level dynamic registration of visual features, determining the position information of the robot, and combining the surgical resection path planned preoperatively, joint control instructions are generated to achieve real-time dynamic compensation and error correction.

Benefits of technology

It shortens the preparation time for intraoperative navigation, ensures the high consistency of the robotic arm movement trajectory and the lesion resection path, reduces the risk of trajectory deviation caused by registration deviation, and improves the accuracy and safety of the operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120241258A_ABST
    Figure CN120241258A_ABST
Patent Text Reader

Abstract

The invention provides a control system and method for an oral surgery robot, and relates to the field of intelligent control, and the method comprises the steps: firstly obtaining thin-layer CT scanning data covering neck and throat regions, constructing a throat three-dimensional model, then collecting a surgery view three-dimensional image through a 3D endoscope, and carrying out the 3D image collection; performing semantic-level dynamic registration based on visual features on the three-dimensional model of the throat, determining pose information of the oral surgical robot, outputting a next target pose point of the oral surgical robot in combination with the pose information and a pre-operative planned surgical resection path, and finally determining the target pose point of the oral surgical robot on the basis of the target pose point. And generating a joint control instruction of the oral surgical robot. Therefore, the navigation preparation time in the operation is shortened, the high consistency of the motion track of the mechanical arm and the focus resection path is ensured, and the track deviation risk caused by registration deviation of a traditional system can be fundamentally solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent control, and more specifically, to a control system and method for a transoral surgical robot. Background Art

[0002] The control technology of transoral surgical robots is a key link in promoting minimally invasive surgery towards higher precision and less trauma. In transoral surgeries such as laryngeal cancer and vocal cord lesions, the surgical area is located in the narrow oropharyngeal cavity, with complex anatomical structures and physiological movement interferences such as breathing and swallowing, which pose control requirements at the millimeter or even sub-millimeter level for the positioning accuracy and dynamic navigation of surgical instruments.

[0003] Currently, most transoral surgical robots used clinically adopt a master-slave teleoperation mode. The core challenge of its control system lies in how to achieve precise motion control of a multi-degree-of-freedom robotic arm in a narrow and enclosed space. Traditional control methods rely on preoperative static image navigation combined with optical or electromagnetic tracking technologies. However, in actual surgeries, dynamic factors such as patient body position changes, soft tissue deformation, and respiratory movements will cause navigation reference drift, significantly reducing the spatial matching accuracy between the end effector of the robotic arm and the target anatomical structure. Problems commonly existing in existing systems, such as insufficient joint clearance compensation and cumulative errors in simplified kinematic models, will cause repeated positioning errors of the robotic arm, and are extremely likely to cause accidental injury to healthy tissues or residual lesions during the resection of small laryngeal lesions. In addition, the monocular endoscope images used in traditional visual servo systems lack depth information and require artificial marking points for spatial registration during the registration process, which not only prolongs the surgical time but also increases the risk of robotic arm motion trajectory deviation due to registration errors.

[0004] Therefore, an optimized control solution for transoral surgical robots is desired. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a control system and method for a transoral surgical robot.

[0006] According to one aspect of the present application, there is provided a control system for an oral surgery robot, comprising: a CT scan data acquisition module for acquiring thin-slice CT scan data covering the neck and larynx regions; a larynx three-dimensional model construction module for constructing a three-dimensional model of the larynx based on the thin-slice CT scan data covering the neck and larynx regions; a stereoscopic image acquisition module for acquiring a stereoscopic image of the surgical field using a 3D endoscope; a pose information determination module for determining the pose information of the oral surgery robot based on the visual registration between the stereoscopic image of the surgical field and the three-dimensional model of the larynx, wherein the pose information determination module is configured to: perform semantic-level dynamic registration based on visual features on the stereoscopic image of the surgical field and the three-dimensional model of the larynx to obtain the pose information of the oral surgery robot; a next target pose point determination module for outputting the next target pose point of the oral surgery robot based on the pose information of the oral surgery robot and the preoperatively planned surgical resection path; and a joint control module for generating a joint control instruction for the oral surgery robot based on the next target pose point of the oral surgery robot.

[0007] According to another aspect of the present application, there is provided a control method for an oral surgery robot, comprising: acquiring thin-slice CT scan data covering the neck and larynx regions; constructing a three-dimensional model of the larynx based on the thin-slice CT scan data covering the neck and larynx regions; acquiring a stereoscopic image of the surgical field using a 3D endoscope; determining the pose information of the oral surgery robot based on the visual registration between the stereoscopic image of the surgical field and the three-dimensional model of the larynx, including: performing semantic-level dynamic registration based on visual features on the stereoscopic image of the surgical field and the three-dimensional model of the larynx to obtain the pose information of the oral surgery robot; outputting the next target pose point of the oral surgery robot based on the pose information of the oral surgery robot and the preoperatively planned surgical resection path; and generating a joint control instruction for the oral surgery robot based on the next target pose point of the oral surgery robot.

[0008] Compared with the prior art, the control system and method for an oral surgery robot provided by the present application first acquire thin-slice CT scan data covering the neck and larynx regions and construct a three-dimensional model of the larynx, then acquire a stereoscopic image of the surgical field using a 3D endoscope, perform semantic-level dynamic registration based on visual features on it and the three-dimensional model of the larynx to determine the pose information of the oral surgery robot, then combine this pose information with the preoperatively planned surgical resection path to output the next target pose point of the oral surgery robot, and finally generate a joint control instruction for the oral surgery robot based on this target pose point. In this way, it helps to shorten the intraoperative navigation preparation time and ensure a high degree of consistency between the movement trajectory of the robotic arm and the lesion resection path, and can fundamentally solve the risk of trajectory deviation caused by registration deviation in the traditional system. Description of the Drawings

[0009] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0010] In the accompanying drawings: Figure 1 It is a system block diagram of a control system for an oral surgery robot according to an embodiment of the present application.

[0011] Figure 2 It is a block diagram of a pose information determination module in a control system for an oral surgery robot according to an embodiment of the present application.

[0012] Figure 3 It is a block diagram of a visual perception registration and coding unit in a control system for an oral surgery robot according to an embodiment of the present application.

[0013] Figure 4 It is a flowchart of a control method for an oral surgery robot according to an embodiment of the present application. Detailed implementation manners

[0014] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0015] The control technology of oral surgery robots is the key to promoting the development of minimally invasive surgery with high precision and less trauma. In oral surgeries such as laryngeal cancer and vocal cord lesions, since the surgical area is located in the narrow oropharyngeal cavity, the anatomical structure is complex and affected by physiological movements, there are extremely high control requirements for the positioning accuracy of surgical instruments and dynamic navigation.

[0016] Currently, most clinical oral surgery robots adopt a master-slave teleoperation mode. The core challenge of its control system is to achieve precise motion control of a multi-degree-of-freedom robotic arm in a narrow and enclosed space. Traditional control methods rely on preoperative static image navigation and tracking technology, but dynamic factors such as patient body position changes will cause the navigation reference to drift, reducing the spatial matching accuracy between the end effector of the robotic arm and the target anatomical structure. At the same time, existing systems have problems such as insufficient joint clearance compensation, which causes repeated positioning errors of the robotic arm and is likely to cause accidental injury to healthy tissues or residual lesions. In addition, the monocular endoscope images of traditional visual servo systems lack depth information, and the registration relies on artificial marking points, which not only prolongs the operation time but also increases the risk of deviation of the robotic arm motion trajectory.

[0017] Based on this, the present application proposes a control system for an oral surgery robot.Figure 1 It is a system block diagram of a control system for an oral surgery robot according to an embodiment of the present application. As Figure 1 shown, in the control system 100 of the oral surgery robot, it includes: a CT scan data acquisition module 110, configured to acquire thin-layer CT scan data covering the neck and laryngeal regions; a laryngeal three-dimensional model construction module 120, configured to construct a laryngeal three-dimensional model based on the thin-layer CT scan data covering the neck and laryngeal regions; a stereoscopic image acquisition module 130, configured to acquire a stereoscopic image of the surgical field of view using a 3D endoscope; a pose information determination module 140, configured to determine the pose information of the oral surgery robot based on the visual registration between the stereoscopic image of the surgical field of view and the laryngeal three-dimensional model; a next target pose point determination module 150, configured to output the next target pose point of the oral surgery robot based on the pose information of the oral surgery robot and the pre-operative planned surgical resection path; and a joint control module 160, configured to generate a joint control instruction for the oral surgery robot based on the next target pose point of the oral surgery robot.

[0018] Specifically, this technical solution effectively addresses the core challenges faced by traditional oral surgery robots, such as navigation reference drift, motion error accumulation, and visual registration defects, by integrating a multi-modal data fusion and real-time dynamic compensation mechanism. For the navigation reference drift problem caused by intraoperative soft tissue deformation and respiratory movement, the system establishes a dynamically updated spatial mapping relationship through the deep registration of the laryngeal three-dimensional anatomical model constructed by high-resolution thin-layer CT and the real-time stereoscopic image of the 3D endoscope, enabling the end effector of the robotic arm to perceive the tissue deformation amount in real time and automatically correct the pose offset. Compared with the monocular vision system that relies on artificial marker points, the binocular disparity data provided by the stereoscopic image acquisition module not only realizes markerless depth perception, but also shortens the intraoperative registration time through the automatic registration of the three-dimensional model and the real-time stereoscopic image, reducing the risk of trajectory deviation caused by manual intervention. At the motion control level, the next target pose point determination module constructs a closed-loop control model with error prediction and compensation functions by fusing the pre-operative planned path and real-time pose feedback data, and can dynamically correct the cumulative error caused by the joint clearance and kinematic model simplification of the robotic arm. When the joint control module performs inverse kinematics calculation based on the dynamically generated pose points, by introducing a joint clearance compensation algorithm, it significantly suppresses the error amplification effect of the serial robotic arm in multi-degree-of-freedom motion. This intelligent real-time registration technology not only shortens the intraoperative navigation preparation time, but also ensures a high degree of consistency between the motion trajectory of the robotic arm and the lesion resection path, and can fundamentally solve the risk of trajectory deviation caused by registration deviation in traditional systems.

[0019] In the embodiment of the present application, the CT scan data acquisition module 110 is configured to acquire thin-slice CT scan data covering the neck and larynx regions. It should be understood that the anatomical structures of the neck and larynx are complex, containing numerous important organs, blood vessels, nerves and other tissues, and during transoral surgery, the surgical area is located within the narrow oropharyngeal cavity, and the surrounding structures are closely adjacent. Through thin-slice CT scanning, high-resolution image data can be obtained, which can clearly show the various tissue layers, lesion sites and their relationships with the surrounding structures in the neck and larynx. These detailed information helps to construct an accurate three-dimensional model of the larynx, and then can better realize subsequent surgical planning and precise control of the robot.

[0020] In the embodiment of the present application, the laryngeal three-dimensional model construction module 120 is used to construct a laryngeal three-dimensional model based on the thin-slice CT scan data covering the neck and laryngeal regions. Correspondingly, considering that the target area of transoral surgery is in the narrow oropharyngeal cavity with complex anatomical structures, traditional imaging data is difficult to comprehensively and accurately display the fine structures of the larynx. The thin-slice CT scan data covering the neck and laryngeal regions can provide high-resolution tomographic images, which contain rich anatomical details. By using these data to construct a laryngeal three-dimensional model, the two-dimensional tomographic images can be converted into intuitive three-dimensional structures, enabling doctors to more clearly and comprehensively understand the morphology, position of the larynx, and the spatial relationships among various tissues. At the same time, during the surgical process, there are physiological movements such as breathing and swallowing, as well as dynamic factors such as patient position changes and soft tissue deformation, which can cause navigation fiducial drift. The laryngeal three-dimensional model can be dynamically updated and registered in combination with the real-time surgical situation, helping the transoral surgical robot to adjust its pose in real time and overcome the limitations of traditional control methods that rely on preoperative static image navigation. Specifically, in a specific example of the present application, constructing a laryngeal three-dimensional model based on the thin-slice CT scan data covering the neck and laryngeal regions can be achieved in the following manner: First, import the thin-slice CT scan data in DICOM format into three-dimensional reconstruction software, unify the gray-scale range of images in different layers through gray-scale normalization processing, and use the median filtering algorithm to suppress the noise of the images, removing high-frequency noise generated during the scanning process due to equipment vibration or patient breathing to ensure the uniformity of image gray-scale values. Then perform image segmentation to extract larynx-related tissues. Use the threshold segmentation algorithm combined with manual interaction to delimit the initial region of interest, separate background tissues such as neck skin, muscles, and bones, and use the region growing algorithm combined with morphological operations (such as dilation and erosion) for fine segmentation of structures such as vocal cords, arytenoid cartilages, and cricoid cartilages, and identify the boundaries of soft tissues such as the mucosal layer through the gradient-based edge detection algorithm to establish two-dimensional masks for each anatomical structure. Then perform interlayer interpolation on the two-dimensional mask sequence along the Z-axis direction, use the cubic spline interpolation algorithm to fill in the detailed information between adjacent layers, construct continuous three-dimensional volume data, and use the marching cubes algorithm to perform surface reconstruction on the volume data to generate an initial three-dimensional model represented by triangular mesh surfaces. This algorithm traverses the eight vertices of each voxel and determines whether the vertex belongs to the surface of the target tissue according to a preset threshold to generate corresponding triangular patches. After that, perform geometric optimization on the initial three-dimensional model, use the Laplacian smoothing algorithm to reduce the sharp edges and corners on the mesh surface, automatically detect and fill local missing regions caused by CT scan slice thickness or segmentation errors through the hole repair algorithm, and use the non-uniform rational B-spline surface fitting technology to smooth the model surface to ensure the geometric continuity and surface smoothness of each anatomical structure.Finally, coordinate system registration is performed. Taking the anterior median line of the patient as the Y-axis and the horizontal cross-section as the XY plane, a patient-specific coordinate system is established. The anatomical landmark points (such as the upper edge of the thyroid cartilage and the midpoint of the cricoid cartilage arch) of the three-dimensional model are aligned with this coordinate system to generate a three-dimensional model containing the complete anatomical structures of the neck and larynx, providing an accurate spatial reference for subsequent surgical planning and visual registration.

[0021] In the embodiment of the present application, the stereoscopic image acquisition module 130 is used to acquire stereoscopic images of the surgical field by using a 3D endoscope. It should be understood that traditional monocular endoscopes can only provide two-dimensional planar images, lacking depth information, which makes it difficult for surgeons to accurately judge the spatial distance between the instrument and the lesion in the narrow oropharyngeal cavity. However, 3D endoscopes can obtain stereoscopic images with depth information, thus completely restoring the three-dimensional structure of the surgical scene. Specifically, the stereoscopic images acquired by the 3D endoscope can provide rich spatial information, including the depth, shape, position of the tissues in the surgical area, and their relative relationships, etc. These information helps to accurately judge the situation of the surgical site. That is, through the stereoscopic images acquired by the 3D endoscope, the model can more clearly understand the fine structure and lesion characteristics of the surgical area, such as the specific shape, boundary of laryngeal lesions, and the demarcation from the surrounding normal tissues, etc., and can provide real-time environmental visual feedback for robot control.

[0022] In the embodiment of the present application, the pose information determination module 140 is used to determine the pose information of the transoral surgical robot based on the visual registration between the stereoscopic images of the surgical field and the three-dimensional model of the larynx. It should be understood that the stereoscopic images of the surgical field provide real-time visual information of the surgical site, but lack the overall anatomical reference; while the three-dimensional model of the larynx shows detailed anatomical structures, but lacks real-time dynamic information. By performing visual registration between the two, the real-time surgical scene can be combined with the accurate anatomical model to achieve the effective fusion of multi-modal visual information, providing a more comprehensive and accurate basis for determining the robot pose. However, existing registration algorithms mostly use shallow feature matching based on edge features or surface curvature, which are difficult to capture biological specific features such as vocal cord mucosal folds and blood vessel textures. When encountering non-rigid deformations caused by respiratory movements, the registration error will be transmitted to the end of the robotic arm through the kinematic chain, resulting in a spatial positioning deviation of millimeters.

[0023] Based on this, in the process of determining the pose information of the transoral surgical robot through visual registration between the stereoscopic image of the surgical field and the three-dimensional model of the larynx, the technical concept of this application is to achieve deep semantic fusion of multi-modal visual data through the collaborative processing of three-dimensional hollow convolution coding and feature dynamic alignment. First, multi-scale feature extraction is performed on the stereoscopic image of the surgical field and the three-dimensional model of the larynx respectively. While retaining biological feature details such as mucosal folds and blood vessel branches, the receptive field is expanded through dilated convolution to capture the spatial topological relationship of the anatomical structure. Subsequently, a visual feature semantic registration mechanism is introduced to establish a cross-modal feature similarity field at the semantic level, effectively overcoming the interference of tissue displacement caused by respiratory movement and achieving elastic registration under non-rigid deformation. Finally, the registered feature map is converted into the pose parameters of the robotic arm joint space. This processing method breaks through the limitations of traditional shallow feature matching, not only achieving precise capture of deep biological features, but also being able to adaptively compensate for the elastic deformation of tissues during surgery, enabling the spatial positioning of the end effector of the robotic arm to always maintain an accurate correspondence with the dynamic anatomical structure, significantly reducing the transmission and accumulation of motion errors in the robotic arm kinematic chain, and providing reliable real-time navigation guarantee for microsurgical operations in narrow cavities.

[0024] Specifically, in the embodiment of this application, the pose information determination module is used to: perform semantic-level dynamic registration based on visual features on the stereoscopic image of the surgical field and the three-dimensional model of the larynx to obtain the pose information of the transoral surgical robot. More specifically, Figure 2 FIG. is a block diagram of the pose information determination module in the control system of the transoral surgical robot according to the embodiment of this application. As Figure 2 shown, the pose information determination module 140 includes: a visual feature extraction unit 141, which is used to perform visual feature extraction on the stereoscopic image of the surgical field and the three-dimensional model of the larynx to obtain a visual feature encoded map of the surgical field and a visual feature encoded map of the three-dimensional model of the larynx; a visual perception registration coding unit 142, which is used to perform visual feature semantic-level registration based on feature structure guidance on the visual feature encoded map of the surgical field and the visual feature encoded map of the three-dimensional model of the larynx to obtain a visual perception registration encoded map; a pose information generation unit 143, which is used to obtain the pose information of the transoral surgical robot based on the visual perception registration encoded map.

[0025] In an embodiment of the present application, the visual feature extraction unit 141 is configured to perform visual feature extraction on the surgical field stereo image and the three-dimensional model of the larynx to obtain a visual feature encoded map of the surgical field and a visual feature encoded map of the three-dimensional model of the larynx. Specifically, in an embodiment of the present application, the visual feature extraction unit is configured to: perform visual feature extraction based on three-dimensional dilated convolution encoding on the surgical field stereo image and the three-dimensional model of the larynx to obtain the visual feature encoded map of the surgical field and the visual feature encoded map of the three-dimensional model of the larynx. Correspondingly, considering that the complex anatomical structure in the laryngeal cavity has multi-scale characteristics, including both micron-scale wrinkle textures on the surface of the vocal cord mucosa and centimeter-scale spatial topological relationships between the thyroid cartilage and the cricoid cartilage. Existing shallow feature extraction methods based on edge detection or curvature calculation are difficult to effectively represent the spatial topological relationships of these microscopic anatomical markers. Especially when non-uniform deformation occurs on the mucosal surface due to respiratory movement, the rigid features extracted by traditional methods are prone to feature point drift or mis-matching, resulting in a continuous shift of the registration benchmark. Based on this, the present application performs visual feature extraction based on three-dimensional dilated convolution encoding on the surgical field stereo image and the three-dimensional model of the larynx to obtain a visual feature encoded map of the surgical field and a visual feature encoded map of the three-dimensional model of the larynx. Specifically, for the surgical field stereo image, the three-dimensional dilated convolution synchronously extracts the microscopic features of the mucosal texture (such as the direction of the folds and the bifurcation points of capillaries) and the spatial distribution features of macroscopic anatomical markers (such as the vocal process and the arytenoid cartilage) through convolution kernels with different dilation rates. Its dilated sampling characteristic can expand the feature perception range without losing spatial resolution. For the three-dimensional model of the larynx, the three-dimensional convolution kernel performs multi-level feature abstraction along the body data axis, which can not only retain the rigid contour features of the cartilage structure but also capture the elastic deformation trend of the mucosal layer and the muscular layer. This dual-channel feature encoding mechanism constructs a feature expression containing biological specific features and anatomical structure associations in a unified three-dimensional space coordinate system, laying an anatomical feature foundation for subsequent cross-modal registration.

[0026] In an embodiment of the present application, the visual perception registration encoding unit 142 is configured to perform semantic-level registration of visual features based on feature structure guidance on the visual feature encoded map of the surgical field and the visual feature encoded map of the three-dimensional model of the larynx to obtain a visual perception registration encoded map. Specifically, Figure 3 It is a block diagram of the visual perception registration encoding unit in the control system of the transoral surgical robot according to an embodiment of the present application. As Figure 3As shown in the figure, the visual perception registration encoding unit 142 includes: a local feature structural similarity search and alignment subunit 1421, which is used to perform local feature structural similarity search and alignment on the visual feature encoding map of the surgical field and the visual feature encoding map of the three-dimensional model of the larynx to obtain a set of phase-aligned {local detail encoding matrix of visual features of the surgical field, local encoding matrix of visual features of the three-dimensional model of the larynx} feature pairs; an attention weight calculation subunit 1422, which is used to perform feature joint dynamic attention perception on each phase-aligned {local detail encoding matrix of visual features of the surgical field, local encoding matrix of visual features of the three-dimensional model of the larynx} feature pair in the set of phase-aligned {local detail encoding matrix of visual features of the surgical field, local encoding matrix of visual features of the three-dimensional model of the larynx} feature pairs to obtain a set of visual perception registration joint perception attention weights; a visual perception registration encoding feature generation subunit 1423, which is used to perform gated significant aggregation on the set of phase-aligned {local detail encoding matrix of visual features of the surgical field, local encoding matrix of visual features of the three-dimensional model of the larynx} feature pairs based on the set of visual perception registration joint perception attention weights to obtain the visual perception registration encoding map.

[0027] It should be understood that the visual feature encoding map of the surgical field and the visual feature encoding map of the three-dimensional model of the larynx extract features from different modal data respectively, but these features need to be fused at a higher level to achieve the effective combination of the two modal information. Traditional registration methods rely on manually designed edge features or geometric contour matching. Such shallow features are prone to feature drift in complex scenarios such as dynamic changes in mucosal folds and local occlusion of vascular textures. Especially in the periodic deformation caused by respiratory movement, the failure of the rigid registration assumption will lead to a systematic misalignment of the feature space mapping relationship, resulting in an insurmountable semantic gap between the soft tissue morphology captured in real time in the stereoscopic image and the anatomical benchmark of the three-dimensional model. More critically, the misalignment of cross-modal features (the light field information of the stereoscopic image and the voxel data of the three-dimensional model) in the spatial expression dimension and semantic abstraction level makes it difficult for traditional Euclidean distance-based similarity metrics to accurately capture the essential association of biological features. Based on this, the present application performs feature structure-guided visual feature semantic-level registration on the visual feature encoding map of the surgical field and the visual feature encoding map of the three-dimensional model of the larynx to obtain the visual perception registration encoding map.

[0028] Specifically, first, the feature map generated by three-dimensional hollow convolution encoding is decoupled, and the global features are decomposed into a set of feature encoding matrices containing local biological characteristics. This decoupling is not a simple spatial segmentation, but through learnable feature channel recombination, encapsulating biological specific information such as mucosal microstructures and vascular branching patterns into discrete semantic units. Subsequently, based on the feature phase similarity metric, matching pairs are dynamically searched in the decoupled feature set. The phase alignment degree is evaluated by analyzing the distribution similarity of feature vectors, which can penetrate apparent interferences such as illumination changes and tissue fluid specular reflection and reach the essential structural consistency of biological features. This dynamic search mechanism breaks through the limitation of fixed spatial correspondence relationships. When the laryngeal tissue undergoes elastic deformation, a cross-temporal and cross-spatial semantic mapping can still be established through the inherent correlation of feature phases. In this way, by registering at the semantic level, an accurate correspondence relationship can be established between the features of the surgical field of view and the three-dimensional laryngeal model, forming a unified semantic space, enabling effective information exchange and fusion of data in the two modalities in this space.

[0029] Specifically, in the embodiment of the present application, the local feature structure similarity search and alignment subunit is used to: perform feature decoupling on the visual feature encoding map of the surgical field of view and the visual feature encoding map of the three-dimensional laryngeal model to obtain a set of local detail encoding matrices of the visual features of the surgical field of view and a set of local encoding matrices of the visual features of the three-dimensional laryngeal model. This process can be represented by the formula: ; ; where, is the visual feature encoding map of the surgical field of view, is the visual feature encoding map of the three-dimensional laryngeal model, is for performing the feature decoupling operation, , , and are respectively the first, the second, the th and the th local detail encoding matrices of the visual features of the surgical field of view in the set of local detail encoding matrices of the visual features of the surgical field of view, , , and are respectively the first, the second, the th and the th local encoding matrices of the visual features of the three-dimensional laryngeal model in the set of local encoding matrices of the visual features of the three-dimensional laryngeal model; Based on the feature phase alignment degree between any two surgical field visual feature local detail coding matrices and laryngeal three-dimensional model visual feature local coding matrices in the set of surgical field visual feature local detail coding matrices and the set of laryngeal three-dimensional model visual feature local coding matrices, perform feature structure similarity search alignment on the set of surgical field visual feature local detail coding matrices and the set of laryngeal three-dimensional model visual feature local coding matrices to obtain the set of {surgical field visual feature local detail coding matrix, laryngeal three-dimensional model visual feature local coding matrix} feature pairs with phase alignment. This process can be expressed by the formula: ; ; where, is the transpose operation, is to calculate the Frobenius norm, is and the feature phase alignment degree between them, is to return the value corresponding to the maximum value, is the position to find the maximum approximate matching value in the set of laryngeal three-dimensional model visual feature local coding matrices.

[0030] It should be understood that traditional methods rely on global features or shallow handcrafted features and are difficult to handle complex scenarios such as dynamic deformation of laryngeal mucosa and blood vessel occlusion. It is easy to cause feature drift and semantic gap due to the failure of the rigid registration assumption. Through feature decoupling, the surgical field visual feature coding map and the laryngeal three-dimensional model visual feature coding map are decomposed into a set of semantic units that focus on local biological characteristics (such as mucosal microstructures and blood vessel branches), effectively solving the problems of computational redundancy and semantic confusion of high-dimensional global features. This process is not a simple spatial segmentation, but through learnable channel recombination, the anatomical specific information is encapsulated into discrete local coding matrices, enhancing the adaptability to elastic deformation. That is, the set of surgical field visual feature local detail coding matrices and the set of laryngeal three-dimensional model visual feature local coding matrices obtained by feature decoupling achieve semantic dimensionality reduction from global coarse-grained to local fine-grained, can provide local feature primitives with clear structure and independent operation for subsequent phase alignment, and at the same time reduce the sensitivity to spatial misalignment.

[0031] Accordingly, due to the misalignment of spatial dimensions and semantic levels between cross-modal features (light field images and voxel models), traditional Euclidean distances cannot capture their essential correlations. By measuring the distribution similarity between the local detail encoding matrix of the visual features in the surgical field of view and the local encoding matrix of the visual features of the three-dimensional model of the larynx through phase alignment, the apparent interferences such as tissue specular reflection and illumination changes can be penetrated, and the structural consistency of biometric features can be directly evaluated. Specifically, the dynamic search mechanism abandons the rigid constraints of fixed spatial mapping, and adaptively matches the feature pairs with the highest phase alignment degree in the set of local detail encoding matrices of the visual features in the decoupled surgical field of view and the set of local encoding matrices of the visual features of the three-dimensional model of the larynx, so as to establish cross-temporal and cross-spatial semantic correlations even when the laryngeal tissue undergoes periodic deformations. That is to say, this step breaks through the geometric contour dependence of traditional registration, establishes cross-modal mapping through the internal consistency of feature phases, significantly reduces systematic misalignments caused by elastic deformations, and lays a foundation for constructing a unified semantic space.

[0032] Specifically, in the embodiment of the present application, the attention weight calculation sub-unit includes: a secondary visual perception registration trace metric value calculation sub-unit, configured to perform matrix trace metric on each of the phase-aligned {local detail encoding matrix of visual features in the surgical field of view, local encoding matrix of visual features of the three-dimensional model of the larynx} feature pairs to obtain a set of visual perception registration trace metric values; a secondary visual perception registration joint perception attention weight calculation sub-unit, configured to perform normalization processing based on the sigmoid function on the set of visual perception registration trace metric values to obtain the set of visual perception registration joint perception attention weights.

[0033] It should be understood that although phase alignment establishes a preliminary correspondence relationship, the importance and interaction patterns of different feature pairs need to be further modeled. Traditional weight assignment methods cannot adapt to complex biometric scenarios, while feature joint dynamic attention perception processing can dynamically learn the interaction relationships between feature pairs. Specifically, the network not only assigns weights, but also analyzes the semantic contribution intensity of feature pairs (such as the topological matching degree between blood vessel branches and model voxels), and suppresses the noise introduced by local occlusion or deformation. That is to say, this step endows the model with the ability to selectively focus on key biomarkers (such as vocal cord edges, lesion areas), enhances the discriminability of feature expressions, and at the same time improves the interpretability of the model through weight visualization, providing an importance prior for subsequent aggregation.

[0034] More specifically, in the embodiment of the present application, the secondary visual perception registration trace metric value calculation sub-unit is configured to: perform matrix trace metric on each of the phase-aligned {local detail encoding matrix of visual features in the surgical field of view, local encoding matrix of visual features of the three-dimensional model of the larynx} feature pairs to obtain a set of visual perception registration trace metric values, which can be expressed by the following formula: ; Where, is the trace metric value of the matrix, is and the visual perception registration trace metric value between.

[0035] More specifically, in the embodiment of the present application, the visual perception registration joint perception attention weight calculation secondary subunit is used to: perform normalization processing based on the sigmoid function on the set of the visual perception registration trace metric values to obtain the set of the visual perception registration joint perception attention weights, which can be expressed by the following formula: ; wherein, is the normalization function, is and the visual perception registration joint perception attention weight between.

[0036] Preferably, in another example of the present application, the visual perception registration joint perception attention weight calculation secondary subunit is used to: perform manifold optimization based on the bidirectional connectivity cardinality on the set of the visual perception registration trace metric values based on the set of the phase alignment {surgical field visual feature local detail coding matrix, laryngeal three-dimensional model visual feature local coding matrix} feature pairs to obtain the set of the optimized visual perception registration trace metric values; perform normalization processing based on the sigmoid function on the set of the optimized visual perception registration trace metric values to obtain the set of the visual perception registration joint perception attention weights. This process is expressed by the formula: ; ; ; ; ; ; ; ; wherein, is the transpose matrix of, is i.e., the corresponding phase alignment laryngeal three-dimensional model visual feature local coding matrix, is the eigenvalue at the th position in, is the eigenvalue at the The eigenvalue of a position, and , , is to calculate and the distance between is to calculate and the distance between is a preset threshold value, is to calculate the number of satisfied conditions, is the distance connectivity cardinality, is the distance connectivity cardinality, is the exponential function value with the natural constant as the base, is dot product by position, is the optimized matrix, is the optimized matrix, is the optimized visual perception registration trace metric value, is and the visual perception registration joint perception attention weight between.

[0037] Specifically, here, for each and for which matrix trace metric operations are performed, let and , and , , by calculating the number of eigenvalues of the and the distance and distance the matrix and the two-way connectivity cardinality between and , that is, the topological steady state measure between the local detail coding matrix of the visual features of the surgical field for phase alignment and the local coding matrix of the visual features of the three-dimensional model of the larynx.

[0038] Then, in the discrete manifold representations of the phase-aligned surgical field visual feature local detail encoding matrix and the phase-aligned laryngeal three-dimensional model visual feature local encoding matrix, the trace metric operation of the phase-aligned surgical field visual feature local detail encoding matrix and the phase-aligned laryngeal three-dimensional model visual feature local encoding matrix is embedded and optimized through manifold compactification driven by the bidirectional connectivity number, so as to improve the calculation accuracy of the joint matrix trace metric on the basis of enhancing the robustness of the far-distance correlation connectivity. That is, based on the initial and calculated , and In the case of, substitute , , , the initial and for optimization, and then calculate The optimized trace metric.

[0039] Specifically, in the embodiment of the present application, the visual perception registration coding feature generation subunit is used to: based on the set of visual perception registration joint perception attention weights, perform gated significant aggregation on the set of {surgical field visual feature local detail encoding matrix, laryngeal three-dimensional model visual feature local encoding matrix} feature pairs after phase alignment to obtain the visual perception registration coding map, which can be expressed by the following formula: ; ; wherein, and are the first joint perception weight matrix and the second joint perception weight matrix, is subtraction by position points, is addition by position points, is and The joint visual perception registration feature joint perception matrix after combination, that is, the th visual perception registration feature joint perception matrix in the set of visual perception registration feature joint perception matrices, , and are respectively the 1st, 2nd and th visual perception registration feature joint perception matrices in the set of visual perception registration feature joint perception matrices, is the visual perception registration coding map.

[0040] It should be understood that traditional aggregation methods (such as averaging or splicing) are prone to weakening the detailed differences of key biological structures (such as mucosal fold curvature, vascular topology) during cross-modal feature fusion. Especially in the scenarios of dynamic deformation or local occlusion, noise interference will contaminate the fusion results. Based on the gating mechanism of feature joint attention weights, the high-confidence biological consistency regions (such as the vocal cord edge features that remain stable during deformation) in the phase-aligned feature pairs are highlighted through non-linear modulation, while suppressing the low-confidence responses caused by respiratory movement or tissue specular reflection. That is to say, gated significant aggregation uses the attention weight as a "biological semantic filter" for feature contribution, while retaining the complementary advantages of the light field image details (such as mucosal surface optical flow information) and the three-dimensional model voxel structure (such as the spatial topology of blood vessel branches), eliminating the redundancy and conflict between cross-modal features. That is, the finally generated visual perception registration encoded map realizes the balance of information density and robustness of feature expression through dynamic weighting of the gating function, significantly narrowing the semantic gap between the intraoperative real-time tissue morphology and the pre-built model, and providing a high-fidelity and interpretable fusion feature benchmark for surgical navigation.

[0041] In the embodiment of the present application, the pose information generation unit 143 is configured to obtain the pose information of the transoral surgical robot based on the visually perceived registration encoded map. Specifically, in the embodiment of the present application, the pose information generation unit is configured to: input the visually perceived registration encoded map into a pose information calculator based on a decoder to obtain the pose information of the transoral surgical robot. It should be understood that the visually perceived registration encoded map is obtained by performing visual feature semantic-level registration on the visual feature encoded map of the surgical field of view and the visual feature encoded map of the three-dimensional model of the larynx. It integrates the real-time information of the surgical site and the three-dimensional model information of the laryngeal anatomical structure. However, this information exists in the form of an encoded map and cannot be directly used to control the pose of the transoral surgical robot. It needs to be further processed and transformed through a special decoder and pose information calculator to be transformed into pose information that the robot can understand and use. The pose information calculator based on the decoder can convert this data in the form of an encoded map into specific parameters related to the robot pose, such as position, attitude, etc., to achieve the mapping from the image feature space to the robot motion space. By inputting the visually perceived registration encoded map into the pose information calculator based on the decoder, through a series of calculations and processes, the accurate pose information of the transoral surgical robot can be extracted from the encoded map, including parameters such as the position coordinates of the robot in the surgical space and the attitude angles of each joint. These pose information can accurately reflect the actual state of the robot in the surgical area and provide an accurate basis for subsequent motion control. Specifically, the pose information completely describes the spatial state of the robot end effector in the current surgical space, specifically including six degrees of freedom kinematic parameters: three translational components and three rotational components. The translational components accurately calibrate the three-dimensional coordinate position (X, Y, Z axes) of the end of the robotic arm in the patient coordinate system, and their values directly reflect the spatial distance between the tip of the surgical instrument and the target anatomical structure (such as the edge of the vocal cord lesion). These data are usually presented with millimeter-level accuracy. For example, in a laryngeal cancer resection surgery, the positioning error is required to be no more than 0.3 millimeters to ensure that the instrument can accurately reach the predetermined resection boundary. The rotational components, in the form of Euler angles, define the spatial orientation of the end effector. This includes the pitch angle, yaw angle, and roll angle of the instrument, and these angle parameters determine the cutting direction of the contact surface between the surgical instrument (such as a laser scalpel or pliers) and the tissue. Through the accurate acquisition of the pose information, the transoral surgical robot can achieve precise positioning and operation in a complex surgical environment, improving the success rate and safety of the surgery.

[0042] In summary, the pose information determination module 140 is clearly described. First, it extracts multi-scale features from the stereoscopic image of the surgical field and the three-dimensional model of the larynx respectively. While retaining the biological feature details such as mucosal folds and blood vessel branches, it expands the receptive field through dilated convolution to capture the spatial topological relationship of the anatomical structure. Subsequently, a visual feature semantic registration mechanism is introduced to establish a cross-modal feature similarity field at the semantic level, effectively overcoming the interference of tissue displacement caused by respiratory movement and realizing elastic registration under non-rigid deformation. Finally, the registered feature map is converted into the pose parameters of the robotic arm joint space. In this way, it effectively breaks through the limitations of traditional shallow feature matching, not only realizes the accurate capture of deep biological features, but also can adaptively compensate for the elastic deformation of tissues during the operation, enabling the spatial positioning of the end effector of the robotic arm to always maintain an accurate correspondence with the dynamic anatomical structure, significantly reducing the transmission and accumulation of motion errors in the robotic arm kinematic chain, and providing a reliable real-time navigation guarantee for microsurgical operations in narrow cavities.

[0043] In the embodiment of the present application, the next target pose point determination module 150 is configured to output the next target pose point of the transoral surgical robot based on the pose information of the transoral surgical robot and the preoperatively planned surgical resection path. It should be understood that the pose information of the transoral surgical robot reflects its current actual position and posture in the surgical space, which is real-time status information. The preoperatively planned surgical resection path is an ideal operation path formulated based on the patient's preoperative examination data and the doctor's surgical plan. Based on the combination of the two, the robot can consider both the current actual position and follow the predetermined surgical plan when performing the operation, ensuring that the operation proceeds as expected. The obtained output of the next target pose point of the transoral surgical robot can provide a clear direction and target for the movement of the robot. This target pose point is determined based on a comprehensive consideration of the current pose and the surgical resection path, enabling the robot to gradually approach the target lesion along the predetermined surgical path and ensuring the orderly progress of the surgical operation.

[0044] The following is a detailed description of a specific implementation process of "outputting the next target pose point of the transoral surgical robot based on the pose information of the transoral surgical robot and the preoperatively planned surgical resection path": First, the pre-operative planned surgical resection path is the basis for the entire operation. It is generated based on the patient's thin-slice CT scan data, the three-dimensional reconstruction results of the lesion, and the clinical surgical plan. During the planning, doctors use medical image processing software, such as Mimics, 3D Slicer, etc., to carefully segment the CT data, thereby accurately extracting the three-dimensional spatial boundaries of the lesion. This boundary is presented in the form of a triangular mesh model or a polygonal surface, clearly marking the boundary coordinates of the tumor in all directions, including the head-tail end, left-right sides, and deep-shallow layers. At the same time, according to oncology principles, in order to ensure complete resection of the lesion and prevent recurrence, a certain safety distance is extended outside the lesion boundary, usually 5-10 mm, and the specific value is determined according to the pathological type of the tumor. This extended area forms a three-dimensional resection area containing a safety margin, and its surface is the motion target surface of the end effector of the robotic arm. In addition, the pre-operative planned path also includes key anatomical structure avoidance points. In the larynx, structures such as the vocal cords, the area where the recurrent laryngeal nerve travels, and the thyroid cartilage plate are all important functional structures. During the planning, their spatial positions need to be specifically marked, and a spherical avoidance area with a radius of 5 mm is generated centered on these positions to ensure that the motion path of the robotic arm does not touch these key parts and avoid unnecessary damage to the patient. The path segmentation control points are also an important part of the pre-operative planning. The surgical process is generally divided into three stages: "entry path", "resection path", and "exit path". The entry path starts from the oral cavity entrance and goes all the way to the starting point of the safety distance outside the lesion; the resection path is planned point by point along the surface of the safety margin and is the key path directly acting on the lesion resection; the exit path returns from the resection end point to the initial position. Each stage contains a series of ordered control points, which not only have accurate three-dimensional coordinates (X, Y, Z), but also the attitude angles of the end effector, namely the pitch angle, yaw angle, and roll angle, as well as the motion speed parameters, providing detailed guidance for the movement of the robot.

[0045] Spatial coordinate system registration and pose mapping are the key links to achieve precise operation. First, a unified coordinate system needs to be established. Based on the anatomical reference of the patient's neck, the midpoint of the posterior edge of the cricoid cartilage arch is defined as the origin, the X-axis points horizontally to the right, the Y-axis points towards the head along the sagittal plane of the human body, and the Z-axis is perpendicular to the XY plane and points ventrally. Both the pre-operative planned path and the real-time pose of the intraoperative robot need to be transformed into this unified coordinate system for subsequent calculations and operations. The pose information of the robotic arm end effector includes the position vector (P_x, P_y, P_z) in the global coordinate system and the attitude matrix (R_x, R_y, R_z) represented by Euler angles, which accurately describe the position and attitude of the end in space.

[0046] To better process the preoperative planning path, it is necessary to parameterize it. For the sequence of three-dimensional control points in the resection path, a cubic spline interpolation algorithm is usually used to fit the continuous trajectory, which can ensure smooth movement between adjacent control points and avoid jerks or mutations during the movement of the robot. For complex surfaces, such as the arcuate resection area on the surface of the vocal cords, the NURBS surface parameterization method is used to represent the path as the mapping relationship between the surface parameters (u, v) and the spatial coordinates. At the same time, each control point is also attached with motion constraint conditions. For example, the speed of entering the path is limited within a certain range, generally ≤10mm / s. The resection path has extremely high requirements for position accuracy, and the error needs to be ≤0.5mm. Also, there are strict requirements for the angle between the end effector and the tissue surface. For example, the laser knife needs to be perpendicular to the resection surface to ensure the accuracy of the surgical operation.

[0047] The fusion process of the real-time pose and the planned path is the core step to ensure the accurate execution of the operation by the robot. During the operation, it is necessary to determine the position of the current pose of the robot in the planned path. By using the nearest neighbor search algorithm, calculate the Euclidean distance between the current position of the robot's end and all control points in the planned path, so as to find the point P_near with the shortest distance, and determine which stage of the path it is currently in. For the curved resection path, the shortest distance from the point to the surface, that is, the projection method, is used to determine the projection point of the current pose on the surface, and calculate the progress parameter s of this point along the surface parameters (u, v). The value range of s is 0≤s≤1, where 0 represents the starting point of the path and 1 represents the ending point of the path.

[0048] If the deviation between the current pose and the planned path exceeds the preset threshold, generally 1mm, the path correction algorithm will be triggered. At this time, taking the current position as the new starting point, a sub-path from the current pose to the subsequent control points is regenerated in the remaining planned path. During this process, the influence of intraoperative soft tissue deformation should be fully considered. Through the real-time registration of the three-dimensional model of the larynx and the 3D endoscope image, the safety margin parameters in the planned path are dynamically adjusted. For example, when it is detected that the vocal cords have a displacement within 5° due to respiratory movement, the control point coordinates in this area are adjusted proportionally according to the deformation direction to ensure that the surgical operation is always carried out within a safe and effective range.

[0049] When generating the target pose points, various factors need to be comprehensively considered. In terms of path segmentation, when entering the path stage and moving from the initial position to the starting point of the safe distance outside the lesion, the A* algorithm is used to search for the optimal path in the three-dimensional model of the neck anatomical structure constructed before surgery, avoiding obstacles such as the mandible and the root of the tongue. An intermediate target pose point is generated every 5 mm to ensure that the end effector moves in a uniform straight line close to the lesion area. In the resection path stage, for a planar resection area, target points are generated in a grid pattern, such as at 2 mm × 2 mm intervals. The attitude angle of each point should ensure that the tool plane of the end effector is parallel to the resection plane. For a curved surface area, sampling is performed at equal arc lengths along the surface parameter lines. The sampling interval is adjusted according to the size of the lesion, generally 1 - 3 mm. At the same time, attitude interpolation is used to ensure that the tool direction is always perpendicular to the tangent plane on the curved surface. In the exit path stage, returning from the resection end point to the initial position along the original path, the generation of the target pose points is symmetric to the entry path. It is necessary to avoid passing through the resected tissue area repeatedly. The reverse path planning algorithm is used to ensure that the deviation of the motion trajectory from the entry path is ≤ 0.3 mm.

[0050] At the same time, when generating the target pose points, the kinematic constraints of the robotic arm must be satisfied, including the joint angle range, the end effector speed limit, and the acceleration limits of each joint. For target poses that exceed the constraints, optimization is performed through the damping least squares method during inverse kinematics solution. The joint angles are adjusted to the feasible region while ensuring the end effector position accuracy. The time-optimal control algorithm is also used to allocate the motion time between adjacent target pose points to ensure that the robotic arm can smoothly transition during the acceleration and deceleration phases and avoid vibration errors caused by sudden speed changes.

[0051] When the target pose points fall into the avoidance area of key anatomical structures, such as the spherical area of the recurrent laryngeal nerve, local path replanning is triggered. Taking the boundary of the avoidance area as a constraint condition, a safety boundary with a radius of 5 mm is generated around it. The artificial potential field method is used to guide the end effector to bypass, and the target poses that meet the avoidance conditions are recalculated to ensure that the minimum distance from the avoidance area is ≥ 3 mm.

[0052] After generating the target pose points, format conversion of the pose parameters is also required. They are converted into an input format recognizable by the robot controller, such as Denavit - Hartenberg parameters, and a motion mode identifier is added, such as linear motion, joint space motion, etc. Before outputting the target pose points, the joint angles corresponding to this pose are calculated through forward kinematics to check whether they exceed the working space and joint limits of the robotic arm. If there is an overlimit situation, the path planning module is returned to adjust the target pose. At the same time, the Kalman filter algorithm is used to smooth the continuously generated target pose points to filter out high-frequency errors caused by intraoperative vibrations or sensor noise and ensure the stability of the pose sequence.

[0053] In the embodiment of the present application, the joint control module 160 is configured to generate a joint control instruction for the transoral surgical robot based on the next target pose point of the transoral surgical robot. Correspondingly, the next target pose point of the transoral surgical robot describes the desired position and orientation of the end effector of the robot in the surgical space, but the actual movement of the robot is achieved by the rotation or movement of each joint. Therefore, it is necessary to convert the target pose point into specific motion parameters of each joint, that is, the joint control instruction, so that the robot can complete the action as expected. In particular, the joint control instruction is a specific signal for controlling the movement of each joint of the robot. By sending these instructions to the joint drive system of the robot, the joints of the robot can move in a predetermined manner, thereby driving the end effector of the robot to accurately reach the next target pose point and realizing the positioning and action execution of the surgical operation.

[0054] The following is a detailed elaboration of a specific implementation process of "generating a joint control instruction for the transoral surgical robot based on the next target pose point of the transoral surgical robot": First, perform kinematic parameter analysis to convert the spatial information of the pose point into input data that can be processed by the robot kinematic model. The target pose point includes two parts: the position coordinates and the attitude matrix. The position coordinates define the target position of the end effector in the global coordinate system (for example, when removing vocal cord polyps, the target position is X = 5mm, Y = 120mm, Z = 15mm), and the attitude matrix represents the spatial orientation of the end through Euler angles (for example, the laser knife needs to be at a 75° pitch angle with the mucosal surface). During the analysis process, the system needs to verify whether these parameters meet the requirements of the safety margin and avoidance area planned before the operation, such as checking whether the Z coordinate is greater than the critical value of the recurrent laryngeal nerve area to ensure that the target pose is within the safe operation space. For a curved resection path (such as the vocal cord mucosal surface), it is also necessary to convert the surface parameters (u, v) into Cartesian coordinates and ensure the position accuracy meets the millimeter-level requirements through the NURBS surface fitting algorithm.

[0055] Next, it enters the inverse kinematics solution process, which is a crucial step in converting the end - effector pose into joint angles. Transoral surgical robots mostly adopt multi - degree - of - freedom serial manipulators, and their kinematic models are established based on the Denavit - Hartenberg (D - H) parameter method. This method constructs the transformation relationship from the joint space to the end - effector space by defining parameters such as the link length, twist angle, and joint offset of each joint. The goal of inverse kinematics solution is to deduce the angle values of each joint based on the end - effector pose. There may be multiple feasible solutions (kinematic redundancy) in this process, and an optimization algorithm is needed to select the optimal solution. Optimization strategies usually include avoiding joint limit positions, minimizing the change in joint angles, or reducing energy consumption. For example, when there are multiple combinations of shoulder rotation angles to achieve the same end - effector position, the combination with the smallest angle change is preferentially selected to avoid violent movement of the manipulator. During the solution process, it is necessary to continuously verify whether the joint angles are within the physically allowed range (such as the elbow joint rotation angle limit is - 90° to 90°). If it exceeds the limit, a multi - solution selection mechanism is triggered to ensure the feasibility of the solution result.

[0056] To address the positioning error caused by joint clearances in traditional robotic systems, a joint clearance compensation mechanism needs to be introduced when generating commands. Before the operation, precise calibration is used to obtain the clearance parameters of each joint (such as the angle deviation range caused by gear transmission clearances), and a non - linear error model is established to clarify the response differences of the joint during forward and reverse movements. For example, if there is a 0.3° clearance when a certain joint moves forward and a 0.2° clearance when it moves backward, the system automatically compensates for this deviation when calculating the target joint angle, increasing the target angle for forward movement by 0.3° to offset the displacement error caused by the clearance. During the operation, combined with the real - time pose feedback of the 3D endoscope, the compensation parameters are dynamically adjusted: when there is a deviation between the actual end - effector position and the target position, the joint clearance error is deduced through a closed - loop control algorithm, and the compensation amount is updated in real - time to ensure that the positioning accuracy of the manipulator end meets the surgical requirements (such as the error does not exceed 0.3 mm during laryngeal cancer resection).

[0057] Motion trajectory planning is the core link in generating joint control instructions, which requires converting discrete target pose points into continuous joint motion curves. According to the path segmentation requirements planned before surgery (such as the entry path speed ≤ 10 mm / s and the resection path speed ≤ 5 mm / s), the motion process is decomposed into an acceleration segment, a constant-speed segment, and a deceleration segment, and the angle-time curves of each joint are generated through trapezoidal speed planning or cubic spline interpolation algorithms. For a multi-degree-of-freedom robotic arm, it is necessary to ensure the coordinated motion of each joint. For example, when removing vocal cord polyps, the wrist joint and the elbow joint need to rotate at a specific ratio to keep the laser knife head always perpendicular to the mucosal surface. During the trajectory planning process, the dynamic constraints of the robotic arm, such as the maximum acceleration limit, also need to be considered to avoid vibration errors caused by sudden speed changes and ensure smooth and impact-free motion. The generated trajectory curve contains the target angles, motion speeds, and acceleration parameters of each joint, and these parameters need to be converted into an instruction format recognizable by the robot controller, which usually includes the joint number, the target angle value, speed and acceleration parameters, and the motion mode (such as linear motion or joint space motion).

[0058] After the instructions are generated, multiple verifications are required to ensure safety and feasibility. First, a workspace check is performed. The theoretical position of the end effector is calculated through forward kinematics to verify whether it is within the reachable range of the robotic arm, avoiding instructions that cause the robotic arm to exceed its physical limits. Second, a joint limit check is carried out. The target angles of each joint are checked one by one to see if they are within the allowable range. If the angle of a certain joint exceeds the limit (such as the shoulder joint rotation angle exceeding 180°), the inverse kinematics module is returned to solve again and other feasible solutions are selected. The avoidance area check link combines the key structure positions planned before surgery, and the joint motion trajectory is deduced from the end position to ensure that the end does not enter the avoidance area (such as the recurrent laryngeal nerve area) during the motion process. Finally, real-time feedback correction is performed. The current joint state of the robotic arm is obtained (such as reading the current angle through an encoder), and the joint motion increment from the current state to the target state is calculated. If the increment exceeds the safety threshold (such as the single-joint angle change rate exceeding 60° / s), the trajectory planning parameters are adjusted to complete the motion in stages to avoid the robotic arm from jerking.

[0059] In summary, the control system 100 of the transoral surgical robot according to the embodiments of the present application is elucidated. First, it acquires thin-slice CT scan data covering the neck and laryngeal regions and constructs a three-dimensional laryngeal model. Then, it uses a 3D endoscope to collect a stereoscopic image of the surgical field, performs semantic-level dynamic registration based on visual features with the three-dimensional laryngeal model to determine the pose information of the transoral surgical robot. Next, in combination with this pose information and the preoperatively planned surgical resection path, it outputs the next target pose point of the transoral surgical robot. Finally, based on this target pose point, it generates joint control instructions for the transoral surgical robot. In this way, it helps to shorten the intraoperative navigation preparation time and ensure a high degree of consistency between the robotic arm movement trajectory and the lesion resection path, and can fundamentally solve the trajectory deviation risk caused by registration deviation in the traditional system.

[0060] Figure 4 FIG. is a flowchart of a control method for a transoral surgical robot according to an embodiment of the present application. As Figure 4 shown, in the control method of the transoral surgical robot, it includes: S110, acquiring thin-slice CT scan data covering the neck and laryngeal regions; S120, constructing a three-dimensional laryngeal model based on the thin-slice CT scan data covering the neck and laryngeal regions; S130, using a 3D endoscope to collect a stereoscopic image of the surgical field; S140, determining the pose information of the transoral surgical robot based on visual registration between the stereoscopic image of the surgical field and the three-dimensional laryngeal model; S150, outputting the next target pose point of the transoral surgical robot based on the pose information of the transoral surgical robot and the preoperatively planned surgical resection path; S160, generating joint control instructions for the transoral surgical robot based on the next target pose point of the transoral surgical robot.

[0061] Here, those skilled in the art can understand that the specific operations of each step in the above control method of the transoral surgical robot have been described in detail in the description of the control system of the transoral surgical robot above with reference to Figures 1 to 3 and thus, the repeated description thereof will be omitted.

[0062] In summary, the control method of the transoral surgical robot based on the embodiments of the present application is elucidated. First, thin-slice CT scan data covering the neck and larynx regions is acquired and a three-dimensional laryngeal model is constructed. Then, a stereoscopic image of the surgical field is collected using a 3D endoscope, and semantic-level dynamic registration based on visual features is performed between the stereoscopic image and the three-dimensional laryngeal model to determine the pose information of the transoral surgical robot. Next, in combination with this pose information and the preoperatively planned surgical resection path, the next target pose point of the transoral surgical robot is output. Finally, based on this target pose point, joint control instructions for the transoral surgical robot are generated. In this way, it helps to shorten the intraoperative navigation preparation time and ensure a high degree of consistency between the robotic arm movement trajectory and the lesion resection path, and can fundamentally solve the risk of trajectory deviation caused by registration deviation in traditional systems.

Claims

1. A control system for an oral surgery robot, characterized in that, Comprising: A CT scan data acquisition module, configured to acquire thin-slice CT scan data covering the neck and larynx regions; A larynx three-dimensional model construction module, configured to construct a larynx three-dimensional model based on the thin-slice CT scan data covering the neck and larynx regions; A stereoscopic image acquisition module, configured to acquire a stereoscopic image of the surgical field using a 3D endoscope; A pose information determination module, configured to determine the pose information of the transoral surgical robot based on the visual registration between the stereoscopic image of the surgical field and the larynx three-dimensional model, wherein the pose information determination module is configured to: perform semantic-level dynamic registration based on visual features on the stereoscopic image of the surgical field and the larynx three-dimensional model to obtain the pose information of the transoral surgical robot; A next target pose point determination module, configured to output the next target pose point of the transoral surgical robot based on the pose information of the transoral surgical robot and the preoperatively planned surgical resection path; A joint control module, configured to generate joint control instructions for the transoral surgical robot based on the next target pose point of the transoral surgical robot.

2. The control system of the oral surgery robot according to claim 1, characterized in that, The pose information determination module includes: A visual feature extraction unit, configured to perform visual feature extraction on the stereoscopic image of the surgical field and the larynx three-dimensional model to obtain a visual feature encoded map of the surgical field and a visual feature encoded map of the larynx three-dimensional model; A visual perception registration encoding unit, configured to perform semantic-level registration of visual features based on feature structure guidance on the visual feature encoded map of the surgical field and the visual feature encoded map of the larynx three-dimensional model to obtain a visual perception registration encoded map; A pose information generation unit, configured to obtain the pose information of the transoral surgical robot based on the visual perception registration encoded map.

3. The control system of the oral surgery robot according to claim 2, wherein The visual feature extraction unit is configured to: perform visual feature extraction based on three-dimensional hollow convolution encoding on the stereoscopic image of the surgical field and the larynx three-dimensional model to obtain the visual feature encoded map of the surgical field and the visual feature encoded map of the larynx three-dimensional model.

4. The control system of the oral surgery robot according to claim 2, characterized in that, The visual perception registration encoding unit includes: A local feature structure similarity search alignment sub-unit, configured to perform local feature structure similarity search alignment on the visual feature encoded map of the surgical field and the visual feature encoded map of the larynx three-dimensional model to obtain a set of phase alignment {surgical field visual feature local detail encoding matrix, larynx three-dimensional model visual feature local encoding matrix} feature pairs; An attention weight calculation sub-unit, configured to perform feature joint dynamic attention perception on each phase alignment {surgical field visual feature local detail encoding matrix, larynx three-dimensional model visual feature local encoding matrix} feature pair in the set of phase alignment {surgical field visual feature local detail encoding matrix, larynx three-dimensional model visual feature local encoding matrix} feature pairs to obtain a set of visual perception registration joint perception attention weights; The visual perception registration coding feature generation subunit is used to perform gated significant aggregation on the set of {surgical field visual feature local detail coding matrix, laryngeal three-dimensional model visual feature local coding matrix} feature pairs based on the set of visual perception registration joint perception attention weights to obtain the visual perception registration coding map.

5. The control system of the oral surgery robot according to claim 4, wherein The local feature structure similarity search alignment subunit is used to: Decouple the features of the surgical field visual feature coding map and the laryngeal three-dimensional model visual feature coding map to obtain a set of surgical field visual feature local detail coding matrices and a set of laryngeal three-dimensional model visual feature local coding matrices; Based on the feature phase alignment degree between any two surgical field visual feature local detail coding matrices and laryngeal three-dimensional model visual feature local coding matrices in the set of surgical field visual feature local detail coding matrices and the set of laryngeal three-dimensional model visual feature local coding matrices, perform feature structure similarity search alignment on the set of surgical field visual feature local detail coding matrices and the set of laryngeal three-dimensional model visual feature local coding matrices to obtain the set of {surgical field visual feature local detail coding matrix, laryngeal three-dimensional model visual feature local coding matrix} feature pairs with phase alignment.

6. The control system of the oral surgery robot according to claim 5, characterized in that, The attention weight calculation subunit includes: The visual perception registration trace metric value calculation secondary subunit is used to perform matrix trace metric on each {surgical field visual feature local detail coding matrix, laryngeal three-dimensional model visual feature local coding matrix} feature pair with phase alignment to obtain a set of visual perception registration trace metric values; The visual perception registration joint perception attention weight calculation secondary subunit is used to perform normalization processing based on the sigmoid function on the set of visual perception registration trace metric values to obtain the set of visual perception registration joint perception attention weights.

7. The control system of the oral surgery robot according to claim 6, wherein The visual perception registration joint perception attention weight calculation secondary subunit is used to: Perform manifold optimization based on the bidirectional connected cardinality on the set of visual perception registration trace metric values based on the set of {surgical field visual feature local detail coding matrix, laryngeal three-dimensional model visual feature local coding matrix} feature pairs with phase alignment to obtain a set of optimized visual perception registration trace metric values; Perform normalization processing based on the sigmoid function on the set of optimized visual perception registration trace metric values to obtain the set of visual perception registration joint perception attention weights.

8. The control system of the oral surgery robot according to claim 7, characterized in that The pose information generation unit is used to: input the visual perception registration coding map into a pose information calculator based on a decoder to obtain the pose information of the transoral surgical robot.

9. A control method for an oral surgery robot, characterized in that, It includes: Obtain thin-slice CT scan data covering the neck and laryngeal regions; Based on the thin-slice CT scan data covering the neck and laryngeal regions, construct a laryngeal three-dimensional model; Use a 3D endoscope to collect stereoscopic images of the surgical field; Based on the visual registration between the stereoscopic image of the surgical field and the three-dimensional model of the larynx, determine the pose information of the transoral surgical robot, including: performing semantic-level dynamic registration based on visual features on the stereoscopic image of the surgical field and the three-dimensional model of the larynx to obtain the pose information of the transoral surgical robot; Based on the pose information of the transoral surgical robot and the preoperatively planned surgical resection path, output the next target pose point of the transoral surgical robot; Based on the next target pose point of the transoral surgical robot, generate the joint control instruction of the transoral surgical robot.

10. The control method of the oral surgery robot according to claim 9, characterized in that, Performing semantic-level dynamic registration based on visual features on the stereoscopic image of the surgical field and the three-dimensional model of the larynx to obtain the pose information of the transoral surgical robot, including: Performing visual feature extraction on the stereoscopic image of the surgical field and the three-dimensional model of the larynx to obtain a visual feature encoding map of the surgical field and a visual feature encoding map of the three-dimensional model of the larynx; Performing semantic-level registration of visual features based on feature structure guidance on the visual feature encoding map of the surgical field and the visual feature encoding map of the three-dimensional model of the larynx to obtain a visual perception registration encoding map; Based on the visual perception registration encoding map, obtain the pose information of the transoral surgical robot.

Citation Information

Patent Citations

  • Textile state detection and analysis system

    CN119887788A

  • Remote automatic acquisition system and method for pressure gauge data

    CN119942509A

  • Laser hole processing method and system for printed circuit board

    CN119952309A

  • Power transformation line alarm device based on real-time monitoring

    CN120294499A

  • Infrared thermal imaging-based intelligent search and rescue method and device for fire personnel

    CN120856961A

Cited By

  • Power transformation line alarm device based on real-time monitoring

    CN120294499A

  • In-vitro magnetic target positioning and navigation method and system for throat

    CN120420086A

  • Spatial positioning method, device and system based on grid coding

    CN121647817A

  • Tumor incision edge design method based on local skin image

    CN122115200A