Intelligent control method, system and equipment for oral cranio-maxillofacial osteotomy robot, medium and robot
By fusing a large language model with a 3D image segmentation model, autonomous planning and real-time control of the oral and maxillofacial surgical robot were achieved, solving the problems of inaccurate positioning and unstable operation in traditional systems, and improving the intelligence and safety of the surgery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-04-14
AI Technical Summary
Current oral and maxillofacial surgery suffers from problems such as intraoperative positioning deviation, uncertain cutting path, and insufficient operational stability. Traditional surgical robot systems lack intelligent understanding and feedback capabilities, making it difficult to flexibly adjust to complex tasks.
A framework integrating a large language model and a 3D image segmentation model is adopted. The osteotomy path is generated through natural language task instructions, and multimodal information is combined for real-time perception and control to achieve autonomous planning and execution.
It improves the accuracy and safety of surgical path planning, lowers the operational threshold, and enhances the intelligence and clinical applicability of the robotic system, especially significantly improving the segmentation efficiency and accuracy in multi-site combined osteotomy procedures.
Smart Images

Figure CN121845748A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of surgical robot control technology, and in particular to an autonomous generation method for robotic arm instruction planning. Background Technology
[0002] Oral and maxillofacial surgery is one of the most challenging surgical fields due to the complex anatomical structures of the surgical area, limited operating space, proximity to important nerves and blood vessels, and the involvement of facial function and aesthetics. Traditional oral and maxillofacial surgery relies on the surgeon's experience and visual intuition. In scenarios involving the resection of complex lesions and precise reconstruction of bony structures, it often faces problems such as intraoperative positioning errors, uncertain cutting paths, and insufficient operational stability. The introduction of digital surgical technologies, such as preoperative image modeling, navigation systems, and 3D printing, has improved preoperative planning and intraoperative support capabilities to some extent. Preoperative modeling based on image segmentation relies on professionals manually annotating structures using medical imaging software (such as Mimics and 3D Slicer), including multiple areas such as the maxilla, mandible, zygomatic bone, and chin, to perform precise surgical planning. Subsequently, the planning file is exported as data in different formats for printing personalized surgical guides or uploaded to navigation software systems to improve surgical accuracy. This process is time-consuming and requires a high level of operational experience. How to integrate these technologies more efficiently and intelligently to achieve a closed loop of preoperative planning, intraoperative operation, and real-time feedback remains a key focus for the industry.
[0003] In recent years, the successful application of surgical robot systems, represented by the da Vinci system, in fields such as neurosurgery and urology has demonstrated advantages such as minimally invasive procedures, high precision, and multiple degrees of freedom. However, due to the lack of compatible professional control modules, the application of existing surgical robot systems in complex bony procedures of the oral and maxillofacial region remains limited. Current research on craniofacial surgical robots has achieved a preliminary combination of image guidance and force feedback mechanisms. For example, a robot platform combining an infrared optical navigation system and a six-dimensional force sensor can assist in intraoperative three-dimensional positioning and force monitoring, enhancing the system's ability to perceive anatomical variations and surgical risks. Through preoperative CBCT image three-dimensional reconstruction and surgical plan input, the intraoperative real-time registration and navigation system can guide the surgical arm to execute a preset path; while the force feedback system provides resistance prompts to the surgeon, assisting in judging the cutting depth and tissue stiffness. Although these systems significantly improve the stability and safety of surgical execution, their control path planning and interaction methods rely on complex programming presets, are limited to button menus and trajectory point settings, and lack intelligent understanding and feedback capabilities for complex tasks. Since equipping every robot-assisted surgery with a computer science expert is impractical, and for deformity reconstructive procedures that require flexible adjustments based on patient needs, frequent algorithm debugging in conjunction with diverse execution sequences and dynamically changing clinical scenarios would generate significant additional time consumption, posing a very high technical barrier for clinicians. Therefore, it is necessary to integrate and construct an integrated autonomous planning, reasoning, and control system supported by multimodal data to further improve the intelligence and clinical applicability of oral and maxillofacial surgical robots.
[0004] With the rapid development of Large Language Models (LLMs) technology, in addition to their outstanding performance in natural language understanding and interactive feedback, research has demonstrated their strong potential in task planning and multimodal information processing fusion, providing crucial support for introducing cognitive intelligence into robotic systems. However, general-purpose LLMs still have shortcomings in understanding professional medical terminology, expressing geometric constraints of surgical pathways, and dynamic decision-making, especially in clinical pathway planning and risk avoidance judgment, requiring customized training based on professional knowledge graphs and medical scenarios. To address the challenges of combining multi-task instance segmentation and semantic guidance tasks, researchers are gradually exploring a framework for integrating LLMs with 3D image segmentation models. In this framework, by designing a training process that includes semantic prompts for surgical procedures, large language models can not only assist in task description understanding and region guidance, but also enhance the ability to express semantic differences among multiple instances through dual-channel embedding of "image + language". For example, the model can accurately locate relevant bone regions based on the text prompt "Maxillary Lefort I osteotomy area", achieving structural segmentation of different osteotomy types in the same 3D space. This fusion mechanism is particularly suitable for planning multi-site osteotomy procedures in the maxillofacial region, which can significantly improve annotation efficiency and segmentation accuracy, providing a precise basis for robotic surgical path generation.
[0005] In intelligent robot control, previous research has attempted to use natural language task instructions as input, leveraging LLMs for task parsing and subtask decomposition to automatically generate operational steps or path planning information, and then driving the mechanical system to execute them through a control interface. This architecture can construct continuous task chains through interactive natural language prompts and historical execution records, improving the robot's understanding of the task scenario and the consistency of its execution. In some technical solutions, researchers have introduced code generation and evaluation mechanisms, or adopted dual-mode models for action planning and strategy verification, improving the accuracy and execution safety of semantic-action mapping through multi-layered interactive iteration. These explorations have preliminarily verified the feasibility of language models in robot instruction generation, but achieving closed-loop control in high-precision, high-safety medical environments still faces challenges such as data sparsity, semantic ambiguity, and insufficient interpretability.
[0006] To achieve autonomous execution and environmental adaptability of robotic systems for complex tasks, current research focuses on constructing an integrated closed-loop optimization system encompassing perception, understanding, and control. By integrating multimodal information such as vision, language, force feedback, spatial depth, and structured data, the system can dynamically adjust its strategies and execution paths, and achieve error self-correction. Some technical solutions propose employing a multimodal feature vector weighting mechanism and a dynamic environmental scene model to maintain high robustness of the control strategy under different perceptual conditions. Simultaneously, the combination of reinforcement learning and large-scale model inference can continuously optimize action sequence generation and execution accuracy, improving the generalization and fault tolerance of the strategy. Furthermore, collaborative processing mechanisms for path planning and semantic feedback are also under development. For example, SLAM methods are introduced, combining map information and language commands into the model to generate high-quality path planning schemes and achieve dynamic adjustments. Although these systems have achieved initial results in service robots, mobile platforms, and home assistants, their application in high-risk, real-time surgical procedures still lacks sufficient validation, particularly in handling intraoperative perception uncertainties, ensuring the safety and controllability of model output, and optimizing the human-machine interface. Further in-depth research is needed in these areas. Summary of the Invention
[0007] To address the aforementioned problems in the prior art, the first aspect of this application provides an intelligent control method for an oral craniofacial osteotomy robot, comprising:
[0008] Step S1: Image Input: Acquire whole-head CT images of the surgical subject and perform preprocessing;
[0009] Step S2: Structure Recognition: Using a large language model to input semantic instructions, and using a pre-built cross-modal data conversion interface and a pre-built instance segmentation model to perform instance segmentation, an anatomical structure mask matching the original CT image space of the surgical area corresponding to a specific osteotomy procedure type is automatically output and mapped to the semantic coordinates in the surgical procedure language description in real time. The large language model generates text instructions based on natural language prompts and related task instructions. The text instructions include the robotic arm's movements, the target object and its target position.
[0010] Step S3: Osteotomy Path Generation: First, the site information is detected and acquired through the navigation system. At the same time, the spatial coordinates obtained from real-time observation of the patient's space are image encoded based on binocular visual tracing, and the text command sequence is text encoded. Then, the results of image encoding and text encoding are weighted and input into the Gaussian mixture model. The Gaussian mixture model is decoded in combination with the site information and loss is calculated using real annotations. Finally, the osteotomy path trajectory points under the optical tracking system are predicted and generated.
[0011] Step S4: Navigation space coordinate transformation: The osteotomy path trajectory points under the optical tracking system are calculated to obtain the osteotomy path trajectory under the robot coordinate system through the coordinate transformation matrix.
[0012] Further, in step S2: structure identification includes the following steps:
[0013] Step S2.1: Constructing a procedure-structure aligned dataset: Collect whole-head CT images of several patients, and perform normalization, artifact removal, and voxel resampling preprocessing. Use medical image software to annotate the anatomical structures of the procedure-related regions, and create instance masks for the surgical areas of specific osteotomy procedures to form a complete procedure-structure aligned dataset. Each type of surgery's procedure-structure aligned dataset includes a reconstructed model, procedure label, instance index, and three-dimensional coordinate information. The osteotomy procedure is selected from any one of the following: zygomatic L-shaped osteotomy, maxillary Lefort I-shaped osteotomy, mandibular sagittal split osteotomy, genioplasty, high condylar resection, and segmental mandibular osteotomy.
[0014] Step S2.2: Construct an instance segmentation model based on the formula-structure alignment dataset: The model adopts an encoder-decoder structure, with 4 encoding modules and 4 decoding modules. Residual connections are used to enhance the ability of shallow spatial information to supplement deep abstract semantics. During model training, the input 3D image is set as... The label mask is The network learning objective is to generate a prediction mask. ,in This represents a parameterized 3D U-Net function, optimized using the Tversky loss function:
[0015] ;
[0016] in: and Let i be the i-th voxel representing the true value and the predicted value, respectively; , ;
[0017] Training was performed using the Adam optimizer with an initial learning rate of The gradient decays to half its original value every 50 rounds, with a maximum iteration count of 300 rounds; each training round employs a gradient accumulation mechanism with a batch size of 1. ;
[0018] Four enhancement strategies—random rotation, random scaling, mirror flipping, and Gaussian noise perturbation—are employed to dynamically combine with set probabilities and are simultaneously applied to the image and corresponding label mask of each training sample. Perturbation samples are dynamically generated during training.
[0019] (1) In terms of spatial angle perturbation, a three-dimensional Euler angle rotation mechanism is used to randomly rotate the input image voxel data;
[0020] The 3D input image is The rotation angles are as follows:
[0021] ,
[0022] in, ,
[0023] Angle rotation operation is defined as:
[0024] ;
[0025] This operation simulates the changes in the patient's head posture during surgery, improving the model's ability to recognize changes in osteotomy orientation.
[0026] (2) Introduce anisotropic scaling strategy in terms of structural scale: Let the scaling factor be ,in The scaling perturbation operation is defined as follows:
[0027] , ;
[0028] By adjusting the bone scale to simulate the physiological structural differences between individuals, the model's segmentation ability under different bone sizes is enhanced.
[0029] (3) To address the issue of left-right structural symmetry, a three-dimensional mirror flipping mechanism is introduced: In In terms of probability in three dimensions Perform an axial mirror flip operation; the transformation expression is:
[0030] , ;
[0031] This enhancement method improves the model's classification neutrality on symmetric structures and reduces the risk of overfitting to skewed structures in the training set;
[0032] (4) Introduce Gaussian noise perturbation at the signal level This is used to simulate signal drift caused by dose changes and differences in reconstruction algorithms during CT scanning; the perturbation follows a normal distribution.
[0033] , ;
[0034] in The noise intensity is expressed in original grayscale values; this mechanism improves the model's stability and response consistency to input images under different imaging conditions.
[0035] During the training process, the above enhancement strategies are dynamically combined with set probabilities and applied simultaneously to the image and the corresponding label mask; the complete enhancement process is expressed as follows:
[0036] ;
[0037] The operators correspond to Gaussian noise addition, mirror flip, scaling perturbation and angle rotation operations in turn, ensuring that the spatial geometry and semantic structure remain consistent in the perturbed image;
[0038] Step S2.3: Based on the constructed cross-modal data structure conversion module, the labeled data and the large language model are adapted, and the anatomical structure labels and surgical language descriptions are embedded and fused. The cross-modal data structure conversion module supports "osteotomy type + image region + structural semantics" as joint input.
[0039] The construction methods for the cross-modal data structure transformation module include:
[0040] 1) Establishment of a medical prompt database: Collect a large number of real surgical records and standard surgical procedures, and extract structural terms;
[0041] 2) Tag mapping mechanism: Map each structural tag to a standard technical description language, and adjust the token structure through multiple rounds of iteration to improve its semantic embedding quality in LLM;
[0042] 3) Embedding fusion and task adaptation: Using "osteotomy type + image region + anatomical structure semantics" as joint input, a structured embedding vector is generated to guide 3D image segmentation;
[0043] 4) Mask alignment mechanism: The final output of the model is a mask map that is spatially registered with the original CT image, which supports direct mapping to the navigation space or subsequent path encoding;
[0044] Step S2.4: The application of the instance segmentation model, the large language model, and the cross-modal data structure conversion module collaboratively completes the instance segmentation task of medical images: The large language model is used as a semantic instruction and collaborates with the instance segmentation model to perform instance segmentation to output an instance structure mask of the surgical area that matches the space of the original CT image and is mapped to the semantic coordinates in the surgical language description in real time. The large language model autonomously generates text instructions based on natural language prompts and relevant task instructions.
[0045] Furthermore, in step S3, a multilayer perceptron is introduced to perform Gaussian mixture modeling on the fused features of the image and language coding vectors, thereby achieving compressed representation and sampling generation of the surgical path; the input image features Depend on The network acts as an image encoder to extract boundary features; the Transformer architecture extracts features from the generated text instructions. Based on the known tracking tags, the three-dimensional coordinates of the target object. The trajectory is tracked and identified in the coordinate system of the optical tracking system; furthermore, a Gaussian mixture model based on a multilayer perceptron is used to model the trajectory distribution, which is based on features formed by concatenating visual latent coding and text latent coding. To study;
[0046] The loss function of this GMM is:
[0047] ;
[0048] ;
[0049] ;
[0050] Among them, parameter set Depend on Extracted from the input.
[0051] To predict the distribution of trajectory points, This represents a Gaussian distribution.
[0052] Further, in step S4, based on the predicted trajectory points... In the coordinate system of the optical tracking system, it can be achieved through the coordinate transformation matrix. The trajectory corresponding to the robot's coordinate system is calculated as follows: The transformation matrix is also obtained by the optical tracking system.
[0053] Furthermore, the intelligent control method for the oral craniofacial osteotomy robot based on a large language model further includes: Step S5: Osteotomy path execution: The robotic arm executes operation instructions on the surgical object according to the osteotomy path, and the robotic arm is tracked in real time through the surgical navigation system to ensure that its operation is consistent with the surgical planning results;
[0054] The target pose matrix of the robotic arm is: ;
[0055] in, : The target pose matrix of the robotic arm;
[0056] The transformation matrix from the reference frame coordinate system R to the robot arm base coordinate system O;
[0057] The transformation matrix from the image guiding coordinate system I to the reference coordinate system R;
[0058] : The transformation matrix from the image guidance coordinate system I to the surgical tool coordinate system S;
[0059] : The transformation matrix from the surgical instrument coordinate system S to the robotic arm end effector coordinate system E.
[0060] Furthermore, the intelligent control method of the oral craniofacial osteotomy robot based on the large language model further includes step S6: force feedback control: a real-time force sensing mechanism is introduced at the robot execution end to realize the integrated collaborative control strategy of sensing-control-constraint, and a speed adjustment law based on the contact reaction force between the lever input and the bone surface is constructed, supplemented by a virtual impedance feedback model established by the danger zone, and the end speed and posture are dynamically adjusted according to the intraoperative state.
[0061] In the force feedback mechanism cooperative control mode, the robot adjusts its speed and direction of motion based on real-time force signals:
[0062] Speed control law: ;
[0063] in, Main control input force, For bone surface contact reaction force, The adjustment coefficient is used to control the end effector's movement speed in a proportional manner to the tactile feedback.
[0064] Virtual reaction force modeling: The system constructs a virtual impedance field when approaching the dangerous anatomical area, and the virtual force is calculated as follows: ;
[0065] in, and Represent the danger membership functions for position and velocity, respectively. ;
[0066] The final control model introduces damping. With inertial parameters Construct a complete feedback loop:
[0067] ;
[0068] in, , For feedback gain coefficient, For virtual reaction force, , These are terminal acceleration and velocity, respectively.
[0069] The second aspect of this application provides an intelligent control system for an oral craniofacial osteotomy robot, which includes:
[0070] The image input module is used to acquire and preprocess whole-head CT images of the surgical subject.
[0071] The structure recognition module is used to input semantic instructions with the help of a large language model, and to perform instance segmentation in collaboration with a pre-built cross-modal data conversion interface and a pre-built instance segmentation model. It automatically outputs an anatomical structure mask that matches the original CT image space of the surgical area corresponding to the selected osteotomy procedure type and is mapped to the semantic coordinates in the surgical procedure language description in real time. The large language model generates text instructions based on natural language prompts and related task instructions. The text instructions include the robotic arm's movements, the target object and its target position.
[0072] The osteotomy path generation module first detects and acquires site information through the navigation system, and simultaneously performs image encoding on the spatial coordinates obtained from real-time observation of the patient's space based on binocular visual tracing, and performs text encoding on the text command sequence; then, the results of image encoding and text encoding are weighted and input into the Gaussian mixture model, which decodes the data by combining the site information and uses real annotations for loss calculation, and finally predicts and generates the osteotomy path trajectory points under the optical tracking system;
[0073] The navigation space coordinate transformation module is used to calculate the osteotomy path trajectory in the robot coordinate system by using a coordinate transformation matrix to transform the osteotomy path trajectory points under the optical tracking system.
[0074] A third aspect of this application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the above-described intelligent control method for an oral craniofacial osteotomy robot.
[0075] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described intelligent control method for an oral craniofacial osteotomy robot.
[0076] The fifth aspect of this application provides an oral craniofacial osteotomy robot, which includes surgical instruments, a robotic arm, and a control module; the control module is used to execute the above-described intelligent control method for the oral craniofacial osteotomy robot; the robotic arm is used to hold the surgical instruments and perform osteotomy operations on the surgical object according to the control instructions output by the control module.
[0077] It is worth noting that the large language models in this application include, but are not limited to, ChatGPT-4o and ChatGPT5.
[0078] Compared with existing technologies, the above technical solution has the following technical advantages:
[0079] This invention proposes an integrated method for medical image 3D model instance segmentation and surgical path semantic generation. Based on multimodal data alignment, this method integrates medical images, surgical procedure text information, and intraoperative spatial coordinates to automate the process from structure recognition to path encoding. It aims to enhance intelligence and applicability, providing a highly interpretable and controllable task execution foundation for intelligent osteotomy robot systems. This invention innovatively integrates semantic task guidance, 3D structure recognition, and controllable path encoding technologies, providing an efficient and intelligent control framework for maxillofacial osteotomy and reduction. It possesses significant technological innovation value and broad clinical application prospects in the fields of medical artificial intelligence and surgical robots.
[0080] Compared with existing technologies, this invention achieves full autonomy and intelligence in the design of the path planning and control system for maxillofacial surgical robots, from data construction, structure recognition, path compression coding, navigation mapping to force feedback closed-loop control. It constructs a complete integrated execution framework of "semantic parsing—structure segmentation—path generation—dynamic control," possessing the following significant technical advantages:
[0081] (1) The semantic-driven multimodal intelligent segmentation system breaks through the bottleneck of traditional structural recognition methods: Existing methods mostly rely on training segmentation networks with a single image channel, lacking the ability to understand surgical intent and language task instructions. This invention introduces a general large language model (such as ChatGPT-4o or ChatGPT5) for the first time to construct a cross-modal interface, integrates "osteotomy type + image region + structural semantics" into a joint prompt, and completes the generation of instance masks for complex structures with the 3D U-Net architecture, solving the problems of boundary overlap and semantic conflict in joint osteotomy of multiple parts.
[0082] (2) The GMM path compression and high-order navigation coordinate transformation mechanism improve the robot's execution efficiency and posture accuracy: The system introduces a multilayer perceptron to perform Gaussian mixture modeling on the fusion features of image and language encoding vectors, realizing the compressed representation and sampling generation of the surgical path, effectively reducing path complexity and redundant point noise. Combined with the five-level coordinate system transformation matrix (image space—navigation space—reference coordinate system—instrument space—end effector) estimated based on SVD, this invention constructs a stable and high-precision navigation space mapping framework. In simulated surgical experiments, the Le Fort I osteotomy path positioning error is controlled within ±0.65mm, which is better than the traditional path transformation method based on manual registration (error of about 1.5mm), improving the accuracy by about 40.9%.
[0083] (3) Introducing a force feedback control model to enhance the active safety and human-machine collaboration of the system: Introducing a real-time force sensing mechanism into the robot control module and constructing a speed regulation law based on the reaction force between the input of the control lever and the contact force of the bone surface, supplemented by a virtual impedance feedback model established by the dangerous area, the system can dynamically adjust the end speed and posture according to the intraoperative cutting depth, tissue hardness and other states.
[0084] (4) The system has a strong closed-loop process, a user-friendly interface, and high scalability, providing a solid foundation for clinical deployment: The planning system proposed in this invention integrates all aspects from image input, surgical procedure selection, structural recognition, path encoding, spatial mapping to execution control into a graphical interactive platform. It supports multiple rounds of preoperative modifications, real-time intraoperative feedback, and postoperative trajectory review, greatly reducing the operational threshold. Compared to traditional surgical navigation systems that require manual setting of path points and parameter adjustment, this system can automatically generate path plans based on text semantic understanding, improving work efficiency.
[0085] In summary, this invention significantly outperforms existing technologies in terms of image semantic processing efficiency, path generation accuracy, security, and ease of operation. It shows broad application prospects, especially in multi-site combined osteotomy procedures, intervention of complex anatomical structures, and operations in high-risk adjacent areas, providing a practical and feasible technical solution for intelligent oral and maxillofacial surgery. Attached Figure Description
[0086] Figure 1 System architecture for intelligent control method of oral craniofacial osteotomy robot;
[0087] Figure 2 A flowchart of an intelligent control method for an oral craniofacial osteotomy robot;
[0088] Figure 3 A flowchart of a segmentation model of anatomical structures in maxillofacial surgical procedures based on a 3D U-Net network.
[0089] Figure 4 A framework for image and text information fusion and trajectory generation;
[0090] Figure 5 This is a modular structure diagram of the intelligent control module for the oral craniofacial osteotomy robot.
[0091] Figure 6 This is a schematic diagram illustrating the working principle of the oral craniofacial osteotomy robot. Detailed Implementation
[0092] The advantages of the present invention are further illustrated below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that the following detailed description is illustrative rather than restrictive and should not be construed as limiting the scope of protection of the present invention.
[0093] Example 1: Intelligent Control Method for Oral and Maxillofacial Osteotomy Robot
[0094] This embodiment provides a method for automatically generating and executing osteotomy paths in an oral and maxillofacial surgical robot, incorporating 3D instance segmentation, navigation space coordinate transformation, and force feedback control. The aim is to improve the level of intelligence and applicability. The system architecture is as follows: Figure 1 As shown. Specific implementation methods are as follows.
[0095] like Figure 2 As shown, the intelligent control method of the oral craniofacial osteotomy robot in this embodiment includes steps S1-S6:
[0096] Step S1: Image Input:
[0097] Full-head CT images of the surgical subject were acquired and preprocessed. All image data were acquired in DICOM format and, after importation, underwent normalization, artifact removal, and voxel resampling preprocessing to ensure data consistency and usability.
[0098] Step S2: Structure Identification:
[0099] Using a large language model to input semantic instructions, and leveraging a pre-built cross-modal data conversion interface and a pre-built instance segmentation model to perform instance segmentation, the system automatically outputs an anatomical structure mask that matches the original CT image space of the surgical area corresponding to a specific osteotomy procedure type. This mask is then mapped in real time to the semantic coordinates in the surgical procedure description. The large language model generates text instructions based on natural language prompts and relevant task instructions. These text instructions include the robotic arm's movements, the target object, and its target location.
[0100] The construction process of the instance segmentation model is as follows:
[0101] Step S2.1: Construct the technique-structure alignment dataset:
[0102] To support model training and evaluation, this invention first constructs a high-quality surgical structure image dataset. The data comes from 500 clinical patients undergoing maxillofacial reconstructive surgery. All patients were ethically approved and anonymized, and the images are high-resolution craniofacial CT data (1mm slice thickness, 0.7mm pixel). The dataset is divided into training, testing, and validation sets in an 8:1:1 ratio. All image data were acquired in DICOM format and preprocessed after import, including normalization, artifact removal, and voxel resampling, to ensure data consistency and usability. During surgical structure annotation, medical imaging software such as Mimics and 3D Slicer were used. Voxel-level annotations of the anatomical structures in the surgical-related regions were completed for each patient's image, and instance masks for the following four types of surgical regions were created:
[0103] 1) Zygomatic bone L-shaped osteotomy: Mark the body and arch of the zygomatic bone, and draw a standard L-shaped osteotomy path;
[0104] 2) Maxillary Lessor I osteotomy: The osteotomy range is marked based on the lower edge of the piriform foramen and the levels of the zygomatic alveolar strut and pterygomaxillary strut;
[0105] 3) Sagittal split osteotomy of the mandible: the sagittal plane boundary area between the medial and lateral plates of the mandibular ramus;
[0106] 4) Mentoplasty: Symmetrically mark the cortical and cancellous bone tissue of the chin, mark the morphology and structure of the mental foramen, and delineate the standard osteotomy boundaries.
[0107] Each data category includes the reconstructed model, technique labels, instance index, and 3D coordinate information. The datasets are uniformly converted to a structured format (NIfTI + JSON), supporting adaptation to model input and training interfaces.
[0108] This dataset has the advantages of semantic diversity and spatial consistency.
[0109] Step S2.2: Construct an instance segmentation model based on 3D U-Net using the formula-structure alignment dataset:
[0110] To achieve accurate identification of complex craniofacial structures and surgically guided osteotomy path delineation, this invention employs an image-driven 3D instance segmentation method, integrated with a language-guided mechanism to assist in structural semantic expression. The system adopts a collaborative architecture of "image module + language model module," where image segmentation uses 3D U-Net as the main algorithm network structure, and ChatGPT-4o is selected as the semantic instruction generation and interaction core, as detailed above. The two work together to achieve accurate structural identification, consistent semantic fusion, and unified control of spatial coordinate response. This system can perform the following tasks:
[0111] 1) Automatic structural segmentation of key bone structures such as the cheekbone, maxilla, mandible, and chin;
[0112] 2) Automatically determine the corresponding boundaries and expected osteotomy path based on surgical prompts (such as "L-shaped osteotomy of the zygomatic bone");
[0113] 3) The segmentation results are reverse-mapped to the coordinate semantics in the technical language description to achieve structure-language alignment;
[0114] 4) Simultaneously present the structural mask, language prompts, and segmentation confidence assessment in the view interaction interface.
[0115] The training and optimization process of the 3D image instance segmentation model based on 3D U-Net is as follows: Figure 3 As shown, the model employs an encoder-decoder structure, with four encoding modules and four decoding modules. Residual connections are used to enhance the ability of shallow spatial information to supplement deep abstract semantics. The input is a resampled, normalized CT image patch, and the output is a surgically related anatomical structure mask.
[0116] In model training, let the input 3D image be... The label mask is The network's learning objective is to generate a predictive mask. ,in This represents the parameterized 3D U-Net function. To improve sensitivity to imbalanced classes, the Tversky loss function is used for optimization.
[0117] ;
[0118] in: and Let i be the i-th voxel representing the true value and the predicted value, respectively; , The model's tolerance for false negatives in small structures is enhanced; the loss function adopts a weighted adjustment strategy for false detection and false negative, improving the recognition accuracy of areas such as the thin layer of the zygomatic bone and the mandibular nerve foramen.
[0119] Training was performed using the Adam optimizer with an initial learning rate of The gradient decays to half its original value every 50 rounds, with a maximum iteration count of 300 rounds. Each training round employs a gradient accumulation mechanism with a batch size of 1. This ensures efficient use of video memory.
[0120] To improve the generalization ability of the 3D U-Net model in segmenting complex craniofacial structures and its robustness to clinical variations, various spatial and signal-level data augmentation techniques were employed, primarily including random rotation, random scaling, mirror flipping, and Gaussian noise perturbation. These augmentation strategies were applied probabilistically to each training sample, and perturbation samples were dynamically generated during training, thereby enhancing the model's adaptability to anatomical differences, pose variations, and image noise in different cases.
[0121] In terms of spatial angle perturbation, a three-dimensional Euler angle rotation mechanism is used to randomly rotate the input image voxel data.
[0122] The 3D input image is The rotation angles are respectively ,in The rotation operation is defined as follows:
[0123] ;
[0124] This operation simulates the changes in the patient's head posture during surgery, improving the model's ability to recognize changes in osteotomy orientation.
[0125] Secondly, an anisotropic scaling strategy is introduced in terms of structural scale. Let the scaling factor be... ,in The scaling perturbation operation is defined as follows:
[0126] , ;
[0127] By adjusting the bone scale to simulate the physiological structural differences between individuals, the model's segmentation ability under different bone sizes is enhanced.
[0128] Third, to address the issue of left-right structural symmetry, a three-dimensional mirror flipping mechanism is introduced. In terms of probability in three dimensions Perform an axial mirroring operation; the transformation expression is:
[0129] , ;
[0130] This enhancement method improves the model's classification neutrality on symmetrical structures such as the left and right cheekbones and the mandibular ramus, reducing the risk of overfitting to skewed structures in the training set.
[0131] Finally, Gaussian noise perturbation is introduced at the signal level to simulate signal drift caused by dose changes and differences in reconstruction algorithms during CT scanning. The perturbation follows a normal distribution.
[0132] , ;
[0133] in The noise level is expressed in original grayscale values. This mechanism improves the model's stability and response consistency to input images under different imaging conditions.
[0134] During the training process, the above enhancement strategies are dynamically combined with set probabilities and applied simultaneously to the image and the corresponding label mask. The complete enhancement process can be expressed as follows:
[0135] ;
[0136] The operators correspond to Gaussian noise addition, mirror flipping, scaling perturbation, and angle rotation operations, respectively, ensuring that the spatial geometry and semantic structure remain consistent in the perturbed image. This enhancement module is built on the MONAI framework and integrated into the model training pipeline, significantly improving the stability and reliability of the network in complex anatomical scenes.
[0137] The model was trained and validated on a self-constructed dataset of 500 case structural annotations, achieving an average DICE coefficient of 0.89 and a recall of 0.92. The instance segmentation model in this application performs exceptionally well, particularly in areas such as the zygomatic arch and mandibular ramus. Multi-site segmentation analysis revealed that the model's recognition accuracy in areas where multiple structures intersect, such as the mandibular ramus and zygomatic arch, is significantly superior to traditional threshold segmentation methods. The trained model was deployed on the Python-PyTorch platform and interfaced with upstream and downstream systems to achieve dynamic path callbacks and semantic interpretation of structural recognition results. This model possesses high practical value in the preprocessing stage of osteotomy navigation, providing a structural foundation for subsequent path encoding and navigation control.
[0138] Step S2.3: Construct a cross-modal data structure conversion module to adapt labeled data and large language models, and embed and fuse anatomical structure labels with surgical language descriptions.
[0139] To achieve an interface for adapting labeled data to the large language model chatGPT4o, a cross-modal data structure conversion module was established, which embeds and fuses anatomical structure labels with surgical procedure descriptions. This enables the large language model to have instance segmentation capabilities. This module supports using "osteotomy type + image region + structural semantics" as joint input and outputs a structural mask registered with the original image, significantly improving the accuracy of multi-task recognition.
[0140] The cross-modal data conversion module is constructed in the following ways:
[0141] 1) Establishment of a medical Prompt database: Collect a large number of real surgical records and standard surgical procedures, and extract high-frequency structural terms such as zygomatic osteotomy, LeFort I osteotomy, mandibular splitting, and chin shaping;
[0142] 2) Tag mapping mechanism: Map each structural tag (such as "right zygomatic bone osteotomy") to a standard surgical procedure description language, and improve its semantic embedding quality in LLM by adjusting the token structure through multiple rounds of iteration;
[0143] 3) Embedding Fusion and Task Adaptation: Using "osteotomy type + image region + anatomical structure semantics" as joint input, structured embedding vectors are generated to guide 3D image segmentation. This approach demonstrates good label consistency and spatial structure coordination in multi-task learning scenarios.
[0144] 4) Mask alignment mechanism: The final output mask map of the model is spatially registered with the original CT image, and supports direct mapping to the navigation space or subsequent path encoding module.
[0145] Step S2.4: The instance segmentation model, the large language model, and the cross-modal data structure conversion module are used to collaboratively complete the instance segmentation task of medical images.
[0146] The chatGPT4o model, trained as described above, collaborates with a 3D vision module to perform instance segmentation of medical images. The model can independently identify key bone structures such as the zygomatic bone, maxilla, mandible, and chin, and automatically generate osteotomy boundaries and target volume regions based on surgical prompts. Segmentation results are mapped in real-time to semantic coordinates in the surgical procedure description. The surgical sequence can be simplified to the movement and positioning of specific surgical instruments, such as moving a reciprocating saw to the starting point of the marked osteotomy line or retracting it to a safe position. chatGPT4o generates text responses based on natural language prompts and relevant task instructions, including the robotic arm's movements, the target object, and its location. Orderly and reasonable responses are considered successful osteotomy sub-task instructions. Approved text instructions are translated and decoded downstream into program code that can be executed by the robotic arm.
[0147] Step S3: Osteotomy path generation:
[0148] First, the location information is obtained through the navigation system, and the real-time spatial observation results of the patient are image encoded and the text command sequence is text encoded. Then, the results of image encoding and text encoding are weighted and input into the Gaussian mixture model. The Gaussian mixture model is decoded by combining the location information and loss is calculated using real annotations. Finally, the osteotomy path trajectory points under the optical tracking system are predicted and generated.
[0149] To achieve the effective conversion of segmentation results into executable robot paths, this invention introduces a method based on the aforementioned GMM after the image segmentation module to extract image spatial coordinates, align them with text information, and perform path compression encoding.
[0150] Specifically, this application innovatively designs an algorithmic framework to predict and generate trajectory points based on textual instructions and visual observation information, such as... Figure 4 As shown. Input image features Depend on The network acts as an image encoder to extract boundary features; simultaneously, the Transformer architecture extracts features from the generated text instructions. Based on the known tracking tags, the three-dimensional coordinates of the target object are... The trajectory is tracked and identified within the coordinate system of the optical tracking system. Furthermore, a Gaussian Mixture Model (GMM) based on a Multilayer Perceptron (MLP) is used to model the trajectory distribution. This model is based on features formed by concatenating visual latent coding and textual latent coding. To learn.
[0151] The loss function of this GMM is:
[0152] ;
[0153] ;
[0154] ;
[0155] Among them, parameter set Depend on Extracted from the input.
[0156] To predict the distribution of trajectory points, This represents a Gaussian distribution.
[0157] Step S4: Navigation space coordinate transformation:
[0158] The actual robotic arm execution trajectory = predicted planned trajectory · navigation pose transformation relationship. Specifically, based on the predicted trajectory points... In the coordinate system of the optical tracking system, it can be achieved through the coordinate transformation matrix. The trajectory corresponding to the robot's coordinate system is calculated as follows: The transformation matrix is also obtained by the optical tracking system of the surgical navigation system.
[0159] A multi-level coordinate system transformation method was used to map the trajectory in the preoperative image space to the intraoperative robot control space. This work establishes the geometric and semantic foundation for subsequent navigation execution and force feedback control.
[0160] The system involves multiple coordinate systems, including the robotic arm base coordinate system O, the robotic arm end effector coordinate system E, the patient image guidance coordinate system I, the reference frame coordinate system R, and the surgical tool coordinate system S. The transformation matrix between the robotic arm end effector and its base is called the robotic arm posture matrix. The surgical navigation subsystem reconstructs a 3D image using the patient's preoperative CT imaging data, thereby establishing an image guidance space. Spatial registration methods are used to determine the coordinate transformation between the reference frame coordinate system and the image guidance coordinate system. Then, the spatial relationship between the surgical tools and the patient area is displayed in real time within the image guidance space. Under the guidance of the surgical navigation subsystem, robot-assisted surgical operations require managing multiple spatial coordinate transformations, i.e., achieving multi-spatial coordinate registration. The purpose of registration is to guide the robotic arm to adjust its target posture using patient image data and determine the required transformation matrix. This represents the transformation from coordinate system E to coordinate system O. Once the target pose is determined based on the patient's image-guided data, the robotic arm is controlled to precisely align the surgical tools on the end effector along the surgical path.
[0161] To enable real-time tracking of the robotic arm via a surgical navigation system and ensure its operation aligns with the surgical plan, a target pose matrix for the robotic arm is required. It is a key component of system design. Key transformation matrices include:
[0162] : The transformation matrix from the image guiding coordinate system I to the reference coordinate system R;
[0163] Transformation matrix from reference frame coordinate system R to robot arm base coordinate system O;
[0164] The transformation matrix from the image guidance coordinate system I to the surgical tool coordinate system S;
[0165] : The transformation matrix from the surgical instrument coordinate system S to the robotic arm end effector coordinate system E;
[0166] : The pose matrix required by the robotic arm.
[0167] Each transformation matrix T consists of a rotation matrix and a translation matrix. To compute these transformation matrices, this study employs a point-based least squares fitting method, particularly singular value decomposition (SVD).
[0168] The transformation relationship between the image guidance space and the robotic arm base is as follows:
[0169] ;
[0170] ;
[0171] The reference frame and robotic arm base are fixed, therefore they can be directly obtained. Spatial transformation matrix It is the spatial transformation matrix that maps the image guidance space to the reference frame space, determined by a combination of point registration and out-of-plane registration methods. By moving the surgical instrument and the robotic arm end effector to the same spatial position and determining the coordinate transformation relationship between the tip of the instrument and the center of the robotic arm end effector, the following can be obtained: By using an optical tracker to track reflective marks on surgical instruments, one can obtain... .
[0172] Therefore, the posture coordinates of the robotic arm in the binocular vision navigation coordinate system can be calculated as follows: .
[0173] Step S5: Osteotomy path execution:
[0174] The robot's robotic arm executes operational instructions on the surgical subject according to the osteotomy path. The surgical navigation system tracks the robotic arm in real time and ensures that its operation is consistent with the surgical plan.
[0175] Step S6: Force Feedback Control
[0176] A real-time force sensing mechanism is introduced at the robot's execution end to realize an integrated collaborative control strategy of perception, control, and constraint. A speed regulation law based on the contact reaction force between the lever input and the bone surface is constructed, supplemented by a virtual impedance feedback model established from the danger zone, to dynamically adjust the end-effector speed and posture according to the intraoperative state.
[0177] Considering the complex structure of the maxillofacial region and the high-risk factors such as proximity to important nerves and blood vessels, this invention introduces a force feedback mechanism at the robot's execution end to achieve an integrated collaborative control strategy of perception, control, and constraint. Control modes are divided into H mode (manual), R mode (automatic), and HR collaborative mode.
[0178] In collaborative mode, the robot adjusts its speed and direction of motion based on real-time force signals.
[0179] Speed control law: ;
[0180] in, Main control input force, For bone surface contact reaction force, This is the adjustment coefficient. The control objective is to make the end-effector movement speed proportional to the tactile feedback.
[0181] Virtual reaction force modeling: The system constructs a virtual impedance field when approaching the dangerous anatomical area, and the virtual force is calculated as follows: ;
[0182] in, and Represent the danger membership functions for position and velocity, respectively. .
[0183] The final control model introduces damping. With inertial parameters Construct a complete feedback loop:
[0184] ;
[0185] in, , For feedback gain coefficient, For virtual reaction force, , These are terminal acceleration and velocity, respectively.
[0186] Example 2 Intelligent Control System for Oral and Maxillofacial Osteotomy Robot
[0187] like Figure 5 As shown, this embodiment provides an intelligent control system for an oral craniofacial osteotomy robot, which includes: an image input module, a structure recognition module, an osteotomy path generation module, a navigation space coordinate transformation module, and a force feedback control module.
[0188] The image input module is used to acquire and preprocess whole-head CT images of the surgical subject.
[0189] The structure recognition module is used to input semantic instructions with the help of a large language model, and to perform instance segmentation in collaboration with a pre-built cross-modal data conversion interface and a pre-built instance segmentation model. It automatically outputs an anatomical structure mask that matches the original CT image space of the surgical area corresponding to the selected osteotomy procedure type, and is mapped to the semantic coordinates in the surgical procedure language description in real time. The large language model generates text instructions based on natural language prompts and related task instructions. The text instructions include the robotic arm's movements, the target object and its target position.
[0190] The structure recognition module is used to implement the following steps 1-4:
[0191] Step 1: Constructing a procedure-structure aligned dataset: Collect whole-head CT images of several patients, and perform normalization, artifact removal, and voxel resampling preprocessing. Use medical image software to annotate the anatomical structures of the procedure-related regions, and create instance masks for the surgical areas of specific osteotomy procedures. This constitutes a complete procedure-structure aligned dataset. Each type of surgery's procedure-structure aligned dataset includes a reconstructed model, procedure label, instance index, and 3D coordinate information. The osteotomy procedure is selected from any one of the following: zygomatic L-shaped osteotomy, maxillary Lefort I-shaped osteotomy, mandibular sagittal split osteotomy, and genioplasty.
[0192] Step 2: Construct an instance segmentation model based on the formula-structure alignment dataset: The model adopts an encoder-decoder structure, with 4 encoding modules and 4 decoding modules. Residual connections are used to enhance the ability of shallow spatial information to supplement deep abstract semantics. During model training, the input 3D image is set as... The label mask is The network learning objective is to generate a prediction mask. ,in This represents a parameterized 3D U-Net function, optimized using the Tversky loss function:
[0193] ;
[0194] in: and Let i be the i-th voxel representing the true value and the predicted value, respectively; , ;
[0195] Training was performed using the Adam optimizer with an initial learning rate of The gradient decays to half its original value every 50 rounds, with a maximum iteration count of 300 rounds; each training round employs a gradient accumulation mechanism with a batch size of 1. ;
[0196] Four enhancement strategies—random rotation, random scaling, mirror flipping, and Gaussian noise perturbation—are employed to dynamically combine with set probabilities and are simultaneously applied to the image and corresponding label mask of each training sample. Perturbation samples are dynamically generated during training.
[0197] (1) In terms of spatial angle perturbation, a three-dimensional Euler angle rotation mechanism is used to randomly rotate the input image voxel data;
[0198] The 3D input image is The rotation angles are as follows:
[0199] ,
[0200] in, ,
[0201] Angle rotation operation is defined as:
[0202] ;
[0203] This operation simulates the changes in the patient's head posture during surgery, improving the model's ability to recognize changes in osteotomy orientation.
[0204] (2) Introduce anisotropic scaling strategy in terms of structural scale: Let the scaling factor be ,in The scaling perturbation operation is defined as follows:
[0205] , ;
[0206] By adjusting the bone scale to simulate the physiological structural differences between individuals, the model's segmentation ability under different bone sizes is enhanced.
[0207] (3) To address the issue of left-right structural symmetry, a three-dimensional mirror flipping mechanism is introduced: In In terms of probability in three dimensions Perform an axial mirror flip operation; the transformation expression is:
[0208] , ;
[0209] This enhancement method improves the model's classification neutrality on symmetric structures and reduces the risk of overfitting to skewed structures in the training set;
[0210] (4) Introduce Gaussian noise perturbation at the signal level This is used to simulate signal drift caused by dose changes and differences in reconstruction algorithms during CT scanning; the perturbation follows a normal distribution.
[0211] , ;
[0212] in The noise intensity is expressed in original grayscale values; this mechanism improves the model's stability and response consistency to input images under different imaging conditions.
[0213] During the training process, the above enhancement strategies are dynamically combined with set probabilities and applied simultaneously to the image and the corresponding label mask; the complete enhancement process is expressed as follows:
[0214] ;
[0215] The operators correspond to Gaussian noise addition, mirror flip, scaling perturbation and angle rotation operations in turn, ensuring that the spatial geometry and semantic structure remain consistent in the perturbed image;
[0216] Step 3: Based on the constructed cross-modal data structure conversion module, the labeled data and the large language model are adapted, and the anatomical structure labels and surgical language descriptions are embedded and fused. The cross-modal data structure conversion module supports "osteotomy type + image region + structural semantics" as joint input.
[0217] The construction methods for the cross-modal data structure transformation module include:
[0218] 1) Establishment of a medical prompt database: Collect a large number of real surgical records and standard surgical procedures, and extract structural terms;
[0219] 2) Tag mapping mechanism: Map each structural tag to a standard technical description language, and adjust the token structure through multiple rounds of iteration to improve its semantic embedding quality in LLM;
[0220] 3) Embedding fusion and task adaptation: Using "osteotomy type + image region + anatomical structure semantics" as joint input, a structured embedding vector is generated to guide 3D image segmentation;
[0221] 4) Mask alignment mechanism: The final output of the model is a mask map that is spatially registered with the original CT image, which supports direct mapping to the navigation space or subsequent path encoding;
[0222] Step 4: The application of instance segmentation model, large language model and cross-modal data structure conversion module to complete the instance segmentation task of medical image: The large language model is used as semantic instruction and works with the instance segmentation model to perform instance segmentation to output the instance structure mask of the surgical area that matches the space of the original CT image and is mapped to the semantic coordinates in the surgical language description in real time. The large language model autonomously generates text instructions based on natural language prompts and relevant task instructions.
[0223] The osteotomy path generation module first detects and obtains site information through the navigation system, and simultaneously performs image encoding on the real-time observation results of the patient's space and text encoding on the text command sequence. Then, the results of image encoding and text encoding are weighted and input into the Gaussian mixture model. The Gaussian mixture model combines the site information for decoding and uses real annotations for loss calculation, and finally predicts and generates the osteotomy path trajectory points under the optical tracking system.
[0224] The osteotomy path generation module introduces a multilayer perceptron to perform Gaussian mixture modeling on the fused features of image and language encoding vectors, achieving compressed representation and sampling generation of the surgical path; the input image features... Depend on The network acts as an image encoder to extract boundary features; the Transformer architecture extracts features from the generated text instructions. Based on the known tracking tags, the three-dimensional coordinates of the target object. The trajectory is tracked and identified in the coordinate system of the optical tracking system; furthermore, a Gaussian mixture model based on a multilayer perceptron is used to model the trajectory distribution, which is based on features formed by concatenating visual latent coding and text latent coding. To study;
[0225] The loss function of this GMM is:
[0226] ;
[0227] ;
[0228] ;
[0229] Among them, parameter set Depend on Extracted from the input.
[0230] To predict the distribution of trajectory points, This represents a Gaussian distribution.
[0231] The navigation space coordinate transformation module is used to transform the osteotomy path trajectory points under the optical tracking system into the osteotomy path trajectory under the robot coordinate system through a multi-level coordinate system transformation method. The robotic arm executes operation commands on the surgical object according to the osteotomy path. The surgical navigation system tracks the robotic arm in real time and ensures that its operation is consistent with the surgical planning results.
[0232] The navigation spatial coordinate transformation module transforms the predicted trajectory points. In the coordinate system of the optical tracking system, it can be achieved through the coordinate transformation matrix. The trajectory corresponding to the robot's coordinate system is calculated as follows: The transformation matrix is also obtained by the optical tracking system.
[0233] A multi-level coordinate system transformation method was used to map the trajectory in the preoperative image space to the intraoperative robot control space. This work establishes the geometric and semantic foundation for subsequent navigation execution and force feedback control.
[0234] The system involves multiple coordinate systems, including the robotic arm base coordinate system O, the robotic arm end effector coordinate system E, the patient image guidance coordinate system I, the reference frame coordinate system R, and the surgical tool coordinate system S. The transformation matrix between the robotic arm end effector and its base is called the robotic arm posture matrix. The surgical navigation subsystem reconstructs a 3D image using the patient's preoperative CT imaging data, thereby establishing an image guidance space. Spatial registration methods are used to determine the coordinate transformation between the reference frame coordinate system and the image guidance coordinate system. Then, the spatial relationship between the surgical tools and the patient area is displayed in real time within the image guidance space. Under the guidance of the surgical navigation subsystem, robot-assisted surgical operations require managing multiple spatial coordinate transformations, i.e., achieving multi-spatial coordinate registration. The purpose of registration is to guide the robotic arm to adjust its target posture using patient image data and determine the required transformation matrix. This represents the transformation from coordinate system E to coordinate system O. Once the target pose is determined based on the patient's image-guided data, the robotic arm is controlled to precisely align the surgical tools on the end effector along the surgical path.
[0235] To enable real-time tracking of the robotic arm via a surgical navigation system and ensure its operation aligns with the surgical plan, a target pose matrix for the robotic arm is required. It is a key component of system design. Key transformation matrices include:
[0236] The transformation matrix from the image guiding coordinate system I to the reference coordinate system R;
[0237] : The transformation matrix from the reference frame coordinate system R to the robot arm base coordinate system O;
[0238] The transformation matrix from the image guidance coordinate system I to the surgical tool coordinate system S;
[0239] : The transformation matrix from the surgical instrument coordinate system S to the robotic arm end effector coordinate system E;
[0240] : The pose matrix required by the robotic arm.
[0241] Each transformation matrix T consists of a rotation matrix and a translation matrix. To compute these transformation matrices, this study employs a point-based least squares fitting method, particularly singular value decomposition (SVD).
[0242] The transformation relationship between the image guidance space and the robotic arm base is as follows:
[0243] ;
[0244] ;
[0245] The reference frame and robotic arm base are fixed, therefore they can be directly obtained. Spatial transformation matrix It is the spatial transformation matrix that maps the image guidance space to the reference frame space, determined by a combination of point registration and out-of-plane registration methods. By moving the surgical instrument and the robotic arm end effector to the same spatial position and determining the coordinate transformation relationship between the tip of the instrument and the center of the robotic arm end effector, the following can be obtained: By using an optical tracker to track reflective marks on surgical instruments, one can obtain... .
[0246] Therefore, the robot arm's posture can be calculated as follows: .
[0247] A real-time force sensing mechanism is introduced at the robot's actuator end to realize an integrated sensing-control-constraint collaborative control strategy. A velocity regulation law based on the contact reaction force between the lever input and the bone surface is constructed, supplemented by a virtual impedance feedback model established from the danger zone, dynamically adjusting the end effector speed and posture according to the intraoperative state. In the force feedback mechanism collaborative control mode, the force feedback control module adjusts the motion speed and direction based on real-time force signals.
[0248] Speed control law: ;
[0249] in, Main control input force, For bone surface contact reaction force, The adjustment coefficient is used to control the end effector's movement speed in a proportional manner to the tactile feedback.
[0250] Virtual reaction force modeling: The system constructs a virtual impedance field when approaching the dangerous anatomical area, and the virtual force is calculated as follows: ;
[0251] in, and Represent the danger membership functions for position and velocity, respectively. ;
[0252] The final control model introduces damping. With inertial parameters Construct a complete feedback loop:
[0253] ;
[0254] in, , For feedback gain coefficient, For virtual reaction force, , These are terminal acceleration and velocity, respectively.
[0255] This system supports a graphical user interface, allowing surgeons to view the osteotomy path, estimate bone density, and receive feedback force values in real time. Thirty experiments were conducted on a 3D-printed craniofacial model to verify the system's repeatability and stability along the Le Fort I osteotomy path, with a force feedback acquisition frequency of 200Hz. Results showed a path accuracy error of 0.65 mm, meeting clinical requirements.
[0256] Example 3: Computer Equipment
[0257] A computer device is characterized by comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed by the processor, the computer program implements the intelligent control method for a craniofacial osteotomy robot as described in Embodiment 1. For simplicity, further details are omitted here. It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The memory can include read-only memory and random access memory (RAM), providing instructions and data to the processor. A portion of the memory can also include non-volatile random access memory (RAM). For example, the memory can also store device type information.
[0258] Example 4: Computer-readable storage medium
[0259] A computer-readable storage medium is characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the intelligent control method of the oral craniofacial osteotomy robot as described in Embodiment 1. For simplicity, further details are omitted here. The computer-readable storage medium can be a tangible device capable of holding and storing instructions used by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusion structures storing instructions thereon, and any suitable combination thereof. The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0260] Example 5: Oral and Maxillofacial Osteotomy Robot
[0261] like Figure 6 As shown, this embodiment provides an oral craniofacial osteotomy robot, which includes surgical instruments, a robotic arm, and a control module. The control module is used to execute the intelligent control method of the oral craniofacial osteotomy robot as described in Embodiment 1, which will not be elaborated further here for simplicity. The robotic arm is used to hold the surgical instruments and perform osteotomy operations on the surgical subject according to the control commands output by the control module.
[0262] It should be noted that the embodiments of the present invention have better implementability and are not intended to limit the present invention in any way. Any person skilled in the art may use the above-disclosed technical content to change or modify it into equivalent effective embodiments. However, any modifications or equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of the technical solution of the present invention.
Claims
1. An intelligent control method for an oral craniofacial osteotomy robot, characterized in that, include: Step S1: Image Input: Acquire whole-head CT images of the surgical subject and perform preprocessing; Step S2: Structure Recognition: Using a large language model to input semantic instructions, and using a pre-built cross-modal data conversion interface and a pre-built instance segmentation model to perform instance segmentation, an anatomical structure mask matching the original CT image space of the surgical area corresponding to a specific osteotomy procedure type is automatically output and mapped to the semantic coordinates in the surgical procedure language description in real time. The large language model generates text instructions based on natural language prompts and related task instructions. The text instructions include the robotic arm's movements, the target object and its target position. Step S3: Osteotomy Path Generation: First, the site information is detected and acquired through the navigation system. At the same time, the spatial coordinates obtained from real-time observation of the patient's space are image encoded based on binocular visual tracing, and the text command sequence is text encoded. Then, the results of image encoding and text encoding are weighted and input into the Gaussian mixture model. The Gaussian mixture model is decoded in combination with the site information and loss is calculated using real annotations. Finally, the osteotomy path trajectory points under the optical tracking system are predicted and generated. Step S4: Navigation space coordinate transformation: The osteotomy path trajectory points under the optical tracking system are calculated to obtain the osteotomy path trajectory under the robot coordinate system through the coordinate transformation matrix.
2. The intelligent control method for the oral craniofacial osteotomy robot as described in claim 1, characterized in that, In step S2: Structure identification includes the following steps: Step S2.1: Constructing a procedure-structure aligned dataset: Collect whole-head CT images of several patients, and perform normalization, artifact removal, and voxel resampling preprocessing. Use medical image software to annotate the anatomical structures of the procedure-related regions, and create instance masks for the surgical areas of specific osteotomy procedures to form a complete procedure-structure aligned dataset. Each type of surgery's procedure-structure aligned dataset includes a reconstructed model, procedure label, instance index, and three-dimensional coordinate information. The osteotomy procedure is selected from any one of the following: zygomatic L-shaped osteotomy, maxillary Lefort I-shaped osteotomy, mandibular sagittal split osteotomy, genioplasty, high condylar resection, and segmental mandibular osteotomy. Step S2.2: Construct an instance segmentation model based on the formula-structure alignment dataset: The model adopts an encoder-decoder structure, with 4 encoding modules and 4 decoding modules. Residual connections are used to enhance the ability of shallow spatial information to supplement deep abstract semantics. During model training, the input 3D image is set as... The label mask is The network learning objective is to generate a prediction mask. ,in This represents a parameterized 3D U-Net function, optimized using the Tversky loss function: ; in: and Let i be the i-th voxel representing the true value and the predicted value, respectively; , ; Training was performed using the Adam optimizer with an initial learning rate of The gradient decays to half its original value every 50 rounds, with a maximum iteration count of 300 rounds; each training round employs a gradient accumulation mechanism with a batch size of 1. ; Four enhancement strategies—random rotation, random scaling, mirror flipping, and Gaussian noise perturbation—are employed to dynamically combine with set probabilities and are simultaneously applied to the image and corresponding label mask of each training sample. Perturbation samples are dynamically generated during training. (1) In terms of spatial angle perturbation, a three-dimensional Euler angle rotation mechanism is used to randomly rotate the input image voxel data; The 3D input image is The rotation angles are as follows: , in, , Angle rotation operation is defined as: ; This operation simulates the changes in the patient's head posture during surgery, improving the model's ability to recognize changes in osteotomy orientation. (2) Introduce anisotropic scaling strategy in terms of structural scale: Let the scaling factor be ,in The scaling perturbation operation is defined as follows: , ; By adjusting the bone scale to simulate the physiological structural differences between individuals, the model's segmentation ability under different bone sizes is enhanced. (3) To address the issue of left-right structural symmetry, a three-dimensional mirror flipping mechanism is introduced: In In terms of probability in three dimensions Perform an axial mirror flip operation; the transformation expression is: , ; This enhancement method improves the model's classification neutrality on symmetric structures and reduces the risk of overfitting to skewed structures in the training set; (4) Introduce Gaussian noise perturbation at the signal level This is used to simulate signal drift caused by dose changes and differences in reconstruction algorithms during CT scanning; the perturbation follows a normal distribution. , ; in The noise intensity is expressed in original grayscale values; this mechanism improves the model's stability and response consistency to input images under different imaging conditions. During the training process, the above enhancement strategies are dynamically combined with set probabilities and applied simultaneously to the image and the corresponding label mask; the complete enhancement process is expressed as follows: ; The operators correspond to Gaussian noise addition, mirror flip, scaling perturbation and angle rotation operations in turn, ensuring that the spatial geometry and semantic structure remain consistent in the perturbed image; Step S2.3: Based on the constructed cross-modal data structure conversion module, the labeled data and the large language model are adapted, and the anatomical structure labels and surgical language descriptions are embedded and fused. The cross-modal data structure conversion module supports "osteotomy type + image region + structural semantics" as joint input. The construction methods for the cross-modal data structure transformation module include: 1) Establishment of a medical prompt database: Collect a large number of real surgical records and standard surgical procedures, and extract structural terms; 2) Tag mapping mechanism: Map each structural tag to a standard technical description language, and adjust the token structure through multiple rounds of iteration to improve its semantic embedding quality in LLM; 3) Embedding fusion and task adaptation: Using "osteotomy type + image region + anatomical structure semantics" as joint input, a structured embedding vector is generated to guide 3D image segmentation; 4) Mask alignment mechanism: The final output of the model is a mask map that is spatially registered with the original CT image, which supports direct mapping to the navigation space or subsequent path encoding; Step S2.4: The application of the instance segmentation model, the large language model, and the cross-modal data structure conversion module collaboratively completes the instance segmentation task of medical images: The large language model is used as a semantic instruction and collaborates with the instance segmentation model to perform instance segmentation to output an instance structure mask of the surgical area that matches the space of the original CT image and is mapped to the semantic coordinates in the surgical language description in real time. The large language model autonomously generates text instructions based on natural language prompts and relevant task instructions.
3. The intelligent control method for the oral craniofacial osteotomy robot as described in claim 1, characterized in that, In step S3, a multilayer perceptron is introduced to perform Gaussian mixture modeling on the fused features of the image and language coding vectors, thereby achieving compressed representation and sampling generation of the surgical path; the input image features Depend on The network acts as an image encoder to extract boundary features; Features of text instructions extracted by the Transformer architecture ; Based on the known tracking tags, the three-dimensional coordinates of the target object. The trajectory is tracked and identified in the coordinate system of the optical tracking system; furthermore, a Gaussian mixture model based on a multilayer perceptron is used to model the trajectory distribution, which is based on features formed by concatenating visual latent coding and text latent coding. To study; The loss function of this GMM is: ; ; ; Among them, parameter set Depend on Extracted from the input. To predict the distribution of trajectory points, This represents a Gaussian distribution.
4. The intelligent control method for the oral craniofacial osteotomy robot as described in claim 1, characterized in that, In step S4, based on the predicted trajectory points In the coordinate system of the optical tracking system, it can be achieved through the coordinate transformation matrix. The trajectory corresponding to the robot's coordinate system is calculated as follows: The transformation matrix is also obtained by the optical tracking system.
5. The intelligent control method for the oral craniofacial osteotomy robot as described in claim 1, characterized in that, Further steps include: Step S5: Osteotomy path execution: The robotic arm executes operation instructions on the surgical object according to the osteotomy path, and the robotic arm is tracked in real time through the surgical navigation system to ensure that its operation is consistent with the surgical planning results; The target pose matrix of the robotic arm is: ; in, : The target pose matrix of the robotic arm; The transformation matrix from the reference frame coordinate system R to the robot arm base coordinate system O; The transformation matrix from the image guiding coordinate system I to the reference coordinate system R; : The transformation matrix from the image guidance coordinate system I to the surgical tool coordinate system S; : The transformation matrix from the surgical instrument coordinate system S to the robotic arm end effector coordinate system E.
6. The intelligent control method for the oral craniofacial osteotomy robot as described in claim 1, characterized in that, Further steps include S6: Force feedback control: Introducing a real-time force sensing mechanism at the robot's execution end to realize an integrated collaborative control strategy of sensing-control-constraint, and constructing a speed regulation law based on the reaction force between the lever input and the bone surface, supplemented by a virtual impedance feedback model established from the danger zone, dynamically adjusting the end-effector speed and posture according to the intraoperative state; In the force feedback mechanism cooperative control mode, the robot adjusts its speed and direction of motion based on real-time force signals: Speed control law: ; in, Main control input force, For bone surface contact reaction force, The adjustment coefficient is used to control the end effector's movement speed in a proportional manner to the tactile feedback. Virtual reaction force modeling: The system constructs a virtual impedance field when approaching the dangerous anatomical area, and the virtual force is calculated as follows: ; in, and Represent the danger membership functions for position and velocity, respectively. ; The final control model introduces damping. With inertial parameters Construct a complete feedback loop: ; in, , For feedback gain coefficient, For virtual reaction force, , These are terminal acceleration and velocity, respectively.
7. An intelligent control system for an oral and maxillofacial osteotomy robot, characterized in that, include: The image input module is used to acquire and preprocess whole-head CT images of the surgical subject. The structure recognition module is used to input semantic instructions with the help of a large language model, and to perform instance segmentation in collaboration with a pre-built cross-modal data conversion interface and a pre-built instance segmentation model. It automatically outputs an anatomical structure mask that matches the original CT image space of the surgical area corresponding to the selected osteotomy procedure type and is mapped to the semantic coordinates in the surgical procedure language description in real time. The large language model generates text instructions based on natural language prompts and related task instructions. The text instructions include the robotic arm's movements, the target object and its target position. The osteotomy path generation module first detects and obtains site information through the navigation system, and simultaneously performs image encoding on the real-time observation results of the patient space and text encoding on the text command sequence. Then, the results of image encoding and text encoding are weighted and input into the Gaussian mixture model. The Gaussian mixture model decodes the site information and uses real annotations to calculate the loss, and finally predicts and generates the osteotomy path trajectory points under the optical tracking system. The navigation space coordinate transformation module is used to calculate the osteotomy path trajectory in the robot coordinate system by using a coordinate transformation matrix to transform the osteotomy path trajectory points under the optical tracking system.
8. A computer device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the intelligent control method for an oral craniofacial osteotomy robot as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the intelligent control method for the oral craniofacial osteotomy robot as described in any one of claims 1 to 6.
10. A craniofacial osteotomy robot, characterized in that, The oral craniofacial osteotomy robot includes surgical instruments, a robotic arm, and a control module; the control module is used to execute the intelligent control method of the oral craniofacial osteotomy robot as described in any one of claims 1 to 6; the robotic arm is used to hold the surgical instruments and perform osteotomy operations on the surgical object according to the control instructions output by the control module.