Instrument pose control method and system based on four-direction pose optical acquisition

By constructing a dynamic three-dimensional digital twin through four-way pose optical acquisition and multi-source data fusion, combined with spatiotemporal networks and adaptive impedance control, the problems of difficult instrument position judgment and operational tremor in traditional minimally invasive surgery are solved, achieving high-precision, safe and accurate instrument pose control.

CN122005083APending Publication Date: 2026-05-12ANHUI MIDU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI MIDU INTELLIGENT TECH CO LTD
Filing Date
2026-02-27
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In traditional minimally invasive surgery, the surgeon's experience and manual manipulation are relied upon. The two-dimensional images provided by the endoscope lack depth information, making it difficult to determine the position of the instruments relative to the patient's anatomical structures, which poses a risk of collision. Furthermore, prolonged surgery can lead to tremors that affect surgical precision and increase the probability of complications.

Method used

A four-way pose optical acquisition method is adopted to construct a dynamic three-dimensional digital twin through multi-source data fusion. Combined with spatiotemporal network and adaptive impedance control, the relative positional relationship between instruments and anatomical structures is reconstructed in real time to generate a safety boundary field. The instrument operation is assisted by AR guidance and impedance controller.

Benefits of technology

It achieves high-precision, forward-looking, and individualized control of instrument positioning, reduces the risk of collision between instruments and key anatomical structures, improves the safety and precision of surgery, and reduces the impact of surgeon fatigue and tremors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122005083A_ABST
    Figure CN122005083A_ABST
Patent Text Reader

Abstract

The invention discloses an instrument pose control method and system based on four-direction pose optical acquisition. The method comprises the following steps: synchronously obtaining multi-source initial data; reconstructing a dynamic three-dimensional digital twinborn body in real time in a virtual space; outputting an operation intention prediction result through a space-time network, resolving to obtain a predicted pose sequence of the instrument in a future preset time window, and accurately sensing a real-time relative position relationship between the instrument and the key anatomical structure; constructing a dynamically updated security boundary field; when it is predicted that the instrument will invade the safety boundary field at the future moment or the real-time force feedback data are abnormal, the self-adaptive impedance controller is started, impedance parameters of the self-adaptive impedance controller are dynamically adjusted, and an auxiliary guiding force instruction or a resistance instruction is generated. Through deep coupling of multi-modal sensing, digital twinborn reconstruction, space-time prediction, dynamic boundary construction, adaptive control and online learning, the safety, accuracy and accessibility of a complex surgical operation are fundamentally improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical device technology, and in particular to a device pose control method and system based on four-way pose optical acquisition. Background Technology

[0002] In the field of surgery, especially minimally invasive surgery, the precise control of instrument positioning is directly related to the safety and effectiveness of the procedure. Traditional minimally invasive surgery relies on the surgeon's experience and hand dexterity, using an endoscope to observe the surgical area and manipulate instruments to complete the operation. However, this approach has several limitations: firstly, the two-dimensional images provided by the endoscope lack depth information, making it difficult for the surgeon to accurately determine the relative position of the instruments to the patient's anatomical structures, increasing the risk of instruments colliding with critical anatomical structures (such as blood vessels and nerves); secondly, prolonged surgery can lead to hand fatigue and tremors, affecting surgical precision and increasing the probability of surgical complications. Summary of the Invention

[0003] The purpose of this invention is to provide a device pose control method and system based on four-way pose optical acquisition, which can solve the above-mentioned problems existing in the prior art.

[0004] To achieve the above objectives, this application adopts the following technical solution: On the one hand, a method for controlling the pose of a device based on four-way pose optical acquisition is provided, which includes: Step S10: Simultaneously acquire multi-source initial data, perform fusion processing on the multi-source initial data, and generate a multi-modal fusion data package; the multi-source initial data includes: two-dimensional RGB images and three-dimensional point cloud data of the surgical area and instruments synchronously acquired by four optical acquisition devices along orthogonal directions, nine-axis inertial data of the instruments acquired by the micro inertial measurement unit, and force feedback data during the interaction between the instruments and tissues acquired by the six-dimensional force / torque sensor mounted on the end of the instruments; Step S20: Based on the multimodal fusion data package, a dynamic three-dimensional digital twin is reconstructed in real time in virtual space. The dynamic three-dimensional digital twin includes surgical instruments, patient anatomical structures, and the interaction between the two. The physics engine is invoked to simulate the mechanical interaction process between the instruments and virtual soft tissue within the dynamic three-dimensional digital twin, and the mechanical interaction simulation results are output. Step S30: Input the state sequence of the dynamic three-dimensional digital twin, the historical motion trajectory of the instrument, and the real-time force feedback data into the pre-trained spatiotemporal network; output the operation intention prediction result through the spatiotemporal network, and simultaneously calculate the predicted pose sequence of the instrument within a future preset time window, and accurately perceive the real-time relative positional relationship between the instrument and key anatomical structures. Step S40: Based on the predicted pose sequence obtained in step S30 and the perceived real-time relative position relationship, a dynamically updated safety boundary field is constructed. The safety boundary field is used to define the safe range of the movement of the device. Step S50: When it is predicted that the device will intrude into the safety boundary field in the future, or when there is an anomaly in the real-time force feedback data, the adaptive impedance controller is activated and its impedance parameters are dynamically adjusted to generate auxiliary guiding force commands or resistance commands; the auxiliary guiding force commands or resistance commands work together with the doctor's operating force to jointly drive the device to perform posture adjustment.

[0005] Preferably, in step S10, the synchronous acquisition of the four optical acquisition devices is achieved through hardware-triggered synchronization, and the two-dimensional RGB image and three-dimensional point cloud data are preprocessed including denoising, distortion correction and coordinate system one.

[0006] Preferably, in step S20, the mechanical model of the virtual soft tissue is constructed based on the finite element method, and the initial values ​​of the parameters of the mechanical model are generated based on the patient's preoperative medical imaging data, including CT images or MRI images.

[0007] Preferably, after step S50, step S60 is further included: During the surgery, the deviation between the predicted interactive force of the digital twin and the actual force feedback sensor data is continuously compared. When the deviation exceeds a preset threshold, a lightweight online learning loop is triggered to fine-tune the tissue mechanics parameters of some layers of the spatiotemporal network or the physical engine using the current surgical data, so that the digital twin can dynamically adapt to the individualized tissue characteristics of the current patient.

[0008] Preferably, the spatiotemporal network adopts an encoder-decoder structure, wherein the encoder is used to process the state sequence of the digital twin, and the decoder is used to output a predicted sequence based on the real-time operation signal of the instrument as the query vector and the most relevant anatomical structure and instrument component for the current operation.

[0009] Preferably, in step S50, the impedance parameter adjustment of the adaptive impedance controller includes: When a critical anatomical structure is predicted to be approaching, the damping of the guidance direction is progressively increased, and a virtual force field channel is rendered in the display module of the virtual space using augmented reality. When high-frequency vibrations or sudden increases are detected in the force feedback data, the stiffness of the guidance direction is instantly reduced and a high-frequency filter is activated to actively suppress unstable operation.

[0010] Preferably, in step S60, a meta-learning optimizer pre-trained on a large amount of offline surgical data is used to obtain initial values ​​of model parameters that can quickly adapt to new tasks, thereby enabling the fine-tuning process to converge quickly with limited real-time surgical data.

[0011] On the other hand, this disclosure also provides a device pose control system based on four-way pose optical acquisition for implementing the method described above, comprising: The multimodal sensing module includes four high-frame-rate optical cameras, an IMU sensor built into the instrument, and a six-dimensional force / torque sensor at the end of the instrument. The digital twin computing engine module is used to run the physics engine and real-time 3D reconstruction algorithms; The intelligent sensing and prediction module has a built-in pre-trained spatiotemporal network; A collaborative control module is used to implement a forward-looking adaptive impedance control algorithm; and The human-machine collaborative interaction module includes a force feedback device and an AR head-mounted display device, which are used to receive doctor's instructions and visualize the digital twin and safety boundaries.

[0012] Preferably, the system supports collaborative operation of multiple instruments, and the digital twin computing engine module independently constructs a sub-digital twin for each instrument, and manages the motion rules and anti-collision logic between instruments through a central coordinator.

[0013] Preferably, the instrument in the human-machine collaborative interaction module can simulate the tactile sensation of virtual boundaries, the force changes of tissue incision, and the tactile feedback of the interaction between the instrument and virtual objects such as sutures, according to the instructions of the adaptive impedance controller.

[0014] The beneficial effects of this application are as follows: By employing a four-way orthogonal arrangement of optical cameras, combined with the instrument's built-in IMU and end-effector force sensor, a multimodal sensing network with no blind spots was constructed. Tightly coupled fusion of multi-source data enabled continuous, high-precision tracking of the instrument's six-degree-of-freedom pose, and simultaneously acquired its interaction with the tissue's mechanical state, providing the system with comprehensive and reliable real-world data mapping.

[0015] Meanwhile, the digital twin, reconstructed in real time based on multimodal data, is a dynamic virtual mirror of the actual surgical scene. Integrating a personalized physics engine based on the finite element method, the twin can simulate the biomechanical interactions between instruments and soft tissue, providing a "digital testbed" with both geometric and physical realism for prediction and control.

[0016] By analyzing the state and historical data of digital twins through spatiotemporal networks, the future position and operational intent of instruments can be predicted. A dynamic safety boundary field constructed using this predictive information extends safety monitoring from the current moment to a future time window, enabling risk warnings. Combined with an AR visualization channel, this forms a multi-sensory integrated intelligent guidance system, significantly optimizing the doctor's operational intuition and fluency while ensuring safety limits.

[0017] Through the deep coupling of multimodal perception, digital twin reconstruction, spatiotemporal prediction, dynamic boundary construction, adaptive control, and online learning, a self-iteratory and continuously optimized intelligent control closed loop is formed. This transforms surgical instruments from passive tools requiring full physician control into intelligent decision-making systems capable of understanding scenarios, predicting risks, and providing compliant assistance, fundamentally improving the safety, precision, and accessibility of complex surgical procedures. Attached Figure Description

[0018] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.

[0019] Figure 1 This is a schematic flowchart of a device pose control method based on four-way pose optical acquisition according to an embodiment of this application; Figure 2 This is a schematic diagram of a device pose control system based on four-way pose optical acquisition according to an embodiment of this application.

[0020] In the picture: 100. Multimodal perception module; 110. Digital twin computing engine module; 120. Intelligent perception and prediction module; 130. Collaborative control module; 140. Human-machine collaborative interaction module. Detailed Implementation

[0021] To make the technical problems solved by this application, the technical solutions adopted, and the technical effects achieved clearer, the technical solutions of the embodiments of this application are further described in detail below. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] In the description of this application, unless otherwise expressly specified and limited, the terms "connected," "linked," and "fixed" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0023] In this application, unless otherwise expressly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature being directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature being directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.

[0024] Please see Figure 1 and Figure 2 This disclosure provides a method and system for instrument pose control based on four-way pose optical acquisition. This method achieves high-precision, forward-looking, and individualized pose control of surgical instruments through four-way orthogonal optical acquisition, multi-source data fusion, digital twin construction, spatiotemporal network prediction, dynamic safety boundary construction, and adaptive impedance control, thereby improving the safety and effectiveness of surgery.

[0025] Furthermore, the method provided in this disclosure is based on a modular instrument pose control system using four-way pose optical acquisition. The modules communicate with each other via high-speed data lines or networks to ensure real-time and reliable data transmission. The system includes a multimodal perception module 100, a digital twin computing engine module 110, an intelligent perception and prediction module 120, a collaborative control module 130, and a human-machine collaborative interaction module 140.

[0026] In one embodiment, the multimodal sensing module 100 is the core of the system's data acquisition, responsible for synchronously acquiring various types of data during the surgical process, providing data support for subsequent digital twin construction, pose prediction, and control decisions. This module includes four high-frame-rate optical cameras, an instrument-embedded IMU sensor, and a six-dimensional force / torque sensor at the instrument's end effector.

[0027] For example, four high frame rate optical cameras are installed in an orthogonal layout, with the specific positions determined according to the type of surgery and the characteristics of the surgical area. For instance, in laparoscopic surgery, the four optical cameras can be installed in the front, back, left, and right directions of the operating table to ensure that image data of the surgical area and instruments can be acquired simultaneously from four orthogonal perspectives.

[0028] It is important to note that the parameters of the optical camera must meet the surgical requirements. It is preferable to use an industrial-grade high frame rate camera with a frame rate of no less than 120fps and a resolution of no less than 1920×1080 to ensure that the rapid movement of instruments can be clearly captured.

[0029] For example, an IMU sensor is built into the shaft of the surgical instrument to collect nine-axis inertial data, including three-axis acceleration, three-axis angular velocity, and three-axis magnetic field strength. The IMU sensor has a sampling rate of no less than 200Hz to ensure accurate capture of subtle motion changes in the instrument. A six-dimensional force / torque sensor is installed at the end of the instrument to collect force feedback data during the interaction between the instrument and patient tissue, including forces along the X, Y, and Z axes and torques about the X, Y, and Z axes. The sensor's range and accuracy are determined according to the type of surgery; for example, in neurosurgery, a six-dimensional force / torque sensor with a range of 0-50N and an accuracy of 0.01N can be selected.

[0030] Furthermore, to ensure time synchronization of multi-source data, the multimodal sensing module 100 adopts a hardware-triggered synchronization method, controlling four optical cameras, an IMU sensor, and a six-dimensional force / torque sensor to simultaneously acquire data via a synchronization trigger signal. The acquired raw data is transmitted to the data preprocessing unit through a high-speed data transmission interface for preprocessing operations such as noise reduction, distortion correction, coordinate system unification, and data format conversion.

[0031] In one embodiment, the digital twin computing engine module 110 can be deployed on a high-speed edge server, responsible for real-time reconstruction of a dynamic 3D digital twin based on multimodal fusion data packets, and simulating the mechanical interaction process between instruments and virtual soft tissue through a physics engine. This module includes a real-time 3D reconstruction algorithm module, a physics engine module, and a data storage module.

[0032] For example, the real-time 3D reconstruction algorithm module employs a deep learning-based 3D reconstruction algorithm, combining 2D RGB images and 3D point cloud data acquired by four optical acquisition devices, as well as the patient's preoperative medical imaging data (such as CT or MRI images), to reconstruct a dynamic 3D digital twin in virtual space in real time. During the reconstruction process, the patient's anatomical structure information (such as organs, blood vessels, nerves, etc.) is first extracted from the preoperative medical imaging data using an image segmentation algorithm to construct an initial 3D model of the anatomical structure; then, based on the optical data acquired in real time during the operation, the initial 3D model is dynamically updated to ensure that the digital twin remains consistent with the actual surgical scene.

[0033] For example, the physics engine module uses the finite element method to construct a virtual soft tissue mechanical model to simulate the mechanical interaction between the instrument and the virtual soft tissue. The initial values ​​of the mechanical model's parameters are generated based on the patient's preoperative medical imaging data; for example, the density and elastic modulus of different tissues are determined by analyzing the grayscale distribution of CT images. During the surgery, the physics engine module calculates the interaction forces between the instrument and the virtual soft tissue in real time based on the force feedback data and instrument pose data in the multimodal fusion data package, and outputs the mechanical interaction simulation results.

[0034] For example, the data storage module is used to store various types of data during the surgical procedure, including raw acquired data, preprocessed data, multimodal fusion data packages, digital twin state data, and mechanical interaction simulation results. The data storage can employ a distributed storage architecture to ensure data security and scalability, while supporting rapid data retrieval and providing data support for subsequent online learning and model optimization.

[0035] In one embodiment, the intelligent sensing and prediction module 120 serves as the system's decision-making unit, responsible for predicting operational intentions, calculating the predicted pose sequence of instruments, and sensing the real-time relative positional relationship between instruments and key anatomical structures through a spatiotemporal network. This module mainly includes a spatiotemporal network module, a data preprocessing module, and a prediction result output module.

[0036] For example, the spatiotemporal network module employs an encoder-decoder structure. The encoder processes the state sequence of the digital twin, while the decoder uses the real-time operation signals of the instruments as query vectors and outputs a predicted sequence based on the most relevant anatomical structures and instrument components for the current operation. The pre-training process of the spatiotemporal network is based on a large amount of offline surgical data, including surgical videos of different surgical types and patients, instrument motion trajectory data, force feedback data, etc. Through pre-training, the spatiotemporal network can learn the instrument motion patterns under different surgical operation scenarios and the mapping relationship between operational intentions and pose sequences.

[0037] Furthermore, the data preprocessing module is used to preprocess the data input to the spatiotemporal network, including operations such as data normalization, data augmentation, and sequence length unification. Specifically, the state sequence of the digital twin, the historical motion trajectory of the device, and the real-time force feedback data are converted into a unified data format and normalized to ensure that the data are within the same numerical range, avoiding the impact of data scale differences on the network's prediction accuracy. At the same time, data augmentation techniques (such as time axis shifting and noise addition) can be used to improve the network's generalization ability. Finally, sequence data of different lengths are unified into a fixed length to meet the input requirements of the spatiotemporal network.

[0038] Furthermore, the prediction result output module is used to parse and post-process the output results of the spatiotemporal network to obtain the operation intention prediction result, the predicted pose sequence of the instrument within a future preset time window, and the real-time relative positional relationship between the instrument and key anatomical structures. The length of the preset time window can be adjusted according to the type of surgery and the difficulty of the operation. For example, in delicate operation scenarios, the preset time window length can be set to 0.5 seconds to 1 second to ensure that the movement trend of the instrument can be predicted in a timely manner.

[0039] In one embodiment, the collaborative control module 130 is responsible for constructing a dynamic safety boundary field based on the output of the intelligent sensing and prediction module 120, and generating auxiliary guiding force commands or resistance commands to achieve proactive safety control of the device. This module includes a safety boundary field construction module, an adaptive impedance control module, and a command generation and output module.

[0040] For example, the safety boundary field construction module constructs a dynamically updated safety boundary field based on the predicted pose sequence output by the intelligent sensing and prediction module 120 and the real-time relative positional relationship between the instrument and the key anatomical structure. The construction of the safety boundary field can be achieved using a distance field function, extending outwards from the surface of the key anatomical structure by a preset safety distance to form a safety boundary region. The size of the safety distance is determined according to the importance of the key anatomical structure and the risk level of the surgical operation. For example, for high-risk anatomical structures such as blood vessels and nerves, the safety distance can be set to 5 mm to 10 mm; for low-risk tissues such as muscle and fat, the safety distance can be set to 2 mm to 5 mm. During the surgery, the safety boundary field is dynamically adjusted with the movement of the instrument and the updates of the digital twin, ensuring that the safe range of instrument movement is always accurately defined.

[0041] For example, the adaptive impedance controller module is the core of the collaborative control module 130. It is responsible for dynamically adjusting the impedance parameters based on the motion state and force feedback data of the instrument, and generating auxiliary guiding force commands or resistance commands. The impedance parameters of the adaptive impedance controller include stiffness, damping, and inertia. The adjustment strategy of these parameters is adaptively determined based on preset control rules and the surgical scenario.

[0042] Specifically, when it is predicted that the instrument is approaching a critical anatomical structure, the damping in the guidance direction is gradually increased to reduce the speed of the instrument. At the same time, a virtual force field channel is rendered in the virtual space display module using augmented reality to guide the doctor to manipulate the instrument to a safe area. When high-frequency tremors or sudden increases are detected in the force feedback data, the stiffness in the guidance direction is instantly reduced and a high-frequency filter is activated to actively suppress unstable operation and avoid surgical risks caused by the doctor's hand tremors or misoperation.

[0043] For example, the instruction generation and output module converts the auxiliary guiding force or resistance instruction generated by the adaptive impedance controller into a control signal that the instrument actuator can recognize, and transmits it to the instrument actuator through a high-speed communication interface. Simultaneously, this module is also responsible for coordinating and integrating the doctor's operating force signal with the auxiliary instruction to ensure that the auxiliary instruction is consistent with the doctor's operating intention, thereby achieving human-machine collaborative control.

[0044] In one embodiment, the human-machine collaborative interaction module 140 is used to realize the interaction between the doctor and the system, including the collection of doctor's operation instructions, the visualization of system feedback information, and force feedback interaction. This module mainly includes surgical instruments with force feedback function, AR head-mounted display device, and operation console.

[0045] Specifically, surgical instruments with force feedback function are the core execution components for interaction between the doctor and the system. Based on auxiliary commands output by the collaborative control module 130, they can simulate tactile sensations of virtual boundaries, force changes during tissue incision, and tactile feedback from interactions between the instrument and virtual objects such as sutures. The force feedback actuator of the instrument can be driven by an electric servo motor, using a torque sensor to detect the doctor's operating force in real time and transmit the force signal to the collaborative control module 130, achieving a synergistic effect between the doctor's operating force and the system's auxiliary commands.

[0046] For example, AR head-mounted displays are used to overlay information such as digital twins, safety boundary fields, and predicted instrument pose sequences onto the doctor's field of vision in an augmented reality manner, enabling the doctor to intuitively understand the relative positional relationship between the instruments and the patient's anatomical structures and their future movement trends. The AR head-mounted display device has a display resolution of no less than 2560×1440 and a field of view of no less than 100° to ensure a clear and comprehensive display of surgical information; simultaneously, the device supports binocular stereoscopic display, providing the doctor with depth perception and further improving the accuracy of the operation.

[0047] Furthermore, the control console is used to configure system parameters, monitor the surgical process, and perform emergency control. Doctors can use the control console to set system parameters such as preset time window length, safety distance, and initial impedance parameters. Simultaneously, the control console displays key data during the surgery in real time, such as instrument position, force feedback data, and the status of the digital twin, facilitating monitoring of surgical progress. In addition, the control console is equipped with an emergency stop button; in case of an emergency, the doctor can press the emergency stop button to immediately stop the movement of the instruments and ensure surgical safety.

[0048] Based on the above system, please refer to Figure 1 This disclosure provides a method for controlling the posture of a device based on four-way pose optical acquisition, comprising the following steps: Step S10: Simultaneously acquire multi-source initial data and perform fusion processing on the multi-source initial data to generate a multimodal fusion data package. The multi-source initial data includes: two-dimensional RGB images and three-dimensional point cloud data of the surgical area and instruments synchronously acquired by four optical acquisition devices along orthogonal directions; nine-axis inertial data of the instruments acquired by the micro inertial measurement unit; and force feedback data during the interaction between the instruments and tissues acquired by the six-dimensional force / torque sensor mounted on the end of the instrument.

[0049] Specifically, before the surgery begins, the various acquisition devices of the multimodal sensing module 100 need to be initialized and their parameters configured.

[0050] First, the four high frame rate optical cameras are calibrated, including intrinsic and extrinsic parameter calibration. Intrinsic parameter calibration can be performed using the Zhang Zhengyou calibration method, which obtains the camera's intrinsic parameter matrix (focal length, principal point coordinates, etc.) and distortion coefficients by photographing a checkerboard calibration board. Extrinsic parameter calibration uses a multi-camera joint calibration method to determine the relative positional relationship and attitude between the four cameras, ensuring that the coordinate systems of the four cameras can be accurately unified to the world coordinate system.

[0051] Secondly, the IMU sensor undergoes zero-drift calibration and temperature compensation to eliminate inherent system errors. The six-dimensional force / torque sensor is calibrated to determine parameters such as sensitivity and zero-point offset, ensuring the accuracy of force feedback data. Finally, data acquisition parameters are configured, including the frame rate and resolution of the optical camera, the sampling rate of the IMU and six-dimensional force / torque sensors, and the communication parameters of the data transmission interface.

[0052] Furthermore, after the surgery begins, a hardware trigger signal initiates the synchronous acquisition of multi-source data. This hardware trigger signal is generated by a synchronization controller and simultaneously sent to four optical cameras, an IMU sensor, and a six-dimensional force / torque sensor, controlling each device to begin data acquisition at the same time. The four optical cameras synchronously acquire two-dimensional RGB images and three-dimensional point cloud data of the surgical area and instruments along orthogonal directions; the IMU sensor synchronously acquires nine-axis inertial data of the instruments (three-axis acceleration, three-axis angular velocity, and three-axis magnetic field strength); and the six-dimensional force / torque sensor synchronously acquires force feedback data (forces along the X, Y, and Z axes and torques around the X, Y, and Z axes) during the interaction between the instruments and patient tissues. The acquired raw data is transmitted in real-time to the data preprocessing unit via a high-speed data transmission interface.

[0053] Furthermore, the data preprocessing unit performs a series of preprocessing operations on the collected raw data to improve data quality and usability. Specifically: First, denoising is performed, including using a Gaussian filtering algorithm to denoise the 2D RGB image and eliminate Gaussian noise; using a median filtering algorithm to denoise the 3D point cloud data and inertial data and eliminate salt-and-pepper noise and impulse noise; and using a Kalman filtering algorithm to denoise the force feedback data and smooth data fluctuations.

[0054] Then distortion correction is performed, including correcting the distortion of the two-dimensional RGB image based on the distortion coefficients obtained from the optical camera calibration, eliminating the impact of lens distortion on image quality, and ensuring that the image can accurately reflect the actual scene.

[0055] Next, coordinate system one is implemented, which involves unifying the 3D point cloud data acquired by the four optical cameras, the inertial data acquired by the IMU sensor, and the force feedback data acquired by the six-dimensional force / torque sensor into the world coordinate system. Specifically, the point cloud data in the camera coordinate system is transformed to the world coordinate system through the extrinsic parameter matrix of the optical cameras; the inertial data in the IMU coordinate system is transformed to the world coordinate system through the calibration relationship between the IMU sensor and the optical cameras; and the force feedback data in the sensor coordinate system is transformed to the world coordinate system through the installation relationship between the six-dimensional force / torque sensor and the instrument.

[0056] Finally, data format conversion is performed, including converting the preprocessed data into a unified data format (such as JSON or binary format) to facilitate subsequent fusion processing and transmission.

[0057] Furthermore, a multimodal data fusion algorithm based on Bayesian estimation can be used to fuse the preprocessed optical data, inertial data, and force feedback data to generate a multimodal fusion data packet. The Bayesian estimation algorithm can fully utilize the prior information and posterior probabilities of various data types to achieve optimal data fusion. The specific fusion process includes: First, establish observation models for each data type, including observation models for optical data, inertial data, and force feedback data, to describe the mapping relationship between each data type and the instrument's pose and interaction state.

[0058] Secondly, prior probabilities are calculated, including calculating the prior probability distribution of observations of various data types based on offline training data.

[0059] Furthermore, posterior probability estimation includes calculating the posterior probability distribution of the device pose and interaction state based on Bayes' theorem, combined with the observations and prior probabilities of various data types.

[0060] Finally, the fusion results are output, including the mean of the posterior probability distribution as the fusion result, generating a multimodal fusion data package containing the device's precise pose data, motion state data, and interaction force data with the tissue.

[0061] Step S20: Based on the multimodal fusion data package, reconstruct a dynamic 3D digital twin in real time in virtual space. The dynamic 3D digital twin includes surgical instruments, patient anatomical structures, and the interaction between the two; call the physics engine to simulate the mechanical interaction process between the instruments and virtual soft tissue within the dynamic 3D digital twin, and output the mechanical interaction simulation results.

[0062] Specifically, by integrating four-way optical, inertial, and force data, surgical instruments and patient anatomy are dynamically reconstructed in three dimensions in virtual space, accurately mapping their spatial and interactive relationships. This creates a virtual mirror highly synchronized with the actual surgical process, providing a unified and accurate foundation for contextual awareness for the entire system. Simultaneously, by invoking a physics engine (such as one based on the finite element method) to simulate the mechanical interactions between instruments and virtual soft tissue, real-time mechanical results, including tissue deformation, stress distribution, and virtual interactive forces, can be output. This transcends simple geometric reproduction, endowing the digital twin with physical properties, enabling it to predict the consequences of physical interactions and providing a simulation basis for force feedback that conforms to real biomechanical characteristics.

[0063] It is important to note that the constructed dynamic and computable digital twin serves as the core data source and testing ground for subsequent steps, including operational intent recognition, future pose prediction, and safety boundary calculation. Its included physical model allows predictions to extend beyond kinematic trajectories to encompass the effects of mechanical interactions, thereby making the prediction and control of the entire system more forward-looking and accurate.

[0064] Step S30: Input the state sequence of the dynamic three-dimensional digital twin, the historical motion trajectory of the instrument, and the real-time force feedback data into the pre-trained spatiotemporal network; output the operation intention prediction result through the spatiotemporal network, and simultaneously calculate the predicted pose sequence of the instrument within the future preset time window, and accurately perceive the real-time relative positional relationship between the instrument and key anatomical structures.

[0065] Understandably, the spatiotemporal network, taking the dynamic 3D digital twin state sequence, the instrument's historical motion trajectory, and real-time force feedback data as input, combined with the surgical scene motion patterns learned through pre-training, can directly output the predicted results of the surgeon's operational intentions, while simultaneously calculating the predicted instrument pose sequence within a preset future time window. Compared to traditional passive control schemes based solely on real-time pose, this step can predict the instrument's motion trend in advance, providing a forward-looking decision-making basis for subsequent dynamic safety boundary field construction and adaptive impedance control, fundamentally reducing the probability of the instrument intruding into dangerous areas. Simultaneously, step S30, relying on the high-fidelity 3D model of the digital twin and combining precise pose information after multi-source data fusion, can calculate the relative position, shortest distance, and posture relationship between the instrument and key anatomical structures in real time, overcoming the limitations of traditional optical acquisition being susceptible to occlusion and inertial measurement exhibiting drift. Furthermore, the feature extraction and fusion capabilities of the spatiotemporal network can effectively filter out noise interference in multi-source data, significantly improving the accuracy of relative position perception and providing reliable positioning references for surgical operations.

[0066] Step S40: Based on the predicted pose sequence obtained in step S30 and the perceived real-time relative position relationship, a dynamically updated safety boundary field is constructed. The safety boundary field is used to define the safe range of the instrument's movement.

[0067] Specifically, by predicting pose sequences, the system can anticipate the future movement risks of instruments before physical contact occurs, enabling it to "prevent problems before they arise" and significantly advancing the timing of safety controls. It's important to note that integrating the predicted trajectory with real-time spatial relationships to construct a field essentially involves comprehensively calculating the doctor's operational intentions, the dynamic characteristics of the instruments, and the patient's individualized anatomical structure within a unified mathematical framework. This allows the system not only to "see" its current position but also to "understand" the interaction between movement trends and the environment, significantly enhancing the system's overall intelligence.

[0068] Step S50: When it is predicted that the device will intrude into the safety boundary field in the future, or when there is an anomaly in the real-time force feedback data, the adaptive impedance controller is activated and its impedance parameters are dynamically adjusted to generate auxiliary guiding force commands or resistance commands. The auxiliary guiding force commands or resistance commands work together with the doctor's operating force to jointly drive the device to perform posture adjustment.

[0069] Specifically, in step S50, the impedance parameter adjustment of the adaptive impedance controller includes: When a critical anatomical structure is predicted to be approaching, the damping of the guidance direction is progressively increased, and a virtual force field channel is rendered in the virtual space display module using augmented reality. When high-frequency vibrations or sudden increases are detected in the force feedback data, the stiffness of the guidance direction is instantly reduced and a high-frequency filter is activated to actively suppress unstable operation.

[0070] Understandably, when predicting approach to critical structures, the "progressive increase in damping of the guiding direction" makes the doctor feel as if they are gradually entering a virtual force field with viscous resistance when operating instruments. This force feedback is continuous and predictable, allowing the doctor to intuitively perceive the direction and degree of approaching risk and adjust their operation naturally and smoothly, avoiding the interruption and discomfort caused by traditional abrupt stops or rigid barriers. Simultaneously, combined with "augmented reality rendering of the virtual force field channel," a recommended safe path is visually formed. The synergistic effect of force and visual guidance constructs a multi-sensory fusion guidance experience, greatly reducing the doctor's cognitive load and improving the intuition and efficiency of finding paths and avoiding obstacles in complex anatomical environments.

[0071] It is important to note that the "high-frequency filter activation" for "high-frequency tremors" is equivalent to adding selective attenuation to the control loop. This effectively filters out high-frequency noise introduced by physiological hand tremors in doctors, significantly improving the operational stability and positioning accuracy of the instrument's end effector. Simultaneously, for "suddenly increased" forces, the "instantaneous reduction of guide direction stiffness" is employed. This is equivalent to instantly switching the "hard" boundary to a "soft" buffer when the system senses a sudden impact, absorbing excessive kinetic energy and preventing tissue scratches, tears, or instrument slippage caused by excessive force, thus enhancing the system's safety tolerance in the face of unexpected operations.

[0072] Furthermore, the impedance parameters in this scheme are dynamically reconstructed based on predicted information (whether it is close to the critical structure) and real-time status (whether the force feedback is abnormal), enabling the system to proactively adapt to different surgical stages and operational situations with optimal impedance characteristics.

[0073] In some implementations, step S60 is also included.

[0074] Specifically, step S60 includes: During the surgery, the deviation between the predicted interactive force of the digital twin and the actual force feedback sensor data is continuously compared. When the deviation exceeds a preset threshold, a lightweight online learning loop is triggered to fine-tune the tissue mechanical parameters of some layers of the spatiotemporal network or the physical engine using the current surgical data, so that the digital twin can dynamically adapt to the individualized tissue characteristics of the current patient.

[0075] In step S60, a meta-learning optimizer pre-trained on a large amount of offline surgical data is used to obtain initial values ​​of model parameters that can quickly adapt to new tasks, thereby enabling the fine-tuning process to converge quickly with limited real-time surgical data.

[0076] Understandably, by continuously comparing the deviation between predicted and measured forces, the system can automatically perceive the differences between the current patient's tissue biomechanical properties (such as elasticity and viscosity) and those of the general model or preoperative predictions. After triggering fine-tuning, the system uses real-time intraoperative data to calibrate the core model (spatiotemporal network or physical parameters), ensuring that the behavior of the digital twin dynamically converges to the actual physical properties of the current patient. Therefore, this step effectively overcomes model inaccuracies caused by individual patient differences and varying pathological states, ensuring the reliability of the prediction, simulation, and control foundations.

[0077] It is important to note that, based on step S60, the system no longer relies on the initial model, but has the ability to adapt to dynamic changes and uncertainties, which greatly improves the reliability and generalizability of the technology in complex real clinical environments.

[0078] In summary, this disclosure provides a method and system for instrument pose control based on four-way pose optical acquisition. By using four orthogonally arranged optical cameras, combined with the instrument's built-in IMU and end-effector force sensor, a multimodal sensing network with no blind spots is constructed. Tight coupling and fusion of multi-source data enables continuous, high-precision tracking of the instrument's six-degree-of-freedom pose, and simultaneously acquires its interaction with the tissue's mechanical state, providing the system with comprehensive and reliable real-world data mapping.

[0079] Meanwhile, the digital twin, reconstructed in real time based on multimodal data, is a dynamic virtual mirror of the actual surgical scene. Integrating a personalized physics engine based on the finite element method, the twin can simulate the biomechanical interactions between instruments and soft tissue, providing a "digital testbed" with both geometric and physical realism for prediction and control.

[0080] By analyzing the state and historical data of digital twins through spatiotemporal networks, the future position and operational intent of instruments can be predicted. A dynamic safety boundary field constructed using this predictive information extends safety monitoring from the current moment to a future time window, enabling risk warnings. Combined with an AR visualization channel, this forms a multi-sensory integrated intelligent guidance system, significantly optimizing the doctor's operational intuition and fluency while ensuring safety limits.

[0081] Based on this, this solution forms a self-iteratory and continuously optimized intelligent control closed loop through the deep coupling of multimodal perception, digital twin reconstruction, spatiotemporal prediction, dynamic boundary construction, adaptive control, and online learning. It transforms surgical instruments from passive tools requiring full physician control into intelligent decision-making systems capable of understanding the scenario, predicting risks, and providing compliant assistance, fundamentally improving the safety, precision, and accessibility of complex surgical procedures.

[0082] In the description herein, it should be understood that the terms "upper," "lower," "left," "right," and other orientations or positional relationships are used only for ease of description and simplification of operation, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first" and "second" are used merely for descriptive distinction and have no special meaning.

[0083] In the description of this specification, references to terms such as "an embodiment," "example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example.

[0084] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style of the specification is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0085] The technical principles of this application have been described above with reference to specific embodiments. These descriptions are merely for explaining the principles of this application and should not be construed as limiting the scope of protection of this application in any way. Based on this explanation, those skilled in the art can readily conceive of other specific embodiments of this application without inventive effort, and these embodiments will all fall within the scope of protection of this application.

Claims

1. A method for controlling the pose of a device based on four-way pose optical acquisition, characterized in that, include: Step S10: Synchronously acquire multi-source initial data, perform fusion processing on the multi-source initial data, and generate a multi-modal fusion data packet; The multi-source initial data includes: two-dimensional RGB images and three-dimensional point cloud data of the surgical area and instruments synchronously acquired by four optical acquisition devices in orthogonal directions; nine-axis inertial data of the instruments acquired by the micro inertial measurement unit; and force feedback data of the interaction between the instruments and tissues acquired by the six-dimensional force / torque sensor mounted on the end of the instruments. Step S20: Based on the multimodal fusion data package, a dynamic three-dimensional digital twin is reconstructed in real time in virtual space. The dynamic three-dimensional digital twin includes surgical instruments, patient anatomical structures, and the interaction between the two. The physics engine is invoked to simulate the mechanical interaction process between the instruments and virtual soft tissue within the dynamic three-dimensional digital twin, and the mechanical interaction simulation results are output. Step S30: Input the state sequence of the dynamic three-dimensional digital twin, the historical motion trajectory of the instrument, and the real-time force feedback data into the pre-trained spatiotemporal network; output the operation intention prediction result through the spatiotemporal network, and simultaneously calculate the predicted pose sequence of the instrument within a future preset time window, and accurately perceive the real-time relative positional relationship between the instrument and key anatomical structures. Step S40: Based on the predicted pose sequence obtained in step S30 and the perceived real-time relative position relationship, a dynamically updated safety boundary field is constructed. The safety boundary field is used to define the safe range of the movement of the device. Step S50: When it is predicted that the device will intrude into the safety boundary field in the future, or when there is an anomaly in the real-time force feedback data, the adaptive impedance controller is activated and its impedance parameters are dynamically adjusted to generate auxiliary guiding force commands or resistance commands; the auxiliary guiding force commands or resistance commands work together with the doctor's operating force to jointly drive the device to perform posture adjustment.

2. The instrument pose control method based on four-way pose optical acquisition according to claim 1, characterized in that, In step S10, the synchronous acquisition of the four optical acquisition devices is achieved through hardware-triggered synchronization, and the two-dimensional RGB image and three-dimensional point cloud data are preprocessed, including denoising, distortion correction and coordinate system one.

3. The instrument pose control method based on four-way pose optical acquisition according to claim 1, characterized in that, In step S20, the mechanical model of the virtual soft tissue is constructed based on the finite element method, and the initial values ​​of the parameters of the mechanical model are generated based on the patient's preoperative medical imaging data, including CT images or MRI images.

4. The instrument pose control method based on four-way pose optical acquisition according to claim 1, characterized in that, Following step S50, step S60 is also included: During the surgery, the deviation between the predicted interactive force of the digital twin and the actual force feedback sensor data is continuously compared. When the deviation exceeds a preset threshold, a lightweight online learning loop is triggered to fine-tune the tissue mechanics parameters of some layers of the spatiotemporal network or the physical engine using the current surgical data, so that the digital twin can dynamically adapt to the individualized tissue characteristics of the current patient.

5. The instrument pose control method based on four-way pose optical acquisition according to claim 1 or 4, characterized in that, The spatiotemporal network adopts an encoder-decoder structure. The encoder is used to process the state sequence of the digital twin, and the decoder is used to output a predicted sequence based on the real-time operation signal of the instrument as the query vector and the most relevant anatomical structure and instrument component for the current operation.

6. The instrument pose control method based on four-way pose optical acquisition according to claim 1, characterized in that, In step S50, the impedance parameter adjustment of the adaptive impedance controller includes: When a critical anatomical structure is predicted to be approaching, the damping of the guidance direction is progressively increased, and a virtual force field channel is rendered in the display module of the virtual space using augmented reality. When high-frequency vibrations or sudden increases are detected in the force feedback data, the stiffness of the guidance direction is instantly reduced and a high-frequency filter is activated to actively suppress unstable operation.

7. The instrument pose control method based on four-way pose optical acquisition according to claim 4, characterized in that, In step S60, a meta-learning optimizer pre-trained on a large amount of offline surgical data is used to obtain initial values ​​of model parameters that can quickly adapt to new tasks, thereby enabling the fine-tuning process to converge quickly with limited real-time surgical data.

8. A device pose control system based on four-way pose optical acquisition for implementing the method as described in any one of claims 1 to 7, characterized in that, include: The multimodal sensing module includes four high-frame-rate optical cameras, an IMU sensor built into the instrument, and a six-dimensional force / torque sensor at the end of the instrument. The digital twin computing engine module is used to run the physics engine and real-time 3D reconstruction algorithms; The intelligent sensing and prediction module has a built-in pre-trained spatiotemporal network; A collaborative control module is used to implement a forward-looking adaptive impedance control algorithm; and The human-machine collaborative interaction module includes a force feedback device and an AR head-mounted display device, which are used to receive doctor's instructions and visualize the digital twin and safety boundaries.

9. The instrument pose control system based on four-way pose optical acquisition according to claim 8, characterized in that, The system supports collaborative operation of multiple instruments. The digital twin computing engine module independently constructs a sub-digital twin for each instrument and manages the motion rules and anti-collision logic between instruments through a central coordinator.

10. The instrument pose control system based on four-way pose optical acquisition according to claim 8, characterized in that, The instrument in the human-machine collaborative interaction module can simulate the tactile sensation of virtual boundaries, the force changes of tissue incision, and the tactile feedback of the interaction between the instrument and virtual objects such as sutures, according to the instructions of the adaptive impedance controller.