Surgical robot resection navigation method based on multi-modal image registration
Through deep learning networks, cross-modal features are extracted and non-rigid deformation fields are predicted, combined with incremental calculations and real-time update strategies, the problems of soft tissue deformation and multi-modal image registration in robot-assisted surgery are solved, and high-precision and real-time navigation information are achieved, improving the accuracy and safety of the surgery.
Patent Information
- Application Number
- CN202510528186.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
In robot-assisted surgery, complex deformation of soft tissues and organs during the operation due to gravity, pneumoperitoneal pressure and other reasons leads to significant deviations between preoperative planning and actual situation during the operation, affecting navigation accuracy. At the same time, the significant differences between multimodal images make it difficult to establish accurate correspondence and increase navigation complexity.
Deep learning network is used to extract cross-modal invariant features from preoperative medical images and intraoperative images, establish initial correspondence, and predict initial non-rigid deformation field. Track the position of surgical robot tools in real time, update the non-rigid deformation field in real time through incremental computing strategies, generate the deformed three-dimensional model, and map the tool position into the model to generate navigation instructions.
Accurate estimation and real-time compensation of soft tissue intraoperative deformation is achieved, the accuracy and safety of navigation is improved, and the preoperative planning can be more accurately mapped to the currently deformed organs, enhancing the robustness and real-time nature of the surgical robot navigation system.
Smart Images

Figure CN120070423A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image data processing, and particularly to a surgical robot resection navigation method based on multimodal image registration. Background Art
[0002] In recent years, Robot-Assisted Minimally Invasive Surgery (RAMIS) has been widely applied and developed in the field of surgery. Compared with traditional open surgery and conventional laparoscopic surgery, a surgical robot system (such as the da Vinci surgical system) can provide surgeons with a three-dimensional high-definition magnified view, flexible operating wrists, and the ability to filter out hand tremors, thus enabling more precise and stable operations in a limited surgical space, reducing patient trauma and shortening the postoperative recovery time.
[0003] In precise resection surgeries such as tumor resection, in order to maximize the resection of diseased tissues while maximizing the preservation of healthy organ functions and avoiding damage to important blood vessels, nerves and other critical structures (i.e., the Organs at Risk, abbreviated as OARs), the precision and safety of the surgery are of utmost importance. For this reason, Image-Guided Navigation (IGN) technology has been introduced into robot-assisted surgery. Its core idea is to accurately map the rich anatomical information and surgical plans (such as tumor boundaries, safety margins, OARs positions) contained in preoperative high-resolution medical images (such as magnetic resonance imaging MRI, computed tomography CT, etc.) into the surgical field of view, providing real-time navigation information for doctors and assisting them in precisely operating the surgical robot.
[0004] A typical image-guided navigation process usually includes: 1) preoperative use of high-resolution images (MRI, CT, etc.) for lesion segmentation, three-dimensional model reconstruction and surgical path planning; 2) intraoperative registration and alignment of the patient's real-time position with preoperative data through specific methods (such as body surface marker points, intraoperative imaging, etc.); 3) real-time tracking of the position of surgical instruments and fusion display of the relationship between the instruments and anatomical structures and planned paths on the navigation interface.
[0005] However, when applying existing image-guided navigation technology to robot-assisted resection surgeries of soft tissue organs (such as the liver, kidney, brain, prostate, etc.), serious technical challenges are faced: I. Soft Tissue Intraoperative Deformation Problem: Different from hard tissues such as bones, soft tissue organs will undergo significant and non-linear deformation (Deformation) during surgery due to gravity, pneumoperitoneum pressure (laparoscopic surgery), patient respiratory movement, traction, extrusion, and even resection operations of surgical instruments. This deformation will cause significant deviations between the pre-planned tumor boundaries, resection paths, and OARs positions and the actual intraoperative situation, seriously affecting the navigation accuracy, and may even lead to incomplete resection or damage to key structures. Therefore, simply rigidly aligning the preoperative static model to the intraoperative situation is far from enough, and accurate estimation and compensation of soft tissue deformation must be carried out.
[0006] II. Challenges in Multimodal Image Registration: There are significant differences between the high-resolution images relied on for preoperative planning (such as MRI providing excellent soft tissue contrast, CT providing good bone and density information, and PET providing functional information) and the imaging modalities that are available or commonly used intraoperatively (such as intraoperative ultrasound iUS providing real-time internal structure information but with low resolution and artifacts, and endoscopes providing surface color texture information but lacking depth and internal information). These differences are reflected in aspects such as imaging principles, image intensity distributions, resolutions, signal-to-noise ratios, geometric distortions, and artifacts. How to establish accurate correspondence relationships between these multimodal images with vastly different appearances and characteristics and achieve robust registration (especially non-rigid registration for compensating deformation) is the core difficulty in image-guided navigation. Traditional methods based on feature point or surface matching often have insufficient accuracy due to difficult feature extraction or sparse features; methods based on image intensity (such as mutual information method) are vulnerable to modal differences and intensity changes.
[0007] III. Requirements for Real-time and Robustness: Surgical navigation information must be provided to doctors in real-time or near real-time to effectively guide surgical operations. This means that complex non-rigid registration algorithms need to have high enough computational efficiency while meeting clinical accuracy requirements. In addition, the surgical environment is complex and variable, and the navigation system must have a certain degree of robustness to image noise, artifacts, occlusion, and unexpected tissue movement.
[0008] Existing surgical navigation methods still have deficiencies in simultaneously solving the above problems such as large soft tissue deformation, multimodal data fusion, and real-time robust registration, and are difficult to meet the urgent need for high-precision navigation in robot-assisted precise resection surgery.
[0009] Therefore, there is an urgent need to develop a new method that can effectively fuse multimodal image information, accurately estimate and compensate soft tissue intraoperative deformation in real-time, so as to provide accurate and reliable navigation information for surgical robots, and improve the precision and safety of robot-assisted resection surgery. Summary of the Invention
[0010] This application proposes a surgical robot resection navigation method based on multimodal image registration, which can accurately estimate and real-time compensate for intraoperative deformation of soft tissues, so as to provide accurate and reliable navigation information for the surgical robot.
[0011] According to an embodiment of the present application, a surgical robot resection navigation method based on multimodal image registration is proposed. The method includes: Obtain preoperative medical images of the patient that at least include a first-modal image and a second-modal image, and reconstruct a preoperative three-dimensional model including a target lesion area and organs at risk (OARs) based on the preoperative image data; Obtain intraoperative images of at least one modality different from the preoperative medical images of the target organ area of the patient; Use a first deep learning network to extract cross-modal invariant features from the preoperative medical images and intraoperative images, establish at least one set of initial correspondence relationships, and use a second deep learning network to predict an initial non-rigid deformation field based on the preoperative medical images, intraoperative images, and the extracted initial correspondence relationships. The non-rigid deformation field is used to represent the geometric deformation of the target organ from the preoperative state to the intraoperative state; During the surgical process, real-time track the position of the end effector of the surgical robot, and based on the real-time obtained intraoperative images and surgical operation information, adopt an incremental calculation strategy to update the non-rigid deformation field in real time; Apply the currently updated non-rigid deformation field to the preoperative three-dimensional model to generate a deformed three-dimensional model, and the state of the target organ in the deformed three-dimensional model matches the state of the target organ in the real-time intraoperative images; Map the position of the real-time tracked end effector of the surgical robot to the coordinate system of the deformed three-dimensional model; Generate navigation instructions based on the mapped position of the tool and the target lesion area and organs at risk of the deformed three-dimensional model to guide the surgical robot to perform a resection operation.
[0012] In some embodiments, the first deep learning network is a Siamese Network structure with a shared-weight encoder for mapping image patches of different modalities to the same feature space. The first deep learning network is trained by the following Contrastive Loss function: , where Y is the label indicating whether the input preoperative image patch and intraoperative image patch belong to the same anatomical position. Y = 0 indicates that the image patch pair is similar, Y = 1 indicates that the image patch pair is dissimilar, D_W represents the Euclidean distance between the two feature vectors output by the Siamese network encoder, and m is a preset boundary parameter, and m is a positive number.
[0013] In some embodiments, the second deep learning network has an encoder-decoder structure and incorporates an attention mechanism, which is used to adaptively focus on regions corresponding to encoder features or cross-modal features that contribute more to the current position deformation prediction during the decoding stage.
[0014] In some embodiments, the second deep learning network is trained using the following loss function:
[0015] where \(L_{gradient\_similarity}\) is an image gradient similarity metric term used to measure the similarity between the intraoperative image obtained by applying the predicted non-rigid deformation field to the preoperative medical image and the actual intraoperative image; \(L_{smoothness}\) is a deformation field smoothness regularization term used to penalize non-smooth changes in the predicted non-rigid deformation field; \(L_{physics}\) is a physical consistency regularization term determined based on a predefined physical deformation model; \(L_{correspondence}\) is an initial correspondence supervision term used to supervise the learning of the second deep learning network using the initial correspondence; and \(\alpha\), \(\beta\), \(\gamma\), and \(\delta\) are weight factors, \(\alpha\), \(\beta\), \(\gamma\) are preset factors, and \(\delta\) is adaptively adjusted according to the confidence of the initial correspondence.
[0016] In some embodiments, the image gradient similarity metric term \(L_{gradient\_similarity}\) is determined according to the following formula:
[0017] where \(\int\) represents integration over the entire target organ region \(\Omega\), \(X\) represents a point in the preoperative medical image, represents the corresponding point obtained by mapping the preoperative point \(X\) to the intraoperative coordinate system through the non-rigid deformation field, represents the gradient of the preoperative medical image after deformation at the point in the intraoperative coordinate system, represents the gradient of the intraoperative image at the point in the intraoperative coordinate system, and \(\|\cdot\|^2\) represents the square of the L2 norm of \(\cdot\).
[0018] In some embodiments, the predefined physical deformation model is a hyperelastic model based on the finite element method, and the model parameters are determined based on the preoperative medical image. The physical consistency regularization term is determined according to the following formula: , where \(\int\) represents integration over the entire target organ region \(\Omega\), is the deformation gradient tensor, , \(I\) represents the identity tensor, represents the predicted non-rigid deformation field, represents the gradient of the non-rigid deformation field, which is a strain energy metric function.
[0019] In some embodiments, the incremental calculation strategy is used to re-calculate the non-rigid deformation field within the affected area and its immediate neighboring areas, and the affected area is determined according to the following method: Based on the current position of the end effector of the surgical robot being tracked in real time, a first area is determined within a preset spatial range around the tool tip; The local image intensity differences between consecutive intraoperative image data frames are calculated, and an area where the local image intensity difference exceeds a preset difference threshold is identified as the second area; Using a physical deformation model, a third area where significant stress is generated is estimated based on the interaction force between the end effector of the surgical robot and the body tissue; The union of the first area, the second area, and the third area is determined as the affected area.
[0020] In some embodiments, re-calculating the non-rigid deformation field within the affected area and its immediate neighboring areas using the incremental calculation strategy includes: Fixing the deformation values at the boundary of the affected area as boundary conditions; Inside the affected area, the non-rigid deformation field parameters are rapidly updated by performing a finite number of local optimization iterations based on gradient descent to minimize a loss function within the affected area that includes image similarity, deformation field smoothness, physical consistency, and correspondence compliance, and using the current global non-rigid deformation field as the initial state; The iterative update further includes an error accumulation control mechanism, including: periodically, or when it is detected that the accumulated error exceeds a preset error threshold, performing global non-rigid deformation field optimization to re-minimize the loss function used to train the second deep learning network over the entire target organ area.
[0021] In some embodiments, the first modality is a magnetic resonance imaging (MRI) T 2 -weighted image, the second modality is a CT angiography (CTA) image, and the intraoperative image is a three-dimensional intraoperative ultrasound (iUS) image.
[0022] According to the surgical robot resection navigation method based on multi-modal image registration proposed in this application, by introducing innovative technical solutions, the key technical bottlenecks existing in existing robot-assisted soft tissue surgery navigation are effectively overcome, and significant beneficial effects are achieved: I. By adopting a non-rigid registration strategy that combines deep learning feature extraction, physical model constraints, and real-time intraoperative image information, this application can accurately estimate and compensate for the complex deformations that occur in soft tissues during surgery. The first deep learning network is used to extract cross-modal invariant features and establish initial correspondences, providing reliable driving information for deformation estimation. At the same time, the second deep learning network predicts the deformation field and constrains it by introducing a physical consistency regularization term based on a biomechanical model, ensuring the physical rationality and smoothness of the predicted deformation. This data-driven and model-driven combination significantly improves the registration accuracy, enabling preoperative planning (such as tumor boundaries and critical structures) to be accurately mapped onto the currently deformed organ, laying the foundation for precise resection.
[0023] II. This application effectively solves the registration problem of significant modal differences between preoperative high-resolution images (such as MRI / CT) and intraoperative real-time images (such as iUS). The first deep learning network (such as a Siamese network) is specifically designed to learn the shared feature representations between cross-modal images, overcoming the limitations of traditional methods that are difficult to establish reliable correspondences under different imaging principles and image characteristics. In addition, the design of the image gradient similarity metric term and other terms in the loss function further enhances the robustness to modal differences, achieving more reliable multi-modal data fusion.
[0024] III. The incremental calculation strategy and error accumulation control mechanism proposed in this application endow the navigation system with the ability to update in real-time dynamically and maintain accuracy over a long period. By concentrating computational resources on the local area affected by surgical operations, rapid iterative updates of the non-rigid deformation field are achieved, enabling timely response to the continuous changes of intraoperative tissues. At the same time, the global optimization mechanism triggered periodically or based on an error threshold effectively prevents the error accumulation that may be brought about by multiple local updates, ensuring the accuracy and reliability of navigation information throughout the surgical process.
[0025] By providing high-precision, real-time updated, and robust navigation information to modal differences, this application can significantly improve the accuracy and safety of robot-assisted resection surgery. Doctors can more clearly and confidently identify the boundaries of target lesions, safe resection margins, and key anatomical structures that need to be avoided, thereby guiding the surgical robot to perform more precise cutting operations. This helps to achieve more complete lesion resection, maximize the preservation of healthy tissues, reduce the risk of intraoperative damage to adjacent important blood vessels and nerves, and is expected to improve the surgical outcomes and prognosis of patients. At the same time, the technical solutions proposed in this application, especially the fast prediction and incremental update strategies based on deep learning, have good potential for computational efficiency and are suitable for clinical practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.
[0027] Figure 1 The flowchart of a surgical robot resection navigation method based on multimodal image registration according to an embodiment of the present application is shown.
[0028] Figure 2(a) shows an example diagram of a preoperative MRI T2-weighted image according to an example of the present application; Figure 2(b) shows an example diagram of a preoperative CT image according to an example of the present application; Figure 2(c) shows an example diagram of a preoperative three-dimensional model according to an example of the present application.
[0029] Figure 3 The schematic diagram of a multi-planar reconstruction (MPR) navigation interface based on three-dimensional intraoperative ultrasound (3D iUS) data according to an example of the present application is shown.
[0030] Figure 4 The schematic diagram of a deformed three-dimensional model according to an example of the present application is shown.
[0031] Figure 5 The schematic diagram of a surgical robot resection navigation interface according to an example of the present application is shown.
[0032] Figure 6 It is the schematic diagram of the structure of an electronic device shown in at least one embodiment of the present application. Detailed implementation manners
[0033] Here, the exemplary embodiments will be described in detail, and their examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0034] The embodiments of the present application can be applied to a computer system / server, which can operate together with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with a computer system / server include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0035] A computer system / server may be described in the general context of computer system-executable instructions, such as program modules, executed by the computer system. Generally, program modules may include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server may be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules may be located on local or remote computing system storage media including storage devices.
[0036] Figure 1 The flowchart of a surgical robot resection navigation method based on multimodal image registration according to an embodiment of the present application is shown. Through advanced image processing and registration techniques, this method can fuse preoperative data containing detailed anatomical information and surgical plans with the dynamically changing organ state during the operation in real time and accurately, providing precise navigation guidance for the surgical robot. As shown in the figure, the method includes steps 1 to 7.
[0037] Step 1: Obtain preoperative medical images of the patient that at least include a first-modal image and a second-modal image, and reconstruct a three-dimensional model including the target lesion area and organs at risk (OARs) based on the preoperative image data.
[0038] Step 1 is the preparation stage. Before the operation, detailed images of the patient's target organ are obtained using at least two different high-resolution imaging techniques (for example, MRI provides soft tissue details, and CTA shows vascular structures). Subsequently, through image segmentation and three-dimensional reconstruction techniques, a digital three-dimensional anatomical model is generated that includes the precise location and shape of the tumor, as well as the distribution of important surrounding blood vessels, nerves, and other dangerous structures.
[0039] The fusion of multimodal images ensures the comprehensiveness of information and overcomes the limitations of a single modality.
[0040] Step 2: Obtain at least one intraoperative image of the target organ area of the patient that is in a different modality from the preoperative medical images.
[0041] During the operation, real-time image information that can reflect the current "true state" of the organ can be obtained. Since preoperative imaging devices are usually not available during the operation, this step uses one or more imaging modalities suitable for the operation (for example, three-dimensional intraoperative ultrasound iUS) to obtain real-time images of the target organ area. This intraoperative modality is usually significantly different from the preoperative modality.
[0042] The intraoperative image provides the "target map" for subsequent registration, that is, the actual shape and position of the organ after being affected by operation-related factors (such as pneumoperitoneum, gravity, and manipulation), and is the basis for subsequent deformation estimation and compensation.
[0043] Step 3: Using a first deep learning network, extract cross-modal invariant features from preoperative medical images and intraoperative images, establish at least one set of initial correspondence relationships, and use a second deep learning network to predict an initial non-rigid deformation field based on the preoperative medical images, intraoperative images, and the extracted initial correspondence relationships. The non-rigid deformation field is used to represent the geometric deformation of the target organ from the preoperative state to the intraoperative state.
[0044] Step 3 is the core computational step for achieving preoperative and intraoperative data alignment. First, use a specially trained first deep learning network (such as a Siamese network) to overcome the modality differences, automatically find and match common anatomical feature points in preoperative and intraoperative images, and establish preliminary spatial correspondence relationships. Then, use a second deep learning network to comprehensively predict an initial non-rigid deformation field by combining the original image information and the extracted correspondence relationships. The deformation field mathematically describes the spatial distortion and displacement required for the organ to "deform" from the preoperative scan state to the current intraoperative initial state.
[0045] Through a two-step deep learning strategy, two major problems of cross-modal registration and complex non-rigid deformation estimation are effectively overcome. The first network focuses on robust feature matching, and the second network is responsible for generating accurate and physically reasonable initial deformation estimates using these matching information and global image information, laying a foundation for subsequent real-time navigation.
[0046] In some embodiments, the first deep learning network has a Siamese Network structure with a shared-weight encoder for mapping image patches of different modalities to the same feature space. The first deep learning network is trained using the following Contrastive Loss function: , where Y is the label indicating whether the input preoperative image patch and intraoperative image patch belong to the same anatomical location. Y = 0 indicates that the image patch pair is similar, Y = 1 indicates that the image patch pair is dissimilar, D_W represents the Euclidean distance between the two feature vectors output by the Siamese network encoder, and m is a preset boundary parameter, and m is a positive number.
[0047] This embodiment can effectively learn invariant feature representations between cross-modal images (such as MRI / CT and iUS) by adopting a Siamese network and contrastive loss, thus significantly improving the robustness and accuracy of establishing reliable initial correspondence relationships under different imaging principles and image characteristics.
[0048] In some embodiments, the second deep learning network has an encoder-decoder structure and incorporates an attention mechanism, which is used to adaptively focus on regions corresponding to encoder features or cross-modal features that contribute more to the current position deformation prediction during the decoding stage.
[0049] In this embodiment, the attention mechanism is integrated into the encoder-decoder network for non-rigid deformation field prediction to enable it to fuse multi-source information. Based on this network structure combined with the attention mechanism, the accuracy and robustness of non-rigid deformation field prediction can be significantly improved. By intelligently focusing on and fusing key information from the encoder and cross-modal correspondences, the network can more accurately capture and align important anatomical structures. This is particularly helpful for improving the registration effect in regions with poor image quality or complex deformations.
[0050] In some embodiments, the second deep learning network is trained using the following loss function:
[0051] where \(L_{gradient\_similarity}\) is an image gradient similarity metric term used to measure the similarity between the intraoperative image obtained by applying the predicted non-rigid deformation field to the preoperative medical image and the actual intraoperative image; \(L_{smoothness}\) is a deformation field smoothness regularization term used to penalize non-smooth changes in the predicted non-rigid deformation field; \(L_{physics}\) is a physical consistency regularization term determined based on a predefined physical deformation model; \(L_{correspondence}\) is an initial correspondence supervision term used to utilize the initial correspondence to supervise the learning of the second deep learning network; and \(\alpha\), \(\beta\), \(\gamma\), and \(\delta\) are weight factors, where \(\alpha\), \(\beta\), and \(\gamma\) are preset factors, and \(\delta\) is adaptively adjusted according to the confidence of the initial correspondence.
[0052] According to the composite loss function of this embodiment, a collaborative optimization framework is formed through a specific combination of four key constraint terms: image gradient similarity (\(L_{gradient\_similarity}\)), deformation field smoothness (\(L_{smoothness}\)), physical consistency based on the physical model (\(L_{physics}\)), and initial correspondence supervision (\(L_{correspondence}\)). Moreover, this loss function introduces an adaptive adjustment mechanism for the weight \(\delta\) of the correspondence supervision term, enabling the network to intelligently evaluate and utilize the reliability of the initial correspondence information from the first network.
[0053] Training the second deep learning network based on the above composite loss function can significantly improve the comprehensive quality of predicting the non-rigid deformation field. By optimizing multiple objectives simultaneously, it ensures that the prediction results are not only aligned at the image level (gradient similarity), but also meet physical rationality (smoothness and physical consistency), and fully utilizes the known feature correspondence information. In particular, the adaptive weight adjustment mechanism enhances the robustness of the system to possible output noise or incorrect correspondences of the first deep learning network. This balanced and robust optimization strategy makes the generated deformation field more accurate, stable and more in line with clinical practice, significantly improving the reliability of subsequent surgical navigation.
[0054] In some embodiments, the image gradient similarity metric term L_gradient_similarity is determined according to the following formula:
[0055] where ∫ represents integration over the entire target organ region Ω, X represents a point in the preoperative medical image, represents the corresponding point obtained by mapping the preoperative point X to the intraoperative coordinate system through the non-rigid deformation field, represents the gradient at the point in the intraoperative coordinate system after the preoperative medical image is deformed, represents the gradient at the point in the intraoperative image in the intraoperative coordinate system, and || *||² represents the square of the L2 norm of *.
[0056] This embodiment adopts a specific image gradient similarity metric term L_gradient_similarity to evaluate the matching degree between the deformed preoperative image and the intraoperative image. By calculating and minimizing the squared L2 norm difference between the gradient vector fields of the two images, rather than directly comparing the image intensity values. This gradient-based strategy can effectively overcome the registration difficulties caused by the inherent huge differences in intensity distribution, contrast and scale between the preoperative high-resolution images (such as MRI / CT) and the intraoperative real-time ultrasound (iUS) images.
[0057] Therefore, by focusing on the gradient information that can better reflect the anatomical structure edges and textures for alignment, this similarity metric term significantly improves the robustness and accuracy of cross-modal image registration, enabling the system to more reliably capture the correspondences of key anatomical structures, providing a more effective supervision signal for generating an accurate non-rigid deformation field, and ultimately improving the navigation accuracy.
[0058] In some embodiments, the predefined physical deformation model is a hyperelastic model based on the finite element method, and the model parameters are determined based on the preoperative medical image. The physical consistency regularization term is determined according to the following formula: , where, ∫ represents the integration over the entire target organ region Ω, is the deformation gradient tensor, , I represents the identity tensor, represents the predicted non-rigid deformation field, represents the gradient of the non-rigid deformation field, is the strain energy metric function.
[0059] In this embodiment, by introducing a physically consistent regularization term L_physics based on biomechanical principles to constrain the prediction of the non-rigid deformation field, the integral of the biomechanical strain energy corresponding to the predicted deformation (calculated according to a predefined hyperelastic model that can reflect the characteristics of the target soft tissue) over the entire organ region is incorporated into the loss function for minimization.
[0060] This constraint based on the physical model is significantly better than simple smoothness or volume preservation requirements, ensuring that the predicted deformation is not only mathematically continuous but also physically consistent with the true mechanical behavior of soft tissues (e.g., the ability to resist excessive stretching or unreasonable compression). Therefore, through this embodiment, it is possible to effectively avoid generating unrealistic distortions or deformations, making the finally predicted deformation field more realistic and reliable, and significantly improving the accuracy of the navigation system in simulating the actual state of the intraoperative organ.
[0061] L_smoothness can adopt a standard diffusion regularization term, such as penalizing the sum of the squares of the gradients of the deformation displacement field.
[0062] L_correspondence can adopt an initial correspondence supervision term, such as calculating the mapping error between corresponding point pairs.
[0063] Step 4, during the surgical process, the position of the end effector of the surgical robot is tracked in real time, and based on the intraoperative images and surgical operation information obtained in real time, an incremental calculation strategy is adopted to update the non-rigid deformation field in real time.
[0064] This application introduces a real-time update mechanism to adapt to the continuous dynamic deformation of soft tissues during surgery. On the one hand, the system accurately obtains the real-time spatial position of the surgical robot tool through a tracking system (such as an optical tracking system, an electromagnetic tracking system, image / vision-based tracking, etc.); on the other hand, according to the newly obtained intraoperative images and the perceived surgical operations (such as cutting, pulling), an efficient incremental calculation method is used to quickly recalculate and update the non-rigid deformation field, and this update mainly focuses on the affected local areas.
[0065] In Step 4, by updating the deformation field in real time, the navigation system can dynamically adapt to the continuous changes of the organ during the surgical process, avoiding the accumulation of navigation errors caused by using outdated deformation information, ensuring the timeliness and accuracy of the navigation information, and effectively balancing the accuracy and computational efficiency through an incremental calculation strategy.
[0066] In some embodiments, the incremental calculation strategy is used to define recalculating the non-rigid deformation field within the affected region and its immediate neighboring regions, and the affected region is determined according to the following method: Based on the current position of the end effector of the surgical robot being tracked in real time, a first region is determined within a preset spatial range around the tool tip; Calculate the local image intensity difference between consecutive intraoperative image data frames, and identify the region where the local image intensity difference exceeds a preset difference threshold as the second region; Using a physical deformation model, estimate a third region where significant stress is generated based on the interaction force between the end effector of the surgical robot and the body tissue; Determine the union of the first region, the second region, and the third region as the affected region.
[0067] To cope with the continuous dynamic changes of soft tissues during surgery, the present application introduces an incremental calculation strategy to update the non-rigid deformation field in real time. According to this embodiment, the incremental calculation strategy efficiently concentrates computational resources on the determined affected region and its neighboring regions. This embodiment dynamically and intelligently determines the currently most-needed updated deformation region (i.e., the affected region) by fusing three information sources (real-time tool position, intensity difference of consecutive image frames, and stress region estimated based on a physical model). The union of these three information can comprehensively capture local deformations caused by surgical operations or physiological factors.
[0068] Combining incremental update with intelligent region determination enables the navigation system to quickly respond to local changes during surgery. While significantly improving the real-time performance and computational efficiency of deformation field updates, by accurately defining the affected region, it ensures the pertinence and effectiveness of the update, so that the navigation system can continuously provide precise guidance information highly matched to the current organ state.
[0069] In some embodiments, recalculating the non-rigid deformation field within the affected region and its immediate neighboring regions using the incremental calculation strategy includes: Fix the deformation values at the boundaries of the affected region as boundary conditions; Inside the affected region, quickly update the non-rigid deformation field parameters by performing a finite number of local optimization iterations based on gradient descent to minimize a loss function that includes image similarity, deformation field smoothness, physical consistency, and correspondence compliance within the affected region, and use the current global non-rigid deformation field as the initial state; The iterative update further includes an error accumulation control mechanism, including: performing global non-rigid deformation field optimization periodically or when it is detected that the accumulated error exceeds a preset error threshold, so as to re-minimize the loss function for training the second deep learning network over the entire target organ region.
[0070] This embodiment adopts local optimization iteration based on gradient descent to quickly fine-tune the non-rigid deformation field parameters inside the affected area, and uses the current global field as the starting point for optimization, ensuring the continuity of the update, and further refining the specific execution method of incremental calculation and the long-term accuracy guarantee method. In particular, this embodiment introduces an error accumulation control mechanism to correct the possible accumulated deviation in the incremental update process through periodic or error threshold-triggered global optimization.
[0071] Through the above combination strategy including local rapid optimization and global correction, an effective balance between real-time performance and long-term accuracy is achieved. Local optimization ensures a rapid response to intraoperative dynamics, while the error control mechanism prevents the gradual drift of navigation accuracy, ensuring the continuous reliability and accuracy of navigation information throughout the surgical process and improving the navigation robustness in complex long-term surgical scenarios.
[0072] Step 5: Apply the currently updated non-rigid deformation field to the preoperative three-dimensional model to generate a deformed three-dimensional model, and the state of the target organ in the deformed three-dimensional model matches the state of the target organ in the real-time intraoperative image.
[0073] Apply the non-rigid deformation field containing the latest dynamic changes obtained in the previous step to the initially established high-precision preoperative three-dimensional model for information fusion. The deformed three-dimensional model not only retains the detailed preoperative anatomical information but also accurately reflects the true shape and position of the current intraoperative organ in three dimensions.
[0074] Step 6: Map the position of the end effector of the surgically robot being tracked in real time into the coordinate system of the deformed three-dimensional model.
[0075] Accurately locate the real-time position information of the surgical robot tool tracked in Step 4 into the space of the deformed three-dimensional model generated in Step 5 through coordinate transformation, thereby establishing an accurate correspondence between the physical world (robot tool) and the digital world (deformed model).
[0076] Step 7: Generate navigation instructions based on the mapped position of the tool, the target lesion area, and the risk structures to be avoided in the deformed three-dimensional model to guide the surgical robot to perform resection operations.
[0077] Step 7 is the final output stage of the navigation system. The system combines the precise position of the current tool on the deformed model and information such as the tumor boundary, safety margin, and important blood vessels contained in the model to generate intuitive navigation instructions, transforming the results of complex technical processing into clinically useful decision-making support information for doctors. Navigation instructions (such as visual cues, distance warnings, etc.) directly assist doctors in judging the relationship between the tool and key structures, guiding them to control the robot for safe and precise resection, thereby improving the surgical accuracy and safety.
[0078] Through innovative multi-modal registration and dynamic update strategies, this application can accurately compensate for the complex deformation of soft tissues during surgery, significantly improving the navigation accuracy. The above solution effectively integrates pre-operative high-resolution images and intra-operative real-time image information, overcomes the registration difficulties brought by modal differences, and ensures the timeliness and long-term accuracy of navigation information throughout the surgical process through a real-time incremental update and error control mechanism. This solution can provide more reliable and precise guidance for surgical robots, helping to improve the safety and effectiveness of complex soft tissue resection surgeries.
[0079] Next, a specific example of robot-assisted liver tumor resection is used to elaborate in detail on an application example of this solution.
[0080] First step: Pre-operative data acquisition and three-dimensional reconstruction.
[0081] In this example, the patient needs to undergo at least two modalities of medical image scans before surgery. Specifically, magnetic resonance imaging (MRI) scans are performed to obtain the MRI T2-weighted image as shown in Figure 2(a). T2WI can clearly show the tumor lesions in the liver parenchyma. At the same time, contrast-enhanced CT scans are performed to obtain the CT angiography (CTA) image as shown in Figure 2(b), which is used to accurately display the arterial, portal vein, and hepatic vein systems inside the liver.
[0082] After obtaining the original DICOM - format MRI T2WI and CTA image data, image segmentation is performed using medical image - processing software (such as the open - source software 3D Slicer or ITK - SNAP, or dedicated commercial software). Based on the segmentation results, a high - precision preoperative 3D model is reconstructed using surface rendering or volume rendering techniques, as shown in Figure 2(c). This model visually presents the spatial relationships among the liver, tumors, and blood vessels. Different colors are used for the liver parenchyma itself (such as the light - green and yellow regions in the figure) to distinguish different liver segments or regions, and the target tumor lesions are clearly marked (such as the solid red structure embedded in the yellow region in Figure 2(c)). At the same time, the main structures at risk of avoidance (OARs), especially the intra - hepatic vascular system (such as the portal vein and hepatic vein branches), are also distinguished and displayed in eye - catching colors (such as the cyan and purple tubular structures visible in Figure 2(c)). This model is stored in a standard format (such as STL or VTK), and virtual surgical planning is performed by surgeons on this model to determine the expected resection boundary and safety margin.
[0083] Step 2: Intra - operative image acquisition.
[0084] At the beginning of the operation, after the patient is anesthetized, pneumoperitoneum is established laparoscopically, and robotic surgical instruments and an endoscope are inserted. In this example, three - dimensional intraoperative ultrasound (3D iUS) is used as the intraoperative real - time imaging modality, as Figure 3 shown. An ultrasound probe that is spatially - located and tracked and connected to a navigation system (for example, operated by an assistant or integrated on a robotic arm) can be used to scan the liver surface. Through specific scanning methods (such as sector scanning, linear translation scanning) and reconstruction algorithms (integrated in the ultrasound device or navigation workstation), three - dimensional ultrasound image volume data of the current liver region are obtained in real - time or near - real - time. This iUS data reflects the current actual shape and internal echo structure of the liver under the influence of pneumoperitoneum establishment, organ exposure, and possible surgical operations.
[0085] Step 3: Real - time tracking of surgical robot tools.
[0086] During the operation, an optical tracking system (e.g., an infrared binocular camera system installed on the ceiling of the operating room, such as the NDI Polaris Spectra) is used to track in real time the reflective marker balls (Markers) fixed on the end tool of the surgical robot (e.g., an electrosurgical knife, a harmonic scalpel, or a grasping forceps). Through the principle of triangulation, the navigation system can accurately calculate the three-dimensional position and orientation of the robot tool (especially its tip or point of action) in the coordinate system of the tracking system at a high frequency (e.g., >20Hz). Through the pre-calibration of the robot-tracker-patient space, the position of the tool can be converted in real time to the patient image coordinate system and the subsequent deformed model coordinate system. Figure 3 Shows a schematic diagram of a multi-planar reconstruction (MPR) navigation interface based on three-dimensional intraoperative ultrasound (3D iUS) data. Figure 3 Shows the ultrasound images on the standard acquisition plane and any reconstructed plane, and the position of the surgically instrument being tracked in real time is fused and displayed. Step 4: Prediction of the initial non-rigid deformation field.
[0087] This step is the core of achieving the accurate alignment between the preoperative plan and the intraoperative reality. It includes the following two sub-steps d1 and d2.
[0088] Sub-step d1, using the first deep learning network to extract cross-modal features and establish correspondences.
[0089] In this example, the network adopts a Siamese Network structure, which includes two convolutional neural network (CNN) encoders with shared weights.
[0090] Training phase (offline): Use a dataset containing paired preoperative image patches (extracted from MRI T2WI, e.g., 32x32x32 voxels) and intraoperative iUS image patches for training. The image patch pairs are labeled Y (Y = 0 similar, Y = 1 dissimilar) according to whether they represent the same anatomical location (e.g., vascular bifurcation points, tumor boundary feature points, etc. annotated by experts). The network is trained through the contrast loss function optimized as above. The training objective is to make the cross-modal image patches from the same anatomical location close in the feature space and far apart at different locations.
[0091] Application phase (intraoperative): Input the currently acquired intraoperative 3D iUS image and preoperative image (MRI / CTA) into the trained Siamese Network. By sliding a window or sampling key regions in the two images and calculating the feature vector distance, several pairs of cross-modal image patches with high feature similarity (small D_W) can be found. The central points of the positions of these high-similarity pairs constitute the initial correspondence. The most reliable part (e.g., 10 - 30 pairs) can be selected through a confidence score (such as based on the D_W value) for subsequent steps.
[0092] Sub-step d2: Predict the initial non-rigid deformation field φ using the second deep learning network.
[0093] In this example, the network adopts a structure based on an encoder-decoder (such as 3D U-Net) and incorporates an attention mechanism (specifically, using the Attention Gate module on the skip connection path of the decoder to focus on features from the encoder and d1).
[0094] The second deep learning network simultaneously receives the original preoperative image data (such as MRI T2WI, and possibly vascular information (from CTA)), the original intraoperative 3D iUS data, and the initial correspondence information obtained from step d1 (for example, it can be encoded as a sparse supervision signal or a conditional input).
[0095] The network outputs the predicted dense non-rigid deformation field φ, which represents the mapping relationship p = φ(X) from the preoperative coordinates X to the intraoperative coordinates p.
[0096] In training and inference, a predefined physical deformation model is adopted to provide physical consistency constraints. In this example, the model is a hyperelastic model based on the finite element method (FEM) (for example, using the Neo-Hookean constitutive law). The mesh of the model is generated according to the reconstructed preoperative three-dimensional liver model, and its material parameters (such as Young's modulus, Poisson's ratio) can be personalized based on the magnetic resonance elastography (MRE) data that may be obtained preoperatively for the patient, or set according to the average biomechanical properties of the liver.
[0097] The second deep learning network is trained (offline) using the above optimized total loss function L_total.
[0098] The weight values α, β, γ are preset factors, and δ is adaptively adjusted according to the confidence of the initial correspondence. For example, correspondences with high confidence contribute more to δ.
[0099] Initial prediction (intraoperative): Input the current preoperative and intraoperative images and correspondences into the trained second deep learning network, and quickly perform forward propagation to obtain an initial non-rigid deformation field.
[0100] Step 5: Update the non-rigid deformation field in real time.
[0101] Due to surgical operations (such as retraction, cutting) or physiological factors (such as residual motion caused by breathing), the liver will be further deformed, and it is necessary to update the deformation field in real time.
[0102] Trigger condition: Based on newly acquired 3D iUS frames and / or detected surgical operation information (such as robotic tool force feedback, position changes).
[0103] The affected area can be determined first. For example, calculate the area within a preset radius (e.g., 1 - 2 cm) around the tool tip (the first area); calculate the absolute difference in voxel intensity between the current iUS frame and the previous frame, and identify the area where the difference is greater than a threshold (e.g., gray value difference > 10) (the second area); use a simplified physical model or empirical rules to estimate the area where significant stress / strain is generated based on the contact force between the robotic tool and the liver (if available) or the known cutting position (the third area).
[0104] The final affected area is defined as the union of these three areas.
[0105] Fix the value of the current global deformation field at the boundary of the affected area as the boundary condition. The local optimization starts with the current global non - rigid deformation field as the initial state.
[0106] Inside the affected area, perform a finite number (e.g., 5 - 15 times) of local optimization iterations based on gradient descent (e.g., using the Adam optimizer) to quickly update the deformation field parameters within this area (e.g., if φ is represented by a deformable mesh, update the mesh node positions; if φ is the direct output of the network, optimize a small correction field added to the network output).
[0107] This local optimization iteration aims to optimize the deformation field parameters to improve image matching, smoothness, and physical consistency, etc. within the affected area. Its objective function is formally similar to L_total during training, but is only calculated within this local area.
[0108] The system monitors the cumulative error (e.g., periodically re - evaluate the global L_total, or track the drift of key points), and periodically (e.g., every 30 seconds or 50 incremental updates), or when it detects that the cumulative error exceeds a preset threshold, perform a global non - rigid deformation field optimization. This can be achieved by re - running a full inference of the second deep learning network (sub - step d2) to refresh the entire deformation field, or by performing several global optimization iterations over the entire organ area.
[0109] Through the above steps, the continuously updated non - rigid deformation field is obtained.
[0110] Step 6: Apply the deformation field to generate the deformed model.
[0111] Apply the current updated non - rigid deformation field to the preoperative 3D model. Specifically, for each vertex V on the model pre , calculate its corresponding intraoperative position V through the deformation field (interpolation)post , a deformed three-dimensional model is generated. This deformed model can match the current state of the target organ (the true shape of the liver) reflected by the intraoperative iUS images obtained in real time in terms of shape, position, and internal structure (the relative positions of tumors and blood vessels). Figure 4 Shows a schematic diagram of the deformed three-dimensional model according to an example of the present application.
[0112] Step 7, Tool mapping and navigation instruction generation The position of the end tool of the surgical robot being tracked in real time is precisely mapped into the coordinate system of the generated deformed three-dimensional model through coordinate transformation to obtain a virtual representation of the tool on the model. This virtual tool (such as Figure 5 the green structure shown in) can reflect the position and posture of the real surgical instrument in real time.
[0113] On the display interface of the navigation system, the deformed three-dimensional model, key anatomical structures, and the mapped virtual tool are fused and displayed. The specific visualization style of this navigation interface (such as color scheme, transparency, rendering details) may vary depending on the navigation software platform or display settings used. In Figure 5 the shown display interface, the deformed liver model is shown in semi-transparent dark red as a whole for observing the internal structure. The internal key vascular system is clearly presented. For example, the main vascular branches are shown in blue, and another group of blood vessels are shown in purple. The target lesion area (tumor) is represented as a prominent, solid red area, clearly showing its position and approximate shape inside the liver.
[0114] Based on the real-time position of the virtual tool and the information of the target lesion area and risk structures to be avoided (such as blood vessels) included in the deformed three-dimensional model, navigation instructions are generated.
[0115] Figure 5 Visual guidance and quantification information are also presented in the lower right corner, and a list of distances (Pointer distance) from the tool tip to predefined key structures can be displayed in real time.
[0116] When the virtual tool approaches important blood vessels or other structures that need to be avoided (for example, enters a preset dangerous range), a clear warning (such as turning red) can be given in the distance list.
[0117] The surgeon can guide the surgical robot to perform precise resection operations according to these navigation instructions that integrate three-dimensional anatomical information, real-time tool position, and quantification data. For example, precisely stripping tissue, or making cuts at positions that meet safety requirements from both the tumor and important blood vessels.
[0118] Through the above steps, the method described in this embodiment can provide high-precision image-guided navigation for robot-assisted liver tumor resection, effectively compensate for intraoperative soft tissue deformation, fuse multi-modal image information, and ensure the real-time performance and robustness of the navigation, thereby assisting surgeons to achieve safer and more precise surgical operations.
[0119] Figure 6 For the electronic device provided by at least one embodiment of the present application, the device includes a memory and a processor. The memory is used to store computer instructions that can run on the processor, and the processor is used to implement the surgical robot resection navigation method based on multi-modal image registration described in any embodiment or implementation manner of the present application when executing the computer instructions.
[0120] At least one embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the surgical robot resection navigation method based on multi-modal image registration described in any embodiment or implementation manner of the present application.
[0121] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0122] The various embodiments in this specification are described in a progressive manner. The same or similar parts among the various embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiments of the data processing device, since it is basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0123] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0124] Although this specification contains many specific implementation details, these should not be construed as limiting the scope of any invention or the scope of what is claimed, but rather as mainly describing the features of specific embodiments of a particular invention. Certain features described in multiple embodiments within this specification can also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may operate in certain combinations as described above and even be initially claimed as such, one or more features from a claimed combination can in some cases be removed from that combination, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination.
[0125] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially, or that all illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0126] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the drawings are not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0127] The above description is only the preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification should be included within the scope protected by one or more embodiments of this specification.
Claims
1. A surgical robot resection navigation method based on multimodal image registration, characterized in that: The method comprises: Acquire a preoperative medical image of the patient including at least a first modality image and a second modality image, and reconstruct a preoperative three-dimensional model including a target lesion area and a risk structure to be avoided based on the preoperative image data; Acquiring at least one intraoperative image of the target organ region of the patient in a different modality than the preoperative medical image; Using a first deep learning network, extracting cross-modal invariant features from preoperative medical images and intraoperative images, establishing at least one set of initial correspondences, and using a second deep learning network, predicting an initial non-rigid deformation field based on the preoperative medical images, the intraoperative images, and the extracted initial correspondences, the non-rigid deformation field being used to represent the geometric deformation of the target organ from the preoperative state to the intraoperative state; During the operation, the position of the surgical robot's end tool is tracked in real time, and based on the real-time acquired intraoperative images and surgical operation information, an incremental calculation strategy is used to update the non-rigid deformation field in real time; Applying the currently updated non-rigid deformation field to the preoperative three-dimensional model to generate a deformed three-dimensional model, wherein the state of the target organ in the deformed three-dimensional model matches the state of the target organ in the real-time intraoperative image; Mapping the real-time tracked position of the surgical robot end tool into the coordinate system of the deformed three-dimensional model; Based on the mapped tool positions and the target lesion area and risk structures to be avoided of the deformed three-dimensional model, navigation instructions are generated to guide the surgical robot to perform the resection operation.
2. The method according to claim 1, characterized in that The first deep learning network is a twin network structure with an encoder with shared weights, which is used to map image patches of different modalities to the same feature space. The first deep learning network is trained using the following contrast loss function: , Among them, Y is the label of whether the input preoperative image block and intraoperative image block belong to the same anatomical position, Y = 0 means that the image block pair is similar, Y = 1 means that the image block pair is not similar, D_W represents the Euclidean distance between the two feature vectors output by the twin network encoder, m is the preset boundary parameter, and m is a positive number.
3. The method according to claim 1, characterized in that The second deep learning network has an encoder-decoder structure and is combined with an attention mechanism, wherein the attention mechanism is used to adaptively focus on areas corresponding to encoder features or cross-modal features that contribute more to deformation prediction of the current position during the decoding stage.
4. The method according to claim 3, characterized in that The second deep learning network is trained with the following loss function: , Among them, L_gradient_similarity is the image gradient similarity measurement term, which is used to measure the similarity between the intraoperative image obtained after applying the predicted non-rigid deformation field to the preoperative medical image and the actual intraoperative image; L_smoothness is the deformation field smoothness regularization term, which is used to penalize the non-smooth changes in the predicted non-rigid deformation field; L_physics is the physical consistency regularization term determined based on the predefined physical deformation model; L_correspondence is the initial correspondence supervision term, which is used to use the initial correspondence to supervise the learning of the second deep learning network; α, β, γ and δ are weight factors, α, β, γ are preset factors, and δ is adaptively adjusted according to the confidence of the initial correspondence.
5. The method according to claim 1, characterized in that: The image gradient similarity measure L_gradient_similarity is determined according to the following formula: , Where ∫ represents the integral over the entire target organ region Ω, X represents a point in the preoperative medical image, It means that the preoperative point X is mapped to the corresponding point in the intraoperative coordinate system through the non-rigid deformation field. Represents the midpoint of the preoperative medical image after deformation in the intraoperative coordinate system The gradient at Indicates the midpoint of the intraoperative image in the intraoperative coordinate system The gradient at , || *||² represents the square of the L2 norm of *.
6. The method according to claim 4, characterized in that The predefined physical deformation model is a hyperelastic model based on the finite element method, and the model parameters are determined based on the preoperative medical image. The physical consistency regularization term is determined according to the following formula: , where ∫ represents the integration over the entire target organ area Ω, is the deformation gradient tensor, , I represents the unit tensor, represents the predicted non-rigid deformation field, represents the gradient of the non-rigid deformation field, is the strain energy measurement function.
7. The method according to claim 4, characterized in that The incremental calculation strategy is used to recalculate the non-rigid deformation field within the affected area and its immediate neighboring area, and the affected area is determined according to the following method: According to the current position of the end tool of the surgical robot tracked in real time, determining a preset spatial range around the tip of the tool as a first area; Calculating the local image intensity difference between consecutive intraoperative image data frames, and identifying a region where the local image intensity difference exceeds a preset difference threshold as a second region; Using a physical deformation model, the third region where significant stress is generated is estimated based on the interaction force between the surgical robot end tool and the target organ tissue; A union of the first region, the second region, and the third region is determined as an affected region.
8. The method according to claim 7, characterized in that The incremental calculation strategy is used to recalculate the non-rigid deformation field in the affected area and its immediate vicinity, including: Fix the deformation value at the boundary of the affected area as the boundary condition; Within the affected area, the non-rigid deformation field parameters are quickly updated by performing a limited number of local optimization iterations based on gradient descent to minimize a loss function that includes image similarity, deformation field smoothness, physical consistency, and correspondence within the affected area, and the current global non-rigid deformation field is used as the initial state; The iterative update also includes an error accumulation control mechanism, including: periodically or when it is detected that the accumulated error exceeds a preset error threshold, performing global non-rigid deformation field optimization to re-minimize the loss function used to train the second deep learning network over the entire target organ area.
9. The method according to claim 1, characterized in that: The first modality is a magnetic resonance imaging (MRI) T2-weighted image, the second modality is a CT angiography (CTA) image, and the intraoperative image is a three-dimensional intraoperative ultrasound (iUS) image.
Citation Information
Patent Citations
Gradient distribution-based non-rigid medical image registration method
CN108053431A
Real-time non-rigid registration method and system for surgical navigation image based on deep learning
CN116485850A
Spine three-dimensional visualization surgical navigation system based on mixed reality
CN116492052A
Three-dimensional imaging reconstruction system in minimally invasive interventional operation based on adaptive neural network
CN119027585A
Real-time registration method from preoperative model to intraoperative point cloud based on deep learning
CN119090929A
Cited By
Auxiliary three-dimensional imaging system and method for complex chest trauma operation
CN120284470A
A complex chest trauma surgery auxiliary stereoscopic imaging system and method
CN120284470B
Anesthesia retardation target positioning method and system based on multi-source data
CN120411248A
Medical equipment automatic control positioning method based on multi-modal data
CN120436780A
Medical equipment automated control and positioning system based on multimodal data
CN120436780B