Surgical robot resection navigation method based on multi-modal image registration

By combining deep learning and physical models to create a multimodal image registration method, the problems of deformation and modal differences in robot-assisted soft tissue resection surgery were solved, enabling real-time and accurate navigation information updates and improving surgical precision and safety.

CN120070423BActive Publication Date: 2026-01-02BEIJING ROSSUM ROBOT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510528186.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2026-01-02
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Existing surgical navigation methods are insufficient in robot-assisted soft tissue resection surgery due to difficulties in addressing intraoperative soft tissue deformation, multimodal image registration, and real-time and robust requirements, resulting in inadequate navigation accuracy and failure to meet the needs of high-precision navigation.

Method used

A deep learning-based multimodal image registration method is adopted, which uses Siamese networks and encoder-decoder structures to extract cross-modal invariant features. Combined with physical model constraints, the non-rigid deformation field is updated in real time. Through incremental calculation strategy and error accumulation control mechanism, accurate compensation of soft tissue deformation and real-time updating of navigation information are achieved.

Benefits of technology

It significantly improves the accuracy and safety of surgical robot navigation, enabling accurate identification of lesion boundaries and avoidance structures, achieving more thorough lesion resection, reducing the risk of damage to important structures, and adapting to practical clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070423B_ABST
    Figure CN120070423B_ABST
Patent Text Reader

Abstract

Disclosed is a surgical robot resection navigation method based on multi-modal image registration. The method comprises: acquiring preoperative multi-modal medical images of a patient and reconstructing a three-dimensional model containing a lesion and a dangerous structure; acquiring real-time intraoperative images; using a first deep learning network to extract cross-modal invariant features from the preoperative and intraoperative images and establish an initial correspondence relationship; using a second deep learning network, based on the preoperative / intraoperative images and the initial correspondence relationship, and combining physical model constraints, to predict an initial non-rigid deformation field; tracking a surgical tool in real time, and using an incremental calculation strategy to update the non-rigid deformation field in real time; applying the updated deformation field to the preoperative three-dimensional model to generate a deformed model; mapping the tool position to the deformed model and generating navigation instructions. The application can accurately compensate for soft tissue deformation, effectively fuse multi-modal image information, and dynamically update navigation information in real time, significantly improving the accuracy and safety of robotic surgery.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image data processing, and in particular to a surgical robot resection navigation method based on multi-modal image registration. BACKGROUND

[0002] In recent years, robot-assisted minimally invasive surgery (RAMIS) has been widely applied and developed in the field of surgery. Compared with traditional open surgery and conventional laparoscopic surgery, a surgical robot system (such as the Da Vinci surgical system) can provide a three-dimensional high-definition magnified field of view, a flexible operating wrist, and the ability to filter out hand tremors for the surgeon, thereby enabling more precise and stable operation in a limited surgical space, reducing patient trauma, and shortening postoperative recovery time.

[0003] In precise resection surgeries such as tumor resection, in order to maximize the resection of diseased tissue while maximizing the preservation of healthy organ function and avoiding damage to critical structures such as blood vessels, nerves, and other key structures (i.e., Organs at Risk, OARs), the precision and safety of the surgery are crucial. For this reason, image-guided navigation (IGN) technology has been introduced into robot-assisted surgery. The core idea is to accurately map the rich anatomical information contained in preoperative high-resolution medical images (such as magnetic resonance imaging MRI, computed tomography CT, etc.) and surgical planning (such as tumor boundaries, safe resection margins, and OARs locations) into the surgical field, providing real-time navigation information for the surgeon to assist in precise operation of the surgical robot.

[0004] A typical image-guided navigation process generally includes: 1) preoperatively segmenting lesions, reconstructing three-dimensional models, and planning surgical paths using high-resolution images (MRI, CT, etc.); 2) intraoperatively aligning the real-time position of the patient with the preoperative data through a specific method (such as body surface marker points, intraoperative imaging, etc.); 3) tracking the position of surgical instruments in real time and displaying the relationship between the instruments and the anatomical structures, as well as the planned path, on the navigation interface.

[0005] However, when applying existing image-guided navigation technology to robot-assisted soft tissue organ (such as liver, kidney, brain, prostate, etc.) resection surgery, there are serious technical challenges:

[0006] I. Soft tissue deformation in surgery: Unlike hard tissues such as bones, soft tissue organs undergo significant and nonlinear deformation during surgery due to gravity, pneumoperitoneum pressure (laparoscopic surgery), patient respiratory motion, pulling by surgical instruments, compression, and even the operation of cutting itself. This deformation can cause significant deviations between preoperatively planned tumor boundaries, resection paths, and OARs positions and the actual intraoperative situation, severely affecting navigation accuracy and even potentially leading to incomplete resection or damage to critical structures. Therefore, simply rigidly aligning preoperative static models to intraoperative conditions is far from sufficient, and accurate estimation and compensation of soft tissue deformation are necessary.

[0007] II. Multi-modal image registration challenges: There are significant differences between preoperative planning-dependent high-resolution images (such as MRI providing excellent soft tissue contrast, CT providing good bone and density information, and PET providing functional information) and intraoperative real-time available or commonly used imaging modalities (such as intraoperative ultrasound iUS providing real-time internal structure information but with lower resolution and artifacts, and endoscopy providing surface color texture information but lacking depth and internal information). These differences are reflected in imaging principles, image intensity distribution, resolution, signal-to-noise ratio, geometric distortion, and artifacts. How to establish an accurate correspondence between these multi-modal images with different appearances and characteristics and achieve robust registration (especially non-rigid registration that compensates for deformation) is a core difficulty in image-guided navigation. Traditional methods based on feature points or surface matching often lack accuracy due to difficulty in feature extraction or sparsity of features; methods based on image intensity (such as mutual information) are easily affected by modal differences and intensity changes.

[0008] III. Real-time and robustness requirements: Surgical navigation information must be provided to the surgeon in real-time or near real-time to effectively guide surgical operations. This means that complex non-rigid registration algorithms need to meet clinical accuracy requirements while having high enough computational efficiency. In addition, the surgical environment is complex and variable, and the navigation system must have certain robustness to image noise, artifacts, occlusions, and unexpected tissue motion.

[0009] Existing surgical navigation methods still have deficiencies in simultaneously solving the problems of large soft tissue deformation, multi-modal data fusion, and real-time robust registration, and are difficult to meet the urgent need for high-precision navigation in robot-assisted precise resection surgery.

[0010] Therefore, there is an urgent need to develop a new method that can effectively fuse multi-modal image information, accurately estimate and compensate for soft tissue intraoperative deformation, and thus provide accurate and reliable navigation information for surgical robots to improve the precision and safety of robot-assisted resection surgery. SUMMARY

[0011] The application provides a surgical robot resection navigation method based on multi-modal image registration, which can accurately estimate and compensate for soft tissue deformation in operation in real time, thereby providing accurate and reliable navigation information for a surgical robot.

[0012] According to one embodiment of the application, a surgical robot resection navigation method based on multi-modal image registration is provided, and the method comprises:

[0013] Obtaining preoperative medical images of a patient, which at least contain first modal images and second modal images, and reconstructing a preoperative three-dimensional model containing a target lesion region and organs at risk (OARs) based on the preoperative image data;

[0014] Obtaining at least one intraoperative image of a target organ region of the patient, which is different from the preoperative medical images in modal;

[0015] Using a first deep learning network to extract cross-modal invariant features from the preoperative medical images and the intraoperative images, establish at least one initial correspondence relationship, and using a second deep learning network to predict an initial non-rigid deformation field based on the preoperative medical images, the intraoperative images and the extracted initial correspondence relationship, the non-rigid deformation field being used to represent the geometric deformation of the target organ from a preoperative state to an intraoperative state;

[0016] In the operation process, the position of the end tool of the surgical robot is tracked in real time, and based on the real-time acquired intraoperative images and operation information, an incremental calculation strategy is adopted to update the non-rigid deformation field in real time;

[0017] The current updated non-rigid deformation field is applied to the preoperative three-dimensional model to generate a deformed three-dimensional model, and the state of the target organ in the deformed three-dimensional model matches the state of the target organ in the real-time intraoperative images;

[0018] The position of the real-time tracked end tool of the surgical robot is mapped into the coordinate system of the deformed three-dimensional model;

[0019] Based on the mapped position of the tool and the target lesion region and the organs at risk of the deformed three-dimensional model, a navigation instruction is generated to guide the surgical robot to perform a resection operation.

[0020] In some embodiments, the first deep learning network is a Siamese Network structure with a shared weight encoder for mapping image blocks of different modalities to the same feature space, and the first deep learning network is trained by the following contrastive loss function:

[0021] ,

[0022] wherein Y is a label indicating whether the input preoperative image patch and the intraoperative image patch belong to the same anatomical location, Y = 0 indicates that the image patch pair is similar, Y = 1 indicates that the image patch pair is not similar, D W represents the Euclidean distance between the two feature vectors output by the twin network encoder, m is a preset boundary parameter, and m is a positive number.

[0023] In some embodiments, the second deep learning network has an encoder-decoder structure and incorporates an attention mechanism for adaptively focusing on the regions corresponding to the encoder features or cross-modal features that contribute more to the current position deformation prediction in the decoding stage.

[0024] In some embodiments, the second deep learning network is trained by the following loss function:

[0025]

[0026] wherein L_gradient_similarity is an image gradient similarity measurement term for measuring the similarity between the intraoperative image obtained by applying the predicted non-rigid deformation field to the preoperative medical image and the actual intraoperative image; L_smoothness is a deformation field smoothness regularization term for penalizing non-smooth changes in the predicted non-rigid deformation field; L_physics is a physics consistency regularization term determined based on a pre-defined physical deformation model; L_correspondence is an initial correspondence supervision term for supervising the learning of the second deep learning network using the initial correspondence; and α, β, γ, and δ are weight factors, α, β, and γ are preset factors, and δ is adaptively adjusted according to the confidence of the initial correspondence.

[0027] In some embodiments, the image gradient similarity measurement term L_gradient_similarity is determined according to the following formula:

[0028]

[0029] wherein ∫ represents integration over the entire target organ region Ω, X represents a point in the preoperative medical image, represents the corresponding point of the preoperative point X mapped into the intraoperative coordinate system by the non-rigid deformation field, represents the gradient of the point in the preoperative medical image deformed in the intraoperative coordinate system, represents the gradient of the point in the intraoperative image in the intraoperative coordinate system, and || *|| 2 represents the square of the L2 norm of *.

[0030] In some embodiments, the predefined physical deformation model is a hyperelastic model based on finite element method, and the model parameters are determined based on preoperative medical images, and the physical consistency regularization term is determined according to the following formula:

[0031] ,

[0032] where ∫ represents integration over the entire target organ region Ω, is a deformation gradient tensor, , I represents a unit tensor, represents a predicted non-rigid deformation field, represents a gradient of the non-rigid deformation field, is a strain energy function.

[0033] In some embodiments, the incremental calculation strategy is used to redefine the non-rigid deformation field in the affected region and its direct adjacent region, and the affected region is determined according to the following method:

[0034] According to the current position of the end tool of the real-time tracking surgical robot, a first region within a preset spatial range around the tool tip is determined;

[0035] Calculate the local image intensity difference between consecutive intraoperative image data frames, and identify a second region where the local image intensity difference exceeds a preset difference threshold;

[0036] Using the physical deformation model, estimate a third region where significant stress is generated according to the interaction force between the end tool of the surgical robot and the body tissue;

[0037] The union of the first region, the second region and the third region is determined as the affected region.

[0038] In some embodiments, redefining the non-rigid deformation field in the affected region and its direct adjacent region using the incremental calculation strategy includes:

[0039] Fix the deformation value at the boundary of the affected region as a boundary condition;

[0040] Within the affected region, quickly update the non-rigid deformation field parameters by performing a limited number of local optimization iterations based on gradient descent to minimize a loss function including image similarity, deformation field smoothness, physical consistency and correspondence compliance within the affected region, and take the current global non-rigid deformation field as the initial state;

[0041] The iterative update further comprises an error accumulation control mechanism, comprising: periodically, or when detecting that accumulated error exceeds a preset error threshold, performing global non-rigid deformation field optimization to re-minimize a loss function for training the second deep learning network over the entire target organ region.

[0042] In some embodiments, the first modality is a magnetic resonance imaging (MRI) T2-weighted image, the second modality is a CT angiography (CTA) image, and the intraoperative image is a three-dimensional intraoperative ultrasound (iUS) image.

[0043] According to the surgical robot resection navigation method based on multi-modal image registration provided in the present application, by introducing the innovative technical solution, the key technical bottlenecks existing in the prior art of robot-assisted soft tissue surgery navigation are effectively overcome, and remarkable beneficial effects are achieved:

[0044] Firstly, the present application can accurately estimate and compensate for the complex deformation of soft tissue during the surgery by adopting a non-rigid registration strategy combining deep learning feature extraction, physical model constraint and real-time intraoperative image information. The first deep learning network is used to extract cross-modal invariant features and establish an initial correspondence, providing reliable driving information for deformation estimation, while the second deep learning network is used to predict the deformation field, and a physical consistency regularization term based on a biomechanical model is introduced to constrain it, ensuring the physical reasonableness and smoothness of the predicted deformation. This combination of data-driven and model-driven methods significantly improves the registration accuracy, enabling accurate mapping of preoperative plans such as tumor boundaries and critical structures onto the current deformed organ, laying the foundation for precise resection.

[0045] Secondly, the present application effectively solves the registration problem of significant modality differences between preoperative high-resolution images (such as MRI / CT) and intraoperative real-time images (such as iUS). The first deep learning network (such as a twin network) is specifically used to learn the shared feature representation between cross-modal images, overcoming the limitations of traditional methods in establishing reliable correspondence under different imaging principles and image characteristics. In addition, the design of the image gradient similarity measurement term in the loss function further enhances the robustness to modality differences, enabling more reliable multi-modal data fusion.

[0046] Thirdly, the incremental calculation strategy and error accumulation control mechanism proposed in the present application give the navigation system the ability to update in real time and maintain accuracy over a long period. By concentrating computing resources on the local area affected by surgical operations, the non-rigid deformation field is updated quickly and iteratively, enabling timely response to the continuous changes of intraoperative tissue. At the same time, the global optimization mechanism triggered periodically or based on error threshold, effectively prevents error accumulation caused by multiple local updates, ensuring the accuracy and reliability of the navigation information throughout the surgery.

[0047] By providing high-precision, real-time updated, and robust to modality difference navigation information, the present application can significantly improve the accuracy and safety of robot-assisted resection surgery. Doctors can more clearly and confidently identify target lesion boundaries, safe margins, and key anatomical structures that need to be avoided, thereby guiding the surgical robot to perform more precise cutting operations. This helps to achieve more complete lesion resection, maximally preserve healthy tissue, reduce the risk of intraoperative damage to adjacent important blood vessels and nerves, and is expected to improve the surgical effect and prognosis of patients. At the same time, the technical solutions proposed in the present application, especially the rapid prediction and incremental update strategy based on deep learning, have good computational efficiency potential and are suitable for clinical practical application. BRIEF DESCRIPTION OF DRAWINGS

[0048] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present specification and serve to explain the principles of the present specification.

[0049] Figure 1 A flowchart of a surgical robot resection navigation method based on multi-modal image registration according to an embodiment of the present application is shown.

[0050] FIG. 2(a) shows an example of a preoperative MRI T2-weighted image according to an example of the present application; FIG. 2(b) shows an example of a preoperative CT image according to an example of the present application; and FIG. 2(c) shows an example of a preoperative three-dimensional model according to an example of the present application.

[0051] Figure 3 A schematic diagram of a multi-planar reconstruction (MPR) navigation interface based on three-dimensional intraoperative ultrasound (3D iUS) data according to an example of the present application is shown.

[0052] Figure 4 A schematic diagram of a deformed three-dimensional model according to an example of the present application is shown.

[0053] Figure 5 A schematic diagram of a surgical robot resection navigation interface according to an example of the present application is shown.

[0054] Figure 6 is a structural schematic diagram of an electronic device according to at least one embodiment of the present application. DETAILED DESCRIPTION

[0055] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The description of the exemplary embodiments is intended to apply to any embodiment of the application, unless specifically stated otherwise. It is to be understood that other embodiments can be utilized, and structural or procedural changes can be made, without departing from the scope of the present application. Furthermore, where possible, like reference numbers can be used in the figures and can refer to the same or similar elements. The exemplary embodiments described herein are not meant to be limiting but rather are illustrative of the application.

[0056] Embodiments of the present application can be applied to a computer system / server, which can be operational with numerous other general purpose or special purpose computing system environments or configurations. Examples of well- known computing systems, environments, and / or configurations that can be suitable for use with computer system / server include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices, and the like.

[0057] Computer system / server can be described in the general context of computer system-executable instructions, such as program modules, being executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like, that perform particular tasks or implement particular abstract data types. Computer system / server can operate in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media including memory storage devices.

[0058] Figure 1 A flowchart of a surgical robot resection navigation method based on multi-modal image registration is shown according to an embodiment of the present application. The method provides accurate navigation guidance for surgical robots by fusing preoperative data containing detailed anatomical information and surgical planning with intraoperative dynamically changing organ states in real time and accurately through advanced image processing and registration techniques. As shown in the figure, the method includes steps 1-7.

[0059] Step 1, obtain preoperative medical images of a patient containing at least first modality images and second modality images, and reconstruct a target lesion region and OARs based on the preoperative image data.

[0060] Step 1 is the preparation phase. Before surgery, detailed images of the patient's target organ are acquired using at least two different high-resolution imaging techniques (e.g., MRI provides soft tissue details, CTA shows vascular structures). Subsequently, through image segmentation and three-dimensional reconstruction techniques, a digital three-dimensional anatomical model is generated, which contains the precise location, shape of the tumor, and the distribution of important blood vessels, nerves, and other dangerous structures around it.

[0061] The fusion of multi-modal images ensures the comprehensiveness of information, overcoming the limitations of a single modality.

[0062] Step 2, acquire at least one intraoperative image of the patient's target organ region in a different modality from the preoperative medical image.

[0063] During the operation, real-time image information reflecting the current "true state" of the organ can be acquired. Since the preoperative imaging equipment cannot be used during the operation, this step uses one or more imaging modalities suitable for intraoperative imaging (e.g., three-dimensional intraoperative ultrasound iUS) to acquire real-time images of the target organ region. This intraoperative modality is usually significantly different from the preoperative modality.

[0064] Intraoperative images provide a "target map" for subsequent registration, that is, the actual shape and position of the organ after being affected by factors related to surgery (such as pneumoperitoneum, gravity, operation), which serves as the basis for subsequent deformation estimation and compensation.

[0065] Step 3, using a first deep learning network, extracting cross-modal invariant features from the preoperative medical image and the intraoperative image, establishing at least one initial correspondence relationship, and using a second deep learning network, based on the preoperative medical image, the intraoperative image and the extracted initial correspondence relationship, predicting an initial non-rigid deformation field, the non-rigid deformation field representing the geometric deformation of the target organ from the preoperative state to the initial state in the operation.

[0066] Step 3 is the core computing step to achieve preoperative and intraoperative data alignment. First, a specially trained first deep learning network (such as a twin network) is used to overcome modality differences, automatically find and match common anatomical feature points in preoperative and intraoperative images, and establish a preliminary spatial correspondence relationship. Then, using a second deep learning network, combining the original image information and the extracted correspondence relationship, an initial non-rigid deformation field is comprehensively predicted. The deformation field mathematically describes the spatial distortion and displacement required for the organ to "deform" from the preoperative scan state to the current initial state in the operation.

[0067] Through the two-step deep learning strategy, the two major difficulties of cross-modal registration and complex non-rigid deformation estimation are effectively overcome. The first network focuses on robust feature matching, and the second network is responsible for generating an accurate and physically reasonable initial deformation estimate using this matching information and global image information, laying the foundation for subsequent real-time navigation.

[0068] In some embodiments, the first deep learning network is a Siamese Network structure with a shared-weight encoder for mapping image patches of different modalities to the same feature space, and the first deep learning network is trained by a contrastive loss function as follows:

[0069] ,

[0070] where Y is the label of whether the input preoperative image patch and intraoperative image patch belong to the same anatomical location, Y = 0 represents that the image patch pair is similar, Y = 1 represents that the image patch pair is not similar, D W represents the Euclidean distance between the two feature vectors output by the Siamese Network encoder, and m is a preset boundary parameter, which is a positive number.

[0071] This embodiment can effectively learn the invariant feature representation between cross-modality images (such as MRI / CT and iUS) by using a Siamese Network and a contrastive loss, thereby significantly improving the robustness and accuracy of establishing a reliable initial correspondence relationship under different imaging principles and image characteristics.

[0072] In some embodiments, the second deep learning network has an encoder-decoder structure and combines an attention mechanism, which is used to adaptively focus on the encoder features or cross-modality feature corresponding regions that contribute more to the current position deformation prediction in the decoding stage.

[0073] This embodiment integrates the attention mechanism into the encoder-decoder network for non-rigid deformation field prediction to enable it to fuse multi-source information. Based on this network structure combined with the attention mechanism, the accuracy and robustness of non-rigid deformation field prediction can be significantly improved. By intelligently focusing and fusing key information from the encoder and cross-modality correspondence, the network can more accurately capture and align important anatomical structures. This is especially helpful in improving the registration effect in areas with poor image quality or complex deformation.

[0074] In some embodiments, the second deep learning network is trained by the following loss function:

[0075]

[0076] L gradient similarity, L smoothness, L physics, and L correspondence, and the adaptive adjustment mechanism of the weight δ of the correspondence supervision term L correspondence. The loss function is used to train the second deep learning network.

[0077] According to the composite loss function of the embodiment, through the specific combination of the four key constraint terms of the image gradient similarity (L gradient similarity), the deformation field smoothness (L smoothness), the physical consistency based on the physical model (L physics), and the initial correspondence supervision (L correspondence), a collaborative optimization framework is formed. Moreover, the loss function introduces an adaptive adjustment mechanism for the weight δ of the correspondence supervision term, so that the network can intelligently evaluate and utilize the reliability of the initial correspondence information from the first network.

[0078] Based on the above composite loss function for training the second deep learning network, the comprehensive quality of the predicted non-rigid deformation field can be significantly improved. By optimizing multiple objectives simultaneously, it ensures that the prediction result not only aligns on the image level (gradient similarity), but also meets the physical rationality (smoothness and physical consistency), and fully utilizes the known feature correspondence information. In particular, the adaptive weight adjustment mechanism enhances the robustness of the system to possible noise or error correspondence output by the first deep learning network. This balanced and robust optimization strategy makes the generated deformation field more accurate, stable and consistent with clinical practice, significantly improving the reliability of subsequent surgical navigation.

[0079] In some embodiments, the image gradient similarity measure term L gradient similarity is determined according to the following formula:

[0080]

[0081] where ∫ represents integration over the entire target organ region Ω, X represents a point in the preoperative medical image, represents the corresponding point in the intraoperative coordinate system after the preoperative point X is mapped by the non-rigid deformation field, represents the gradient of the point in the intraoperative coordinate system after the preoperative medical image is deformed. denotes the gradient of the intra-operative image at the point denotes the square of the L2-norm of *.

[0082] The embodiment adopts a specific image gradient similarity measure term L_gradient_similarity to evaluate the matching degree of the deformed pre-operative image and the intra-operative image, by calculating and minimizing the squared L2-norm difference between the gradient vector fields of the two images, instead of directly comparing the image intensity values. This gradient-based strategy can effectively overcome the registration difficulty caused by the huge difference in intensity distribution, contrast and scale between the pre-operative high-resolution images (such as MRI / CT) and the intra-operative real-time ultrasound (iUS) images.

[0083] Therefore, by focusing on the gradient information that can better reflect the edges and textures of anatomical structures for alignment, this similarity measure term significantly improves the robustness and accuracy of cross-modality image registration, enabling the system to more reliably capture the correspondence of key anatomical structures, providing a more effective supervision signal for generating accurate non-rigid deformation fields, and ultimately improving the navigation accuracy.

[0084] In some embodiments, the pre-defined physical deformation model is a hyperelastic model based on the finite element method, and the model parameters are determined based on the pre-operative medical images, and the physical consistency regularization term is determined according to the following formula:

[0085] ,

[0086] where ∫ denotes integration over the entire target organ region Ω, is the deformation gradient tensor, , I denotes the unit tensor, denotes the predicted non-rigid deformation field, denotes the gradient of the non-rigid deformation field, is the strain energy function.

[0087] The embodiment constrains the prediction of the non-rigid deformation field by introducing a physical consistency regularization term L_physics based on biomechanical principles, and integrates the biomechanical strain energy corresponding to the predicted deformation (calculated according to a pre-defined hyperelastic model that can reflect the characteristics of the target soft tissue) over the entire organ region into the loss function for minimization.

[0088] The physical model-based constraint, which is significantly superior to simple smoothness or volume preservation requirements, ensures that the predicted deformation is not only mathematically continuous, but also physically consistent with the real mechanical behavior of soft tissue (e.g., the ability to resist excessive stretching or unreasonable compression). Therefore, through this implementation, it is possible to effectively avoid generating unrealistic distortions or deformations, so that the final predicted deformation field is more realistic and credible, significantly improving the accuracy of the navigation system in simulating the actual state of the organ during surgery.

[0089] L_smoothness can adopt a standard diffusion regularization term, such as penalizing the sum of squares of the gradient of the deformation displacement field.

[0090] L_correspondence can adopt an initial correspondence supervision term, such as calculating the mapping error between pairs of corresponding points.

[0091] Step 4, during the operation, the position of the end tool of the surgical robot is tracked in real time, and based on the real-time acquired intraoperative image and surgical operation information, an incremental calculation strategy is adopted to update the non-rigid deformation field in real time.

[0092] The present application introduces a real-time updating mechanism to adapt to the continuous dynamic deformation of soft tissue during surgery. On the one hand, the system accurately acquires the real-time spatial position of the surgical robot tool through a tracking system (such as an optical tracking system, an electromagnetic tracking system, an image / vision-based tracking system, etc.); on the other hand, according to the newly acquired intraoperative image and the perceived surgical operation (such as cutting, pulling), an efficient incremental calculation method is adopted to quickly recalculate and update the non-rigid deformation field, which is mainly concentrated in the affected local area.

[0093] Step 4 updates the deformation field in real time, so that the navigation system can dynamically adapt to the continuous change of the organ during the operation process, avoiding the accumulation of navigation errors caused by the use of outdated deformation information, ensuring the timeliness and accuracy of the navigation information, and effectively balancing the accuracy and calculation efficiency through the incremental calculation strategy.

[0094] In some embodiments, the incremental calculation strategy is used to recalculate the non-rigid deformation field in the affected area and its direct adjacent area, and the affected area is determined according to the following method:

[0095] According to the current position of the end tool of the surgical robot tracked in real time, the first area within a preset spatial range around the tool tip is determined;

[0096] Calculate the local image intensity difference between consecutive intraoperative image data frames, and identify the area where the local image intensity difference exceeds a preset difference threshold as the second area;

[0097] a third region of significant stress is estimated based on the interaction force between the surgical robot end-effector and the body tissue using a physical deformation model;

[0098] a union of the first region, the second region and the third region is determined as the affected region.

[0099] To cope with the continuous dynamic changes of soft tissue during surgery, an incremental computation strategy is introduced to update the non-rigid deformation field in real time. According to the present embodiment, the incremental computation strategy efficiently concentrates the computational resources on the determined affected region and its neighboring regions. The present embodiment dynamically and intelligently determines the deformation region that needs to be updated most (i.e. the affected region) by fusing three information sources (real-time tool position, intensity difference of consecutive image frames and stress region estimated based on physical model), the union of which can comprehensively capture the local deformation caused by surgical operation or physiological factors.

[0100] Combining incremental update and intelligent region determination, the navigation system can quickly respond to local changes during surgery, while significantly improving the real-time performance and computational efficiency of deformation field update, by accurately defining the affected region, ensuring the pertinence and effectiveness of the update, so that the navigation system can continuously provide accurate guidance information that matches the current organ state.

[0101] In some embodiments, the re-computing the non-rigid deformation field in the affected region and its direct neighboring regions using the incremental computation strategy comprises:

[0102] fixing the deformation values at the boundary of the affected region as boundary conditions;

[0103] inside the affected region, quickly updating the non-rigid deformation field parameters by performing a limited number of local optimization iterations based on gradient descent to minimize a loss function including image similarity, deformation field smoothness, physical consistency and correspondence compliance in the affected region, and taking the current global non-rigid deformation field as the initial state;

[0104] The iterative update further includes an error accumulation control mechanism, including: periodically, or when it is detected that the accumulated error exceeds a preset error threshold, performing global non-rigid deformation field optimization to re-minimize the loss function used to train the second deep learning network over the entire target organ region.

[0105] The embodiment adopts local optimization iteration based on gradient descent to fine-tune the non-rigid deformation field parameters within the affected area quickly, and takes the current global field as the starting point of optimization, ensuring the continuity of the update and further refining the specific execution mode and long-term precision guarantee mode of incremental calculation. In particular, the embodiment introduces an error accumulation control mechanism, which corrects the deviation that may accumulate during incremental updating through periodic or error threshold-based global optimization.

[0106] Through the above combination strategy of local rapid optimization and global correction, an effective balance between real-time performance and long-term accuracy is achieved. Local optimization ensures a quick response to intraoperative dynamics, while the error control mechanism prevents gradual drift of navigation accuracy, ensuring the continuous reliability and accuracy of navigation information throughout the entire surgical process, and improving the navigation robustness in complex long-time surgical scenarios.

[0107] Step 5: Apply the current updated non-rigid deformation field to the preoperative three-dimensional model to generate a deformed three-dimensional model, in which the target organ state matches that in the real-time intraoperative image.

[0108] Apply the non-rigid deformation field containing the latest dynamic changes obtained in the previous step to the initially established high-precision preoperative three-dimensional model for information fusion. The deformed three-dimensional model not only retains the detailed anatomical information of the preoperative model, but also accurately reflects the current intraoperative organ shape and position.

[0109] Step 6: Map the position of the real-time tracked surgical robot end tool to the coordinate system of the deformed three-dimensional model.

[0110] Through coordinate transformation, the real-time position information of the surgical robot tool tracked in step 4 is accurately positioned in the space of the deformed three-dimensional model generated in step 5, thereby establishing an accurate correspondence between the physical world (robot tool) and the digital world (deformed model).

[0111] Step 7: Based on the mapped tool position and the target lesion area and risk avoidance structure in the deformed three-dimensional model, generate navigation instructions to guide the surgical robot to perform resection operations.

[0112] Step 7 is the final output link of the navigation system. The system combines the accurate position of the current tool on the deformed model with the tumor boundary, safety margin, important blood vessels, and other information contained in the model to generate intuitive navigation instructions, converting complex technical processing results into useful clinical decision support information for doctors. Navigation instructions (such as visual cues, distance warnings, etc.) directly assist doctors in judging the relationship between the tool and critical structures, guiding them to control the robot for safe and precise resection, thereby improving surgical precision and safety.

[0113] The application can accurately compensate for the complex deformation of soft tissues during surgery through an innovative multi-modal registration and dynamic updating strategy, significantly improving navigation accuracy. The above scheme effectively integrates preoperative high-resolution image and intraoperative real-time image information, overcomes the registration difficulties caused by modal differences, and ensures the timeliness and long-term accuracy of navigation information during the entire surgical process through real-time incremental updating and error control mechanism. The scheme can provide more reliable and accurate guidance for surgical robots, and help improve the safety and effectiveness of complex soft tissue resection surgery.

[0114] An application example of the scheme will be described in detail below in conjunction with an example of robot-assisted liver tumor resection.

[0115] Step 1: Preoperative data acquisition and three-dimensional reconstruction.

[0116] In this example, the patient needs to undergo at least two modalities of medical image scanning before surgery. Specifically, magnetic resonance imaging (MRI) scanning is performed to obtain the MRI T2 weighted image as shown in Figure 2 (a). T2WI can clearly show the tumor lesions in the liver parenchyma. At the same time, enhanced CT scanning is performed to obtain the CT angiography (CTA) image as shown in Figure 2 (b), which is used to accurately display the arterial, portal vein and hepatic vein system inside the liver.

[0117] After obtaining the original DICOM format MRI T2WI and CTA image data, image segmentation is performed using medical image processing software (such as open source software 3D Slicer or ITK-SNAP, or dedicated commercial software). Based on the segmentation results, high-precision preoperative three-dimensional models are reconstructed using surface rendering (Surface Rendering) or volume rendering (Volume Rendering) techniques, as shown in Figure 2 (c). The model presents the spatial relationship between the liver, tumor and blood vessels in a visual manner. The liver parenchyma itself is distinguished by different colors (such as the light green and yellow areas in the figure) to distinguish different liver segments or regions, and the target tumor lesions are clearly marked (such as the solid red structure embedded in the yellow area in Figure 2 (c)). At the same time, the main risk avoidance structures (OARs), especially the intraparenchymal vascular system (such as the portal vein and hepatic vein branches), are also displayed using eye-catching colors (such as the blue and purple tubular structures visible in Figure 2 (c)). The model is stored in a standard format (such as STL or VTK), and a virtual surgery plan is made on the model by the surgeon to determine the expected resection boundary and safety margin.

[0118] Step 2: Intraoperative image acquisition.

[0119] The procedure starts after the patient is anesthetized, a pneumoperitoneum is established under laparoscopy, and the robotic surgical instruments and endoscope are inserted. In this example, three-dimensional intraoperative ultrasound (3D iUS) is used as the intraoperative real-time imaging modality, as shown in Figure 3 An ultrasound probe (e.g., operated by an assistant or integrated on a robotic arm) connected to the navigation system and spatially tracked can be used to perform a scan on the liver surface. Through specific scanning methods (e.g., fan-shaped scanning, linear translation scanning) and reconstruction algorithms (integrated in the ultrasound device or the navigation workstation), three-dimensional ultrasound image volume data of the current liver region is obtained in real-time or near real-time. This iUS data reflects the current actual shape and internal echo structure of the liver under the influence of the establishment of pneumoperitoneum, organ exposure, and possible surgical operations.

[0120] Step 3: Real-time tracking of the surgical robotic tool.

[0121] During the procedure, the optical tracking system (e.g., an infrared binocular camera system mounted on the operating room ceiling, such as NDI Polaris Spectra) is used to track the reflective marker balls (Markers) fixed on the end tool of the surgical robot (e.g., an electrotome, an ultrasonic knife, or a grasper). Through the principle of triangulation, the navigation system can accurately calculate the three-dimensional position and attitude of the robotic tool (especially its tip or action point) in the tracking system coordinate system at a high frequency (e.g., >20 Hz). Through the pre-performed robot-tracker-patient spatial calibration, the position of the tool can be converted into the patient image coordinate system and the subsequent deformation model coordinate system in real time. Figure 3 A multi-planar reconstruction (MPR) navigation interface schematic based on three-dimensional intraoperative ultrasound (3D iUS) data is shown. Figure 3 The ultrasound images on the standard acquisition plane and the arbitrary reconstruction plane are displayed, and the position of the real-time tracked surgical tool is fused and displayed

[0122] Step 4: Initial non-rigid deformation field prediction.

[0123] This step is the core of realizing the accurate alignment of preoperative planning and intraoperative reality. It includes the following two sub-steps d1 and d2.

[0124] Sub-step d1, cross-modal feature extraction and correspondence establishment using a first deep learning network.

[0125] In this example, the network adopts a Siamese Network structure, containing two convolutional neural network (CNN) encoders sharing weights.

[0126] Training phase (offline): Training is performed using a dataset containing pairs of preoperative image patches (extracted from MRI T2WI, e.g. 32x32x32 voxels) and intraoperative iUS image patches. The image patch pairs are assigned a label Y (Y=0 similar, Y=1 dissimilar) according to whether they represent the same anatomical location (e.g. a bifurcation point of a blood vessel, a tumor boundary landmark, etc. as annotated by an expert). The network is trained by optimizing the contrast loss function as above. The training goal is to let the cross-modality image patches from the same anatomical location be close in feature space, and those from different locations be far apart.

[0127] Application phase (intraoperative): The current acquired intraoperative 3D iUS image and preoperative image (MRI / CTA) are input into the trained twin network. By sliding a window or sampling key regions in both images and calculating the feature vector distance, several pairs of cross-modality image patches with high feature similarity (D_W small) can be found. The location center points of these high-similarity pairs constitute the initial correspondence. A confidence score (e.g. based on D_W value) can be used to filter out the most reliable part (e.g. 10-30 pairs) for the subsequent steps.

[0128] Sub-step d2, predict the initial non-rigid deformation field φ using a second deep learning network.

[0129] In this example, the network adopts an encoder-decoder (e.g. 3D U-Net) based structure, combined with an attention mechanism (specifically, an Attention Gate module is used on the skip connection path of the decoder, focusing on features from the encoder and d1).

[0130] The second deep learning network receives the original preoperative image data (e.g. MRI T2WI, and possibly vessel information (from CTA)), the original intraoperative 3D iUS data, and the initial correspondence information from step d1 (e.g. which can be encoded as a sparse supervision signal or conditional input) simultaneously.

[0131] The network outputs a predicted dense non-rigid deformation field φ, which represents the mapping from preoperative coordinates X to intraoperative coordinates p = φ(X).

[0132] In training and inference, a pre-defined physical deformation model is adopted to provide physical consistency constraints. In this example, the model is a hyperelastic model based on the finite element method (FEM) (e.g. using the Neo-Hookean constitutive law). The mesh of the model is generated according to the reconstructed preoperative three-dimensional liver model, and its material parameters (e.g. Young's modulus, Poisson's ratio) can be set individually based on the patient's preoperative magnetic resonance elastography (MRE) data, or set according to the average biomechanical properties of the liver.

[0133] The second deep learning network is trained (offline) by the above optimized total loss function L_total.

[0134] The weight values a, b, g are preset factors, and d is adaptively adjusted according to the confidence of the initial correspondence.

[0135] Initial prediction (intraoperative): input the current preoperative, intraoperative image and the correspondence into the trained second deep learning network, and quickly forward propagate to obtain an initial non-rigid deformation field.

[0136] Step 5: Real-time update of the non-rigid deformation field.

[0137] Since surgical operations (such as pulling, cutting) or physiological factors (such as residual motion caused by breathing) will cause the liver to deform further, it is necessary to update the deformation field in real time.

[0138] Triggering condition: based on the newly acquired 3D iUS frame and / or detected surgical operation information (such as robot tool force feedback, position change).

[0139] The affected area can be determined first. For example, calculate the area within a preset radius (e.g. 1-2 cm) around the tool tip (first area); calculate the absolute difference in voxel intensity between the current iUS frame and the last frame, and identify the area where the difference is greater than a threshold (e.g. gray value difference > 10); use a simplified physical model or empirical rule to estimate the area where significant stress / strain is generated based on the contact force between the robot tool and the liver (if available) or the known cutting position (third area).

[0140] The final affected area is defined as the union of the three areas.

[0141] Fix the value of the current global deformation field at the boundary of the affected area as a boundary condition. Start local optimization with the current global non-rigid deformation field as the initial state.

[0142] Within the affected area, perform a limited number (e.g. 5-15) of local optimization iterations based on gradient descent (e.g. using the Adam optimizer) to quickly update the deformation field parameters within the area (e.g. if φ is represented by a deformable mesh, update the mesh node positions; if φ is directly output by the network, optimize a small correction field added to the network output).

[0143] The local optimization iteration aims to optimize the deformation field parameters to improve image matching, smoothness, and physical consistency within the affected area, etc. Its objective function is similar in form to L_total during training, but only calculates in the local area.

[0144] The system monitors the accumulated error (e.g., periodically re-evaluates the global L_total, or tracks the drift of the key points), and performs a global non-rigid deformation field optimization periodically (e.g., every 30 seconds or 50 incremental updates), or when it detects that the accumulated error exceeds a pre-set threshold. This can be done by re-running a complete second deep learning network inference (sub-step d2) to refresh the entire deformation field, or performing a few steps of global optimization iterations over the entire organ region.

[0145] By the above steps, a continuously updated non-rigid deformation field is obtained.

[0146] Step 6: Apply the deformation field to generate a deformed model.

[0147] The current updated non-rigid deformation field is applied to the preoperative 3D model. Specifically, for each vertex V pre on the model, its corresponding intraoperative position V post is computed via the deformation field (interpolated), thus generating a deformed 3D model. This deformed model can match the current target organ state (the real shape of the liver) reflected by the real-time acquired intraoperative iUS images in terms of shape, position, and internal structure (relative position of tumor, blood vessels). Figure 4 A schematic diagram of a deformed 3D model according to an example of the present application is shown.

[0148] Step 7: Tool mapping and navigation instruction generation

[0149] The position of the real-time tracked surgical robot end tool is accurately mapped into the coordinate system of the generated deformed 3D model via coordinate system transformation, obtaining a virtual representation of the tool on the model. This virtual tool (e.g., the green structure shown in Figure 5 ) can reflect the real-time position and pose of the real surgical instrument.

[0150] On the display interface of the navigation system, the deformed 3D model, key anatomical structures, and the mapped virtual tool are displayed in fusion. The specific visualization style (e.g., color scheme, transparency, rendering details) of this navigation interface can vary depending on the navigation software platform or display settings used. In the display interface shown in Figure 5 , the deformed liver model is displayed as a whole in a semi-transparent dark red color to observe the internal structure. The key blood vessel system inside is clearly presented, for example, the main blood vessel branches are displayed in blue, and another group of blood vessels are displayed in purple. The target lesion area (tumor) is represented as a prominent, solid red area, clearly showing its position and approximate shape inside the liver.

[0151] Based on the real-time position of the virtual tool and the information of the target lesion region and the risk-avoiding structure (blood vessels, etc.) contained in the deformed three-dimensional model, the navigation instructions are generated.

[0152] Figure 5 The lower right corner also presents visual guidance and quantitative information, which can display a list of tool tip distances to predefined key structures (Pointer distance) in real time.

[0153] When the virtual tool approaches important blood vessels or other structures that need to be avoided (for example, enters the preset danger range), an explicit warning (for example, turns red) can be given in the distance list.

[0154] Based on these navigation instructions that integrate three-dimensional anatomical information, real-time tool position, and quantitative data, the surgeon can guide the surgical robot to perform precise resection operations, such as accurately stripping tissues or cutting at a location that meets the safety requirements for both the tumor and the important blood vessels.

[0155] Through the above steps, the method described in the embodiment can provide high-precision image-guided navigation for robot-assisted liver tumor resection, effectively compensate for intraoperative soft tissue deformation, integrate multi-modal image information, and ensure the real-time and robustness of navigation, thereby assisting the surgeon to achieve safer and more precise surgical operations.

[0156] Figure 6 The electronic device provided for at least one embodiment of the present application includes a memory and a processor. The memory is used to store computer instructions executable on the processor. The processor is used to implement the multi-modal image registration-based surgical robot resection navigation method described in any embodiment or implementation manner of the present application when executing the computer instructions.

[0157] At least one embodiment of the present application also provides a computer readable storage medium having a computer program stored thereon. The program is executed by a processor to implement the multi-modal image registration-based surgical robot resection navigation method described in any embodiment or implementation manner of the present application.

[0158] Those skilled in the art should understand that one or more embodiments of the present specification can be provided as a method, a system or a computer program product. Therefore, one or more embodiments of the present specification can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects.

[0159] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the data processing device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0160] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0161] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0162] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0163] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0164] The above description is only the preferred embodiment of one or more embodiments of the specification, and is not used to limit one or more embodiments of the specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of the specification should be included in the protection range of one or more embodiments of the specification.

Claims

1. A surgical robotic resection navigation method based on multi-modal image registration, characterized by, The method comprises: acquiring preoperative medical images of a patient containing at least first and second modality images, and reconstructing a preoperative three-dimensional model containing a target lesion region and a risk structure to be avoided based on the preoperative medical images; acquiring at least one intraoperative image of a target organ region of the patient in a different modality from the preoperative medical images; extracting cross-modality invariant features from the preoperative medical images and the intraoperative images using a first deep learning network, establishing at least one initial correspondence relationship, and using a second deep learning network to predict an initial non-rigid deformation field based on the preoperative medical images, the intraoperative images, and the extracted initial correspondence relationship, the non-rigid deformation field being used to represent the geometric deformation of the target organ from a preoperative state to an intraoperative state; during the operation, tracking the position of the end tool of the surgical robot in real time, and based on the real-time acquired intraoperative images and operation information, using an incremental calculation strategy to update the non-rigid deformation field in real time; applying the current updated non-rigid deformation field to the preoperative three-dimensional model to generate a deformed three-dimensional model, the state of the target organ in the deformed three-dimensional model matching the state of the target organ in the real-time intraoperative images; mapping the position of the real-time tracked end tool of the surgical robot to the coordinate system of the deformed three-dimensional model; based on the mapped position of the tool and the target lesion region and the risk structure to be avoided of the deformed three-dimensional model, generating navigation instructions to guide the surgical robot to perform a resection operation; wherein the first deep learning network is a twin network structure with a shared weight encoder for mapping image blocks in different modalities to the same feature space, and the first deep learning network is trained by the following contrast loss function: L_contrastive = (1 - Y) * (1 / 2) * (D_W) 2 + Y * (1 / 2) * (max(0, m - D_W)) 2 , wherein Y is a label indicating whether the input preoperative image block and the intraoperative image block belong to the same anatomical position, Y=0 indicates that the image block pair is similar, Y=1 indicates that the image block pair is not similar, D_W represents the Euclidean distance between the two feature vectors output by the twin network encoder, and m is a preset boundary parameter, which is a positive number.

2. The method of claim 1, wherein, The second deep learning network has an encoder-decoder structure and combines an attention mechanism, which is used to adaptively focus on the regions corresponding to the encoder features or cross-modality features that contribute more to the current position deformation prediction in the decoding stage.

3. The method of claim 2, wherein, The second deep learning network is trained by the following loss function: L_total=α*L_gradient_similarity+β*L_smoothness+γ*L_physics+δ*L_correspondence, Wherein, L_gradient_similarity is an image gradient similarity measure term, used to measure the similarity between the intraoperative image obtained by applying the predicted non-rigid deformation field to the preoperative medical image and the actual intraoperative image; L_smoothness is a deformation field smoothness regularization term, used to penalize the non-smooth changes in the predicted non-rigid deformation field; L_physics is a physics consistency regularization term determined based on a predefined physical deformation model; L_correspondence is an initial correspondence supervision term, used to supervise the learning of the second deep learning network using the initial correspondence; and α, β, γ and δ are weight factors, α, β, γ are preset factors, and δ is adaptively adjusted according to the confidence of the initial correspondence.

4. The method of claim 1, wherein, The image gradient similarity measure term L_gradient_similarity is determined according to the following formula: where ∫ denotes integration over the entire target organ region Ω, X denotes a point in the pre-operative medical image, denotes the corresponding point in the intra-operative coordinate system to which the pre-operative point X is mapped via the non-rigid deformation field, denotes the gradient of the deformed pre-operative medical image at point in the intra-operative coordinate system, denotes the gradient of the intra-operative image at point in the intra-operative coordinate system, ||*||2denotes the square of the L2-norm of *.

5. The method of claim 3, wherein, The predefined physical deformation model is a hyperelastic model based on the finite element method, and the parameters of the physical deformation model are determined based on the preoperative medical image. The physics consistency regularization term L_physics is determined according to the following formula: where ∫ denotes integration over the entire target organ region Ω, is the deformation gradient tensor, I denotes the identity tensor, denotes the predicted non-rigid deformation field, denotes the gradient of the non-rigid deformation field, is the strain energy functional.

6. The method of claim 3, wherein, The incremental calculation strategy is used to redefine the non-rigid deformation field within the affected area and its direct adjacent area, and the affected area is determined according to the following method: According to the current position of the end tool of the real-time tracking surgical robot, a first area within a preset spatial range around the tool tip is determined; The local image intensity difference between consecutive intraoperative image data frames is calculated, and an area where the local image intensity difference exceeds a preset difference threshold is identified as a second area; Using the physical deformation model, a third area where significant stress is generated is estimated according to the interaction force between the end tool of the surgical robot and the target organ tissue; The union of the first area, the second area and the third area is determined as the affected area.

7. The method of claim 6, wherein, The incremental calculation strategy for redefining the non-rigid deformation field within the affected area and its direct adjacent area includes: Fixing the deformation values at the boundaries of the affected area as boundary conditions; Within the affected area, the non-rigid deformation field parameters are quickly updated by performing a limited number of local optimization iterations based on gradient descent to minimize a loss function including image similarity, deformation field smoothness, physics consistency and correspondence compliance, and taking the current global non-rigid deformation field as the initial state; The iterative update further includes an error accumulation control mechanism, including periodically or when it is detected that the accumulated error exceeds a preset error threshold, performing global non-rigid deformation field optimization to re-minimize the loss function for training the second deep learning network over the entire target organ area.

8. The method of claim 1, wherein, The first modality is a magnetic resonance imaging (MRI) T2 weighted image, the second modality is a CT angiography (CTA) image, and the intraoperative image is a three-dimensional intraoperative ultrasound (iUS) image.

Citation Information

Patent Citations

  • Gradient distribution-based non-rigid medical image registration method

    CN108053431A

  • Real-time non-rigid registration method and system for surgical navigation image based on deep learning

    CN116485850A