Method and system for model and organ registration to assist navigation

By combining Transformer stereo matching, NeRF deep supervision and finite element mechanical modeling, the problems of insufficient stereo reconstruction accuracy and organ deformation in medical surgical navigation are solved, and high-precision, real-time alignment of virtual organs and real organs is achieved, improving the accuracy and real-time performance of surgical navigation.

CN120411186BActive Publication Date: 2025-09-09THE FIRST AFFILIATED HOSPITAL OF MEDICAL COLLEGE OF XIAN JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510896530.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-09-09
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

Existing technologies in medical surgical navigation have problems such as insufficient stereoscopic reconstruction accuracy, organ deformation affecting navigation accuracy, and limited mixed reality registration accuracy. In particular, it is difficult to achieve real-time, high-precision alignment of virtual organs with real organs in minimally invasive surgical environments.

Method used

A method combining Transformer stereo matching of binocular laparoscopic images, NeRF deep supervision, finite element mechanics modeling and graph neural network is used to achieve accurate alignment between virtual organ models and real organs through self-supervised optimization and real-time deformation prediction.

Benefits of technology

The accuracy and real-time performance of stereo matching are significantly improved, ensuring high-precision alignment of virtual organ models with real organs, and meeting the real-time and precision requirements of surgical navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411186B_ABST
    Figure CN120411186B_ABST
Patent Text Reader

Abstract

The present invention relates to electronic digital data processing and discloses a method and system for aligning models with organs to assist navigation, thereby improving accuracy and real-time performance. The method includes: using a Transformer architecture to perform feature matching along the epipolar direction of the left and right color images, retaining only disparity candidates whose inter-pixel matching probability meets a preset confidence threshold, constructing a sparse probability disparity cost volume, performing regularized smoothing on each disparity cost volume and converting it into a normalized probability distribution, then calculating the expectation of the probability distribution as a disparity estimate, and then converting the disparity estimate of each pixel into depth information to obtain an initial depth map; performing self-supervised optimization on binocular depth estimation based on NeRF; and performing three-dimensional matching on the shape and position of a virtual organ model based on the two-dimensional information observed in the left or right color image combined with the depth information of each pixel, so that the virtual organ model is aligned with the current morphology of the real organ.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for aligning a model with an organ to assist navigation. Background Art

[0002] Currently, mixed reality (MR), deep learning, organ biomechanical modeling, and stereoscopic vision technologies are increasingly being used in medical surgical navigation. MR technology can fuse digital models with real-world images, providing surgeons with intuitive 3D visualization of anatomical structures, helping to improve surgical precision and reduce cognitive load. For example, in anterior cruciate ligament reconstruction surgery, the use of MR visualization of a pre-operative 3D model reconstructed by computed tomography (CT) significantly improved femoral tunnel positioning accuracy (the MR group achieved a positioning error of approximately 0.22±0.16mm, significantly lower than the 2.57±0.30mm achieved with traditional methods). MR also supports human-computer interfaces such as voice control and gesture interaction, enabling contactless access and manipulation of information during surgery. Deep learning has recently demonstrated remarkable performance in medical image processing, enabling automated feature extraction and organ recognition from intraoperative images, such as laparoscopic images. However, deep learning models can only extract and analyze features from existing intraoperative images and are unable to effectively incorporate information such as preoperative patient testing. Organ biomechanical modeling uses physical models to simulate the deformation behavior of three-dimensional organs after being stretched or compressed, but it cannot adapt to intraoperative images such as two-dimensional laparoscopy in real time. Binocular laparoscopic stereoscopic imaging can obtain real-time depth information of the surgical field and obtain a three-dimensional reconstruction of the organ surface, but it only provides doctors with a three-dimensional visual experience and cannot provide additional information. However, the integration of these technologies for real-time intraoperative navigation still has many limitations and deficiencies.

[0003] Limitations of stereo laparoscopy: Traditional binocular stereo matching algorithms rely on hand-crafted feature matching and are susceptible to interference from complex factors in the surgical environment. Minimally invasive surgery (MIS) environments are subject to smoke, liquid (blood), instrument occlusion, and varying lighting conditions, all of which pose challenges for stereo matching. Uneven lighting and reflections can cause image contrast variations, making it difficult for traditional algorithms to find robust correspondences in areas with weak or highly repetitive textures. Consequently, stereo reconstruction accuracy is suboptimal in these scenarios, with depth maps often exhibiting noise and holes. In particular, many areas of organ surfaces during laparoscopic surgery lack significant features, making them difficult for traditional methods to match, resulting in insufficient depth reconstruction accuracy. Furthermore, laparoscopic camera and organ motion introduce dynamic variations in the disparity range, making the fixed disparity search ranges predefined by conventional stereo algorithms incapable of handling rapid or large deformations. Furthermore, to meet the real-time demands of intraoperative applications, algorithms must complete computations within a limited timeframe, often forcing a compromise between accuracy and speed. Algorithms implemented by traditional CPUs often cannot meet real-time requirements. Although GPU parallel acceleration speeds up processing, more optimized algorithms are still needed to further reduce latency.

[0004] Organ deformation and registration challenges: Organs (such as the liver and kidneys) undergo continuous deformation during surgery due to breathing, traction, and other factors. This deformation impairs the accuracy of navigation guided by preoperative images, as preoperative images and intraoperative organ morphology are inconsistent. Traditional surgical navigation relies heavily on preoperative images for intraoperative reference, but organ displacement and deformation limit this approach to rigid or nearly rigid anatomical structures (such as bones). To address this issue, biomechanical models must be integrated with intraoperative data to update the preoperative model in real time to match the intraoperative situation. Biomechanical models, such as the finite element method (FEM), are considered promising approaches for simulating organ deformation based on mechanical properties and for performing elastic deformation correction on preoperative images. However, sophisticated FEM modeling is computationally complex, making direct intraoperative application prohibitively expensive. Furthermore, sufficient intraoperative data is required to constrain and drive the model (e.g., intraoperative point clouds of organ surfaces or positional changes of markers). Current laparoscopic stereoscopic reconstruction provides a means of acquiring surface point clouds, but using sparse / semi-dense point clouds to drive high-dimensional non-rigid models remains a challenge. Relying solely on manual registration is not only time-consuming but can also lead to misalignment between the real and the virtual. Previous studies have shown that manually adjusting the position of holographic models to align with real organs in MR navigation carries the risk of model misalignment and poor accuracy. Therefore, more intelligent algorithms are needed to autonomously achieve the difficult-to-accurate manual registration between models and organs.

[0005] Current shortcomings of mixed reality surgical navigation: Although MR has shown great potential in surgical navigation, the current application also has limited registration accuracy: many MR navigation systems rely on optical or manual initialization registration, and there may be accumulated errors in the position correspondence between the holographic image and the real anatomy. Especially when the organ changes dynamically, if the initial registration is not accurate, significant deviations may occur during the operation. For example, experiments have found that the stability of using some MR glasses for spatial registration is poor, and the average positioning error between different users can reach about 6mm, which is far from meeting the requirements of delicate surgery. Therefore, improving the accuracy of MR registration is an urgent problem to be solved. Summary of the Invention

[0006] The present invention aims to disclose a method and system for aligning a model with an organ to assist navigation, so as to improve accuracy and real-time performance.

[0007] To achieve the above-mentioned purpose, the method disclosed in the present invention for aligning a model with an organ to assist navigation includes:

[0008] Step S1: convert the parallax of the color images on the left and right sides of the binocular laparoscope into depth information to obtain an initial depth map;

[0009] The left and right color images are first processed through a multi-scale CNN (Convolutional Neural Network) to extract feature pyramids. Image features at different scales are then fused to obtain multi-channel feature representations for the left and right eyes. A Transformer architecture based on epipolar constraints is then used to implement self-attention and cross-attention mechanisms. The self-attention mechanism calculates pixel-level association weights along the epipolar direction in the feature sequence of the same view. The cross-attention mechanism performs feature matching along the epipolar direction on the left and right color images, retaining only disparity candidates whose inter-pixel matching probability meets a preset confidence threshold. A sparse probabilistic disparity cost volume is then constructed. Each disparity cost volume is then regularized and smoothed, converted into a normalized probability distribution, and the expectation of this probability distribution is calculated as the disparity estimate.

[0010] Step S2: Perform self-supervised optimization of binocular depth estimation based on the trained NeRF (Neural Radiance Fields) model. The training process of the NeRF model specifically includes:

[0011] Step S21: The initial depth map is input as a depth prior together with the corresponding left or right single-side color image into the NeRF model. The NeRF model synthesizes a new image from a preset perspective using volume rendering technology, and compares the synthesized new image with the input left or right single-side color image to calculate the reprojection error.

[0012] Step S22: Process pixel depth and color rendering in the same rendering mode, replace the color of each sampling point with the corresponding depth position, divide it by the total weight to obtain the depth rendering, and then map each depth rendering calculated value to the corresponding pixel to obtain the depth rendering value of each pixel; then calculate the depth consistency loss between the depth rendering value inferred by the rendering and the disparity estimation value;

[0013] Step S23: Adopting a strategy of alternating training of the NeRF model and the Transformer architecture to gradually optimize the depth estimation results, iteratively running the fixed NeRF model to update the Transformer architecture to reduce the pixel color reprojection error and the depth consistency error, and then fixing the Transformer architecture to update the NeRF model to improve the fitting accuracy of the laparoscopic view geometric consistency, ultimately minimizing the weighted sum of the reprojection error and the depth consistency loss;

[0014] Step S3: Perform three-dimensional matching based on the two-dimensional information observed in the left or right color image combined with the depth information of each pixel and the shape and position of the virtual organ model to align the virtual organ model with the current morphology of the real organ.

[0015] Preferably, in step S1, in the process of calculating a sparse probabilistic disparity cost volume that meets a confidence threshold based on the matching probability between pixels, a sparse attention strategy is introduced to limit the disparity search to a matching probability range greater than the confidence threshold.

[0016] Preferably, in step S3, it specifically includes:

[0017] Step S31: combining the physical simulation capabilities of finite element analysis with the machine learning capabilities of graph neural networks to achieve accurate registration of non-rigid organ deformations;

[0018] Step S32: using a visual SLAM (Simultaneous Localization and Mapping) algorithm to estimate the position of the laparoscopic camera in real time;

[0019] Step S33: After obtaining the organ deformation and camera pose, the virtual organ model is superimposed on the surgical field image through mixed reality technology.

[0020] Preferably, the step S31 specifically includes:

[0021] Step S311: The finite element model calculates the deformation distribution of each part of the organ under the action of external forces based on the biomechanical properties and boundaries of the organ, providing physically realistic deformation results;

[0022] Step S312: Discretely represent the organ as a graph structure, and input the mesh node displacements and topological connections obtained by finite element calculation into a GNN (Graph Neural Network) model;

[0023] Step S313, the GNN model uses organ grid nodes as graph nodes and grid topology as graph connections, and predicts the real-time deformation of the organ by learning the correlation and deformation pattern between nodes. The correlation includes the mechanical correlation between nodes learned by the message passing mechanism, and the mechanical correlation includes the mapping relationship from the boundary after displacement formed by the displacement or force on the organ surface to the displacement field of the internal node.

[0024] Preferably, the step S32 specifically includes:

[0025] Step S321: The SLAM algorithm extracts and matches feature points in the left or right color image frame by frame to minimize the reprojection error and obtains the direction and position trajectory of the camera in the surgical environment. At the same time, a dense point cloud of the surgical area is constructed so that the camera positioning is constrained by the anatomical structure.

[0026] Step S322: converting the virtual organ model to the camera coordinate system according to the camera positioning information for alignment rendering;

[0027] Step S323: By minimizing the alignment error between the virtual image displayed by the MR device and the real image, the posture and position of the virtual organ model are adjusted in real time to ensure that the virtual organ model is spatially aligned with the actual endoscopic left or right side color image.

[0028] Preferably, the method of the present invention further comprises: using Kalman filtering to fuse the camera pose solved by visual SLAM with the high-frequency acceleration, angular velocity and EM (Electromagnetic Tracking System) positioning results measured by an IMU (Inertial Measurement Unit), where the IMU and EM are different sensors deployed on the laparoscope lens.

[0029] Preferably, the step S33 specifically includes:

[0030] Step S331: Obtain the internal and external parameters of the laparoscopic camera calibration, wherein the internal parameters include focal length, principal point, and distortion coefficient, and the external parameters include the position in the surgical coordinate system; set the parameters of the virtual camera according to the calibration results of the internal and external parameters so that the perspective of the rendered virtual model is consistent with that of the real laparoscopic camera;

[0031] Step S332: transforming the virtual organ model to the surgical coordinate system, and adjusting the size and position of the virtual organ model according to the actual anatomical proportions to ensure that the virtual organ model is spatially aligned with the real organ;

[0032] Step S333: During the rendering process, lens illumination matching is performed, distortion correction of the virtual organ image is performed according to the lens characteristics of the laparoscope, and the lighting effect of the surgical light source on the virtual organ model is simulated; and the dense point cloud or key anatomical features reconstructed by SLAM are used to dynamically adjust the position and scale of the virtual organ model during the operation.

[0033] Preferably, the process of dynamically adjusting the position and scale of the virtual organ model during surgery specifically includes:

[0034] Step S3331: extract a series of three-dimensional key contour vertices evenly distributed along the organ edge or the fitted contour curve from the anatomical surface of the reconstructed virtual model. , The vertex index of the contour is used to obtain the posture of the virtual organ model in the mixed reality coordinate system. ,in, is the rotation matrix of the model, is the three-dimensional translation matrix of the model, each three-dimensional contour point , calculate the projection of each contour point on the image plane ,in, is the homogeneous depth of the projected point, is the two-dimensional coordinate of the contour vertex in the image plane; is the transpose symbol;

[0035] Step S3332: Using the edge contour of the organ in the actual image as a reference, obtain the real contour pixel point set For each projection point , find each projection point in the true contour set The nearest neighbor point on the , calculates the alignment error of the nearest neighbor point, takes minimizing the sum of squared deviations of all contour points as the optimization goal, and finds the optimal rigid body fine-tuning transformation.

[0036] To achieve the above objectives, the present invention also discloses a system for aligning models with organs to assist navigation, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the computer program.

[0037] The present invention has the following beneficial effects:

[0038] 1. Introducing the Transformer into stereo matching realizes a new global context-aware epipolar matching mechanism. Traditional stereo matching is usually limited by local correlation, but the Transformer architecture based on this invention can establish pixel correspondences on a global scale, significantly improving matching accuracy. At the same time, after feature matching along the epipolar direction of the left and right color images, only disparity candidates with inter-pixel matching probabilities that meet a preset confidence threshold are retained to construct a sparse probabilistic disparity cost volume. Each disparity cost volume is then regularized and smoothed and converted into a normalized probability distribution. The expectation of this probability distribution is then calculated as the disparity estimate, significantly improving the real-time performance of the disparity calculation.

[0039] 2. Introducing NeRF as a depth supervision signal, depth optimization is transformed into a differentiable new perspective synthesis problem, optimizing depth results from the perspective of global photometric consistency. Under the global supervision of NeRF, the shortcomings of traditional stereo matching in low-texture or high-gloss areas are effectively compensated, making depth estimation in tricky areas equally accurate. At the same time, the joint optimization of the depth consistency constraint and the NeRF-Transformer stereo network ensures that the depth map meets the consistency of the reprojected photos while maintaining geometric accuracy; the innovative strategy of integrating neural rendering with stereo vision greatly improves the details and reliability of depth estimation, achieving higher accuracy.

[0040] 3. Based on the two-dimensional information observed in the left or right color image, the depth information of each pixel is combined with the shape and position of the virtual organ model to perform three-dimensional matching, ensuring the accuracy of the alignment of the virtual organ model with the current morphology of the real organ.

[0041] The present invention will be described in further detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0043] Figure 1 It is a flowchart of the method for aligning a model with an organ to assist navigation disclosed in Example 1 of the present invention. DETAILED DESCRIPTION

[0044] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. However, the present invention can be implemented in many different ways as defined and covered by the claims.

[0045] Example 1

[0046] This embodiment discloses a method for real-time, flexible intraoperative organ registration based on the fusion of binocular laparoscopic reconstruction and mixed reality. The system integrates binocular stereo vision, neural radiance field (NeRF), finite element modeling, SLAM tracking, and mixed reality (MR) rendering, ensuring registration accuracy while meeting real-time surgical requirements. The overall process is as follows: First, 3D reconstruction is performed using color images from the left and right sides of the binocular laparoscope. Second, NeRF is introduced for self-supervised optimization of depth estimation, improving reconstruction accuracy and reducing reliance on large amounts of calibration data. Organ deformation prediction is then constructed by combining a finite element model (FEM) with a graph neural network (GNN) to achieve physical deformation of the organ model during surgery. Next, an improved visual SLAM algorithm is used to obtain the laparoscopic camera pose. Combined with a pre-registered preoperative model, the virtual model is accurately aligned with the endoscopic image. Finally, optimized rendering overlays the corrected virtual organ model on the intraoperative image in real time, providing realistic augmented reality navigation support. The following describes the technical details of each functional module, along with the corresponding mathematical models and performance data.

[0047] 1. Optimization of the binocular Transformer stereo reconstruction module.

[0048] This module uses an improved Transformer network to complete high-precision stereo matching and depth estimation of color images of the left and right sides of the laparoscope.

[0049] A1. Embedding layer improvements: The left and right laparoscopic images (or the color images of the left and right sides, where "side" can also be called "eye" and will not be further explained) are first processed through a convolutional neural network with shared weights (the left and right laparoscopic images of the binocular laparoscope are input into two parallel network branches that use exactly the same set of convolution kernel parameters) to extract a multi-scale feature pyramid.

[0050] Preferably, this embodiment can also encode the position of each pixel and the disparity range estimated based on the intrinsic parameters of the two cameras into the feature to fuse the camera intrinsic parameters and epipolar constraint information. It can be expressed in function form as follows: for each pixel coordinate in the left image , and its corresponding normalized coordinates With the assumed disparity prior Mapped together as position vector , embedded into the feature space through a linear transformation. This embedding preserves pixel geometry and disparity priors, helping the Transformer fully exploit spatial relationships during subsequent processing, improving matching accuracy and robustness. Compared to traditional CNN features that lack explicit geometric encoding, the Transformer, combined with sequential position embedding, captures global image dependencies. This makes it more robust when finding corresponding points in areas with weak or repetitive textures.

[0051] The estimated disparity range is the preset possible depth range for each pixel. By default, the depth of each pixel is selected within the estimated disparity range. This can be obtained by first obtaining the laparoscope camera parameters from the laparoscope manufacturer, primarily the baseline length b and focal length f. The depth range is set based on the typical working distance between the camera and human tissue (default: [Zmin, Zmax] = [2 cm, 10 cm]). The depth limits are then mapped to the disparity limits using the formula d = bf / Z, where d is the formula for calculating the disparity limits. This results in the estimated disparity range being [dmin, dmax] = [bf / Zmax, bf / Zmin]. Finally, [dmin, dmax] is discretized into {dmin, dmin+Δd, dmin+2×Δd, …, dmax}; Δd can be adjusted according to the computing power and the expected stereo matching accuracy. The default value is (dmax-dmin) / 50, which means it can be discretized into 50+1=51 estimated disparity range intervals. It is assumed that the disparity d is a specific value within the discretized estimated disparity range.

[0052] A2. Attention mechanism optimization: In the Transformer matching module, this embodiment uses alternating self-attention and cross-attention layers to process the left and right feature sequences. Multi-head self-attention is calculated by Long-range associations are established along the epipolar lines within each visual view (including the left and right color images) to capture image context. Cross-attention then uses the query from one visual view to calculate matching correlations with the keys and values ​​of the other visual view, achieving feature alignment across the views. This alternating attention mechanism overcomes the limitations of fixed-length cost windows and allows for matching searches along the entire epipolar line (i.e., allowing each pixel to "look out" to the entire horizon and associate with more macroscopic textures and structures, thereby achieving both local accuracy and global consistency during matching). This context-awareness helps the model find more reliable matching points in areas with weak or repeated textures by combining more distinct features from a distance, thus making it more adaptable to large parallax and occluded areas. Furthermore, to reduce the computational overhead of global attention, this embodiment introduces a sparse attention strategy: attention weights are only calculated at possible matching locations, significantly reducing computational complexity. For example, by limiting the disparity search to a set of candidates with high match confidence (i.e., a range with a matching probability greater than a preset confidence threshold) rather than the entire range, the matching complexity can be reduced from the original [number of matches to be matched]. Reduced to about (in, is the number of pixels, is the parallax search range, is the number of candidates per pixel); this ensures that the algorithm can run in real time while maintaining accuracy.

[0053] A3. Cost volume construction and regularization: Based on the matching results output by Transformer, this embodiment explicitly constructs a dense cost volume to fuse matching evidence. Unlike the traditional method of traversing all disparities pixel by pixel to calculate the matching cost, this embodiment uses the matching probability output by the attention layer to directly form the cost volume, retaining only the disparity positions with high confidence, thereby reducing redundant calculations. Represents pixels Parallax The matching probability can be calculated only when Larger Calculate the cost volume value at , while other disparities are ignored. The initial cost volume is locally regularized and smoothed via 3D convolution to filter out noise and gross errors. During processing, the residual module based on the Transformer architecture and the local smoothing properties of convolution inherent in its feedforward network, along with the Transformer's self-attention aggregation mechanism, model long-range depth-dependent correlations in the cost volume, thereby improving the accuracy and stability of disparity estimation and further optimizing the global consistency of the cost distribution.

[0054] A4. Disparity regression and depth output: After obtaining the regularized cost volume, a sub-pixel disparity map can be obtained through Softmax regression. Specifically, for each pixel Calculate the normalized matching distribution ; Then the disparity estimate is taken as the expected , which can be interpolated to sub-pixel. Then, according to the intrinsic parameters of the camera calibration, the disparity is converted into dense depth , for example, for parallel binocular cameras, there is a relationship . The output depth map is aligned with the real geometry at the sub-pixel level, providing accurate three-dimensional information for subsequent steps. Experimental calibration shows that compared with the baseline network without Transformer, this module reduces the average depth error by about 20%, and exhibits excellent generalization performance on both open datasets and clinical data. This benefit comes from the fact that the improved Transformer in this embodiment has stronger global matching capabilities: compared with traditional algorithms that assume a fixed disparity range, Transformer stereo matching is not limited by a preset range, can adaptively match larger depth differences and identify occluded areas, thereby significantly improving the reconstruction completeness and accuracy.

[0055] 2. Deep supervision and optimization strategy based on NeRF.

[0056] In order to further improve the robustness of reconstruction and reduce the dependence on real calibration data, this embodiment integrates the Neural Radiance Field (NeRF) technology to perform self-supervised optimization of binocular depth estimation. NeRF can learn implicit 3D scene representations through a small number of RGB images from different perspectives, and render realistic images from any new perspective. This embodiment takes advantage of this feature and combines the depth predicted by the binocular network with the NeRF rendering process during the training phase to introduce cross-perspective reprojection consistency constraints. The specific implementation is as follows: First, the initial depth map output by the binocular Transformer stereo reconstruction module is input into the NeRF model together with the left or right single-side color image (if the left view is used as the reference and the right view is used as the matching map, the depth map of the left view will eventually be generated. If the right view is used as the reference and the left view is used as the matching map, the algorithm translation vector parameters need to be adjusted accordingly). NeRF represents the scene as a radiance field function , where p is the spatial coordinate, is the sight direction, output color and body density For each ray shot by the camera , sample the view at a point or local area in space according to the predicted depth, and pass the volume rendering integral Transformed into discrete form, we get the volume rendering equation of NeRF To calculate the composite color of the light at another perspective:

[0057] ;

[0058] in, represents the sampling point index along the ray, is the distance between adjacent sampling points, Indicates the Transmittance in front of the point. Formula, just change the color of each sampling point Replace with its depth position , and then divided by the total weight, we get the depth rendering , it is only necessary to render the depth of the light corresponding to the pixel (i, j), that is, , and then normalize it to get the depth rendering of the pixel .

[0059] ;

[0060] Then, the rendered color is compared with the actual color of the corresponding pixel from another perspective to construct the reprojection loss: .in, Represents a set of sampled rays (e.g., all pixels or randomly sampled pixels). At the same time, this embodiment introduces depth consistency loss to constrain the depth predicted by the stereo network. Depth derived from volume rendering with NeRF Keep it consistent, i.e. add The comprehensive optimization goal is to minimize ( is the trade-off coefficient), and backpropagation is used to simultaneously update the stereo matching network and NeRF parameters, enabling them to converge collaboratively. In this way, NeRF provides a self-supervisory signal source for the stereo network, leveraging information inherent in the image to optimize depth estimation. Compared to supervised training that relies solely on limited real-world depth data, this strategy can fully utilize unlabeled laparoscopic videos for training, alleviating data scarcity and approaching physically realistic geometric solutions.

[0061] Furthermore, this embodiment employs an iterative fusion optimization strategy: alternately optimizing the binocular Transformer stereo reconstruction network and the NeRF model. Initially, the binocular Transformer network refines the predicted depth interval by increasing the number of disparity candidates D', obtaining an accurate depth result as the training ground truth. NeRF is then trained to fit this depth. Next, the binocular Transformer network is supervised using new-view images synthesized by NeRF to continue training for fine matching when the number of disparity candidates D is small. After updating the depth, NeRF is trained again, repeating this cycle to gradually improve depth accuracy. Through this iterative fusion optimization strategy, the binocular Transformer stereo reconstruction network + NeRF joint optimization process in this embodiment ensures that the final binocular laparoscope image depth map is primarily contributed by stereo matching in textured areas, resulting in rich high-frequency details. In contrast, NeRF compensates for this in textureless or reflective areas, resulting in more accurate edge contours. Furthermore, the binocular Transformer network accurately outputs precise depth even when the number of disparity candidates D is small, reducing the algorithm computational complexity from O(N·D') to O(N·D), reducing the computing power burden and computation time for real-time flexible registration during surgery. Experimental results show that after introducing NeRF self-supervision, the average depth error of reconstruction is reduced from approximately 3.0mm to approximately 2.0mm, a reduction of approximately 33%. The improvement in depth quality is even more significant under conditions of varying illumination or on surfaces with low texture. For example, the depth error at the edges of small blood vessels is reduced by approximately 30%. Furthermore, the addition of NeRF constraints reduces the network's demand for training data, allowing it to maintain stable performance even with limited samples. This fusion strategy effectively improves reconstruction accuracy and generalization.

[0062] In this embodiment, NeRF's volume rendering essentially represents pixel color as an integral along a camera ray—that is, color is the expected value of the weighted cumulative volume density along the line of sight. Pixel color is equivalent to the weighted average of the colors at each point along the ray, where the weight is proportional to the transmittance-density product at that point. The calculation of pixel depth follows the same structure as color rendering: if the "color" of each sampling point is replaced by the depth at that point, the rendering result is the expected depth for that ray. Therefore, the weight of a sampling point's contribution to the final color also reflects the probability distribution of the interaction between the ray and the object at that location. NeRF's color rendering and depth rendering are mathematically unified: both can be viewed as weighted integrals of attributes along the ray: one attribute is color, the other is depth. This provides a rigorous mathematical basis for calculating depth using the same weights as color, namely, depth is the expected position of the density distribution along the ray.

[0063] 3. Deformation registration module integrating finite element calculation and GNN.

[0064] In response to the non-rigid deformation of organs during surgery, this embodiment establishes a physics-driven nonlinear finite element (FEM) model and combines it with a graph neural network (GNN) to achieve high-speed deformation prediction, so as to align the pre-operative organ model with the intraoperative state in real time.

[0065] C1. Nonlinear finite element modeling of organs: The three-dimensional model of the organ before surgery is used as the initial reference configuration, and its coordinates are defined as , the coordinates after deformation are , then the displacement field is The deformation is expressed using Lagrangian description, and the deformation gradient tensor is defined as: , where I is the second-order unit tensor that retains the original configuration baseline. hour, ,when hour, When considering large deformation, the Green–Lagrange strain tensor is introduced:

[0066] ;

[0067] Its expansion includes the first-order and second-order terms of displacement, which can accurately describe the deformation under large rotation and large strain. Assuming that the organ is an isotropic nonlinear elastic body, a suitable strain energy function can be selected To characterize its elastic mechanical properties. For example, under the St.Venant-Kirchhoff model, the strain energy in the form of a quadratic polynomial is: ;in, Characterizes the resistance of a controlled material to volume changes, Characterizes the shear modulus, both are Lamé constants (both related to Young's modulus and Poisson's ratio Related); tr() is the trace operator, which represents the sum of all main diagonal elements; The volume energy term corresponds to the energy contribution stored by human organs during volume changes (expansion or compression); is the shear energy term, which corresponds to the energy stored in the human body during shape change (shear, torsion, anisotropic stretching). Based on this strain energy, the second-order Piola-Kirchhoff stress can be calculated: . Further through Mapped to first-order Piola-Kirchhoff stress , used to calculate the internal nodal forces of the finite element. Considering that most organs are approximately incompressible, this embodiment adds a volume penalty term to the material model or uses the form of decomposition of strain energy to meet the approximate The volume constraint can be removed to simulate the deformation behavior of organs under compression and other working conditions more realistically. The above nonlinear finite element modeling provides physical priors for registration, making the deformation calculation conform to the laws of mechanics. Solving the deformation of organs can be transformed into an energy function minimization problem: find the displacement field The total energy of the system Reaching the minimum, that is, satisfying the equilibrium equation The total energy includes elastic deformation energy, work done by external forces, and registration constraints, etc.:

[0068] ;

[0069] in, is the body force density (such as gravity), is the surface force, and the last term is the registration error term. Represents the coordinates of the real-time reconstructed surface points (or feature points), The displacement of the corresponding point on the preoperative model The position under action, is the weight parameter, represents the total number of corresponding point pairs used for registration between the model and the organ, Represents the three-dimensional volume integral domain of the organ model. By minimizing this objective function, the displacement field that best matches the simulated deformed model surface with the reconstructed surface can be solved. . Since the optimization is nonlinear, this embodiment can adopt the Newton-Raphson iterative solution and use implicit time integration to maintain the numerical stability of the solution. In order to meet the real-time requirements, this embodiment performs GPU parallel acceleration on the solution process and introduces an adaptive solution strategy: the finite element mesh is adaptively refined according to the stress / strain distribution of the current deformation field, the mesh density is increased in the high strain area, and a coarser mesh is maintained in the area with gentle deformation, so as to reduce the global calculation amount while ensuring local accuracy. This adaptive strategy effectively reduces the degrees of freedom of each iterative solution step, thereby improving computational efficiency.

[0070] C2. Graph Neural Network-Assisted Deformation Prediction: Despite the aforementioned optimizations, real-time nonlinear FEM solutions still have considerable computational complexity. To this end, this embodiment innovatively introduces a GNN as a proxy model for the FEM. Specifically, a large amount of offline finite element simulation data (covering various organ anatomical boundaries and external forces) is utilized, combined with a dense point cloud, to compare the differences between the current organ surface and the initial model surface. Based on this difference, the surface displacement field is extracted and used as a constraint for the GNN prediction. Next, a GNN model is trained to directly learn the mapping from the displaced (organ anatomical) boundaries (e.g., the displacement or force applied to the organ surface) to the displacement fields of internal nodes. This GNN uses organ mesh nodes as graph nodes and the mesh topology as the graph connections, employing a message passing mechanism to learn the mechanical dependencies between nodes. After training, the GNN model can instantly predict the approximate deformation of the organ when given a new (organ anatomical) boundary, significantly accelerating computation. Previous studies have shown that a well-trained GNN model can simulate an 80-second organ compression deformation in only about 47 seconds, while a high-precision FEM simulation would take 15 hours. However, the average error in node position predictions remains within 1mm. Clearly, GNNs can accelerate organ mechanics simulations by hundreds of times while maintaining accuracy.

[0071] In this embodiment, the deformation field predicted by the GNN serves as the initial solution for the FEM iteration, employing a "GNN + FEM two-stage calculation" model: the GNN first rapidly estimates the displacement field, and then the FEM simulation corrects the details. This combination ensures real-time performance while preserving physical accuracy, significantly reducing the number of Newton iterations and accelerating convergence. Furthermore, due to the GNN's self-learning capabilities, the system can gradually fine-tune model parameters based on actual deformation data during each surgery, achieving an "increasing accuracy with use" effect.

[0072] C3. Flexible Organ Registration and Accuracy Evaluation: Through the aforementioned FEM physical modeling and GNN-accelerated prediction, this embodiment can align the preoperative organ model deformation to the endoscopically reconstructed organ surface in real time during surgery. Registration errors primarily arise from inaccurate deformation prediction and noise in the reconstructed surface.

[0073] Preferably, the present embodiment adopts the surface distance error as the evaluation index of the registration accuracy, that is, the average distance error between the virtual model surface and the real reconstructed surface point cloud after calculation. Experimental results show that the method of the present embodiment can control the error at a low level under various working conditions. For example, for the simulated liver compression deformation scene, if only rigid registration is used, the deviation between the model surface and the real surface exceeds 10mm, while the method of the present embodiment reduces the average surface distance error to within 2mm through FEM deformation compensation. Even in the case of a lack of feature matching points in local areas, the introduced elastic regularization term ensures that the deformation field is smooth and reasonable, and no local excessive distortion occurs.

[0074] This embodiment also compares an elastic registration algorithm with only geometric constraints and the physical mechanics fusion method of this embodiment. When there is no external force deformation, the accuracy of the two is comparable; but when there is external force pressing, the error of the pure geometric method increases to more than 4mm, while the method of this embodiment still controls the error within 2.5mm under the same circumstances due to the introduction of FEM physical constraints. This shows that this embodiment is superior to the existing technical methods in both registration accuracy and stability. To meet real-time requirements, the depth estimation network is deployed on the Nvidia RTX 4090 GPU, and CUDA parallel computing is used to control the computational complexity of the disparity estimation step to within 1.5ms; the SLAM module eliminates low-information feature points to speed up matching; the finite element solution is limited to 50 iterations per step and hot-started using the previous frame result. Through the above methods, and thanks to the good initial solution and GPU acceleration provided by GNN, the method of this embodiment can complete a complete deformation solution update within 100ms, achieving a registration refresh rate close to 10Hz; for smaller-scale deformation updates, the calculation frame rate can be increased to above 20Hz, thus basically meeting the real-time requirements of laparoscopic surgery. In summary, intraoperative flexible organ registration based on the combination of deep learning and biomechanical models ensures high-precision alignment of virtual models and real organs in complex surgical environments, providing doctors with a reliable anatomical positioning reference.

[0075] 4. SLAM camera tracking and virtual-reality fusion optimization.

[0076] This module aims to accurately locate the laparoscopic camera's position and stably integrate virtual content into the real surgical field. Optionally, this embodiment can be improved upon the ORB-SLAM2 algorithm (a classic SLAM algorithm) to accommodate the specific needs of laparoscopic surgery.

[0077] D1. Visual SLAM tracking optimization: In response to the rapid movement, out-of-view, occlusion and other situations that may occur in surgical scenes, this embodiment adds multiple constraints to the SLAM front-end and back-end to improve robustness. For example, in the feature tracking stage, a motion model prediction based on local organ rigidity is added to reduce the loss of feature matching during rapid movement; in the back-end optimization, the reconstructed dense surface point cloud is integrated as a geometric constraint, and participates in the camera pose optimization together with the sparse feature points, so as to keep the camera positioning from diverging in texture-poor areas. Camera pose Obtained by minimizing the reprojection error: ,in is a 3D point on the map, is the 2D feature observed in the current frame, is the projection function. After introducing the dense point cloud constraint, the optimization objective also includes , represents the point of real-time reconstruction, is a pre-stored organ surface model, and K is the set of all points in the outermost layer of S, so that the camera positioning is constrained by the anatomical structure. These improvements significantly improve the robustness of SLAM in difficult situations. Experiments show that in in vivo and in vitro tests, the improved SLAM of this embodiment can still maintain continuous tracking and is not prone to losing positioning when encountering rapid lens movement, temporary out-of-field view, and partial occlusion. Compared with the original ORB-SLAM2, the tracking loss rate of this embodiment is reduced by about 50%, and the mean square error of the camera pose estimation is reduced by about 20%. Since only a small number of optimization constraints are added, the algorithm processing frame rate remains at the real-time level (above 30Hz), which can keep up with the frame rate output of laparoscopic video.

[0078] D2. Hybrid Tracking and Multi-Sensor Fusion: To further enhance tracking robustness, the system of this embodiment supports a hybrid tracking mode combining optical, inertial, and electromagnetic positioning. A miniature IMU and electromagnetic positioning (EM) sensor are installed near the laparoscope lens. Kalman filtering is used to fuse the camera pose calculated by visual SLAM with the high-frequency acceleration / angular velocity measured by the IMU and the EM positioning results. This fusion process combines the advantages of vision and external sensing: vision provides high-precision relative motion estimation, the IMU provides high-frequency continuous motion compensation, and the EM provides a reference to counteract magnetic field occlusion drift. This results in a smooth and accurate 6-DoF camera trajectory. Animal experiments have demonstrated that with hybrid tracking enabled, the system maintains stable registration using the IMU / EM even when the laparoscope lens is obscured by blood for up to 2 seconds. Once vision is restored, the virtual overlay position shows no significant jumps. This demonstrates that multi-sensor fusion significantly enhances the system's robustness under extreme conditions. Compared with pure visual tracking, the fusion solution reduces posture error by nearly 80% during short-term visual failure (position drift is significantly reduced during occlusion), providing higher positioning reliability during clinical surgery.

[0079] 5. Mixed reality rendering and visual calibration.

[0080] After acquiring highly accurate organ deformation and camera pose, this embodiment uses mixed reality technology to accurately overlay the virtual model onto the surgical field image. To ensure the accuracy and realism of virtual-reality fusion, this embodiment optimizes two aspects: online calibration and alignment, and photorealistic rendering.

[0081] E1. Online calibration and alignment refinement: Initially, the preoperative virtual model is aligned with the laparoscope coordinate system through offline calibration. However, due to equipment errors and human body movements during the operation, there may be slight deviations in direct superposition. To this end, the system binds the preoperative virtual model to the MR Hand Tracking module, and the surgeon can achieve preliminary manual alignment with the laparoscope field of view by dragging, rotating and scaling with gestures; next, this embodiment adopts an online calibration optimization strategy: using the dense point cloud or key anatomical features reconstructed by SLAM, the position and scale of the virtual model are automatically and dynamically fine-tuned during the operation. For example, first, a series of three-dimensional key contour vertices evenly distributed along the edge of the organ or the fitted contour curve are extracted from the anatomical surface of the reconstructed virtual model. , obtain the posture of the virtual organ model in the mixed reality coordinate system , where R is the rotation matrix of the model, is the three-dimensional translation matrix of the model, each three-dimensional contour point , calculate its projection on the image plane ,in, is the homogeneous depth of the projected point; is the transposition symbol. Using the edge contour of the organ in the endoscope image as a reference, the true contour pixel point set is obtained , for each projection point , find its true contour set The nearest neighbor on , calculate the alignment error of this point , the sum of squared deviations of all contour points is taken as the optimization target , by adjusting the model pose In the rotation matrix and translation matrix parameters, search for a minimum in the rigid body transformation space The optimal model pose , thereby obtaining the optimal fine-tuning transformation , in order to find the optimal rigid body fine-tuning transformation, the Applied to the virtual model, the projection contour of the virtual organ can be precisely aligned with the image contour of the real organ. , first expand to homogeneous coordinates ; Then multiply the updated transformation matrix on the right , discard the last dimension and get the corrected three-dimensional coordinates Through these operations, the overlay error can be reduced to the sub-pixel level. Experiments have shown that after online calibration, the average deviation between the virtual model's anatomical structure and the endoscopic image feature points can be controlled within 0.5 pixels, reducing the projection error by approximately 40% compared to the uncalibrated state. This precise calibration ensures the precise overlap of virtual augmentation information with the real organ anatomy, providing the surgeon with reliable navigation guidance.

[0082] E2. Real-time rendering and enhanced realism: In terms of graphics rendering, this embodiment is optimized for medical scenarios to improve the realism of mixed reality. It uses a physically based rendering (PBR) shading model, using the physical material properties of organs (such as diffuse reflectance coefficient) , Highlight Specular Coefficient , subsurface scattering coefficient, etc.) to perform lighting calculations to more realistically represent the texture of the organs. Ambient occlusion and global illumination are considered in the rendering to ensure that the virtual organs are consistent with the surrounding lighting conditions. Specifically, the diffuse and specular components are calculated simultaneously for each rendered pixel: ,in Describes the cosine of the angle between the incident light and the surface normal, Describes the specular falloff ( is the reflection direction, is the sight direction, is the highlight index). By adjusting Parameters such as gloss and reflective properties are used to match the laparoscope image, and the estimated ambient light intensity is used to simulate the subsurface scattering effect of the organ, so that the rendered virtual organ blends in with the surrounding real organs in color and brightness. The rendering results are superimposed and output frame by frame with the real-time laparoscope video, and finally presented on the MR device or monitor for the surgeon to view and interact with. Thanks to GPU acceleration, the rendering pipeline of this system takes no more than 10ms per frame, and the overall rendering frame rate can reach 60Hz, corresponding to a visual delay of less than 20ms. In addition, combined with the above-mentioned tracking and calibration optimization, the total delay of the system from capturing the image to presenting the virtual overlay is controlled at around 50ms, which is far below the threshold of the human eye's perceptible delay. This means that the virtual model can be superimposed synchronously and stably as the endoscope moves, without obvious lag or jitter, ensuring the accuracy and immersion of the MR overlay during surgery.

[0083] In summary, this embodiment constructs a complete intraoperative real-time organ registration method, organically integrating binocular laparoscopic 3D reconstruction, NeRF self-supervised optimization, FEM physical modeling, SLAM tracking, and mixed reality rendering. The synergistic effect of multiple key technologies achieves high levels of depth estimation accuracy, deformation registration error, and real-time performance. For example, compared to a baseline method without these improvements, this embodiment improves reconstruction accuracy by approximately 60%, reduces registration error by more than half, and increases processing frame rate by more than twofold. In vitro animal experiments have demonstrated that the system can achieve stable and reliable virtual-real alignment in the complex and ever-changing laparoscopic surgical environment, with an average registration error of only 2-3 mm, meeting clinical accuracy requirements. Furthermore, it exhibits excellent real-time performance, achieving a processing speed of 20–30 Hz under typical hardware conditions, and keeping mixed reality rendering latency within tens of milliseconds, without disrupting the surgical workflow. Both theoretical analysis and experimental results demonstrate that the method of this embodiment offers significant advantages in the field of organ navigation and registration, providing physicians with intuitive and accurate 3D anatomical positioning and decision support, thereby enhancing surgical safety and accuracy.

[0084] In summary, the method disclosed in this embodiment for aligning models with organs to assist navigation is as follows: Figure 1 As shown, its core contents mainly include:

[0085] Step S1: convert the parallax of the color images on the left and right sides of the binocular laparoscope into depth information to obtain an initial depth map.

[0086] The left and right color images are first processed through a multi-scale CNN (Convolutional Neural Network) to extract feature pyramids, and image features at different scales are fused to obtain multi-channel feature representations for the left and right eyes. The self-attention and cross-attention mechanisms are implemented using a Transformer architecture based on epipolar constraints. The self-attention mechanism calculates pixel-level association weights along the epipolar direction in the feature sequence of the same view. The cross-attention mechanism performs feature matching along the epipolar direction on the left and right color images, retaining only disparity candidates whose inter-pixel matching probability meets a preset confidence threshold. A sparse probabilistic disparity cost volume is constructed, and each disparity cost volume is then regularized and smoothed and converted into a normalized probability distribution. The expectation of this probability distribution is then calculated as the disparity estimate.

[0087] Step S2: Perform self-supervised optimization on binocular depth estimation based on the trained NeRF model. The training process of the NeRF model specifically includes:

[0088] Step S21: Input the initial depth map as a depth prior together with the corresponding left or right single-side color image into the NeRF model. The NeRF model synthesizes a new image from a preset perspective through volume rendering technology, and compares the synthesized new image with the input left or right single-side color image to calculate the reprojection error.

[0089] Step S22: Process pixel depth and color rendering in the same rendering mode, replace the color of each sampling point with the corresponding depth position, divide it by the total weight to obtain the depth rendering, and then correspond each depth rendering calculation value to the corresponding pixel to obtain the depth rendering value of each pixel; then calculate the depth consistency loss between the depth rendering value inferred by the rendering and the disparity estimation value.

[0090] Step S23: Adopt the strategy of alternating training of NeRF model and Transformer architecture to gradually optimize the depth estimation results, iteratively run the fixed NeRF model to update the Transformer architecture to reduce the pixel color reprojection error and depth consistency error, and fix the Transformer architecture to update the NeRF model to improve the fitting accuracy of the geometric consistency of the laparoscopic view, and finally minimize the weighted sum of the reprojection error and depth consistency loss.

[0091] Step S3: Perform three-dimensional matching based on the two-dimensional information observed in the left or right color image combined with the depth information of each pixel and the shape and position of the virtual organ model to align the virtual organ model with the current morphology of the real organ.

[0092] Preferably, in step S1, in the process of calculating a sparse probabilistic disparity cost volume that meets a confidence threshold based on the matching probability between pixels, a sparse attention strategy is introduced to limit the disparity search to a matching probability range greater than the confidence threshold.

[0093] Preferably, in step S3, it specifically includes:

[0094] Step S31: Combine the physical simulation capabilities of finite element analysis with the machine learning capabilities of graph neural networks to achieve accurate registration of non-rigid deformations of organs.

[0095] Step S32: Use a visual SLAM algorithm to estimate the position of the laparoscope camera in real time.

[0096] Step S33: After obtaining the organ deformation and camera pose, the virtual organ model is superimposed on the surgical field image through mixed reality technology.

[0097] Preferably, the step S31 specifically includes:

[0098] Step S311: The finite element model calculates the deformation distribution of each part of the organ under the action of external forces based on the biomechanical properties and boundaries of the organ (usually anatomical boundaries discernible to the naked eye), providing physically realistic deformation results.

[0099] Step S312: Discretely represent the organ as a graph structure, and input the grid node displacement and topological connection obtained by finite element calculation into the GNN model;

[0100] Step S313, the GNN model uses organ grid nodes as graph nodes and grid topology as graph connections, and predicts the real-time deformation of the organ by learning the correlation and deformation pattern between nodes. The correlation includes the mechanical correlation between nodes learned by the message passing mechanism, and the mechanical correlation includes the mapping relationship from the boundary after displacement formed by the displacement or force on the organ surface to the displacement field of the internal node.

[0101] Preferably, the step S32 specifically includes:

[0102] Step S321: The SLAM algorithm extracts and matches the feature points in the left or right color image frame by frame to obtain the direction and position trajectory of the camera in the surgical environment, and simultaneously constructs a dense point cloud of the surgical area so that the camera positioning is constrained by the anatomical structure.

[0103] Step S322: Convert the virtual organ model to the camera coordinate system according to the camera positioning information for alignment rendering. The virtual organ model is usually reconstructed by preoperative CT or MRI.

[0104] Step S323: By minimizing the alignment error between the virtual image displayed by the MR device and the real image, the posture and position of the virtual organ model are adjusted in real time to ensure spatial alignment of the virtual organ model with the actual endoscopic color image of the left or right side. This step can be specifically implemented by pre-calibrating the MR display coordinate system and the surgical coordinate system using optical positioning methods to initially align the virtual model projection with the real image. Then, an iterative closest point (ICP) algorithm is used to minimize the alignment error between the virtual organ model and the laparoscopic color image with pixel-by-pixel depth information during the MR device overlay and alignment process. The rigid body transformation of the virtual model is updated for each frame to ensure spatial alignment of the virtual organ model with the actual endoscopic color image of the left or right side.

[0105] Preferably, the method of the present invention further comprises: using Kalman filtering to fuse the camera pose solved by visual SLAM with the high-frequency acceleration, angular velocity and EM positioning results measured by IMU, wherein the IMU and EM are different sensors deployed on the laparoscope lens.

[0106] Preferably, the step S33 specifically includes:

[0107] Step S331: Obtain the internal and external parameters of the laparoscopic camera calibration, where the internal parameters include focal length, principal point, and distortion coefficient, and the external parameters include the position in the surgical coordinate system; set the parameters of the virtual camera according to the calibration results of the internal and external parameters so that the rendered virtual model perspective is consistent with the real laparoscopic camera.

[0108] Step S332: transform the virtual organ model into the surgical coordinate system, and adjust the size and position of the virtual organ model according to the actual anatomical proportions to ensure that the virtual organ model is spatially aligned with the real organ.

[0109] Step S333: During the rendering process, lens illumination matching is performed, distortion correction of the virtual organ image is performed according to the lens characteristics of the laparoscope, and the lighting effect of the surgical light source on the virtual organ model is simulated; and the dense point cloud or key anatomical features reconstructed by SLAM are used to dynamically adjust the position and scale of the virtual organ model during the operation.

[0110] It is worth noting that in the prior art, some keywords overlap with those in the above steps, such as the method for establishing a disparity prediction model and depth estimation method for endoscopic images disclosed in domestic patent application document CN113435573A, which involves the relationship between epipolar lines and disparity depth. In the method for generating large parallax new perspective images based on NERF disclosed in patent application document CN118037915A, the large parallax new perspective images are generated based on predicted depth information. However, the causal relationship between overlapping keywords in such prior art differs from the scenario in this embodiment in many essential ways, specifically:

[0111] The epipolar lines mentioned in patent application document CN118037915A are only used to filter the data set to ensure that the constructed data set is a data set that has been corrected by the epipolar lines. This is essentially different from the present invention, which performs feature matching of the left and right images based on the epipolar lines, and then filters disparity candidates based on the matching results to construct a sparse probabilistic disparity cost volume, an intermediate variable. Although both are used for disparity calculation, the present invention uses probability volumes and threshold screening to reduce the search space, while patent application document CN118037915A directly calculates the convolution cost within the full disparity range and does not use a similar discretized disparity cost volume structure. The technical implementation method and optimization focus are different.

[0112] Similarly, there is an essential difference in motivation between the patent application document CN113435573A, which generates a new perspective image with large parallax based on predicted depth information, and the present embodiment, which optimizes the initial depth information in a self-supervised manner based on synthesized new images (in other words, although both mention volume rendering, the volume rendering in the present invention is used for view synthesis to assist depth optimization, while the volume rendering in CN118037915A is used to directly generate new perspective images under large parallax. The functions and contexts are different, so the application scenarios and technical effects are not equivalent). In addition, the attention mechanism in the patent application document CN113435573A is used for stereo matching under multi-view information fusion, which is different from the matching strategy of the present invention, in which self-attention calculates pixel-level correlation in the same view feature sequence, and cross-attention matches features between the left and right views along the epipolar direction, so the technical effects are also inconsistent.

[0113] In other words, this embodiment filters low-confidence matches through self-attention and cross-attention along the polar direction, and converts the matching probability into the expected disparity; then uses the NeRF model to perform self-supervision optimization on the depth map; finally, combines the shape and position information of the virtual organ model to perform three-dimensional matching; its core purpose is to achieve high-precision real-time navigation and registration. Among them, the Transformer matching architecture combined with confidence screening can effectively filter out false matches and improve the accuracy of the initial depth map; the NeRF-based self-supervision optimization further makes the depth estimation more consistent with the real view geometry, enhancing the accuracy and stability of the navigation results. Although some of the existing technologies are of reference significance for the implementation of the discrete technical points of the present invention, the overall technology cannot enable those skilled in the art to obtain relevant technical inspiration on how to convert and apply it to provide navigation and registration of models and organs in clinical applications such as surgery.

[0114] Example 2

[0115] This embodiment discloses a system for aligning a model with an organ to assist navigation, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method of the above embodiment is implemented.

[0116] In summary, the methods and systems disclosed in the embodiments of the present invention respectively have at least the following beneficial effects:

[0117] 1. Introducing the Transformer into stereo matching realizes a new global context-aware epipolar matching mechanism. Traditional stereo matching is usually limited by local correlation, but the Transformer architecture based on this invention can establish pixel correspondences on a global scale, significantly improving matching accuracy. At the same time, after feature matching along the epipolar direction of the left and right color images, only disparity candidates with inter-pixel matching probabilities that meet a preset confidence threshold are retained to construct a sparse probabilistic disparity cost volume. Each disparity cost volume is then regularized and smoothed and converted into a normalized probability distribution. The expectation of this probability distribution is then calculated as the disparity estimate, significantly improving the real-time performance of the disparity calculation.

[0118] 2. Introducing NeRF as a depth supervision signal, depth optimization is transformed into a differentiable new perspective synthesis problem, optimizing depth results from the perspective of global photometric consistency. Under the global supervision of NeRF, the shortcomings of traditional stereo matching in low-texture or high-gloss areas are effectively compensated, making depth estimation in tricky areas equally accurate. At the same time, the joint optimization of the depth consistency constraint and the NeRF-Transformer stereo network ensures that the depth map meets the consistency of the reprojected photos while maintaining geometric accuracy; the innovative strategy of integrating neural rendering with stereo vision greatly improves the details and reliability of depth estimation, achieving higher accuracy.

[0119] 3. Based on the two-dimensional information observed in the left or right color image, the depth information of each pixel is combined with the shape and position of the virtual organ model to perform three-dimensional matching, ensuring the accuracy of the alignment of the virtual organ model with the current morphology of the real organ.

[0120] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for registering a model with an organ to assist navigation, characterized in that: include: Step S1, converting the parallax of the color images on the left and right sides of the binocular laparoscope into depth information to obtain an initial depth map; specifically comprising: The left and right color images are first processed through a multi-scale CNN to extract feature pyramids, and the image features of different scales are fused to obtain multi-channel feature representations for the left and right eyes. The self-attention and cross-attention mechanisms are implemented using a Transformer architecture based on epipolar constraints. The self-attention mechanism calculates pixel-level association weights along the epipolar direction in the feature sequence of the same view. The cross-attention mechanism performs feature matching along the epipolar direction on the left and right color images, retaining only disparity candidates whose inter-pixel matching probability meets a preset confidence threshold. A sparse probabilistic disparity cost volume is constructed, and each disparity cost volume is then regularized and smoothed and converted into a normalized probability distribution. The expectation of this probability distribution is then calculated as the disparity estimate. Step S2: self-supervised optimization of binocular depth estimation based on the trained NeRF model. The training process of the NeRF model specifically includes: Step S21: The initial depth map is input as a depth prior together with the corresponding left or right single-side color image into the NeRF model. The NeRF model synthesizes a new image from a preset perspective using volume rendering technology, and compares the synthesized new image with the input left or right single-side color image to calculate the reprojection error. Step S22: Process pixel depth and color rendering in the same rendering mode, replace the color of each sampling point with the corresponding depth position, divide it by the total weight to obtain the depth rendering, and then map each depth rendering calculated value to the corresponding pixel to obtain the depth rendering value of each pixel; then calculate the depth consistency loss between the depth rendering value inferred by the rendering and the disparity estimation value; Step S23: Adopting a strategy of alternating training of the NeRF model and the Transformer architecture to gradually optimize the depth estimation results, iteratively running the fixed NeRF model to update the Transformer architecture to reduce the pixel color reprojection error and the depth consistency error, and then fixing the Transformer architecture to update the NeRF model to improve the fitting accuracy of the laparoscopic view geometric consistency, ultimately minimizing the weighted sum of the reprojection error and the depth consistency loss; Step S3: Perform three-dimensional matching based on the two-dimensional information observed in the left or right color image combined with the depth information of each pixel and the shape and position of the virtual organ model to align the virtual organ model with the current morphology of the real organ.

2. The method for model and organ registration to assist navigation according to claim 1, characterized in that: In step S1, in the process of calculating a sparse probabilistic disparity cost volume that meets a confidence threshold based on the matching probability between pixels, a sparse attention strategy is introduced to limit the disparity search to a matching probability range greater than the confidence threshold.

3. The method for model and organ registration to assist navigation according to claim 1, characterized in that: In step S3, it specifically includes: Step S31: combining the physical simulation capabilities of finite element analysis with the machine learning capabilities of graph neural networks to achieve accurate registration of non-rigid organ deformations; Step S32: using a visual SLAM algorithm to estimate the position of the laparoscope camera in real time; Step S33: After obtaining the organ deformation and camera pose, the virtual organ model is superimposed on the surgical field image through mixed reality technology.

4. The method for assisting navigation by aligning a model with an organ according to claim 3, wherein: The step S31 specifically includes: Step S311: The finite element model calculates the deformation distribution of each part of the organ under the action of external forces based on the biomechanical properties and boundaries of the organ, providing physically realistic deformation results; Step S312: Discretely represent the organ as a graph structure, and input the grid node displacement and topological connection obtained by finite element calculation into the GNN model; Step S313, the GNN model uses organ grid nodes as graph nodes and grid topology as graph connections, and predicts the real-time deformation of the organ by learning the correlation and deformation pattern between nodes. The correlation includes the mechanical correlation between nodes learned by the message passing mechanism, and the mechanical correlation includes the mapping relationship from the boundary after displacement formed by the displacement or force on the organ surface to the displacement field of the internal node.

5. The method for assisting navigation by aligning a model with an organ according to claim 3, wherein: The step S32 specifically includes: Step S321: The SLAM algorithm extracts and matches feature points in the left or right color image frame by frame to minimize the reprojection error and obtains the direction and position trajectory of the camera in the surgical environment. At the same time, a dense point cloud of the surgical area is constructed so that the camera positioning is constrained by the anatomical structure. Step S322: converting the virtual organ model to the camera coordinate system according to the camera positioning information for alignment rendering; Step S323: By minimizing the alignment error between the virtual image displayed by the MR device and the real image, the posture and position of the virtual organ model are adjusted in real time to ensure that the virtual organ model is spatially aligned with the actual endoscopic left or right side color image.

6. The method for assisting navigation by aligning a model with an organ according to claim 5, wherein: Also includes: The Kalman filter is used to fuse the camera pose solved by visual SLAM with the high-frequency acceleration, angular velocity and EM positioning results measured by the IMU. The IMU and EM are different sensors deployed on the laparoscope lens.

7. The method for assisting navigation by aligning a model with an organ according to claim 6, wherein: The step S33 specifically includes: Step S331: Obtain the internal and external parameters of the laparoscopic camera calibration, wherein the internal parameters include focal length, principal point, and distortion coefficient, and the external parameters include the position in the surgical coordinate system; set the parameters of the virtual camera according to the calibration results of the internal and external parameters so that the perspective of the rendered virtual model is consistent with that of the real laparoscopic camera; Step S332: transforming the virtual organ model to the surgical coordinate system, and adjusting the size and position of the virtual organ model according to the actual anatomical proportions to ensure that the virtual organ model is spatially aligned with the real organ; Step S333: During the rendering process, lens illumination matching is performed, distortion correction of the virtual organ image is performed according to the lens characteristics of the laparoscope, and the lighting effect of the surgical light source on the virtual organ model is simulated; and the dense point cloud or key anatomical features reconstructed by SLAM are used to dynamically adjust the position and scale of the virtual organ model during the operation.

8. The method for assisting navigation by aligning a model with an organ according to claim 7, wherein: The process of dynamically adjusting the position and scale of the virtual organ model during surgery specifically includes: Step S3331: extract a series of three-dimensional key contour vertices evenly distributed along the organ edge or the fitted contour curve from the anatomical surface of the reconstructed virtual model. , The vertex index of the contour is used to obtain the posture of the virtual organ model in the mixed reality coordinate system. ,in, is the rotation matrix of the model, is the three-dimensional translation matrix of the model, each three-dimensional contour point , calculate the projection of each contour point on the image plane ,in, is the homogeneous depth of the projected point, is the two-dimensional coordinate of the contour vertex in the image plane; is the transpose symbol; Step S3332: Using the edge contour of the organ in the actual image as a reference, obtain the real contour pixel point set For each projection point , find each projection point in the true contour set The nearest neighbor point on the , calculates the alignment error of the nearest neighbor point, takes minimizing the sum of squared deviations of all contour points as the optimization goal, and finds the optimal rigid body fine-tuning transformation.

9. A system for registering a model with an organ to assist navigation, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • NERF-based large-parallax new-view-angle image generation method

    CN118037915A

  • Parallax prediction model building method and depth estimation method of endoscope image

    CN113435573A

  • Sparse reconstruction method and device based on three-dimensional scene

    CN117953151A