Multi-mode large model assisted laparoscope soft tissue registration surgical navigation method and system

By fusion of preoperative and intraoperative image features through multimodal large models and combining RANSAC and TPS deformation models, high-precision and real-time registration of soft tissues in laparoscopic surgery is achieved, solving the problems of limited accuracy and high calculation delay in traditional methods.

CN119941813APending Publication Date: 2025-05-06THE FIRST AFFILIATED HOSPITAL OF MEDICAL COLLEGE OF XIAN JIAOTONG UNIV

Patent Information

Application Number
CN202510420876.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In laparoscopic surgery, it is difficult for the prior art to achieve high-precision and real-time registration of soft tissues, especially in multimodal data fusion and large deformation scenarios of soft tissues. Traditional methods have problems of limited accuracy and high calculation delay.

Method used

The multimodal large model-assisted laparoscopic soft tissue registration method is used to fuse the multimodal features of preoperative CT/MRI images and intraoperative laparoscopic images through the Transformer model, and combine the RANSAC and TPS deformation models to achieve cross-modal matching and non-rigid deformation registration.

Benefits of technology

It realizes automatic, high-precision, real-time registration of multi-organ soft tissues under multimodal and multivariable conditions, reducing power consumption and computing time, and providing trusted navigation and positioning support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941813A_ABST
    Figure CN119941813A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal large model assisted laparoscope soft tissue registration surgical navigation method and system. The method comprises the following steps: acquiring a preoperative CT (Computed Tomography) or MRI (Magnetic Resonance Imaging) image and an intraoperative laparoscope image, and converting to obtain a multi-modal feature sequence; inputting the obtained multi-modal feature sequence into a large model for fusion to obtain a cross-modal matching matrix; random sampling is carried out by using an RANSAC random sampling consensus algorithm, an error between corresponding point sets of two data is measured by using an Euclidean distance, an analytical solution is obtained through SVD decomposition, and a feature matching error is introduced to obtain an optimal rigid body transformation pose; performing non-rigid deformation local fitting adjustment on the optimal rigid body transformation pose residual error by adopting a parameterized deformation model in combination with learning prediction deformation, and performing loop execution to obtain a soft tissue organ real-time registration image; automatic, high-timeliness and high-precision real-time registration of soft tissues is realized under the multi-mode and multi-variable conditions, and reliable navigation and positioning support is ensured to be provided in a complex operation environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of surgical navigation and tracking, and in particular relates to a multi-modal large model-assisted laparoscopic soft tissue registration surgical navigation method and system. Background Art

[0002] In laparoscopic surgery, information from a single modality is often insufficient to provide reliable constraints: for example, it is difficult to determine spatial depth based on laparoscopic two-dimensional images alone, and it is difficult to find corresponding anatomical structures in the intraoperative field of view based on CT models alone. Some existing methods attempt to align preoperative and intraoperative data by manually selecting points or matching based on contours and edges, but manual labeling is time-consuming and susceptible to subjective influences; edge / intensity-based methods (such as 2D / 3D registration using organ contour coincidence or mutual information) have limited accuracy in soft tissue scenarios, especially when internal tissues lack obvious features or undergo large deformations. Therefore, laparoscopic surgical navigation is required to match and register preoperative imaging data (such as three-dimensional anatomical information provided by CT and MRI) with intraoperative laparoscopic data (two-dimensional / three-dimensional images obtained in real time). Preoperative CT / MRI and intraoperative laparoscopic images belong to different imaging modalities, and the direct correspondence is unclear. How to extract and match features with physical correspondence is a difficult problem. At the same time, the shape of organs under laparoscopy will change dynamically due to factors such as traction and breathing. The shape and position of organs during surgery may be quite different from those before surgery. This results in the inability of traditional rigid registration methods (such as the classic ICP algorithm) to cover this elastic deformation when directly used on soft tissues, and it is difficult to achieve ideal results. It is necessary to introduce a non-rigid registration mechanism to deal with local differences. Moreover, due to the limitation of the endoscopic viewing angle, the target organs are often partially visible or even temporarily completely invisible during surgery. In addition, laparoscopic images are reflected light images, which have problems such as illumination changes and instrument reflections. The texture presented is essentially different from the transmission imaging of CT / MRI, which increases the difficulty of multi-view alignment.

[0003] Another key issue is real-time computing performance. The registration algorithm needs to be fast enough to keep up with the progress of the surgery, while being insensitive to local information loss caused by occlusion, blood, etc. and initial alignment errors, and maintaining stable and reliable output to avoid interrupting intraoperative operations (1080p dual-channel video), which means that each frame needs to be processed within 33 milliseconds to meet the real-time requirement of 30 FPS. Some current methods are still not fully optimized on hardware such as GPU / TPU / FPGA, resulting in high algorithm latency, which is not conducive to a smooth augmented reality (AR) navigation experience.

[0004] Existing registration algorithms are often only applicable to a single imaging mode. Monocular cameras only provide 2D images, while binocular / depth cameras provide 3D point clouds. They cannot adapt to the anatomical differences of various organs and are only applicable to specific scenarios or rely on manual adjustments by doctors.

[0005] To solve the above difficulties, the field of surgical navigation urgently needs a matching and registration technology solution that can adapt to large deformation of soft tissue, integrate multi-source information and be automatic and robust. Summary of the invention

[0006] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a multi-modal large model-assisted laparoscopic soft tissue registration surgical navigation method and system, which can integrate multi-source information, realize automatic, high-precision real-time registration of multi-organ soft tissues under multi-modal and multi-variable conditions, realize low-power and high-speed registration operations of the system, and ensure reliable navigation and positioning support in complex surgical environments.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions: A multi-modal large model-assisted laparoscopic soft tissue registration surgery navigation method comprises the following steps: Step 1: Obtain preoperative CT or MRI images and intraoperative laparoscopic images, and convert the image data into intermediate feature representations that can be processed by the Transformer model to obtain a multimodal feature sequence; Step 2: Input the multimodal feature sequence obtained in step 1 into the fusion layer of the Transformer large model to fuse the features of different modalities and infer the corresponding relationship to obtain a cross-modal matching matrix; Step 3: Using the cross-modal matching matrix obtained in step 2, the RANSAC random sampling consensus algorithm is used for random sampling, and the error between the corresponding point sets of the two data is measured by Euclidean distance. The analytical solution is obtained through SVD decomposition and the feature matching error is introduced for rigid registration to obtain the optimal rigid body transformation posture; Step 4: Use the parameterized thin plate spline TPS deformation model and combine it with Transformer learning to predict deformation to perform local fitting adjustment of the best rigid body transformation pose residual obtained in step 3 to obtain a soft tissue organ registration image; Execute steps 1 to 4 in a loop in real time to obtain a real-time registration image of soft tissue organs; Based on the comprehensive loss function, an end-to-end strategy is used to train the optimal model.

[0008] The present invention also has the following technical features: Preferably, the data conversion method described in step 1 includes: Perform image segmentation on preoperative CT or MRI to obtain a 3D reconstructed image of the target organ, extract the key anatomical landmarks of the organ and their coordinates in the 3D model, use the network to extract the markers representing the surface / volume features of the preoperative organ, and add spatial location information; For intraoperative laparoscopic images obtained by binocular laparoscopes or systems with depth sensing, a point cloud of the organ surface depicting the three-dimensional shape of the current canvas surface of the organ in the scene is obtained through stereo matching or structured light reconstruction. , using point cloud network to point cloud Calculate local geometric features at each point , then embed the point features into a fixed-length feature vector, and add the three-dimensional coordinate position encoding of the point Or the grid index, forming a mark with location information; Two-dimensional video frames of intraoperative laparoscopic images obtained by monocular laparoscope , select key frame images for registration, and use the depth estimation network to predict the depth map to obtain rough point cloud or depth information as an aid; use convolutional neural network to image Extracting multi-scale visual feature maps , expand the feature map into a series of feature vector identifiers, each of which corresponds to a local area on the image and attaches its two-dimensional position code; Through full connection mapping and normalization processing, the feature vector identifiers of preoperative and intraoperative images have intermediate feature representations with the same dimension, and a multimodal feature sequence is obtained.

[0009] Preferably, the fusion of different modal features and the inference process of the corresponding relationship described in step 2 include: Transformer performs multi-head self-attention calculations on each modality to enhance local structural features; Adopt Transformer’s multi-head cross-attention mechanism to establish cross-modal associations; The cross-modal matching matrix is ​​obtained through the multi-level fusion iteration of Transformer.

[0010] Preferably, the process of solving the optimal rigid body transformation posture described in step 3 includes: Preoperative point collection Intraoperative point set For a matching pair, the mean square error of the rigid registration is defined as: , in, is the rotation matrix, is the translation vector, It is a matching relationship. ; Calculate matching point set and The centroid of the data is centered and the covariance matrix is ​​constructed: ,right Perform singular value decomposition ; The optimal rotation matrix is , translated to ; Will Applied to the intraoperative point set, preliminary alignment with the preoperative model was achieved; Introducing the feature map learned by Transformer , so that each point Both have high-dimensional descriptions and , define the comprehensive error: , Among them, the parameters Weigh the importance of geometric distance and feature distance by minimizing , so that the spatial distance and feature differences of the corresponding points are reduced, thus obtaining a more accurate .

[0011] Preferably, the thin plate spline TPS deformation model described in step 4 represents the non-rigid deformation as an affine transformation plus a radial basis function deformation, and the formula is as follows: in, for Affine matrix, is the translation vector, is the selected control point position, , is the TPS radial basis kernel function, is the weight to be sought, Represents the spatial position coordinates of any point in the organ model reconstructed by preoperative 3D CT or MRI to be registered; Introducing bending energy To penalize severe deformation; weight the registration error and bending penalty to form the total energy , taking its partial derivative = 0, we can get The closed-form solution of The process of Transformer learning to predict deformation described in step 4 includes: Using neural network to predict Move to The displacement vector , equivalently, predicts the deformation field make , the preoperative point Mapped to the corresponding position during surgery.

[0012] Preferably, the comprehensive loss function is composed of a weighted sum of cross entropy loss, alignment error and regularization constraint: For the cross-modal matching matrix described in step 2, by comparing the real correspondences generated by manual annotation or simulation, a cross entropy loss is applied to increase the attention weight of the correct correspondence , suppressing false correspondences so that the model learns to output accurate matching pairs; For the output rigid body transformation or non-rigid deformation The MSE loss is used to measure the difference between the predicted and true parameters; or the predicted transformation is applied to the preoperative image feature vector, and the Chamfer distance with the intraoperative laparoscopic image feature vector is calculated as the loss. The Chamfer distance is defined as: ; For non-rigid deformation parameters, add elastic energy constraints based on physical models ; The final loss function The formula is as follows: , in, Represents large model parameters Predicted non-rigid deformation transformation, is the matching correspondence inferred by the model, is the regularization term, For its weight.

[0013] The present invention also protects a multimodal large model-assisted laparoscopic soft tissue registration surgical navigation system, including a multimodal image acquisition module, an image preprocessing module, a deep learning model fusion and registration calculation module, a calculation acceleration module, and an intraoperative navigation display and control module; The computing acceleration module includes parallel GPU, TPU and FPGA units, wherein the GPU unit optimizes convolution and matrix operations using the CUDA parallel computing framework to accelerate the calculation of convolutional networks and Transformer attention modules; the TPU unit is used to optimize the structure of deep learning matrix operations, improve the efficiency of large-scale matrix multiplication and tensor operations, and thus accelerate model reasoning; the FPGA unit maps the convolutional neural network and attention mechanism into digital circuits that can be parallel pipelined, realizes concurrent execution of each layer of the model, and further improves the reasoning speed and reduces power consumption by customizing low-precision computing units; The multimodal image acquisition module is used to obtain preoperative CT or MRI images and intraoperative laparoscopic images; The image preprocessing module is used to perform segmentation, normalization and feature extraction on the collected multimodal images; The deep learning model fusion and registration calculation module includes several layers of Transformer encoders and decoders based on a multi-head attention mechanism, which are used to perform rigid registration and non-rigid registration on the pre-processed multimodal image data, and obtain the final matching relationship and deformation parameter estimation after several feedforward networks and normalization; The intraoperative navigation display and control module, the control module stores executable instructions, and when the executable instructions are executed, the multi-modal large model-assisted laparoscopic soft tissue registration surgical navigation method as described above is implemented, and the registration image is displayed on the navigation display module.

[0014] Preferably, the system further comprises an augmented reality overlay display module; The augmented reality overlay display module projects the real-time registration images of multiple organs onto the surgical area through an adaptive projection algorithm, and adjusts the viewing angle and depth perception of the projection in real time.

[0015] Compared with the prior art, the present invention has the following technical effects: In order to enhance the robustness and versatility of the algorithm, the navigation method of the present invention introduces a deep learning large model represented by Transformer to process data of different modalities. The large capacity and self-attention mechanism of Transformer enable it to learn the soft tissue deformation law and cross-modal correspondence from massive training samples, and can extract more discriminative high-dimensional feature descriptions, and establish a soft correspondence between the preoperative 3D model and the intraoperative image through the attention mechanism, so as to find the possible corresponding anatomical structure pairs, and embed the global three-dimensional anatomical structure provided by CT / MRI into the laparoscopic image features through multimodal feature fusion to improve the matching accuracy, and use the current morphology provided by the real-time laparoscope image to correct the position and shape of the CT model, and can globally associate the corresponding areas of the preoperative and intraoperative data. Even if there is local occlusion or deformation, the match can be found through global information, so that the registration result fits the actual surgical situation, which can significantly reduce the uncertainty caused by texture loss or perspective change when relying on a single modality; The navigation method of the present invention integrates the non-rigid deformation modeling strategy, so that the registration is not limited to rigid body transformation, but can also adapt to the local elastic changes of soft tissue. By introducing the deformation field prediction mechanism in the deep network, the deformation alignment process is integrated into the end-to-end learning, so that the algorithm can automatically learn the mapping relationship of soft tissue deformation from the initial state to the intraoperative state, directly output the deformation parameters of the organ model or the coordinates after deformation, and realize the elastic alignment of the preoperative model. At the same time, regular constraints with physical significance are added to ensure that the predicted deformation is reasonable and feasible, thereby improving the registration accuracy and avoiding singular solutions; enhancing the adaptability of the algorithm to large deformation of soft tissue, and realizing a more general organ registration scheme; combining the soft tissue registration method of the large model with global modeling and multimodal context, it can still maintain stable matching performance in the face of unknown deformation situations; The navigation method of the present invention utilizes the end-to-end training framework of deep learning to optimize the feature extraction-matching-registration links in a unified manner. A comprehensive loss function is designed during training to enable the model to learn feature matching and pose solving at the same time. After sufficient training, only one forward calculation is required during reasoning to output the matching and registration results, which is equivalent to using a neural network to approximate the ICP iterative optimization process, greatly improving the registration efficiency and reducing the sensitivity to the initial pose. In addition, through transfer learning and data enhancement, the model can also be generalized to different patients and organs: training on massive real and simulated surgical data allows the model to see various possible organ deformations and imaging changes, thereby improving the robustness to unknown scenes; Through the innovative combination of large model + multimodal data fusion, a registration method from global to local, from rigid to non-rigid is provided, giving full play to the ability of deep learning to fit complex patterns and the accuracy of classical geometric algorithms: the large model provides "perceptual intelligence" for soft tissue anatomical changes, and classical registration provides physically reliable constraint solutions. The organic combination of the two makes the system both "smart" and "robust"; it can automatically overcome various uncertainties in soft tissue registration, achieve higher registration accuracy and robustness, and provide reliable support for laparoscopic surgical navigation; The navigation system of the present invention combines GPU parallel computing, model compression and dedicated hardware acceleration strategies, optimizes key steps such as Transformer attention calculation and nearest neighbor search through CUDA parallelism, accelerates the reasoning speed, and utilizes the parallel computing capabilities of FPGA / TPU to achieve low-power, high-speed edge reasoning, thus ensuring that the algorithm of the present invention can achieve near real-time alignment on a high-performance workstation and meet clinical use requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic flow chart of the laparoscopic multi-soft tissue organ matching and registration surgical navigation method of the present invention with multi-modal large model fusion; Figure 2 is a schematic diagram of the optimization details of the registration algorithm of the present invention; Figure 3 It is a comparative schematic diagram of the optimization of the registration calculation process of the present invention; Figure 4 It is a schematic diagram of the algorithm acceleration deployment process of the present invention; Figure 5 It is a schematic diagram of the structure of the accelerated deployment module of the system of the present invention; Figure 6 It is the image display of the registration process of soft tissue organs in laparoscopic surgery of the present invention; Figure 7 The present invention is an image display of the tracking process of soft tissue organs in laparoscopic surgery. DETAILED DESCRIPTION

[0017] The specific contents of the present invention are further explained in detail below in conjunction with embodiments.

[0018] Embodiment 1 This embodiment provides a multi-modal large model-assisted laparoscopic soft tissue registration surgery navigation method. Figure 1 As shown, the following steps are included: Step 1: Obtain preoperative CT or MRI images and intraoperative laparoscopic images, and convert the image data into intermediate feature representations that can be processed by the Transformer model to obtain a multimodal feature sequence; Perform image segmentation on preoperative CT or MRI to obtain a three-dimensional reconstructed image of the target organ; extract key anatomical landmarks of the organ (such as organ edges and vascular bifurcation points) and their coordinates in the three-dimensional model, use the network to extract markers representing the surface / volume features of the preoperative organ, and add spatial location information; Select the feature extraction method based on the representation form of preoperative CT or MRI images. For example, if the image is a point cloud or triangular mesh, use the PointNet series network to extract the features of each node. If the image is a voxel mesh, use 3D CNN or input the voxel data into Transformer to extract voxel features. Regardless of the form, a set of identifiers representing the surface / volume features of the preoperative organ is obtained, with additional spatial position information, such as 3D coordinates or voxel indices.

[0019] For intraoperative laparoscopic images obtained by binocular laparoscopes or systems with depth sensing, a point cloud of the organ surface depicting the three-dimensional shape of the current canvas surface of the organ in the scene is obtained through stereo matching or structured light reconstruction. , using PointNet++ point cloud network to point cloud Calculate local geometric features at each point These features reflect the shape features of the point neighborhood, such as curvature and normal vector, and then embed the point features into a fixed-length feature vector, while adding the three-dimensional coordinate position encoding of the point Or the grid index, forming a mark with location information; Two-dimensional video frames of intraoperative laparoscopic images obtained by monocular laparoscope , select key frame images for registration, and use the depth estimation network to predict the depth map to obtain rough point cloud or depth information as an aid; use convolutional neural network (such as ResNet) to image Extracting multi-scale visual feature maps ,These features encode the texture, edge and other information in the image, and then expand the feature map into a series of feature vector identifiers, each of which corresponds to a local area on the image and is attached with its two-dimensional position code; Through full connection mapping and normalization processing, the feature vector identifiers of different scales and modalities of preoperative and intraoperative images are mapped to a unified feature space and dimension to form an intermediate feature representation with the same dimension. For example, the initial camera pose provided by the surgical navigation and positioning system is aligned, and the size scale is unified and the color is normalized. At the same time, position encoding (optional sine function encoding or learnable position vector) is used to attach the position of each feature vector identifier so that the model can perceive the spatial position relationship corresponding to the feature in subsequent processing. Finally, the output is a feature vector identifier sequence containing multimodal features of the laparoscope and preoperative model, that is, a multimodal feature sequence.

[0020] Step 2: Input the multimodal feature sequence obtained in step 1 into the fusion layer of the Transformer large model to fuse the features of different modalities and infer the corresponding relationship to obtain a cross-modal matching matrix; Transformer first performs multi-head self-attention calculations on each modality to strengthen local structural features. For example, in the feature vector identification sequence of laparoscopic images, self-attention combines the context of the entire image to update each local feature so that it contains global texture and edge information. In the feature vector identification sequence of point cloud, self-attention aggregates the geometric information of neighboring points to strengthen local structural features. After self-attention, the obtained features are more robust and semantically rich, laying the foundation for cross-modal matching.

[0021] The multi-head cross-attention mechanism of Transformer is used to establish cross-modal associations, allowing the two modal features to "talk" to each other; for example, using laparoscopic features as the query , with preoperative CT features as key and value pairs , calculate the attention weight to converge the corresponding Information to superior.

[0022] For laparoscopic modality Features and preoperative modality Features , attention score The dot product calculation by softmax normalization is obtained: in, is the feature dimension, is the number of preoperative characteristics.

[0023] Attention weight matrix Reflects laparoscopic features Preoperative characteristics When When it is larger, it can be regarded as the network believes (the spatial position of the intraoperative feature) and (Preoperative feature position) There is a correspondence. Through multi-head cross attention, the model matches features from different angles (shape, color, anatomical structure and other subspaces), which improves the reliability of association to find the true correspondence.

[0024] The cross-modal matching matrix is ​​obtained through the multi-level fusion iteration of Transformer.

[0025] The features output by each layer will incorporate more information from the other modality. layer, the laparoscopic features are gradually embedded in the global anatomical context of the CT model, and the CT features also absorb the current morphological features observed by the laparoscope. After sufficient cross-fusion, the Transformer outputs the multimodal context features of each modality. In particular, for the feature sequence on the laparoscopic side, it already contains clues to the corresponding preoperative CT features - this is reflected in the final attention matrix of the Transformer. Each laparoscopic token will pay close attention to certain preoperative tokens, thus forming a soft correspondence matrix These soft correspond It is actually a prediction of cross-modal matching: it gives each pair of points / pixels a confidence score that they belong to a corresponding relationship Note that soft correspondence allows for one-to-many or many-to-one relationships, which is a modeling of soft tissue conditions (since soft tissue may have missing or new points, the correspondence is not strictly one-to-one). By setting a threshold or performing a bipartite graph matching algorithm, the matrix Extract the set of matching pairs with the highest confidence , as the constraint basis for subsequent precise registration. In general, the large model automatically completes feature matching through the above process, replacing the manual or nearest neighbor-based matching steps in traditional ICP, and realizing intelligent and robust correspondence inference.

[0026] Step 3: Using the cross-modal matching matrix obtained in step 2, the RANSAC random sampling consensus algorithm is used for random sampling, and the error between the corresponding point sets of the two data is measured by Euclidean distance. The analytical solution is obtained through SVD decomposition and the feature matching error is introduced for rigid registration to obtain the optimal rigid body transformation posture. ,like Figure 2 As shown; The Euclidean distance is used to measure the error between two corresponding data point sets; Preoperative point collection Intraoperative point set For a matching pair, the mean square error of the rigid registration is defined as: , in, is the rotation matrix, is the translation vector, It is a matching relationship. ; The above quadratic error can be optimized and the analytical solution can be obtained through SVD decomposition; the matching point set is calculated and The centroid of the data is centered and the covariance matrix is ​​constructed: ,right Perform singular value decomposition ; The optimal rotation matrix is , translated to ; Will Applied to the intraoperative point set, preliminary alignment with the preoperative model was achieved; Unlike ICP, this paper not only relies on geometric error when solving the transformation, but also incorporates feature matching error to improve accuracy; it introduces the feature mapping learned by Transformer , so that each point Both have high-dimensional descriptions and , define the comprehensive error: , Among them, the parameters Weigh the importance of geometric distance and feature distance by minimizing , so that the spatial distance and feature differences of the corresponding points are reduced, thus obtaining a more accurate Equivalently, this is equivalent to the deep network providing a similarity measure to assist in the registration during the matching process. The Transformer's attention mechanism has minimized the feature differences of the corresponding points, that is, giving high weights to the correct corresponding points, which is equivalent to letting Therefore, the optimization It can better take into account appearance changes and counteract the interference of soft tissue deformation on registration.

[0027] In order to avoid the influence of wrong matching point pairs on the solution, the present invention preferably uses the RANSAC random sampling consensus algorithm. It repeatedly randomly samples a small number of matching pairs from a large number of soft matching pairs, calculates the corresponding rigid body transformation, and evaluates the number of internal points that are adapted. Finally, a set of matching with the most internal points (the smallest error) is selected to obtain the As a global rigid body solution, RANSAC can effectively eliminate the interference of mismatched pairs and make the rigid body solution more robust.

[0028] Through the above steps, the preoperative model has been rotated and translated so that it is aligned with the intraoperative data in terms of rough rigidity. This provides a unified reference framework for the next step of processing soft tissue elasticity.

[0029] Step 4: Use the parameterized thin plate spline TPS deformation model and combine it with Transformer learning to predict deformation to perform non-rigid registration on the best rigid body transformation pose residual obtained in step 3 by local fitting adjustment of non-rigid deformation to obtain a soft tissue organ registration image; The TPS deformation model represents non-rigid deformation as affine transformation plus radial basis function deformation, and the formula is as follows: in, for Affine matrix, is the translation vector, is the selected control point position, , is the TPS radial basis kernel function, is the weight to be sought, Represents the spatial position coordinates of any point in the organ model reconstructed by preoperative 3D CT or MRI to be registered; Introducing bending energy To penalize severe deformation; weight the registration error and bending penalty to form the total energy , taking its partial derivative = 0, we can get The closed-form solution of Mapping can finely stretch / shrink the preoperative model to align it to the intraoperative point cloud to minimize local errors.

[0030] Use the Transformer large model to directly learn non-rigid deformation and treat it as a regression problem. After the above matching, the Transformer can output each matching pair , train a neural network to predict from Move to The displacement vector , equivalently, predicts the deformation field make , the preoperative point Mapped to the corresponding position during surgery. By training on a large amount of data with real deformation, the network can automatically learn complex nonlinear deformations without presetting a specific function form. This method fully utilizes the expressive power of large models: for example, multi-head attention can infer how a certain part needs to be displaced to align from the global anatomical structure (even if the part is compressed or stretched during surgery). The distance between the registration points is used as the loss in training to update the network parameters. make Approximate the real deformation. Ultimately, the network learns to predict any new input and can output complex deformation mappings in seconds. In this embodiment, the learned deformation can be combined with the model-based deformation: for example, a preliminary deformation is first given by a neural network, and then the smoothness is fine-tuned using the TPS model, or vice versa, TPS is roughly matched first and then the network is fine-tuned for the residual, so as to take into account both physical priors and data-driven flexibility. After the non-rigid registration stage, the preoperative model (such as a CT point cloud) will accurately fit the soft tissue shape observed during the operation. This rigid + non-rigid two-stage scheme complies with the "rough first, fine later" principle of medical image registration, ensuring that local deformations are corrected while the global position is aligned, such as Figure 6 shown.

[0031] After the above steps, the algorithm finally outputs the spatial transformation after registration. The output format is slightly different for different laparoscopic modes: If it is binocular / 3D point cloud mode, the output non-rigid transformation Will act on the preoperative CT point cloud On the top, we get the point cloud Aligned point clouds (or organ surface meshes). This alignment result can be used for collision detection and distance measurement in surgical navigation, and can overlay preoperative planning information (such as tumor location) to the corresponding position in the intraoperative field of view to assist doctors in locating the target.

[0032] If it is a monocular 2D image mode, the output is mainly the camera extrinsic pose (and possible organ deformation parameters ). Using this camera position, the preoperative 3D model can be projected onto the intraoperative laparoscopic image to achieve augmented reality overlay display. For example, the general outline of the tumor inside the liver, the direction of important blood vessels and other information from CT scans can be displayed in real time on the endoscopic image, so that doctors can navigate the operation without seeing the internal structure.

[0033] Since the surgical process is dynamic, the methods of steps one to four are executed in real time in a loop: each time a new frame of laparoscopic data is acquired, the above-mentioned matching and registration is performed with the preoperative model to obtain a new alignment result, thereby continuously updating the enhanced display interface. If the registration results of adjacent frames do not change much, the results of the previous frame can be used as a priori to integrate into the processing of the current frame (for example, adding the registration pose encoding of the previous frame to the Transformer input, or using the previous frame results in the iterative optimization initial value). This temporal fusion improves the stability and continuity of the registration, allowing the navigation display to transition smoothly and reduce fluctuations, and ultimately obtaining a real-time registered image of soft tissue organs; In order to make the above steps work together, an end-to-end training strategy is adopted to learn the model parameters based on a comprehensive loss function.

[0034] The comprehensive loss function consists of the weighted sum of cross entropy loss, alignment error and regularization constraint: For the cross-modal matching matrix output in step 2, by comparing the real correspondences generated by manual annotation or simulation, a cross entropy loss is applied to increase the attention weight of the correct correspondence , suppressing false correspondences so that the model learns to output accurate matching pairs; For the output rigid body transformation or non-rigid deformation , (known transformations generated from phantom experiments with sensors or computer simulations) can directly use MSE loss to measure the difference between the predicted and true parameters; or apply the predicted transformation to the preoperative image feature vector (point cloud), calculate the alignment error (such as mean square distance or Chamfer distance) with the intraoperative image feature vector (point cloud) as the loss, and back propagate to promote more accurate transformation. Chamfer distance is defined as: ; It provides a measure of the proximity of the overall point set when there is no exact corresponding label, which can be used for unsupervised registration optimization. Further combined with deep features, it makes it possible to calculate Chamfer distance in feature space, encouraging point clouds to align in feature space.

[0035] For non-rigid deformation parameters, add TPS smooth constraints or elastic energy constraints based on physical models , to avoid unreasonable topological deformation. This ensures that the deformation fits the data while maintaining anatomical rationality.

[0036] The final loss function The formula is as follows: , in, Represents large model parameters Predicted non-rigid deformation transformation, is the matching correspondence inferred by the model, is the regularization term, For its weight.

[0037] By minimizing To train the model parameters, the model learns both feature matching and deformation estimation. This training is equivalent to "integrating" the ICP iteration process into network learning. The final model can output the result in one step, greatly improving the reasoning efficiency. More importantly, because the loss incorporates deep feature similarity and prior constraints, the trained model is insensitive to initial misalignment and noise, and the registration results are more robust.

[0038] like Figure 3As shown in the figure, the difference between the traditional ICP algorithm and the large model registration method of the present invention: the left side shows that the traditional ICP method needs to iterate and cycle between the two steps of "nearest point matching" and "calculating rigid body transformation" for many times until the error converges; the right side shows that the method of the present invention uses the Transformer large model to predict the corresponding relationship at one time and directly solve the spatial transformation, so that the registration result can be output without iteration. By eliminating the repeated iteration process, the present invention greatly improves the registration speed, and meets the real-time requirements while ensuring high accuracy.

[0039] The data required for training can come from clinically registered and annotated data, or a large amount of simulated data can be used (such as using simulated laparoscopes to generate registration pairs with known transformations). In addition, the generalization of the model is improved through data enhancement (random deformation, occlusion, illumination changes, etc.) and transfer learning (pre-training on simulated data and fine-tuning on real data). For example, the model is exposed to the position changes of various organs caused by instrument traction and respiratory movement, and learns to match correspondence in the deformation-invariant feature space; real surgical videos are used to add noise training so that the model can still extract effective feature matches in the presence of smoke and blood. Transformer's self-attention gives the model a global vision, and even if there is local occlusion or missing, the corresponding relationship can be inferred through information from other areas. A fully trained model can adapt to new patient anatomical differences, reflecting the powerful learning and generalization capabilities of large models.

[0040] To ensure that the above method can be run in real time during surgery, this embodiment adopts multiple optimization strategies: ① GPU parallel computing: GPU acceleration optimization is performed for time-consuming links such as attention calculation and nearest neighbor search. The process of calculating the attention matrix by dot product is regarded as matrix multiplication, which is calculated in parallel in the CUDA kernel and normalized by Softmax to quickly obtain all At the same time, the nearest neighbor query in the ICP iteration is assigned to the GPU thread for parallel execution, and the shared memory is used to reduce repeated distance calculations; when necessary, a spatial index structure (such as kd tree or Octree) is built on the GPU to further accelerate high-dimensional point matching. Matrix operations in rigid body solutions (such as SVD decomposition) are completed in batches using efficient libraries such as cuBLAS / cuSolver. These measures make full use of the parallel capabilities of the GPU for both deep feature matching and geometric calculations, greatly shortening the processing time of single-frame registration.

[0041] ②Model inference optimization: The trained Transformer model is optimized through tools such as TensorRT. Specifically, it includes: fusing adjacent network layers to reduce memory access overhead (for example, merging the linear transformation and Softmax of multi-head attention into one CUDA kernel for execution); using 16-bit or even 8-bit low-precision numerical inference to improve computing throughput. While keeping the accuracy loss within a controllable range, low-precision calculations can increase the inference speed several times. Combined with GPU parallelism and inference optimization, near real-time registration speed can be achieved on high-end GPUs, making it possible to process multiple frames of data per second.

[0042] ③ In some portable or edge scenarios, it can also be deployed to FPGA or edge TPU for operation. Through high-level synthesis (HLS), the matrix multiplication and other operations of Transformer are implemented as custom circuits on FPGA, which can obtain low-latency and low-power reasoning capabilities. At the same time, the large model is pruned, distilled and quantized to compress the model volume to adapt it to the resource limitations of FPGA / TPU. For example, the main matching network is quantized to 8 bits and mapped to Google Edge TPU for operation, so that each frame feature extraction and matching can be completed within tens of milliseconds. In this way, the key functions of this algorithm can be achieved even on small surgical equipment without GPU. If necessary, a cloud-end collaborative approach can also be adopted: perform lightweight preliminary alignment on the edge device, upload the intermediate results to the cloud / server to use TPU to complete complex deep calculations, and then return the refined results, thereby reducing the edge burden without sacrificing accuracy. In addition, the present invention further reduces power consumption by dynamic frame rate adjustment at the software layer: when the changes in the alignment results of consecutive frames are very small, the forward frequency of the large model is reduced, and only fast tracking is performed, thereby saving computing power, such as Figure 7 In summary, these optimization strategies ensure that the method of the present invention runs smoothly under the clinical real-time requirements, and provide feasibility guarantee for the actual deployment of the system.

[0043] Through the above steps, the laparoscopic soft tissue multimodal matching and registration method provided in this embodiment is fully realized. Its process runs through all links from data acquisition, feature extraction, intelligent matching, to transformation solution, deformation correction, result output and accelerated deployment, and clearly shows how to gradually solve the core problems in soft tissue registration.

[0044] Embodiment 2 This embodiment provides a multi-modal large model-assisted laparoscopic soft tissue registration surgery navigation system. Figure 4 and Figure 5 As shown, it includes a multimodal image acquisition module, an image preprocessing module, a deep learning model fusion and registration calculation module, a calculation acceleration module, and an intraoperative navigation display and control module; The computing acceleration module includes parallel GPU, TPU and FPGA units. The GPU unit uses the CUDA parallel computing framework to optimize convolution and matrix operations, accelerating the calculation of convolutional networks and Transformer attention modules; the TPU unit is used to optimize the structure of deep learning matrix operations, improve the efficiency of large-scale matrix multiplication and tensor operations, and thus accelerate model reasoning; the FPGA unit maps the convolutional neural network and attention mechanism into digital circuits that can be parallelized, realizing concurrent execution of each layer of the model, and further improving the reasoning speed and reducing power consumption by customizing low-precision computing units; The multimodal image acquisition module is used to obtain preoperative CT or MRI images and intraoperative laparoscopic images; The image preprocessing module is used to segment, normalize and extract features of the acquired multimodal images; The deep learning model fusion and registration calculation module includes several layers of Transformer encoders and decoders based on the multi-head attention mechanism, which are used to perform rigid and non-rigid registration on the pre-processed multimodal image data. After several feed-forward networks and normalization, the final matching relationship and deformation parameter estimation are obtained. The intraoperative navigation display and control module stores executable instructions, which, when executed, implement the multimodal large model-assisted laparoscopic soft tissue registration surgical navigation method and display the registration image on the navigation display module.

[0045] The system also includes an augmented reality overlay display module; The augmented reality overlay display module projects real-time multi-organ registration images into the surgical area through an adaptive projection algorithm, and adjusts the projection viewing angle and depth perception in real time.

[0046] Combining GPU parallel computing, model compression and dedicated hardware acceleration strategies, key steps such as Transformer attention calculation and nearest neighbor search are optimized through CUDA parallelism to accelerate the reasoning speed, and the parallel computing capabilities of FPGA / TPU are used to achieve low-power, high-speed edge reasoning, thus ensuring that the algorithm of the present invention can achieve near real-time alignment on high-performance workstations and meet clinical use requirements.

[0047] Although the embodiments disclosed in the present invention are as above, the contents described are only embodiments adopted for facilitating the understanding of the present invention and are not intended to limit the present invention. Any technician in the technical field to which the present invention belongs can make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in the present invention, but the protection scope of the present invention shall still be subject to the scope defined in the attached claims.

Claims

1. A multi-modal large model-assisted laparoscopic soft tissue registration surgical navigation method, characterized in that: The following steps are involved: Step 1: Obtain preoperative CT or MRI images and intraoperative laparoscopic images, and convert the image data into intermediate feature representations that can be processed by the Transformer model to obtain a multimodal feature sequence; Step 2: Input the multimodal feature sequence obtained in step 1 into the fusion layer of the Transformer large model to fuse the features of different modalities and infer the corresponding relationship to obtain a cross-modal matching matrix; Step 3: Using the cross-modal matching matrix obtained in step 2, the RANSAC random sampling consensus algorithm is used for random sampling, and the error between the corresponding point sets of the two data is measured by Euclidean distance. The analytical solution is obtained through SVD decomposition and the feature matching error is introduced for rigid registration to obtain the optimal rigid body transformation posture; Step 4: Use the parameterized thin plate spline TPS deformation model and combine it with Transformer learning to predict deformation to perform local fitting adjustment of the best rigid body transformation pose residual obtained in step 3 to obtain a soft tissue organ registration image; Execute steps 1 to 4 in a loop in real time to obtain a real-time registration image of soft tissue organs; Based on the comprehensive loss function, an end-to-end strategy is used to train the optimal model.

2. The multimodal large model-assisted laparoscopic soft tissue registration surgical navigation method according to claim 1, characterized in that: The data conversion method described in step 1 includes: Perform image segmentation on preoperative CT or MRI to obtain a 3D reconstructed image of the target organ, extract the key anatomical landmarks of the organ and their coordinates in the 3D model, use the network to extract the markers representing the surface / volume features of the preoperative organ, and add spatial location information; For intraoperative laparoscopic images obtained by binocular laparoscopes or systems with depth sensing, a point cloud of the organ surface depicting the three-dimensional shape of the current canvas surface of the organ in the scene is obtained through stereo matching or structured light reconstruction. , using point cloud network to point cloud Calculate local geometric features at each point , then embed the point features into a fixed-length feature vector, and add the three-dimensional coordinate position encoding of the point Or the grid index, forming a mark with location information; Two-dimensional video frames of intraoperative laparoscopic images obtained by monocular laparoscope , select key frame images for registration, and use the depth estimation network to predict the depth map to obtain rough point cloud or depth information as an aid; use convolutional neural network to image Extracting multi-scale visual feature maps , expand the feature map into a series of feature vector identifiers, each of which corresponds to a local area on the image and attaches its two-dimensional position code; Through full connection mapping and normalization processing, the feature vector identifiers of preoperative and intraoperative images have intermediate feature representations with the same dimension, and a multimodal feature sequence is obtained.

3. The multimodal large model-assisted laparoscopic soft tissue registration surgical navigation method according to claim 1, characterized in that: The fusion of different modal features and the inference process of the corresponding relationship described in step 2 include: Transformer performs multi-head self-attention calculations on each modality to enhance local structural features; Adopt Transformer’s multi-head cross-attention mechanism to establish cross-modal associations; The cross-modal matching matrix is ​​obtained through the multi-level fusion iteration of Transformer.

4. The multimodal large model-assisted laparoscopic soft tissue registration surgical navigation method according to claim 1, characterized in that: The process of solving the optimal rigid body transformation posture described in step 3 includes: Preoperative point collection Intraoperative point set For a matching pair, the mean square error of the rigid registration is defined as: , in, is the rotation matrix, is the translation vector, It is a matching relationship. ; Calculate matching point set and The centroid of the data is centered and the covariance matrix is ​​constructed: ,right Perform singular value decomposition ; The optimal rotation matrix is , translated to ; Will Applied to the intraoperative point set, preliminary alignment with the preoperative model was achieved; Introducing the feature mapping learned by Transformer , so that each point Both have high-dimensional descriptions and , define the comprehensive error: , Among them, the parameters Weigh the importance of geometric distance and feature distance by minimizing , so that the spatial distance and feature differences of the corresponding points are reduced, thus obtaining a more accurate .

5. The multimodal large model-assisted laparoscopic soft tissue registration surgical navigation method according to claim 1, characterized in that: The thin plate spline TPS deformation model described in step 4 expresses the non-rigid deformation as an affine transformation plus a radial basis function deformation, and the formula is as follows: in, for Affine matrix, is the translation vector, is the selected control point position, , is the TPS radial basis kernel function, is the weight to be sought, Represents the spatial position coordinates of any point in the organ model reconstructed by preoperative 3D CT or MRI to be registered; Introducing bending energy To penalize severe deformation; weight the registration error and bending penalty to form the total energy , taking its partial derivative = 0, we can get The closed-form solution of The process of Transformer learning to predict deformation described in step 4 includes: Using neural network to predict Move to The displacement vector , equivalently, predicts the deformation field make , the preoperative point Mapped to the corresponding position during surgery.

6. The multimodal large model-assisted laparoscopic soft tissue registration surgical navigation method according to claim 1, characterized in that: The comprehensive loss function is composed of the weighted sum of cross entropy loss, alignment error and regularization constraint: For the cross-modal matching matrix described in step 2, by comparing the real correspondences generated by manual annotation or simulation, a cross entropy loss is applied to increase the attention weight of the correct correspondence , suppressing false correspondences so that the model learns to output accurate matching pairs; For the output rigid body transformation or non-rigid deformation use MSE The loss measures the difference between the predicted and true parameters; or the predicted transformation is applied to the preoperative image feature vector, and the Chamfer distance with the intraoperative laparoscopic image feature vector is calculated as the loss. The Chamfer distance is defined as: ; For non-rigid deformation parameters, add elastic energy constraints based on physical models ; The final loss function The formula is as follows: , in, Represents large model parameters Predicted non-rigid deformation transformation, is the matching correspondence inferred by the model, is the regularization term, For its weight.

7. A multi-modal large model-assisted laparoscopic soft tissue registration surgical navigation system, characterized in that: It includes multimodal image acquisition module, image preprocessing module, deep learning model fusion and registration calculation module, calculation acceleration module, intraoperative navigation display and control module; The computing acceleration module includes parallel GPU, TPU and FPGA units, wherein the GPU unit optimizes convolution and matrix operations using the CUDA parallel computing framework to accelerate the calculation of convolutional networks and Transformer attention modules; the TPU unit is used to optimize the structure of deep learning matrix operations, improve the efficiency of large-scale matrix multiplication and tensor operations, and thus accelerate model reasoning; the FPGA unit maps the convolutional neural network and attention mechanism into digital circuits that can be parallel pipelined, realizes concurrent execution of each layer of the model, and further improves the reasoning speed and reduces power consumption by customizing low-precision computing units; The multimodal image acquisition module is used to obtain preoperative CT or MRI images and intraoperative laparoscopic images; The image preprocessing module is used to perform segmentation, normalization and feature extraction on the collected multimodal images; The deep learning model fusion and registration calculation module includes several layers of Transformer encoders and decoders based on a multi-head attention mechanism, which are used to perform rigid registration and non-rigid registration on the pre-processed multimodal image data, and obtain the final matching relationship and deformation parameter estimation after several feedforward networks and normalization; The intraoperative navigation display and control module, the control module stores executable instructions, and when the executable instructions are executed, the multimodal large model-assisted laparoscopic soft tissue registration surgical navigation method as described in any one of claims 1 to 6 is implemented, and the registration image is displayed on the navigation display module.

8. The multimodal large model-assisted laparoscopic soft tissue registration surgical navigation system as claimed in claim 7, characterized in that: Also included is an augmented reality overlay display module; The augmented reality overlay display module projects the real-time registration images of multiple organs onto the surgical area through an adaptive projection algorithm, and adjusts the viewing angle and depth perception of the projection in real time.

Citation Information

Patent Citations

  • Unsupervised registration method and system for pelvic reduction surgical navigation

    CN119515942A

  • Enhanced method for correcting data for deformations during image guided procedures

    US20140037161A1

  • Registration method for multimodality image-guided radiotherapy

    WO2024169341A1

  • Pet imaging method and apparatus based on optical flow registration, and device and storage medium

    WO2025035380A1

  • Surgical navigation positioning system and method

    WO2025059980A1

Cited By

  • Operation auxiliary system and method based on image processing

    CN120203763A

  • Image fusion and target positioning method, system, device and medium

    CN120563496A

  • An image fusion and target positioning method, system, device and medium

    CN120563496B

  • Virtual reality and reality image fusion navigation method and system for temporal bone surgery

    CN120788730A

  • Anesthesia puncture visual navigation system combining augmented reality and three-dimensional reconstruction

    CN120827435A