Surgical robot real-time image registration method based on multi-modal feature fusion

The method addresses real-time multi-modal image registration challenges by using lightweight feature extractors and cross-modal attention to adapt to varying surgical conditions, improving precision and robustness for surgical robots.

CN120318286AActive Publication Date: 2025-07-15BEIJING JISHUITAN HOSPITAL

Patent Information

Application Number
CN202510804299.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

In real-time feedback surgical robot application scenarios, the existing multimodal surgical image registration technology faces problems such as geometric distortion, image attribute differences, noise and artifacts, resulting in insufficient registration accuracy and robustness. The calculation of traditional methods is time-consuming and difficult to meet clinical needs. Deep learning methods have calculation bottlenecks and insufficient generalization capabilities in complex models.

Method used

The real-time image registration method of surgical robot based on multimodal feature fusion is adopted. Through the lightweight feature extractor dynamic adaptive adjustment, semantic feature conversion to target latent space, cross-modal attention mechanism and uncertainty modulation, image registration is combined with deep neural network to achieve high accuracy, robustness and real-timeness.

Benefits of technology

It significantly improves the accuracy and robustness of multimodal image registration in complex surgical environments, meets the real-time requirements of surgical navigation, reduces dependence on external marking points, and provides more reliable intraoperative guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318286A_ABST
    Figure CN120318286A_ABST
Patent Text Reader

Abstract

The invention discloses a surgical robot real-time image registration method based on multi-modal feature fusion. The method comprises the following steps: acquiring different modal medical images before and during an operation; respectively extracting depth features by using a first lightweight feature extractor and a second lightweight feature extractor, wherein the second extractor performs dynamic adjustment according to the real-time quality evaluation of the intraoperative image; mapping the depth features to the same submerged space to obtain submerged space representation through a semantic feature conversion module comprising a projection network of a pre-training segmentation model and contrast loss training; fusing the submerged space representation by using a cross-modal attention mechanism of uncertainty modulation to obtain a fusion feature; and based on the fusion features, the dense displacement field is regressed through a deep learning network to realize image real-time registration. According to the method, the precision and robustness of multi-modal image registration in a complex operation environment are remarkably improved, the real-time navigation requirement is met, and reliable guidance is provided for surgeons.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image data processing, and particularly to a real-time image registration method for surgical robots based on multimodal feature fusion. Background Art

[0002] The emergence and development of image-guided surgery (IGS) have significantly improved the accuracy and safety of surgeries. Among them, image registration technology plays a crucial role. By precisely aligning medical images at different times (such as pre-operative and intra-operative) or of different modalities (such as CT, MRI, ultrasound, endoscopy), it can provide surgeons with more comprehensive lesion information, clearer anatomical structures, and more reliable surgical navigation, which is of great significance for achieving minimally invasive surgery, reducing complications, and improving patient prognosis. Each of the commonly used imaging modalities in clinical practice has its own advantages. For example, CT can provide high-resolution anatomical structures, MRI has excellent contrast for soft tissues, while ultrasound (US) and endoscopy can provide real-time feedback. Multimodal image registration overcomes the limitations of a single modality by fusing this complementary information.

[0003] However, existing multimodal surgical image registration technologies still face many challenges, especially in the application scenarios of surgical robots that require real-time feedback. First, factors such as inherent geometric distortions between different modality images, differences in image attributes (such as intensity, texture, resolution), as well as noise and artifacts seriously affect the accuracy and robustness of registration. Especially in soft tissue surgeries, problems such as organ deformation and brain tissue drift caused by respiration and instrument operation make precise alignment particularly difficult. Second, the surgical process has extremely high requirements for the real-time performance of the registration algorithm. Traditional optimization-based methods often consume a lot of computing time and are difficult to meet clinical needs. Although deep learning methods have improved in speed, complex models may still face computing bottlenecks, and problems such as their "black box" nature, dependence on large-scale high-quality labeled datasets, and insufficient generalization ability also limit their clinical application and the establishment of doctor trust.

[0004] In addition, how to effectively fuse heterogeneous features from different modalities, extract and align semantic information with clear anatomical significance, rather than simply relying on apparent intensity or texture, is the key to improving the registration quality. In existing technologies, some methods attempt to convert images into modality-neutral representations, but they may perform poorly in regions with subtle texture differences but clear semantic meanings, and do not fully consider the adaptive impact of image quality or uncertainty on the fusion process, which is also a major obstacle to its wide clinical application. Therefore, there is an urgent clinical need for a multimodal image registration solution that can respond to various complex intraoperative situations in real time, accurately, and robustly, and can perform intelligent and interpretable feature fusion. Summary of the Invention

[0005] The present application proposes a real-time image registration method for surgical robots based on multi-modal feature fusion, which is used to solve the problem that it is difficult to balance high registration accuracy, strong robustness and high real-time performance in the complex situation of intraoperative images with variable quality, noise and artifact interference, and soft tissue deformation. In particular, there are deficiencies in the existing methods in effectively fusing the deep semantic information of different modal images to guide registration.

[0006] According to one aspect of the present application, a real-time image registration method for surgical robots based on multi-modal feature fusion is proposed. The method includes: Obtain preoperative medical images of the first modality and intraoperative medical images of the second modality; Extract first deep features from the obtained preoperative medical images using a first lightweight feature extractor, and extract second deep features from the obtained intraoperative medical images using a second lightweight feature extractor, where the second lightweight feature extractor is dynamically and adaptively adjusted according to the real-time image quality evaluation result of the intraoperative medical images; Use a semantic feature conversion module to convert the first deep feature and the second deep feature into the same target latent space, and obtain a first latent space representation and a second latent space representation respectively. The semantic feature conversion module includes a pre-trained segmentation model and a learning-based projection network trained by a contrast loss function. The segmentation model is used to extract high-level anatomical features based on the first deep feature and the second deep feature, and the learning-based projection network is used to map the high-level anatomical features of different modalities into the target latent space; Use a cross-modal attention mechanism to fuse the first latent space representation and the second latent space representation to obtain a fused feature; Estimate registration parameters for aligning the preoperative medical images and the intraoperative medical images based on the fused feature, so as to register the preoperative medical images and the intraoperative medical images in real time.

[0007] In some embodiments, the preoperative medical images of the first modality are computed tomography (CT) images or magnetic resonance imaging (MRI) images, and the intraoperative medical images of the second modality are three-dimensional ultrasound (3D US) images or stereoscopic endoscope video images.

[0008] In some embodiments, the first lightweight feature extractor is a pre-trained three-dimensional convolutional neural network.

[0009] In some embodiments, when the intraoperative medical images of the second modality are three-dimensional ultrasound images, the backbone network of the second lightweight feature extractor adopts a hybrid network architecture including a shallow three-dimensional convolutional neural network and a Transformer encoder layer, and the Transformer encoder layer is applied to the patch embedding of the feature map of the shallow three-dimensional convolutional neural network.

[0010] In some embodiments, when the intraoperative medical image of the second modality is a stereoscopic endoscope video image, the backbone network of the second lightweight feature extractor adopts a compact Vision Transformer (ViT) architecture with reduced number of layers and attention heads.

[0011] In some embodiments, the image quality assessment result is obtained through an auxiliary lightweight convolutional neural network branch in the second lightweight feature extractor. The input of the lightweight convolutional neural network branch is the intraoperative medical image, and the output is the image quality assessment result, which includes some or all of signal-to-noise ratio, contrast, and degree of artifacts.

[0012] In some embodiments, the second lightweight feature extractor performs dynamic adaptive adjustment according to the real-time image quality assessment result of the intraoperative medical image, including: According to the image quality assessment result, adjust the filter parameters of the second lightweight feature extractor through a gating mechanism, or selectively activate or deactivate the corresponding network layers or attention heads of the second lightweight feature extractor.

[0013] In some embodiments, the pre-trained segmentation model is the TotalSegmentator model, and the feature maps after the end of each downsampling stage in its encoder path are used as the high-level anatomical features.

[0014] In some embodiments, the learning-based projection network is a shallow multi-layer perceptron (MLP) or a small autoencoder structure, and the contrast loss function is the InfoNCE loss defined according to the following formula: , where is the representation of the th sample from the first modality in the target latent space, is the representation of the corresponding positive sample from the second modality in the target latent space, and represent negative samples, is the batch size, denotes calculating the and cosine similarity, is the temperature hyperparameter.

[0015] In some embodiments, the cross-modal attention mechanism further includes the following uncertainty modulation: Estimate the feature uncertainty related to the second latent space representation before fusion, and adjust the contribution weights of the first latent space representation and the second latent space representation in the fusion process based on the estimated feature uncertainty.

[0016] In some embodiments, the feature uncertainty includes aleatoric uncertainty and epistemic uncertainty, wherein the aleatoric uncertainty is obtained by modifying the last layer of the second lightweight feature extractor to predict the probability distribution parameters of the second depth features, and the epistemic uncertainty is estimated by applying the Monte Carlo Dropout method to the second lightweight feature extractor or by using the deep ensemble method.

[0017] In some embodiments, the uncertainty modulation is achieved by: Taking the th feature vector in the first latent space representation as the query , taking the th feature vector in the second latent space representation as the key after linear transformation , calculating the original similarity score between the query and the key ; Multiplying the original similarity score by the weight to obtain the adjusted similarity score , , is a hyperparameter greater than zero, is the feature uncertainty corresponding to the th feature vector in the second latent space representation; Using the adjusted similarity score to calculate the attention weight, which is used to perform a weighted sum on the value vector obtained by another linear transformation of the feature vectors in the second latent space representation, so as to obtain the context-related feature vector for the query ; Combining the context-related feature vector for the query and the query through a preset fusion operation to generate the fused feature.

[0018] In some embodiments, the uncertainty modulation is achieved by: Taking the th feature vector in the first latent space representation as the query , taking the th feature vector in the second latent space representation as the key after linear transformation , calculating the attention weight between the query and the key ; According to the The feature uncertainty corresponding to the feature vector , for the th feature vector in the second latent space representation, the value vector obtained by another linear transformation is scaled to obtain an adjusted value vector ; In the attention mechanism, the calculated attention weights are used to perform a weighted sum on the adjusted value vector to obtain a context-related feature vector for the query ; The context-related feature vector for the query is combined with the query through a preset fusion operation to generate a fused feature.

[0019] In some embodiments, based on the fused feature, registration parameters for aligning the preoperative medical image and the intraoperative medical image are estimated, including: Inputting the fused feature into a deep neural network with a U-Net architecture of VoxelMorph-like, and using the deep neural network to regress a dense displacement field as the registration parameter. Among them, the deep neural network is optimized using a composite loss function, which includes a semantic similarity loss and a regularization loss. The semantic similarity loss is obtained by calculating the similarity between the feature obtained by deforming the first latent space representation according to the estimated dense displacement field in the target latent space and the second latent space representation, and the regularization loss is obtained by calculating the norm of the gradient of the estimated dense displacement field or its bending energy term.

[0020] The real-time image registration method for surgical robots based on multi-modal feature fusion proposed in this application can significantly improve the accuracy and robustness of multi-modal image registration in complex surgical environments through dynamic adaptive feature extraction, transformation to a modality-neutral latent space rich in anatomical significance, uncertainty-modulated feature fusion, and semantic-based registration optimization. This method can not only meet the strict requirements of surgical navigation for real-time performance, but also provide more reliable intraoperative guidance for surgeons through deeper semantic understanding and adaptation to changes in data quality, while reducing the dependence on external markers. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.

[0022] Figure 1 Shows a flowchart of a real-time image registration method for surgical robots based on multi-modal feature fusion according to an embodiment of this application.

[0023] Figure 2 shows a block diagram of a real-time image registration system for a surgical robot based on multi-modal feature fusion according to an exemplary embodiment of the present application. Detailed implementation manners

[0024] In order to enable those skilled in the art to better understand and implement the technical solutions of the present application, the following will combine the accompanying drawings and preferred embodiments to elaborate in detail on the real-time image registration method for a surgical robot based on multi-modal feature fusion proposed by the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0025] Refer to Figure 1 , a real-time image registration method for a surgical robot based on multi-modal feature fusion provided by an embodiment of the present application, the method includes the following steps 101 to 105.

[0026] Step 101, obtain a preoperative medical image of the first modality and an intraoperative medical image of the second modality.

[0027] According to this embodiment, first, medical image data from two different sources for registration are obtained.

[0028] The preoperative medical image of the first modality is usually image data with high image quality and resolution obtained before surgery, used for surgical planning and as a reference for registration. In some implementation manners, the preoperative medical image of the first modality may be a computed tomography (CT) image or a magnetic resonance imaging (MRI) image. For example, preoperative CT images (such as 512x512xZ slices, 1mm isotropic resolution) or T1-weighted MRI images containing the target lesion area can be acquired.

[0029] The intraoperative medical image of the second modality is image data obtained in real-time or near real-time during the surgery, used to reflect the current anatomical structure and surgical progress. In some implementation manners, the intraoperative medical image of the second modality may be a three-dimensional ultrasound (3D US) image or a stereoscopic endoscope video image. For example, during the surgery, volume data (such as 200x150x100 voxels, 5-10Hz) can be acquired in real-time through a tracked three-dimensional ultrasound transducer, or continuous image pairs (such as 1920x1080 resolution, 30Hz) can be acquired through a stereoscopic endoscope.

[0030] After obtaining the images of the first modality and the second modality, data synchronization and preprocessing can also be performed as needed. In particular, for the intraoperative medical images of the second modality obtained in the present application (such as three-dimensional ultrasound or stereoscopic endoscope video), their temporal and spatial synchronization with the surgical robot system or an external tracking system (such as an optical tracking system) can be achieved or assisted in the following ways.

[0031] Precise timestamp marking can be performed on the image data stream collected in real time during the operation and the status data stream of the surgical robot / tracking system (such as robot joint angles, end effector positions, or spatial positioning information of the ultrasound probe / endoscope) to achieve time synchronization. Where feasible, a hardware trigger mechanism (for example, triggering an image acquisition once the tracking system positioning is completed) or software methods such as the high-precision Network Time Protocol (NTP) can be used to ensure the strict temporal correspondence of these data from different sources.

[0032] Spatial calibration can be performed on the intraoperative imaging device (such as an ultrasound probe or an endoscope) to determine the relative position relationship between its imaging coordinate system and the physical space or tracking markers. Spatial calibration can also be performed on the surgical robot to ensure the alignment of the images with the surgical instruments or anatomical structures in the same physical space.

[0033] In addition, according to some embodiments of the present application, a bilateral filter can also be used to despeckle the intraoperative three-dimensional ultrasound images to improve the image quality; gamma correction or defogging processing can be performed on the stereoscopic endoscope video images to enhance the image clarity. Intensity normalization (for example, Z-score normalization) and image resampling (for example, unifying to an isotropic voxel spacing) can be performed on all image data (including preoperative images and preprocessed intraoperative images), which helps to improve the stability and accuracy of subsequent feature extraction and registration.

[0034] Step 102, extracting first-depth features from the obtained preoperative medical images using a first lightweight feature extractor, and extracting second-depth features from the obtained intraoperative medical images using a second lightweight feature extractor, wherein the second lightweight feature extractor is dynamically adaptively adjusted according to the real-time image quality evaluation result of the intraoperative medical images.

[0035] In this step, different lightweight feature extractors can be used to extract depth features for images of different modalities, and a dynamic adaptive mechanism is particularly introduced for the feature extraction of intraoperative images.

[0036] In some embodiments, the first lightweight feature extractor can be a pre-trained three-dimensional convolutional neural network (CNN). For example, a 3D variant of a compact ResNet or EfficientNet pre-trained on a large medical image dataset can be used to extract multi-scale first-depth features from preoperative CT or MRI images.

[0037] In some embodiments, in order to better adapt to the real-time nature and variable quality of intraoperative images, a second lightweight feature extractor can be designed specifically according to different modalities of intraoperative medical images.

[0038] When the intraoperative medical image of the second modality is a three-dimensional ultrasound image, the backbone network of the second lightweight feature extractor can adopt a hybrid network architecture including a shallow three-dimensional convolutional neural network and Transformer encoder layers (for example, it can be called a hybrid CNN-Transformer architecture), and the Transformer encoder layers are applied to the patch embeddings of the feature maps of the shallow three-dimensional convolutional neural network. The second lightweight feature extractor designed according to this can efficiently extract local texture and structural features of ultrasound images and better overcome noise and artifact problems.

[0039] In some examples, this hybrid CNN-Transformer architecture can first adopt a shallow (for example, including 3-4 convolutional modules) three-dimensional CNN. This CNN module is responsible for efficiently extracting local three-dimensional textures and basic structural features from the input raw three-dimensional ultrasound volume data, such as the echo pattern of tissues, the cross-section of small blood vessels, etc. Its design can draw on the concept of lightweight CNN, for example, using a smaller convolutional kernel (such as 3x3x3), strided convolution, or depthwise separable convolution to optimize parameters and calculations. In order to capture longer-range spatial dependencies and global context information, which is important for accurately understanding the overall contour and relative position of organs in noisy and blurred ultrasound images, this hybrid network architecture can further apply one or more small Transformer encoder layers to the feature maps output from the CNN module. For example, the three-dimensional feature maps output by the CNN can be segmented into patches, linearly embedded (patch embeddings) into sequence data, and position encoding is added, and then this sequence is input into the Transformer encoder. The self-attention mechanism of the Transformer can calculate the mutual dependencies between different patches in the sequence, thereby capturing the global context. And this hybrid network architecture fully considers the requirements of lightweight characteristics to better meet the real-time requirements during the operation.

[0040] When the intraoperative medical image in the second modality is a stereoscopic endoscope video image, the backbone network of the second lightweight feature extractor preferably adopts a compact Vision Transformer (ViT) architecture with reduced number of layers and number of attention heads. The second lightweight feature extractor designed according to this can efficiently process the features of endoscopic video frames while meeting the real-time requirements of the surgical scenario.

[0041] In some examples, this compact Vision Transformer (ViT) architecture can, through input processing and patch embedding, for each frame of the input stereoscopic endoscope video (or one frame in the left and right views), first segment it into a series of two-dimensional image patches of a fixed size, non-overlapping or partially overlapping, and then linearly flatten these two-dimensional image patches and embed them into a lower-dimensional vector space through a learnable linear projection layer (embedding layer) to form a patch embedding sequence. To retain the spatial position information of the patches, learnable positional encodings are usually added to the patch embedding.

[0042] Then, the sequence processed by patch embedding and positional encoding is input into the compact Transformer encoder. This encoder can be stacked by multiple Transformer layers. To achieve "compactness", compared with the standard ViT model, the number of layers of the Transformer encoder adopted here (for example, using fewer Transformer layers) and the number of attention heads in each multi-head self-attention (MHSA) module can be reduced to significantly reduce the number of model parameters and computational complexity, so as to adapt to the requirements of real-time processing. The MHSA module in each Transformer layer allows the model to learn the dependencies between different regions (patches) of the image in parallel in different representation subspaces, so as to capture the global context information and important features of the image. Even if the number of attention heads is reduced, the compact model according to this embodiment can still effectively extract key information.

[0043] Each Transformer layer usually can also include one or more fully connected feed-forward network layers for performing non-linear transformation on the output of the self-attention mechanism. The final output of the Transformer encoder (for example, the sequence of all patch feature representations, or an aggregated global feature representation) after appropriate post-processing, such as dimension adjustment or format unification to ensure a compatible structure with the features extracted in the case of three-dimensional ultrasound, constitutes the second depth feature.

[0044] This compact ViT architecture optimized for endoscopic videos can effectively capture both global and local visual features within image frames while maintaining low computational overhead, and the features it extracts are of great significance for understanding information such as endoscopic scenes, organ surfaces, and surgical instruments.

[0045] According to the second lightweight feature extractor of the present application, a dynamic adaptive adjustment mechanism is integrated, which can dynamically adjust the behavior of the backbone network according to the real-time image quality evaluation results, so as to extract more robust and effective depth features under various imaging conditions to cope with the real-time changing image quality of intraoperative medical images.

[0046] In some embodiments, the image quality evaluation result is obtained through an auxiliary lightweight convolutional neural network (CNN) branch in the second lightweight feature extractor.

[0047] This auxiliary CNN branch can directly receive the original intraoperative medical images (e.g., current three-dimensional ultrasound volume data frames or stereoscopic endoscopic video frames, or their fast downsampled versions) as input, and its output is the real-time image quality evaluation result, which is usually a vector, and each component of the vector corresponds to the quantization value of different image quality metrics.

[0048] In some examples, the image quality evaluation result may include some or all of the signal-to-noise ratio (SNR), image contrast, degree of artifacts (such as acoustic shadows in ultrasound images, smoke or specular highlights in endoscopic images), etc.

[0049] In some examples, this auxiliary CNN branch itself can be designed to be very lightweight. For example, it can be composed of a few convolutional layers (such as 3x3x3 convolutional kernels, with the stride controlling the downsampling speed), activation functions (such as ReLU), a global average pooling layer, and one or more fully connected layers, and finally output the above-mentioned quality evaluation vector to quickly and accurately evaluate the overall image quality of the current frame.

[0050] In some embodiments, the backbone network of the second lightweight feature extractor (such as a hybrid CNN-Transformer architecture, a compact ViT architecture, etc.) can be dynamically adaptively adjusted according to the real-time image quality evaluation result (such as output by the above-mentioned auxiliary CNN branch). For example, the filter parameters of the second lightweight feature extractor can be adjusted through a gating mechanism, or the corresponding network layers or attention heads of the second lightweight feature extractor can be selectively activated or deactivated.

[0051] For the CNN-based part (e.g., the shallow 3D CNN module in a hybrid CNN-Transformer architecture), the filter weights of the convolutional layer can be dynamically adjusted or different preset filter banks can be selected according to the image quality assessment results (e.g., low signal-to-noise ratio or high artifacts). This can be achieved through a gating unit that receives real-time image quality assessment results and outputs a control signal to modulate the filter behavior.

[0052] For networks containing Transformer layers (e.g., the Transformer encoder module in a hybrid CNN-Transformer architecture, or a compact ViT architecture), certain layers of the Transformer or the attention heads inside it can be dynamically activated or "pruned" according to real-time image quality assessment results. For example, when the image quality is very poor, it can be designed to dynamically reduce the number of active Transformer layers or the number of attention heads among them, or in some cases, if the overall network structure allows, the computational path can be controlled through a gating mechanism to preferentially utilize the shallower or less computationally intensive parts. The gating signal is generated by passing the real-time image quality assessment results through a small gating unit (e.g., one or more fully connected layers followed by a Sigmoid activation function), and the output gating value (ranging from 0 to 1) can be directly multiplied by the output of the corresponding attention head or used to scale the output of the entire network layer.

[0053] Through these adjustments, the second lightweight feature extractor can fully utilize its capabilities to extract fine features when the image quality is good, and reduce the sensitivity to noise and artifacts when the image quality deteriorates, extracting more robust macroscopic structure information, thereby improving the quality and stability of the second depth features as a whole.

[0054] Step 103: Use the semantic feature conversion module to convert the first depth feature and the second depth feature into the same target latent space, obtaining a first latent space representation and a second latent space representation respectively. The semantic feature conversion module includes a pre-trained segmentation model and a learning-based projection network trained with a contrastive loss function. The segmentation model is used to extract high-level anatomical features based on the first depth feature and the second depth feature, and the learning-based projection network is used to map high-level anatomical features of different modalities into the target latent space.

[0055] The purpose of this step is to map depth features from different modalities with different characteristics into a common, more abstract, and anatomically meaningful latent space for subsequent comparison and fusion.

[0056] First, both the first depth feature and the second depth feature extracted in step 102 are processed through a semantic abstraction process. Here, a pre-trained segmentation model can be utilized. In some embodiments, the pre-trained segmentation model can be the TotalSegmentator model, and the feature maps after the end of each downsampling stage in its encoder path are used as the high-level anatomical features. According to this embodiment, high-level information with clear anatomical significance, such as distinguishing different tissue types, organ boundaries, etc., can be further refined from the depth features, rather than just low-level textures or edges.

[0057] Then, a learning-based projection network can be used to map these high-level anatomical (semantic) features from different modalities to the same target latent space. The learning-based projection network can be a shallow multi-layer perceptron (MLP) or a small autoencoder structure. To ensure that the corresponding anatomical structures from different modalities are mapped to similar representations in this latent space, while different structures are mapped to more distant points, the projection network can be trained using a contrastive loss function. In some embodiments, the InfoNCE loss defined by the following formula can be adopted: , where, is the representation of the th sample from the first modality in the target latent space, is the representation of the corresponding positive sample from the second modality in the target latent space, and represent negative samples, is the batch size, denotes calculating and 's cosine similarity, is the temperature hyperparameter, which is usually a preset and tunable parameter.

[0058] Through the above process, the original modality-specific depth features are converted into modality-invariant (or modality-neutral), anatomically semantic-rich latent space representations, namely the first latent space representation and the second latent space representation.

[0059] Step 104, using a cross-modal attention mechanism, fuse the first latent space representation and the second latent space representation to obtain a fused feature.

[0060] In this step, the feature representations from different modalities but already in the same target latent space are effectively fused to combine their respective advantageous information.

[0061] Feature fusion is performed using a cross-modal attention mechanism according to an embodiment of the present application. The cross-modal attention mechanism allows features of one modality (e.g., the latent space representation of preoperative CT) to attend to features of another modality (e.g., the latent space representation of intraoperative ultrasound), and vice versa, thereby learning the correlation between them and identifying complementary information.

[0062] In some embodiments, the cross-modal attention mechanism further includes uncertainty modulation. Specifically, before fusion, the feature uncertainty associated with the second latent space representation (i.e., features derived from intraoperative images) can be estimated, and based on the estimated uncertainty, the contribution weights of the first latent space representation and the second latent space representation in the fusion process are dynamically adjusted, such that the contribution weight of the latent space representation with higher uncertainty is reduced in the fusion process.

[0063] In some embodiments, the feature uncertainty may include aleatoric uncertainty and epistemic uncertainty.

[0064] Aleatoric uncertainty generally reflects the inherent noise, ambiguity, or information incompleteness of the data itself. For example, in intraoperative ultrasound images, due to the limitations of the imaging principle, speckle noise, acoustic shadows, or unclear tissue boundaries may occur, which will all result in the second depth features extracted having aleatoric uncertainty.

[0065] In some embodiments, it can be achieved by modifying the network structure of the second lightweight feature extractor, particularly its last layer (i.e., the output layer), so that it not only predicts the second depth features themselves but also predicts the probability distribution parameters of these features. For example, assuming that each element of the second depth feature follows an independent Gaussian distribution, then the output layer of the second lightweight feature extractor can be designed to predict two values simultaneously for each feature dimension: the mean ( ) of the Gaussian distribution of the feature in that dimension and the variance ( ). The predicted mean (μ) can be regarded as the best estimate of the feature, i.e., the second depth feature itself. The predicted variance (σ²) can directly quantify the uncertainty brought by factors such as the inherent noise of the data for that feature element. The larger the variance, the higher the uncertainty of that feature element.

[0066] If the second depth feature is a feature map (e.g., the output of a CNN), then a mean and variance can be predicted for each channel of each pixel / voxel on the feature map, thus forming an uncertainty map of the same size as the feature map. To enable the network to learn to predict meaningful variances, a specific loss function such as Negative Log-Likelihood (NLL) can be used for training. This loss function can penalize the difference between the predicted mean and the target value, and encourage the model to predict larger variances where the data points are inherently more ambiguous or noisy. In this way, the second lightweight feature extractor can output the second depth feature and its accompanying pixel / voxel-level aleatoric uncertainty map.

[0067] Epistemic uncertainty reflects the uncertainty of the model itself due to reasons such as insufficient training data, imperfect model structure, or incomplete convergence of parameters, and it can represent how confident the model is in its predictions. Epistemic uncertainty is usually higher when the model encounters data not seen in the training set or data with large differences. In some embodiments, epistemic uncertainty can be estimated by applying Monte Carlo Dropout to the second lightweight feature extractor or using a deep ensemble method.

[0068] Monte Carlo Dropout (MC Dropout) is a regularization technique that can be used in neural network training. It prevents overfitting by randomly deactivating a portion of neurons during each forward pass. The core idea of MC Dropout is to keep the Dropout activation during the inference (testing) phase. For example, for the same intraoperative medical image input, multiple (e.g., typically dozens to hundreds of times) forward passes can be performed through the second lightweight feature extractor (with its internal Dropout layer kept on during inference). Since different neurons are randomly deactivated during each pass, multiple slightly different second depth feature outputs will be obtained. The statistical characteristics of these outputs (e.g., calculating their variance or standard deviation in each feature dimension) can be used as a measure of epistemic uncertainty. Similarly, if the output is a feature map, an epistemic uncertainty map can be obtained.

[0069] The Deep Ensembles method is achieved by training multiple (e.g., M) second lightweight feature extractor models with the same structure but different initializations (or trained using different data subsets, different training orders). For example, for the same intraoperative medical image input, it can be inferred through these M independent models respectively to obtain M different second deep feature outputs. The consistency or difference between the outputs of these M models (e.g., calculating their variances in each feature dimension) can be used to quantify the epistemic uncertainty. The greater the difference in model outputs, the higher the epistemic uncertainty.

[0070] In practical applications, those skilled in the art can combine the estimated aleatoric uncertainty map and epistemic uncertainty map as needed (e.g., simple addition or more complex combination methods) to obtain a comprehensive feature uncertainty.

[0071] In some embodiments, the uncertainty modulation can be achieved in the following way: taking the th feature vector in the first latent space representation as the query , taking the th feature vector in the second latent space representation as the key after linear transformation , calculating the original similarity score between the query and the key ; multiplying the original similarity score by the weight , , is a hyperparameter greater than zero, is the feature uncertainty corresponding to the th feature vector in the second latent space representation, to obtain the adjusted similarity score ; using the adjusted similarity score to calculate the attention weight, and then this attention weight is used to perform a weighted sum on the value vectors corresponding to the feature vectors in the second latent space representation, to obtain a context-related feature vector for the query ; combining the context-related feature vector for the query and the query through a preset fusion operation to generate a fused feature.

[0072] The core idea of this embodiment is to adjust the original similarity or correlation between different features according to the feature uncertainty before calculating the final attention weight.

[0073] In the cross-modal attention mechanism, there is a sequence of queries (Q) and a sequence of keys (K). According to this embodiment, the feature vectors in the first latent space representation (derived from preoperative images) can be defined as the query sequence , and the feature vectors in the second latent space representation (derived from intraoperative images) are respectively subjected to different linear transformations to obtain the key sequence and the value sequence .

[0074] To integrate the influence of uncertainty, the original similarity score between a query (one query vector in ) and a key (one key vector in ) can be calculated, for example, through dot product . And using the feature uncertainty corresponding to the key , a modulation weight is generated , where is a hyperparameter greater than zero. The modulation weight is then multiplied by the original similarity score to obtain the adjusted similarity score . Based on the adjusted similarity score, the final attention weights

[0075] can be calculated through the Softmax function . After obtaining the attention weights considering uncertainty , these weights can be used to perform weighted summation on the corresponding value vectors to obtain the context-related feature vector

[0076] for the current query . The context-related feature vector can be regarded as the reliability-weighted information extracted from the second modality for the first modality

[0077] In some other embodiments, the uncertainty modulation is achieved by the following method: taking the -th feature vector in the first latent space representation as the query , and taking the The eigenvectors are used as keys after linear transformation , and the attention weights between them are calculated; according to the feature uncertainty corresponding to the th eigenvector in the second latent space representation , the original value vector obtained by another linear transformation of the th eigenvector in the second latent space representation is scaled to obtain an adjusted value vector ; in the attention mechanism, the adjusted value vector is weighted and summed using the attention weights to obtain a context-related eigenvector for the query ; the context-related eigenvector for the query is combined with the query through a preset fusion operation to generate a fused feature.

[0078] The core idea of this embodiment is to directly adjust the amount of information of intraoperative features in the attention mechanism according to the degree of uncertainty of intraoperative features (i.e., scale their value vectors), so as to reduce the overall contribution of unreliable intraoperative features when fusing with preoperative feature information in the subsequent process.

[0079] Similarly, according to this embodiment, the eigenvectors in the first latent space representation (derived from preoperative images) can be defined as a query sequence , and at the same time, the eigenvectors in the second latent space representation (derived from intraoperative images) are respectively linearly transformed to obtain a key sequence and a value sequence .

[0080] On this basis, the calculation of the attention weights can follow the standard method, that is, through the similarity calculation (such as dot product followed by Softmax normalization) between a query (one of the query vectors in ) and a key (one of the key vectors in ), and this weight reflects the attention degree of the query of the first modality to each part of the second modality.

[0081] Different from the previous embodiment, in this embodiment, the influence of uncertainty mainly acts on the value vector. According to the estimated feature uncertainty , the original value vector is scaled to obtain an adjusted value vector . For example, it can be scaled by , where γ is a scaling factor. When the feature uncertainty is high, the value of is close to 1 approaches 0, thus significantly reducing the modulus of; conversely, when the uncertainty is low, will approach its original value .

[0082] When calculating the final output of the attention mechanism, the previously calculated attention weights act on the value vector adjusted by uncertainty for weighted summation, thereby obtaining the context-related feature vector for the current query , such as . In this way, even if some intraoperative features obtain high attention weights, if their own uncertainty is very high, their actual contribution in the weighted summation will also be limited due to the attenuation of the value vector itself.

[0083] Finally, the context-related feature vector can be combined with its corresponding original query vector through a preset fusion operation, such as concatenating the two and then processing through a linear layer, or performing element-wise addition, or a more complex gating mechanism can also be used to dynamically control their fusion ratio.

[0084] The above two modulation methods can both achieve the purpose of reducing the influence of intraoperative features with high uncertainty in attention calculation, and at the same time ensure that the information of the two modalities is effectively and reliably combined to obtain more robust and reliable fused features.

[0085] Step 105, estimate the registration parameters for aligning the preoperative medical image and the intraoperative medical image based on the fused features, so as to register the preoperative medical image and the intraoperative medical image in real time.

[0086] In this step, the high-quality fused features obtained in the previous step are used to drive the estimation of the registration parameters, and finally the alignment of the images is achieved.

[0087] In some embodiments, first, the fused features output from the previous step can be used as input and fed into a specially designed deep neural network. This deep neural network preferably adopts a U-Net architecture similar to VoxelMorph, which has been proven to be able to effectively learn complex spatial transformation relationships in the field of image registration. After this network is trained, its task is to regress a dense displacement field (DDF), and this DDF constitutes the registration parameters for aligning the two groups of images.

[0088] To ensure that this deep neural network can accurately learn the correct deformation, its training process can be optimized using a composite loss function. This composite loss function can include two main parts: semantic similarity loss and regularization loss.

[0089] The semantic similarity loss is calculated in the target latent space to measure the alignment degree between the deformed preoperative image features and the intraoperative image features. More specifically, it can be obtained by calculating the similarity between the features obtained by deforming the first latent space representation (corresponding to the original preoperative image) according to the dense displacement field predicted by the current network and the second latent space representation (the feature representation corresponding to the original intraoperative image). This similarity can be quantified by calculating the L2 distance or negative cosine similarity between them. By comparing in the semantically rich latent space, more meaningful anatomical structure alignment can be promoted.

[0090] Meanwhile, the regularization loss is used to impose constraints on the predicted dense displacement field for smoothness or physical rationality, which can usually be achieved by calculating the norm of the spatial gradient of the dense displacement field or introducing a bending energy term that simulates physical deformation (such as elastic deformation). The regularization loss helps to prevent the generation of unrealistic and overly distorted deformations, ensuring the smoothness and anatomical rationality of the registration result.

[0091] By minimizing this composite loss function, the deep neural network is driven to learn to generate a dense displacement field that can accurately align the images and has good physical properties, thereby achieving high-precision real-time image registration.

[0092] In some embodiments, the registration process can be performed in a hierarchical manner from coarse to fine. For example, first, the global rigid or affine transformation parameters are estimated, and then more refined non-rigid deformation estimation is performed using features at different scales of the fusion module. This helps to improve the robustness and efficiency of the registration.

[0093] Finally, the obtained registration parameters (such as DDF) are applied to the preoperative medical image to transform it into the coordinate system of the intraoperative medical image, thereby achieving real-time registration of the two.

[0094] In some embodiments, after registration, the deformed preoperative image (or its segmentation model) can be superimposed and displayed with the real-time intraoperative image (such as an endoscopic video), or presented on an augmented reality (AR) head-mounted device to provide enhanced visualization and navigation capabilities for surgeons. The deformation process of the image is usually accelerated on the GPU using efficient interpolation algorithms (such as B-spline interpolation or trilinear interpolation) to meet the real-time requirements.

[0095] The embodiment of the present application provides an end-to-end process that combines dynamic adaptive feature extraction, semantic feature to modal-neutral latent space conversion, uncertainty-modulated adaptive cross-modal attention fusion, and registration parameter estimation based on semantic loss, providing a multi-modal surgical image registration method with high precision, high robustness, and meeting real-time requirements. This method can more effectively utilize the complementary information of different modal images and adapt to the complex intraoperative environment, providing reliable navigation support for surgical robots and having important clinical application prospects.

[0096] To more intuitively understand the system implementation of the surgical robot real-time image registration method based on multi-modal feature fusion proposed in the present application, reference can be made to Figure 2 , which schematically shows the system structure block diagram of an exemplary embodiment of the present application.

[0097] As Figure 2 shown, the system first obtains the first-modal preoperative medical image (such as CT or MRI) and the second-modal intraoperative medical image (such as US or endoscope).

[0098] The first-modal preoperative medical image is fed into the first lightweight feature extractor, which (as described above, can be a pre-trained three-dimensional convolutional neural network) extracts the first depth feature therefrom.

[0099] The second-modal intraoperative medical image goes through a more complex processing path. The intraoperative medical image is fed into the backbone network of the second lightweight feature extractor (as described above, can be a CNN-Transformer hybrid architecture for ultrasound or a compact ViT architecture for endoscope) to extract the second depth feature. At the same time, the second-modal intraoperative medical image is also input into the real-time image quality assessment module (auxiliary CNN branch), which outputs the image quality assessment result. This assessment result is fed back to the backbone network of the second lightweight feature extractor for dynamic adaptive adjustment to ensure that robust second depth features can be extracted under different image qualities.

[0100] Subsequently, the first depth feature and the second depth feature respectively extracted by the first and second lightweight feature extractors are both fed into the semantic feature conversion module. Inside this module, first, a pre-trained segmentation model is used to extract more anatomically significant high-level anatomical features from the input modal anisotropic depth features. Then, the high-level anatomical features are processed by a learning-based projection network (which can be trained by a contrast loss function) and are respectively mapped to the same target latent space to form the first latent space representation and the second latent space representation.

[0101] Next, enter the feature fusion stage. The second latent space representation (derived from intraoperative images) is fed into the uncertainty estimation module, which estimates and outputs the feature uncertainty associated with this latent space representation. Subsequently, the first latent space representation, the second latent space representation, and the estimated feature uncertainty are jointly used as inputs and fed into the cross-modal attention mechanism module. This attention mechanism utilizes the feature uncertainty to modulate the fusion process (in the way of modulating the attention score or value vector as described above), effectively combines the two latent space representations, and finally outputs the fused features.

[0102] These fused features are then input into the registration parameter estimation network (such as the U-Net architecture like VoxelMorph). Based on the input fused features, this network learns and regresses the registration parameters (usually the dense displacement field DDF) for aligning the preoperative and intraoperative images.

[0103] Finally, these estimated registration parameters are fed into the real-time registration module. This module uses these parameters to transform the preoperative medical image into the coordinate system of the intraoperative medical image, thereby outputting the aligned image / deformation field, achieving the real-time registration of the two, and providing support for surgical navigation.

[0104] Through Figure 2 the system structure shown, this application can systematically and efficiently implement the complete process from multi-modal image input to real-time registration output, where each module works collaboratively. In particular, the fusion mechanism of dynamic adaptive feature extraction and uncertainty modulation ensures the accuracy and robustness of registration in a complex surgical environment.

[0105] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0106] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the data processing device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

[0107] Although this specification contains many specific implementation details, these should not be construed as limiting the scope of any invention or the scope of what is claimed, but rather are mainly used to describe the features of specific embodiments of a particular invention. Certain features described in multiple embodiments in this specification can also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may operate in certain combinations as described above and even be initially claimed as such, one or more features from a claimed combination can in some cases be removed from that combination, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination.

[0108] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of the various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0109] The above description is only a preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope protected by one or more embodiments of this specification.

Claims

1. A real-time image registration method for surgical robots based on multi-modal feature fusion, characterized in that The method includes: Obtaining preoperative medical images of a first modality and intraoperative medical images of a second modality; Extracting first depth features from the obtained preoperative medical images using a first lightweight feature extractor, and extracting second depth features from the obtained intraoperative medical images using a second lightweight feature extractor, wherein the second lightweight feature extractor is dynamically adaptively adjusted according to the real-time image quality assessment result of the intraoperative medical images; Using a semantic feature conversion module to convert the first depth features and the second depth features into the same target latent space, respectively obtaining a first latent space representation and a second latent space representation, the semantic feature conversion module includes a pre-trained segmentation model and a learning-based projection network trained by a contrastive loss function, the segmentation model is used to extract high-level anatomical features based on the first depth features and the second depth features, and the learning-based projection network is used to map high-level anatomical features of different modalities into the target latent space; Using a cross-modal attention mechanism to fuse the first latent space representation and the second latent space representation to obtain a fused feature; Estimating registration parameters for aligning the preoperative medical images and the intraoperative medical images based on the fused features to register the preoperative medical images and the intraoperative medical images in real time.

2. The method according to claim 1, wherein The preoperative medical images of the first modality are computed tomography (CT) images or magnetic resonance imaging (MRI) images, and the intraoperative medical images of the second modality are three-dimensional ultrasound (3DUS) images or stereoscopic endoscopic video images.

3. The method according to claim 1, characterized in that The first lightweight feature extractor is a pre-trained three-dimensional convolutional neural network.

4. The method according to claim 2, wherein When the intraoperative medical images of the second modality are three-dimensional ultrasound images, the backbone network of the second lightweight feature extractor adopts a hybrid network architecture including a shallow three-dimensional convolutional neural network and a Transformer encoder layer, and the Transformer encoder layer is applied to the patch embedding of the feature map of the shallow three-dimensional convolutional neural network.

5. The method according to claim 2, wherein When the intraoperative medical images of the second modality are stereoscopic endoscopic video images, the backbone network of the second lightweight feature extractor adopts a compact vision Transformer (ViT) architecture with reduced number of layers and attention heads.

6. The method according to claim 1, characterized in that, The image quality assessment result is obtained through an auxiliary lightweight convolutional neural network branch in the second lightweight feature extractor. The input of the lightweight convolutional neural network branch is the intraoperative medical images, and the output is the image quality assessment result. The image quality assessment result includes some or all of signal-to-noise ratio, contrast, and artifact degree.

7. The method according to claim 1, wherein The second lightweight feature extractor is dynamically adaptively adjusted according to the real-time image quality assessment result of the intraoperative medical images, including: According to the image quality assessment result, adjusting the filter parameters of the second lightweight feature extractor through a gating mechanism, or selectively activating or deactivating the corresponding network layers or attention heads of the second lightweight feature extractor.

8. The method according to claim 1, wherein The pre-trained segmentation model is the TotalSegmentator model, and the feature maps after the end of each downsampling stage in its encoder path are used as the high-level anatomical features.

9. The method according to claim 1, wherein The learning-based projection network is a shallow multi-layer perceptron (MLP) or a small autoencoder structure, and the contrast loss function is the InfoNCE loss defined by the following formula: , Among them, is the representation of the th sample from the first modality in the target latent space, is the representation of the corresponding positive sample from the second modality in the target latent space, and represent negative samples, is the batch size, denotes the calculation of and cosine similarity, is the temperature hyperparameter.

10. The method according to claim 1, wherein The cross-modal attention mechanism further includes the following uncertainty modulation: Estimate the feature uncertainty related to the second latent space representation before fusion, and adjust the contribution weights of the first latent space representation and the second latent space representation during the fusion process based on the estimated feature uncertainty.

11. The method according to claim 10, wherein: The feature uncertainty includes aleatoric uncertainty and epistemic uncertainty. Among them, the aleatoric uncertainty is obtained by modifying the last layer of the second lightweight feature extractor to predict the probability distribution parameters of the second depth feature, and the epistemic uncertainty is estimated by applying the Monte Carlo Dropout method to the second lightweight feature extractor or by using the deep ensemble method.

12. The method according to claim 10, wherein The uncertainty modulation is achieved in the following manner: Use the th eigenvector in the first latent space representation as the query , and use the th eigenvector in the second latent space representation as the key after linear transformation . Calculate the raw similarity score between the query and the key ; Multiply the original similarity score by the weight to obtain the adjusted similarity score , , where is a hyperparameter greater than zero, is the feature uncertainty corresponding to the Using the adjusted similarity score Calculate attention weights, which are used to perform weighted summation on the value vectors obtained by another linear transformation of the feature vectors in the second latent space representation to obtain a context-related feature vector for the query ; The context-related feature vector for the query and the query are combined through a preset fusion operation to generate the fused feature.

13. The method according to claim 10, wherein The uncertainty modulation is achieved in the following manner: Use the th eigenvector in the first latent space representation as the query , and use the th eigenvector in the second latent space representation as the key after linear transformation , calculate the attention weights between the query and the key ; According to the feature uncertainty corresponding to the th eigenvector in the second latent space representation , scale the value vector obtained by another linear transformation of the th eigenvector in the second latent space representation to obtain an adjusted value vector ; In the attention mechanism, the calculated attention weights are used to perform a weighted sum on the adjusted value vectors to obtain a context-related feature vector for the query ; The context-related feature vector for the query is combined with the query through a preset fusion operation to generate a fused feature.

14. The method according to claim 1, characterized in that, Based on the fused features, estimate the registration parameters for aligning the preoperative medical image and the intraoperative medical image, including: Input the fused features into a deep neural network with a U-Net architecture of the VoxelMorph class, and use the deep neural network to regress the dense displacement field as the registration parameter. Among them, the deep neural network is optimized using a composite loss function, which includes a semantic similarity loss and a regularization loss. The semantic similarity loss is obtained by calculating the similarity between the features obtained by deforming the first latent space representation according to the estimated dense displacement field in the target latent space and the second latent space representation, and the regularization loss is obtained by calculating the norm of the gradient of the estimated dense displacement field or its bending energy term.

Citation Information

Patent Citations

  • Workpiece surface topography generation method and device based on multi-modal image generation

    CN116977652A

  • Method and device for repairing degraded image of under-screen camera, computer equipment and storage medium

    CN118261826A

  • Cross-modal medical image registration method based on submerged space diffusion model

    CN119494863A

  • Intelligent pulmonary nodule grading method and system based on multi-modality feature fusion

    WO2025020719A1

  • Image processing method and apparatus, and computer device, storage medium and program product

    WO2025044485A1

Cited By

  • Surgical robot dynamic compensation method based on multi-mode real-time 4D digital twinning and related device

    CN120899400A

  • Operation quality evaluation method and system based on deep learning

    CN122290889A

  • A Deep Learning-Based Surgical Quality Assessment Method and System

    CN122290889B