Unsupervised medical image registration method based on Mamba-CNN (Convolutional Neural Network) coding feature fusion

Through the Mamba-CNN encoding feature fusion method, the problems of incomplete feature extraction and high computational complexity in existing medical image registration are solved, and high-precision and lightweight medical image registration are achieved, which is suitable for single-modal and multimodal image registration.

CN120259385APending Publication Date: 2025-07-04CHENGDU UNIV OF INFORMATION TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510403883.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing medical image registration methods have incomplete feature extraction, high computational complexity, and difficult to balance registration accuracy and lightweighting, especially in large-scale and complex 3D image processing, which have problems with computing efficiency bottlenecks and model scalability.

Method used

Unsupervised medical image registration method based on Mamba-CNN encoding feature fusion is adopted. Through the parallel structure of Mamba branch encoder and CNN branch encoder, combined with the attention feature fusion module, global and local features are extracted and fused respectively to reduce the computational complexity and improve the registration accuracy.

Benefits of technology

It realizes the simultaneously capture of global and local features in high-resolution medical image registration, reduces computing overhead, improves registration accuracy, and has good lightweighting and wide adaptability. It is suitable for single-modal and multimodal medical image registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259385A_ABST
    Figure CN120259385A_ABST
Patent Text Reader

Abstract

The invention discloses an unsupervised medical image registration method based on Mamb-CNN coding feature fusion, and belongs to the technical field of medical image registration, and the method comprises the steps: obtaining a spliced image; respectively inputting the spliced image into a Mamba branch encoder and a CNN branch encoder, and respectively outputting a global feature map and a local feature map of the spliced image; inputting the global feature map and the local feature map of the spliced image into an attention feature fusion module, and outputting a weighted fusion feature map; and inputting the weighted fusion feature map into a decoder, and outputting a deformation field for aligning the moving image and the fixed image. According to the method, the global features and the local features of the medical image can be captured at the same time, registration precision and lightweight unsupervised medical image registration are taken into consideration, and the perception ability of the model to global and local information is enhanced, so that the registration precision is improved, the model parameter quantity and computing resource consumption are greatly reduced, and a better lightweight effect is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image registration, and particularly relates to an unsupervised medical image registration method based on Mamba-CNN encoded feature fusion. Background Art

[0002] Medical image registration is a key technology in medical image processing, mainly used for image alignment under different times, devices or imaging modes, so as to better perform tasks such as disease diagnosis and surgical navigation. Traditional image registration methods mostly rely on classical algorithms, such as the demons algorithm, LDDMM, etc. These methods have a solid foundation in mathematical theory. However, with the increase in the scale of medical image data, these algorithms have bottlenecks in computational efficiency and are difficult to handle large-scale and complex 3D images. More importantly, such methods usually rely on iterative optimization, with large computational overhead and are not easily extended to complex clinical application scenarios. In addition, it is very difficult to obtain the gold standard data of medical images. These problems limit their efficiency and wide application.

[0003] To overcome these limitations, in recent years, unsupervised registration methods based on deep learning have developed rapidly, such as VoxelMorph and its derivative models. These models predict the image deformation field through a convolutional neural network (CNN) and are optimized through the similarity loss function between images, avoiding the dependence on gold standard data. However, such models usually only rely on the CNN to extract local features and lack the capture of global features. To alleviate this problem, LKU-Net increases its receptive field and multi-scale feature extraction ability by integrating parallel multi-scale convolutional blocks, but still cannot capture global features.

[0004] Models such as TransMorph, TransMatch, and ModeT replace traditional convolutional blocks with Transformer blocks in the encoder to capture global features and improve registration accuracy. Based on TransMorph, MFCTrans further enhances registration accuracy by using an attention mechanism at skip connections to fuse features at different levels and introducing residual connections. However, the time complexity of Transformer, which is positively related to the image size when processing images, poses a challenge. Especially in dense prediction tasks such as medical image registration, the high computational overhead has a significant impact, which requires the model to be lightweight. Therefore, existing technologies improve registration accuracy by introducing a linear wavelet self-attention module in the model to capture rich high-frequency information, and reduce the computational overhead to a certain extent through its unique linear calculation method. Existing technologies alleviate the computational complexity problem through a multi-window MLP module and can capture global features at full resolution. In addition, Transformer cannot extract local features, such as the capture of edge and corner features, resulting in insufficient registration accuracy on small organs. Existing technologies attempt to capture both types of features by directly cascading a Transformer block after the CNN block, but since the features processed by CNN are already highly localized, this affects the ability of Transformer to capture long-range dependencies. Summary of the Invention

[0005] An object of the present invention is to address the above deficiencies in the prior art and provide an unsupervised medical image registration method based on Mamba-CNN encoded feature fusion to solve the problems of incomplete feature extraction, high computational complexity, and difficulty in balancing registration accuracy and lightweight in existing medical image registration.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] An unsupervised medical image registration method based on Mamba-CNN encoded feature fusion, which includes the following steps:

[0008] S1. Obtain a moving image and a fixed image, and perform splicing processing on the moving image and the fixed image to obtain a spliced image;

[0009] S2. Input the spliced image into the Mamba branch encoder and the CNN branch encoder in the unsupervised medical image registration model respectively, and output the global feature map and the local feature map of the spliced image;

[0010] S3. Input the global feature map and the local feature map of the spliced image into the attention feature fusion module in the unsupervised medical image registration model, and output a weighted fusion feature map;

[0011] S4. Input the weighted fusion feature map into the decoder in the unsupervised medical image registration model, and output the deformation field for aligning the moving image and the fixed image.

[0012] Further, in S1, the splicing process of the moving image and the fixed image includes:

[0013] x in = Concat(Moving, Fixed)

[0014] where x in is the spliced image; Concat() is the splicing operation; Moving is the moving image; Fixed is the fixed image.

[0015] Further, in S2, the Mamba branch encoder includes three levels of Mamba encoders, and the outputs of the three levels of Mamba encoders are:

[0016]

[0017] where are the feature maps output after being processed by the Mamba encoders of the 1st, 2nd, and 3rd layers respectively; ML1, ML2, and ML3 are the Mamba encoders of the 1st, 2nd, and 3rd layers respectively; PE1, PE2, and PE3 are the patch embedding layers of the 1st, 2nd, and 3rd layers respectively; C1 is the number of channels output by the 1st layer Mamba encoder; D, H, and W are the height, width, and length of the feature map respectively.

[0018] Further, the spliced image enters the Mamba encoder, and the spliced image is serialized based on the patch embedding layer, and sine positional embeddings are added to the serialized data to obtain the first feature tensor; the first feature tensor is respectively input into two data streams in the Mamba encoder;

[0019] In the first data stream, the first feature tensor undergoes a linear transformation through a linear layer to adjust the feature dimension; the linearly transformed first feature tensor is passed to a one-dimensional convolutional layer to extract the local feature tensor; the local feature tensor undergoes a non-linear mapping through an activation function layer, and finally enters the SSM layer to capture the global feature and output the processed global feature tensor;

[0020] In the second data stream, the first feature tensor undergoes a linear transformation through a linear layer to adjust the feature dimension; the linearly transformed first feature tensor enters the activation function layer and undergoes a non-linear mapping to enhance the feature expression; the enhanced first feature tensor is input into a gating mechanism, and the gating mechanism receives the first feature tensor and a control signal, and selectively retains or suppresses some information in the first feature tensor through element-wise operations, and finally outputs the screened feature tensor;

[0021] Element - wise multiply the global feature tensor output from the first data stream and the filtered feature tensor output from the second data stream to obtain a fused feature tensor; project the fused feature tensor linearly through a linear layer to output a first multi - dimensional feature tensor.

[0022] Further, the zero - order hold method is used to discretize the linear ordinary differential equation in the SSM layer:

[0023]

[0024] where A is the state - transition matrix, A = e ΔA ; B and C are projection parameter matrices, B=(ΔA) -1 (e ΔA - I)ΔB, C = C; Δ is the time - scale parameter, I is the identity matrix; t is the index of the token; h t is the state at time t; h t-1 is the state at time t - 1; x t is the input feature vector at time t; y t is the prediction or processing result at time t;

[0025] According to the discretized linear ordinary differential equation, obtain the serialized representation structure of the Mamba model:

[0026]

[0027] where y is the output sequence; x is the input sequence; * is the convolution operation; is the structured convolution kernel; L is the length of the sequence in the model.

[0028] Further, in S2, the spliced image is subjected to two pre - convolution and down - sampling processes to obtain an initial input feature map x with a resolution of cnn_in , and the initial input feature map x cnn_in is input into the CNN - branch encoder. The CNN - branch encoder includes three - level CNN encoders, and the outputs of the three - level CNN encoders are:

[0029]

[0030] where start_channel is a hyperparameter; CL1, CL2, and CL3 are the CNN encoders of the first, second, and third layers respectively; are the feature maps output after being processed by the CNN encoders of the first, second, and third layers respectively; C2 is the number of channels output by the first - layer CNN encoder.

[0031] Further, in S3, the attention feature fusion module includes a channel attention sub-module and a spatial attention sub-module;

[0032] The feature map output by each layer of the Mamba encoder in the Mamba branch encoder and the feature map output by each layer of the CNN encoder in the CNN branch encoder The concatenated feature map X is used as the input to the channel attention sub-module and the spatial attention sub-module, where i = 1, 2, or 3;

[0033] The outputs of the channel attention sub-module and the spatial attention sub-module are the channel-weighted feature map F i CA and the spatial-weighted feature map F i SA ;

[0034] The channel-weighted feature map F i CA and the spatial-weighted feature map F i SA After fusion, a weighted fusion feature map is obtained:

[0035] out = (1 + Sigmoid(F i CA + F i SA )) ⊙ X

[0036] where out is the weighted fusion feature map; Sigmoid is the activation function; ⊙ is the Hadamard multiplication.

[0037] Further, the channel attention sub-module is an ECA attention module;

[0038] The feature F of each layer of the Mamba encoder in the Mamba branch encoder i Mamba and the feature F of each layer of the CNN encoder in the CNN branch encoder i CNN The concatenated feature map is input into the ECA attention module;

[0039] The ECA attention module performs global average pooling on the input feature map to extract the global information Y of each channel, passes the global information Y to the parameterized one-dimensional convolutional layer, outputs the intermediate information Y', and after the intermediate information Y' is processed by the Sigmoid activation function, it is converted into a channel weight map, and the channel weight map is multiplied by the original input feature map X using the Hadamard multiplication operation to generate the channel-weighted feature map:

[0040] F i CA= F i CA_weight ⊙X

[0041] Among them, F i CA_weight is the channel weight map.

[0042] Furthermore, the feature F of each layer of the Mamba encoder in the Mamba branch encoder i Mamba and the feature F of each layer of the CNN encoder in the CNN branch encoder i CNN After splicing, the feature map is input into the spatial attention sub-module;

[0043] The spatial attention sub-module calculates the average value and the maximum value of the input feature map in each channel, respectively obtaining the mean output and the maximum value output. Stack the mean output and the maximum value output to form a two-channel feature map. Process this two-channel feature map through a 3D convolutional layer with a convolutional kernel size of 3 or 7 to extract spatial information and regenerate a single-channel feature map. The single-channel feature map generates a spatial weight map through the Sigmoid activation function; Use the Hadamard multiplication operation to multiply the spatial weight map by the original input feature map X to generate the spatially weighted feature map F i SA :

[0044] F i SA = F i SA_weight ⊙X

[0045] Among them, F i SA_weight is the spatial weight map.

[0046] Furthermore, the loss function of the unsupervised medical image registration model is:

[0047]

[0048] Among them, is the normalized cross-correlation coefficient; is the smoothness regularization penalty term; is the weighted total loss function; NCC is an index for image matching and registration; Moving i is the i-th moving image; is the i-th fixed image; is the warping operator that applies the deformation field to the moving graphic; φ i is the estimated deformation field; Θ is the model parameter; is the first-order gradient calculation smoothness obtained by finite difference; N is the number of image pairs for training; λ is a hyperparameter; v iis the gradient value of the deformation field; is the square of the second norm.

[0049] The unsupervised medical image registration method based on Mamba-CNN encoded feature fusion provided by the present invention has the following beneficial effects:

[0050] 1. The present invention can simultaneously capture the global features and local features of medical images, and balance the registration accuracy and lightweight unsupervised medical image registration. The present invention overcomes the limitation of existing methods that cannot extract global and local features simultaneously through a parallel Mamba branch encoder and CNN branch encoder structure. Among them, the Mamba encoder is responsible for extracting global features, and the CNN encoder is responsible for extracting local features. The two are effectively fused through an attention feature fusion module (AFFM), enhancing the model's perception ability of global and local information, thereby improving the registration accuracy.

[0051] 2. The linear time complexity of the Mamba encoder of the present invention significantly reduces the computational overhead. Especially in the high-resolution medical image registration task, its computational complexity is much lower than that of the traditional Transformer architecture. Through experimental verification, while maintaining the accuracy, this model significantly reduces the number of model parameters and computational resource consumption, achieving a good lightweight effect.

[0052] 3. The present invention realizes the effective fusion of the features extracted by the Mamba and CNN encoders by using the attention feature fusion module (AFFM), especially improving the effectiveness of feature expression in both the spatial and channel dimensions. This design effectively solves the fusion problem of existing models when dealing with multi-scale and multi-resolution features, making the model perform more stably in the registration task and the registration result more accurate.

[0053] 4. The two-stream parallel encoder architecture of the present invention has high adaptability. It can not only be applied to the registration of single-modal medical images, but also be extended to the registration task of multi-modal images. The present invention has good scalability and can further improve the performance of specific tasks by replacing or expanding the encoder module. In addition, the flexibility of the model of the present invention makes it applicable to different types of medical images and imaging devices, and has a wide range of application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is the flowchart of the unsupervised medical image registration method based on Mamba-CNN encoded feature fusion of the present invention.

[0055] Figure 2 is the architecture diagram of the unsupervised medical image registration model of the present invention.

[0056] Figure 3This is the structural diagram of the Mamba encoder of the present invention.

[0057] Figure 4 This is the structural diagram of the attention feature fusion module of the present invention. Specific embodiments

[0058] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

[0059] Example 1

[0060] This embodiment provides an unsupervised medical image registration method based on Mamba-CNN encoding feature fusion, which can overcome the defects of existing medical image registration methods in feature extraction, computational complexity, and model lightweighting, and provide an efficient, accurate, and lightweight registration solution for clinical medical image analysis. Refer to Figure 1 , and it specifically includes the following contents:

[0061] Step S1: Obtain a moving image and a fixed image, and perform stitching processing on the moving image and the fixed image to obtain a stitched image;

[0062] The stitching processing of the moving image and the fixed image in this embodiment includes:

[0063] x in = Concat(Moving, Fixed)

[0064] where x in is the stitched image; Concat() is the stitching operation; Moving is the moving image; Fixed is the fixed image.

[0065] Step S2: Input the stitched image into the Mamba branch encoder and the CNN branch encoder in the unsupervised medical image registration model respectively, and output the global feature map and the local feature map of the stitched image;

[0066] The unsupervised medical image registration model of this embodiment, as Figure 2As shown in the figure, it includes a parallel dual-encoder backbone network composed of a Mamba branch encoder and a CNN branch encoder. Among them, the Mamba branch encoder utilizes advanced sequence processing techniques to effectively extract global features and capture long-range dependencies. The CNN branch encoder, on the other hand, accurately captures local detail features through convolution operations. By combining these two encoders in a parallel network, the unsupervised medical image registration model can simultaneously extract global and local features, thereby achieving higher registration accuracy. This parallel dual-encoder design not only overcomes the limitations of a single encoder model but also enables the model to have higher robustness and adaptability when processing complex 3D medical images.

[0067] Specifically, referring to Figure 2 , the Mamba branch encoder in this embodiment includes three levels of Mamba encoders. The input of the first layer of the Mamba encoder is x obtained by splicing the moving image and the fixed image in , and its resolution is 2×D×H×W, where 2 is the number of channels.

[0068] In this embodiment, to ensure a fair comparison with other models, the hyperparameters of the Mamba encoder are set the same as those of TransMorph. Based on this, the outputs of the three levels of Mamba encoders in this embodiment are:

[0069]

[0070] Among them, are the feature maps output after being processed by the Mamba encoders of the first, second, and third layers respectively; ML1, ML2, and ML3 are the Mamba encoders of the first, second, and third layers respectively; PE1, PE2, and PE3 are the patch embedding layers of the first, second, and third layers respectively; C1 is the number of channels output by the first-layer Mamba encoder; D, H, and W are the height, width, and length of the feature map respectively.

[0071] Referring to Figure 3 , the Mamba encoder in this embodiment is divided into two data streams;

[0072] After the spliced image enters the Mamba encoder, it needs to be serialized; as can be seen from Figure 3 , the serialization is achieved through the patch embedding layer. Specifically, the input spliced image is first divided into several fixed-size image blocks (patches). Specifically, the spliced image is serialized according to a patch size of 4X4. Each patch is unfolded into a one-dimensional vector and projected into a low-dimensional representation vector through the patch embedding layer. This serialization process ensures that the local information of the image is effectively encoded during the feature extraction stage and provides a unified input format for subsequent processing.

[0073] Subsequently, sine position embeddings are added to the serialized features to explicitly encode the spatial position information of each patch; this process adds the position information vector generated by the sine function element-wise to the patch embedding features to preserve the global spatial structure of the input image in the feature representation, thereby obtaining the first feature tensor.

[0074] The Mamba encoder is divided into two data streams;

[0075] In the first data stream, the first feature is linearly projected through a multi-layer perceptron, and after linear projection, it is sent to a one-dimensional convolutional layer. After passing through the activation layer, the first feature tensor enters the SSM layer for processing; specifically, the linear layer of the first data stream receives the first feature tensor and performs a linear transformation on the first feature tensor to adjust the feature dimension; the linearly transformed first feature tensor is passed to the one-dimensional convolutional layer to extract local feature patterns; subsequently, the first feature tensor is non-linearly mapped through an activation function layer (such as ReLU or Sigmoid) to enhance the feature expression ability; finally, the activated tensor enters the SSM layer, and its structure is used to capture global features and output the processed global feature tensor.

[0076] In the second data stream, a gating mechanism is introduced to selectively propagate or forget data; specifically, this data stream first linearly projects the input features through a linear layer to adjust the feature dimension; the projected feature tensor enters the activation layer and is non-linearly mapped to enhance the feature expression ability; subsequently, the feature tensor is input into the gating mechanism, which receives the input features and control signals and selectively retains or suppresses some information in the feature tensor through element-wise operations, and finally outputs the screened feature tensor.

[0077] Finally, the output tensors of the two data streams are element-wise multiplied for feature fusion; specifically, the global feature tensor generated after the first data stream is processed by the SSM layer is element-wise multiplied by the feature tensor screened by the gating mechanism of the second data stream to achieve deep fusion at the feature level and obtain the fused feature tensor; the fused feature tensor is further projected through a linear layer to adjust the dimension and expression of the features, and finally outputs the first multi-dimensional feature tensor. Here, the first multi-dimensional feature tensor outputs the feature map through patch merging

[0078] Among them, SSM is a variant of the Recurrent Neural Network (RNN) and is widely used in fields such as signal processing, automatic control, and economics. SSM processes one-dimensional input sequences through intermediate hidden states The mapping of From a mathematical perspective, SSM is usually expressed as a linear ordinary differential equation:

[0079]

[0080] Among them, h′(t) is the first derivative of the state at time t; x(t) is the input feature vector at time t; is the state transition matrix; and are the projection parameter matrices.

[0081] Given the continuous-time characteristics of the SSM, there are significant challenges in directly applying it in discrete environments such as deep learning. To integrate the model into the deep learning framework, the above ODE needs to be converted into a discrete function. Ensuring the alignment of the model with the sampling rate of the underlying signal in the input data is crucial for achieving efficient computation. Specifically, considering the input i.e., a vector of length L in the signal flow, by introducing the time-scale parameter Δ, the continuous matrices A and B are discretized using the Zero-Order Hold (ZOH) method:

[0082]

[0083] Among them, A is the state transition matrix, A = e ΔA ; B and C are the projection parameter matrices, B = (ΔA) -1 (e ΔA - I)ΔB, C = C; Δ is the time-scale parameter, I is the identity matrix; t is the index of the token; h t is the state of the system at time t; h t-1 is the state of the system at time t - 1; x t is the input feature vector at time t; y t is the prediction or processing result at time t;

[0084] According to the discretized linear ordinary differential equation, the serialized representation structure of the Mamba model is obtained:

[0085]

[0086] Among them, y is the output sequence; x is the input sequence; * is the convolution operation; is the structured convolution kernel; L is the length of the sequence in the model.

[0087] In this embodiment, the Mamba encoder designs a simple selection mechanism that continuously adjusts the SSM parameters Δ, as well as matrices A, B, and C according to the input x to build the correlation between the parameters and the input, similar to the gates in the Long Short-Term Memory (LSTM). Thanks to the input correlation mechanism and various hardware-aware technologies, Mamba can perform remote modeling with fewer computing resources compared to Transformer.

[0088] In subsequent experiments, in this embodiment, by serializing a single 3D data sample with a resolution of 160×192×224 using patches of size 4, approximately 110,000 tokens can be obtained. In this case, the linear computational complexity of Mamba can significantly reduce the computational pressure compared to the O(n 2 ) complexity of Transformer, and at the same time, it can also well maintain the association ability of long sequences, which makes Mamba a powerful challenger to Transformer. Therefore, by introducing the Mamba module into the registration module, the performance is effectively improved and the computational cost is reduced.

[0089] For the CNN branch encoder, in order to maintain light weight, each CNN encoder block consists of only two pairs of sequentially arranged basic convolutional blocks and PReLU activation functions, where the convolutional kernel size of the convolutional block is 3. To facilitate subsequent multi-scale feature fusion, the feature size output by each layer of the CNN encoder needs to be the same as that of the Mamba encoder in the same layer. Therefore, it is necessary to first perform two pre-convolutions and downsamplings on x in to obtain the initial input feature map x with a resolution of cnn_in . The initial input feature map x cnn_in is input into the CNN branch encoder. The CNN branch encoder includes three levels of CNN encoders, and the outputs of the three levels of CNN encoders are:

[0090]

[0091] where start_channel is a hyperparameter, and the default setting is 8; CL1, CL2, and CL3 are the CNN encoders of the first, second, and third layers respectively; are the feature maps output after being processed by the CNN encoders of the first, second, and third layers respectively; C2 is the number of channels output by the first-layer CNN encoder, that is, 4×start_channel.

[0092] Step S3: Input the global feature map and local feature map of the spliced image into the attention feature fusion module in the unsupervised medical image registration model, and output a weighted fusion feature map;

[0093] In the traditional U-Net structure, features are generally directly passed through a concatenation operation after encoding. However, in the unsupervised medical image registration model proposed in this embodiment, directly concatenating two types of features may lead to conflicts, especially when different types of features have different distributions and scales. This can make it difficult for the model to effectively utilize these features during training, thereby affecting performance. In addition, the concatenation operation does not screen and optimize features, which may result in the existence of a large amount of redundant information, increasing the computational complexity and potentially affecting the model learning efficiency.

[0094] Based on this, in order to effectively integrate the features extracted by the CNN and Mamba branches, an attention feature fusion module AFFM is designed, and the structure is as Figure 4 shown. This module includes a channel attention sub-module and a spatial attention sub-module;

[0095] The feature map output by each layer of the Mamba encoder in the Mamba branch encoder and the feature map output by each layer of the CNN encoder in the CNN branch encoder

[0096] After being concatenated, the resulting feature map X is used as the input to the channel attention sub-module and the spatial attention sub-module, where i = 1, 2, or 3.

[0097] For the channel attention sub-module;

[0098] The channel attention sub-module adopts the ECA attention module, which consists of global average pooling, one-dimensional convolution, and a Sigmoid activation function. i CA_weight During specific operations, the ECA attention module first performs global average pooling on the input feature map to extract the global information Y of each channel. Subsequently, the global information Y is passed to a parameterized one-dimensional convolutional layer, which is responsible for implementing local cross-channel interactions, and the output intermediate information Y′ is obtained. The size of the convolutional kernel determines the range of cross-channel interactions. This process does not require dimensionality reduction to maintain the richness of features. Then, after being processed by the Sigmoid activation function, the intermediate information Y′ is converted into a channel weight map F i CA _weight This weight map represents the importance of each channel. Finally, through the Hadamard multiplication operation, F i CA :

[0099] F i CA = Fi CA_weight ⊙X

[0100] For the spatial attention module;

[0101] This module first calculates the average value and the maximum value of each channel of the input feature map to obtain the mean output and the maximum output respectively, and then stacks the mean output and the maximum output to form a two-channel feature map. This feature map is processed by a 3D convolutional layer with a convolutional kernel size of 3 or 7 to extract spatial information and regenerate a single-channel feature map, and then the final spatial weight map F is generated through the Sigmoid activation function i SA_weight , and multiplies the spatial weight map with the input feature map through the Hadamard multiplication operation to generate the final spatially weighted feature map F i SA :

[0102] F i SA = F i SA_weight ⊙X

[0103] Finally, the channel-weighted feature map F i CA and the spatially weighted feature map F i SA are fused to obtain the weighted fusion feature map:

[0104] out = (1 + Sigmoid(F i CA + F i SA )) ⊙X

[0105] where out is the weighted fusion feature map; Sigmoid is the activation function; ⊙ is the Hadamard multiplication.

[0106] The attention feature fusion module AFFM of this embodiment activates the channels that contribute to feature registration through channel attention, suppresses irrelevant or redundant channels, and then adjusts the distribution of features in the image space through spatial attention, finally realizing the effective enhancement and fusion of the Mamba branch features and the CNN branch features, finely controlling the information flow, enhancing the attention to the region to be registered, reducing the attention to the irrelevant regions, providing rich multi-dimensional features for the decoder, and improving the model learning efficiency.

[0107] Step S4, input the weighted fusion feature map into the decoder in the unsupervised medical image registration model, and output the deformation field for aligning the moving image and the fixed image.

[0108] The present invention adopts an encoder-decoder structure in the style of U-Net, and splices the features output by the encoder with the decoder through skip connections to simultaneously fuse low-level detail information and high-level context information, ensuring the accuracy of registration. Refer to Figure 2 , in the decoding stage, the feature map is restored through a series of upsampling operations and spliced with the corresponding feature map output by the encoder through skip connections to generate the final deformation field φ for aligning the input fixed image and moving image.

[0109] The loss function of this embodiment is constructed as follows:

[0110] Variable image registration optimizes an energy function to establish a spatial correspondence between two images, mainly composed of the normalized cross-correlation coefficient measuring the similarity between two images and the smooth regularization penalty term . NCC is an index commonly used in image matching and registration to measure the similarity between two images in a specific region. The smooth regularization penalty term constrains the gradient of the deformation field to ensure the smoothness of the deformation field, thereby avoiding excessive distortion during the image registration process.

[0111] Based on this, the loss function of this embodiment is:

[0112]

[0113] Among them, is the normalized cross-correlation coefficient; is the smooth regularization penalty term; is the weighted total loss function; NCC is an index for image matching and registration; Moving i is the i-th moving image; is the i-th fixed image; is the warping operator that applies the deformation field to the moving graphic; φ i is the estimated deformation field; Θ is the model parameter; is the first-order gradient calculation smoothness obtained by finite difference; N is the number of image pairs for training; λ is a hyperparameter; v i is the gradient value of the deformation field; is the square of the two-norm.

[0114] Although the specific implementation manners of the invention have been described in detail with reference to the accompanying drawings, it should not be construed as a limitation on the protection scope of this patent. Within the scope described in the claims, various modifications and deformations that can be made by those skilled in the art without creative efforts still fall within the protection scope of this patent.

Claims

1. An unsupervised medical image registration method based on Mamba-CNN encoded feature fusion, characterized in that, It includes the following steps: S1. Obtain a moving image and a fixed image, and perform stitching processing on the moving image and the fixed image to obtain a stitched image; S2. Input the stitched image into the Mamba branch encoder and the CNN branch encoder in the unsupervised medical image registration model respectively, and output the global feature map and the local feature map of the stitched image respectively; S3. Input the global feature map and the local feature map of the stitched image into the attention feature fusion module in the unsupervised medical image registration model, and output a weighted fusion feature map; S4. Input the weighted fusion feature map into the decoder in the unsupervised medical image registration model, and output a deformation field for aligning the moving image and the fixed image.

2. The unsupervised medical image registration method based on Mamba-CNN encoded feature fusion according to claim 1, wherein In S1, the stitching processing of the moving image and the fixed image includes: x in = Concat(Moving, Fixed) where x in is the stitched image; Concat() is the stitching operation; Moving is the moving image; Fixed is the fixed image.

3. The unsupervised medical image registration method based on Mamba-CNN encoded feature fusion according to claim 2, wherein, In S2, the Mamba branch encoder includes three levels of Mamba encoders, and the outputs of the three levels of Mamba encoders are: Among them, are the feature maps output after being processed by the Mamba encoders of the first, second, and third layers respectively; ML1, ML2, and ML3 are the Mamba encoders of the first, second, and third layers respectively; PE1, PE2, and PE3 are the patch embedding layers of the first, second, and third layers respectively; C1 is the number of channels output by the first-layer Mamba encoder; D, H, and W are the height, width, and length of the feature map respectively.

4. The unsupervised medical image registration method based on Mamba-CNN encoded feature fusion according to claim 3, wherein: The stitched image enters the Mamba encoder, and the stitched image is serialized based on the patch embedding layer, and sinusoidal positional embeddings are added to the serialized data to obtain a first feature tensor; the first feature tensor is respectively input into two data streams in the Mamba encoder; In the first data stream, the first feature tensor undergoes a linear transformation through a linear layer to adjust the feature dimension; the linearly transformed first feature tensor is passed to a one-dimensional convolutional layer to extract a local feature tensor; the local feature tensor undergoes a non-linear mapping through an activation function layer, and finally enters the SSM layer to capture the global feature and output the processed global feature tensor; In the second data stream, the first feature tensor undergoes a linear transformation through a linear layer to adjust the feature dimension; the linearly transformed first feature tensor enters the activation function layer and undergoes a non-linear mapping to enhance the feature expression; the enhanced first feature tensor is input into a gating mechanism, and the gating mechanism receives the first feature tensor and a control signal, and selectively retains or suppresses part of the information in the first feature tensor through element-wise operations, and finally outputs the filtered feature tensor; The global feature tensor output in the first data stream and the filtered feature tensor output in the second data stream are multiplied element-wise to obtain a fusion feature tensor; the fusion feature tensor undergoes a linear projection through a linear layer to output a first multi-dimensional feature tensor.

5. The unsupervised medical image registration method based on Mamba-CNN encoded feature fusion according to claim 4, characterized in that, The linear ordinary differential equation in the SSM layer is discretized by using the zero-order hold method: where, A is the state transition matrix, A = e ΔA ; B and C are the projection parameter matrices, B = (ΔA) -1 (e ΔA - I)ΔB, C = C; Δ is the time scale parameter, I is the identity matrix; t is the index of the token; h t is the state at time t; h t-1 is the state at time t - 1; x t is the input feature vector at time t; y t is the prediction or processing result at time t; According to the discretized linear ordinary differential equation, a serialized representation structure of the Mamba model is obtained: where y is the output sequence; x is the input sequence; * is the convolution operation; is the structured convolution kernel; L is the length of the sequence in the model.

6. According to the unsupervised medical image registration method based on Mamba-CNN encoding feature fusion in claim 3, in said S2, the spliced ​​image is pre-convolved and down-sampled twice to obtain a resolution of The initial input feature map x cnn_in , the initial input feature map x cnn_in Input to the CNN branch encoder, the CNN branch encoder includes three levels of CNN encoders, and the output of the three levels of CNN encoders is: Among them, start_channel is a hyperparameter; CL1, CL2, and CL3 are the CNN encoders of the first, second, and third layers respectively; They are the feature maps output after being processed by the CNN encoders of the first, second, and third layers respectively; C2 is the number of channels output by the first-layer CNN encoder.

7. According to the unsupervised medical image registration method based on Mamba-CNN encoded feature fusion described in claim 1, in S3, the attention feature fusion module includes a channel attention sub-module and a spatial attention sub-module; The feature maps of the outputs of each layer of the Mamba encoder in the Mamba branch encoder and the feature maps of the outputs of each layer of the CNN encoder in the CNN branch encoder The concatenated feature map X is used as the input to the channel attention sub-module and the spatial attention sub-module, where i = 1, 2, or 3; The outputs of the channel attention sub-module and the spatial attention sub-module are the channel-weighted feature map and the spatial-weighted feature map The channel-weighted feature map and the spatial-weighted feature map are fused to obtain the weighted fusion feature map: Among them, out is the weighted fusion feature map; Sigmoid is the activation function; ⊙ is the Hadamard multiplication.

8. According to the unsupervised medical image registration method based on Mamba-CNN encoded feature fusion described in claim 7, the channel attention sub-module is an ECA attention module; Features of each layer of the Mamba encoder in the Mamba branch encoder and features of each layer of the CNN encoder in the CNN branch encoder The concatenated feature map is input into the ECA attention module; The ECA attention module performs global average pooling on the input feature map to extract the global information Y of each channel, transfers the global information Y to a parameterized one-dimensional convolutional layer, and outputs the intermediate information Y'. After the intermediate information Y' is processed by the Sigmoid activation function, it is converted into a channel weight map. The channel weight map is multiplied by the original input feature map X using the Hadamard multiplication operation to generate a channel-weighted feature map: Among them, It is a channel weight map.

9. The method for unsupervised medical image registration based on Mamba-CNN encoded feature fusion according to claim 7, wherein the features of each layer of the Mamba encoder in the Mamba branch encoder and the features of each layer of the CNN encoder in the CNN branch encoder The spliced feature map is input into the spatial attention sub-module; The spatial attention sub-module calculates the average value and the maximum value of the input feature map in each channel, respectively obtaining the mean output and the maximum output. The mean output and the maximum output are stacked to form a two-channel feature map. This two-channel feature map is processed by a 3D convolutional layer with a convolutional kernel size of 3 or 7 to extract spatial information and regenerate a single-channel feature map. The single-channel feature map generates a spatial weight map through the Sigmoid activation function; the spatial weight map is multiplied by the original input feature map X using the Hadamard multiplication operation to generate a spatially weighted feature map Among them, is a spatial weight map.

10. The loss function of the unsupervised medical image registration model according to claim 1 for the unsupervised medical image registration method based on Mamba-CNN encoded feature fusion is: Among them, is the normalized cross - correlation coefficient; is the smoothness regularization penalty term; is the weighted total loss function; NCC is an index for image matching and registration; Moving i is the \(i\) - th moving image; is the \(i\) - th fixed image; is the warping operator that applies the deformation field to the moving image; \(\varphi\) i is the estimated deformation field; \(\Theta\) is the model parameter; \(\nabla\) is the first - order gradient calculation for smoothness obtained by finite differences; \(N\) is the number of image pairs for training; \(\lambda\) is the hyperparameter; \(v\) i is the gradient value of the deformation field; is the square of the \(L_2\) norm.

Citation Information

Cited By

  • Photoacoustic image enhancement method and device combining Mama and CNN (Convolutional Neural Network)

    CN120707580A

  • Image processing method and system and storage medium

    CN120931669A