An image registration method based on cross-neighbor attention enhancement mechanism

CN118447061BActive Publication Date: 2026-09-22WUHAN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410471133.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-18
Publication Date
2026-09-22
Estimated Expiration
2044-04-18

AI Technical Summary

Technical Problem

但是,由于Transformer受限于其特征空间分辨率,无法充分提取局部信息

Benefits of technology

[0062](1)、使用三分支编码器,通过两个多尺度特征张量提取模块分别提取浮动图像和固定图像的多尺度特征,通过四阶段Swin-Transformer模块提取浮动图像和固定图像拼接结果的多尺度特征,从而在浮动图像和固定图像的多尺度的特征之间建立显式的映射关系,有效弥补Swin-Transformer难以充分提取位移场方向信息的不足;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118447061B_ABST
    Figure CN118447061B_ABST
Patent Text Reader

Abstract

The application discloses an image registration method based on a cross-neighbor attention enhancement mechanism, first, the fixed image and the floating image are extracted by a multi-scale feature tensor extraction module to extract feature information with local detail dependency, and the splicing result of the fixed image and the floating image is extracted by a four-stage Swin-Transformer module to extract feature information with long-distance dependency; secondly, different feature information obtained is input into a multi-scale displacement fusion module based on cross-neighbor attention, and a cross-neighbor attention enhancement mechanism is adopted to improve registration accuracy and improve detail information; then, multi-scale displacement field features of different feature information are fused to obtain an accurate dense displacement field, so that the image registration accuracy is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of digital image processing technology, specifically relating to an image registration method based on a cross-neighborhood attention enhancement mechanism, applicable to various modal imaging data. Background Technology

[0002] Image registration is a classic problem and technical challenge in image processing. Its purpose is to compare or fuse multimodal or temporal images of the same object acquired under different conditions, such as images from different acquisition devices, at different acquisition times, or from different shooting angles, to ensure a one-to-one spatial correspondence. Image registration is often a preprocessing step in image information fusion and has significant practical value in computer vision. For example, in computer vision, registration can be used for video analysis, pattern recognition, motion tracking, and recognition. In materials mechanics, registration is typically used to study mechanical properties; by acquiring information from different sensors (such as shape and temperature), it can obtain data such as strain fields and temperature fields.

[0003] Traditional image registration methods mostly rely on manually designed features to establish spatial correspondences between image pairs. However, these methods are based on iterative optimization strategies, requiring iterative solutions for each new image pair, resulting in significant time consumption in practical applications. To improve registration efficiency and address the lack of realistic deformation fields in learning-based registration algorithms, unsupervised image registration algorithms based on deep learning have become the mainstream. However, registration methods based on convolutional neural networks (CNNs) are limited by the size of the receptive field of the convolutional kernel, failing to consider long-distance spatial dependencies in images and thus hindering the effective extraction of global information. Constructing a visual transformer (Transformer) can acquire long-distance dependency information through self-attention, mitigating some of the shortcomings of CNNs. However, the Transformer is limited by its feature space resolution, preventing the full extraction of local information. Furthermore, single-stream input modes struggle with explicit feature matching between floating and stationary images, leading to slow model iteration and optimization speeds, and difficulty in quickly capturing the orientation information of the displacement field during initial iterations. Summary of the Invention

[0004] The purpose of this invention is to address the aforementioned problems in the existing technology by providing an image registration method based on a cross-neighborhood attention enhancement mechanism.

[0005] The above-mentioned objectives of the present invention are achieved by the following technical means:

[0006] An image registration method based on a cross-neighborhood attention enhancement mechanism includes the following steps:

[0007] Step 1: Obtain multiple images from the public dataset to form a training image set. Take two images from the training image set as a pair as a group of training samples. In each group of training samples, take one image as a floating image and the other image as a fixed image.

[0008] Step 2: Build a registration network model. The registration network model includes a three-branch encoder, a multi-scale displacement fusion module based on cross-neighborhood attention, and a decoder module. The three-branch encoder includes a four-stage Swin-Transformer module and a multi-scale feature tensor extraction module.

[0009] Step 3: Input the floating image into the multi-scale feature tensor extraction module to obtain the floating image feature tensor F at multiple scales. k m For k = 0,...,5, a fixed image is input into a multi-scale feature tensor extraction module to obtain fixed image feature tensors F at multiple scales. v f ,i=0,...,5,k and v are both scale numbers;

[0010] Step 4: Stitch the floating image and the fixed image together, then encode the stitched image and input it into the four-stage Swin-Transformer module to obtain the ST output features F of each stage. j st j is the stage number, j = 0,...,3;

[0011] Step 5: Convert the floating image feature tensor F 5 m Fixed image feature tensor F 5 f and ST output feature F 3 st These are respectively used as the first, second, and third inputs to a multi-scale displacement fusion module based on cross-neighborhood attention to obtain the displacement field increment. As the increment of the output displacement field Increment of displacement field Add the initial displacement field to obtain the displacement field φ1, where the initial displacement field is a set value;

[0012] Step 6: Upsample the displacement field φ1 so that the width, height, and depth of the displacement field φ1 are consistent with the floating image feature tensor F. 4 m The width, height, and depth of the floating image feature tensor F are the same. 4 m Perform the twisting to obtain the twisted feature tensor F. 4 mwTwisted feature tensor F 4 mw The width, height, and depth in the dimensions of the floating image feature tensor F 4 m The width, height, and depth of the dimension are the same, and then the distorted feature tensor F is... 4 mw Fixed image feature tensor F 4 f ST output feature F 2 st and ST output feature F 3 st These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 1 D As a feature of the output decoder;

[0013] Step 7: Convert the distorted feature tensor F 4 mw Fixed image feature tensor F 4 f and decoder output feature F 1 D These are respectively used as the first, second, and third inputs to a multi-scale displacement fusion module based on cross-neighborhood attention, to obtain the displacement field increment. As the increment of the output displacement field The displacement field φ1 is upsampled and compared with the displacement field increment. The displacement field φ2 is obtained by adding them together;

[0014] Step 8: Upsample the displacement field φ2 so that its width, height, and depth are correlated with the floating image feature tensor F. 3 m The width, height, and depth of the floating image feature tensor F are the same. 3 m Perform the twisting to obtain the twisted feature tensor F. 3 mw Twisted feature tensor F 3 mw The width, height, and depth in the dimensions of the floating image feature tensor F 3 m The width, height, and depth of the distorted feature tensor F are the same. 3 mw Fixed image feature tensor F 3 f ST output features F 1 stand decoder output feature F 1 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 2 D As a feature of the output decoder;

[0015] Step 9: Convert the distorted feature tensor F 3 mw Fixed image feature tensor F 3 f and decoder output feature F 2 D These are respectively used as the first, second, and third inputs to a multi-scale displacement fusion module based on cross-neighborhood attention, to obtain the displacement field increment. As the increment of the output displacement field The displacement field φ2 is upsampled and then compared with the displacement field increment. The displacement field φ3 is obtained by adding them together;

[0016] Step 10: Upsample the displacement field φ3 so that the width, height, and depth of the displacement field φ3 are related to the floating image feature tensor F. 2 m The width, height, and depth of the floating image feature tensor F are the same. 2 m Perform the twisting to obtain the twisted feature tensor F. 2 mw Twisted feature tensor F 2 mw The width, height, and depth in the dimensions of the floating image feature tensor F 2 m The width, height, and depth of the distorted feature tensor F are the same. 2 mw Fixed image feature tensor F 2 f ST output feature F 0 st and decoder output feature F 2 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 3 D As a feature of the output decoder;

[0017] Step 11: Convert the distorted feature tensor F 2 mw Fixed image feature tensor F2 f and decoder output feature F 3 D These are respectively used as the first, second, and third inputs to a multi-scale displacement fusion module based on cross-neighborhood attention, to obtain the displacement field increment. As the increment of the output displacement field The displacement field φ3 is upsampled and compared with the displacement field increment. The displacement field φ4 is obtained by adding them together;

[0018] Step 12: Upsample the displacement field φ4 so that the width, height, and depth of the displacement field φ4 are related to the floating image feature tensor F. 1 m The width, height, and depth of the floating image feature tensor F are the same. 1 m Perform the twisting to obtain the twisted feature tensor F. 1 mw The twisted feature tensor F 1 mw Fixed image feature tensor F 1 f ST output features F 0 st and decoder output feature F 3 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 4 D As output decoder features, the decoder output feature F 4 D Perform convolution processing to obtain the displacement field increment. After upsampling, the displacement field φ4 is compared with the displacement field increment. The displacement field φ5 is obtained by adding them together;

[0019] Step 13: Upsample the displacement field φ5 so that the width, height, and depth of the displacement field φ5 are related to the floating image feature tensor F. 0 m The width, height, and depth of the floating image feature tensor F are the same. 0 m Perform the twisting to obtain the twisted feature tensor F. 0 mw The twisted feature tensor F 0 mw Fixed image feature tensor F 0 f ST output features F 0st and decoder output feature F 4 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 5 D As output decoder features, the decoder output feature F 5 D Perform convolution processing to obtain the displacement field increment. After upsampling, the displacement field φ5 is compared with the displacement field increment. The displacement field φ6 is obtained by adding the two images together. The floating image is then distorted by the displacement field φ6 to obtain the target distorted image.

[0020] Step 14: Minimize the total loss function With the goal of training the registration network model using the training samples in step 1, the model parameters are saved after training to obtain the optimized registration network model.

[0021] Total loss function Based on the following formula:

[0022]

[0023] In the formula, I f It is a fixed image, I m It is a floating image, where φ is the displacement field φ6. It is the similarity loss between the fixed image and the distorted image of the target, where λ is the weighting coefficient. It is the smoothing and regularization loss of the displacement field φ6;

[0024] Step 15: Input the floating image and the fixed image to be processed into the optimization registration network model to obtain the optimized target distortion image after the floating image is deformed.

[0025] As described above, step 3 specifically includes the following steps:

[0026] Step 3.1: Input the floating image into the multi-scale feature tensor extraction module. The multi-scale feature tensor extraction module performs multiple feature extraction operations on the floating image. Each feature extraction operation includes convolution and average pooling to obtain the floating image feature tensor F at multiple scales. k m Floating image feature tensor F k m The dimensions include the floating image feature tensor F. k m The number of channels, floating image feature tensor F k m The width of the floating image feature tensor Fk m The high, and floating image feature tensor F k m The depth;

[0027] Step 3.2: Input the fixed image into the multi-scale feature tensor extraction module. The multi-scale feature tensor extraction module performs multiple feature extraction operations on the fixed image. Each feature extraction operation includes convolution and average pooling to obtain the fixed image feature tensor F at multiple scales. v f Fixed image feature tensor F v f The dimensions include a fixed image feature tensor F. v f The number of channels, with a fixed image feature tensor F v f The width of the fixed image feature tensor F v f The high, and the fixed image feature tensor F v f The depth.

[0028] As described above, step 4 specifically includes the following steps:

[0029] Step 4.1: After concatenating the floating image and the fixed image along the channel dimension, the images are encoded to obtain a block feature map;

[0030] Step 4.2: After linear mapping, the block feature maps are transformed into feature map tensors. These feature map tensors are then input into the four-stage Swin-Transformer module to obtain the ST output features F for each stage. j st ST output feature F j st The dimensions include ST output features F j st The number of channels, ST output features F j st The width of ST output feature F j st The high, and ST output feature F j st The depth.

[0031] As mentioned above, the dimensions of both the floating image and the fixed image are (1, w, h, d), where 1 is the number of channels of the feature of the floating image or the fixed image, w is the width of the floating image or the fixed image, h is the height of the floating image or the fixed image, and d is the depth of the floating image or the fixed image.

[0032] In step 3.1, the floating image feature tensor F 0m ~Floating image feature tensor F 5 m The dimensions are (2, w, h, d), (4, w / 2, h / 2, d / 2), (8, w / 4, h / 4, d / 4), (16, w / 8, h / 8, d / 8), (32, w / 16, h / 16, d / 16), (32, w / 32, h / 32, d / 32);

[0033] In step 3.2, the image feature tensor F is fixed. 0 f ~Fixed image feature tensor F 5 f The dimensions are (2, w, h, d), (4, w / 2, h / 2, d / 2), (8, w / 4, h / 4, d / 4), (16, w / 8, h / 8, d / 8), (32, w / 16, h / 16, d / 16), (32, w / 32, h / 32, d / 32);

[0034] In step 4.2, the ST output feature F 0 st ~ST output characteristics F 3 st The dimensions are (c, w / 4, h / 4, d / 4), (2c, w / 8, h / 8, d / 8), (4c, w / 16, h / 16, d / 16), and (8c, w / 32, h / 32, d / 32);

[0035] Where c is the number of channels in the feature map tensor.

[0036] As described above, the first, second, and third inputs are fed into a multi-scale displacement fusion module based on cross-neighborhood attention to obtain the output displacement field increment. Specifically, the process includes the following:

[0037] First, the second input is sequentially processed through linear mapping, dimensional expansion, and transpose to obtain the query vector Q. The first input is sequentially processed through linear mapping, dimensional expansion, and transpose to obtain the intermediate feature tensor G. The width, height, and depth dimensions of the intermediate feature tensor G are then filled to obtain the intermediate feature tensor E. Finally, the intermediate feature tensor E is divided into blocks with a step size of 1 and a size of 3×3×3 to obtain the key vector K.

[0038] Secondly, the query vector Q is multiplied by the key vector K to obtain the neighborhood attention NA;

[0039] Next, construct a value vector V, transpose V, and multiply it with the neighborhood attention NA to obtain an intermediate feature tensor F. Compress the dimensions of the intermediate feature tensor F to obtain the local displacement field increment. The global displacement field increment is obtained by convolution processing the third input.

[0040] Finally, the local displacement field increment and global displacement field increment The output displacement field increment is obtained by weighting using the following formula.

[0041]

[0042] In the formula, W 1 and W 2 These are the local displacement field increments. and global displacement field increment The weight.

[0043] As mentioned above, the dimension of the query vector Q is (W1, H1, D1, 1, Am), where W1, H1, and D1 are the width, height, and depth of the query vector Q, respectively; the values ​​of W1, H1, and D1 are the width, height, and depth of the second input, respectively; 1 is the vector length; m is the vector subspace dimension; and A is the number of subspaces.

[0044] The intermediate feature tensor E has dimensions (W2+2, H2+2, D2+2, 1, Am), where W2, H2, and D2 are the width, height, and depth of the first input dimension, respectively; and 1 is the vector length.

[0045] The key vector K has dimensions (W3, H3, D3, 27, Am), where W3, H3, and D3 are the width, height, and depth of the key vector K, respectively; the values ​​of W3, H3, and D3 are the width, height, width, and depth of the first input, respectively; and 27 is the total vector length of the 27 blocks obtained after processing with a size of 3×3×3.

[0046] The dimensions of the neighborhood attention NA are (W4, H4, D4, 27, 1), where W4, H4, and D4 are the width, height, and depth of the neighborhood attention NA, respectively; the values ​​of W4, H4, and D4 are the width, height, and depth of the query vector Q or key vector K, respectively; 27 is the total vector length of the 27 blocks obtained after processing with a size of 3×3×3; and 1 is the attention for each block.

[0047] The dimensions of the value vector V are (W5, H5, D5, 27, 3), where W5, H5, and D5 are the width, height, and depth of the value vector V, respectively. The width, height, and depth of the value vector V are the same as the width, height, and depth of the neighborhood attention NA. 27 is the vector length; 3 is the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image.

[0048] The intermediate feature tensor F has dimensions (W6, H6, D6, 3, 1), where W6, H6, and D6 are the width, height, and depth of the intermediate feature tensor F, respectively; the values ​​of W6, H6, and D6 are the width, height, and depth of the value vector V or the neighborhood attention NA, respectively; and 3 represents the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image.

[0049] Local displacement field The dimensions are (W7, H7, D7, 3), where W7, H7, and D7 are the local displacement fields. The width, height, and depth of the intermediate feature tensor F are represented by the values ​​of W7, H7, and D7, respectively; 3 represents the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image.

[0050] Global displacement field The dimensions are (W8, H8, D8, 3), where W8, H8, and D8 are the global displacement field increments, respectively. The width, height, and depth of the third input dimension are represented by W8, H8, and D8, respectively; 3 represents the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image.

[0051] The number A of subspaces is 8, 8, 4, and 2 in steps 5, 7, 9, and 11, respectively.

[0052] As described above, the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input are input to the decoder module to obtain the output decoder features. The specific process includes the following steps:

[0053] First, the inputs of the first decoder, the second decoder, and the fourth decoder are concatenated and then upsampled to obtain the output upsampled features;

[0054] Then, the output upsampled features and the third decoder input are concatenated and processed in two rounds. Each round of processing includes convolution, ReLU activation, and instance normalization to obtain the output decoder features.

[0055] Similarity loss between fixed and distorted images as described above Locally normalized cross-correlation loss LNCC(I) f ,I m ,φ), Locally Normalized Cross-Correlation Loss LNCC(I f ,I m ,φ) is based on the following formula:

[0056]

[0057] In the formula, The floating image is located in a local window of size n, centered at voxel p. 3 The average voxel value; The image is fixed in a local window of size n centered at position p voxels. 3 Mean voxel value; p i Centered on voxel p, with a range of n 3 The position of the neighboring voxels, where i is the position number, i∈n 3 ; This is the target distortion image obtained by applying a displacement field φ6 to the floating image. The target distorted image at position p i voxel values; The target distorted image is located at voxel p, with a local window size of n. 3 Average voxel value; I f (p i ) is a fixed image at position p i The voxel value; Ω is the range of voxel positions in a fixed image or a distorted image of a target.

[0058] As described above, the smoothing regularization loss of the displacement field Based on the following formula:

[0059]

[0060] In the formula, u(p) is the spatial gradient of the displacement field φ6. Ω is the change in displacement of adjacent voxels of voxel p, and Ω is the range of voxel positions in a fixed image or a distorted image of a target.

[0061] Compared with the prior art, the present invention has the following advantages:

[0062] (1) Using a three-branch encoder, the multi-scale features of floating and fixed images are extracted by two multi-scale feature tensor extraction modules respectively, and the multi-scale features of the splicing result of floating and fixed images are extracted by a four-stage Swin-Transformer module. Thus, an explicit mapping relationship is established between the multi-scale features of floating and fixed images, which effectively makes up for the lack of Swin-Transformer in fully extracting displacement field direction information.

[0063] (2) The multi-scale displacement fusion module based on cross-neighborhood attention can obtain displacement fields with long-distance dependence and displacement fields with local detail dependence, realize the fusion of multi-scale displacement field information, and obtain a more accurate deformation field.

[0064] (3) The displacement field incremental superposition scheme is adopted to realize the gradual learning process of displacement field from low resolution to high resolution, which improves the convergence speed of the registration model and also improves the deformation effect of images with large-scale deformation. Attached Figure Description

[0065] Figure 1 This is a flowchart of the present invention;

[0066] Figure 2 This is a schematic diagram of the network structure of the registration model of the present invention;

[0067] Figure 3 This is a schematic diagram of the multi-scale displacement fusion module based on cross-neighborhood attention according to the present invention;

[0068] Figure 4 This is a schematic diagram of the decoder module of the present invention. Detailed Implementation

[0069] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to embodiments. The embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0070] Example 1:

[0071] An image registration method based on a cross-neighborhood attention enhancement mechanism includes the following steps:

[0072] Step 1: Obtain multiple images from a public dataset (such as the LPBA40 public dataset) to form a training image set. Take two images from the training image set as a pair, and use one image from each pair as the floating image and the other as the fixed image. Both the floating and fixed images have dimensions (1, w, h, d), where 1 is the number of feature channels in the floating or fixed image, w is the width of the floating or fixed image, h is the height of the floating or fixed image, and d is the depth of the floating or fixed image.

[0073] Step 2: Build a registration network model. The registration network model includes a three-branch encoder, a multi-scale displacement fusion module based on cross-neighborhood attention, and a decoder module. The three-branch encoder includes a four-stage Swin-Transformer (ST, a transformer network based on a moving window) module and two multi-scale feature tensor extraction modules.

[0074] The number of attention mechanism layers in each stage of the four-stage Swin-Transformer module are 2, 2, 4, and 2, respectively.

[0075] A three-branch encoder is used to extract multi-scale features of floating and fixed images through two multi-scale feature tensor extraction modules, and then a four-stage Swin-Transformer module is used to extract multi-scale features of the spliced ​​floating and fixed images. This establishes an explicit mapping relationship between the multi-scale features of floating and fixed images, effectively making up for the Swin-Transformer's inability to fully extract displacement field direction information.

[0076] Step 3: Input the floating image into one of the multi-scale feature tensor extraction modules to obtain the floating image feature tensors F at multiple scales. k m For k = 0,...,5, the fixed image is input into another multi-scale feature tensor extraction module to obtain fixed image feature tensors F at multiple scales. v f Where i = 0,...,5, and k and v are scale numbers, the specific steps include:

[0077] Step 3.1: Input the floating image into one of the multi-scale feature tensor extraction modules. The multi-scale feature tensor extraction module performs multiple feature extraction operations on the floating image with input dimensions (1, w, h, d). Each feature extraction operation includes one convolution and one average pooling, to obtain the floating image feature tensor F at multiple scales. k m k = 0,...,5, floating image feature tensor F k mThe dimension is (floating image feature tensor F) k m The number of channels, floating image feature tensor F k m The width of the floating image feature tensor F k m High, floating image feature tensor F k m (depth); floating image feature tensor F 0 m ~Floating image feature tensor F 5 m The dimensions are (2, w, h, d), (4, w / 2, h / 2, d / 2), (8, w / 4, h / 4, d / 4), (16, w / 8, h / 8, d / 8), (32, w / 16, h / 16, d / 16), (32, w / 32, h / 32, d / 32);

[0078] Step 3.2: Input the fixed image into another multi-scale feature tensor extraction module. The multi-scale feature tensor extraction module performs multiple feature extraction operations on the fixed image with input dimensions (1, w, h, d). Each feature extraction operation includes one convolution and one average pooling to obtain a fixed image feature tensor F at multiple scales. v f i = 0,...,5, with a fixed image feature tensor F v f The dimension is (fixed image feature tensor F) v f The number of channels, with a fixed image feature tensor F v f The width of the fixed image feature tensor F v f High, fixed image feature tensor F v f (depth); fixed image feature tensor F 0 f ~Fixed image feature tensor F 5 f The dimensions are (2,w,h,d), (4,w / 2,h / 2,d / 2), (8,w / 4,h / 4,d / 4), (16,w / 8,h / 8,d / 8), (32,w / 16,h / 16,d / 16), and (32,w / 32,h / 32,d / 32).

[0079] Step 4: Stitch the floating image and the fixed image together, then encode the stitched image and input it into the four-stage Swin-Transformer module to obtain the ST output features F of each stage. jst j = 0,...,3, where j is the stage number, specifically including the following steps;

[0080] Step 4.1: After concatenating the floating image and the fixed image along the channel dimension, the images are encoded to obtain patch feature maps;

[0081] Step 4.2: After linear mapping, the block feature maps are transformed into feature map tensors. These feature map tensors are then input into the four-stage Swin-Transformer module to obtain the ST output features F for each stage. j st j = 0,...,3, ST output feature F j st The dimension is (ST output feature F) j st The number of channels, ST output features F j st The width of ST output feature F j st High, ST output feature F j st (depth); ST output feature F 0 st ~ST output characteristics F 3 st The dimensions are (c, w / 4, h / 4, d / 4), (2c, w / 8, h / 8, d / 8), (4c, w / 16, h / 16, d / 16), and (8c, w / 32, h / 32, d / 32), respectively, where c is the number of channels of the feature map tensor.

[0082] Step 5: Convert the floating image feature tensor F 5 m Fixed image feature tensor F 5 f and ST output feature F 3 st These are respectively used as the first, second, and third inputs to a multi-scale displacement fusion module based on cross-neighborhood attention to obtain the displacement field increment. As the increment of the output displacement field Increment of displacement field Adding the initial displacement field to the displacement field φ1, in this embodiment, the initial displacement field is a set value, such as 0, and specifically includes the following steps:

[0083] Step 5.1: The second input is sequentially processed through linear mapping, dimension expansion, and transpose to obtain the query vector Q. The dimensions of the query vector Q are (W1, H1, D1, 1, Am), where W1, H1, and D1 are the width, height, and depth of the query vector Q, respectively. The values ​​of W1, H1, and D1 are the width, height, and depth of the second input, respectively. In this step, the values ​​of W1, H1, and D1 are fixed image feature tensors F. 5 f The dimensions of the vector are width, height, and depth, i.e., w / 32, h / 32, and d / 32; 1 is the vector length; m is the dimension of the vector subspace; A is the number of subspaces, which is 8 in step 5.

[0084] The first input is then processed sequentially through linear mapping, dimensional expansion, and transpose to obtain an intermediate feature tensor G. The dimensions of the intermediate feature tensor G are then padded to (W2+2, H2+2, D2+2, 1, Am) to obtain an intermediate feature tensor E, where W2, H2, and D2 are the width, height, and depth of the first input, respectively. In this step, the values ​​of W2, H2, and D2 are respectively the values ​​of the floating image feature tensor F. 5 m The dimensions of the vector are width, height, and depth, i.e., w / 32, h / 32, and d / 32; 1 is the vector length.

[0085] Finally, the intermediate feature tensor E is divided into blocks with a step size of 1 and a size of (3×3×3) to obtain the key vector K. The dimensions of the key vector K are (W3, H3, D3, 27, Am), where W3, H3, and D3 are the width, height, and depth of the key vector K, respectively. The values ​​of W3, H3, and D3 are the width, height, and depth of the first input, respectively. In this step, the values ​​of W3, H3, and D3 are respectively the values ​​of the floating image feature tensor F. 5 m The dimensions of the vector are width, height, and depth, i.e., w / 32, h / 32, and d / 32; 27 is the vector length (i.e., the total vector length of the 27 blocks obtained after processing into blocks of size (3×3×3));

[0086] Step 5.2: Multiply the query vector Q and the key vector K to obtain the neighborhood attention NA. The dimensions of the neighborhood attention NA are (W4, H4, D4, 27, 1), where W4, H4, and D4 are the width, height, and depth of the dimensions of the neighborhood attention NA, respectively. The values ​​of W4, H4, and D4 are the width, height, and depth of the dimensions of the query vector Q (or key vector K). In this step, the values ​​of W4, H4, and D4 are w / 32, h / 32, and d / 32, respectively; 27 is the total vector length of the 27 blocks obtained after processing with a size of (3×3×3); and 1 is the attention for each block.

[0087] Step 5.3: Construct a value vector V with dimensions (W5, H5, D5, 27, 3), where W5, H5, and D5 are the width, height, and depth of the value vector V, respectively. The width, height, and depth of the value vector V are the same as those of the neighborhood attention NA. In this step, the values ​​of W5, H5, and D5 are w / 32, h / 32, and d / 32, respectively; 27 is the vector length; and 3 is the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image.

[0088] The value vector V is then transposed and multiplied by the neighborhood attention NA to obtain the intermediate feature tensor F. The dimensions of the intermediate feature tensor F are (W6, H6, D6, 3, 1), where W6, H6, and D6 are the width, height, and depth of the intermediate feature tensor F, respectively. The values ​​of W6, H6, and D6 are the width, height, and depth of the value vector V (or the neighborhood attention NA), respectively. In this step, the values ​​of W6, H6, and D6 are w / 32, h / 32, and d / 32, respectively; 3 represents the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image; and 1 represents the additional dimension.

[0089] Then, the dimensions in the intermediate feature tensor F are compressed to obtain the local displacement field increment. Local displacement field increment The dimensions are (W7, H7, D7, 3), where W7, H7, and D7 are the local displacement field increments, respectively. The width, height, and depth of the intermediate feature tensor F are represented by the values ​​of W7, H7, and D7, respectively. In this step, the values ​​of W7, H7, and D7 are w / 32, h / 32, and d / 32, respectively. 3 represents the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image.

[0090] Finally, the global displacement field increment is obtained by convolution processing the third input. Global displacement field increment The dimensions are (W8, H8, D8, 3), where W8, H8, and D8 are the global displacement field increments, respectively. The width, height, and depth of the third input dimension are represented by W8, H8, and D8, respectively. In this step, the values ​​of W8, H8, and D8 are used to output the feature F. 3 st The dimensions are width, height, and depth, i.e., w / 32, h / 32, and d / 32; 3 represents the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image.

[0091] Step 5.4: Increment the local displacement field and global displacement field increment The weighted sum is used to obtain the output displacement field increment. Output displacement field increment Adding this to the initial displacement field yields the displacement field φ1. In this embodiment, the initial displacement field is a set value, such as 0. The output displacement field increment is then calculated. Based on the following formula:

[0092]

[0093] Among them, W 1 and W 2 These are the local displacement field increments. and global displacement field increment The weights, W 1 and W 2 In step 5, the value can be set to 0.5.

[0094] Step 6: Upsample the displacement field φ1 so that the width, height, and depth of the displacement field φ1 are consistent with the floating image feature tensor F. 4 m The width, height, and depth of the floating image feature tensor F are the same. 4 m Perform the twisting to obtain the twisted feature tensor F. 4 mw Then the distorted feature tensor F 4 mw Fixed image feature tensor F 4 f ST output features F 2 st and ST output feature F 3 stThese four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 1 D As the output decoder feature, the decoder output feature F 1 D The width, height, and depth of the dimension are w / 16, h / 16, and d / 16 respectively, and the specific steps include:

[0095] Step 6.1: After upsampling the displacement field φ1, apply the feature tensor F of the floating image. 4 m Perform the twisting to obtain the twisted feature tensor F. 4 mw Twisted feature tensor F 4 mw The width, height, and depth in the dimensions of the floating image feature tensor F 4 m The width, height, and depth are the same in all dimensions;

[0096] Step 6.2: Concatenate the inputs of the first decoder, the second decoder, and the fourth decoder, and then upsample the result to obtain the output upsampled feature;

[0097] Step 6.3: After concatenating the output upsampled features and the input of the third decoder, the features are processed in two rounds. Each round of processing includes convolution, ReLU activation, and instance normalization to obtain the output decoder features.

[0098] Step 7: Convert the distorted feature tensor F 4 mw Fixed image feature tensor F 4 f and decoder output feature F 1 D The first, second, and third inputs are respectively fed into a multi-scale displacement fusion module based on cross-neighborhood attention to obtain displacement field increments. As the increment of the output displacement field The displacement field φ1 is upsampled and compared with the displacement field increment. The displacement field φ2 is obtained by adding them together;

[0099] In step 7, the values ​​of W1, H1, and D1 in the dimension of the query vector Q are respectively the values ​​of the fixed image feature tensor F. 4 f The width, height, and depth of the dimension are w / 16, h / 16, and d / 16, respectively; the number of subspaces A in the dimension of the query vector Q is 8 in step 7.

[0100] The values ​​of W2, H2, and D2 in the dimensions of the intermediate feature tensor E in step 7 are respectively the values ​​of the warped feature tensor F. 4 mw The dimensions of width, height, and depth, i.e., w / 16, h / 16, and d / 16;

[0101] The values ​​of W3, H3, and D3 in the dimension of the key vector K in step 7 are respectively the values ​​of the warped feature tensor F. 4 mw The dimensions of width, height, and depth, i.e., w / 16, h / 16, and d / 16;

[0102] The values ​​of W4, H4, and D4 in the dimensions of the neighborhood attention NA in step 7 are w / 16, h / 16, and d / 16, respectively.

[0103] In step 7, the values ​​of W5, H5, and D5 in the dimension of the value vector V are w / 16, h / 16, and d / 16, respectively; the 3 in the dimension of the value vector V represents the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image.

[0104] In step 7, the values ​​of W6, H6, and D6 in the intermediate feature tensor F are w / 16, h / 16, and d / 16, respectively; the 3 in the intermediate feature tensor F represents the displacement of the voxels of the floating image in the width, height, and depth directions of the floating image.

[0105] Local displacement field increment The values ​​of W7, H7, and D7 in the dimensions are w / 16, h / 16, and d / 16 respectively in step 7; local displacement field increment The dimension 3 represents the displacement of the voxels of the floating image in the width, height, and depth directions of the floating image.

[0106] Global displacement field increment The values ​​of W8, H8, and D8 in the dimension of F in step 7 are respectively the decoder output features F 1 D The dimensions of width, height, and depth, i.e., w / 16, h / 16, and d / 16; global displacement field increment. The dimension 3 represents the displacement of the voxels of the floating image in the width, height, and depth directions of the floating image.

[0107] W 1 and W 2 In step 7, the value can be set to 0.5.

[0108] Step 8: Upsample the displacement field φ2 so that its width, height, and depth are correlated with the floating image feature tensor F. 3 m The width, height, and depth of the floating image feature tensor F are the same. 3 m Perform the twisting to obtain the twisted feature tensor F. 3 mw Twisted feature tensor F 3 mw The width, height, and depth in the dimensions of the floating image feature tensor F 3 m The width, height, and depth of the distorted feature tensor F are the same. 3 mw Fixed image feature tensor F 3 f ST output features F 1 st and decoder output feature F 1 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 2 D As the output decoder feature, the decoder output feature F 2 D The width, height, and depth of the dimension are w / 8, h / 8, and d / 8, respectively.

[0109] Step 9: Convert the distorted feature tensor F 3 mw Fixed image feature tensor F 3 f and decoder output feature F 2 D The first, second, and third inputs are respectively fed into a multi-scale displacement fusion module based on cross-neighborhood attention to obtain displacement field increments. As the increment of the output displacement field The displacement field φ2 is upsampled and then compared with the displacement field increment. The displacement field φ3 is obtained by adding them together;

[0110] In step 9, the values ​​of W1, H1, and D1 in the dimension of the query vector Q are respectively the values ​​of the fixed image feature tensor F. 3 f The width, height, and depth of the dimension are w / 8, h / 8, and d / 8, respectively; the number of subspaces A in the dimension of the query vector Q is 4 in step 9.

[0111] The values ​​of W2, H2, and D2 in the dimensions of the intermediate feature tensor E in step 9 are respectively the values ​​of the warped feature tensor F. 3 mw The dimensions of width, height, and depth, i.e., w / 8, h / 8, and d / 8;

[0112] The values ​​of W3, H3, and D3 in the dimension of the key vector K in step 9 are respectively the values ​​of the warped feature tensor F. 3 mw The dimensions of width, height, and depth, i.e., w / 8, h / 8, and d / 8;

[0113] The values ​​of W4, H4, and D4 in the dimensions of the neighborhood attention NA in step 9 are w / 8, h / 8, and d / 8, respectively.

[0114] In step 9, the values ​​of W5, H5, and D5 in the dimension of the value vector V are w / 8, h / 8, and d / 8, respectively; the 3 in the dimension of the value vector V represents the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image.

[0115] In step 9, the values ​​of W6, H6, and D6 in the intermediate feature tensor F are w / 8, h / 8, and d / 8, respectively; the 3 in the intermediate feature tensor F represents the displacement of the voxels of the floating image in the width, height, and depth directions of the floating image.

[0116] Local displacement field increment The values ​​of W7, H7, and D7 in the dimensions are w / 8, h / 8, and d / 8 respectively in step 9; local displacement field increment The dimension 3 represents the displacement of the voxels of the floating image in the width, height, and depth directions of the floating image.

[0117] Global displacement field increment The values ​​of W8, H8, and D8 in the dimension of F in step 9 are respectively the decoder output features F 2 D The dimensions of width, height, and depth, i.e., w / 8, h / 8, and d / 8; global displacement field increment. The dimension 3 represents the displacement of the voxels of the floating image in the width, height, and depth directions of the floating image.

[0118] W 1 and W 2 In step 9, the value can be set to 0.5.

[0119] Step 10: Upsample the displacement field φ3 so that the width, height, and depth of the displacement field φ3 are related to the floating image feature tensor F. 2m The width, height, and depth of the floating image feature tensor F are the same. 2 m Perform the twisting to obtain the twisted feature tensor F. 2 mw Twisted feature tensor F 2 mw The width, height, and depth in the dimensions of the floating image feature tensor F 2 m The width, height, and depth of the distorted feature tensor F are the same. 2 mw Fixed image feature tensor F 2 f ST output feature F 0 st and decoder output feature F 2 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 3 D As the output decoder feature, the decoder output feature F 3 D The width, height, and depth of the dimension are w / 4, h / 4, and d / 4, respectively.

[0120] Step 11: Convert the distorted feature tensor F 2 mw Fixed image feature tensor F 2 f and decoder output feature F 3 D The first, second, and third inputs are respectively fed into a multi-scale displacement fusion module based on cross-neighborhood attention to obtain displacement field increments. As the increment of the output displacement field The displacement field φ3 is upsampled and compared with the displacement field increment. The displacement field φ4 is obtained by adding them together;

[0121] Wherein, the values ​​of W1, H1, and D1 in the dimension of the query vector Q in step 11 are respectively the values ​​of the fixed image feature tensor F. 2 f The dimensions of the query vector Q are width, height, and depth, i.e., w / 4, h / 4, and d / 4; the number of subspaces A in the dimension of the query vector Q is 2 in step 11.

[0122] The values ​​of W2, H2, and D2 in the dimensions of the intermediate feature tensor E in step 11 are respectively the values ​​of the warped feature tensor F. 2mw The dimensions of width, height, and depth, i.e., w / 4, h / 4, and d / 4;

[0123] The values ​​of W3, H3, and D3 in the dimension of the key vector K in step 11 are respectively the values ​​of the warped feature tensor F. 2 mw The dimensions of width, height, and depth, i.e., w / 4, h / 4, and d / 4;

[0124] The values ​​of W4, H4, and D4 in the dimensions of the neighborhood attention NA in step 11 are w / 4, h / 4, and d / 4, respectively.

[0125] The values ​​of W5, H5, and D5 in the dimension of the value vector V in step 11 are w / 4, h / 4, and d / 4, respectively; the 3 in the dimension of the value vector V represents the displacement of the voxel of the floating image in the width, height, and depth directions of the floating image.

[0126] In step 11, the values ​​of W6, H6, and D6 in the intermediate feature tensor F are w / 4, h / 4, and d / 4, respectively; the 3 in the intermediate feature tensor F represents the displacement of the voxels of the floating image in the width, height, and depth directions of the floating image.

[0127] Local displacement field increment The values ​​of W7, H7, and D7 in the dimensions are w / 4, h / 4, and d / 4 respectively in step 11; local displacement field increment The dimension 3 represents the displacement of the voxels of the floating image in the width, height, and depth directions of the floating image.

[0128] Global displacement field increment The values ​​of W8, H8, and D8 in the dimension of F in step 11 are respectively the decoder output features F 3 D The dimensions of width, height, and depth, i.e., w / 4, h / 4, and d / 4; global displacement field increment. The dimension 3 represents the displacement of the voxels of the floating image in the width, height, and depth directions of the floating image.

[0129] W 1 and W 2 In step 11, the value can be set to 0.5.

[0130] Step 12: Upsample the displacement field φ4 so that the width, height, and depth of the displacement field φ4 are related to the floating image feature tensor F. 1 m The width, height, and depth of the floating image feature tensor F are the same. 1m Perform the twisting to obtain the twisted feature tensor F. 1 mw The twisted feature tensor F 1 mw Fixed image feature tensor F 1 f ST output features F 0 st and decoder output feature F 3 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 4 D As output decoder features, the decoder output feature F 4 D Perform convolution processing to obtain the displacement field increment. After upsampling, the displacement field φ4 is compared with the displacement field increment. The displacement field φ5 is obtained by adding them together.

[0131] Step 13: Upsample the displacement field φ5 so that the width, height, and depth of the displacement field φ5 are related to the floating image feature tensor F. 0 m The width, height, and depth of the floating image feature tensor F are the same. 0 m Perform the twisting to obtain the twisted feature tensor F. 0 mw The twisted feature tensor F 0 mw Fixed image feature tensor F 0 f ST output features F 0 st and decoder output feature F 4 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 5 D As output decoder features, the decoder output feature F 5 D Perform convolution processing to obtain the displacement field increment. After upsampling, the displacement field φ5 is compared with the displacement field increment. The displacement field φ6 is obtained by adding the two images together. The floating image is then distorted by the displacement field φ6 to obtain the target distorted image.

[0132] This invention employs an incremental superposition method of displacement field to achieve a gradual learning process of displacement field from low resolution to high resolution, thereby improving the convergence speed of the registration model and enhancing the deformation effect of images with large-scale deformation.

[0133] Step 14: Minimize the total loss function (total loss function) Including similarity loss between fixed and distorted images and the smoothing loss of the displacement field With the goal of training the registration network model using the training samples from step 1, and saving the model parameters after training, an optimized registration network model is obtained. This process includes the following steps:

[0134] Total loss function Based on the following formula:

[0135]

[0136] Among them, I f It is a fixed image, I m It is a floating image, where φ is the displacement field φ6. It is the similarity loss between the fixed image and the distorted image of the target, where λ is the weighting coefficient. It is the smoothing and regularization loss of the displacement field φ6.

[0137] The similarity loss between the fixed and distorted images is the locally normalized cross-correlation loss (LNCC(I)). f ,I m ,φ), Locally Normalized Cross-Correlation Loss LNCC(I f ,I m ,φ) is based on the following formula:

[0138]

[0139] in, This indicates that the floating image has a local window of size n centered at the voxel at position p. 3 The average voxel value; This indicates that the image is fixed with a local window of size n centered at voxel position p. 3 Mean voxel value; p i Centered on the voxel at position p, with a range of n 3 The position of the neighboring voxels (i is the position number, i∈n) 3 ). This represents the target distortion image obtained after applying a displacement field φ6 to the floating image. This indicates that the target distorts the image at position p. i voxel values; This indicates that the target distorted image is located at voxel p, with a local window size of n. 3 Average voxel value; I f (p i ) indicates that the image is fixed at position p i The voxel value; Ω is the range of voxel positions in a fixed image or a distorted image of a target.

[0140] Smoothing regularization loss of displacement field Based on the following formula:

[0141]

[0142] Where u(p) is the spatial gradient of the displacement field φ6. Ω is the change in displacement of adjacent voxels of voxel p, and Ω is the range of voxel positions in a fixed image or a distorted image of a target.

[0143] Step 15: Input the floating image and the fixed image to be processed into the optimization registration network model to obtain the optimized target distortion image after the floating image is deformed.

[0144] Through steps 1-15, the method of the present invention can effectively fuse displacement fields with long-distance dependency information and displacement fields with local detail dependency information, fuse multi-scale displacement field information, and obtain a more accurate deformation field (i.e., displacement field φ6).

[0145] Table 1 shows the quantitative results obtained using different registration methods (VoxelMorph, TransMorph, TransMatch, and the method of this invention). The data is the LPBA40 public dataset. DSC refers to the Dice Similarity Coefficient. A higher DSC value indicates better registration performance.

[0146] Table 1 shows the quantitative results obtained using different registration methods.

[0147] DSC 64.3±3.5 67.0±3.3 68.7±1.8 72.4±1.4

[0148] It should be noted that the embodiments described in this invention are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains can make various modifications or additions to the described embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.

Claims

1. An image registration method based on a cross-neighborhood attention enhancement mechanism, characterized in that, Includes the following steps: Step 1: Obtain multiple images from the public dataset to form a training image set. Take two images from the training image set as a pair as a group of training samples. In each group of training samples, take one image as a floating image and the other image as a fixed image. Step 2: Build a registration network model. The registration network model includes a three-branch encoder, a multi-scale displacement fusion module based on cross-neighborhood attention, and a decoder module. The three-branch encoder includes a four-stage Swin-Transformer module and a multi-scale feature tensor extraction module. Step 3: Input the floating image into the multi-scale feature tensor extraction module to obtain the floating image feature tensor F at multiple scales. k m For k=0,...,5, a fixed image is input into a multi-scale feature tensor extraction module to obtain fixed image feature tensors F at multiple scales. v f k and v are both scale numbers, v=0,...,5; Step 4: Stitch the floating image and the fixed image together, then encode the stitched image and input it into the four-stage Swin-Transformer module to obtain the ST output features F of each stage. j st j is the stage number, j=0,...,3; Step 5: Convert the floating image feature tensor F 5 m Fixed image feature tensor F 5 f and ST output feature F 3 st These are respectively used as the first, second, and third inputs to a multi-scale displacement fusion module based on cross-neighborhood attention to obtain the displacement field increment. As the increment of the output displacement field Increment of displacement field Adding it to the initial displacement field yields the displacement field. The initial displacement field is a set value; Step 6: Displacement field Upsampling makes the displacement field The width, height, and depth in the dimensions of the floating image feature tensor F 4 m The width, height, and depth of the floating image feature tensor F are the same. 4 m Perform the twisting to obtain the twisted feature tensor F. 4 mw Twisted feature tensor F 4 mw The width, height, and depth in the dimensions of the floating image feature tensor F 4 m The width, height, and depth of the dimension are the same, and then the distorted feature tensor F is... 4 mw Fixed image feature tensor F 4 f ST output features F 2 st and ST output feature F 3 st These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 1 D As a feature of the output decoder; Step 7: Convert the distorted feature tensor F 4 mw Fixed image feature tensor F 4 f and decoder output feature F 1 D These are respectively used as the first, second, and third inputs to a multi-scale displacement fusion module based on cross-neighborhood attention to obtain the displacement field increment. As the increment of the output displacement field , displacement field After upsampling, and compared with the displacement field increment Adding them together yields the displacement field. ; Step 8: Displacement field After upsampling, the displacement field is... The width, height, and depth in the dimensions of the floating image feature tensor F 3 m The width, height, and depth of the floating image feature tensor F are the same. 3 m Perform the twisting to obtain the twisted feature tensor F. 3 mw Twisted feature tensor F 3 mw The width, height, and depth in the dimensions of the floating image feature tensor F 3 m The width, height, and depth of the distorted feature tensor F are the same. 3 mw Fixed image feature tensor F 3 f ST output features F 1 st and decoder output feature F 1 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 2 D As a feature of the output decoder; Step 9: Convert the distorted feature tensor F 3 mw Fixed image feature tensor F 3 f and decoder output feature F 2 D These are respectively used as the first, second, and third inputs to a multi-scale displacement fusion module based on cross-neighborhood attention to obtain the displacement field increment. As the increment of the output displacement field , displacement field After upsampling, and compared with the displacement field increment Adding them together yields the displacement field. ; Step 10: Displacement field Upsampling makes the displacement field The width, height, and depth in the dimensions of the floating image feature tensor F 2 m The width, height, and depth of the floating image feature tensor F are the same. 2 m Perform the twisting to obtain the twisted feature tensor F. 2 mw Twisted feature tensor F 2 mw The width, height, and depth in the dimensions of the floating image feature tensor F 2 m The width, height, and depth of the distorted feature tensor F are the same. 2 mw Fixed image feature tensor F 2 f ST output feature F 0 st and decoder output feature F 2 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 3 D As a feature of the output decoder; Step 11: Convert the distorted feature tensor F 2 mw Fixed image feature tensor F 2 f and decoder output feature F 3 D These are respectively used as the first, second, and third inputs to a multi-scale displacement fusion module based on cross-neighborhood attention to obtain the displacement field increment. As the increment of the output displacement field , displacement field After upsampling, and compared with the displacement field increment Adding them together yields the displacement field. ; Step 12: Displacement field Upsampling makes the displacement field The width, height, and depth in the dimensions of the floating image feature tensor F 1 m The width, height, and depth of the floating image feature tensor F are the same. 1 m Perform the twisting to obtain the twisted feature tensor F. 1 mw The twisted feature tensor F 1 mw Fixed image feature tensor F 1 f ST output features F 0 st and decoder output feature F 3 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 4 D As output decoder features, the decoder output feature F 4 D Perform convolution processing to obtain the displacement field increment. Displacement field After upsampling, and compared with the displacement field increment Adding them together yields the displacement field. ; Step 13: Displacement field Upsampling makes the displacement field The width, height, and depth in the dimensions of the floating image feature tensor F 0 m The width, height, and depth of the floating image feature tensor F are the same. 0 m Perform the twisting to obtain the twisted feature tensor F. 0 mw The twisted feature tensor F 0 mw Fixed image feature tensor F 0 f ST output features F 0 st and decoder output feature F 4 D These four inputs, respectively, are used as the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input, and are fed into the decoder module to obtain the decoder output feature F. 5 D As output decoder features, the decoder output feature F 5 D Perform convolution processing to obtain the displacement field increment. Displacement field After upsampling, and compared with the displacement field increment Adding them together yields the displacement field. The floating image is passed through the displacement field Distort to obtain a distorted image of the target image; Step 14: Minimize the total loss function With the goal of training the registration network model using the training samples in step 1, the model parameters are saved after training to obtain the optimized registration network model. Total loss function Based on the following formula: In the formula, It is a fixed image. It is a floating image. It is a displacement field , It is the similarity loss between the fixed image and the distorted image of the target. These are weighting coefficients. It is a displacement field Smoothing regularization loss; Step 15: Input the floating image and the fixed image to be processed into the optimization registration network model to obtain the optimized target distortion image after the floating image is deformed.

2. The image registration method based on the cross-neighborhood attention enhancement mechanism according to claim 1, characterized in that, Step 3 specifically includes the following steps: Step 3.1: Input the floating image into the multi-scale feature tensor extraction module. The multi-scale feature tensor extraction module performs multiple feature extraction operations on the floating image. Each feature extraction operation includes convolution and average pooling to obtain the floating image feature tensor F at multiple scales. k m Floating image feature tensor F k m The dimensions include the floating image feature tensor F. k m The number of channels, floating image feature tensor F k m The width of the floating image feature tensor F k m The high, and floating image feature tensor F k m The depth; Step 3.2: Input the fixed image into the multi-scale feature tensor extraction module. The multi-scale feature tensor extraction module performs multiple feature extraction operations on the fixed image. Each feature extraction operation includes convolution and average pooling to obtain the fixed image feature tensor F at multiple scales. v f Fixed image feature tensor F v f The dimensions include a fixed image feature tensor F. v f The number of channels, with a fixed image feature tensor F v f The width of the fixed image feature tensor F v f The high, and the fixed image feature tensor F v f The depth.

3. The image registration method based on the cross-neighborhood attention enhancement mechanism according to claim 2, characterized in that, Step 4 specifically includes the following steps: Step 4.1: After concatenating the floating image and the fixed image along the channel dimension, the images are encoded to obtain a block feature map; Step 4.2: After linear mapping, the block feature maps are transformed into feature map tensors. These feature map tensors are then input into the four-stage Swin-Transformer module to obtain the ST output features F for each stage. j st ST output feature F j st The dimensions include ST output features F j st The number of channels, ST output features F j st The width of ST output feature F j st The high, and ST output feature F j st The depth.

4. The image registration method based on the cross-neighborhood attention enhancement mechanism according to claim 3, characterized in that, The dimensions of both the floating image and the fixed image are (1, w, h, d), where 1 is the number of channels of the feature of the floating image or the fixed image, w is the width of the floating image or the fixed image, h is the height of the floating image or the fixed image, and d is the depth of the floating image or the fixed image. In step 3.1, the floating image feature tensor F 0 m ~Floating image feature tensor F 5 m The dimensions are (2, w, h, d), (4, w / 2, h / 2, d / 2), (8, w / 4, h / 4, d / 4), (16, w / 8, h / 8, d / 8), (32, w / 16, h / 16, d / 16), (32, w / 32, h / 32, d / 32); In step 3.2, the image feature tensor F is fixed. 0 f ~Fixed image feature tensor F 5 f The dimensions are (2, w, h, d), (4, w / 2, h / 2, d / 2), (8, w / 4, h / 4, d / 4), (16, w / 8, h / 8, d / 8), (32, w / 16, h / 16, d / 16), (32, w / 32, h / 32, d / 32); In step 4.2, the ST output feature F 0 st ~ST Output Feature F 3 st The dimensions are (c, w / 4, h / 4, d / 4), (2c, w / 8, h / 8, d / 8), (4c, w / 16, h / 16, d / 16), and (8c, w / 32, h / 32, d / 32); Where c is the number of channels in the feature map tensor.

5. The image registration method based on the cross-neighborhood attention enhancement mechanism according to claim 3, characterized in that, The first, second, and third inputs are fed into a multi-scale displacement fusion module based on cross-neighborhood attention to obtain the output displacement field increment. ,specific Includes the following processes: First, the second input is sequentially processed through linear mapping, dimensional expansion, and transpose to obtain the query vector Q. The first input is sequentially processed through linear mapping, dimensional expansion, and transpose to obtain the intermediate feature tensor G. The width, height, and depth dimensions of the intermediate feature tensor G are then filled to obtain the intermediate feature tensor E. Finally, the intermediate feature tensor E is divided into blocks with a step size of 1 and a size of 3×3×3 to obtain the key vector K. Secondly, the query vector Q is multiplied by the key vector K to obtain the neighborhood attention NA; Next, construct a value vector V, transpose V, and multiply it with the neighborhood attention NA to obtain an intermediate feature tensor F. Compress the dimensions of the intermediate feature tensor F to obtain the local displacement field increment. The global displacement field increment is obtained by convolution processing the third input. ; Finally, the local displacement field increment and global displacement field increment The output displacement field increment is obtained by weighting using the following formula. : = In 1 × + In 2 × In the formula, W 1 and W 2 These are the local displacement field increments. and global displacement field increment The weight.

6. The image registration method based on the cross-neighborhood attention enhancement mechanism according to claim 5, characterized in that, The dimension of the query vector Q is (W1, H1, D1, 1, Am), where W1, H1, and D1 are the width, height, and depth of the query vector Q, respectively; the values ​​of W1, H1, and D1 are the width, height, and depth of the second input, respectively; 1 is the vector length; m is the vector subspace dimension; and A is the number of subspaces. The intermediate feature tensor E has dimensions (W2+2, H2+2, D2+2, 1, Am), where W2, H2, and D2 are the width, height, and depth of the first input dimension, respectively. The dimension of the key vector K is (W3, H3, D3, 27, Am), where W3, H3, and D3 are the width, height, and depth of the key vector K, respectively; the values ​​of W3, H3, and D3 are the width, height, width, and depth of the first input, respectively; and 27 is the total vector length of the 27 blocks obtained after processing with a size of 3×3×3. The dimensions of the neighborhood attention NA are (W4, H4, D4, 27, 1), where W4, H4, and D4 are the width, height, and depth of the neighborhood attention NA, respectively; the values ​​of W4, H4, and D4 are the width, height, and depth of the query vector Q or key vector K, respectively; 27 is the total vector length of the 27 blocks obtained after processing with a size of 3×3×3; and the attention value of each block is 1. The dimensions of the value vector V are (W5, H5, D5, 27, 3), where W5, H5, and D5 are the width, height, and depth of the value vector V, respectively. The width, height, and depth of the value vector V are the same as the width, height, and depth of the neighborhood attention NA. 27 is the vector length. The displacement of the voxels of the floating image in the width, height, and depth directions of the floating image is 3. The intermediate feature tensor F has dimensions (W6, H6, D6, 3, 1), where W6, H6, and D6 are the width, height, and depth of the intermediate feature tensor F, respectively. The values ​​of W6, H6, and D6 are the width, height, and depth of the value vector V or the neighborhood attention NA, respectively. The voxels of the floating image are displaced by 3 in the width, height, and depth directions of the floating image. Local displacement field The dimensions are (W7, H7, D7, 3), where W7, H7, and D7 are the local displacement fields. The width, height, and depth of the intermediate feature tensor F are represented by the values ​​of W7, H7, and D7, respectively. The voxels of the floating image are displaced by 3 in the width, height, and depth directions of the floating image. Global displacement field The dimensions are (W8, H8, D8, 3), where W8, H8, and D8 are the global displacement field increments, respectively. The width, height, and depth of the third input dimension are represented by W8, H8, and D8, respectively; the voxels of the floating image are displaced by 3 in the width, height, and depth directions of the floating image. The number A of subspaces is 8, 8, 4, and 2 in steps 5, 7, 9, and 11, respectively.

7. The image registration method based on the cross-neighborhood attention enhancement mechanism according to claim 6, characterized in that, The process of inputting the first decoder input, the second decoder input, the third decoder input, and the fourth decoder input into the decoder module to obtain the output decoder features specifically includes the following steps: First, the inputs of the first decoder, the second decoder, and the fourth decoder are concatenated and then upsampled to obtain the output upsampled features; Then, the output upsampled features and the third decoder input are concatenated and processed in two rounds. Each round of processing includes convolution, ReLU activation, and instance normalization to obtain the output decoder features.

8. The image registration method based on the cross-neighborhood attention enhancement mechanism according to claim 7, characterized in that, The similarity loss between the fixed image and the distorted image For locally normalized cross-correlation loss Locally normalized cross-correlation loss Based on the following formula: In the formula, Is it a floating image in position? Centered on a voxel, the local window size is n 3 The average voxel value; Is the image fixed in position? Centered on a voxel, the local window size is n 3 The average voxel value; Based on position Centered on a voxel, with a range of n 3 The position of the neighboring voxels, where i is the position number. n 3 ; Applying a displacement field to a floating image The resulting distorted image of the target The target distorts the image at the location voxel values; The target distorts the image in position Centered on a voxel, the local window size is n 3 The average voxel value; Is the image fixed in position? voxel values; It is the range of voxel positions in a fixed image or a distorted image of a target.

9. The image registration method based on the cross-neighborhood attention enhancement mechanism according to claim 8, characterized in that, The smoothing regularization loss of the displacement field Based on the following formula: In the formula, It is a displacement field Spatial gradient, yes The change in displacement between adjacent voxels of a voxel. It is the range of voxel positions in a fixed image or a distorted image of a target.

Citation Information

Patent Citations

  • Disparity map enhancement method based on RGB and DVS image fusion in high dynamic range scene

    CN112396562A

  • Image registration method based on Swin Transform and CNN double-branch coupling

    CN115082293A