Medical image parallel registration network method based on variable window features
By adopting variable window feature fusion and spatial transformation modules in the medical image registration network, the problem of insufficient image registration accuracy and efficiency in the prior art is solved, and more efficient and accurate image registration is achieved.
Patent Information
- Application Number
- CN202411837125.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing medical image registration methods have shortcomings in accuracy and efficiency, especially the method based on image features is time-consuming, the method based on grayscale is poorly robust, and the method based on deep learning is difficult to represent the global and local features of floating images and fixed images at the same time.
A medical image parallel registration network based on variable window features is adopted. Through dual input and output modules, window scale division modules, window feature fusion modules and spatial transformation modules, the rough features of floating images and fixed images are extracted, and the registration images are generated through variable window feature fusion and spatial transformation.
The spatial dependence of local and global features of the medical image registration network is improved, the registration accuracy and efficiency are improved, and the grayscale images and noise-containing images can be better processed before and after enhancement.
Smart Images

Figure CN119991754A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to medical image registration technology, and is mainly applied to the field of medical radiotherapy. Specifically, it relates to a medical image registration network method based on variable window features. Background Art
[0002] Medical image registration technology is an important technology in medical image processing. It can align medical images of different times or different modalities (such as CT, MRI, etc.) to achieve accurate comparison and analysis. Medical image registration technology plays an important role in medical diagnosis, surgical planning, treatment planning, etc.
[0003] With the rapid development of image registration technology and computer technology, common registration methods in medical image registration include: image feature-based methods, image grayscale-based methods and deep learning-based methods. Image feature-based methods use feature points or feature regions in the image for matching, such as SIFT, SURF, ORB and other algorithms. Grayscale-based registration methods use square difference, absolute difference, mutual information and other methods to measure grayscale similarity. Deep learning-based methods usually use technologies such as convolutional neural networks (CNN) or generative adversarial networks (GAN), which can automatically extract all feature information in medical image data sets, and can be quickly trained and predicted, with high accuracy and high efficiency. However, medical image registration methods based on image features require manual extraction and matching of feature points, which takes a lot of time and effort. Image grayscale-based methods do not consider the spatial dependency between pixels, and the similarity measure is less robust for grayscale images before and after enhancement or images with noise. Existing deep learning-based methods cannot simultaneously represent the global and local features of floating images and fixed images, and cannot provide sufficient registration accuracy.
[0004] In order to solve the problems existing in the above registration algorithm, a medical image registration network based on variable window features is considered to improve the spatial dependency of local features and global features and improve the registration accuracy of the network. Summary of the invention
[0005] Purpose of the invention: The purpose of the present invention is to provide a medical image parallel registration network method based on variable window features.
[0006] Technical solution: The medical image parallel registration network based on variable window features described in the present invention includes a dual-input and output module, a window scale division module, a window feature fusion module, and a space transformation module.
[0007] Step (1): using two U-shaped channels in parallel to connect the floating image and the fixed image to form a dual-channel input;
[0008] Step (2): In the encoder stage, the two inputs of step (1) are respectively used to extract the rough features of the input single image using two window modules, and each window module is composed of a window division module and a search domain division module;
[0009] Step (3): The window division module includes two window division methods: window division and window region division. Window division uses direct division to divide an image into multiple windows of the same size, and each window is called a basic window. The region window division provides a corresponding but larger search window for each basic window.
[0010] Step (4): In the feature fusion module, a set of windows of different sizes is generated using step (3) to learn the relationship between two windows of different sizes, working by calculating the weight (weighting) of the search window relative to the base window. These weights represent the importance of the information in the search window to the base window. The weighted search window is then added to the base window through a residual connection, which allows the base window to fuse information from the search window;
[0011] Step (5): Finally, the output of step (4) is processed to output the final feature, i.e., the deformation field. The deformation field is applied to the floating image, and the registration image is generated through STN, and the similarity is calculated with the fixed image. The spatial transformation module is constructed by including a dissimilarity term related to the local cross-correlation (CC) and a smooth regularization term φ through the objective loss function. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A method flow chart of an embodiment of the present invention;
[0013] Figure 2 This is a model architecture diagram of the present invention;
[0014] Figure 3 It is a schematic diagram of window division of the present invention;
[0015] Figure 4 It is a schematic diagram of cross-feature fusion of the present invention; DETAILED DESCRIPTION
[0016] The invention discloses a medical image parallel registration network method based on variable window features, and the invention is further described below in conjunction with the accompanying drawings and embodiments.
[0017] The invention discloses a medical image parallel registration network method based on variable window features, comprising a dual-input and output module, a window scale division module, a window feature fusion module and a space transformation module.
[0018] First, the registration network includes the following steps:
[0019] Step (1): Preprocess the two input images as follows: 1. Image resampling: To ensure that the two images have the same spatial resolution and dimension, resample the images. For images of the heart, lungs, etc., the size is unified to (x, x, x-16), where x is generally 144 or 208. 2. Grayscale standardization: Different imaging devices or settings may cause the grayscale values of the images to be different. The images are processed uniformly and the grayscale value channels are standardized to [0, 1]. 3. Affine preprocessing: Each pair of moving images and fixed images is pre-aligned by affine transformation using the SyN alignment method. The preprocessed floating images and fixed images are input into two neural networks for feature extraction to form dual-channel input.
[0020] Step (2): using two window modules to extract rough features of the input single image from the two inputs of step (1), each window module is composed of a window division module and a window feature fusion module;
[0021] Step (3): The window division module includes two division methods, namely window division and search area division. The window division adopts a direct division method to divide an image into n windows of the same size. Each window is called a basic window w bi , where i∈[1,n]. The search area division is to provide a corresponding limited range search area w for each basic window si , where i∈[1,n], thus avoiding excessive resource consumption caused by global search. When the image is sent to the window division module, it will be divided into n windows of size h×w×d. At the same time, the search area division generates n search areas of size λh×λw×λd through a sliding window with a stride of 1, where λ is the area coefficient, and the size of the corresponding search area is controlled by controlling its size. Since the two input images are of the same size and the basic window must correspond to the search area window one by one, it is necessary to fill it before dividing the search area to ensure that the search area of the corresponding scale can also be generated at the boundary.
[0022] Step (4): In the window feature fusion module, the windows and search domain sets with different sizes are generated using step (3) to learn the relationship between two windows of different sizes, and perform relationship search from a base window to its corresponding search window. It uses an attention mechanism similar to ViT and SwinTransformer, where the attention calculation formula is as follows:
[0023]
[0024] Among them, Q, K, V are the query matrix, key matrix and value matrix extracted from the image features respectively, and d is the query matrix.
[0025] The dimension of the query matrix is changed. At the same time, it is improved by using the cross transformation method to replace the original multi-head self-attention mechanism, where Q is calculated by the basic window and K, V` are calculated by the search area. Let f a and f b Represents a pair of features in the base window and search domain, with sizes V×C and λ respectively. 3 V×C, where V=h×w×d, the process of generating query, key, and value matrices is as follows:
[0026]
[0027] where i∈[1,n], δ represents the function that reshapes the feature map into the desired form, {w q ,w k ,w v} represents three parameter-independent linear mapping functions, and multi-head operations are applied. The cross transformer block includes a multi-head cross attention module based on different windows and a 2-layer MLP with GELU nonlinearity. Each module is followed by a residual connection, and a LayerNorm (LN) layer is applied to ensure the effectiveness of each layer.
[0028] Step (5): Finally, the output of step (4) is processed to output the final feature, i.e., the deformation field. The deformation field is applied to the floating image, and the registration image is generated by STN, and the similarity is calculated with the fixed image. The objective loss function includes a dissimilarity term related to the local cross-correlation (CC) and a smooth regularization term φ, forming an objective function optimization module. The spatial transformation network (STN) adopts a three-dimensional transformation function, and its bilinear interpolation is defined as follows:
[0029]
[0030] Where p is a voxel, G(φ(p)) represents the 8 neighborhoods of φ(p), is the spatial transformation function.
[0031] The objective loss function includes a dissimilarity term related to the local cross-correlation (CC) and a smooth regularization term φ, which is defined as follows:
[0032]
[0033] Where λ is a hyperparameter, Ω represents the 3D voxel, and CC(A,B) is defined as follows:
[0034]
[0035] where v i Select a patch size of 9×9×9, and denote the local averages of A(v) and B(v), respectively.
Claims
1. A medical image parallel registration network method based on variable window features, characterized in that: The variable window feature fusion parallel network includes a dual input and output module, a window scale division module, a window feature fusion module, and a space transformation module; The specific steps of the registration method are as follows: Step (1): using two U-shaped channels in parallel to connect the floating image and the fixed image to form a dual-channel input; Step (2): In the encoder stage, the two inputs of step (1) are respectively used to extract the rough features of the input single image using two window modules, and each window module consists of a window division module and a feature fusion module; Step (3): The window division module includes two window division methods: window division and window region division. Window division uses direct division to divide an image into multiple windows of the same size, and each window is called a basic window. The region window division provides a corresponding but larger search window for each basic window. Step (4): In the feature fusion module, a set of windows of different sizes is generated using step (3) to learn the relationship between two windows of different sizes, working by calculating the weight (weighting) of the search window relative to the base window. These weights represent the importance of the information in the search window to the base window. The weighted search window is then added to the base window through a residual connection, which allows the base window to fuse information from the search window; Step (5): Finally, the output of step (4) is processed to output the final feature, i.e., the deformation field. The deformation field is applied to the floating image, and the registration image is generated through STN, and the similarity is calculated with the fixed image. The spatial transformation module is constructed by including a dissimilarity term related to the local cross-correlation (CC) and a smooth regularization term φ through the objective loss function.
2. According to the medical image parallel registration network based on variable window features as described in claim 1, it is characterized in that: In step (1), the floating image and the fixed image used as input are CT images, MRI images, or a combination of several of them.
3. According to the medical image registration network based on variable window features described in claim 1, it is characterized in that: In step (2), the window division module includes two window division methods: window division and window region division. Window division adopts direct division to divide an image into multiple windows of the same size, and each window is called a basic window. The region window division provides a corresponding but larger search window for each basic window.
4. According to the medical image registration network based on variable window features described in claim 1, it is characterized in that: In step (3), a set of windows with different sizes is generated using the base window and the region window divided in step (2) to learn the relationship between two windows of different sizes. Inspired by the visual attention mechanism (VIT), the feature fusion module is constructed using the visual attention mechanism. Compared with VIT, this module retains the multi-head self-attention mechanism, which helps to improve the registration network's attention to global information. At the same time, this paper makes some improvements to the traditional VIT to form a new feature fusion module. First, the redundant position embedding is removed after patch segmentation and embedding; second, the convolution module is used to replace the original patch segmentation embedding and linear mapping to directly calculate the weight matrix, which reduces the computational complexity and better meets the needs of 3D image registration; finally, an additional output path is designed. The information flow in the double-layer features is connected by using a convolution module or a deconvolution module with a stride of 2. The input feature X is processed by three different convolution modules to obtain the Query matrix Q, Key matrix K and Value matrix V respectively. The expressions are as follows: Q=W Q ·X;K=W K ·X;V=W V ·X (1) The matrices Q, K, V are reshaped into a set of three-dimensional flat patches: Where B is the mini-batch, C is the number of channels, (P, P, P) is the resolution of each volume patch, and N = HWD / P 3 Represents the number of patches after the resolution of the input feature (H, W, D) is split. Then, by splitting the heads in the embedding channel and swapping the order of the axes, it can be converted into a 4D vector: Where k represents the number of heads, d k =(P 3 C) / k represents the number of embedding channels each head has. Matrix multiplication, scaling, and softmax operations will be performed on Q, K, and V to produce the output Y. The query Q is projected from the base window, while the key K and value V are projected from the search window and multiplied by three parameter-independent linear mapping functions. The window feature fusion applies a multi-head operation. The cross transformer block includes a multi-head cross attention module based on different windows and a 2-layer MLP with GELU nonlinearity. Each module is followed by a residual connection, and a LayerNorm (LN) layer is applied to ensure the effectiveness of each layer.
5. According to the medical image parallel registration network based on variable window attention as described in claim 1, it is characterized in that ,In step (5), the output of step (4) is processed and the final feature, i.e., the deformation field, is output. It is then applied to the floating image, and the registered image is generated through STN, and the loss is calculated with the fixed image.
6. According to claim 6, a medical image registration network based on multi-level connections of Transformer, characterized in that ,The spatial transformation network (STN) adopts a three-dimensional transformation function, and its bilinear interpolation is defined as follows: Where p is a voxel, G(φ(p)) represents the 8 neighborhoods of φ(p), is the spatial transformation function. The objective loss function includes a dissimilarity term related to the local cross-correlation (CC) and a smooth regularization term φ, which is defined as follows: Where λ is a hyperparameter, Ω represents the 3D voxel, and CC(A,B) is defined as follows: where v i Select a patch size of 9×9×9, and denote the local averages of A(v) and B(v), respectively.
Citation Information
Patent Citations
Multi-size window Transform network cloth image registration method and system based on cross attention
CN116934820A
Unsupervised medical image registration method and system based on multiple channels and residual attention mechanism
CN118314167A