Image registration method and system and electronic equipment

By using the Swin-UNet neural network model for image registration, the problems of low registration accuracy and efficiency in the prior art are solved, efficient and accurate image registration is achieved, and the differential homoembryonic characteristics and reversibility of the registration domain are ensured.

CN120107320APending Publication Date: 2025-06-06SHENZHEN COMEN MEDICAL INSTR
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510115673.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has low registration accuracy and efficiency in image registration, and ignores the topological protection and reversibility of the registration domain.

Method used

Image registration is performed using the Swin-UNet neural network model. This model uses symmetric settings of the encoding module and the decoding module to generate registration images using multi-level feature extraction and linear mapping to ensure the differential homoembryonic characteristics of the registration domain.

Benefits of technology

It significantly improves the accuracy and efficiency of image registration, and ensures the differential homoembryonic characteristics and reversibility of the registration results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107320A_ABST
    Figure CN120107320A_ABST
Patent Text Reader

Abstract

The invention provides an image registration method, an image registration system and electronic equipment, which can realize image registration more efficiently and accurately. The method comprises the following steps: acquiring a source image and a target image, processing the source image and the target image, and generating input data of a Swindow-UNet neural network model; performing multi-level feature extraction on the input data by using a plurality of first feature extraction layers in a coding module of the Swindow-UNet neural network model to generate feature representation data; performing feature extraction on the feature representation data by using a plurality of second feature extraction layers in a decoding module of a Swindow-UNet neural network model to generate content feature data; the plurality of second feature extraction layers and the plurality of first feature extraction layers in the coding module are symmetrically arranged; and performing linear mapping based on the content feature data to determine registration domain data, and performing spatial transformation on the registration domain data to generate a registration image. According to the invention, the precision and efficiency of image registration can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an image registration method, system and electronic equipment. Background Art

[0002] In recent years, research on deformable medical image registration at home and abroad can be divided into two categories: traditional registration methods (non-deep learning-based methods) and deep learning-based methods. Traditional registration methods are mainly used to solve the optimization of deformable space. A common method is to use the displacement vector transformation domain (differential homeomorphism domain or registration domain, differential homeomorphism domain is a concept in mathematical differential geometry, which means that there is a reversible smooth mapping, and the inverse mapping is also smooth), mainly including elastic-type model (an image model or deformation model based on elasticity principle), b-splines (B-splines), statistical parameter mapping, Demons (demon algorithm) or others. Registration methods based on deep learning are mainly divided into two methods: using convolutional neural networks to predict the differential homeomorphism transformation domain required for registration; using deep learning networks to calculate the similarity measure of a pair of images, and iteratively optimizing them with traditional registration methods. According to the type of deep learning, it can be divided into two categories: registration based on supervised learning and registration based on unsupervised learning. However, the current related technologies have poor registration accuracy and efficiency for image registration. Summary of the invention

[0003] In view of this, the embodiments of the present invention provide an image registration method, system and electronic device, which can realize image registration more efficiently and accurately.

[0004] In one aspect, an embodiment of the present invention provides an image registration method, comprising:

[0005] Obtain a source image and a target image, process the source image and the target image, and generate input data for a Swin-UNet neural network model;

[0006] Perform multi-level feature extraction on the input data using multiple first feature extraction layers in the encoding module of the Swin-UNet neural network model to generate feature representation data;

[0007] The feature representation data is extracted using multiple second feature extraction layers in the decoding module of the Swin-UNet neural network model to generate content feature data; the multiple second feature extraction layers in the decoding module are symmetrically arranged with the multiple first feature extraction layers in the encoding module;

[0008] Linear mapping is performed based on content feature data to determine registration domain data, and spatial transformation is performed on the registration domain data to generate a registration image.

[0009] The embodiment of the present invention further provides an image registration system, comprising:

[0010] A data input module, used for acquiring a source image and a target image, and processing the source image and the target image to generate input data for a Swin-UNet neural network model;

[0011] An encoding feature extraction module, used to perform multi-level feature extraction on input data using multiple first feature extraction layers in the encoding module of the Swin-UNet neural network model to generate feature representation data;

[0012] A decoding feature extraction module, used to extract features from feature representation data using multiple second feature extraction layers in the decoding module of the Swin-UNet neural network model to generate content feature data; the multiple second feature extraction layers in the decoding module are symmetrically arranged with the multiple first feature extraction layers in the encoding module; and

[0013] The registration image generation module is used to perform linear mapping based on the content feature data to determine the registration domain data, and perform spatial transformation on the registration domain data to generate a registration image.

[0014] An embodiment of the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the image registration method according to the first aspect is implemented.

[0015] It can be seen from the above that the image registration method, system and electronic device provided by the embodiments of the present invention have the following beneficial technical effects:

[0016] A Swin-UNet neural network model is proposed to align the source image with the target image. The Swin-UNet neural network model is a 3D-U type encoding-decoding network based on Transformer. This model combines the powerful capabilities of Swin Transformer and the local information capture characteristics of Unet. Therefore, using the Swin-UNet neural network model for deformable image registration can significantly improve the accuracy and efficiency of image registration compared with the registration methods in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:

[0018] Figure 1 A schematic diagram of an image registration method provided by one or more optional embodiments of the present invention is shown;

[0019] Figure 2A schematic diagram showing the structure of a Swin-UNet neural network model in an image registration method provided by one or more optional embodiments of the present invention;

[0020] Figure 3 A schematic diagram of a method for generating feature representation data in an image registration method provided by one or more optional embodiments of the present invention is shown;

[0021] Figure 4 A schematic diagram showing a method for generating content feature data in an image registration method provided by one or more optional embodiments of the present invention is shown;

[0022] Figure 5 A schematic diagram of the structure of SwinTransformer in an image registration method provided by one or more optional embodiments of the present invention is shown;

[0023] Figure 6 A schematic diagram of a method for generating a registration image in an image registration method provided by one or more optional embodiments of the present invention is shown;

[0024] Figure 7 A schematic diagram of a framework for optimizing the registration of a Swin-UNet neural network model in an image registration method provided by one or more optional embodiments of the present invention is shown;

[0025] Figure 8 A qualitative schematic diagram of experimental results provided by one or more optional embodiments of the present invention is shown; wherein, Figure 8 (a) is the source image, 8(b) is the target image, 8(c)-8(f) are the registered images, transformation fields and Jacobian of the transformations obtained by VoxelMorph (MSE), VoxelMorph-diff, DeepFLASH, SYMNet and Swin-VoxelMorph methods respectively;

[0026] Fig. 9 A schematic diagram of Dice scores of different brain tissues provided by one or more optional embodiments of the present invention;

[0027] Fig.10 A schematic diagram of the structure of an image registration system provided by one or more optional embodiments of the present invention is shown;

[0028] Fig.11 A schematic diagram of the structure of an electronic device provided by one or more optional embodiments of the present invention is shown. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0030] In recent years, research on deformable medical image registration at home and abroad can be divided into two categories: based on traditional registration methods (based on non-deep learning methods) and based on deep learning methods. Traditional registration methods are mainly used to solve the optimization of deformable space. A common method is to use the displacement vector transformation domain (diffeomorphism domain or registration domain, which is a concept in mathematical differential geometry, referring to the existence of a reversible smooth mapping, and the inverse mapping is also smooth), mainly including elastic-type models, b-splines, statistical parameter mapping, Demons or other discrete methods. Since the differential homeomorphism domain (registration domain) must guarantee some excellent properties, such as topological invariance, a large number of advanced methods have been developed for differential homeomorphism transformation, such as large diffusion pure distance metric mapping (LDDMM), diffeomorphic anatomical registration through exponentiated Lie Algebra (DARTEL), differential homeomorphism Demons and symmetric normalization (SyN). Generally, these methods require a lot of time and computing resources to register a pair of images. Some recent GPU-based iterative algorithms use these frameworks to develop faster algorithms. At the same time, deep learning-based registration methods have significantly improved the running speed and are mainly divided into two methods: (1) using convolutional neural networks to predict the differential homeomorphic transformation domain required for registration; (2) using deep learning networks to calculate the similarity measure of a pair of images and iteratively optimize them using traditional registration methods. According to the type of deep learning, it can be divided into two categories: supervised learning-based registration and unsupervised learning-based registration.

[0031] The models used in related technologies ignore the use of attention mechanisms to resolve long-range inter-image correlations in embedding learning, which limits the method's ability to recognize the communication of meaningful semantic information about anatomical structures; they also ignore the topology preservation and reversibility of the registration domain, although they can achieve fast image registration. The topology preservation and reversibility of the registration domain mean that the registration transformation itself is smooth, there is a reversible smooth mapping, and the inverse mapping is also smooth. The above problems lead to poor registration accuracy and efficiency in image registration in related technologies.

[0032] In view of the above problems, the purpose of an embodiment of the present invention is to propose an image registration method, which uses a neural network model with a symmetrical unsupervised learning framework to perform feature learning on the source image and the target image to minimize the dissimilarity between the images and simultaneously estimate the forward and reverse transformations, thereby achieving image registration more efficiently and accurately.

[0033] Based on the above objectives, in one aspect, an embodiment of the present invention provides an image registration method.

[0034] like Figure 1 As shown, one or more optional embodiments of the present invention provide an image registration method, comprising:

[0035] S1: Obtain a source image and a target image, process the source image and the target image, and generate input data for the Swin-UNet neural network model.

[0036] The purpose of image registration is to use a certain method to optimally map one or more images (partially) to the target image. Usually, one image (source image, Moving Image) can be mapped to another image (target image, Fixed Image) to obtain a registered image pair (Moved Image). First, the source image M and the target image F can be obtained, and the source image and the target image can be processed using a neural network model.

[0037] refer to Figure 2 As shown in FIG. 1 , it is a schematic diagram of the structure of the Swin-UNet neural network model. The source image data and the target image data can be processed using the PatchPartition processing layer. Specifically, after the source image and the target image are acquired, the source image and the target image are pixel-matrixed to generate source matrix data and target matrix data corresponding to the source image and the target image, and then the source matrix data and the target matrix data are block-partitioned (PatchPartition) to obtain block matrix data, which is used as input data for processing using the Swin-UNet neural network model.

[0038] In some optional implementations, the source image M and the target image F can be concatenated into a 2-channel 3D image as the input of the Patch Partition layer, where the input 2-channel 3D image is divided into non-overlapping patches of size 4x4x4 and the input is transformed into a sequence embedding. Through this separation method, the feature dimension of each patch is 128.

[0039] The overall data dimension of the source image M and the target image F spliced ​​into a 2-channel 3D image data can be expressed as:

[0040] D×H×W×2

[0041] The overall data dimension after Patch Partition layer block segmentation can be expressed as:

[0042]

[0043] That is, after block segmentation, input data with a feature dimension of 128 is generated.

[0044] S2: Use multiple first feature extraction layers in the encoding module of the Swin-UNet neural network model to perform multi-level feature extraction on the input data to generate feature representation data.

[0045] A plurality of first feature extraction layers may be set in the encoding module (Encoder) of the Swin-UNet neural network model, and the plurality of first feature extraction layers are used to perform multi-level feature extraction on the input data, thereby determining hierarchical feature representation data.

[0046] like Figure 3 As shown, in an image registration method provided by one or more optional embodiments of the present invention, a plurality of first feature extraction layers in an encoding module of a Swin-UNet neural network model are used to perform multi-level feature extraction on input data to generate feature representation data, including:

[0047] S201: In the encoding module of the Swin-UNet neural network model, a linear embedding layer is used to map the input data into preset dimensional data.

[0048] refer to Figure 2 As shown, the input data can first be processed by a linear embedding layer to perform feature mapping, thereby mapping the input data into data of a preset dimension (characterized as C).

[0049] The linear embedding layer can map the feature dimension of the input matrix data to any dimension (represented as C). The overall data dimension after the feature dimension conversion by the linear embedding layer can be expressed as:

[0050]

[0051] It is understandable that the corresponding characteristic dimensions of the generated preset dimension data can be flexibly set according to actual conditions.

[0052] Embodiments of the present invention provide a learnable linear embedding layer to predict the registration domain using low-level spatial features and enhancement of high-level correlations.

[0053] S202: Perform multi-level feature extraction on preset dimensional data using a plurality of first feature extraction layers to generate feature representation data.

[0054] The first feature extraction layer may include a Swin transformation layer (Swin-Transformer) and a feature fusion layer (PatchMerging). When the first feature extraction layer is used to extract features from input data (specifically, it may be preset dimensional data obtained by mapping the input data using a linear embedding layer), the Swin-Transformer layer is first used to perform feature representation learning on the input data (preset dimensional data), and the feature fusion layer is used to downsample and increase the dimension of the data generated by the feature representation learning.

[0055] refer to Figure 2 As shown, in some optional embodiments, three first feature extraction layers are provided in the encoding module (Encoder) of the Swin-UNet neural network model. Each first feature extraction layer includes a Swin transformation layer (Swin-Transformer) and a feature fusion layer (Patch Merging). Through multiple first feature extraction layers, feature extraction is performed on the preset dimensional data step by step and the data dimension is increased, so that hierarchical feature representation data can be obtained.

[0056] After being processed by the three first feature extraction layers, the overall data dimensions of the corresponding output data can be expressed as:

[0057]

[0058] Through the multi-level feature extraction processing of the three first feature extraction layers in the encoding module (Encoder) of the Swin-UNet neural network model, hierarchical feature representation data can be generated.

[0059] S3: Use multiple second feature extraction layers in the decoding module of the Swin-UNet neural network model to extract features from the feature representation data to generate content feature data; the multiple second feature extraction layers in the decoding module are symmetrically arranged with the multiple first feature extraction layers in the encoding module.

[0060] In some optional implementations, a Swin transformation layer (Swin-Transformer) can also be set between the encoding module (Encoder) and the decoding module (Decoder) of the Swin-UNet neural network model as a connection module (Bottleneck) between the encoding module (Encoder) and the decoding module (Decoder).

[0061] A plurality of second feature extraction layers may be set in the decoding module (Decoder) to extract content from the hierarchical feature representation data output by the Encoder module. The plurality of second feature extraction layers in the Decoder module are symmetrically arranged with the plurality of first feature extraction layers in the Encoder module.

[0062] like Figure 4 As shown, in an image registration method provided by one or more optional embodiments of the present invention, multiple second feature extraction layers in a decoding module of a Swin-UNet neural network model are used to extract features from feature representation data to generate content feature data, including:

[0063] S301: In a decoding module of the Swin-UNet neural network model, multiple second feature extraction layers are used to extract features from feature representation data to generate content feature data.

[0064] The second feature extraction layer may include a Swin transformation layer (Swin-Transformer) and a feature expansion layer (PatchExpanding). When the second feature extraction layer is used to extract features from feature representation data, the PatchExpanding layer is first used to upsample and reduce the dimension of the feature representation data, and the Swin-Transformer layer is used to perform feature representation learning on the data processed by the feature expansion layer. Among them, the Patch Expanding layer is used to reorganize the feature maps of adjacent dimensions into a larger feature map with a 2x upsampling resolution.

[0065] refer to Figure 2As shown, in some optional embodiments, three second feature extraction layers are provided in the decoding module (Decoder) of the Swin-UNet neural network model. Each second feature extraction layer includes a feature expansion layer (PatchExpanding) and a Swin transformation layer (Swin-Transformer). Through multiple second feature extraction layers, the feature representation data is dimensionally reduced step by step and content feature data is extracted.

[0066] After being processed by the three second feature extraction layers, the overall data dimensions of the corresponding output data can be expressed as:

[0067]

[0068] The embodiment of the present invention provides a novel symmetric unsupervised deep learning framework Swin-VoxelMorph, in which the symmetry is mainly reflected in the first feature extraction layer and the second feature extraction layer.

[0069] It should be further explained that in the Swin-UNet neural network model, a skip connection (SkipConnection) is used to perform data fusion between the multiple first feature extraction layers of the encoding module (Encoder) and the multiple second feature extraction layers of the decoding module (Decoder), thereby compensating for the loss of spatial information caused by downsampling.

[0070] Based on the symmetrical arrangement structure of multiple second feature extraction layers in the Decoder module and multiple first feature extraction layers in the Encoder module, along the data processing transmission direction, the content feature data extracted and determined by the first second feature extraction layer in the Decoder module is fused with the feature representation data extracted and determined by the third first feature extraction layer in the Encoder module using a jump link, which can be expressed as:

[0071] Skip Connection

[0072] 1 / 16

[0073] The second second feature extraction layer is jump-linked to the second first feature extraction layer, and the third second feature extraction layer is jump-linked to the first first feature extraction layer, which can be expressed as:

[0074]

[0075] The corresponding content feature data of the second feature extraction layer is fused with the corresponding feature representation data of the first feature extraction layer through skip connections, so as to make up for the spatial information loss caused by multiple downsampling in the encoder module in the decoder module. In Swin-Unet, the encoder part adopts the Swin Transformer structure with a shift window to extract context information. Each Swin Block contains a local perception layer and a global perception layer, and uses local windows and global windows for self-attention calculation respectively, realizing the fusion of global and local context information. The decoder part adopts a structure based on the symmetric Swin Transformer to perform patch upsampling operations to restore the spatial resolution of the feature map. At the same time, the upsampling module restores the encoder's feature map to its original size through upsampling operations and fuses it with the corresponding skip connection features, which helps to restore details and edges.

[0076] In addition, Swin-Unet uses skip-connections to simultaneously learn local and global semantic features. This structure enables the model to perform well in medical image segmentation tasks.

[0077] The deep learning framework provided by the embodiment of the present invention is based on Swin-Unet and includes a multi-layer (for example, 3 layers) encoder-decoder with skip connections.

[0078] like Figure 5 As shown, in an image registration method provided by one or more optional embodiments of the present invention, the Swin Transformer layer is composed of a shifted window based MSA with two layers of multi-layer perceptrons (MLP). A normalization (LayerNorm) layer is used before each multi-head self-attention mechanism (MSA) module and each multi-layer perceptron (MLP), and a residual connection is used after each MSA and MLP.

[0079] S302: Perform feature mapping recovery on content feature data using the block extension layer.

[0080] After the content feature data is extracted and determined by multiple second feature extraction layers, the feature map of the content feature data is restored using the block expansion layer (PatchExpanding). The block expansion layer can perform 4x upsampling to restore the resolution of the feature map, thereby restoring the content feature data to the same dimension (DxWxH) as the source image and the target image.

[0081] S4: Perform linear mapping based on the content feature data to determine the registration domain data, and perform spatial transformation on the registration domain data to generate a registration image.

[0082] like Figure 6 As shown, in an image registration method provided by one or more optional embodiments of the present invention, linear mapping is performed based on content feature data to determine registration domain data, and spatial transformation is performed on the registration domain data to generate a registration image, including:

[0083] S401: Processing the content feature data after feature mapping restoration using a linear mapping layer to generate registration domain data.

[0084] refer to Figure 2 As shown in FIG. 1 , for the content feature data determined by the Decoder layer, the linear projection layer (LinearProjection) can be used to process the content feature data after feature mapping is restored to generate the registration domain data φ. The linear projection layer is used to perform up-sampling feature estimation on the content feature data to generate the registration domain data. The registration domain data includes the forward registration domain and the reverse registration domain (φ MF and φ FM ).

[0085] S402: Using a differentiable spatial transformation function to perform spatial transformation processing on the registration domain data to generate a registration image.

[0086] For the registration domain data, a differentiable spatial transformation function can be used to perform spatial transformation (SpatialTransform) on it to generate a registered image. Corresponding to the source registration domain and the target registration domain (φ MF and φ FM ), and the source registration image and the target registration image (φ MF (M),φ FM (F)).

[0087] Image registration method, Swin-UNet neural network model is proposed to process source image and target image. Swin-UNet neural network model is a 3D-U type encoding-decoding network based on Transformer, and SwinTransformer module is used for deformable medical image registration. The encoding module is used to extract multi-level feature content learning of source image and target image data, and the decoding module symmetrically set with the encoding module is used to learn multi-level content features. Low-level spatial features and high-level related enhancements are used to predict the registration domain, which can minimize the dissimilarity between images and estimate the forward and reverse transformations at the same time. Therefore, the Swin-UNet neural network model can be used to efficiently and accurately predict and determine the registration domain data, and the registered image is obtained after the spatial transformation of the registration domain data. The registration result can guarantee the optimal differential homeomorphism characteristics.

[0088] One or more optional embodiments of the present invention provide an image registration method, which, after generating a registered image using a Swin-UNet neural network model, also performs registration optimization on the Swin-UNet neural network model based on registration domain data and the registered image.

[0089] The optimization problem of registration can be expressed as:

[0090] L(I 1 ,I 2 ,φ)=L sim (I 1 ,I 2 (φ))+L reg (φ)

[0091] Among them, I 1 ,I 2 represent the two images to be registered, φ represents the registration domain data (such as the differential homeomorphism domain), L sim represents the image similarity function, L reg (φ) represents the regularization process applied to the registration domain data.

[0092] The optimization problem of image registration aims to minimize the registration domain data (φ) while maintaining a smooth displacement vector transformation domain.

[0093] The loss function for registration optimization of the Swin-UNet neural network model is:

[0094] L total =L sim-pair +λ 1 L smooth +λ 2 L Jdet +λ 3 L cons-pair

[0095] Among them, L sim-pair is the similarity term, which is used to measure the minimum mean square error between the two images to be registered.

[0096] Similarity Item:

[0097] L sim-pair (F,M,φ)=L sim (F,M(φ MF ))+L sim (M,F(φ FM ))

[0098] Among them, F, M represent the target image and source image respectively, φ MF represents the registration domain obtained by registering M to F, φ FM represents the registration domain obtained by registering F to M;

[0099] In the embodiment of the present invention, a differentiable space transformation function may be used to align M and F respectively.

[0100] L smooth is a smoothing term used to measure the regularization of the diffeomorphic domain.

[0101] Smooth Item:

[0102]

[0103] Where p represents the voxel position, Ω represents the three-dimensional voxel, Represents gradient calculation;

[0104] L Jdet is a local directional consistency restriction term, which is used to impose local directional consistency restrictions on the estimated φ to ensure topology-preserving transformation.

[0105] Local directional consistency constraints:

[0106]

[0107] Where J(φ) represents the Jacobian of the registration domain, σ represents the ReLU function, and N represents the total number of pixels;

[0108] L cons-pair is a reversible consistency constraint, which means that the estimated bidirectional transformation between image pairs should share the same path, which means that the combination of forward and reverse transformations should be the unit cell or close to the unit cell.

[0109] Reversible consistency restrictions:

[0110]

[0111] Among them, φ 0 represents the unit orthogonal lattice, represents the combination of the forward registration domain and the backward registration domain;

[0112] λ 1 ,λ 2 ,λ 3 Represents adjustment parameters, which can be fine-tuned using grid search.

[0113] The loss function provided by the embodiment of the present invention is symmetric and reversible.

[0114] like Figure 7 , which is a schematic diagram of a framework for optimizing the registration of a Swin-UNet neural network model in an image registration method provided by one or more optional embodiments of the present invention.

[0115] refer to Figure 7 As shown in FIG. 1 , the overall process of the image registration method includes:

[0116] The source image M (Moving Image) and the target image F (Fixed Image) are used as the input data of the Swin-UNet neural network model, and the Swin-UNet neural network model is used to predict and generate the registration domain data φ MF and φ FM , and then through spatial transformation (Spatial Transform) processing, the corresponding registration image pair φ is generated MF (M),φ FM (F). Then, based on M, F, φ MF ,φ FM ,φ MF (M),φ FM (F) is used to calculate and determine the loss function of the Swin-UNet neural network model, thereby optimizing and updating the Swin-UNet neural network model.

[0117] Loss function:

[0118] L total =L sim-pair +λ 1 L smooth +λ 2 L Jdet +λ 3 L cons-pair

[0119] First, we can compare M and φ FM (F) to determine L sim (M,F(φ FM )), compare F and φ MF (M) to determine Lsim (F,M(φ MF )), and then calculate and determine L sim-pair ; For φ MF With φ FM , and the corresponding gradient operation can be performed to determine L smooth ; For φ MF With φ FM , respectively determine the corresponding Jacobian determinant, and we can determine L Jdet ; For φ MF With φ FM , we can also calculate and determine L cons-pair .

[0120] In the image registration method, the objective function for optimizing the Swin-UNet neural network model includes directional and reversible consistency restrictions, and the loss function comprehensively measures four aspects: similarity, smoothness, directional consistency and reversible consistency. The Swin-UNet neural network model is trained and optimized based on the objective function and the loss function, which can ensure the topological preservation and reversible consistency of the predicted transformation of the Swin-UNet neural network model, thereby determining the efficiency and accuracy of the results of image registration using the Swin-UNet neural network model.

[0121] In some optional embodiments, the performance of the image registration method is verified.

[0122] Among them, data set and preprocessing:

[0123] It was validated on two datasets, ADNI and PPMI, including 1961 T1-weighted brain magnetic resonance imaging (MRI). They were divided into 1569, 196 and 196 (8:1:1) volumes for training, validation and test sets, respectively. The standard preprocessing steps of structural brain MRI include skull removal, resampling and affine space normalization using FreeSurfer, and cropping to 160x192x224. Segmentation maps including 29 anatomical structures can be obtained with FreeSurfer and used for evaluation.

[0124] Evaluation criteria: The Dice score is used to measure the anatomical structure overlap results, and the number of Jacobian determinants ≤ 0 reflects the differential homeomorphism of the registration domain.

[0125] Experimental settings: Based on Pytorch and using Adam optimizer, the learning rate is 0.0001. Set epochs (all samples in the training set are trained once as one epoch) to 1500, batch size to 1, and the step size of each epoch to 100. Use grid search to fine-tune the parameter λ 1 =0.01,λ 2 =1000,λ 3 =10.

[0126] The experimental results are shown in Table 1. Figure 8 and Fig. 9 In Table 1, Affine only means that only affine transformation is used, ANTs SyN (CC) means a registration tool, VoxelMorph (MSE), VoxelMorph-diff, DeepFLASH, and SYMNet are all image registration methods in the related art, Swin-VoxelMorph is a registration method provided by an embodiment of the present invention, Avg.Dice means the average dice coefficient, GPU sec means the running time of GPU, and CPU sec means the running time of CPU. Figure 8 Schematic diagram for comparing the registration of two magnetic resonance (MR) slices using different registration methods, where 8(a) is the source image, 8(b) is the target image, 8(c)-8(f) are the registered images, transformation fields, and Jacobian determinants of the transformations obtained by VoxelMorph (MSE), VoxelMorph-diff, DeepFLASH, SYMNet, and Swin-VoxelMorph methods, respectively. Fig. 9 The Dice scores of ANTs SyN, VoxelMorph (MSE), SYMNet, and Swin-VoxelMorph on anatomical structures are shown. The left and right hemispheres are merged into one structure for visualization. The figure contains the brainstem (BS), thalamus (Th), cerebellar cortex (CblmC), lateral ventricle (LV), cerebellar white matter (CblmWM), lentiform nucleus (Pu), cerebral white matter (CeblWM), deep ventral nucleus (VDC), caudate nucleus (Ca), globus pallidus (Pa), hippocampus (Hi), third ventricle (3V), fourth ventricle (4V), amygdala (Am), cerebrospinal fluid (CSF), cerebral cortex (CeblC), and choroid plexus (CP).

[0127] Table 1

[0128]

[0129] It can be seen from the experimental verification results that, compared with several other image registration methods, the image registration method provided by the embodiment of the present invention achieves the optimal average Dice score and produces the best differential homeomorphism registration domain (fewer non-positive Jacobian determinants), that is, the registration result of the embodiment of the present invention can guarantee the optimal differential homeomorphism characteristics.

[0130] It should be noted that the method of one or more embodiments of the present invention may be performed by a single device, such as a computer or server. The method of this embodiment may also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices may only perform one or more steps in the method of one or more embodiments of the present invention, and the multiple devices may interact with each other to complete the method.

[0131] It should be noted that the above describes specific embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0132] Based on the same purpose, corresponding to any of the above-mentioned embodiments and methods, an embodiment of the present invention further provides an image registration system.

[0133] refer to Fig.10 , image registration system, including:

[0134] A data input module is used to obtain a source image and a target image, process the source image and the target image, and generate input data for a Swin-UNet neural network model;

[0135] An encoding feature extraction module, used to perform multi-level feature extraction on input data using multiple first feature extraction layers in the encoding module of the Swin-UNet neural network model to generate feature representation data;

[0136] A decoding feature extraction module, used to extract features from feature representation data using multiple second feature extraction layers in the decoding module of the Swin-UNet neural network model to generate content feature data; the multiple second feature extraction layers in the decoding module are symmetrically arranged with the multiple first feature extraction layers in the encoding module; and

[0137] The registration image generation module is used to perform linear mapping based on the content feature data to determine the registration domain data, and perform spatial transformation on the registration domain data to generate a registration image.

[0138] In an image registration system provided by one or more optional embodiments of the present invention, the encoding feature extraction module is also used to map the input data into preset dimensional data using a linear embedding layer in the encoding module of the Swin-UNet neural network model; and perform multi-level feature extraction on the preset dimensional data using multiple first feature extraction layers to generate feature representation data.

[0139] In an image registration system provided by one or more optional embodiments of the present invention, the decoding feature extraction module is also used to extract features from feature representation data using multiple second feature extraction layers in the decoding module of the Swin-UNet neural network model to generate content feature data; and to restore feature mapping of the content feature data using a block expansion layer.

[0140] In an image registration system provided by one or more optional embodiments of the present invention, the first feature extraction layer includes a Swin conversion layer and a feature fusion layer, and the second feature extraction layer includes a Swin conversion layer and a feature expansion layer. The encoding feature extraction module is also used to use the Swin conversion layer to perform feature representation learning on the input data; and use the feature fusion layer to downsample and increase the dimension of the data generated by the feature representation learning. The decoding feature extraction module is also used to use the feature expansion layer to upsample and reduce the dimension of the feature representation data; and use the Swin conversion layer to perform feature representation learning on the data processed by the feature expansion layer.

[0141] In an image registration system provided by one or more optional embodiments of the present invention, the Swin conversion layer is composed of a shifted window based MSA with two layers of multi-layer perceptrons (MLP); a normalization (LayerNorm) layer is used before each multi-head self-attention mechanism (MSA) module and each multi-layer perceptron (MLP), and a residual connection is used after each MSA and MLP.

[0142] In an image registration system provided by one or more optional embodiments of the present invention, the registration image generation module is also used to use a linear mapping layer to process the content feature data after feature mapping is restored to generate registration domain data; and use a differentiable spatial transformation function to perform spatial transformation processing on the registration domain data to generate a registration image.

[0143] In an image registration system provided by one or more optional embodiments of the present invention, the decoding feature extraction module is also used to jump-link the data determined by the extraction of multiple second feature extraction layers with the data determined by the extraction of multiple first feature extraction layers, so as to fuse the content feature data with the feature representation data.

[0144] One or more optional embodiments of the present invention provide an image registration system, further comprising a registration optimization module. The registration optimization module is used to perform registration optimization on the Swin-UNet neural network model based on the registration domain data and the registration image. The registration optimization function used by the registration optimization module is as follows:

[0145] L(I 1 ,I 2 ,φ)=L sim (I 1 ,I 2 (φ))+L reg (φ)

[0146] Among them, I 1 ,I 2 Represent the two images to be registered, φ represents the registration domain data, L sim represents the image similarity function, L reg (φ) represents the regularization process applied to the registration domain data.

[0147] The loss function used by the registration optimization module for registration optimization of the Swin-UNet neural network model is as follows:

[0148] L total =L sim-pair +λ 1 L smooth +λ 2 L Jdet +λ 3 L cons-pair

[0149] Among them, L sim-pair is the similarity term, similarity term:

[0150] L sim-pair (F,M,φ)=L sim (F,M(φ MF ))+L sim (M,F(φ FM ))

[0151] Among them, F, M represent the target image and source image respectively, φ MF represents the registration domain obtained by registering M to F, φ FM represents the registration domain obtained by registering F to M;

[0152] L smooth is a smooth term, a smooth term:

[0153]

[0154] Where p represents the voxel position, Ω represents the three-dimensional voxel, Represents gradient calculation;

[0155] L Jdet is a local directional consistency restriction item, a local directional consistency restriction item:

[0156]

[0157] Where J(φ) represents the Jacobian of the registration domain, σ represents the ReLU function, and N represents the total number of pixels;

[0158] L cons-pair is a reversible consistency restriction item. The reversible consistency restriction item is:

[0159]

[0160] Among them, φ 0 represents the unit orthogonal lattice, φ MF °φ FM represents the combination of the forward registration domain and the backward registration domain;

[0161] λ 1 ,λ 2 ,λ 3 Indicates adjustment parameters.

[0162] For the convenience of description, the above system is described by dividing the functions into various modules. Of course, when implementing one or more embodiments of the present invention, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0163] The system of the above embodiment is used to implement the corresponding method in the above embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0164] Fig.11 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0165] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.

[0166] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solution provided by the embodiment of the present invention is implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0167] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0168] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0169] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0170] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present invention, and does not necessarily include all the components shown in the figure.

[0171] The electronic device of the above embodiment is used to implement the corresponding method in the above embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0172] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute the image registration method of any of the above embodiments.

[0173] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0174] The computer instructions stored in the storage medium of the above embodiments are used to enable a computer to execute the image registration method of any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be described in detail here.

[0175] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memory.

[0176] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0177] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0178] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0179] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0180] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0181] Each embodiment of the present invention is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0182] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. In line with the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present invention as above, which are not provided in detail for the sake of simplicity.

[0183] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0184] One or more embodiments of the present invention are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present invention should be included in the scope of protection of this disclosure.

Claims

1. An image registration method, characterized in that: The method comprises: Acquire a source image and a target image, process the source image and the target image, and generate input data for a Swin-UNet neural network model; Perform multi-level feature extraction on the input data using a plurality of first feature extraction layers in the encoding module of the Swin-UNet neural network model to generate feature representation data; Using a plurality of second feature extraction layers in the decoding module of the Swin-UNet neural network model to extract features from the feature representation data to generate content feature data; the plurality of second feature extraction layers in the decoding module are symmetrically arranged with the plurality of first feature extraction layers in the encoding module; Linear mapping is performed based on the content feature data to determine registration domain data, and spatial transformation is performed on the registration domain data to generate a registration image.

2. The image registration method according to claim 1, characterized in that: The input data is subjected to multi-level feature extraction using a plurality of first feature extraction layers in the encoding module of the Swin-UNet neural network model to generate feature representation data, including: In the encoding module of the Swin-UNet neural network model, the input data is mapped into preset dimensional data using a linear embedding layer; The plurality of first feature extraction layers are used to perform multi-level feature extraction on the preset dimensional data to generate feature representation data.

3. The image registration method according to claim 1, characterized in that: The feature representation data is subjected to feature extraction using multiple second feature extraction layers in the decoding module of the Swin-UNet neural network model to generate content feature data, including: In a decoding module of the Swin-UNet neural network model, a plurality of the second feature extraction layers are used to perform feature extraction on the feature representation data to generate content feature data; The block extension layer is used to perform feature mapping recovery on the content feature data.

4. The image registration method according to any one of claims 1 to 3, characterized in that: The first feature extraction layer includes a Swin conversion layer and a feature fusion layer; The method of performing multi-level feature extraction on the input data using multiple first feature extraction layers in the encoding module of the Swin-UNet neural network model includes: Using the Swin conversion layer to perform feature representation learning on the input data; Using the feature fusion layer to downsample and increase the dimension of the data generated by feature representation learning; The second feature extraction layer includes a Swin conversion layer and a feature expansion layer; The step of extracting features from the feature representation data using a plurality of second feature extraction layers in the decoding module of the Swin-UNet neural network model includes: Using the feature expansion layer to upsample the feature representation data and reduce the dimension; The Swin conversion layer is used to perform feature representation learning on the data processed by the feature expansion layer.

5. The image registration method according to claim 4, characterized in that: The Swin conversion layer consists of a moving window based on a multi-head self-attention mechanism with two layers of multi-layer perceptrons; A normalization layer is used before each multi-head self-attention mechanism module and each multi-layer perceptron, and a residual connection is used after each multi-head self-attention mechanism module and multi-layer perceptron.

6. The image registration method according to claim 3, characterized in that: Performing linear mapping based on the content feature data to determine registration domain data, and performing spatial transformation on the registration domain data to generate a registration image, including: Processing the content feature data after feature mapping restoration using a linear mapping layer to generate registration domain data; A differentiable spatial transformation function is used to perform spatial transformation processing on the registration domain data to generate the registration image.

7. The image registration method according to claim 1, characterized in that: When the method uses a plurality of second feature extraction layers in a decoding module of a Swin-UNet neural network model to extract features from the feature representation data, the method further includes: A plurality of data determined by extraction from the second feature extraction layer is jump-linked with a plurality of data determined by extraction from the first feature extraction layer, so as to perform data fusion of the content feature data and the feature representation data.

8. The image registration method according to claim 1, characterized in that: The registration domain data includes a source registration domain and a target registration domain, and the registration image includes a source registration image and a target registration image; After generating the registration image, the method further comprises: The Swin-UNet neural network model is registered and optimized based on the registration domain data and the registration image, and the corresponding registration optimization function is: L(I1,I2,φ)=L sim (I1,I2(φ))+L reg (f) Where I1 and I2 represent the two images to be registered, φ represents the registration domain data, and L sim represents the image similarity function, L reg (φ) represents applying regularization processing to the registration domain data; The loss function for registration optimization of the Swin-UNet neural network model is: L total =L sim-pair +λ1L smooth +λ2L Jdet +λ3L cons-pair Among them, L sim-pair is a similarity term, wherein: L sim-pair (F,M,φ)=L sim (F,M(φ MF ))+L sim (M,F(φ FM )) Among them, F, M represent the target image and the source image respectively, φ MF represents the registration domain obtained by registering the source image M to the target image F, φ FM represents the registration domain obtained by registering the target image F to the source image M; L smooth is a smooth term, the smooth term: Where p represents the voxel position, Ω represents the three-dimensional voxel, Represents gradient calculation; L Jdet is a local directional consistency restriction term, wherein: Where J(φ) represents the Jacobian of the registration domain, σ represents the ReLU function, and N represents the total number of pixels; L cons-pair is a reversible consistency restriction item, wherein: Among them, φ0 represents the unit orthogonal grid, represents the combination of the forward registration domain and the backward registration domain; λ1,λ2,λ3 represent adjustment parameters.

9. An image registration system, characterized in that: The system comprises: A data input module, used to obtain a source image and a target image, process the source image and the target image, and generate input data for a Swin-UNet neural network model; An encoding feature extraction module, used to perform multi-level feature extraction on the input data using multiple first feature extraction layers in the encoding module of the Swin-UNet neural network model to generate feature representation data; A decoding feature extraction module, used to extract features from the feature representation data using multiple second feature extraction layers in the decoding module of the Swin-UNet neural network model to generate content feature data; the multiple second feature extraction layers in the decoding module are symmetrically arranged with the multiple first feature extraction layers in the encoding module; and The registration image generation module is used to perform linear mapping based on the content feature data to determine the registration domain data, and perform spatial transformation on the registration domain data to generate a registration image.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the image registration method according to any one of claims 1 to 8 is implemented.