Image registration method, device, computer device and storage medium
Through the iterative determination of multi-level feature extraction and transform relationship information, combined with the parallel domain adaptation module and the gated fusion module, the problem of insufficient image registration accuracy and speed in the prior art is solved, and efficient and high-precision image registration is achieved.
Patent Information
- Application Number
- CN202111130970.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-09-26
AI Technical Summary
The existing image registration methods combine linear registration and deformation registration, and there are shortcomings in accuracy and speed, especially when the global deformation is too large, the accuracy is lower and the speed is slower.
Through iterative determination of multi-level feature extraction, image registration is performed using the transformation relationship information obtained from multi-level feature extraction, and combining the parallel domain adaptation module and the gated fusion module to improve the accuracy and efficiency of feature extraction.
Improve the accuracy and efficiency of image registration, especially when the global deformation is large.
Smart Images

Figure CN113850853B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of deep learning technology, and more particularly, to an image registration method, apparatus, computer device, and storage medium. Background Art
[0002] Image registration refers to establishing a spatial transformation relationship between parts with the same semantic information in multiple medical images. Imaging examination is a commonly used examination method in the medical field; imaging instruments can be used to image the internal organs of the human body to determine the location of lesions or the development of the disease process; in order to make the results of medical diagnosis more accurate, in many cases, it is necessary to perform image registration on medical images of patients at different times. Summary of the Invention
[0003] Embodiments of the present disclosure provide at least an image registration method, apparatus, computer device, and storage medium.
[0004] In a first aspect, embodiments of the present disclosure provide a registration method, including:
[0005] Performing multi-level feature extraction on a first image to obtain first target feature maps respectively corresponding to the multi-level feature extraction; performing multi-level feature extraction on a second image to obtain second target feature maps respectively corresponding to the multi-level feature extraction;
[0006] For each level of feature extraction, based on the first target feature map, second target feature map corresponding to this level of feature extraction, and the first transformation relationship information corresponding to this level of feature extraction, determining the second transformation relationship information corresponding to this level of feature extraction; wherein, the first transformation relationship information corresponding to this level of feature extraction includes: the second transformation relationship information corresponding to the previous level of feature extraction, or the original transformation relationship information between the first image and the second image; and
[0007] Registering the first image and the second image based on the second transformation relationship information corresponding to the last level of feature extraction.
[0008] In this way, the transformation relationship information is obtained based on multi-level feature extraction, and the first image and the second image are registered using the transformation relationship information determined at the last level, thereby improving the registration accuracy and registration efficiency.
[0009] In an optional implementation, performing multi-level feature extraction on an image to obtain feature maps respectively corresponding to the multi-level feature extraction includes:
[0010] For each level of feature extraction in multi-level feature extraction, determine the first input data and the second input data for this level of feature extraction; the first input data includes: the image, or the encoded feature map corresponding to the next level of feature extraction; the second input data includes: the decoded feature map corresponding to the previous level of feature extraction;
[0011] Perform encoding processing corresponding to this level of feature extraction on the first input data to obtain the encoded feature map corresponding to this level of feature extraction;
[0012] Perform fusion processing on the encoded feature map corresponding to this level of feature extraction and the second input data to obtain a fused feature map;
[0013] Perform decoding processing corresponding to this level of feature extraction on the fused feature map to obtain the decoded feature map corresponding to this level of feature extraction;
[0014] Determine the decoded feature map corresponding to this level of feature extraction as the target feature map corresponding to this level of feature extraction;
[0015] The image includes: the first image and / or the second image;
[0016] The target feature map includes: the first target feature map and / or the second target feature map.
[0017] In this way, through the above multi-level feature extraction method, the features obtained by each level of feature extraction can be transmitted between multi-level feature extractions, further improving the accuracy of determining the transformation relationship between the first image and the second image through the results of multi-level feature extraction.
[0018] In an optional implementation manner, the performing encoding processing corresponding to this level of feature extraction on the first input data to obtain the encoded feature map corresponding to this level of feature extraction includes:
[0019] Perform downsampling processing on the first input data to obtain a downsampled feature map;
[0020] Perform channel attention processing on the downsampled feature map to obtain the attention weights corresponding to each data channel in the downsampled feature map;
[0021] Based on the downsampled feature map and the attention weights, obtain the encoded feature map.
[0022] In this way, by performing channel attention processing on the downsampled feature map to improve the weight features of the encoded feature map for the target channel, certain specific features in the obtained encoded feature map can be made more prominent, improving the accuracy of determining the transformation relationship between the first image and the second image through the results of multi-level feature extraction.
[0023] In an alternative implementation, the channel attention processing of the downsampled feature map to obtain the attention weights corresponding to each data channel in the downsampled feature map includes:
[0024] Performing global average pooling on the downsampled feature map to obtain a first feature submap;
[0025] Based on the first feature submap, determining the candidate attention weights corresponding to each data channel; and based on the first feature submap, determining the feature domain weights of the candidate attention corresponding to each data channel;
[0026] Based on the candidate attention weights corresponding to each data channel and the feature domain weights of the candidate attention corresponding to each data channel, obtaining the attention weights corresponding to each data channel in the downsampled feature map.
[0027] In this way, through the domain-aware mechanism, according to the properties of the feature map, multiple candidate channel weights are combined to obtain the final result, so that when the performance differences between the first image and the second image in the image domain are large, similar features can still be obtained through the feature extractor, facilitating the subsequent registration process.
[0028] In an alternative implementation, the fusion processing of the encoded feature map corresponding to the feature extraction at this level and the second input data to obtain a fused feature map includes:
[0029] Based on the encoded feature map corresponding to the feature extraction at this level and the second input data, obtaining the weights corresponding to each feature point in the encoded feature map corresponding to the feature extraction at this level;
[0030] Based on the weights corresponding to each feature point in the encoded feature map corresponding to the feature extraction at this level and the encoded feature map, obtaining a second feature submap;
[0031] Concatenating the second feature submap and the second input data to obtain the fused feature map.
[0032] In this way, the decoded feature map can be improved in both spatial resolution and semantic information expression.
[0033] In an alternative implementation, the obtaining of the weights corresponding to each feature point in the encoded feature map corresponding to the feature extraction at this level based on the encoded feature map corresponding to the feature extraction at this level and the second input data includes:
[0034] Concatenating the encoded feature map corresponding to the feature extraction at this level and the second input data to obtain a third feature submap;
[0035] Perform convolution processing on the third feature sub - graph to obtain a fourth feature sub - graph, and based on the fourth feature sub - graph, obtain the encoded feature map and the local autocorrelation coefficient of the second input data;
[0036] Based on the local autocorrelation coefficient, obtain the weights corresponding to each feature point in the encoded feature map corresponding to this level of feature extraction.
[0037] In an optional implementation, the eigenvalue of any feature point in the third feature sub - graph represents the autocorrelation coefficient of the image region corresponding to this feature point;
[0038] The obtaining the weights corresponding to each feature point in the encoded feature map corresponding to this level of feature extraction based on the local autocorrelation coefficient includes:
[0039] Use an activation function to perform activation processing on the autocorrelation coefficient to obtain the gating activation values corresponding to each feature point in the encoded feature map; the gating activation values are used to represent the weights corresponding to each feature point in the encoded feature map.
[0040] In an optional implementation, the obtaining the encoded feature map and the local autocorrelation coefficient of the second input data based on the fourth feature sub - graph includes:
[0041] Perform maximum processing on the channel dimension of the fourth feature sub - graph to obtain the maximum value of the channel dimension of the fourth feature sub - graph;
[0042] Perform average processing on the channel dimension of the fourth feature sub - graph to obtain the average value of the channel dimension of the fourth feature sub - graph; and
[0043] Perform splicing processing on the maximum value of the channel dimension and the average value of the channel dimension to obtain a channel - dimension splicing result, and perform convolution and normalization processing on the channel - dimension splicing result to obtain the encoded feature map and the local autocorrelation coefficient of the second input data.
[0044] In an optional implementation, for each level of feature extraction, based on the first target feature map, the second target feature map, and the first transformation relationship information corresponding to this level of feature extraction, determining the second transformation relationship information corresponding to this level of feature extraction includes:
[0045] Based on the first target feature map, the second target feature map, and the first transformation relationship information, determine the transformation residual between this level of feature extraction and the previous level of feature extraction;
[0046] Based on the transformation residual and the first transformation relationship information, determine the second transformation relationship information corresponding to this level of feature extraction.
[0047] In this way, the transformation relationship information can be determined more accurately.
[0048] In an alternative embodiment, the registration method is applied to a pre-trained registration neural network, which includes two branch networks: a feature extraction neural network and a multi-level registration neural network; the feature extraction neural network is used to perform multi-level feature extraction processing on the first image and the second image respectively; the multi-level registration neural network is used to determine the target transformation relationship information between the first image and the second image based on the multi-level features extracted from the first image and the second image, where the target transformation relationship information is used to register the first image and the second image.
[0049] In this way, the first image and the second image can be registered more accurately based on the trained network.
[0050] In a second aspect, an embodiment of the present disclosure further provides an image registration device, including:
[0051] An extraction module, configured to perform multi-level feature extraction on a first image to obtain first target feature maps respectively corresponding to the multi-level feature extraction; and perform multi-level feature extraction on a second image to obtain second target feature maps respectively corresponding to the multi-level feature extraction;
[0052] A determination module, configured to, for each level of feature extraction, determine second transformation relationship information corresponding to the level of feature extraction based on the first target feature map corresponding to the level of feature extraction, the second target feature map, and the first transformation relationship information corresponding to the level of feature extraction; where the first transformation relationship information corresponding to the level of feature extraction includes: the second transformation relationship information corresponding to the previous level of feature extraction, or the original transformation relationship information between the first image and the second image; and
[0053] A registration module, configured to register the first image and the second image based on the second transformation relationship information corresponding to the last level of feature extraction.
[0054] In an alternative embodiment, the extraction module includes:
[0055] An extraction unit, configured to, for each level of feature extraction in the multi-level feature extraction, determine first input data and second input data for the level of feature extraction; the first input data includes: the image, or the encoded feature map corresponding to the next level of feature extraction; the second input data includes: the decoded feature map corresponding to the previous level of feature extraction;
[0056] An encoding unit, configured to perform encoding processing corresponding to the level of feature extraction on the first input data to obtain an encoded feature map corresponding to the level of feature extraction;
[0057] A fusion unit, configured to perform a fusion process on the encoded feature map corresponding to the feature extraction at this level and the second input data to obtain a fused feature map;
[0058] A decoding unit, configured to perform decoding processing corresponding to the feature extraction at this level on the fused feature map to obtain a decoded feature map corresponding to the feature extraction at this level;
[0059] A determination unit, configured to determine the decoded feature map corresponding to the feature extraction at this level as the target feature map corresponding to the feature extraction at this level;
[0060] The image includes: the first image and / or the second image;
[0061] The target feature map includes: the first target feature map and / or the second target feature map.
[0062] In an optional implementation manner, the encoding unit includes:
[0063] A downsampling sub-unit, configured to perform downsampling processing on the first input data to obtain a downsampled feature map;
[0064] A channel attention processing sub-unit, configured to perform channel attention processing on the downsampled feature map to obtain attention weights corresponding to each data channel in the downsampled feature map;
[0065] A first determination sub-unit, configured to obtain the encoded feature map based on the downsampled feature map and the attention weights.
[0066] In an optional implementation manner, the channel attention processing sub-unit is specifically configured to:
[0067] Perform global average pooling processing on the downsampled feature map to obtain a first feature sub-map;
[0068] Based on the first feature sub-map, determine candidate attention weights corresponding to each data channel; and based on the first feature sub-map, determine feature domain weights of candidate attention corresponding to each data channel;
[0069] Based on the candidate attention weights corresponding to each data channel and the feature domain weights of candidate attention corresponding to each data channel, obtain the attention weights corresponding to each data channel in the downsampled feature map.
[0070] In an optional implementation manner, the fusion unit includes:
[0071] A second determination subunit, configured to obtain weights corresponding to each feature point in the encoded feature map corresponding to the feature extraction at this level, based on the encoded feature map corresponding to the feature extraction at this level and the second input data;
[0072] A third determination subunit, configured to obtain a second feature sub-map based on the weights corresponding to each feature point in the encoded feature map corresponding to the feature extraction at this level and the encoded feature map;
[0073] A splicing subunit, configured to splice the second feature sub-map and the second input data to obtain the fused feature map.
[0074] In an optional implementation manner, the second determination subunit is specifically configured to:
[0075] Splice the encoded feature map corresponding to the feature extraction at this level and the second input data to obtain a third feature sub-map;
[0076] Perform convolution processing on the third feature sub-map to obtain a fourth feature sub-map, and based on the fourth feature sub-map, obtain the encoded feature map and the local autocorrelation coefficient of the second input data;
[0077] Based on the local autocorrelation coefficient, obtain weights corresponding to each feature point in the encoded feature map corresponding to the feature extraction at this level.
[0078] In an optional implementation manner, the eigenvalue of any feature point in the third feature sub-map represents the autocorrelation coefficient of the image region corresponding to the feature point;
[0079] The second determination subunit is specifically configured to:
[0080] Use an activation function to perform activation processing on the autocorrelation coefficient to obtain the gated activation values corresponding to each feature point in the encoded feature map; the gated activation values are used to represent the weights corresponding to each feature point in the encoded feature map.
[0081] In an optional implementation manner, the second determination subunit is specifically configured to:
[0082] Perform maximum processing on the channel dimension of the fourth feature sub-map to obtain the maximum value of the channel dimension of the fourth feature sub-map;
[0083] Perform average processing on the channel dimension of the fourth feature sub-map to obtain the average value of the channel dimension of the fourth feature sub-map; and
[0084] Perform splicing processing on the maximum value of the channel dimension and the average value of the channel dimension to obtain a channel dimension splicing result, and perform convolution and normalization processing on the channel dimension splicing result to obtain the encoded feature map and the local autocorrelation coefficient of the second input data.
[0085] In an alternative implementation, the determining module is specifically configured to:
[0086] Determine the transformation residual between this-level feature extraction and the previous-level feature extraction based on the first target feature map, the second target feature map, and the first transformation relationship information;
[0087] Determine the second transformation relationship information corresponding to this-level feature extraction based on the transformation residual and the first transformation relationship information.
[0088] In an alternative implementation, the registration device is applied to a pre-trained registration neural network, and the registration neural network includes two branch networks, a feature extraction neural network and a multi-level registration neural network; the feature extraction neural network is used to perform multi-level feature extraction processing on the first image and the second image respectively; the multi-level registration neural network is used to determine the target transformation relationship information between the first image and the second image based on the multi-level features extracted from the first image and the second image, where the target transformation relationship information is used to register the first image and the second image.
[0089] In a third aspect, an embodiment of the present disclosure further provides a computer device, including: a processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus, and when the machine-readable instructions are executed by the processor, the steps in the first aspect, or any possible implementation manner in the first aspect are executed.
[0090] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps in the first aspect, or any possible implementation manner in the first aspect are executed.
[0091] The image registration method, device, computer device, and storage medium provided by the embodiments of the present disclosure perform multi-level feature extraction on a first image and a second image, and perform transformation based on the extracted multi-level features to obtain the transformation relationship information corresponding to each level, and determine the transformation relationship information corresponding to the last level. Thus, based on the transformation relationship information, the first image and the second image are registered to complete the registration process. In this way, by iteratively determining multi-level features, the transformation relationship information between the first target feature map and the second target feature map is extracted, and the first image and the second image are registered using the transformation relationship information determined at the last level, thereby improving the registration accuracy and registration efficiency.
[0092] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required to be used in the embodiments. The accompanying drawings are incorporated into the specification and constitute a part of this specification. These drawings show embodiments that conform to the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0094] Figure 1 Shows a flowchart of an image registration method provided by an embodiment of the present disclosure;
[0095] Figure 2 Shows a schematic diagram of a specific method for performing multi-level feature extraction on an image provided by an embodiment of the present disclosure;
[0096] Figure 3 Shows a structural example of a feature extractor for performing multi-level feature extraction provided by an embodiment of the present disclosure;
[0097] Figure 4 Shows a specific example of a parallel domain adaptation module provided by an embodiment of the present disclosure;
[0098] Figure 5 Shows a specific example of a gated fusion module provided by an embodiment of the present disclosure;
[0099] Figure 6 Shows a structural example of a registration neural network provided by an embodiment of the present disclosure;
[0100] Figure 7Shows a schematic diagram of an image registration device provided by an embodiment of the present disclosure;
[0101] Figure 8 Shows a specific schematic diagram of an extraction module in the image registration device provided by an embodiment of the present disclosure;
[0102] Figure 9 Shows a specific schematic diagram of an encoding unit in the extraction module provided by an embodiment of the present disclosure;
[0103] Figure 10 Shows a specific schematic diagram of a fusion unit in the extraction module provided by an embodiment of the present disclosure
[0104] Figure 11 Shows a schematic diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners
[0105] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments. The components of the embodiments of the present disclosure described and illustrated herein generally may be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided herein is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0106] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings.
[0107] The term "and / or" in this article merely describes an associated relationship and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this article means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.
[0108] It has been found through research that according to different types of spatial transformations, image registration methods can be divided into two categories: linear transformation and deformation transformation. Generally, linear transformation includes rigid body registration, affine registration, etc., which are applicable to global and large-scale spatial transformations; deformation transformation includes deformation registration based on control points and deformation registration based on dense deformation fields, etc., which are usually local and non-uniform spatial transformations. In addition, according to different implementation methods of the registration process, the registration methods can also be divided into two categories: optimization-based registration and learning-based registration; optimization-based registration is more common in traditional registration methods, which can obtain higher accuracy, but the registration speed is usually slow, especially in non-linear registration; learning-based registration is now often realized through deep learning, which requires a large amount of data for training, has a faster speed during inference, and has higher accuracy in deformation registration.
[0109] During the application process of registration, in order to ensure both the accuracy and speed of registration, linear registration and non-linear registration are usually combined to jointly complete the registration process. For example: first, use the linear registration method of rigid or affine transformation to roughly align the images to be registered, and then use the deformation registration method for fine matching.
[0110] There are mainly three existing implementation schemes for combining linear registration and deformation registration:
[0111] First, both linear registration and deformation registration use optimization-based methods. In this way, a large amount of data is not required for training, but the accuracy of deformation registration is lower than that of learning-based methods, and a large amount of inference time is consumed; second, both linear registration and deformation registration use learning-based methods. This method has the fastest registration speed, but it is necessary to train the neural network related to linear registration and the neural network related to deformation registration respectively; and during inference, the video memory of more than two neural networks needs to be maintained, which poses a relatively high requirement for the equipment during the registration process. In addition, when using the learning-based registration method for linear registration, if the initial displacement between images is very large, the registration accuracy will be inferior to that of the optimization-based registration method. Third, use the optimization-based registration method for linear registration and the learning-based registration method for deformation registration. This scheme is a compromise between the above two schemes, avoiding some of the disadvantages of the above two schemes, and having relatively high accuracy, but the inference speed is still slow, and the registration accuracy is low when the global deformation is too large.
[0112] Based on the above research, the present disclosure provides an image registration method, which determines multi-level features through iteration, extracts the transformation relationship information between the obtained first target feature map and the second target feature map, and uses the transformation relationship information determined at the last level to register the first image and the second image, thereby improving the registration accuracy and registration efficiency.
[0113] For ease of understanding of this embodiment, first, a method for registering an image disclosed in this disclosure embodiment will be introduced in detail. The execution subject of the registration method provided in this disclosure embodiment is generally a computer device with certain computing capabilities, such as a terminal device, a server, or other processing devices. In some possible implementation manners, the registration method can be implemented by a processor calling computer-readable instructions stored in a memory.
[0114] See Figure 1 As shown, it is a flowchart of the method for registering an image provided in this disclosure embodiment. The method includes steps S101 to S103, where:
[0115] S101: Perform multi-level feature extraction on a first image to obtain a first target feature map corresponding to the multi-level feature extraction respectively; perform multi-level feature extraction on a second image to obtain a second target feature map corresponding to the multi-level feature extraction respectively.
[0116] Among them, the first image and the second image may include different images taken of the same object at different times or different angles. For example, when applying the image registration method provided in this disclosure example to the medical field, the first image and the second image may be medical images obtained by taking multiple shots during a single scan of the same lesion of the same patient, or may be different medical images obtained during scans of the same lesion of the patient at different times, which is not limited here.
[0117] In this disclosure embodiment, after the first image and the second image are determined, multi-level feature extraction can be performed on the two images respectively to obtain a first target feature map and a second target feature map corresponding to the multi-level feature extraction respectively. Among them, multi-level feature extraction can, for example, adopt the method of neural network self-learning to realize the gradient backpropagation of multi-level features of the image and extract the high-level semantic features of the image.
[0118] Refer to Figure 2 , Figure 2 It is a schematic diagram of a specific method for performing multi-level feature extraction on an image shown in this disclosure embodiment, including steps S1011 to S1015, where:
[0119] S1011: For each level of feature extraction in the multi-level feature extraction, determine the first input data and the second input data of this level of feature extraction; the first input data includes: the image, or the encoded feature map corresponding to the next level of feature extraction; the second input data includes: the decoded feature map corresponding to the previous level of feature extraction;
[0120] S1012: Perform encoding processing corresponding to the feature extraction of this level on the first input data to obtain an encoded feature map corresponding to the feature extraction of this level;
[0121] S1013: Perform fusion processing on the encoded feature map corresponding to the feature extraction of this level and the second input data to obtain a fused feature map;
[0122] S1014: Perform decoding processing corresponding to the feature extraction of this level on the fused feature map to obtain a decoded feature map corresponding to the feature extraction of this level;
[0123] S1015: Determine the decoded feature map corresponding to the feature extraction of this level as the target feature map corresponding to the feature extraction of this level.
[0124] Wherein, the image includes: the first image and / or the second image;
[0125] The target feature map includes: the first target feature map and / or the second target feature map.
[0126] In this example, the feature extractor can perform feature extraction on the first image and the second image respectively to obtain multiple levels of features corresponding to the two images. The feature extractor is a feature pyramid network, which is divided into two stages: encoding and decoding, and connects the low-level and high-level networks by means of skip connections. The feature extractor receives an image as input data, downsamples (also known as "downsampling") and extracts features layer by layer in the encoding module, and upsamples (also known as "upsampling") and extracts features layer by layer in the decoding module. The output of each decoding module will be fed into the iterative registration network to form a multi-level, coarse-to-fine pyramid feature.
[0127] For the encoding feature extraction process, for the highest level, that is, the coarsest level in the "coarse-to-fine" structure of the feature "pyramid" network, the corresponding input data is the original image, that is, the first image or the second image. For other levels, the corresponding input data is the encoded feature map corresponding to the feature extraction of the subsequent level. Thus, the iterative process of encoding feature extraction can be realized. For the decoding feature extraction process, its lowest level, that is, the finest level in the "coarse-to-fine" structure of the feature "pyramid" network, is the decoded feature map obtained from the corresponding lowest-level encoded feature map. For other levels, the corresponding input data is the decoded feature map corresponding to the feature extraction of the previous level, and finally the feature extraction result for the original image is output. Thus, the iterative process of decoding feature extraction can be realized.
[0128] The multi-level feature extraction described in the embodiments of the present disclosure indicates a multi-level feature extraction process from "fine" to "coarse". Each level of feature extraction includes an encoding process and a decoding process corresponding to that level of feature extraction, that is, from the bottom to the top of the "pyramid" network.
[0129] As Figure 3 shown, it is a structural example of a feature extractor for performing multi-level feature extraction provided by the embodiments of the present disclosure.
[0130] In this example, each level of feature extraction includes an encoding network corresponding to that level of feature extraction and a decoding network corresponding to that level of feature extraction;
[0131] wherein, each encoding network includes: an encoding module and a parallel domain adaptation module.
[0132] The decoding network corresponding to the first-level encoding network includes: a decoding module.
[0133] The decoding networks corresponding to the other encoding networks except the first-level encoding network include: a decoding module and a gated fusion module.
[0134] According to Figure 3 it can be known that the encoding process is a process from "coarse" to "fine", while the decoding process decodes sequentially from "fine" to "coarse". The multi-level feature extraction described in the embodiments of the present disclosure corresponds to the decoding process.
[0135] In this example, it includes: 4-level feature extraction;
[0136] wherein, the multi-level feature extraction refers to the feature extraction from the first level to the fourth level.
[0137] Wherein, the first-level feature extraction includes: an encoding module 1, a parallel domain adaptation module 1, and a decoding module 1;
[0138] The second-level feature extraction includes: an encoding module 2, a parallel domain adaptation module 2, a gated fusion module 2, and a decoding module 2;
[0139] The third-level feature extraction includes: an encoding module 3, a parallel domain adaptation module 3, a gated fusion module 3, and a decoding module 3;
[0140] The fourth-level feature extraction includes: an encoding module 4, a parallel domain adaptation module 4, a gated fusion module 4, and a decoding module 4.
[0141] Wherein, for the encoding network in each level of feature extraction, it is used to perform encoding processing on the first input data corresponding to that level of feature extraction to obtain the encoded feature map corresponding to that level of feature extraction.
[0142] For the decoding network in each level of feature extraction, it is used to fuse the encoded feature map corresponding to the feature extraction at this level and the second input data to obtain a fused feature map, and perform decoding processing corresponding to the feature extraction at this level on the fused feature map to obtain the decoded feature map corresponding to the feature extraction at this level.
[0143] Taking the neural network shown above Figure 3 as an example, in the above step S1011, the first input data and the second input data in the feature extraction process for each level can be determined.
[0144] Among them, for the first-level feature extraction, since there is no previous-level feature extraction, only the first input data needs to be input. The first input data includes: the encoded feature map s1 corresponding to the second-level feature extraction, and the second input data is empty.
[0145] For the second-level feature extraction, the first input data includes: the encoded feature map s2 corresponding to the third-level feature extraction; the second input data includes: the decoded feature map k1 corresponding to the first-level feature extraction.
[0146] For the third-level feature extraction, the first input data includes: the encoded feature map s3 corresponding to the fourth-level feature extraction; the second input data includes: the decoded feature map k2 corresponding to the second-level feature extraction.
[0147] For the fourth-level feature extraction, the first input data includes: the original image; the second input data includes: the decoded feature map k3 corresponding to the third-level feature extraction.
[0148] In the above step S1012, after determining the first input data corresponding to each level in the encoding process, encoding processing can be performed based on the first input data to obtain the encoded feature map corresponding to each level.
[0149] In an embodiment of the present disclosure, the encoding the first input data with the encoding process corresponding to the feature extraction at this level to obtain the encoded feature map corresponding to the feature extraction at this level includes:
[0150] Performing downsampling processing on the first input data to obtain a downsampled feature map;
[0151] Performing channel attention processing on the downsampled feature map to obtain the attention weights corresponding to each data channel in the downsampled feature map;
[0152] Based on the downsampled feature map and the attention weights, the encoded feature map is obtained.
[0153] As Figure 3In the illustrated example, when the first input data is input into the encoding module in each level of feature extraction, the encoding module performs downsampling processing on the first input data to obtain a corresponding downsampled feature map, and inputs the downsampled feature map into the parallel domain adaptation module corresponding to each level of the encoding module. Among them, the parallel domain adaptation module can enhance the expression ability of the encoding network for certain specific features during the encoding stage; for example, if the image includes human organs; then the feature extractor can be trained to enhance the expression ability of the texture features of the organs. After the downsampled feature map is input into the parallel domain adaptation module, the parallel domain adaptation module performs channel attention processing on the downsampled feature map to obtain the attention weights corresponding to each data channel in the downsampled feature map. Specifically, the above process includes:
[0154] Perform global average pooling on the downsampled feature map to obtain a first feature sub-map;
[0155] Based on the first feature sub-map, determine the candidate attention weights corresponding to each data channel; and based on the first feature sub-map, determine the feature domain weights of the candidate attention corresponding to each data channel;
[0156] Based on the candidate attention weights corresponding to each data channel and the feature domain weights of the candidate attention corresponding to each data channel, obtain the attention weights corresponding to each data channel in the downsampled feature map.
[0157] In the embodiments of the present disclosure, the parallel domain adaptation module includes two mechanisms: a channel attention mechanism and a domain perception mechanism. In the parallel domain adaptation module, the channel attention mechanism determines the channel attention weights of multiple channels in the form of grouped convolution, thereby generating multiple candidate channel weights. The domain perception mechanism can combine multiple candidate channel weights according to the properties of the feature map to obtain the final result, so that when the performance differences between the first image and the second image in the image domain are large, similar features can also be obtained through the feature extractor, which is convenient for the subsequent registration process.
[0158] In this embodiment, the parallel domain adaptation module is applied to the encoding stage of multi-level feature extraction. Specifically, during the encoding process from the (i + 1)-th level to the i-th level, the corresponding parallel domain adaptation module is applied. For example, in Figure 3 during the encoding process from the 3rd level to the 2nd level shown in, the parallel domain adaptation module 3 is applied. As Figure 4As shown, it is a specific example of a parallel domain adaptation module provided by an embodiment of the present disclosure. Here, it is assumed that the downsampled feature map corresponding to the encoding module of the i-th layer is represented as (H, C, D, H, W), where N represents the number of feature maps included in one feature extraction (Number of instances in batch), C represents the number of channels of the feature map (Channel), and D, H, W respectively represent the length, width, and height of the feature map (Depth, Height, Width), and B represents the number of branches of channel attention (Branches of attention). After inputting the encoded feature map into this parallel domain adaptation module, after obtaining the downsampled feature map through the encoding module, global average pooling processing (Global Average Pooling) can be performed on the downsampled feature map, that is, dimensionality reduction of the feature map, by taking the average of all pixels in the length, width, and height spatial dimensions to reduce the dimensionality of the data, so that the feature map changes from (N, C, D, H, W) to (N, C), thereby extracting the overall information in the channel dimension to obtain the first feature sub-map. Then, the first feature sub-map is respectively subjected to repeat (Repeat) and flatten (Flatten) processing through two branches. In the first branch, after repeating the first feature sub-map, the first feature sub-map is transformed into the form of (N, BC, 1), where repeating means repeating the dimensionality-reduced data B times, corresponding to B channel attentions. Here, the repeating operation is to change the data format for convenient calculation. Then, the feature map in this form is subjected to B groups of convolutions, activation, and then B groups of convolutions. Among them, the activation function can adopt the rectified linear unit (Rectified Linear Unit, ReLU). Finally, the channel reorganization of the convolved feature map is performed to obtain a feature map of (N, C, B), thereby obtaining the candidate attention weights corresponding to each data channel. In the second branch, after flattening the first feature sub-map, the first feature sub-map is transformed into the form of (N, C), where flattening means rearranging the order of the data. Here, the flattening operation is to change the data format for convenient calculation. Then, the feature map in this form is subjected to fully connected and activation processing. Among them, the ReLU activation function and the excitation function (Softmax) of the deep learning output layer can be respectively adopted to obtain a feature map in the form of (N, B), thereby obtaining the feature domain weights of the candidate attentions corresponding to each data channel. After obtaining the candidate attention weights and feature domain weights corresponding to each data channel, the above two data are subjected to inner product processing to obtain feature data of (N, C). Then, activation processing is performed on this data. Among them, the activation function can be the sigmoid growth curve (Sigmoid), so that the encoded feature map of the i-th level can be obtained.
[0159] After determining the feature domain weights of the candidate attention corresponding to each channel, the attention weights corresponding to each data channel in the downsampled feature map can be determined based on the feature domain weights of the candidate attention corresponding to each data channel. Taking multiple images of the human liver as an example, for the same lesion, due to different scanning and shooting times or different shooting angles, although it is the same position, the presented images are different. Therefore, based on the feature domain weights determined by the domain perception module, the feature domain weights of the candidate attention corresponding to each data channel are proportioned to highlight the weight occupied by the target position, that is, the weight occupied by the liver position, and weaken the weights occupied by other positions, such as muscles, blood, etc. In this way, the target position can be prominently displayed. Thus, even when the display effects of the first image and the second image differ greatly, similar features can be obtained through the feature extractor for registration.
[0160] Here, the channel attention mechanism can adopt a method for determining the channel attention weights in multiple parallel groups. For example, assuming there are 12 channels, usually, the above 12 channels can be convolved separately to obtain the candidate attention weights corresponding to the 12 channels. To improve the processing speed, the 12 channels can be divided into three branch channel attentions, with 4 channels in each branch. In this way, the three branch channel attentions can be convolved simultaneously, improving the processing speed.
[0161] After determining the attention weights corresponding to each data channel in the downsampled feature map, the encoded feature map can be obtained based on the downsampled feature map and the attention weights, that is, the features of each channel in the downsampled feature map are recombined according to the corresponding attention weights to obtain the encoded feature maps corresponding to each level.
[0162] In the above step S1013, as Figure 3 shown, in the decoding process corresponding to each level of feature extraction, in addition to the second input data input to each level of decoding module, the corresponding encoded feature map also participates in the generation process of the decoded feature map.
[0163] Here, since the encoded feature map has a high spatial resolution but a low degree of semantic information expression, and the upsampled feature map of the previous decoding module has a low spatial resolution but a high degree of semantic information expression, therefore, in order to combine the respective advantages of the encoded feature map and the upsampled feature map of the previous decoding module, in the embodiments of the present disclosure, a gated fusion module for improving the fusion effect of high-level low-resolution features and low-level high-resolution features is adopted in the decoding stage.
[0164] Specifically, the fusion process of the encoded feature map corresponding to the feature extraction at this level and the second input data to obtain a fused feature map includes:
[0165] Based on the encoded feature map corresponding to this level of feature extraction and the second input data, obtain the weights corresponding to each feature point in the encoded feature map corresponding to this level of feature extraction;
[0166] Based on the weights corresponding to each feature point in the encoded feature map corresponding to this level of feature extraction and the encoded feature map, obtain a second feature sub-map;
[0167] Concatenate the second feature sub-map and the second input data to obtain the fused feature map.
[0168] Exemplarily, as Figure 3 shown, through a gated fusion module, fuse the encoded feature map of the encoded module data corresponding to this level of decoding module with the second feature data input to this decoding module to obtain the weights occupied by the encoded feature map corresponding to this level of feature extraction, multiply the encoded feature map by this weight to obtain a second feature sub-map, and concatenate (Concatenate) it with the second feature data input to this decoding module, then the fused feature map can be obtained. In this way, after multiplying the encoded feature map from the encoding module by the weight, it is concatenated with the decoded feature map of the decoding module in the channel dimension and sent to the next decoding module, enabling the feature maps from the two sources to be fused more efficiently.
[0169] As Figure 5 shown, the embodiments of the present disclosure also provide a specific example of a gated fusion module. The main function of the gated fusion module is to determine the weights corresponding to each feature point in the encoded feature map. Specifically, the step of obtaining the weights corresponding to each feature point in the encoded feature map corresponding to this level of feature extraction based on the encoded feature map corresponding to this level of feature extraction and the second input data includes:
[0170] Concatenate the encoded feature map corresponding to this level of feature extraction and the second input data to obtain a third feature sub-map;
[0171] Perform convolution processing on the third feature sub-map to obtain a fourth feature sub-map, and based on the fourth feature sub-map, obtain the encoded feature map and the local autocorrelation coefficient with the second input data;
[0172] Based on the local autocorrelation coefficient, obtain the weights corresponding to each feature point in the encoded feature map corresponding to this level of feature extraction.
[0173] Among them, the feature value of any feature point in the third feature sub-map represents the autocorrelation coefficient of the image region corresponding to this feature point.
[0174] After obtaining the third feature sub - map obtained by splicing the encoded feature map corresponding to the feature extraction at this level and the second input data, perform convolution, normalization, and activation processing based on the Rectified Linear Unit (ReLU) on the third feature sub - map to obtain a fourth feature sub - map corresponding to the third feature sub - map. Here, convolution, normalization, and activation processing are an operation pattern in a Convolutional Neural Network (CNN). The number of times the above operations appear represents the depth of the network. The deeper the network, the stronger the expression ability, and at the same time, the larger the number of parameters. Here, in order to improve the expression ability of the network, two or more layers of convolution, normalization, and activation processing can be selected.
[0175] After determining the fourth feature sub - map, based on the fourth feature sub - map, the encoded feature map and the local autocorrelation coefficient of the second input data can be obtained, including:
[0176] Perform maximum processing on the channel dimension of the fourth feature sub - map to obtain the maximum value of the channel dimension of the fourth feature sub - map; perform average processing on the channel dimension of the fourth feature sub - map to obtain the average value of the channel dimension of the fourth feature sub - map; and
[0177] Perform splicing processing on the maximum value of the channel dimension and the average value of the channel dimension to obtain a splicing result in the channel dimension, and perform convolution and normalization processing on the splicing result in the channel dimension to obtain the encoded feature map and the local autocorrelation coefficient of the second input data.
[0178] Among them, both the maximum in the channel dimension and the average in the channel dimension are dimensionality reduction operations for data. They operate based on the channel dimension. The purpose is to reduce the number of parameters and retain spatial features, so that the data changes from (N, C, D, H, W) to (N, 1, C, H, W). Here, both the maximum value of the channel dimension and the average value of the channel dimension can be expressed as (N, 1, C, H, W). After determining the maximum value of the channel dimension and the average value of the channel dimension, the above two values can be spliced to obtain a spliced result (N, 2, C, H, W). Then, convolution and normalization processing can be performed on this spliced result to obtain the encoded feature map and the local autocorrelation coefficient of the second input data.
[0179] After obtaining the local autocorrelation coefficient, the activation function (sigmoid) can be used to activate the autocorrelation coefficient to obtain the gating activation values corresponding to each feature point in the encoded feature map; the gating activation values are used to represent the weights corresponding to each feature point in the encoded feature map. In this way, after multiplying the encoded feature map by the gating activation value obtained through the activation function and then concatenating it with the second input data, the fused feature map can be obtained.
[0180] As Figure 5 In the specific example of the gating fusion module shown, the encoded feature map output by the i-th level encoding module and the decoded feature map of the data after downsampling by the i-th level decoding module are concatenated. Here, concatenation means concatenation in the channel dimension, and it is required that the sizes of other channel dimensions are the same. For example, if the two data are respectively (N, C1, D, H, W) and (N, C2, D, H, W), the size after concatenation is (N, C1 + C2, D, H, W). Then, convolution, normalization, and activation processing are performed on the concatenated feature map, and channel dimension maximum and channel dimension average processing are also performed on it. After obtaining the channel dimension maximum value and the channel dimension average value, the above two data are concatenated, and convolution, normalization, and activation processing are performed again. After multiplication processing, the decoded feature map of the (i + 1)-th level is obtained.
[0181] In the above steps S1014 - S1015, as Figure 3 shown, after obtaining the fused feature map in the decoding stage, decoding processing corresponding to the feature extraction at this level is performed on the fused feature map to obtain the decoded feature map corresponding to the feature extraction at this level, and the decoded feature map corresponding to the feature extraction at this level is determined as the target feature map corresponding to the feature extraction at this level, which is convenient for subsequent inputting the target feature map corresponding to each level into the registration network.
[0182] Thus, in the embodiments of the present disclosure, the method of multi-level feature extraction, that is, the pyramid feature method, has a high degree of neural network reuse and fast iterative registration speed. At the same time, parallel domain adaptation modules and gating fusion modules are used in the encoding and decoding stages to enhance the generalization ability and robustness of the neural network and improve the abstraction ability in the feature extraction process.
[0183] Continuing from the above S101, this registration method further includes:
[0184] S102: For each level of feature extraction, based on the first target feature map, the second target feature map corresponding to the feature extraction at this level, and the first transformation relationship information corresponding to the feature extraction at this level, determine the second transformation relationship information corresponding to the feature extraction at this level.
[0185] Among them, the first transformation relation information corresponding to this level of feature extraction includes: the second transformation relation information corresponding to the previous level of feature extraction, or the original transformation relation information between the first image and the second image.
[0186] Exemplarily, the original transformation relation information is also called the initial deformation relation. When the global deformation is small, the identity transformation can be used as the initial deformation relation, that is, the initial images of the first image and the second image are used as input data; when there is a large global deformation, the initialization methods commonly used in traditional methods can be used, such as: initializing the displacement based on the center of the image or the center of gravity of the gray level, and using the moment estimation of the gray level to initialize the rotation, etc. In this way, the problems of poor accuracy and even field of view (FoV) mismatch when the previous linear registration network encounters too large global deformation can be solved.
[0187] Specifically, for each level of feature extraction, based on the first target feature map, the second target feature map corresponding to this level of feature extraction, and the first transformation relation information corresponding to this level of feature extraction, determining the second transformation relation information corresponding to this level of feature extraction includes:
[0188] Based on the first target feature map, the second target feature map, and the first transformation relation information, determining the transformation residual between this level of feature extraction and the previous level of feature extraction;
[0189] Based on the transformation residual and the first transformation relation information, determining the second transformation relation information corresponding to this level of feature extraction.
[0190] Among them, the iterative registration framework accepts an initial transformation relation in the form of a deformation field and several pairs of feature maps as input (the registration of each stage corresponds to the feature maps of one stage). The iterative registration framework includes multiple stages, and each stage includes a registration module and a combination module. The registration module accepts the feature map pair of this stage and the transformation relation of the previous stage (for the first stage, it is the initial transformation relation) as input, and outputs the residual with respect to the previous transformation relation, that is:
[0191] (1)
[0192] where is the registration module, represents the feature map output by the target image feature extractor in the i-th stage, represents the feature map output by the source image feature extractor in the i-th stage, represents applying the transformation relation, represents the cumulative transformation relation in the i-th stage, represents with respect to The residuals. and The superscript of indicates the direction of the transformation relationship, that is, represents the transformation relationship from the source image to the target image, represents the transformation relationship from the target image to the source image.
[0193] In the embodiments of the present disclosure, both the source image and the target image can be two images with different sources to be registered during the registration process, such as two images from source A and source B. The source image can be the source A image or the source B image, and the target image can also be the source A image or the source B image. In some possible embodiments, if the target direction of registration is determined, for example, registering the image from source A to the image from source B, then the image from source B can be determined as the target image and the image from source A as the source image, which is not limited herein.
[0194] Here, the deformation field refers to that in image registration, a regular spatial grid is added with the predicted deformation field to obtain a sampling grid. Using the sampling grid containing deformation information for the floating image, the resulting image is the deformed image. The size of the deformation field corresponding to a two-dimensional image of size [W, H] is [W, H, 2], where the size of the third dimension is 2, respectively representing the displacements in the x-axis and y-axis directions. Similarly, the size of the deformation field corresponding to a three-dimensional image of size [D, W, H] is [D, W, H, 3], where the size of the third dimension is 3, respectively representing the displacements in the x-axis, y-axis, and z-axis directions.
[0195] The registration module is divided into two types: linear transformation and deformation transformation. They are the same in terms of input and output. Therefore, in application, the registration modules of linear and deformation transformations can be freely combined according to requirements, such as connecting one linear registration module and three deformation transformation modules in series. To ensure that the transformation process is diffeomorphic, the two modules do not directly output , but obtain according to the transformation , that is, the output is The corresponding matrix. Here, by ensuring that the transformation process is diffeomorphic, the number of operations can be reduced. At the same time, since the output is a matrix, this process is reversible, providing a role of verification.
[0196] The specific process of the registration process is as follows:
[0197] The original output of the linear registration module is the forward rotation, scaling, and shearing matrix and the forward translation vector . It can be directly predicted through the network, or the parameters of rotation, scaling, and shearing can be predicted separately and then combined. The reverse linear transformation relationship is obtained by inverting the linear transformation, i.e.:
[0198] (2)
[0199] (3)
[0200] Obtain and b, apply them to the mesh grid, and then subtract the mesh grid to obtain the deformation relationship in the form of a deformation field .
[0201] The original output of the deformation transformation is the forward deformation relationship in the form of a velocity field , and based on this, the deformation relationship in the form of a deformation field is obtained through integration, where the deformation relationship satisfies the following formulas (4) and (5):
[0202] (4)
[0203] (5)
[0204] where represents the application of the deformation relationship.
[0205] The combination module accumulates the of each previous step to obtain the cumulative deformation relationship , that is (note that for forward and reverse deformations, the combination order is opposite):
[0206] , (6)
[0207] , (7)
[0208] where and are the initial deformation relationships, and the final network output is also obtained by accumulating through the combination module.
[0209] Thus, in the embodiments of the present disclosure, after performing deformation initialization, a registration neural network is used to perform multi-level fitting residuals, which not only has the advantage of fast registration of neural network learning, but also can avoid the disadvantage of low accuracy when the global deformation is too large in neural network learning, making the registration fast and accurate.
[0210] S103: Based on the second transformation relationship information extracted corresponding to the last-level feature, register the first image and the second image.
[0211] Specifically, after performing multi-level iterative registration, based on the second transformation relationship information output after the last-level registration, that is, the final transformation information, the first image and the second image can be registered.
[0212] When registering the first image and the second image, for example, the second transformation relationship information corresponding to the last-level feature extraction can be used to perform a transformation process on the first image to obtain the transformed image of the first image, and then the position of the transformed image and the second image are matched one by one to obtain the registration result.
[0213] Alternatively, the second transformation relationship information corresponding to the last-level feature extraction can be used to perform a transformation process on the second image to obtain the transformed image of the second image, and then the position of the transformed image and the first image are matched one by one to obtain the registration result.
[0214] In the embodiments of the present disclosure, the registration method involved is applied to a pre-trained registration neural network, and the registration neural network includes two branch networks: a feature extraction neural network and a multi-level registration neural network; the feature extraction neural network is used to perform multi-level feature extraction processing on the first image and the second image respectively; the multi-level registration neural network is used to determine the target transformation relationship information between the first image and the second image based on the multi-level features extracted from the first image and the second image, where the target transformation relationship information is used to register the first image and the second image.
[0215] See Figure 6 As shown, the embodiments of the present disclosure provide a structural example of a registration neural network. Taking the above Figure 3 as an example, the registration neural network includes: the first, second, third, and fourth feature maps respectively corresponding to the first image and the second image output by the feature extraction neural network, and the above first to fourth feature maps are used as the input data of the registration module in the registration neural network, as Figure 6As shown in the figure, the input data of the first registration module is the first transformation relationship information, that is, the initial transformation relationship information, and the first feature maps respectively from the first image and the second image. Then, the registration data of the first registration module is input into the first combination module to obtain the first-level registration relationship; the input data of the second registration module is the first registration relationship output by the first combination module, and the second feature maps respectively from the first image and the second image. Then, the registration data of the second registration module is input into the second combination module to obtain the second-level registration relationship; the input data of the third registration module is the second registration relationship output by the second combination module, and the third feature maps respectively from the first image and the second image. Then, the registration data of the third registration module is input into the third combination module to obtain the third-level registration relationship; the input data of the fourth registration module is the third registration relationship output by the third combination module, and the third feature maps respectively from the first image and the second image. Then, the registration data of the fourth registration module is input into the fourth combination module to obtain the second transformation relationship information for registering the first image and the second image, that is, the target transformation relationship information.
[0216] The registration method, device, computer device and storage medium provided by the disclosed embodiments perform multi-level feature extraction on the first image and the second image, and perform feature registration based on the extracted multi-level features to obtain the registration relationship between the first image and the second image. In this way, the transformation relationship information between the first target feature map and the second target feature map obtained by iteratively determining multi-level feature extraction is used, and the transformation relationship information determined at the last level is used to register the first image and the second image, thereby improving the registration accuracy and registration efficiency.
[0217] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0218] Based on the same inventive concept, a registration device corresponding to the registration method is also provided in the disclosed embodiments. Since the principle of solving problems by the device in the disclosed embodiments is similar to the above registration method in the disclosed embodiments, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0219] Refer to Figure 7 、 Figure 8 、 Figure 9 、 Figure 10 shown in Figure 7 is a schematic diagram of an image registration device provided by the disclosed embodiments; Figure 8In the image registration device provided by the embodiments of the present disclosure, a specific schematic diagram of the extraction module; Figure 9 In the extraction module provided by the embodiments of the present disclosure, a specific schematic diagram of the encoding unit; Figure 10 In the extraction module provided by the embodiments of the present disclosure, a specific schematic diagram of the fusion unit. As Figure 7 shown, the registration device includes: an extraction module 710, a determination module 720, and a registration module 730, where:
[0220] The extraction module 710 is configured to perform multi-level feature extraction on a first image to obtain first target feature maps respectively corresponding to the multi-level feature extraction; perform multi-level feature extraction on a second image to obtain second target feature maps respectively corresponding to the multi-level feature extraction;
[0221] The determination module 720 is configured to, for each level of feature extraction, based on the first target feature map, the second target feature map corresponding to this level of feature extraction, and the first transformation relationship information corresponding to this level of feature extraction, determine the second transformation relationship information corresponding to this level of feature extraction; where the first transformation relationship information corresponding to this level of feature extraction includes: the second transformation relationship information corresponding to the previous level of feature extraction, or the original transformation relationship information between the first image and the second image; and
[0222] The registration module 730 is configured to register the first image and the second image based on the second transformation relationship information corresponding to the last level of feature extraction.
[0223] In an optional implementation manner, as Figure 8 shown, the extraction module 710 includes:
[0224] An extraction unit 711, configured to, for each level of feature extraction in the multi-level feature extraction, determine the first input data and the second input data for this level of feature extraction; the first input data includes: the image, or the encoded feature map corresponding to the next level of feature extraction; the second input data includes: the decoded feature map corresponding to the previous level of feature extraction;
[0225] An encoding unit 712, configured to perform encoding processing corresponding to this level of feature extraction on the first input data to obtain an encoded feature map corresponding to this level of feature extraction;
[0226] A fusion unit 713, configured to perform fusion processing on the encoded feature map corresponding to this level of feature extraction and the second input data to obtain a fused feature map;
[0227] A decoding unit 714, configured to perform decoding processing corresponding to this level of feature extraction on the fused feature map to obtain a decoded feature map corresponding to this level of feature extraction;
[0228] A determination unit 715 is configured to determine the decoded feature map corresponding to the feature extraction at this level as the target feature map corresponding to the feature extraction at this level;
[0229] The image includes: the first image and / or the second image;
[0230] The target feature map includes: the first target feature map and / or the second target feature map.
[0231] In an alternative embodiment, as Figure 9 shown, the encoding unit 712 includes:
[0232] A downsampling sub-unit 7111 is configured to perform downsampling processing on the first input data to obtain a downsampled feature map;
[0233] A channel attention processing sub-unit 7112 is configured to perform channel attention processing on the downsampled feature map to obtain the attention weights corresponding to each data channel in the downsampled feature map;
[0234] A first determination sub-unit 7113 is configured to obtain the encoded feature map based on the downsampled feature map and the attention weights.
[0235] In an alternative embodiment, the channel attention processing sub-unit 7112 is specifically configured to:
[0236] Perform global average pooling processing on the downsampled feature map to obtain a first feature sub-map;
[0237] Based on the first feature sub-map, determine the candidate attention weights corresponding to each data channel; and based on the first feature sub-map, determine the feature domain weights of the candidate attention corresponding to each data channel;
[0238] Based on the candidate attention weights corresponding to each data channel and the feature domain weights of the candidate attention corresponding to each data channel, obtain the attention weights corresponding to each data channel in the downsampled feature map.
[0239] In an alternative embodiment, as Figure 10 shown, the fusion unit 713 includes:
[0240] A second determination sub-unit 7131 is configured to obtain the weights corresponding to each feature point in the encoded feature map corresponding to the feature extraction at this level based on the encoded feature map corresponding to the feature extraction at this level and the second input data;
[0241] A third determination subunit 7132, configured to obtain a second feature subgraph based on the weights corresponding to each feature point in the encoded feature map corresponding to the feature extraction at this level and the encoded feature map;
[0242] A splicing subunit 7133, configured to splice the second feature subgraph and the second input data to obtain the fused feature map.
[0243] In an optional implementation manner, the second determination subunit 7131 is specifically configured to:
[0244] Splice the encoded feature map corresponding to the feature extraction at this level and the second input data to obtain a third feature subgraph;
[0245] Perform convolution processing on the third feature subgraph to obtain a fourth feature subgraph, and based on the fourth feature subgraph, obtain the encoded feature map and the local autocorrelation coefficient of the second input data;
[0246] Based on the local autocorrelation coefficient, obtain the weights corresponding to each feature point in the encoded feature map corresponding to the feature extraction at this level.
[0247] In an optional implementation manner, the feature value of any feature point in the third feature subgraph represents the autocorrelation coefficient of the image region corresponding to the feature point;
[0248] The second determination subunit 7131 is specifically configured to:
[0249] Use an activation function to perform activation processing on the autocorrelation coefficient to obtain the gating activation values corresponding to each feature point in the encoded feature map; the gating activation values are used to represent the weights corresponding to each feature point in the encoded feature map.
[0250] In an optional implementation manner, the second determination subunit 7131 is specifically configured to:
[0251] Perform maximum processing on the channel dimension of the fourth feature subgraph to obtain the maximum value of the channel dimension of the fourth feature subgraph;
[0252] Perform average processing on the channel dimension of the fourth feature subgraph to obtain the average value of the channel dimension of the fourth feature subgraph; and
[0253] Perform splicing processing on the maximum value of the channel dimension and the average value of the channel dimension to obtain a channel dimension splicing result, and perform convolution and normalization processing on the channel dimension splicing result to obtain the encoded feature map and the local autocorrelation coefficient of the second input data.
[0254] In an optional implementation manner, the determination module 720 is specifically configured to:
[0255] Based on the first target feature map, the second target feature map, and the first transformation relationship information, determine the transformation residual between the feature extraction at this level and the feature extraction at the previous level;
[0256] Based on the transformation residual and the first transformation relationship information, determine the second transformation relationship information corresponding to the feature extraction at this level.
[0257] In an optional implementation manner, the registration device is applied to a pre-trained registration neural network, and the registration neural network includes two branch networks, namely a feature extraction neural network and a multi-level registration neural network; the feature extraction neural network is used to perform multi-level feature extraction processing on the first image and the second image respectively; the multi-level registration neural network is used to determine the target transformation relationship information between the first image and the second image based on the multi-level features extracted from the first image and the second image, where the target transformation relationship information is used to register the first image and the second image.
[0258] In the embodiments of the present disclosure, multi-level feature extraction is performed on the first image and the second image, and feature registration is performed based on the extracted multi-level features to obtain the registration relationship between the first image and the second image. In this way, by iteratively determining the transformation relationship information between the first target feature map and the second target feature map obtained by multi-level feature extraction, and using the transformation relationship information determined at the last level, the first image and the second image are registered, thereby improving the registration accuracy and registration efficiency.
[0259] For the processing flow of each module in the device and the interaction flow between each module, reference may be made to the relevant descriptions in the above method embodiments, which will not be elaborated here.
[0260] Corresponding to Figure 1 the registration method in, the embodiments of the present disclosure further provide a computer device, as Figure 11 shown, which is a schematic structural diagram of the computer device provided by the embodiments of the present disclosure, including:
[0261] A processor 101, a memory 102, and a bus 103; the memory 102 is used to store execution instructions, including an internal memory 1021 and an external memory 1022; here, the internal memory 1021 is also called the main memory, which is used to temporarily store the operation data in the processor 101 and the data exchanged with the external memory 1022 such as a hard disk. The processor 101 exchanges data with the external memory 1022 through the internal memory 1021. When the computer device runs, the processor 101 communicates with the memory 102 through the bus 103, so that the processor 101 executes the following instructions:
[0262] Perform multi-level feature extraction on the first image to obtain first target feature maps respectively corresponding to the multi-level feature extraction; perform multi-level feature extraction on the second image to obtain second target feature maps respectively corresponding to the multi-level feature extraction;
[0263] For each level of feature extraction, based on the first target feature map corresponding to this level of feature extraction, the second target feature map corresponding to this level of feature extraction, and the first transformation relation information corresponding to this level of feature extraction, determine the second transformation relation information corresponding to this level of feature extraction; wherein, the first transformation relation information corresponding to this level of feature extraction includes: the second transformation relation information corresponding to the previous level of feature extraction, or the original transformation relation information between the first image and the second image; and
[0264] Register the first image and the second image based on the second transformation relation information corresponding to the last level of feature extraction.
[0265] The embodiments of the present disclosure further provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the registration method described in the above method embodiments. Wherein, the storage medium may be a volatile or non-volatile computer-readable storage medium.
[0266] The embodiments of the present disclosure further provide a computer program product, which carries program codes, and the instructions included in the program codes can be used to execute the steps of the registration method described in the above method embodiments. For details, refer to the above method embodiments, and details will not be repeated here.
[0267] Wherein, the above computer program product can be specifically implemented in a manner of hardware, software or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0268] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In the several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0269] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0270] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0271] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0272] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting it. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions described in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for registering images, characterized in that, Including: Performing multi-level feature extraction on a first image to obtain first target feature maps respectively corresponding to the multi-level feature extraction; Performing multi-level feature extraction on a second image to obtain second target feature maps respectively corresponding to the multi-level feature extraction; wherein, each level of feature extraction includes an encoding network and a decoding network corresponding to this level of feature extraction. The encoding network is used to perform downsampling on the encoding feature map of the subsequent level of feature extraction to obtain a downsampled feature map, perform channel attention processing on the downsampled feature map to obtain attention weights respectively corresponding to each data channel in the downsampled feature map, and obtain the encoding feature map corresponding to this level of feature extraction based on the downsampled feature map and the attention weights. The decoding network is used to perform decoding processing on the encoding feature map corresponding to this level of feature extraction and the decoding feature map of the previous level of feature extraction to obtain the decoding feature map corresponding to this level of feature extraction. The decoding feature map corresponding to this level of feature extraction is the target feature map corresponding to this level of feature extraction. Performing channel attention processing on the downsampled feature map to obtain attention weights respectively corresponding to each data channel in the downsampled feature map includes: performing global average pooling processing on the downsampled feature map to obtain a first feature sub-map; determining candidate attention weights respectively corresponding to each data channel based on the first feature sub-map; and determining feature domain weights of candidate attention respectively corresponding to each data channel based on the first feature sub-map; obtaining the attention weights respectively corresponding to each data channel in the downsampled feature map based on the candidate attention weights respectively corresponding to each data channel and the feature domain weights of candidate attention respectively corresponding to each data channel; For each level of feature extraction, determining second transformation relationship information corresponding to this level of feature extraction based on the first target feature map, the second target feature map, and the first transformation relationship information corresponding to this level of feature extraction; wherein, the first transformation relationship information corresponding to this level of feature extraction includes: the second transformation relationship information corresponding to the previous level of feature extraction, or the original transformation relationship information between the first image and the second image; and Registering the first image and the second image based on the second transformation relationship information corresponding to the last level of feature extraction.
2. The registration method according to claim 1, wherein Performing multi-level feature extraction on an image to obtain feature maps respectively corresponding to the multi-level feature extraction, including: For each level of feature extraction in the multi-level feature extraction, determining first input data and second input data for this level of feature extraction; the first input data includes: the image, or the encoding feature map corresponding to the subsequent level of feature extraction; the second input data includes: the decoding feature map corresponding to the previous level of feature extraction; Performing encoding processing corresponding to this level of feature extraction on the first input data to obtain the encoding feature map corresponding to this level of feature extraction; Performing fusion processing on the encoding feature map corresponding to this level of feature extraction and the second input data to obtain a fused feature map; Performing decoding processing corresponding to this level of feature extraction on the fused feature map to obtain the decoding feature map corresponding to this level of feature extraction; Extract the corresponding decoded feature map of this level of feature extraction and determine it as the target feature map corresponding to this level of feature extraction; The image includes: the first image and / or the second image; The target feature map includes: the first target feature map and / or the second target feature map.
3. The registration method according to claim 2, wherein The encoding process corresponding to this level of feature extraction on the first input data to obtain the encoding feature map corresponding to this level of feature extraction includes: Perform downsampling processing on the first input data to obtain a downsampled feature map; Perform channel attention processing on the downsampled feature map to obtain the attention weights corresponding to each data channel in the downsampled feature map; Based on the downsampled feature map and the attention weights, obtain the encoding feature map.
4. The registration method according to any one of claims 2-3, characterized in that The fusion process on the encoding feature map corresponding to this level of feature extraction and the second input data to obtain a fusion feature map includes: Based on the encoding feature map corresponding to this level of feature extraction and the second input data, obtain the weights corresponding to each feature point in the encoding feature map corresponding to this level of feature extraction; Based on the weights corresponding to each feature point in the encoding feature map corresponding to this level of feature extraction and the encoding feature map, obtain a second feature sub-map; Stitch the second feature sub-map and the second input data to obtain the fusion feature map.
5. The registration method according to claim 4, wherein The obtaining of the weights corresponding to each feature point in the encoding feature map corresponding to this level of feature extraction based on the encoding feature map corresponding to this level of feature extraction and the second input data includes: Stitch the encoding feature map corresponding to this level of feature extraction and the second input data to obtain a third feature sub-map; Perform convolution processing on the third feature sub-map to obtain a fourth feature sub-map, and based on the fourth feature sub-map, obtain the encoding feature map and the local autocorrelation coefficient with the second input data; Based on the local autocorrelation coefficient, obtain the weights corresponding to each feature point in the encoding feature map corresponding to this level of feature extraction.
6. The registration method according to claim 5, wherein The obtaining of the weights corresponding to each feature point in the encoding feature map corresponding to this level of feature extraction based on the local autocorrelation coefficient includes: Use an activation function to perform activation processing on the autocorrelation coefficient to obtain the weights corresponding to each feature point in the encoding feature map.
7. The registration method according to claim 5, characterized in that The obtaining of the encoding feature map and the local autocorrelation coefficient with the second input data based on the fourth feature sub-map includes: Perform maximum processing on the channel dimension of the fourth feature sub-map to obtain the maximum value of the channel dimension of the fourth feature sub-map; Perform average processing on the channel dimension of the fourth feature sub-map to obtain the average value of the channel dimension of the fourth feature sub-map; and Perform stitching processing on the maximum value of the channel dimension and the average value of the channel dimension to obtain a channel dimension stitching result, and perform convolution and normalization processing on the channel dimension stitching result to obtain the encoding feature map and the local autocorrelation coefficient with the second input data.
8. The registration method according to any one of claims 1-7, characterized in that, For each level of feature extraction, based on the corresponding first target feature map, second target feature map, and first transformation relationship information of this level of feature extraction, determining the second transformation relationship information corresponding to this level of feature extraction includes: Based on the first target feature map, the second target feature map, and the first transformation relationship information, determining the transformation residual between this level of feature extraction and the previous level of feature extraction; Based on the transformation residual and the first transformation relationship information, determining the second transformation relationship information corresponding to this level of feature extraction.
9. The registration method according to any one of claims 1-8, characterized in that, Including: The registration method is applied to a pre-trained registration neural network, and the registration neural network includes a feature extraction neural network and a multi-level registration neural network; The feature extraction neural network is used to perform multi-level feature extraction processing on the first image and the second image respectively; The multi-level registration neural network is used to determine the target transformation relationship information between the first image and the second image based on the multi-level features extracted from the first image and the second image, where the target transformation relationship information is used to register the first image and the second image.
10. An image registration device, characterized in that, Including: An extraction module, configured to perform multi-level feature extraction on a first image to obtain first target feature maps corresponding to the multi-level feature extraction respectively; Performing multi-level feature extraction on a second image to obtain second target feature maps corresponding to the multi-level feature extraction respectively; wherein, each level of feature extraction includes an encoding network and a decoding network corresponding to this level of feature extraction. The encoding network is configured to perform downsampling on the encoding feature map of the subsequent level of feature extraction to obtain a downsampled feature map, perform channel attention processing on the downsampled feature map to obtain attention weights corresponding to each data channel in the downsampled feature map, and obtain the encoding feature map corresponding to this level of feature extraction based on the downsampled feature map and the attention weights. The decoding network is configured to perform decoding processing based on the encoding feature map corresponding to this level of feature extraction and the decoding feature map of the previous level of feature extraction to obtain the decoding feature map corresponding to this level of feature extraction. The decoding feature map corresponding to this level of feature extraction is the target feature map corresponding to this level of feature extraction. Performing channel attention processing on the downsampled feature map to obtain attention weights corresponding to each data channel in the downsampled feature map includes: performing global average pooling processing on the downsampled feature map to obtain a first feature sub-map; based on the first feature sub-map, determining candidate attention weights corresponding to each data channel respectively; and based on the first feature sub-map, determining the feature domain weights of the candidate attention corresponding to each data channel respectively; based on the candidate attention weights corresponding to each data channel respectively and the feature domain weights of the candidate attention corresponding to each data channel respectively, obtaining the attention weights corresponding to each data channel in the downsampled feature map; A determination module, configured to perform feature extraction for each level, and determine second transformation relation information corresponding to the feature extraction at this level based on a first target feature map, a second target feature map corresponding to the feature extraction at this level, and first transformation relation information corresponding to the feature extraction at this level; wherein the first transformation relation information corresponding to the feature extraction at this level includes: second transformation relation information corresponding to the feature extraction at the previous level, or the original transformation relation information between the first image and the second image; and A registration module, configured to register the first image and the second image based on the second transformation relation information corresponding to the feature extraction at the last level.
11. A computer device, characterized in that, Comprising: A processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the registration method according to any one of claims 1 to 9 are performed.
12. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of the registration method according to any one of claims 1 to 9 are performed.
Citation Information
Patent Citations
End-to-end classification method for single-voyage InSAR system based on multistage deep learning network
CN112083422A
Image Processing Method, Image Processing Apparatus, and Device
US20220319155A1