Multimodal image registration method, system and computer device based on feature decoupling
By using a feature-decoupled neural network model in multimodal image registration, the sharing and unique features of images are extracted, and the problem of difficulty in dealing with modal differences is solved in traditional methods, and a high-precision and robust multimodal image registration is achieved.
Patent Information
- Application Number
- CN202510170563.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Traditional multimodal image registration methods are difficult to adapt to the significant differences between different modes, resulting in limited registration accuracy and robustness.
Using a multimodal image registration method based on feature decoupling, a neural network model containing a general information extraction module, a feature decoupling module and a registration parameter estimation module are used to extract the shared feature map and unique feature map of the image, and a preset loss function is used for model training to obtain the image conversion parameter matrix.
The performance of multimodal image registration is improved, especially in the face of large differences between different modes, which shows good robustness and achieve high-precision image registration.
Smart Images

Figure CN119625039B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image data processing, and in particular relates to a multimodal image registration method, system and computer equipment based on feature decoupling. Background Art
[0002] With the development of various sensor technologies, multimodal images acquired by different platforms are increasingly used in various fields. Multimodal images refer to images acquired by different types of sensors (such as optical cameras, infrared cameras, lidar, etc.). These images have different perception mechanisms and can provide complementary environmental information. Multimodal image registration is to match images of different modalities with reference images to achieve the fusion and utilization of multi-source data. It has important applications in scenarios such as remote sensing image analysis, autonomous driving, and robot navigation.
[0003] However, due to the significant differences in perception mechanism, radiation intensity, and geometric characteristics of multimodal images, high-precision multimodal image registration has become a challenging task. Traditional registration methods usually rely on direct feature alignment between images, but for cross-modal images, this method is often difficult to apply due to the obvious differences in features between modalities, thus limiting the further improvement of registration accuracy and system robustness.
[0004] In order to overcome the above challenges, the present invention proposes a multimodal image registration method based on feature decoupling, which can effectively improve the performance of multimodal image registration, especially in the case of large differences between different modalities, and shows good robustness. Summary of the invention
[0005] In view of the above technical problems, the present invention provides a multimodal image registration method, system and computer device based on feature decoupling.
[0006] The technical solution adopted by the present invention to solve the technical problem is:
[0007] A multimodal image registration method based on feature decoupling, the method comprising the following steps:
[0008] S100: acquiring a plurality of multimodal images and resizing the images, dividing the images with the same size after resizing into reference images and images to be registered and forming a plurality of image pairs, thereby obtaining a multimodal image registration dataset with diversity;
[0009] S200: Building a multimodal image registration neural network model based on feature decoupling, the model includes a general information extraction module, a feature decoupling module and a registration parameter estimation module, the feature decoupling module includes a shared feature mapping branch and a unique feature mapping branch;
[0010] S300: Input a pair of multimodal images in the data set to the general information extraction module for processing, and obtain the general feature map of the reference image in the selected image pair and the general feature map of the image to be registered; input the general feature map of the reference image and the image to be registered to the feature decoupling module for processing, and extract the shared feature map and the unique feature map of the reference image and the image to be registered respectively; input the shared feature map of the reference image and the shared feature map of the image to be registered to the registration parameter estimation module for processing, and output the estimated conversion parameter matrix of the selected image pair;
[0011] S400: Calculate the total loss function value of each image pair based on the shared feature map and the unique feature map of the reference image and the image to be registered, the estimated conversion parameter matrix of the image pair and the preset loss function, supervise the training process of the multimodal image registration neural network model, select the network parameters with the minimum loss value to update the multimodal image registration neural network model, and obtain the trained multimodal image registration neural network model;
[0012] S500: Acquire multimodal images to be registered in a real scene and form image pairs to be registered, use a trained multimodal image registration neural network model to process the multimodal image pairs to be registered, obtain an image conversion parameter matrix, and convert the images to be registered in the image pairs according to the image conversion parameter matrix to obtain registered images.
[0013] Preferably, in S100, a plurality of multimodal images are acquired and image size adjustment is performed, and the image size adjustment is specifically as follows:
[0014] (1)
[0015] In the formula, is the pixel value of the original image, is the position after resizing, is the set of neighborhood pixels considered during the resizing process, Is with the field pixel The relevant weight function, The new image after resizing is at position The pixel value of .
[0016] Preferably, the general information extraction module in S200 includes a convolution module, a normalization layer, an activation function layer, and three residual modules connected in sequence, wherein the convolution module includes three convolution layers connected in sequence, the step size of each convolution layer is 2, and the convolution kernel size is 3. 3. The residual module consists of two convolutional layers connected in sequence. The step size of each convolutional layer is 1 and the convolution kernel size is 3. 3. Interlayer pass 1 1. Convolution adjusts the dimension to achieve residual connection; the general feature map of the reference image and the general feature map of the image to be registered in S200 are specifically:
[0017] (2)
[0018] In the formula, represents the general information extraction module, and Respectively represent the reference image and the image to be registered in the selected image pair, and They respectively represent the universal feature map of the reference image in the selected image pair and the universal feature map of the image to be registered.
[0019] Preferably, the shared feature mapping branch includes a convolution layer, a normalization layer, and an activation function layer connected in sequence, wherein the step size of the convolution layer is 1 and the convolution kernel size is 1. 1; The unique feature mapping branch includes a convolutional layer, a normalization layer, an activation function layer, and two residual modules connected in sequence, where the step size of the convolutional layer is 1 and the convolution kernel size is 1 1. The residual module consists of two convolutional layers connected in sequence. The stride of each convolutional layer is 1 and the convolution kernel size is 1. 1; The shared feature map and the unique feature map of the reference image and the image to be registered in S300 are specifically expressed as:
[0020] (3)
[0021] (4)
[0022] In the formula, represents the shared feature map branch, represents a unique feature mapping branch, and They represent the universal feature map of the reference image in the selected image pair and the universal feature map of the image to be registered, and Represent the unique feature map and shared feature map of the reference image respectively, and They represent the unique feature map and shared feature map of the image to be registered respectively.
[0023] Preferably, the registration parameter estimation module includes a neighborhood cost volume layer, a first convolution layer, a normalization layer, an activation function layer, a maximum pooling layer and a second convolution layer connected in sequence, wherein the step size of the first and second convolution layers is 1, and the convolution kernel size is 1. 1; The image conversion parameter matrix in S300 is specifically expressed as:
[0024] (5)
[0025] In the formula, represents the image transformation parameter matrix for the image pair, Representation 1 1 convolutional layer, represents the maximum pooling layer, represents the neighborhood cost volume construction operator, and They represent the shared feature map of the reference image and the shared feature map of the image to be registered respectively.
[0026] Preferably, the loss function includes a similarity loss function, a difference loss function, a reconstruction loss function and a registration true value loss function. The loss function designed in S400 to calculate the total loss function value includes sequentially calculating the similarity loss function value of the shared feature map of the reference image and the shared feature map of the image to be registered, the difference loss function value and the reconstruction loss value of the shared feature map of the reference image and its unique feature map, the difference loss function value and the reconstruction loss value of the shared feature map of the image to be registered and its unique feature map, and the registration true value loss function value of the estimated transformation parameter matrix and the true value transformation parameter matrix.
[0027] Preferably, the loss function designed in S400 calculates the total loss function value as follows:
[0028] (6)
[0029] (7)
[0030] (8)
[0031] (9)
[0032] (10)
[0033] In the formula, Represents the total loss function calculation value, Represents the calculated value of the similarity loss function, Represents the calculated value of the difference loss function, Represents the calculated value of the reconstruction loss function, Represents the calculated value of the registration truth loss function, and Represent the unique feature map and shared feature map of the reference image respectively, and Respectively represent the unique feature map and shared feature map of the image to be registered, is the empirical expectation vector, yes The central moment vector of the sample is represents the Euclidean norm, and Respectively represent the starting point and end point of the interval used for normalization calculation, represents the trace of the matrix, represents the transpose of a matrix, and represent the reference image and the image to be registered, respectively. represents the 1 norm, and are the estimated conversion parameter matrix and the true value conversion parameter matrix respectively.
[0034] Preferably, in S500, the image to be registered in the image pair is transformed according to the image transformation parameter matrix, specifically:
[0035] (11)
[0036] In the formula, represents the image after the image to be registered is registered. represents the image to be registered, Represents the image conversion process, Represents the reference image in the real scene and the image to be registered The image transformation parameter matrix corresponding to the image pair to be registered.
[0037] A multimodal image registration system based on feature decoupling, including a data set determination module, a multimodal image registration neural network model building module, an estimated conversion parameter matrix determination module for selected image pairs, a model training module and an image registration module;
[0038] A data set determination module is used to obtain a number of multimodal images and adjust the image size, divide the images with the same size after adjustment into reference images and images to be registered and form a number of image pairs, so as to obtain a multimodal image registration data set with diversity;
[0039] A multimodal image registration neural network model building module is used to build a multimodal image registration neural network model based on feature decoupling. The model includes a general information extraction module, a feature decoupling module and a registration parameter estimation module. The feature decoupling module includes a shared feature mapping branch and a unique feature mapping branch.
[0040] The module for determining the estimated conversion parameter matrix of the selected image pair is used to input a pair of multimodal images in the data set into the general information extraction module for processing, and obtain the general feature map of the reference image and the general feature map of the image to be registered in the selected image pair; input the general feature map of the reference image and the image to be registered into the feature decoupling module for processing, and respectively extract the shared feature map and the unique feature map of the reference image and the image to be registered; input the shared feature map of the reference image and the shared feature map of the image to be registered into the registration parameter estimation module for processing, and output the estimated conversion parameter matrix of the selected image pair;
[0041] A model training module is used to calculate the total loss function value of each image pair based on the shared feature map and the unique feature map of the reference image and the image to be registered, the estimated conversion parameter matrix of the image pair and the preset loss function, supervise the training process of the multimodal image registration neural network model, select the network parameters with the smallest loss value to update the multimodal image registration neural network model, and obtain a trained multimodal image registration neural network model;
[0042] The image registration module is used to process the multimodal image to be registered using the trained multimodal image registration neural network model to obtain an image conversion parameter matrix, and transform the image to be registered in the image pair according to the image conversion parameter matrix to obtain a registered image.
[0043] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a multimodal image registration method based on feature decoupling when executing the computer program.
[0044] The above-mentioned multimodal image registration method, system and computer device based on feature decoupling obtain several multimodal images and adjust the image size, divide the images with the same size after adjustment into reference images and images to be registered and form several image pairs, obtain a multimodal image registration data set with diversity, build a multimodal image registration neural network model based on feature decoupling, design a loss function to supervise the training process of the multimodal image registration neural network model, use the trained multimodal image registration neural network model to process the multimodal images to be registered, obtain the image conversion parameter matrix, convert the images to be registered in the image pair according to the image conversion parameter matrix, and obtain the registered image. This method can effectively improve the performance of multimodal image registration, especially in the case of large differences between different modalities, and shows good robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a flow chart of a multimodal image registration method based on feature decoupling in one embodiment of the present invention;
[0046] Figure 2 Schematic diagram of the principle of a multimodal image registration method based on feature decoupling in one embodiment of the present invention. DETAILED DESCRIPTION
[0047] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings.
[0048] In one embodiment, Figure 1 As shown, a multimodal image registration method based on feature decoupling comprises the following steps:
[0049] S100: acquiring a plurality of multimodal images and resizing the images, dividing the images with the same size after resizing into reference images and images to be registered and forming a plurality of image pairs, thereby obtaining a multimodal image registration dataset with diversity;
[0050] S200: Building a multimodal image registration neural network model based on feature decoupling, the model includes a general information extraction module, a feature decoupling module and a registration parameter estimation module, the feature decoupling module includes a shared feature mapping branch and a unique feature mapping branch;
[0051] S300: Input a pair of multimodal images in the data set to the general information extraction module for processing, and obtain the general feature map of the reference image in the selected image pair and the general feature map of the image to be registered; input the general feature map of the reference image and the image to be registered to the feature decoupling module for processing, and extract the shared feature map and the unique feature map of the reference image and the image to be registered respectively; input the shared feature map of the reference image and the shared feature map of the image to be registered to the registration parameter estimation module for processing, and output the estimated conversion parameter matrix of the selected image pair;
[0052] S400: Calculate the total loss function value of each image pair based on the shared feature map and the unique feature map of the reference image and the image to be registered, the estimated conversion parameter matrix of the image pair and the preset loss function, supervise the training process of the multimodal image registration neural network model, select the network parameters with the minimum loss value to update the multimodal image registration neural network model, and obtain the trained multimodal image registration neural network model;
[0053] S500: Acquire a multimodal image to be registered, use a trained multimodal image registration neural network model to process the multimodal image to be registered, obtain an image conversion parameter matrix, and convert the image to be registered in the image pair according to the image conversion parameter matrix to obtain a registered image.
[0054] In one embodiment, in S100, a plurality of multimodal images are acquired and image size adjustment is performed, and the image size adjustment is specifically as follows:
[0055] (1)
[0056] In the formula, is the pixel value of the original image, is the position after resizing, is the set of neighborhood pixels considered during the resizing process, Is with the field pixel The relevant weight function, The new image after resizing is at position The pixel value of .
[0057] In one embodiment, the general information extraction module in S200 includes a convolution module, a normalization layer, an activation function layer, and three residual modules connected in sequence, wherein the convolution module includes three convolution layers connected in sequence, the step size of each convolution layer is 2, and the convolution kernel size is 3. 3. The residual module consists of two convolutional layers connected in sequence. The step size of each convolutional layer is 1 and the convolution kernel size is 3. 3. Interlayer pass 1 1. Convolution adjusts the dimension to achieve residual connection; the general feature map of the reference image and the general feature map of the image to be registered in S200 are specifically:
[0058] (2)
[0059] In the formula, represents the general information extraction module, and Respectively represent the reference image and the image to be registered in the selected image pair, and They respectively represent the universal feature map of the reference image in the selected image pair and the universal feature map of the image to be registered.
[0060] Specifically, the general information extraction module is used to extract useful feature maps from the input image and provide high-quality feature representation for subsequent feature decoupling and registration parameter estimation. The convolution module greatly reduces the spatial dimension of the input data through multi-layer convolution operations, while increasing the depth of the feature map to quickly capture the basic feature maps in the input data. The normalization layer is used to accelerate the training process and improve the generalization ability of the model. The activation function layer is used to introduce nonlinear transformations so that the model can learn and represent complex function mappings. The residual module is used to avoid the gradient vanishing / exploding problem during training through residual connections, while enhancing the expressiveness of feature maps.
[0061] In one embodiment, the shared feature map branch includes a convolution layer, a normalization layer, and an activation function layer connected in sequence, wherein the step size of the convolution layer is 1 and the convolution kernel size is 1. 1; The unique feature mapping branch includes a convolutional layer, a normalization layer, an activation function layer, and two residual modules connected in sequence, where the step size of the convolutional layer is 1 and the convolution kernel size is 1 1. The residual module consists of two convolutional layers connected in sequence. The stride of each convolutional layer is 1 and the convolution kernel size is 1. 1; The shared feature map and the unique feature map of the reference image and the image to be registered in S300 are specifically expressed as:
[0062] (3)
[0063] (4)
[0064] In the formula, represents the shared feature map branch, represents a unique feature mapping branch, and They represent the universal feature map of the reference image in the selected image pair and the universal feature map of the image to be registered, and Represent the unique feature map and shared feature map of the reference image respectively, and They represent the unique feature map and shared feature map of the image to be registered respectively.
[0065] Specifically, the feature decoupling module is used to decompose the extracted common feature maps into shared feature maps and unique feature maps so as to better handle the differences between multimodal images.
[0066] In one embodiment, the registration parameter estimation module includes a neighborhood cost volume layer, a first convolution layer, a normalization layer, an activation function layer, a maximum pooling layer and a second convolution layer connected in sequence, wherein the step size of the first and second convolution layers is 1, and the convolution kernel size is 1. 1; The image conversion parameter matrix in S300 is specifically expressed as:
[0067] (5)
[0068] In the formula, represents the image transformation parameter matrix for the image pair, Representation 1 1 convolutional layer, represents the maximum pooling layer, represents the neighborhood cost volume construction operator, and They represent the shared feature map of the reference image and the shared feature map of the image to be registered respectively.
[0069] Specifically, the registration parameter estimation module is used to estimate the registration parameters between multimodal images based on the extracted shared feature maps to achieve precise alignment of the images. The neighborhood cost volume layer is used to calculate the local similarity of the feature maps and generate the neighborhood cost volume. The first convolutional layer is used to adjust the depth of the feature maps to make them suitable for subsequent processing. The normalization layer is used to accelerate the training process and improve the generalization ability of the model. The activation function layer introduces nonlinear transformations. The maximum pooling layer is used to reduce the spatial dimension of the feature map while retaining the most important feature information. The second convolutional layer is used to further adjust the depth of the feature map to provide input for the final registration parameter estimation.
[0070] In one embodiment, the loss function includes a similarity loss function, a difference loss function, a reconstruction loss function and a registration true value loss function. The loss function designed in S400 to calculate the total loss function value includes sequentially calculating the similarity loss function value of the shared feature map of the reference image and the shared feature map of the image to be registered, the difference loss function value and the reconstruction loss value of the shared feature map of the reference image and its unique feature map, the difference loss function value and the reconstruction loss value of the shared feature map of the image to be registered and its unique feature map, and the registration true value loss function value of the estimated transformation parameter matrix and the true value transformation parameter matrix.
[0071] Specifically, the similarity loss is used to align the shared feature maps of the reference image and the image to be registered to ensure consistency between the modalities; the difference loss ensures the independence of shared features and unique features to avoid information confusion; the reconstruction loss ensures the integrity of the modal characteristic information by measuring the error of image reconstruction caused by unique features; and the registration truth loss directly evaluates the difference between the estimated transformation parameters and the true transformation parameters to guide the model to optimize the registration accuracy. These losses work together to improve the accuracy and robustness of the model for multimodal image registration.
[0072] In one embodiment, the loss function is designed in S400 to calculate the total loss function value as follows:
[0073] (6)
[0074] (7)
[0075] (8)
[0076] (9)
[0077] (10)
[0078] In the formula, Represents the total loss function calculation value, Represents the calculated value of the similarity loss function, Represents the calculated value of the difference loss function, Represents the calculated value of the reconstruction loss function, Represents the calculated value of the registration truth loss function, and Represent the unique feature map and shared feature map of the reference image respectively, and Respectively represent the unique feature map and shared feature map of the image to be registered, is the empirical expectation vector, yes The central moment vector of the sample is represents the Euclidean norm, and Respectively represent the starting point and end point of the interval used for normalization calculation, represents the trace of the matrix, represents the transpose of a matrix, and represent the reference image and the image to be registered, respectively. represents the 1 norm, and are the estimated conversion parameter matrix and the true value conversion parameter matrix respectively.
[0079] Specifically, when training the multimodal image registration neural network model, for each pair of images, the model generates an image transformation parameter matrix And calculate the corresponding loss value The smaller the loss value, the better the alignment effect between the transformed image to be matched and the reference image. The model weights are updated by selecting the network parameters corresponding to the minimum loss value, and the network performance is gradually optimized. After training, the obtained multimodal image registration neural network model can accurately estimate the transformation parameter matrix of the image, thereby achieving high-precision multimodal image registration.
[0080] In one embodiment, in S500, the image to be registered in the image pair is transformed according to the image transformation parameter matrix, specifically:
[0081] (11)
[0082] In the formula, represents the image after the image to be registered is registered. represents the image to be registered, Represents the image conversion process, Represents the reference image in the real scene and the image to be registered The image transformation parameter matrix corresponding to the image pair to be registered.
[0083] Specifically, see Figure 2 , Figure 2 Schematic diagram of a multimodal image registration process in a real scene according to an embodiment of the present invention.
[0084] Obtain multimodal images to be registered in real scenes and form image pairs to be registered and , the image pair to be registered is input into the multimodal image registration neural network model based on feature decoupling, and the image conversion parameter matrix corresponding to the image pair to be registered can be obtained Specifically, refer to formula (2), according to the reference image in the image pair to be registered in the real scene and the image to be registered The corresponding general feature map can be obtained and , and then refer to formulas (3) and (4) to obtain the unique feature map and shared feature map of the reference image and , the unique feature map and shared feature map of the image to be registered and , then refer to formula (5) to obtain the image transformation parameter matrix of the image pair to be registered , and then calculate according to formula (11) to obtain the registered image of the image to be registered in the real scene.
[0085] The above-mentioned multimodal image registration method based on feature decoupling obtains several multimodal images and resizes them, divides the images with the same size after resize into reference images and images to be registered and forms several image pairs, obtains a diverse multimodal image registration data set, builds a multimodal image registration neural network model based on feature decoupling, designs a loss function to supervise the training process of the multimodal image registration neural network model, uses the trained multimodal image registration neural network model to process the multimodal images to be registered, obtains the image conversion parameter matrix, converts the images to be registered in the image pair according to the image conversion parameter matrix, and obtains the registered image. This method can effectively improve the performance of multimodal image registration, especially in the case of large differences between different modalities, and shows good robustness.
[0086] A multimodal image registration system based on feature decoupling, comprising a data set determination module, a multimodal image registration neural network model building module, a module for determining an estimated conversion parameter matrix of a selected image pair, a model training module and an image registration module;
[0087] A data set determination module is used to obtain a number of multimodal images and adjust the image size, divide the images with the same size after adjustment into reference images and images to be registered and form a number of image pairs, so as to obtain a multimodal image registration data set with diversity;
[0088] A multimodal image registration neural network model building module is used to build a multimodal image registration neural network model based on feature decoupling. The model includes a general information extraction module, a feature decoupling module and a registration parameter estimation module. The feature decoupling module includes a shared feature mapping branch and a unique feature mapping branch.
[0089] The module for determining the estimated conversion parameter matrix of the selected image pair is used to input a pair of multimodal images in the data set into the general information extraction module for processing, and obtain the general feature map of the reference image and the general feature map of the image to be registered in the selected image pair; input the general feature map of the reference image and the image to be registered into the feature decoupling module for processing, and respectively extract the shared feature map and the unique feature map of the reference image and the image to be registered; input the shared feature map of the reference image and the shared feature map of the image to be registered into the registration parameter estimation module for processing, and output the estimated conversion parameter matrix of the selected image pair;
[0090] A model training module is used to calculate the total loss function value of each image pair based on the shared feature map and the unique feature map of the reference image and the image to be registered, the estimated conversion parameter matrix of the image pair and the preset loss function, supervise the training process of the multimodal image registration neural network model, select the network parameters with the smallest loss value to update the multimodal image registration neural network model, and obtain a trained multimodal image registration neural network model;
[0091] The image registration module is used to process the multimodal image to be registered using the trained multimodal image registration neural network model to obtain an image conversion parameter matrix, and transform the image to be registered in the image pair according to the image conversion parameter matrix to obtain a registered image.
[0092] For the specific definition of a multimodal image registration system based on feature decoupling, please refer to the definition of a multimodal image registration method based on feature decoupling above, which will not be repeated here. Each module in the above-mentioned multimodal image registration system based on feature decoupling can be implemented in whole or in part through software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0093] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a multimodal image registration method based on feature decoupling when executing the computer program.
[0094] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0095] The above is a detailed introduction to the multimodal image registration method, system and computer device based on feature decoupling provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A multimodal image registration method based on feature decoupling, characterized in that: The method comprises the following steps: S100: acquiring a plurality of multimodal images and resizing the images, dividing the images with the same size after resizing into reference images and images to be registered and forming a plurality of image pairs, thereby obtaining a multimodal image registration dataset with diversity; S200: Building a multimodal image registration neural network model based on feature decoupling, the model includes a general information extraction module, a feature decoupling module and a registration parameter estimation module, the feature decoupling module includes a shared feature mapping branch and a unique feature mapping branch; S300: Input a pair of multimodal images in the data set into the general information extraction module for processing, and obtain the general feature map of the reference image in the selected image pair and the general feature map of the image to be registered; input the general feature maps of the reference image and the image to be registered into the feature decoupling module for processing, and extract the shared feature map and the unique feature map of the reference image and the image to be registered respectively; input the shared feature map of the reference image and the shared feature map of the image to be registered into the registration parameter estimation module for processing, and output the estimated conversion parameter matrix of the selected image pair; the registration parameter estimation module includes a neighborhood cost convolution layer, a first convolution layer, a normalization layer, an activation function layer, a maximum pooling layer and a second convolution layer connected in sequence, wherein the step size of the first and second convolution layers is 1, and the convolution kernel size is 1. 1; The image conversion parameter matrix in S300 is specifically expressed as: (5) In the formula, represents the image transformation parameter matrix for the image pair, Representation 1 1 convolutional layer, represents the maximum pooling layer, represents the neighborhood cost volume construction operator, and Respectively represent the shared feature map of the reference image and the shared feature map of the image to be registered; S400: Calculate the total loss function value of each image pair based on the shared feature map and the unique feature map of the reference image and the image to be registered, the estimated conversion parameter matrix of the image pair and the preset loss function, supervise the training process of the multimodal image registration neural network model, select the network parameters with the minimum loss value to update the multimodal image registration neural network model, and obtain the trained multimodal image registration neural network model; S500: Acquire multimodal images to be registered in a real scene and form image pairs to be registered, use a trained multimodal image registration neural network model to process the multimodal image pairs to be registered, obtain an image conversion parameter matrix, and convert the images to be registered in the image pairs according to the image conversion parameter matrix to obtain registered images.
2. The method according to claim 1, characterized in that: In S100, a plurality of multimodal images are acquired and image size adjustment is performed. Specifically, the image size adjustment is as follows: (1) In the formula, is the pixel value of the original image, is the position after resizing, is the set of neighborhood pixels considered during the resizing process, Is with the field pixel The relevant weight function, The new image after resizing is at position The pixel value of .
3. The method according to claim 2, characterized in that The general information extraction module in S200 includes a convolution module, a normalization layer, an activation function layer, and three residual modules connected in sequence. The convolution module includes three convolution layers connected in sequence. The step size of each convolution layer is 2 and the convolution kernel size is 3.
3. The residual module consists of two convolutional layers connected in sequence. The step size of each convolutional layer is 1 and the convolution kernel size is 3.
3. Interlayer pass 1 1. Convolution adjusts the dimension to achieve residual connection; the general feature map of the reference image and the general feature map of the image to be registered in S200 are specifically: (2) In the formula, represents the general information extraction module, and Respectively represent the reference image and the image to be registered in the selected image pair, and They respectively represent the universal feature map of the reference image in the selected image pair and the universal feature map of the image to be registered.
4. The method according to claim 3, characterized in that The shared feature map branch includes a convolutional layer, a normalization layer, and an activation function layer connected in sequence, where the step size of the convolutional layer is 1 and the convolution kernel size is 1. 1; The unique feature mapping branch includes a convolutional layer, a normalization layer, an activation function layer, and two residual modules connected in sequence, where the stride of the convolutional layer is 1 and the convolution kernel size is 1.
1. The residual module consists of two convolutional layers connected in sequence. The stride of each convolutional layer is 1 and the convolution kernel size is 1. 1; The shared feature map and the unique feature map of the reference image and the image to be registered in S300 are specifically expressed as: (3) (4) In the formula, represents the shared feature map branch, represents a unique feature mapping branch, and They represent the universal feature map of the reference image in the selected image pair and the universal feature map of the image to be registered, and Represent the unique feature map and shared feature map of the reference image respectively, and They represent the unique feature map and shared feature map of the image to be registered respectively.
5. The method according to claim 4, characterized in that The loss function includes a similarity loss function, a difference loss function, a reconstruction loss function and a registration true value loss function. The loss function designed in S400 calculates the total loss function value, including sequentially calculating the similarity loss function value of the shared feature map of the reference image and the shared feature map of the image to be registered, the difference loss function value and the reconstruction loss value of the shared feature map of the reference image and its unique feature map, the difference loss function value and the reconstruction loss value of the shared feature map of the image to be registered and its unique feature map, and the registration true value loss function value of the estimated transformation parameter matrix and the true value transformation parameter matrix.
6. The method according to claim 5, characterized in that The loss function designed in S400 calculates the total loss function value as follows: (6) (7) (8) (9) (10) In the formula, Represents the total loss function calculation value, Represents the calculated value of the similarity loss function, Represents the calculated value of the difference loss function, Represents the calculated value of the reconstruction loss function, Represents the calculated value of the registration truth loss function, and Represent the unique feature map and shared feature map of the reference image respectively, and Respectively represent the unique feature map and shared feature map of the image to be registered, is the empirical expectation vector, yes The central moment vector of the sample is represents the Euclidean norm, and Respectively represent the starting point and end point of the interval used for normalization calculation, represents the trace of the matrix, represents the transpose of a matrix, and represent the reference image and the image to be registered respectively, represents the 1 norm, and are the estimated conversion parameter matrix and the true value conversion parameter matrix respectively.
7. The method according to claim 6, characterized in that In S500, the image to be registered in the image pair is transformed according to the image transformation parameter matrix, specifically: (11) In the formula, represents the image after the image to be registered is registered. represents the image to be registered, Represents the image conversion process, Represents the reference image in the real scene and the image to be registered The image transformation parameter matrix corresponding to the image pair to be registered.
8. A multimodal image registration system based on feature decoupling, characterized in that: It includes a data set determination module, a multimodal image registration neural network model building module, a module for determining the estimated conversion parameter matrix of the selected image pair, a model training module and an image registration module; A data set determination module is used to obtain a number of multimodal images and adjust the image size, divide the images with the same size after adjustment into reference images and images to be registered and form a number of image pairs, so as to obtain a multimodal image registration data set with diversity; A multimodal image registration neural network model building module is used to build a multimodal image registration neural network model based on feature decoupling. The model includes a general information extraction module, a feature decoupling module and a registration parameter estimation module. The feature decoupling module includes a shared feature mapping branch and a unique feature mapping branch. The module for determining the estimated transformation parameter matrix of the selected image pair is used to input a pair of multimodal images in the data set into the general information extraction module for processing, and obtain the general feature map of the reference image in the selected image pair and the general feature map of the image to be registered; the general feature map of the reference image and the image to be registered is input into the feature decoupling module for processing, and the shared feature map and the unique feature map of the reference image and the image to be registered are respectively extracted; The shared feature map of the reference image and the shared feature map of the image to be registered are input into the registration parameter estimation module for processing, and the estimated transformation parameter matrix of the selected image pair is output; the registration parameter estimation module includes a neighborhood cost volume layer, a first convolution layer, a normalization layer, an activation function layer, a maximum pooling layer and a second convolution layer connected in sequence, wherein the step size of the first and second convolution layers is 1, and the convolution kernel size is 1. 1; The image conversion parameter matrix is specifically expressed as: (5) In the formula, represents the image transformation parameter matrix for the image pair, Representation 1 1 convolutional layer, represents the maximum pooling layer, represents the neighborhood cost volume construction operator, and Respectively represent the shared feature map of the reference image and the shared feature map of the image to be registered; A model training module is used to calculate the total loss function value of each image pair based on the shared feature map and the unique feature map of the reference image and the image to be registered, the estimated conversion parameter matrix of the image pair and the preset loss function, supervise the training process of the multimodal image registration neural network model, select the network parameters with the smallest loss value to update the multimodal image registration neural network model, and obtain a trained multimodal image registration neural network model; The image registration module is used to process the multimodal image to be registered using the trained multimodal image registration neural network model to obtain an image conversion parameter matrix, and transform the image to be registered in the image pair according to the image conversion parameter matrix to obtain a registered image.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-modal image registration method based on deep learning
CN116523981A
Multi-source remote sensing image registration method based on decoupling feature mutual information
CN118485697A