A multimodal image registration method based on automodal correlation and cross-modal estimation
By constructing a deep learning joint network based on CNN-Transformer, the high-frequency and low-frequency components of the image are decoupled and autocorrelated information is extracted, and the accuracy and robustness of the registration of visible light and thermal infrared images in the prior art are solved, thereby achieving high-precision registration effect in harsh environments.
Patent Information
- Application Number
- CN202510274637.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The existing methods of visible and thermal infrared image registration are difficult to achieve the accuracy and robustness of cross-modal matching estimation when facing significant modal differences and viewing angle changes.
Using a deep learning joint network architecture based on convolution and transformer CNN-Transformer, the high-frequency and low-frequency components of the image are decoupled, and the local and global autocorrelation information in the mode is extracted to realize the matching estimation of cross-modal features by constructing the automodal similarity encoding module and the cross-modal matching estimation module.
It significantly improves the accuracy and robustness of visible light and thermal infrared image registration, and can be registered stably under harsh environmental conditions such as night, rainy days, and foggy days, and has strong applicability and performance.
Smart Images

Figure CN119762559B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an image registration method and to the technical field of cross-modal image registration, and in particular to a multi-modal image registration method based on auto-modal correlation and cross-modal estimation. Background Art
[0002] By comprehensively utilizing the cross-modal complementary information of visible light and thermal infrared images, multispectral machine vision technology has been widely used in fields such as autonomous driving, traffic monitoring, power production, and military counter-terrorism. Existing visible light and thermal infrared image acquisition systems mostly use binocular configurations, and there are inevitably differences in viewing angles when acquiring images; at the same time, due to the differences in physical properties and imaging principles between visible light and thermal infrared images, there are significant modal differences between the two. However, existing visual technology research based on visible light and thermal infrared images, such as target recognition, object tracking, and fault diagnosis, often requires two accurately registered images as input. Therefore, how to accurately establish the feature matching relationship between visible light and infrared images and achieve accurate registration has become the key to promoting the development of multi-spectral / cross-modal vision tasks.
[0003] At present, there are mainly two types of methods for matching estimation and registration of visible light and thermal infrared images: traditional methods based on artificial feature descriptors and deep learning methods based on neural network models. Traditional methods based on artificial feature descriptors generate corresponding feature descriptors by extracting significant feature points or geometric structures in the image, such as corners, contours, edges, etc., and calculate the similarity of the feature descriptors in the two images to determine the matching relationship and obtain the corresponding matching transformation model (affine transformation, perspective transformation, etc.). Although this type of method has achieved certain results in some traditional image registration tasks, it still has certain limitations when facing significant modal differences and perspective changes; especially in the registration of visible light and thermal infrared images, due to the difference in pixel amplitudes and the complexity of geometric transformations between the two, the matching accuracy and robustness of this type of method often cannot meet the needs of practical applications. Deep learning methods based on neural network models mainly convert visible light images into the style of thermal infrared images by constructing a style transfer network, thereby reducing modal differences and simplifying the difficulty of cross-modal matching; then, the feature mapping relationship between images is learned through the convolutional neural network model to generate an optical flow map reflecting pixel-by-pixel displacement estimation to achieve image registration. Although deep learning methods have shown strong adaptability and performance in some complex scenes, there are still significant differences between the generated pseudo-infrared images and real thermal infrared images due to the lack of prior information such as temperature radiation. In addition, the architecture of convolutional neural networks is limited by the size of the convolution kernel and the receptive field, which makes it difficult to achieve high-precision registration when processing visible light and thermal infrared image pairs with large parallax and complex geometric transformations. Therefore, existing visible light and thermal infrared image registration methods still find it difficult to achieve accuracy and robustness in cross-modal matching estimation when faced with significant modal differences and changes in perspective. Summary of the invention
[0004] In order to solve the problems existing in the background technology, the present invention provides a multimodal image registration method based on automodal correlation and cross-modal estimation. The method of the present invention constructs a deep learning joint network architecture based on convolution and transformer CNN (Convolutional Neural Networks)-Transformer, decouples the high-frequency and low-frequency components of the image, and extracts local and global autocorrelation information within the modality respectively to promote the matching estimation of cross-modal features, thereby significantly improving the accuracy and robustness of visible light and thermal infrared image registration.
[0005] The technical solution adopted by the present invention is:
[0006] The multimodal image registration method based on self-modal correlation and cross-modal estimation of the present invention comprises:
[0007] S1: Construct an initial framework model of a deep learning joint network consisting of a sequentially connected image frequency decomposition module, an automodal similarity encoding module, a cross-modal matching estimation module, an optical flow generation and refinement module, and an image transformation module.
[0008] S2: Acquire several registered visible light and thermal infrared road traffic images and perform spatial transformation, simulate and generate several visible light and thermal infrared road traffic images to be registered and construct them as a training set, input the training set into the deep learning joint network initial framework model for training, and then obtain the trained visible light and thermal infrared registration model.
[0009] S3: Obtain the visible light and thermal infrared road traffic images to be registered and input them into the visible light and thermal infrared registration model. After processing, output the visible light road traffic image registered with the thermal infrared road traffic image and display it on the display, completing the multimodal image registration.
[0010] In the step S1, the self-modal similarity encoding module of the initial framework model of the deep learning joint network includes a high-frequency encoder and a low-frequency encoder, and the cross-modal matching estimation module includes a high-frequency estimator and a low-frequency estimator. The visible light and thermal infrared road traffic images input by the initial framework model of the deep learning joint network are first processed by the image frequency decomposition module to output the high-frequency and low-frequency components of the visible light road traffic images and the high-frequency and low-frequency components of the thermal infrared road traffic images. The high-frequency components of the visible light and thermal infrared road traffic images are input into the high-frequency encoder for processing to output the high-frequency local significant feature maps of visible light and thermal infrared, and then input into the high-frequency estimator. The low-frequency components of the visible light and thermal infrared road traffic images are input into the low-frequency encoder for processing and then output as low-frequency global autocorrelation feature maps of visible light and thermal infrared. These are then input into the low-frequency estimator for processing and then output as two low-frequency global cross-modal matching estimation maps. The two high-frequency local cross-modal matching estimation maps and the two low-frequency global cross-modal matching estimation maps are jointly input into the optical flow generation and refinement module for processing and then output as optical flow estimation maps. The optical flow estimation maps are then input into the image transformation module for processing and finally output as visible light road traffic images that have been registered with the thermal infrared road traffic images.
[0011] The image frequency decomposition module includes two image frequency decomposers. The visible light and thermal infrared road traffic images are respectively input into the two image frequency decomposers for processing and decomposition, and then the high-frequency and low-frequency components of the visible light road traffic image and the high-frequency and low-frequency components of the thermal infrared road traffic image are output.
[0012] The high-frequency encoder of the self-modal similarity encoding module includes two parallel high-frequency encoding structures, each of which includes a number of high-frequency local feature extraction units connected in sequence, and the low-frequency encoder includes two parallel low-frequency encoding structures, each of which includes a number of low-frequency global autocorrelation extraction units connected in sequence. The high-frequency components of the visible light and thermal infrared road traffic images are respectively input into the two high-frequency encoding structures of the high-frequency encoder for processing, and the processed outputs are then jointly input into the high-frequency estimator for processing. The low-frequency components of the visible light and thermal infrared road traffic images are respectively input into the two low-frequency encoding structures of the low-frequency encoder for processing, and the processed outputs are then jointly input into the low-frequency estimator for processing.
[0013] The high-frequency local feature extraction unit includes a first conventional convolution layer, a first depth-separable convolution layer, a layer normalization process, a first fully connected layer, a first Gaussian error linear module activation function GELU (Gaussian Error Linear Units) and a second fully connected layer connected in sequence. The initial input of the high-frequency local feature extraction unit is processed in sequence by the first conventional convolution layer, the first depth-separable convolution layer, the first layer normalization process, the first fully connected layer, the first Gaussian error linear module activation function and the second fully connected layer. The output result is then added to the initial input and output as the final output of the high-frequency local feature extraction unit.
[0014] The low-frequency global autocorrelation extraction unit includes an image block position encoding process, a first pyramid pooling module, a first multi-head self-attention module, a layer normalization process and a first inverted bottleneck module which are connected in sequence. The initial input of the low-frequency global autocorrelation extraction unit is processed by the image block position encoding process and the first pyramid pooling module in sequence to obtain a first query vector Query, a first key vector Key and a first value vector Value respectively, and then they are input into the first multi-head self-attention module for processing and then output a global self-attention map. The global self-attention map and the first query vector Query are added to obtain a first addition result. The first addition result is processed by the layer normalization process and the first inverted bottleneck module in sequence. The processed output is then added to the first addition result and output as the final output of the low-frequency global autocorrelation extraction unit.
[0015] The high-frequency estimator of the cross-modal matching estimation module includes a plurality of high-frequency local cross-modal matching estimation units connected in sequence, and the low-frequency estimator includes a plurality of low-frequency global cross-modal matching estimation units connected in sequence.
[0016] The high-frequency local cross-modal matching estimation unit includes a channel dimension splicing process, a second conventional convolution layer, a second depth-separable convolution layer, a layer normalization process, a second fully connected layer, a second Gaussian error linear module activation function GELU, a second fully connected layer and a channel dimension separation process, which are connected in sequence. The visible light and thermal infrared high-frequency local significant feature maps are processed by channel dimension splicing to obtain a first splicing result, and the thermal infrared and visible light high-frequency local significant feature maps are processed by channel dimension splicing to obtain a second splicing result. The first splicing result and the second splicing result are processed by channel dimension splicing and then sequentially processed by the second conventional convolution layer, the second depth-separable convolution layer, the layer normalization process, the second fully connected layer, the second Gaussian error linear module activation function GELU, the second fully connected layer and the channel dimension separation process to output the first and second separation results. The first separation result and the first splicing result are added to obtain a first high-frequency local cross-modal matching estimation map, and the second separation result and the second splicing result are added to obtain a second high-frequency local cross-modal matching estimation map. The two high-frequency local cross-modal matching estimation maps are used as the final output of the high-frequency local cross-modal matching estimation unit.
[0017] The low-frequency global cross-modal matching estimation unit includes a sequentially connected image block position encoding process, a second pyramid pooling module, a third pyramid pooling module, a second multi-head self-attention module, a layer normalization process, a second inverted bottleneck module and a third inverted bottleneck module. The visible light low-frequency global autocorrelation feature map is sequentially processed by the image block position encoding process and the second pyramid pooling module to output a second query vector Query, a second key vector Key and a second value vector Value. The thermal infrared low-frequency global autocorrelation feature map is sequentially processed by the image block position encoding process and the third pyramid pooling module to output a third query vector Query, a third key vector Key and a third value vector Value. Each query vector Query, key vector Key and value vector Value are jointly input into the second multi-head self-attention module for processing. The second query vector Query, the third key vector Ke are After processing, y and the third value vector Value output a first cross-modal attention result, and after processing, the third query vector Query, the second key vector Key and the second value vector Value output a second cross-modal attention result. The first cross-modal attention result and the second query vector Query are added to obtain a second addition result, and the second addition result is sequentially processed by layer normalization, the output after the second inverted bottleneck module, and then added to the second addition result to obtain a first low-frequency global cross-modal matching estimation graph. The second cross-modal attention result and the third query vector Query are added to obtain a third addition result, and the third addition result is sequentially processed by layer normalization, the output after the third inverted bottleneck module, and then added to the third addition result to obtain a second low-frequency global cross-modal matching estimation graph. The two low-frequency global cross-modal matching estimation graphs are used as the final output of the low-frequency global cross-modal matching estimation unit.
[0018] The optical flow generation and refinement module includes a first optical flow estimation layer, a second optical flow estimation layer and a plurality of optical flow refinement units connected in sequence. Two high-frequency local cross-modal matching estimation maps are input into the first optical flow estimation layer for processing and then output a high-frequency optical flow estimation map. Two low-frequency global cross-modal matching estimation maps are input into the second optical flow estimation layer for processing and then output a low-frequency optical flow estimation map. The high-frequency optical flow estimation map and the low-frequency optical flow estimation map are weightedly summed and then processed by each optical flow refinement unit in sequence to output an optical flow estimation map. Finally, in the image transformation module, the visible light road traffic image is spatially transformed according to the optical flow estimation map to obtain a visible light road traffic image that is aligned with the thermal infrared road traffic image.
[0019] In the step S2, for each pair of registered visible light and thermal infrared road traffic images, one or more of the affine transformation model, perspective transformation model and thin plate spline transformation model generated by random parameters are used to perform spatial transformation processing on the thermal infrared road traffic image to generate one or more transformed thermal infrared road traffic images, and each transformed thermal infrared road traffic image and the original visible light road traffic image are combined into a pair of visible light and thermal infrared road traffic images to be registered and constructed as a training set; the spatial transformation is performed on the image block obtained by cropping the central area of the original image.
[0020] When training the initial framework model of the deep learning joint network, the method of generating batch data sequentially is used to input the training set, the gradient weight decay method with adaptive learning rate is used to update the model parameters, and the gradient clipping method is used to adjust the model parameters until the loss function of the initial framework model of the deep learning joint network converges to obtain the trained visible light and thermal infrared registration model.
[0021] The electronic device of the present invention comprises: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method as described above.
[0022] The computer-readable storage medium of the present invention stores program data thereon, and is characterized in that the program data implements the method described above when executed by a processor.
[0023] The beneficial effects of the present invention are:
[0024] 1. The present invention can fully construct the self-modal similarity of visible light and thermal infrared images, improve the accuracy of cross-modal feature matching estimation, and achieve high-precision registration of visible light and thermal infrared images with significant viewing angle differences and modality differences. In addition, the visible light and thermal infrared images can still be stably registered under harsh environmental conditions such as night, rainy days, and foggy days, and it has strong robustness.
[0025] 2. The present invention introduces a high-frequency local feature extraction unit and a low-frequency global autocorrelation extraction unit to extract high-frequency local significant features and low-frequency global autocorrelation features respectively according to frequency specificity, providing a key basis for cross-modal matching estimation; it adopts deep separable convolution and Gaussian error linear unit activation function GELU to effectively reduce the number of model parameters, improve computational efficiency, and enhance the ability to extract complex significant features.
[0026] 3. The present invention introduces a high-frequency local cross-modal matching estimation unit and a low-frequency global cross-modal matching estimation unit, performs high-frequency local cross-modal matching estimation and low-frequency global cross-modal matching estimation respectively according to frequency specificity, and combines automodal correlation with cross-modal matching estimation, thereby greatly improving the accuracy of optical flow estimation.
[0027] 4. The present invention generates a query vector Query of one modality and a key vector Key and a value vector Value of another modality respectively, and performs cross-modal feature attention calculation based on them, thereby expanding the self-attention mechanism of the traditional transformer network layer (TransformerLayer) to a cross-modal attention mechanism, effectively realizing the global regional cross-modal matching estimation of low-frequency image features, and significantly improving the accuracy and robustness of image registration. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a method flow chart of an embodiment of the present invention;
[0029] Figure 2 It is a schematic diagram of the overall structure of the deep learning joint network framework model adopted in the embodiment of the present invention;
[0030] Figure 3 It is a structural schematic diagram of a high-frequency local feature extraction unit in a deep learning joint network framework model of an embodiment of the present invention;
[0031] Figure 4 It is a structural schematic diagram of a low-frequency global autocorrelation extraction unit in a deep learning joint network framework model of an embodiment of the present invention;
[0032] Figure 5 It is a structural schematic diagram of a high-frequency local cross-modal matching estimation unit in a deep learning joint network framework model of an embodiment of the present invention;
[0033] Figure 6 It is a structural schematic diagram of a low-frequency global cross-modal matching estimation unit in a deep learning joint network framework model of an embodiment of the present invention;
[0034] Figure 7 is a diagram of target detection results in different data set scenarios according to an embodiment of the present invention, wherein: Figure 7(a) is a target detection result diagram of the present invention in a daytime scene in the visible light-thermal infrared road traffic image dataset KAIST. Figure 7 (b) is a target detection result diagram of the present invention in a night scene in the visible light-thermal infrared urban image dataset MFNet. Figure 7 (c) is the multi-scenario cross-modal dataset M of the present invention. 3 FD (multi-scenario multi-modality dataset) target detection results in rainy scenes, Figure 7 (d) is the multi-scenario cross-modal dataset M of the present invention. 3 Target detection results in foggy scenes in FD. DETAILED DESCRIPTION
[0035] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] like Figure 1 As shown, the multimodal image registration method based on self-modal correlation and cross-modal estimation of the present invention is specifically as follows:
[0037] First, a deep learning joint network initial framework model is constructed, which is composed of a sequentially connected image frequency decomposition module, an automodal similarity encoding module, a cross-modal matching estimation module, an optical flow generation and refinement module, and an image transformation module. The deep learning joint network initial framework model is constructed based on the overall framework of convolution and transformer CNN-Transformer, such as Figure 2 As shown in the figure, specifically, the framework mainly includes two image frequency decomposers, two cascaded high-frequency local feature extraction units, two cascaded low-frequency global autocorrelation extraction units, two cascaded high-frequency local cross-modal matching estimation units, two cascaded low-frequency global cross-modal matching estimation units, five cascaded optical flow refinement units and an image transformation unit. The two image frequency decomposers process the input visible light and thermal infrared images respectively to obtain four high-frequency and low-frequency image components (high-frequency visible light component, high-frequency thermal infrared component, low-frequency visible light component and low-frequency thermal infrared component), which are then sequentially input into the cascaded two high-frequency local feature extraction units, two low-frequency global autocorrelation extraction units, two high-frequency local cross-modal matching estimation units, two low-frequency global cross-modal matching estimation units, five optical flow refinement units and an image transformation unit to obtain the output results.
[0038] The self-modal similarity encoding module of the deep learning joint network initial framework model includes a high-frequency encoder and a low-frequency encoder, and the cross-modal matching estimation module includes a high-frequency estimator and a low-frequency estimator. The visible light and thermal infrared road traffic images input by the deep learning joint network initial framework model are first processed by the image frequency decomposition module to output the high-frequency and low-frequency components of the visible light road traffic images and the high-frequency and low-frequency components of the thermal infrared road traffic images. The high-frequency components of the visible light and thermal infrared road traffic images are input into the high-frequency encoder for processing and then output the high-frequency local significant feature maps of visible light and thermal infrared, which are then input into the high-frequency estimator for processing. After processing, two high-frequency local cross-modal matching estimation maps are output; the low-frequency components of the visible light and thermal infrared road traffic images are input into the low-frequency encoder for processing and then output as visible light and thermal infrared low-frequency global autocorrelation feature maps, which are then input into the low-frequency estimator for processing and then output as two low-frequency global cross-modal matching estimation maps; the two high-frequency local cross-modal matching estimation maps and the two low-frequency global cross-modal matching estimation maps are jointly input into the optical flow generation and refinement module for processing and then output as optical flow estimation maps; the optical flow estimation maps are then input into the image transformation module for processing and finally output as visible light road traffic images that are aligned with the thermal infrared road traffic images.
[0039] The image frequency decomposition module includes two image frequency decomposers. The visible light and thermal infrared road traffic images are respectively input into the two image frequency decomposers for processing and decomposition, and then the high-frequency and low-frequency components of the visible light road traffic image and the high-frequency and low-frequency components of the thermal infrared road traffic image are output. The image frequency decomposer used in the present invention specifically processes the input visible light image and thermal infrared image respectively through a guided filter (Guided Image Filtering) to extract the corresponding high-frequency component map. Subsequently, the low-frequency component map is calculated and obtained by subtracting the original image from the high-frequency component map. Finally, four image components are output, including a high-frequency visible light image component, a high-frequency thermal infrared image component, a low-frequency visible light image component, and a low-frequency thermal infrared image component.
[0040] The high-frequency encoder of the self-modal similarity coding module includes two parallel high-frequency coding structures, each of which includes two high-frequency local feature extraction units connected in sequence. The low-frequency encoder includes two parallel low-frequency coding structures, each of which includes two low-frequency global autocorrelation extraction units connected in sequence. The high-frequency components of the visible light and thermal infrared road traffic images are respectively input into the two high-frequency coding structures of the high-frequency encoder for processing, and the processed outputs are then jointly input into the high-frequency estimator for processing. The low-frequency components of the visible light and thermal infrared road traffic images are respectively input into the two low-frequency coding structures of the low-frequency encoder for processing, and the processed outputs are then jointly input into the low-frequency estimator for processing.
[0041] like Figure 3 As shown, it is a high-frequency local feature extraction unitM LFE Structure, high-frequency local feature extraction unit M LFE Used to extract local significant features of the high-frequency components of the image. The high-frequency local feature extraction unit includes a first conventional convolution layer, a first depth-separable convolution layer, a layer normalization process, a first fully connected layer, a first Gaussian error linear module activation function GELU, and a second fully connected layer connected in sequence. The initial input of the high-frequency local feature extraction unit is processed in sequence by the first conventional convolution layer, the first depth-separable convolution layer, the first layer normalization process, the first fully connected layer, the first Gaussian error linear module activation function, and the second fully connected layer. The output result is then added to the initial input and output as the final output of the high-frequency local feature extraction unit. In specific implementation, the high-frequency components of the visible light road traffic image I v hf and high frequency components of thermal infrared road traffic images I t hf It is sequentially input into two conventional convolutional layers with convolution kernel sizes of 4×4 and 8×8 respectively, and feature maps with the size of 1 / 4 and 1 / 8 of the original image are obtained; then, it is input into a depth-separable convolutional layer with a convolution kernel size of 7×7, and then after layer normalization, it is sequentially input into the second fully connected layer through the first fully connected layer and Gaussian error linear unit activation function GELU for feature mapping and activation. The obtained features are added to the high-frequency components of the input image to obtain the final unit output, including the visible light high-frequency local salient feature map F v and thermal infrared high-frequency local salient feature map F t .
[0042] High frequency local feature extraction unit M LFE The high-frequency visible light and high-frequency thermal infrared image components are processed separately, and high-frequency local salient features are output. The purpose is to extract features from high-texture salient areas in visible light and thermal infrared images, which not only provides an effective basis for subsequent cross-modal matching estimation, but also provides supplementary information for the low-frequency global autocorrelation extraction unit.
[0043] like Figure 4 As shown, it is a low-frequency global autocorrelation extraction unit M GSCE Structure, low-frequency global autocorrelation extraction unit M GSCEUsed to extract global autocorrelation of low-frequency components of an image. The low-frequency global autocorrelation extraction unit includes a sequentially connected image block position encoding process, a first pyramid pooling module PPM (Pyramid Pooling Module), a first multi-head self-attention module MSA (Multi-Head Self-Attention), a layer normalization process, and a first inverted bottleneck module IRB (Institutional Review Board). The initial input of the low-frequency global autocorrelation extraction unit is processed by the image block position encoding process and the first pyramid pooling module to obtain a first query vector Query, a first key vector Key, and a first value vector Value, respectively, and then they are input into the first multi-head self-attention module for processing and then output a global self-attention map. The global self-attention map and the first query vector Query are added to obtain a first addition result. The first addition result is sequentially processed by the layer normalization process and the first inverted bottleneck module. The processed output is then added to the first addition result and output as the final output of the low-frequency global autocorrelation extraction unit. Low-frequency global autocorrelation extraction unit M GSCE The specific workflow is as follows:
[0044] First, the low-frequency components of the visible light road traffic image are I v lf and low frequency components of thermal infrared road traffic images I t lf To ensure the continuity of features between adjacent image blocks, the present invention uses conventional convolution operations to extract overlapping feature image blocks and encodes them into image block sequences through a fully connected layer. E v(t) Then the image block sequence after position encoding is E v(t) Input the pyramid pooling module PPM, extract the multi-scale feature vector through the pyramid average pooling operation, then sequentially concatenate the extracted multi-scale feature vectors and input them into the fully connected layer to obtain the three output vectors of the pyramid pooling module PPM, which correspond to the query vector (Query) Q v(t) 、Key vector(Key) K v(t) Sum value vector (Value) V v(t), which realizes the feature extraction of large low-texture areas in the low-frequency feature map, and at the same time reduces the computational burden of the subsequent multi-head self-attention module; then the query vector Query, key vector Key and value vector Value output by the pyramid pooling module PPM are input into the multi-head self-attention module MSA to calculate the global self-attention map, as follows:
[0045] φ v(t) = S ( Q v(t) ( K v(t) ) T / ( C K1 ) 1 / 2 )× V v(t)
[0046] in, φ v(t) Represents the global self-attention map of visible light (or thermal infrared) obtained after calculation by the multi-head self-attention module MSA; S ( ) represents the normalized exponential activation function; T represents transpose; C K1 Represents the key vector (Key) K v(t) The number of channels.
[0047] Finally, the global self-attention map φ v(t) After layer normalization, the input is put into the inverted bottleneck module IRB, and feature transformation and nonlinear processing are performed on the global self-attention map to further enhance the feature representation ability, improve the computational efficiency and performance of the model, and finally obtain the low-frequency global autocorrelation feature map of visible light and thermal infrared. Φ v(t) .
[0048] The low-frequency global autocorrelation extraction unit processes the low-frequency visible light and high-frequency thermal infrared components respectively, and constructs a global autocorrelation feature map by combining the high-frequency local significant features output by the high-frequency local feature extraction unit. Its purpose is to provide a basis for subsequent cross-modal estimation by modeling frequency-specific self-modal correlation, and reduce the matching difficulties caused by the significant modal differences between visible light and thermal infrared. In addition, the high-frequency local feature extraction unit uses deep separable convolution and Gaussian error linear unit GELU activation function, which can enhance the ability to extract complex significant features while reducing the number of model parameters and improving computational efficiency; the low-frequency global autocorrelation extraction unit adopts a pyramid pooling strategy to generate multi-scale low-frequency image features, and models the correlation between global low-frequency features through the basic transformer network layer Transformer Layer, thereby significantly reducing the computational burden while improving the model's global feature modeling capabilities.
[0049] The high-frequency local feature extraction unit is used to extract local significant features of the high-frequency image components decomposed by the image frequency decomposer, and obtain high-frequency local significant feature maps of visible light and thermal infrared; the low-frequency global cross-modal matching estimation unit is used to extract global autocorrelation of the low-frequency image components decomposed by the image frequency decomposer, and obtain low-frequency global self-similarity feature maps of visible light and thermal infrared.
[0050] The high frequency estimator of the cross-modal matching estimation module includes two high frequency local cross-modal matching estimation units connected in sequence, such as Figure 5 As shown, it is a high-frequency local cross-modal matching estimation unit M LCCE Structure, high-frequency local cross-modal matching estimation unit M LCCEIt is used to extract local cross-modal matching estimation between visible light and thermal infrared high-frequency local salient feature maps. The high-frequency local cross-modal matching estimation unit includes a channel dimension splicing process, a second conventional convolution layer, a second depth-separable convolution layer, a layer normalization process, a second fully connected layer, a second Gaussian error linear module activation function GELU, a second fully connected layer and a channel dimension separation process, which are connected in sequence. The visible light and thermal infrared high-frequency local significant feature maps are processed by channel dimension splicing to obtain a first splicing result, and the thermal infrared and visible light high-frequency local significant feature maps are processed by channel dimension splicing to obtain a second splicing result. The first splicing result and the second splicing result are processed by channel dimension splicing and then sequentially processed by the second conventional convolution layer, the second depth-separable convolution layer, the layer normalization process, the second fully connected layer, the second Gaussian error linear module activation function GELU, the second fully connected layer and the channel dimension separation process to output the first and second separation results. The first separation result and the first splicing result are added to obtain a first high-frequency local cross-modal matching estimation map, and the second separation result and the second splicing result are added to obtain a second high-frequency local cross-modal matching estimation map. The two high-frequency local cross-modal matching estimation maps are used as the final output of the high-frequency local cross-modal matching estimation unit.
[0051] High-frequency local cross-modal matching estimation unit M LCCE Structure and high-frequency local feature extraction unit M LFE The composition structure of the network layer remains consistent, combining two high-frequency local cross-modal matching estimation units M LCCE The convolution kernel sizes of the conventional convolutional layers in the proposed method are set to 16×16 and 32×32, respectively, to obtain feature maps with the size of 1 / 16 and 1 / 32 of the original image, respectively. In order to extract the local cross-modal correlation between visible light and thermal infrared features, the present invention converts the input visible light high-frequency local salient feature map F v and thermal infrared high-frequency local salient feature map F t The channel dimension is first concatenated and then input into the conventional convolution layer, and the output features are separated in the channel dimension after the second fully connected layer to obtain the first high-frequency local cross-modal matching estimation map F v-t and the second high-frequency local cross-modal matching estimation map F t-v , i.e., high-frequency local cross-modal matching estimation unit M LCCE The final output.
[0052] The high-frequency local cross-modal matching estimation unit processes the local salient feature maps of high-frequency visible light and high-frequency thermal infrared respectively, and outputs the high-frequency local cross-modal matching estimation, which aims to perform matching estimation on the salient features in the visible light and thermal infrared images in a local range and capture the cross-modal correspondence of small displacement features. The high-frequency local feature extraction unit and the high-frequency local cross-modal matching estimation unit use deep separable convolution and Gaussian error linear unit activation function GELU, which effectively reduces the number of model parameters and improves computational efficiency while improving the ability to extract complex salient features.
[0053] The low-frequency estimator of the cross-modal matching estimation module includes two low-frequency global cross-modal matching estimation units connected in sequence. Figure 6 As shown, it is a low-frequency global cross-modal matching estimation unit M GCCE Structure, low-frequency global cross-modal matching estimation unit M GCCE Used to extract low-frequency global autocorrelation feature maps of visible light Φ v And thermal infrared low frequency global autocorrelation feature map Φ tThe global cross-modal matching estimation between them. The low-frequency global cross-modal matching estimation unit includes a sequentially connected image block position encoding process, a second pyramid pooling module, a third pyramid pooling module, a second multi-head self-attention module, a layer normalization process, a second inverted bottleneck module and a third inverted bottleneck module. The visible light low-frequency global autocorrelation feature map is sequentially processed by the image block position encoding process and the second pyramid pooling module to output a second query vector Query, a second key vector Key and a second value vector Value. The thermal infrared low-frequency global autocorrelation feature map is sequentially processed by the image block position encoding process and the third pyramid pooling module to output a third query vector Query, a third key vector Key and a third value vector Value. Each query vector Query, key vector Key and value vector Value are jointly input into the second multi-head self-attention module for processing. The second query vector Query, the third key vector Ke are output. After processing, y and the third value vector Value output a first cross-modal attention result, and after processing, the third query vector Query, the second key vector Key and the second value vector Value output a second cross-modal attention result. The first cross-modal attention result and the second query vector Query are added to obtain a second addition result, and the second addition result is sequentially processed by layer normalization, the output after the second inverted bottleneck module, and then added to the second addition result to obtain a first low-frequency global cross-modal matching estimation graph. The second cross-modal attention result and the third query vector Query are added to obtain a third addition result, and the third addition result is sequentially processed by layer normalization, the output after the third inverted bottleneck module, and then added to the third addition result to obtain a second low-frequency global cross-modal matching estimation graph. The two low-frequency global cross-modal matching estimation graphs are used as the final output of the low-frequency global cross-modal matching estimation unit.
[0054] Low-frequency global cross-modal matching estimation unit M GCCE Low frequency global autocorrelation extraction unit M GSCE The composition is consistent with that of the present invention, except that the low-frequency global autocorrelation extraction unit M GSC The global self-attention mechanism in is extended to cross-modality, generating a query vector from one of the modal features. Q v(t) , another modal feature generates a healthy vector K t(v) With value vector V t(v) , through the cross operation between the generated vectors from the two modalities, the global cross-modal matching estimation calculation is realized, as follows:
[0055] φ v-t(t-v) = S ( Qv(t) ( K t(v) ) T / ( C K2 ) 1 / 2 )× V t(v)
[0056] in, φ v-t(t-v) represents the low-frequency global cross-modal matching estimate between visible light and thermal infrared (or thermal infrared and visible light); C K2 Represents the key vector K t(v) The number of channels.
[0057] The low-frequency global cross-modal matching estimation unit processes the autocorrelation feature maps of low-frequency visible light and low-frequency thermal infrared respectively, capturing the cross-modal correspondence of large displacement features in the global range.
[0058] In the initial framework model of the deep learning joint network, the low-frequency global autocorrelation extraction unit and the low-frequency global cross-modal matching estimation unit use a pyramid pooling strategy to generate multi-scale low-frequency image features, and calculate the global spatial attention mechanism through the basic transformer network layer (Transformer Layer), which significantly reduces the model calculation burden while building the correlation between global low-frequency features. The low-frequency global cross-modal matching estimation unit calculates the cross-modal feature attention score by generating a query vector Query of one modality and a key vector Key and value vector Value of another modality, respectively, and expands the self-attention mechanism of the traditional transformer network layer to a cross-modal attention mechanism, realizing the global regional cross-modal matching estimation of low-frequency image features.
[0059] The high-frequency local cross-modal matching estimation unit is used to perform cross-modal matching estimation in the local area on the high-frequency local significant feature map extracted by the high-frequency local feature extraction unit; the low-frequency global cross-modal matching estimation unit is used to perform cross-modal matching estimation in the global area on the low-frequency global autocorrelation feature map extracted by the low-frequency global cross-modal matching estimation unit.
[0060] The optical flow generation and refinement module includes the first optical flow estimation layer, the second optical flow estimation layer, and five optical flow refinement units connected in sequence. Two high-frequency local cross-modal matching estimation maps are input into the first optical flow estimation layer for processing and then output as high-frequency optical flow estimation maps. Two low-frequency global cross-modal matching estimation maps are input into the second optical flow estimation layer for processing and then output as low-frequency optical flow estimation maps, which are used as the subsequent optical flow refinement units. M FRBThe high-frequency optical flow estimation map and the low-frequency optical flow estimation map are weightedly summed and processed by each optical flow refinement unit in turn to output the optical flow estimation map; finally, in the image transformation module, the visible light road traffic image is spatially transformed according to the optical flow estimation map to obtain the visible light road traffic image that is aligned with the thermal infrared road traffic image.
[0061] Optical flow refinement unit M FRB It is used to perform multi-level size recovery and refinement on the generated initial optical flow estimation map to obtain the output of the optical flow estimation map. First, the feature map is upsampled to twice the input resolution through bilinear interpolation, and then three conventional convolution layers with a convolution kernel size of 3×3 are used to further extract high-dimensional optical flow features, and finally the refined optical flow estimation map is output. The optical flow refinement unit first generates an initial optical flow estimation map based on the feature matching relationship and the corresponding image coordinates to characterize the pixel-by-pixel displacement information in the width and height directions of the image; then, through a multi-level cyclic optimization process, the size of the optical flow estimation map is gradually restored and reconstructed, and finally the optical flow estimation results of the input visible light image and thermal infrared image are obtained. Since the final optical flow estimation map is consistent with the resolution of the original image, it can densely characterize the displacement information of each pixel, thereby significantly improving the registration accuracy and enhancing the applicability and robustness of cross-modal matching estimation under non-rigid transformation conditions.
[0062] The final image transformation unit M IWB Used to refine the input visible light image according to the optical flow unit M FRB The obtained optical flow estimation map is spatially transformed to obtain a visible light image registered with the input thermal infrared image as the final output of the joint network architecture model, as follows:
[0063] I v ' ( x' , y' )= I v ( x + u ( x , y ), y + v ( x , y ))
[0064] in, I v ' and( x' , y') represent the visible light image and its pixel coordinates after registration with the input thermal infrared image, I v and( x , y ) represent the input visible light image and its pixel coordinates respectively; u ( x , y )and v ( x , y ) represent the optical flow estimation graphs I f middle( x , y ) stores the estimated displacements in the horizontal and vertical directions.
[0065] The image transformation unit uses a grid sampling method to map the pixel information of the input visible light image one by one based on the optical flow estimation map, converting the information in the visible light image to the perspective of the thermal infrared image to achieve precise pixel-by-pixel matching.
[0066] After the model is built, several registered visible light and thermal infrared road traffic images are obtained and spatially transformed. For each pair of registered visible light and thermal infrared road traffic images, the thermal infrared road traffic image is spatially transformed using one or more of the affine transformation model, perspective transformation model, and thin plate spline transformation model generated by random parameters to generate one or more transformed thermal infrared road traffic images. Each transformed thermal infrared road traffic image and the original visible light road traffic image are combined into a pair of visible light and thermal infrared road traffic images to be registered and constructed as a training set; the spatial transformation is performed on the image block obtained by cropping the central area of the original image. In the specific implementation, the visible light-thermal infrared road traffic image dataset KAIST is used. The visible light-thermal infrared road traffic image dataset KAIST consists of 1569 pairs of well-registered visible light and thermal infrared image pairs, which are collected under different weather, time, and lighting conditions to ensure the diversity and representativeness of the images, and also include various scenes such as urban roads, traffic signs, pedestrians, and vehicles. Then, the infrared images in the data set are transformed by affine transformation, perspective transformation and thin plate spline transformation through random parameter generation to generate simulated image pairs to be registered, and the true value labels of the optical flow maps are generated accordingly as supervision information for model training. The specific implementation is as follows:
[0067] 1) Randomly select one of the three transformation types and randomly generate transformation parameters based on the selected transformation model.
[0068] 2) Original thermal infrared image I tApply the selected transformation model and randomly generated transformation parameters to obtain the transformed thermal infrared image I t ' .
[0069] 3) Generate the true value label of the optical flow map based on the generated transformation model and transformation parameters I f gt Optical flow map ground truth label I f gt The transformation relationship described by the transformation model and transformation parameters is derived, which contains the specific displacement information of each pixel in the image registration process.
[0070] Based on the division of training and testing data samples provided by KAIST of the visible light and thermal infrared road traffic image dataset, the generated simulated image pairs to be registered are input into the initial framework model of the deep learning network for training, and the corresponding optical flow map true value labels are used as supervision information for supervision, thereby obtaining the visible light and thermal infrared image registration model.
[0071] The initial framework model of the deep learning network includes but is not limited to the following specific operations when training:
[0072] 1) During training, the input of visible light images and infrared images uses a method of sequentially generating batch data, and the batch size is 16.
[0073] 2) The initial framework model of the deep learning network uses the gradient weight decay method with adaptive learning rate to update the model parameters during training, and uses the gradient clipping method to adjust the model parameters, and requires several rounds of training.
[0074] The overall loss function used when training the model L The details are as follows:
[0075] L =∑ k α N-k ‖ I f k - I f gt,k ‖ 1 , k =0, 1, 2, ..., N
[0076] in, α Represents the weight parameter, which is specifically set to 0.9; N Indicates the total number of levels in the optical flow refinement unit, specifically set to 5; If k Represents the first k Optical flow estimation map generated by the layer; I f gt,k Indicates the corresponding k Optical flow ground truth labels at the resolution of the layer optical flow estimation graph; ‖ ‖ 1 Represents the 1-norm.
[0077] The overall loss function of the initial framework model of the deep learning joint network L Convergence, and obtain the trained visible light and thermal infrared registration model.
[0078] After the model training is completed, the visible light and thermal infrared road traffic images to be registered are obtained and input into the visible light and thermal infrared registration model. The visible light and thermal infrared image registration model performs frequency decomposition, self-modal similarity encoding, cross-modal matching estimation, optical flow refinement and spatial transformation on the images to be registered. After the processing is completed, the visible light road traffic image registered with the thermal infrared road traffic image is output and displayed on the display. The image content of the registered visible light corresponds to the input thermal infrared image pixel by pixel, completing the multimodal image registration. Among them, the input image to be registered needs to include both the visible light image and the thermal infrared image.
[0079] like Figure 7 (a) Figure 7 (b) Figure 7 (c) and Figure 7 As shown in (d), the architecture of the present invention is respectively demonstrated on the visible light-thermal infrared road traffic image dataset KAIST, the visible light-thermal infrared city image dataset MFNet and the multi-scene cross-modal dataset M 3 The matching results on the test set in FD include different lighting conditions (such as day and night), different weather conditions (such as rainy and foggy days) and different road scenes. According to the visualization result of the fusion effect of the two images after registration, it is not difficult to find that the algorithm proposed by the architecture of the present invention has high registration accuracy and has excellent registration performance for both high-texture target areas and low-texture background areas. For example, Figure 7 (b) Figure 7 (c) and Figure 7 In (d), the architecture of the present invention can achieve accurate registration of foreground pedestrians, vehicles and background buildings. It can be seen that it has strong registration capabilities for cross-modal image contents of different granularities and displacement degrees, and is highly robust to interference from different environmental factors, such as at night, in rainy days, in foggy days, etc.
[0080] The present invention realizes dense optical flow estimation and registration of visible light and thermal infrared images with significant viewpoint and modality differences by constructing a multimodal image registration method based on self-modal correlation and cross-modal estimation, and can stably register visible light and thermal infrared images under harsh environmental conditions such as rainy days, foggy days, and nights, with strong robustness. The architecture of the present invention can also be used for other multi-sensor visual analysis tasks, and realize high-quality data collection for autonomous driving, traffic monitoring, power production, military counter-terrorism, etc. based on visible light and thermal infrared images, and promote the development of downstream tasks such as target detection, segmentation, and tracking.
[0081] It should be understood by those skilled in the art that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present application may be implemented in various computer languages. The present application is described according to the flowcharts of the methods, systems and computer program products of the embodiments of the present application.
[0082] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the present invention is intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0083] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of equivalent technologies of the present invention, the present application is also intended to include these modifications and variations.
Claims
1. A multimodal image registration method based on self-modal correlation and cross-modal estimation, characterized in that: include: S1: Construct the initial framework model of the deep learning joint network consisting of a sequentially connected image frequency decomposition module, an automodal similarity encoding module, a cross-modal matching estimation module, an optical flow generation and refinement module, and an image transformation module; S2: Acquire several registered visible light and thermal infrared road traffic images and perform spatial transformation, simulate and generate several visible light and thermal infrared road traffic images to be registered and construct them as training sets, input the training sets into the initial framework model of the deep learning joint network for training, and then obtain the trained visible light and thermal infrared registration model; S3: Obtain the visible light and thermal infrared road traffic images to be registered and input them into the visible light and thermal infrared registration model. After processing, output the visible light road traffic image registered with the thermal infrared road traffic image and display it on a display, thereby completing the multimodal image registration. In the step S1, the self-modal similarity encoding module of the initial framework model of the deep learning joint network includes a high-frequency encoder and a low-frequency encoder, and the cross-modal matching estimation module includes a high-frequency estimator and a low-frequency estimator. The visible light and thermal infrared road traffic images input by the initial framework model of the deep learning joint network are first processed by the image frequency decomposition module to output the high-frequency and low-frequency components of the visible light road traffic images and the high-frequency and low-frequency components of the thermal infrared road traffic images. The high-frequency components of the visible light and thermal infrared road traffic images are input into the high-frequency encoder for processing to output the high-frequency local significant feature maps of visible light and thermal infrared, and then input into the high-frequency estimator. After being processed in the optical flow generation and refinement module, two high-frequency local cross-modal matching estimation maps are output; the low-frequency components of the visible light and thermal infrared road traffic images are input into the low-frequency encoder for processing and then output the visible light and thermal infrared low-frequency global autocorrelation feature maps, which are then input into the low-frequency estimator for processing and then output two low-frequency global cross-modal matching estimation maps; the two high-frequency local cross-modal matching estimation maps and the two low-frequency global cross-modal matching estimation maps are jointly input into the optical flow generation and refinement module for processing and then output the optical flow estimation map; the optical flow estimation map is then input into the image transformation module for processing, and finally the visible light road traffic image registered with the thermal infrared road traffic image is output; In the step S2, for each pair of registered visible light and thermal infrared road traffic images, one or more of the affine transformation model, perspective transformation model and thin plate spline transformation model generated by random parameters are used to perform spatial transformation processing on the thermal infrared road traffic image to generate one or more transformed thermal infrared road traffic images, and each transformed thermal infrared road traffic image and the original visible light road traffic image are combined into a pair of visible light and thermal infrared road traffic images to be registered and constructed as a training set; When training the initial framework model of the deep learning joint network, the method of generating batch data sequentially is used to input the training set, the gradient weight decay method with adaptive learning rate is used to update the model parameters, and the gradient clipping method is used to adjust the model parameters until the loss function of the initial framework model of the deep learning joint network converges to obtain the trained visible light and thermal infrared registration model.
2. The multimodal image registration method based on self-modal correlation and cross-modal estimation according to claim 1, characterized in that: The image frequency decomposition module includes two image frequency decomposers. The visible light and thermal infrared road traffic images are respectively input into the two image frequency decomposers for processing and decomposition, and then the high-frequency and low-frequency components of the visible light road traffic image and the high-frequency and low-frequency components of the thermal infrared road traffic image are output.
3. The multimodal image registration method based on automodal correlation and cross-modal estimation according to claim 1, characterized in that: The high-frequency encoder of the self-modal similarity encoding module includes two parallel high-frequency encoding structures, each of which includes a number of high-frequency local feature extraction units connected in sequence, and the low-frequency encoder includes two parallel low-frequency encoding structures, each of which includes a number of low-frequency global autocorrelation extraction units connected in sequence. The high-frequency components of the visible light and thermal infrared road traffic images are respectively input into the two high-frequency encoding structures of the high-frequency encoder for processing, and the processed outputs are then jointly input into the high-frequency estimator for processing. The low-frequency components of the visible light and thermal infrared road traffic images are respectively input into the two low-frequency encoding structures of the low-frequency encoder for processing, and the processed outputs are then jointly input into the low-frequency estimator for processing.
4. The multimodal image registration method based on self-modal correlation and cross-modal estimation according to claim 3, characterized in that: The high-frequency local feature extraction unit includes a first convolution layer, a first depth-separable convolution layer, a layer normalization process, a first fully connected layer, a first Gaussian error linear module activation function GELU, and a second fully connected layer connected in sequence. The initial input of the high-frequency local feature extraction unit is processed in sequence by the first convolution layer, the first depth-separable convolution layer, the first layer normalization process, the first fully connected layer, the first Gaussian error linear module activation function, and the second fully connected layer. The output result is then added to the initial input and output as the final output of the high-frequency local feature extraction unit; The low-frequency global autocorrelation extraction unit includes a position encoding process, a first pyramid pooling module, a first multi-head self-attention module, a layer normalization process and a first inverted bottleneck module which are connected in sequence. The initial input of the low-frequency global autocorrelation extraction unit is processed by the position encoding process and the first pyramid pooling module in sequence to obtain a first query vector Query, a first key vector Key and a first value vector Value respectively, and then they are input into the first multi-head self-attention module for processing and then output a global self-attention map. The global self-attention map and the first query vector Query are added to obtain a first addition result. The first addition result is processed by the layer normalization process and the first inverted bottleneck module in sequence. The processed output is then added to the first addition result and then output as the final output of the low-frequency global autocorrelation extraction unit.
5. The multimodal image registration method based on automodal correlation and cross-modal estimation according to claim 1, characterized in that: The high-frequency estimator of the cross-modal matching estimation module includes a plurality of high-frequency local cross-modal matching estimation units connected in sequence, and the low-frequency estimator includes a plurality of low-frequency global cross-modal matching estimation units connected in sequence; The high-frequency local cross-modal matching estimation unit includes a channel dimension splicing process, a second convolution layer, a second depth-separable convolution layer, a layer normalization process, a second fully connected layer, a second Gaussian error linear module activation function GELU, a second fully connected layer and a channel dimension separation process, which are connected in sequence. The visible light and thermal infrared high-frequency local significant feature maps are processed by channel dimension splicing to obtain a first splicing result, and the thermal infrared and visible light high-frequency local significant feature maps are processed by channel dimension splicing to obtain a second splicing result. The first splicing result and the second splicing result are processed by channel dimension splicing and then sequentially processed by the second convolution layer, the second depth-separable convolution layer, the layer normalization process, the second fully connected layer, the second Gaussian error linear module activation function GELU, the second fully connected layer and the channel dimension separation process to output the first and second separation results. The first separation result and the first splicing result are added to obtain a first high-frequency local cross-modal matching estimation map, and the second separation result and the second splicing result are added to obtain a second high-frequency local cross-modal matching estimation map. The two high-frequency local cross-modal matching estimation maps are used as the final output of the high-frequency local cross-modal matching estimation unit; The low-frequency global cross-modal matching estimation unit includes a position encoding process, a second pyramid pooling module, a third pyramid pooling module, a second multi-head self-attention module, a layer normalization process, a second inverted bottleneck module and a third inverted bottleneck module connected in sequence. The visible light low-frequency global autocorrelation feature map is processed by the position encoding process and the second pyramid pooling module in sequence to output a second query vector Query, a second key vector Key and a second value vector Value. The thermal infrared low-frequency global autocorrelation feature map is processed by the position encoding process and the third pyramid pooling module in sequence to output a third query vector Query, a third key vector Key and a third value vector Value. Each query vector Query, key vector Key and value vector Value are input into the second multi-head self-attention module for processing. The second query vector Query, the third key vector Key and the third value vector Value are output. After processing, the value vector Value outputs a first cross-modal attention result, and after processing, the third query vector Query, the second key vector Key and the second value vector Value output a second cross-modal attention result. The first cross-modal attention result and the second query vector Query are added to obtain a second addition result, and the second addition result is sequentially processed by layer normalization, the output after the second inverted bottleneck module, and then added to the second addition result to obtain a first low-frequency global cross-modal matching estimation graph. The second cross-modal attention result and the third query vector Query are added to obtain a third addition result, and the third addition result is sequentially processed by layer normalization, the output after the third inverted bottleneck module, and then added to the third addition result to obtain a second low-frequency global cross-modal matching estimation graph. The two low-frequency global cross-modal matching estimation graphs are used as the final output of the low-frequency global cross-modal matching estimation unit.
6. The multimodal image registration method based on automodal correlation and cross-modal estimation according to claim 1, characterized in that: The optical flow generation and refinement module includes a first optical flow estimation layer, a second optical flow estimation layer and a plurality of optical flow refinement units connected in sequence. Two high-frequency local cross-modal matching estimation maps are input into the first optical flow estimation layer for processing and then output a high-frequency optical flow estimation map. Two low-frequency global cross-modal matching estimation maps are input into the second optical flow estimation layer for processing and then output a low-frequency optical flow estimation map. The high-frequency optical flow estimation map and the low-frequency optical flow estimation map are weightedly summed and then processed by each optical flow refinement unit in sequence to output an optical flow estimation map. Finally, in the image transformation module, the visible light road traffic image is spatially transformed according to the optical flow estimation map to obtain a visible light road traffic image that is aligned with the thermal infrared road traffic image.
7. An electronic device, characterized in that: include: A memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having program data stored thereon, characterized in that: When the program data is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Transformer substation safe operation monitoring method and system based on OpenCV
CN118485973A
Frequency-adaptive cross-modal pedestrian retrieval method and device
CN118570845A