A semantic alignment method for fusion of infrared image and microwave non-image information
By constructing a unified semantic representation space and an efficient feature fusion mechanism, the problem of difference in signal characteristics and semantic expression in infrared and microwave data fusion is solved, and more accurate target recognition and environmental perception are achieved.
Patent Information
- Application Number
- CN202411799726.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-12-09
AI Technical Summary
The existing infrared and microwave data fusion methods have failed to effectively solve the differences in signal characteristics, resolution and semantic expression of the two data, resulting in information redundancy, noise amplification or semantic misalignment, affecting the fusion effect.
A semantic alignment method based on the fusion of infrared images and microwave non-image information is adopted, and the spatial and semantic features of infrared images are extracted through convolutional neural networks. The Fourier transform extracts microwave frequency domain features, and the twin network and contrast learning method are used to embed features into the shared semantic space. The dynamic weighted cosine similarity is calculated. The Hungarian algorithm performs global optimal matching, and finally feature fusion is performed through wavelet transform fusion and cross-modal attention mechanism.
It effectively solves the differences between infrared images and microwave data in feature expression forms and semantic levels, fully explores the complementary information of the two modes, and improves the target recognition and environmental perception capabilities in complex environments.
Smart Images

Figure CN119274182B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal data processing and fusion, and in particular to a semantic alignment method for fusing infrared images with microwave non-image information. Background Art
[0002] Infrared technology relies on thermal radiation imaging of objects and is suitable for nighttime or complex lighting conditions, but is greatly affected by weather, smoke, rain, snow and other environmental factors, is difficult to penetrate obstacles, and has a single information expression, which can only reflect the thermal characteristics of the surface; microwave technology detects targets through electromagnetic waves, has strong penetration ability, and can work stably in bad weather, but its resolution is low, and the output form is mostly non-image data (such as waveform, spectral information), which is significantly different from infrared images in characteristics and semantic expression. Current infrared and microwave fusion methods mainly focus on the data layer or decision layer, and fail to fully solve the differences in signal characteristics, resolution and semantic expression between the two types of data. The fusion effect is often affected by information redundancy, noise amplification or semantic misalignment.
[0003] With the increasing application demands in complex environment target recognition, multimodal remote sensing, national defense and security, autonomous driving and other fields, a single sensing technology is difficult to meet the needs of diverse scenarios. The fusion of infrared and microwave non-image information can combine the advantages of the two technologies to achieve more accurate target recognition and environmental monitoring. However, existing methods are difficult to fully utilize the complementary information of infrared images and microwave data, and lack an effective semantic alignment mechanism, resulting in poor fusion effects in complex scenes such as multiple obstacles and low visibility. To this end, by constructing a unified semantic space, a semantic alignment method based on the fusion of infrared images and microwave non-image information is developed, which can effectively improve the efficiency of information utilization, overcome the limitations of a single technology, and provide strong support for the development and practical application of multimodal perception technology. Summary of the invention
[0004] In view of the above-mentioned technical deficiencies, the purpose of the present invention is to provide a semantic alignment method for fusing infrared images and microwave non-image information.
[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0006] The present invention provides a semantic alignment method for fusing infrared images and microwave non-image information, comprising the following steps:
[0007] Step 1: Collect infrared images and microwave non-image information fragments of the same scene;
[0008] Step 2: Preprocess the infrared image collected in Step 1 by using non-local mean denoising and adaptive histogram equalization, and preprocess the microwave non-image information by using adaptive Wiener filtering and dynamic range compression plus gamma correction based on logarithmic transformation;
[0009] Step 3: Use convolutional neural network method to extract spatial and semantic features of infrared images, and use Fourier transform method to extract frequency domain features of microwave non-image information;
[0010] Step 4: Build a shared semantic space and use the Siamese Network deep learning model and contrastive learning method to embed infrared image features and microwave image features into the same high-dimensional semantic space.
[0011] Step 5, using dynamically weighted cosine similarity to calculate the similarity between the infrared image feature target and the microwave image feature in semantic representation;
[0012] Step 6: Based on the similarity calculation results in Step 5, use the Hungarian algorithm to perform global optimal matching;
[0013] Step 7. Use the wavelet transform fusion algorithm to perform wavelet decomposition on the infrared image and microwave features respectively, and extract the low-frequency and high-frequency coefficients at different resolutions. Use the weighted average method to fuse the low-frequency coefficients, use the cross-modal attention mechanism to dynamically allocate weights, and combine the deep learning edge detection model to optimize the fusion of high-frequency features;
[0014] Step 8: Post-process the fused image by using a three-dimensional block matching algorithm to denoise the fused image and using a multi-scale pyramid Laplace sharpening algorithm to enhance the image edges and details;
[0015] Step 9. Save the fusion result as a high-precision single-frame image (TIFF / PNG).
[0016] Furthermore, the Step 2 specifically includes:
[0017] Step 2.1-1, perform non-local mean denoising on the collected infrared image. Non-local mean denoising is an image denoising algorithm. For each pixel value, consider similar pixel blocks in the entire image range, and take a weighted average based on the similarity to obtain the denoised pixel value. Represents the intensity value of pixel x in the image after denoising; represents the raw intensity value of pixel y in the image; Represents the similarity weight of pixels x and y, indicating the degree of similarity between x and y; the non-local mean denoising formula is as follows:
[0018] ;
[0019] in, Represents the position coordinate of pixel x; Represents the pixel y position coordinate; Represents the domain of the image, that is, the coordinate set of all pixels in the image;
[0020] Similarity Weight It is calculated based on the similarity between the pixel blocks of x and y. The specific formula is as follows:
[0021] ;
[0022] in, The pixel position The image patch centered at the center; The pixel position The image patch centered at the center; is a filtering parameter used to control the decay speed of the weight;
[0023] Euclidean distance between two image patches The specific formula is as follows:
[0024] ;
[0025] in, and Represents the intensity values of corresponding pixels in two image blocks;
[0026] Step 2.1-2, use adaptive histogram equalization to perform local contrast enhancement on infrared images. By dividing the image into multiple local areas and performing histogram equalization on each area, the local contrast is enhanced, and excessive noise amplification is avoided by limiting the contrast enhancement amplitude (histogram clipping); Representing images The original pixel value at the position, Represents the pixel value after mapping, the formula for histogram equalization is:
[0027] ;
[0028] Where L is the number of gray levels of the image, is the frequency of gray value k, N is the total number of pixels in the image;
[0029] The contrast clipping formula is:
[0030] ;
[0031] in, Represents the clipped histogram; ClipLimit represents the clipping threshold, which limits the maximum frequency of each gray level; Excess is the sum of the clipped frequencies, which are redistributed to all gray levels;
[0032] Step 2.2-1, dynamically adjust the filter strength of microwave non-image information according to the local statistical characteristics of each pixel, use noise power to distinguish noise and signal, use adaptive Wiener filtering to remove random noise while retaining the key characteristics of the target to avoid blurring of important signals. is the pixel value after filtering, is the original pixel value, The mean of the local window, is the variance of the local window, is the noise power, w is the local window defined around the pixel, and the formula for adaptive Wiener filtering is:
[0033] ;
[0034] The formulas for the local mean and local variance are:
[0035] ;
[0036] ;
[0037] in, is the number of pixels in the window;
[0038] The formula for noise power is:
[0039] ;
[0040] Where n represents the total number of pixels in the flat area, r represents the selected flat area, Represents the mean value of pixels in the flat area, which refers to the area in the image with small brightness changes and almost no significant features;
[0041] Step 2.2-2, use the dynamic range compression and gamma correction method based on logarithmic transformation to adjust the contrast of microwave signals, adjust the grayscale distribution of data through nonlinear transformation, highlight the weak signal details of microwave data, and make it easier to match with infrared images in the semantic alignment and fusion process. Represents the compressed pixel value, the logarithmic transformation formula is:
[0042] ;
[0043] The formula for gamma correction is:
[0044] ;
[0045] in, is the gamma value;
[0046] Furthermore, the Step 3 specifically includes:
[0047] Step 3.1, use the convolutional neural network method to extract spatial and semantic features of infrared images. The convolutional neural network extracts features layer by layer, capturing the spatial and semantic information in the infrared image from local to global. represents the convolution result, represents the pixel value of the input image, represents the weight of the convolution kernel, Represents the bias of the convolution kernel, and the convolution formula is:
[0048] ;
[0049] in, Indicates the current position of the convolution kernel, Represents the local offset of the convolution kernel, m and n represent the size of the convolution kernel;
[0050] After operations such as convolutional layers, pooling layers, and activation functions, the features of the infrared image are gradually extracted from local to global, and finally a high-dimensional feature vector is formed. ;
[0051] Step 3.2, use the Fourier transform method to extract the frequency domain features of microwave non-image information. Convert the time series information of microwave data into frequency domain signals through Fourier transform, extract the component amplitudes of microwave signals at different frequencies, highlight the key features, and set Indicates the frequency The signal strength at Indicates yes The signal strength under the condition. The formula of Fourier transform is:
[0052] ;
[0053] Since the features after Fourier transformation are usually high-dimensional and may contain a lot of redundant information, they need to be reduced in dimension and normalized. Finally, the microwave features after Fourier transformation and processing are output as feature vectors. .
[0054] Furthermore, the Step 4 specifically includes:
[0055] The high-dimensional infrared image feature vector and high-dimensional microwave frequency domain feature vector extracted from Step 3 are projected into a shared semantic space using the Siamese Network deep learning model to facilitate modality alignment. The semantic representation alignment within and between modalities is performed by contrasting the loss function. is the embedded representation of the infrared image, indicating the location of the infrared image features in the shared semantic space. is the embedded representation of the microwave image, indicating the location of the microwave non-image information in the shared semantic space. The formula of the twin network is:
[0056] ;
[0057] ;
[0058] set up It represents contrast loss, and the formula of contrast loss is:
[0059] ;
[0060] Among them, Y is a binary label, which is used to guide the optimization of contrast loss, so that the feature distance of positive samples is closer and the feature distance of negative samples is farther; d represents the Euclidean distance, which is used to calculate the straight-line distance between two vectors; m represents the boundary threshold, which is used to control the penalty intensity of the loss function on the sample, ensuring that the distribution of positive and negative samples is reasonable;
[0061] The calculation formula of Euclidean distance is:
[0062] ;
[0063] Furthermore, the Step 5 specifically includes:
[0064] By calculating similarity, we ensure that semantically identical objects are closer in the shared space, using dynamic weighting factors , To flexibly deal with the different significance of infrared and microwave features in different scenes, Represents the dynamically weighted cosine similarity. The formula for dynamically weighted cosine similarity is:
[0065] ;
[0066] in, , They represent the dynamic weight factors of infrared image features and microwave features when calculating similarity, respectively, and are used to adjust the importance of infrared image features and microwave features in similarity calculation.
[0067] Furthermore, the Step 6 specifically includes:
[0068] According to the similarity calculation results in Step 5, the Hungarian algorithm is used to optimize the matching of infrared targets and microwave features globally. By considering the similarity between all targets, the total matching score is maximized. Represents the similarity between infrared features and microwave features. The similarity matrix formula is:
[0069] ;
[0070] Among them, the matrix Each element in Represents the matching score between infrared target i and microwave feature j, and its dimension is , where m is the number of targets in the infrared image and n is the number of microwave features;
[0071] The Hungarian algorithm solves the problem of minimizing weight matching, while the similarity matrix represents "maximizing similarity". Convert to cost matrix ,set up Represents the matching cost, and the formula is:
[0072] ;
[0073] The Hungarian algorithm constructs an initial zero matrix by subtracting the minimum value of each row or column from each row and column of the matrix, finds independent zero elements in the matrix, and matches the target with the feature. If the number of zero elements is not enough to complete all matches, adjust the uncovered elements and repeat the assignment process until the global optimal match is found. In order to solve the problem that the number of infrared targets and microwave features is unequal, the cost matrix Add virtual rows or columns to obtain a set of matching pairs using the Hungarian algorithm , which is used to represent the best match between infrared target i and microwave signature j.
[0074] Furthermore, the Step 7 includes:
[0075] Step 7-1, use the wavelet transform fusion algorithm to perform wavelet decomposition on the infrared image and microwave features respectively, and separate the infrared image and microwave features into low-frequency and high-frequency sub-bands. The low-frequency sub-band retains the global information, and the high-frequency sub-band captures the local details. represents the transformed signal, Represents the wavelet function, and the formula for wavelet transform is:
[0076] ;
[0077] in, is the scale coefficient of the wavelet, and n represents the offset or time step of the wavelet basis function;
[0078] By performing wavelet transform on the infrared image and microwave features, the infrared image and microwave features are decomposed into two parts, the low-frequency sub-band , and high frequency subband , ,
[0079] Step 7-2: Perform weighted average fusion on the low-frequency sub-bands to highlight the global features among multiple modalities. is the low-frequency subband after the final fusion, and the implementation formula is:
[0080] ;
[0081] in, , represents the dynamic weighting coefficients of the infrared and microwave low-frequency sub-bands, which are used to adjust the contribution of infrared and microwave in low-frequency fusion, and ;
[0082] Let SNR denote the signal-to-noise ratio, , The implementation formula is:
[0083] , ;
[0084] The calculation formula of signal-to-noise ratio SNR is:
[0085] ;
[0086] in is the image average, is the image standard deviation.
[0087] Step 7-3: Dynamically assign weights to high-frequency sub-bands using a cross-modal attention mechanism to perform detail enhancement fusion. Represents the fused high frequency subband. The implementation formula is:
[0088] ;
[0089] The cross-modal attention mechanism calculates the weight matrix through a convolutional neural network (CNN) and a sigmoid activation function. , used to control the degree of fusion of each high-frequency sub-band. Weight matrix The implementation formula is:
[0090] ;
[0091] Among them, CNN is a convolutional neural network, which is used to extract the features of high-frequency sub-bands. Sigmoid is an activation function that limits the output of the weight matrix to between [0,1], representing the fusion weight of each sub-band;
[0092] Perform inverse wavelet transform on the fused low-frequency and high-frequency sub-bands to reassemble them into a complete image. Represents inverse wavelet transform. The inverse wavelet transform formula is:
[0093] ;
[0094] Furthermore, the Step 8 specifically includes:
[0095] Using the three-dimensional block matching algorithm, similar blocks are searched in the fused image in units of fixed-size blocks. The mean square error is used to measure the similarity between blocks. Similar blocks are stacked to form 3D data and the denoised 3D data is inversely transformed. The formula for the mean square error is:
[0096] ;
[0097] in, and is the pixel value of the corresponding position of the two blocks;
[0098] Multi-scale pyramid Laplacian sharpening decomposes the image into Laplacian pyramids of different scales to highlight the detailed features of the image, enhance the edge and texture information, and is the image after sharpening, is the original fused image, represents the i-th layer of the Laplace pyramid, and the multi-scale pyramid Laplace sharpening formula is:
[0099] ;
[0100] in, is the enhancement coefficient, which is used to control the degree of enhancement of details at different scales. n is the number of layers of the Laplace pyramid, which determines the multi-scale range and affects the coverage of global and detail features.
[0101] The beneficial effects of the present invention are as follows: the present invention effectively solves the differences in feature expression form and semantic level between infrared images and microwave data by constructing a unified semantic representation space and an efficient feature fusion mechanism, fully mines the complementary information of the two modalities, and improves target recognition and environmental perception capabilities in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0103] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0104] Figure 2 It is a schematic diagram of the process of extracting spatial and semantic features from infrared images using convolutional neural network method;
[0105] Figure 3 Schematic diagram of the process of extracting frequency domain features from microwave non-image information using Fourier transform method;
[0106] Figure 4 A schematic diagram of a flow chart of the Hungarian algorithm used in an embodiment of the present invention;
[0107] Figure 5 Schematic diagram of the flow of the wavelet transform fusion algorithm in the embodiment of the present invention. DETAILED DESCRIPTION
[0108] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0109] Embodiment 1, as Figures 1 to 5 As shown, a semantic alignment method for fusing infrared images and microwave non-image information includes the following steps:
[0110] Step 1: Collect infrared images and microwave non-image information fragments of the same scene;
[0111] Step 2: Preprocess the infrared image collected in Step 1 by using non-local mean denoising and adaptive histogram equalization, and preprocess the microwave non-image information by using adaptive Wiener filtering and dynamic range compression plus gamma correction based on logarithmic transformation;
[0112] Step 3: Use convolutional neural network method to extract spatial and semantic features of infrared images, and use Fourier transform method to extract frequency domain features of microwave non-image information;
[0113] Step 4: Build a shared semantic space and use the twin network deep learning model and contrastive learning method to embed infrared image features and microwave image features into the same high-dimensional semantic space;
[0114] Step 5, using dynamically weighted cosine similarity to calculate the similarity between the infrared image feature target and the microwave image feature in semantic representation;
[0115] Step 6: Based on the similarity calculation results in Step 5, use the Hungarian algorithm to perform global optimal matching;
[0116] Step 7. Use the wavelet transform fusion algorithm to perform wavelet decomposition on the infrared image and microwave features respectively, and extract the low-frequency and high-frequency coefficients at different resolutions. Use the weighted average method to fuse the low-frequency coefficients, use the cross-modal attention mechanism to dynamically allocate weights, and combine the deep learning edge detection model to optimize the fusion of high-frequency features;
[0117] Step 8: Post-process the fused image by using a three-dimensional block matching algorithm to denoise the fused image and using a multi-scale pyramid Laplace sharpening algorithm to enhance the image edges and details;
[0118] Step 9. Save the fusion result as a high-precision single-frame image (TIFF / PNG).
[0119] As a preferred implementation, the Step 2 preprocesses the infrared image collected in Step 1 by using non-local mean denoising and adaptive histogram equalization, and preprocesses the microwave non-image information by using adaptive Wiener filtering and a dynamic range compression plus gamma correction method based on logarithmic transformation, specifically including the following steps:
[0120] Step 2.1-1, perform non-local mean denoising on the collected infrared image. Non-local mean denoising is an image denoising algorithm that considers similar pixel blocks in the entire image range for each pixel value and obtains the denoised pixel value based on the weighted average of the similarity. Assume Represents the intensity value of pixel x in the image after denoising; represents the raw intensity value of pixel y in the image; Represents the similarity weight of pixels x and y, indicating the degree of similarity between x and y; the non-local mean denoising formula is as follows:
[0121] ;
[0122] in, Represents the position coordinate of pixel x; Represents the pixel y position coordinate; Represents the domain of the image, that is, the coordinate set of all pixels in the image;
[0123] Similarity Weight It is calculated based on the similarity between the pixel blocks of x and y. The specific formula is as follows:
[0124] ;
[0125] in, The pixel position The image patch centered at the center; The pixel position The image patch centered at the center; is a filtering parameter used to control the decay speed of the weight;
[0126] Euclidean distance between two image patches The specific formula is as follows:
[0127] ;
[0128] in, and Represents the intensity values of corresponding pixels in two image blocks;
[0129] Step 2.1-2, use adaptive histogram equalization to perform local contrast enhancement on infrared images. By dividing the image into multiple local areas and performing histogram equalization on each area, the local contrast is enhanced, and the excessive amplification of noise is avoided by limiting the contrast enhancement amplitude (histogram clipping). Representing images The original pixel value at the position, Represents the pixel value after mapping, the formula for histogram equalization is:
[0130] ;
[0131] Where L is the number of gray levels of the image, is the frequency of gray value k, N is the total number of pixels in the image;
[0132] The contrast clipping formula is:
[0133] ;
[0134] in, Represents the clipped histogram; ClipLimit represents the clipping threshold, which limits the maximum frequency of each gray level; Excess is the sum of the clipped frequencies, which are redistributed to all gray levels;
[0135] Step 2.2-1, dynamically adjust the filtering strength of microwave non-image information according to the local statistical characteristics of each pixel, use noise power to distinguish noise and signal, use adaptive Wiener filtering to remove random noise while retaining the key characteristics of the target to avoid blurring of important signals; is the pixel value after filtering, is the original pixel value, The mean of the local window, is the variance of the local window, is the noise power, w is the local window defined around the pixel, and the formula for adaptive Wiener filtering is:
[0136] ;
[0137] The formulas for the local mean and local variance are:
[0138] ;
[0139] ;
[0140] in, is the number of pixels in the window;
[0141] The formula for noise power is:
[0142] ;
[0143] Where n represents the total number of pixels in the flat area, r represents the selected flat area, Represents the mean value of pixels in the flat area, which refers to the area in the image with small brightness changes and almost no significant features;
[0144] Step 2.2-2, dynamic range compression and gamma correction based on logarithmic transformation are used to adjust the contrast of microwave signals, and the grayscale distribution of data is adjusted through nonlinear transformation to highlight the weak signal details of microwave data, making it easier to match with infrared images in the process of semantic alignment and fusion. Represents the compressed pixel value, and the formula for logarithmic transformation is:
[0145] ;
[0146] The formula for gamma correction is:
[0147] ;
[0148] in, is the gamma value;
[0149] As a preferred implementation, Step 3 includes the following steps:
[0150] Step 3.1, use the convolutional neural network method to extract spatial and semantic features of infrared images. The convolutional neural network extracts features layer by layer, capturing the spatial and semantic information in the infrared image from local to global. The flowchart is shown in the figure. Figure 2 As shown. represents the convolution result, represents the pixel value of the input image, represents the weight of the convolution kernel, Represents the bias of the convolution kernel, and the convolution formula is:
[0151] ;
[0152] in, Indicates the current position of the convolution kernel, Represents the local offset of the convolution kernel, m and n represent the size of the convolution kernel;
[0153] After operations such as convolutional layers, pooling layers, and activation functions, the features of the infrared image are gradually extracted from local to global, and finally a high-dimensional feature vector is formed. ;
[0154] Step 3.2, use the Fourier transform method to extract the frequency domain features of microwave non-image information; transform the time series information of microwave data into frequency domain signals through Fourier transform, extract the component amplitudes of microwave signals at different frequencies, and highlight the key features. The flow chart is shown in the figure. Figure 3 As shown. Indicates the frequency The signal strength at Indicates yes The signal strength under Fourier transform formula is:
[0155] ;
[0156] Since the features after Fourier transformation are usually high-dimensional and may contain a lot of redundant information, they need to be reduced in dimension and normalized. Finally, the microwave features after Fourier transformation and processing are output as feature vectors. .
[0157] As a preferred implementation, Step 4 includes the following steps:
[0158] The high-dimensional infrared image feature vector and high-dimensional microwave frequency domain feature vector extracted from Step 3 are projected into a shared semantic space using a twin network deep learning model to facilitate modality alignment. The semantic representation alignment within and between modalities is performed by contrasting the loss function. is the embedded representation of the infrared image, indicating the location of the infrared image features in the shared semantic space. is the embedded representation of the microwave image, indicating the location of the microwave non-image information in the shared semantic space. The formula for the twin network embedding network is:
[0159] ;
[0160] ;
[0161] set up It represents contrast loss, and the formula of contrast loss is:
[0162] ;
[0163] Among them, Y is a binary label, which is used to guide the optimization of contrast loss, so that the feature distance of positive samples is closer and the feature distance of negative samples is farther. d represents the Euclidean distance, which is used to calculate the straight-line distance between two vectors. m represents the boundary threshold, which is used to control the penalty intensity of the loss function on the sample, ensuring that the distribution of positive and negative samples is reasonable.
[0164] The calculation formula of Euclidean distance is:
[0165] ;
[0166] As a preferred implementation, Step 5 includes the following steps:
[0167] By calculating similarity, we ensure that semantically identical objects are closer in the shared space, using dynamic weighting factors , To flexibly deal with the different significance of infrared and microwave features in different scenes, It represents the dynamic weighted cosine similarity. The formula of dynamic weighted cosine similarity is:
[0168] ;
[0169] in, , They represent the dynamic weight factors of infrared image features and microwave features when calculating similarity, respectively, and are used to adjust the importance of infrared image features and microwave features in similarity calculation.
[0170] As a preferred implementation, Step 6 includes the following steps:
[0171] According to the similarity calculation results in Step 5, the Hungarian algorithm is used to optimize the matching of infrared targets and microwave features in a global scope, and the total matching score is maximized by considering the similarity between all targets. The flow chart is shown in Figure 4. Represents the similarity between infrared features and microwave features. The similarity matrix formula is:
[0172] ;
[0173] Among them, the matrix Each element in Represents the matching score between infrared target i and microwave feature j, and its dimension is , where m is the number of targets in the infrared image and n is the number of microwave features;
[0174] The Hungarian algorithm solves the problem of minimizing weight matching, while the similarity matrix represents "maximizing similarity". Convert to cost matrix .set up Represents the matching cost, and the formula is:
[0175] ;
[0176] The Hungarian algorithm constructs an initial zero matrix by subtracting the minimum value of each row or column from each row and column of the matrix, finds independent zero elements in the matrix, and matches the target with the feature. If the number of zero elements is not enough to complete all matches, adjust the uncovered elements and repeat the assignment process until the global optimal match is found. In order to solve the problem that the number of infrared targets and microwave features is unequal, the cost matrix Add virtual rows or columns to obtain a set of matching pairs using the Hungarian algorithm , which is used to represent the best match between infrared target i and microwave signature j.
[0177] As a preferred implementation, Step 7 includes the following steps:
[0178] Step 7-1, use the wavelet transform fusion algorithm to perform wavelet decomposition on the infrared image and microwave features respectively, and separate the infrared image and microwave features into low-frequency and high-frequency sub-bands. The low-frequency sub-band retains the global information, and the high-frequency sub-band captures the local details. The process diagram is as follows Figure 5 As shown, represents the transformed signal, Represents the wavelet function. The formula for wavelet transform is:
[0179] ;
[0180] in, is the scale coefficient of the wavelet, and n represents the offset or time step of the wavelet basis function;
[0181] By performing wavelet transform on the infrared image and microwave features respectively, the infrared image and microwave features are decomposed into two parts. , and high frequency subband , ;
[0182] Step 7-2: Perform weighted average fusion on the low-frequency sub-bands to highlight the global features among multiple modalities. is the low-frequency subband after the final fusion, and the implementation formula is:
[0183] ;
[0184] in, , represents the dynamic weighting coefficients of the infrared and microwave low-frequency sub-bands, which are used to adjust the contribution of infrared and microwave in low-frequency fusion, and ;
[0185] Let SNR denote the signal-to-noise ratio, , The implementation formula is:
[0186] , ;
[0187] The calculation formula of signal-to-noise ratio SNR is:
[0188] ;
[0189] in is the image average, is the image standard deviation;
[0190] Step 7-3: Dynamically assign weights to high-frequency sub-bands using a cross-modal attention mechanism to perform detail enhancement fusion. Represents the fused high frequency subband. The implementation formula is:
[0191] ;
[0192] The cross-modal attention mechanism calculates the weight matrix through a convolutional neural network (CNN) and a sigmoid activation function. , used to control the degree of fusion of each high-frequency sub-band, the weight matrix The implementation formula is:
[0193] ;
[0194] Among them, CNN is a convolutional neural network, which is used to extract the features of high-frequency sub-bands. Sigmoid is an activation function that limits the output of the weight matrix to [0,1], representing the fusion weight of each sub-band;
[0195] Perform inverse wavelet transform on the fused low-frequency and high-frequency sub-bands to reassemble them into a complete image. Represents inverse wavelet transform. The inverse wavelet transform formula is:
[0196] ;
[0197] As a preferred implementation, Step 8 includes the following steps:
[0198] Using the three-dimensional block matching algorithm, similar blocks are searched in the fused image in units of fixed-size blocks. The mean square error is used to measure the similarity between blocks. Similar blocks are stacked to form 3D data and the denoised 3D data is inversely transformed. The formula for the mean square error is:
[0199] ;
[0200] in, and is the pixel value of the corresponding position of the two blocks;
[0201] Multi-scale pyramid Laplacian sharpening decomposes the image into Laplacian pyramids of different scales to highlight the detailed features of the image, enhance the edge and texture information, and is the image after sharpening, is the original fused image, represents the i-th layer of the Laplace pyramid, and the multi-scale pyramid Laplace sharpening formula is:
[0202] ;
[0203] in, is the enhancement coefficient, which is used to control the degree of enhancement of details at different scales. n is the number of layers of the Laplace pyramid, which determines the multi-scale range and affects the coverage of global and detail features.
[0204] In summary, the semantic alignment method based on the fusion of infrared image and microwave non-image information can effectively solve the differences in feature expression form and semantic level between infrared image and microwave data by constructing a unified semantic representation space and an efficient feature fusion mechanism, fully explore the complementary information of the two modalities, and enhance the target recognition and environmental perception capabilities in complex environments.
[0205] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A semantic alignment method for fusion of infrared image and microwave non-image information, characterized in that: The steps include: Step 1: Collect infrared images and microwave non-image information fragments of the same scene; Step 2: The infrared image collected in Step 1 is preprocessed by non-local mean denoising and adaptive histogram equalization, and the microwave non-image information is preprocessed by adaptive Wiener filtering and dynamic range compression plus gamma correction based on logarithmic transformation; Step 3: Use convolutional neural network method to extract spatial and semantic features of infrared images, and use Fourier transform method to extract frequency domain features of microwave non-image information; Step 4: Build a shared semantic space, use the twin network deep learning model and combine it with the contrastive learning method to embed the infrared image features and microwave non-image information features into the same high-dimensional semantic space; Step 5, using dynamically weighted cosine similarity to calculate the similarity between the infrared image feature target and the microwave non-image feature in semantic representation; Step 6: Based on the similarity calculation results in Step 5, use the Hungarian algorithm to perform global optimal matching; Step 7: Use the wavelet transform fusion algorithm to perform wavelet decomposition on the infrared image and microwave features respectively, and extract the low-frequency and high-frequency coefficients at different resolutions; use the weighted average method to fuse the low-frequency coefficients, use the cross-modal attention mechanism to dynamically allocate weights and combine the deep learning edge detection model to optimize the fusion of high-frequency features; Step 8: Post-process the fused image by using a three-dimensional block matching algorithm to denoise the fused image and using a multi-scale pyramid Laplace sharpening algorithm to enhance the image edges and details; Step 9. Save the optimized fusion result as a high-precision single-frame image.
2. The semantic alignment method for fusion of infrared image and microwave non-image information as claimed in claim 1, characterized in that: Step 2 specifically includes: Step 2.1-1, perform non-local mean denoising on the collected infrared image, and obtain the denoised pixel value by weighted average of each pixel value according to the similarity. Represents the intensity value of pixel x in the image after denoising; represents the raw intensity value of pixel y in the image; Represents the similarity weight of pixels x and y, indicating the degree of similarity between x and y; the non-local mean denoising formula is as follows: ; in, Represents the position coordinate of pixel x; Represents the pixel y position coordinate; Represents the domain of the image, that is, the coordinate set of all pixels in the image; Similarity Weight It is calculated based on the similarity between the pixel blocks of x and y. The specific formula is as follows: ; in, The pixel position The image patch centered at the center; The pixel position The image patch centered at the center; is a filtering parameter used to control the decay speed of the weight; Euclidean distance between two image patches The specific formula is as follows: ; in, and Represents the intensity values of corresponding pixels in two image blocks; Step 2.1-2, use adaptive histogram equalization to perform local contrast enhancement on infrared images. By dividing the image into multiple local areas and performing histogram equalization on each area, the local contrast is enhanced and excessive noise amplification is avoided by limiting the contrast enhancement amplitude. Representing images The original pixel value at the position, Represents the pixel value after mapping, the formula for histogram equalization is: ; Where L is the number of gray levels of the image, is the frequency of gray value k, N is the total number of pixels in the image; The contrast clipping formula is: ; in, Represents the clipped histogram; ClipLimit represents the clipping threshold, which limits the maximum frequency of each gray level; Excess is the sum of the clipped frequencies, which are redistributed to all gray levels; Step 2.2-1, dynamically adjust the filtering strength of microwave non-image information according to the local statistical characteristics of each pixel, use noise power to distinguish noise and signal, use adaptive Wiener filtering to remove random noise while retaining the key characteristics of the target to avoid blurring of important signals. is the pixel value after filtering, is the original pixel value, The mean of the local window, is the variance of the local window, is the noise power, w is the local window defined around the pixel, and the formula for adaptive Wiener filtering is: ; The formulas for the local mean and local variance are: ; ; in, is the number of pixels in the window; The formula for noise power is: ; Where n represents the total number of pixels in the flat area, r represents the selected flat area, Represents the mean value of pixels in the flat area; Step 2.2-2, the contrast of microwave signals is adjusted by dynamic range compression and gamma correction based on logarithmic transformation, and the grayscale distribution of data is adjusted by nonlinear transformation to highlight the weak signal details of microwave data; Represents the compressed pixel value, the logarithmic transformation formula is: ; The formula for gamma correction is: ; in, is the gamma value.
3. The semantic alignment method for fusion of infrared image and microwave non-image information as claimed in claim 2, characterized in that: Step 3 specifically includes: Step 3.1, use the convolutional neural network method to extract spatial and semantic features of infrared images. The convolutional neural network extracts features layer by layer, capturing the spatial and semantic information in the infrared image from local to global. represents the convolution result, represents the pixel value of the input image, represents the weight of the convolution kernel, Represents the bias of the convolution kernel, and the convolution formula is: ; in, Indicates the current position of the convolution kernel, Represents the local offset of the convolution kernel, m and n represent the size of the convolution kernel; After the convolution layer, pooling layer, and activation function operations, the features of the infrared image are gradually extracted from local to global, and finally a high-dimensional feature vector is formed. ; Step 3.2, use the Fourier transform method to extract the frequency domain features of microwave non-image information. Convert the time series information of microwave data into frequency domain signals through Fourier transform, extract the component amplitudes of microwave signals at different frequencies, highlight the key features, and set Indicates the frequency The signal strength at Indicates yes The signal strength under Fourier transform formula is: ; It is subjected to dimensionality reduction and normalization. Finally, the microwave feature after Fourier transformation and processing is output as a feature vector .
4. The semantic alignment method for fusion of infrared image and microwave non-image information as claimed in claim 3, characterized in that: Step 4 specifically includes: The high-dimensional infrared image feature vector and high-dimensional microwave frequency domain feature vector extracted from Step 3 are projected into a shared semantic space using the SiameseNetwork deep learning model to facilitate modality alignment. The semantic representation alignment within and between modalities is performed by contrasting the loss function. is the embedded representation of the infrared image, indicating the location of the infrared image features in the shared semantic space. It is the embedded representation of the microwave image, indicating the location of the microwave non-image information in the shared semantic space. The formula of the Siamese Network embedding network is: ; ; set up It represents contrast loss, and the formula of contrast loss is: ; Among them, Y is a binary label, d represents the Euclidean distance, and m represents the boundary threshold used to control the penalty intensity of the loss function on the sample; The calculation formula of Euclidean distance is: 。 5. The semantic alignment method for fusion of infrared image and microwave non-image information as claimed in claim 1, characterized in that: Step 5 specifically includes: set up It represents the dynamic weighted cosine similarity. The formula of dynamic weighted cosine similarity is: ; in, , They represent the dynamic weight factors of infrared image features and microwave features when calculating similarity, respectively, and are used to adjust the importance of infrared image features and microwave features in similarity calculation.
6. The semantic alignment method for fusion of infrared image and microwave non-image information as claimed in claim 5, characterized in that: Step 6 specifically includes: According to the similarity calculation results in Step 5, the Hungarian algorithm is used to optimize the matching of infrared targets and microwave features in a global range. Represents the similarity between infrared features and microwave features. The similarity matrix formula is: ; Among them, the matrix Each element in Represents the matching score between infrared target i and microwave feature j, and its dimension is , where m is the number of targets in the infrared image and n is the number of microwave features; The similarity matrix Convert to cost matrix ,set up Represents the matching cost, and the formula is: 。 7. The semantic alignment method for fusion of infrared image and microwave non-image information as claimed in claim 1, characterized in that: Step 7 specifically includes: Step 7-1, use the wavelet transform fusion algorithm to perform wavelet decomposition on the infrared image and microwave features respectively, and separate the infrared image and microwave features into low-frequency and high-frequency sub-bands. The low-frequency sub-band retains the global information, and the high-frequency sub-band captures the local details. represents the transformed signal, Represents the wavelet function, and the formula for wavelet transform is: ; in, is the scale coefficient of the wavelet, and n represents the offset or time step of the wavelet basis function; By performing wavelet transform on the infrared image and microwave features, the infrared image and microwave features are decomposed into two parts, the low-frequency sub-band , and high frequency subband , ; Step 7-2: Perform weighted average fusion on the low-frequency sub-bands to highlight the global features among multiple modalities. is the low-frequency subband after the final fusion, and the implementation formula is: ; in, , represents the dynamic weighting coefficients of the infrared and microwave low-frequency sub-bands, which are used to adjust the contribution of infrared and microwave in low-frequency fusion, and ; Let SNR denote the signal-to-noise ratio, , The implementation formula is: 、 ; The calculation formula of signal-to-noise ratio SNR is: ; in is the image average, is the image standard deviation; Step 7-3: Dynamically assign weights to high-frequency sub-bands using a cross-modal attention mechanism to perform detail enhancement fusion. Represents the fused high-frequency subband, and the implementation formula is: ; The cross-modal attention mechanism calculates the weight matrix through a convolutional neural network and a sigmoid activation function. , used to control the degree of fusion of each high-frequency sub-band, the weight matrix The implementation formula is: ; Among them, CNN is a convolutional neural network, which is used to extract the features of high-frequency sub-bands. Sigmoid is an activation function that limits the output of the weight matrix to between [0,1], indicating the fusion weight of each sub-band. Perform inverse wavelet transform on the fused low-frequency and high-frequency sub-bands to reassemble them into a complete image. Represents the inverse wavelet transform, and the inverse wavelet transform formula is: 。 8. The semantic alignment method for fusion of infrared image and microwave non-image information as claimed in claim 1, characterized in that: Step 8 specifically includes: Using the three-dimensional block matching algorithm, similar blocks are searched in the fused image in units of fixed-size blocks. The mean square error is used to measure the similarity between blocks. Similar blocks are stacked to form 3D data and the denoised 3D data is inversely transformed. The formula for the mean square error is: ; in, and is the pixel value of the corresponding position of the two blocks; Multi-scale pyramid Laplacian sharpening highlights the image's detail features and enhances edge and texture information by decomposing the image into Laplacian pyramids of different scales. is the image after sharpening, is the original fused image, represents the i-th layer of the Laplace pyramid, and the multi-scale pyramid Laplace sharpening formula is: ; in, is the enhancement coefficient, which is used to control the degree of enhancement of details at different scales. n is the number of layers of the Laplace pyramid, which determines the multi-scale range and affects the coverage of global and detail features.
Citation Information
Patent Citations
SAR, infrared and visible light image fusion method
CN105321172A
Rapid infrared and visible light image fusion method based on local and global information interaction, electronic equipment and storage medium
CN118172256A