Visible light infrared fusion target detection method based on characteristic decomposition
By using feature decomposition methods in object detection, the unique features and interactive features of visible light and infrared modes are extracted and fused, the limitations of object detection in complex scenarios are solved, and all-weather high-precision object detection is achieved.
Patent Information
- Application Number
- CN202510071352.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has limitations in object detection in complex scenarios, is susceptible to adverse weather and environmental factors, and the fusion divergence loss function is prone to misdetection.
The visible light infrared fusion object detection method based on feature decomposition is adopted, and the unique feature and interactive features are extracted through the unique feature encoder and interactive feature encoder, and the feature decomposition loss function constraint extraction process is used to integrate local comparison features and gradient enhancement features, remove redundant information in the interactive features, and finally perform fusion feature detection.
High accuracy of object detection in all-weather scenarios is achieved. By fully extracting and retaining unique features of different modes, the influence of redundant information is reduced, and the discriminationability and accuracy of detection is improved.
Smart Images

Figure CN120070852A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and image processing, and particularly relates to a visible light and infrared fusion target detection method based on feature decomposition. Background Art
[0002] In recent years, with the rapid development of computer vision, target detection has been increasingly widely used in fields such as security monitoring, autonomous driving, military reconnaissance, crop monitoring, remote sensing, and pedestrian detection. However, for target detection and recognition in complex scenarios, traditional visible light single-modal target detection methods have limitations and are vulnerable to adverse weather conditions, such as rain, fog, light changes, occlusion, and other environmental factors. Therefore, the target detection method based on visible light and infrared fusion shows great advantages. The infrared modality forms images based on thermal radiation information and can provide the thermal radiation information of the target in the case of insufficient light and occlusion, complementing the information of the visible light modality.
[0003] Beijing Institute of Technology proposed an infrared and visible light target detection method focusing on differential feature perception in its patent document "Infrared and Visible Light Target Detection Method Based on Differential Feature Perception in Adversarial Learning" (Application No.: CN202310655119.1, Publication No.: CN116681994A, Publication Date: September 1, 2023). Its main steps are as follows: (1) Based on the fusion divergence loss function composed of KL and JS divergences, non-shared feature extraction networks are used to extract differential infrared features and visible light features respectively; (2) The IR-Attention module and the RGB-Attention module are used to re-extract the infrared features and visible light features respectively; (3) The F-Attention module is used to fuse the re-extracted differential dual-modal features, pursuing more differential information while retaining the commonalities of the dual-modal features; (4) The RPN is used to perform regression and classification on the fused dual-modal features to complete target detection. However, the still existing deficiency of this method is that this fusion divergence loss function constraint pays more attention to differential unique features, and it is easy to cause misdetection and other situations due to modal differences.
[0004] Kunming University of Science and Technology proposed a fusion object detection method focusing on feature vector optimization and enhancement in its patent document "An Infrared-Visible Light Object Detection Method Based on Deformable Attention Mechanism" (Application No.: CN202311330611.8, Publication No.: CN117078920B, Publication Date: January 23, 2024). This method mainly includes the following steps: (1) Stitch the visible light and infrared feature maps and flatten them into vectors as input to the Transformer encoder, and optimize the global semantic features through the deformable attention mechanism; (2) Arrange the optimized feature vectors in descending order of eigenvalues, and select the top items as prior knowledge vectors to generate content query and coordinate query vectors for the classification branch and the regression branch; (3) Reshape the optimized feature vectors into the shape of the feature map, and generate a two-dimensional Gaussian score map based on the coordinate query vectors to update the feature map information; (4) In the Transformer decoder, use the deformable attention mechanism to perform cross-attention calculations on the content query vectors, coordinate query vectors, and feature maps, and output content prediction and coordinate prediction vectors; (5) Through the linear mapping layer and loss value calculation, compare the prediction vectors with the true values, and optimize the object detection network parameters based on the loss. However, the disadvantage of this method is that it is not easy to pay attention to the unique features with low similarity during the cross-attention calculation for optimizing feature vectors. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a visible light and infrared fusion object detection method based on feature decomposition. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0006] A visible light and infrared fusion object detection method based on feature decomposition includes:
[0007] S100, obtaining the original image, and extracting single-modal features therefrom by using a feature extraction backbone network with non-shared parameters; the single-modal features include visible light modal features and infrared modal features;
[0008] S200, extracting unique features and interaction features from the single-modal features by using a unique feature encoder and an interaction feature encoder, and constraining the extraction process by using a feature decomposition loss function to obtain unique features and interaction features; the feature decomposition loss function includes a first loss function with brightness constraint, a second loss function with gradient constraint, and a third loss function with similarity constraint;
[0009] S300, incorporating local contrast features and gradient enhancement features into the unique features to obtain enhanced unique features; the enhanced unique features include visible light enhanced unique features and infrared enhanced unique features;
[0010] S400, remove redundant information in the interaction features based on local and global features to obtain redundancy-removed interaction features;
[0011] S500, fuse the feature enhancement features and the redundancy-removed interaction features to obtain fused features;
[0012] S600, send the fused features into a detection head for object recognition and detection to obtain object detection results.
[0013] Beneficial effects:
[0014] First, since the present invention learns cross-modal specific features and interaction features by means of feature decomposition, and guides the network to learn the significant target information of the infrared modality and the texture detail information of the visible light modality through the designed loss constraint mechanism, the specific features of different modalities are fully extracted and retained.
[0015] Second, since the present invention utilizes local contrast and gradient priors to enhance the representation of specific features during the forward propagation process, the discriminability of the fused features of the present invention is greatly improved.
[0016] Third, since the present invention utilizes the global and local information of the interaction features between modalities to adaptively remove redundant information from the interaction features, the influence of redundant information in the fused features is reduced in the present invention.
[0017] The present invention will be further described in detail below in conjunction with the drawings and embodiments. Description of the Drawings
[0018] Figure 1 is a schematic flowchart of a visible light-infrared fusion object detection method based on feature decomposition provided by the present invention;
[0019] Figs. 2(a)-(d) are simulation effect diagrams provided by the present invention. Specific Embodiments
[0020] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0021] Refer to Figure 1 as shown, the present invention provides a visible light-infrared fusion object detection method based on feature decomposition, including:
[0022] S100, obtain an original image, and extract single-modal features therefrom by using a feature extraction backbone network with non-shared parameters; the single-modal features include visible light modality features and infrared modality features;
[0023] S200, extract specific features and interaction features from the unimodal features using a specific feature encoder and an interaction feature encoder, and use a feature decomposition loss function to constrain the extraction process to obtain specific features and interaction features; the feature decomposition loss function includes a first loss function for luminance constraint, a second loss function for gradient constraint, and a third loss function for similarity constraint;
[0024] The present invention constructs a feature decomposition loss function based on mutual information and modal characteristics to constrain the learning of visible light and infrared specific features and interaction features. The feature decomposition loss function consists of a first loss function for luminance constraint, a second loss function for gradient constraint, and a third loss function for similarity constraint.
[0025] In a specific embodiment of the present invention, S200 includes:
[0026] S210, construct a feature decomposition loss function;
[0027] In a specific embodiment of the present invention, S210 includes:
[0028] S211, use the infrared modal features and the infrared specific features to be extracted to construct a first loss function for luminance constraint;
[0029] Based on the characteristic that the target luminance of the infrared modality is significant, in this step, a loss function with luminance constraint is adopted for learning the infrared specific features The histogram distribution of the infrared modal features is denoted as The histogram distribution of the infrared specific features is denoted as Then the first loss function is expressed as:
[0030]
[0031] wherein, is the histogram distribution of the infrared modal features, is the histogram distribution of the infrared specific features, and N is the total number of feature map elements.
[0032] S212, use the visible light modal features and the visible light specific features to be extracted to construct a second loss function for gradient constraint;
[0033] Based on the characteristic that the texture details of the visible light modality are rich, in this step, a loss function with gradient constraint is adopted for learning the visible light specific features The gradient representation of the visible light modal features is denoted as The gradient representation of the visible light specific features is denoted as Then the second loss function is expressed as:
[0034]
[0035] Among them, is the gradient representation of the visible light modal feature ; is the gradient representation of the unique feature of visible light , where H is the height of the feature map size and W is the width of the feature map size.
[0036] S213, the unique feature of visible light extracted as needed and the unique feature of infrared The visible light interaction feature to be extracted and the infrared interaction feature are used to construct the third loss function of similarity constraint;
[0037] S213 includes:
[0038] S2131, calculate the unique feature of visible light and the unique feature of infrared The visible light interaction feature and the infrared interaction feature respectively to obtain the mutual information value between them;
[0039] This step measures the similarity between features based on the mutual information value, and further constrains the extraction representation of the unique feature and the interaction feature as:
[0040]
[0041] In the formula, represents the mutual information value between the unique feature of visible light and the unique feature of infrared ; represents the cross entropy between the unique feature of visible light and the unique feature of infrared ; represents the KL divergence between the unique feature of visible light and the unique feature of infrared ; represents the cross entropy between the unique feature of infrared and the unique feature of visible light ; represents the KL divergence between the unique feature of infrared and the unique feature of visible light ; represents the mutual information value between the visible light interaction feature and the infrared interaction feature ; represents the visible light interaction feature Cross-entropy with infrared interaction features ; Denote the KL divergence between visible light interaction features and infrared interaction features , Denote the cross-entropy between infrared interaction features and visible light interaction features ; Denote the KL divergence between infrared interaction features and visible light interaction features ;
[0042] S2132, Using the unique features of visible light and the unique features of infrared Visible light interaction features and infrared interaction features The mutual information value between them is used to construct the third loss function of similarity constraint, expressed as:
[0043]
[0044] where Denote the mutual information value between the unique features of visible light and the unique features of infrared , Denote the mutual information value between visible light interaction features and infrared interaction features ;
[0045] S214, Using the first loss function of brightness constraint, the second loss function of gradient constraint and the third loss function of similarity constraint, construct the feature decomposition loss function, expressed as:
[0046]
[0047] In the formula, λ is the balance loss factor.
[0048] S220, Using the decomposition loss function to constrain the extraction process of the unique feature encoder S 1 and S 2 , and respectively through the unique feature encoder S 1 and S 2 , extract the unique features of visible light and the unique features of infrared from the visible light modality features and the infrared modality features
[0049] S230, Using the decomposition loss function to constrain the interaction feature encoder C 1 and C2 extraction process, and respectively through the interactive feature encoder C 1 and C 2 , extract visible light interaction features and infrared modality features from the said visible light modality features and infrared interaction features
[0050] S300, incorporate local contrast features and gradient enhancement features into the said specific features to obtain enhanced specific features; the enhanced specific features include visible light enhanced specific features and infrared enhanced specific features;
[0051] In a specific embodiment of the present invention, S300 includes:
[0052] S310, for visible light specific features and infrared specific features both obtain visible light local contrast features and infrared local contrast features
[0053] In a specific embodiment of the present invention, S310 includes:
[0054] S311, for visible light specific features and infrared specific features select multiple scales to calculate the local feature contrast values of the feature maps;
[0055] S312, select the maximum local feature contrast value at multiple scales at the corresponding positions of the feature maps as the local contrast features, to obtain visible light local contrast features and infrared local contrast features which are respectively expressed as:
[0056]
[0057] wherein, δ is the sigmoid activation function, sup represents taking the maximum value for each feature map, d i represents the dilation scale in the local contrast calculation process, represents the local contrast feature map based on the visible light local contrast feature i at the dilation scale d , n represents the number of dilation scales d i ; represents the local contrast feature map based on the infrared local contrast feature i at the dilation scale d .
[0058] S320, for the specific features of visible light and the specific features of infrared Both enhance the representation ability of the specific features through differential convolution to obtain the visible light gradient enhanced feature and the infrared gradient enhanced feature
[0059]
[0060] In a specific embodiment of the present invention, S320 includes:
[0061] S321, perform an element-wise multiplication operation on the feature map between the visible light local contrast feature and the visible light gradient enhanced feature and use the specific feature of visible light as a residual connection to obtain the enhanced specific feature of visible light which is expressed as:
[0062]
[0063] S322, perform an element-wise multiplication operation on the feature map between the infrared local contrast feature and the infrared gradient enhanced feature and use the specific feature of infrared as a residual connection to obtain the enhanced specific feature of infrared which is expressed as:
[0064]
[0065] S330, fuse the visible light local contrast feature and the visible light gradient enhanced feature in the specific feature of visible light to obtain the enhanced specific feature of visible light and fuse the infrared local contrast feature and the infrared gradient enhanced feature in the specific feature of infrared to obtain the enhanced specific feature of infrared
[0066] S400, remove the redundant information in the interaction features based on the local features and global features to obtain the redundancy-removed interaction features;
[0067] In a specific embodiment of the present invention, S400 includes:
[0068] S410, using GCNorm convolution and learnable parameters, adaptively select important features from the interaction features for merging to obtain input features, and perform information interaction on the input features at the spatial level and the channel level to obtain local features; the local features are expressed by the formula:
[0069]
[0070] In the formula, M o is the input feature, δ is the sigmoid activation function, FC is the fully connected layer, Relu is the relu activation function, Cat is the feature channel merging operation, β is the learnable parameter, and GC is the GCNorm convolution;
[0071] S420, use a 3×3 depth convolution to process the interaction features to obtain non-local information, and merge the introduced computational divergence with the non-local information through a 1×1 depth convolution to obtain global features; the global features include visible light global features and infrared global features; the visible light global features and infrared global features are respectively expressed by the formula:
[0072]
[0073] Among them, σ 2 is the variance calculation function, Conv is the depth convolution, and DW is the depthwise separable convolution.
[0074] S430, according to the local features and the global features, adopt a progressive iterative method to remove redundancy from the interaction features to obtain redundancy-removed interaction features.
[0075] Adopt a progressive iterative method to remove redundancy from the interaction features to obtain redundancy-removed interaction features Add the local and global features, multiply them with the redundancy-removed feature obtained in the previous iteration, and connect the input interaction features as residuals to obtain the redundancy-removed feature in the current iteration process. The progressive iteration consists of n iterations, and the i-th iteration process can be expressed as:
[0076]
[0077] After n iterations, the redundancy-removed features can be merged to obtain redundancy-removed interaction features The redundancy-removed interaction features are expressed by the formula:
[0078]
[0079] In the formula, is the redundancy-removed interaction feature obtained at the end of the n-th iteration process, is the redundancy-removed interaction feature obtained at the end of the i-th iteration process, The redundancy-removed interaction feature obtained at the end of the (i - 1)-th iteration process, where n is the number of iterations.
[0080] Among them, the redundancy-removed interaction feature includes a visible light redundancy-removed interaction feature and an infrared redundancy-removed interaction feature.
[0081] S500, fuse the feature enhancement feature and the redundancy-removed interaction feature to obtain a fused feature.
[0082] This step uses a 1×1 convolution operation to perform a merging and fusing operation on the enhanced specific feature and the redundancy-removed interaction feature in the channels to obtain a fused feature. It is expressed as:
[0083]
[0084] In the formula, is the visible light enhanced specific feature, is the infrared enhanced specific feature, is the redundancy-removed interaction feature obtained at the end of the n-th iteration process.
[0085] S600, send the fused feature into a detection head for target recognition and detection to obtain a target detection result.
[0086] The present invention provides a visible light and infrared fusion target detection method based on feature decomposition, including: using a specific feature encoder and an interaction feature encoder to extract specific features and interaction features from single-modal features, and using a feature decomposition loss function to constrain the extraction process to obtain specific features and interaction features; incorporating local contrast features and gradient enhancement features into the specific features to obtain enhanced specific features; removing redundant information in the interaction features based on local features and all features; fusing the enhanced specific features and the redundancy-removed interaction features to obtain a fused feature and then sending it into a detection head for target recognition and detection to obtain a target detection result. The present invention obtains specific features and interaction features through feature decomposition, enhances the representation of the specific features respectively, removes redundancy from the interaction features, and obtains a highly discriminative fused feature, so the detection accuracy of target detection can be improved. The present invention has the advantage of high target detection accuracy in all-weather scenarios.
[0087] The following further illustrates the effect of the present invention through simulation experiments:
[0088] 1. Simulation experiment conditions:
[0089] The hardware platform for the simulation experiment of the present invention is: the processor is 13th Gen Intel(R) Core(TM) i9-13900K, the main frequency is 3.00 GHz, the memory is 64 GB, and the GPU is RTX 4090.
[0090] The software platform for the simulation experiment of the present invention is: Windows 10 operating system, python3.8, PyTorch2.4.1, and cuda12.1.
[0091] The simulation experiment of the present invention was carried out on the 3 MFD dataset to verify the proposed algorithm. It contains six types of targets, namely people, buses, cars, motorcycles, lights, and trucks.
[0092] 2. Simulation content and result analysis:
[0093] In the simulation experiment of the present invention, several advanced visible light and infrared fusion object detection methods were selected as comparison algorithms. The mean average precision mAP with an intersection over union of 0.5 50 was used as an evaluation metric to evaluate the performance of the object detection algorithm. The experimental results are shown in Table 1:
[0094] Table 1 Comparison with other algorithms on the 3 MFD dataset
[0095]
[0096]
[0097] Combined with Table 1, it can be seen that the method in this paper achieved an average precision of 87.5%. Compared with the existing best-performing MMFN network, the precision was improved by 1.3%. In the motorcycle category, it was improved by 6.7%, in the person category by 4.4%, and in the truck category by 3.2%, reaching precisions of 80.4%, 87.4%, and 90.6% respectively. This proves that our method can effectively represent and fuse cross-modal complementary information. By enhancing the representation of specific features and removing the redundancy of interactive features, a rich fused feature representation was achieved.
[0098] The effect of the present invention will be further described below in combination with the simulation diagrams:
[0099] Figure 2(a) shows the detection ground truth bounding box of the image in the simulation experiment of the present invention. Figure 2(b) shows the result diagram of detecting the visible light image using the visible light single-modal detection model in the simulation experiment of the present invention. Figure 2(c) shows the result diagram of detecting the infrared image using the infrared single-modal detection model in the simulation experiment of the present invention. Figure 2(d) shows the result diagram of fusing and detecting the dual-channel input of visible light and infrared images using the method of the present invention in the simulation experiment of the present invention.
[0100] As can be seen from Figure 2(b), in a low-light environment, the visible light modality lacking texture information is prone to confusing features and difficult to extract recognizable target features, resulting in some false detections and missed detections (marked by red circles).
[0101] As can be seen from Fig. 2(c), due to its advantages in thermal radiation imaging, the infrared modality makes the targets in low-light environments more obvious than in the visible light modality, but there are still cases of missed detection (marked by red circles).
[0102] As can be seen from Fig. 2(d), the detection results of the present invention exclude false targets and accurately identify and detect all target objects. It proves that the present invention can effectively combine the advantages of the two modalities, eliminate the interference caused by the modality gap, and achieve accurate target detection under all-weather conditions.
[0103] The above simulation experiments show that: The method of the present invention designs a feature decomposition loss function according to the characteristics of the modalities and the correlation constraints between features, so as to constrain the learning representations of cross-modal specific features and interaction features. While fully characterizing the cross-modal complementary information, it achieves the purpose of bridging the modality gap. In addition, in order to enhance the fusion representation, first, the gradient and local contrast prior are combined to enhance the representation of specific features; second, the global and local features are combined to remove redundant information in the interaction features. By enhancing the representation of specific features and removing redundant information in the interaction features, the discriminability of the fusion features is improved. At the same time, the experiments prove that this method improves the accuracy of target detection in all-weather scenarios and is a very practical visible light and infrared fusion target detection method.
[0104] It should be noted that the terms "first" and "second" in the present invention are only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0105] Although the present application has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed present application, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure content, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of cases.
[0106] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A visible light infrared fusion target detection method based on feature decomposition, characterized in that: include: S100, acquiring an original image, and extracting single-modal features from the original image using a feature extraction backbone network that does not share parameters; the single-modal features include visible light modal features and infrared modal features; S200, extracting unique features and interactive features from the unimodal features using a unique feature encoder and an interactive feature encoder, and constraining the extraction process using a feature decomposition loss function to obtain unique features and interactive features; The feature decomposition loss function includes a first loss function of brightness constraint, a second loss function of gradient constraint and a third loss function of similarity constraint; S300, integrating the local contrast feature and the gradient enhancement feature into the unique feature to obtain an enhanced unique feature; The enhanced unique features include visible light enhanced unique features and infrared enhanced unique features; S400, removing redundant information in the interactive features based on local features and global features to obtain de-redundant interactive features; S500, fusing the feature enhancement feature with the de-redundant interaction feature to obtain a fused feature; S600: Send the fused features to a detection head for target recognition and detection to obtain a target detection result.
2. The visible light infrared fusion target detection method based on feature decomposition according to claim 1 is characterized in that: S200 includes: S210, constructing a feature decomposition loss function; S220, using the decomposition loss function to constrain the extraction process of the unique feature encoders S1 and S2, and respectively extracting the visible light modal features through the unique feature encoders S1 and S2 and infrared modal characteristics Extracting the unique features of visible light and infrared characteristics S230, using the decomposition loss function to constrain the extraction process of the interactive feature encoders C1 and C2, and extracting the visible light modal feature from the interactive feature encoders C1 and C2 respectively. and infrared modal characteristics Extracting visible light interaction features Interactive features with infrared 3. The visible light infrared fusion target detection method based on feature decomposition according to claim 2 is characterized in that: S210 includes: S211, using infrared modal features And the infrared unique features that need to be extracted Construct the first loss function of brightness constraint, expressed as: in, is the histogram distribution of infrared modal features, is the histogram distribution of infrared-specific features, and N is the total number of feature map elements; S212, using visible light modal features and the unique features of visible light that need to be extracted Construct the second loss function with gradient constraints, expressed as: in, is the visible light modal feature The gradient of Unique features of visible light The gradient representation of , H is the height of the feature map size, and W is the width of the feature map size; S213, extract the unique features of visible light as needed With infrared characteristics Visible light interaction features that need to be extracted Interaction with infrared features The third loss function for constructing similarity constraints is expressed as: in, Indicates the unique characteristics of visible light With infrared characteristics The mutual information value between Represents the visible light interaction feature Interaction with infrared features The mutual information value between ; S214, constructing a feature decomposition loss function using the first loss function of the brightness constraint, the second loss function of the gradient constraint, and the third loss function of the similarity constraint, expressed as: Where λ is the balance loss factor.
4. The visible light infrared fusion target detection method based on feature decomposition according to claim 3 is characterized in that: S213 includes: S2131, calculate the unique features of visible light With infrared characteristics Visible light interaction features Interaction with infrared features The mutual information value between is expressed as: In the formula, Indicates the unique characteristics of visible light With infrared characteristics The mutual information value between Indicates the unique characteristics of visible light With infrared characteristics The cross entropy between Indicates the unique characteristics of visible light With infrared characteristics The KL divergence between Indicates infrared characteristics Unique features of visible light The cross entropy between Indicates infrared characteristics Unique features of visible light The KL divergence between Represents the visible light interaction feature Interaction with infrared features The mutual information value between ; Represents the visible light interaction feature Interaction with infrared features The cross entropy between Represents the visible light interaction feature Interaction with infrared features The KL divergence between Indicates infrared interaction characteristics Interaction characteristics with visible light The cross entropy between Indicates infrared interaction characteristics Interaction characteristics with visible light The KL divergence between S2132, using the unique characteristics of visible light With infrared characteristics Visible light interaction features Interaction with infrared features The mutual information value between them is used to construct the third loss function of similarity constraint.
5. The visible light infrared fusion target detection method based on feature decomposition according to claim 2 is characterized in that: S300 includes: S310, for the unique characteristics of visible light and infrared characteristics The visible light local contrast features are obtained by multi-scale local contrast calculation. and infrared local contrast features S320, for the unique characteristics of visible light and infrared characteristics Differential convolution is used to enhance the representation ability of unique features and obtain visible light gradient enhancement features. and infrared gradient enhancement features S330, unique features in visible light Fusion of visible light local contrast features and visible light gradient enhancement features Obtaining unique features of visible light enhancement And the unique features in infrared Fusion of infrared local contrast features and infrared gradient enhancement features Obtain infrared enhanced unique features 6. The visible light infrared fusion target detection method based on feature decomposition according to claim 5 is characterized in that: S310 includes: S311, for the unique characteristics of visible light and infrared characteristics Select multiple scales to calculate the local feature contrast value of the feature map; S312, selecting the maximum value of the local feature contrast at multiple scales at the corresponding position of the feature map as the local contrast feature, and obtaining the visible light local contrast feature and infrared local contrast features Respectively expressed as: In the formula, δ is the sigmoid activation function, sup means taking the maximum value for each feature map, and d i represents the dilation scale in the local contrast calculation process, Indicates that the expansion scale d i Based on the local contrast feature of visible light The local contrast feature map, n represents the expansion scale d i The number of Indicates that the expansion scale d i Based on infrared local contrast features The local contrast feature map of .
7. The visible light infrared fusion target detection method based on feature decomposition according to claim 5 is characterized in that: S320 includes: S321, local contrast features of visible light are added to the feature map and visible light gradient enhancement features Perform element-by-element multiplication and convert the unique features of visible light into As a residual connection, obtain the unique features of visible light enhancement It is expressed as: S322, infrared local contrast features are added to the feature map and infrared gradient enhancement features Perform element-by-element multiplication and convert infrared-specific features As a residual connection, obtain infrared enhanced unique features It is expressed as:
8. The visible light infrared fusion target detection method based on feature decomposition according to claim 2 is characterized in that: S400 includes: S410, using GCNorm convolution and learnable parameters, adaptively selecting important features from the interaction features and merging them to obtain input features, and performing information interaction on the input features at the spatial level and the channel level to obtain local features; S420, using 3x3 depth convolution to process the interactive features to obtain non-local information, and combining the introduced computational divergence with the non-local information through 1x1 depth convolution to obtain global features; the global features include visible light global features and infrared global features; S430: According to the local features and the global features, remove redundancy from the interactive features in a progressive iterative manner to obtain de-redundant interactive features.
9. The visible light infrared fusion target detection method based on feature decomposition according to claim 8 is characterized in that: The local features in S410 are expressed as: Where M o is the input feature, δ is the sigmoid activation function, FC is the fully connected layer, Relu is the relu activation function, Cat is the feature channel merging operation, β is the learnable parameter, and GC is the GCNorm convolution; The visible light global features and infrared global features in S420 are expressed by the following formulas: Among them, σ 2 is the variance calculation function, Conv is the depth convolution, and DW is the depth separable convolution; The redundant interaction feature in S430 is expressed as follows: In the formula, is the de-redundant interaction feature obtained at the end of the nth iteration process, is the de-redundant interaction feature obtained at the end of the i-th iteration process, It is the de-redundant interaction feature obtained at the end of the i-1th iteration process, and n is the number of iterations.
10. The visible light infrared fusion target detection method based on feature decomposition according to claim 1 is characterized in that: S500 includes: A 1x1 convolution operation is used to merge and fuse the enhanced unique features and the de-redundant interactive features on the channel to obtain the fused features. It is expressed as: In the formula, It is a unique feature of visible light enhancement. It is a unique feature of infrared enhancement. It is the de-redundant interaction feature obtained at the end of the nth iteration process.
Citation Information
Patent Citations
Infrared visible light target detection method based on differential feature perception in adversarial learning
CN116681994A
Infrared-visible light target detection method based on deformable attention mechanism
CN117078920A
An infrared-visible light target detection method based on deformable attention mechanism
CN117078920B