A battlefield target detection system and method based on visible light and infrared image fusion

By fusing the characteristics of visible light and infrared images in the battlefield target detection system and combining the temperature difference sensing module, the problem of inaccurate target detection in complex battlefield environments is solved, and high-precision all-weather target detection capabilities are achieved, providing a reliable basis for battlefield command decisions.

CN117765359BActive Publication Date: 2025-05-06ARMOR ACADEMY OF CHINESE PEOPLES LIBERATION ARMY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311595532.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-06
Estimated Expiration
2043-11-27

AI Technical Summary

Technical Problem

In complex battlefield environments, it is difficult for the existing technology to accurately detect target characteristics and locations during heavy fog weather or night combat, resulting in inaccurate judgment of battlefield target conditions and unable to provide a reliable basis for war command decisions.

Method used

The battlefield target detection system based on the fusion of visible light and infrared images is adopted, and the features of visible light and infrared images are extracted respectively through a dual-channel feature extraction network, combined with the temperature difference perception module and feature fusion module, the spatial alignment and fusion of visible light and infrared images are realized, and the final fusion feature map is generated to predict the target category and position prediction.

Benefits of technology

It improves the target detection accuracy in complex ground battlefield environments, eliminates the adverse effects of various interference factors, realizes all-weather target detection capabilities, can automatically spatially align multi-spectral image features, perceives the temperature differences between different objects in infrared images, further improves detection accuracy and efficiency, and provides a reliable basis for command decisions in the battlefield.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117765359B_ABST
    Figure CN117765359B_ABST
Patent Text Reader

Abstract

The present invention relates to a battlefield target detection system based on the fusion of visible light and infrared images, including a dual-path feature extraction network, a feature alignment module, a temperature difference perception module, a feature fusion module, a neck network and a task head network, wherein the dual-path feature extraction network is used to extract the features of the visible light image and the features of the infrared image respectively; the feature alignment module is used to spatially align the features of the visible light image and the infrared image; the temperature difference perception module is used to obtain a predicted temperature difference mask; the feature fusion module is used to generate a fusion feature map; the neck network and the task head network are used to complete target category prediction and target position prediction through convolution processing according to the fusion feature map provided by the neck network, so as to obtain the final prediction result. The present invention can better cope with complex ground battlefield environments, eliminate the adverse effects of multiple interference factors, improve detection accuracy and efficiency, and provide a reliable basis for command decision-making in the battlefield.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of military target detection, and in particular to a battlefield target detection system and method based on the fusion of visible light and infrared images. Background Art

[0002] With the continuous development of military technology, battlefield target detection technology has been more and more widely used. For example, in war, the army can identify enemy targets to determine the enemy's strength, military equipment and other information; in anti-terrorism operations, the police can track and combat terrorists by identifying suspected targets; in maritime patrols, the coast guard can identify suspicious ships to combat pirates and illegal fishing boats. With the development of battlefield target detection technology, it can not only help the army win in wars, but also help the police fight terrorism and maintain social stability. Battlefield target detection technology refers to the accurate identification and judgment of targets on the ground, sea or in the air through specific technical means, so as to provide accurate guidance and guarantee for military operations. However, the actual battlefield environment is usually more complex, and there are many interference factors when detecting battlefield targets, such as foggy weather, night operations, etc., which leads to unclear target features and inaccurate target positions detected by information technology, thereby affecting the judgment of target conditions on the battlefield and failing to provide a reliable basis for war command decisions. Summary of the invention

[0003] The present invention intends to provide a battlefield target detection system and method based on the fusion of visible light and infrared images to solve the deficiencies in the prior art. The technical problem to be solved by the present invention is achieved through the following technical solutions.

[0004] A battlefield target detection system based on the fusion of visible light and infrared images comprises a dual-path feature extraction network, a feature alignment module, a temperature difference perception module, a feature fusion module, a neck network and a task head network, wherein the dual-path feature extraction network is used to respectively extract features of visible light images and obtain visible light feature maps, and extract features of infrared images and obtain infrared feature maps; the feature alignment module is used to spatially align features of visible light images and infrared images; the temperature difference perception module is used to obtain a predicted temperature difference mask after convolution processing of the infrared feature map; the feature fusion module is used to fuse the visible light feature map with the infrared feature map in combination with the temperature difference mask and generate a fused feature map; the neck network is used to further fuse the deep fusion feature map and the shallow fusion feature map; the task head network is used to complete target category prediction and target position prediction through convolution processing according to the fusion feature map provided by the neck network, so as to obtain a final prediction result.

[0005] Preferably, the dual-path feature extraction network includes a visible light image feature extraction network and an infrared image feature extraction network. The visible light image feature extraction network is used to extract shallow and deep features of the visible light image and obtain a visible light feature map, and the infrared image feature extraction network is used to extract shallow and deep features of the infrared image and obtain an infrared feature map.

[0006] Preferably, the neck network adopts an FPN network or a PAN network.

[0007] A battlefield target detection method based on visible light and infrared image fusion includes the following steps:

[0008] Step 1, feature extraction: extract the features of the visible light image and obtain a visible light feature map, extract the features of the infrared image and obtain an infrared feature map through a dual-path feature extraction module;

[0009] Step 2: feature alignment: spatially align the features of the visible light image and the infrared image through a feature alignment module;

[0010] Step 3: Temperature difference mask prediction: The infrared feature map is convolved by the temperature difference perception module to obtain a predicted temperature difference mask;

[0011] Step 4: feature fusion: The visible light feature map and the infrared feature map are fused through the feature fusion module to generate a preliminary fused feature map, and then the preliminary fused feature map and the temperature difference mask prediction value are feature weighted by multiplying the corresponding spatial position features to obtain the final fused feature map;

[0012] Step 5: Get the final prediction result. After the fusion feature map is processed by the neck network and the task head network, the target category prediction and target position prediction are obtained, and the final prediction result is obtained after the target bounding box regression processing.

[0013] Preferably, in step 2, the method for spatially aligning the features of the visible light image and the infrared image through the feature alignment module is to perform channel splicing on the visible light image features and the infrared image features to obtain a spliced ​​feature map; after the spliced ​​feature map passes through the convolution block, a bias parameter offset of a feasible variable convolution is obtained, and the bias parameter offset represents the size of the feasible variable convolution kernel; after the infrared feature map passes through the feasible variable convolution, an aligned infrared feature map is obtained.

[0014] Preferably, in step 3, the infrared feature map is subjected to the convolution block of the temperature difference perception module to obtain a predicted temperature difference mask; during the convolution training process, the target area image and the target and background area images are extracted respectively, and the mean and variance of the regional pixels are calculated to obtain the mean u1 and variance σ1 of the target area image, the mean u2 and variance σ2 of the target and background area images, and a single Gaussian model N1 (u1, σ1) of the target area and a single Gaussian model N2 (u2, σ2) of the target and background area are constructed; the KL divergence D of the two single Gaussian models is calculated. KL (N1(u1,σ1)||N2(u2,σ2)); the label value of the temperature difference mask in the target area is calculated, and the predicted value calculation method of the temperature difference mask in the target area is Among them, α is a hyperparameter, D KL ∈[0,+∞), T label ∈[0,1).

[0015] Preferably, in step 4, the calculation formula of the fusion feature map is F fus =(F T +F I )·e β*T , where F fus represents the fusion feature map, F T Represents the visible light feature map, F I represents the infrared feature map, T represents the predicted value of the temperature difference mask, and β is a hyperparameter used to control the degree of feature enhancement.

[0016] Preferably, in step 5, the target bounding box regression processing method is: providing a pair of visible light and infrared images; taking the visible light image as the main state and the infrared image as the auxiliary state, and using the annotation information of the visible light image as the label information; obtaining the temperature difference mask prediction value, the category prediction cls_predict and the target space position prediction reg_predict, and calculating the total loss loss in combination with the label information, the calculation formula is loss=loss_T+loss_cls+loss_reg, wherein loss_T represents the temperature difference mask loss, and the calculation formula is: loss_T=(1-T label )log(1-T)+T label log(T), loss_cls represents category loss, loss_reg represents target position regression loss, loss_cls uses cross entropy loss function, loss_reg uses IOU loss, and after the total loss loss is back-propagated, the network weight parameters are updated.

[0017] Preferably, in step one, visible light image features are extracted by the visible light image feature extraction network in the dual-path feature extraction network, and visible light feature maps of different scales at three levels are obtained respectively, and infrared image features are extracted by the infrared image feature extraction network in the dual-path feature extraction module, and infrared feature maps of different scales at three levels are obtained respectively.

[0018] Preferably, in step 2, the features of the visible light feature map and the infrared feature map at the same level of the dual-path feature extraction network are aligned through a feature alignment module to obtain three aligned infrared feature maps based on the visible light feature map; in step 3, the three aligned infrared feature maps are processed by a temperature difference perception module to obtain predicted values ​​of three temperature difference masks respectively; in step 4, the predicted values ​​of the aligned infrared feature map, the visible light feature map and the temperature difference mask are processed by a feature fusion module to obtain preliminary fused feature maps of three levels, and the final fused feature map is obtained after combining the temperature difference mask processing.

[0019] The present invention provides a battlefield target detection system based on the fusion of visible light and infrared images, which extracts visible light image and infrared image features through convolution blocks, and further completes the spatial alignment and fusion of the two features, thereby improving the target detection accuracy in a complex ground battlefield environment. The battlefield target detection system in the present invention can better cope with the complex ground battlefield environment, eliminate the adverse effects of various interference factors, has all-weather target detection capabilities, can automatically perform spatial alignment of multi-spectral image features, perceive the temperature differences of different objects in infrared images, and further improve detection accuracy and efficiency by fusing features, providing a reliable basis for command decisions in the battlefield. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a structural block diagram of the battlefield target detection system in the present invention;

[0021] Figure 2 It is a flowchart of the workflow of the feature alignment module in the present invention;

[0022] Figure 3 This is a flowchart of the temperature difference sensing module in the present invention;

[0023] Figure 4 This is a flowchart of the workflow of the feature fusion module in the present invention. DETAILED DESCRIPTION

[0024] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0025] Embodiment 1:

[0026] Reference Figure 1 As shown, a battlefield target detection system based on the fusion of visible light and infrared images, the improvement of which lies in: comprising a dual-path feature extraction network, a feature alignment module, a temperature difference perception module, a feature fusion module, a neck network and a task head network, the dual-path feature extraction network is used to extract the features of the visible light image and obtain a visible light feature map, extract the features of the infrared image and obtain an infrared feature map; the feature alignment module is used to spatially align the features of the visible light image and the infrared image; the temperature difference perception module is used to obtain a predicted temperature difference mask after convolution processing on the infrared feature map; the feature fusion module is used to fuse the visible light feature map with the infrared feature map in combination with the temperature difference mask and generate a fused feature map; the neck network is used to further fuse the deep fusion feature map and the shallow fusion feature map, and the task head network is used to complete the target category prediction and the target position prediction through convolution processing according to the fusion feature map provided by the neck network to obtain the final prediction result.

[0027] Furthermore, the dual-path feature extraction network includes a visible light image feature extraction network and an infrared image feature extraction network. The visible light image feature extraction network is used to extract shallow and deep features of the visible light image and obtain a visible light feature map, and the infrared image feature extraction network is used to extract shallow and deep features of the infrared image and obtain an infrared feature map. The visible light image feature extraction network and the infrared image feature extraction network have the same network structure, and extract shallow and deep features of the visible light image and the infrared image, respectively.

[0028] Furthermore, the dual-path feature extraction network adopts a resnet series feature extraction network or a darknet series feature extraction network.

[0029] Furthermore, the neck network adopts an FPN network or a PAN network. The neck network fuses the deep fusion feature map and the shallow fusion feature map, so that the three feature maps sent to the task head network have both shallow texture features and deep semantic features.

[0030] Furthermore, the task head network is composed of multiple convolutions.

[0031] The battlefield target detection system based on the fusion of visible light and infrared images provided in this embodiment extracts visible light image and infrared image features through convolution blocks, and further completes the spatial alignment and fusion of the two features, thereby improving the target detection accuracy in a complex ground battlefield environment. The battlefield target detection system in this embodiment can better cope with the complex ground battlefield environment, eliminate the adverse effects of various interference factors, has all-weather target detection capabilities, can automatically align multi-spectral image features in space, perceive the temperature differences of different objects in infrared images, and further improve the detection accuracy and efficiency by fusing features, providing a reliable basis for command decisions in the battlefield.

[0032] Embodiment 2:

[0033] Reference Figure 1 As shown, a battlefield target detection method based on visible light and infrared image fusion, the improvement of which is that it includes the following steps:

[0034] Step 1, feature extraction: extract the features of the visible light image and obtain a visible light feature map, extract the features of the infrared image and obtain an infrared feature map through a dual-path feature extraction module;

[0035] Step 2: feature alignment: spatially align the features of the visible light image and the infrared image through a feature alignment module;

[0036] Step 3: Temperature difference mask prediction: The infrared feature map is convolved by the temperature difference perception module to obtain a predicted temperature difference mask;

[0037] Step 4: feature fusion: The visible light feature map and the infrared feature map are fused through the feature fusion module to generate a preliminary fused feature map, and then the preliminary fused feature map and the temperature difference mask prediction value are feature weighted by multiplying the corresponding spatial position features to obtain the final fused feature map;

[0038] Step 5: Get the final prediction result. After the fusion feature map is processed by the neck network and the task head network, the target category prediction and target position prediction are obtained, and the final prediction result is obtained after the target bounding box regression processing.

[0039] Furthermore, in step 2, the method for spatially aligning the features of the visible light image and the infrared image through the feature alignment module is to perform channel splicing on the visible light image features and the infrared image features to obtain a spliced ​​feature map; after the spliced ​​feature map passes through the convolution block, the bias parameter offset of the feasible variable convolution is obtained, and the bias parameter offset represents the size of the feasible variable convolution kernel; after the infrared feature map passes through the feasible variable convolution, an aligned infrared feature map is obtained.

[0040] Further, see Figure 3As shown in FIG. 1 , in step 3, the infrared feature map is subjected to the convolution block of the temperature difference perception module to obtain the predicted temperature difference mask; during the convolution training process, the target area image and the target and background area images are extracted respectively, and the mean and variance of the regional pixels are calculated to obtain the mean u1 and variance σ1 of the target area image, the mean u2 and variance σ2 of the target and background area images, and the single Gaussian model N1(u1,σ1) of the target area and the single Gaussian model N2(u2,σ2) of the target and background area are constructed; the KL divergence D of the two single Gaussian models is calculated. KL (N1(u1,σ1)||N2(u2,σ2)); the label value of the temperature difference mask in the target area is calculated, and the predicted value calculation method of the temperature difference mask in the target area is Among them, α is a hyperparameter, D KL ∈[0,+∞), T label ∈[0,1). The predicted temperature difference mask represents the temperature difference between the area represented by the current feature point and the surrounding area. The larger the temperature difference mask value, the greater the temperature difference between the area and the surrounding area, and the greater the possibility of the existence of the target.

[0041] Further, see Figure 4 As shown in step 4, the calculation formula for the fusion feature map is F fus =(F T +F I )·e β*T , where F fus represents the fusion feature map, F T Represents the visible light feature map, F I represents the infrared feature map, T represents the predicted value of the temperature difference mask, and β is a hyperparameter used to control the degree of feature enhancement. In the feature fusion module, the feature map addition method is used to achieve the complementarity of different features. The temperature difference mask acts as a spatial attention on the fusion feature map. In the fusion process, different weights are given to the fusion features at different positions to enhance the features of important areas.

[0042] Furthermore, in step 5, the target bounding box regression processing method is: provide a pair of visible light and infrared images; take the visible light image as the main state and the infrared image as the auxiliary state, and use the annotation information of the visible light image as the label information; obtain the temperature difference mask prediction value, the category prediction cls_predict and the target space position prediction reg_predict, and calculate the total loss loss in combination with the label information, and the calculation formula is loss=loss_T+loss_cls+loss_reg, where loss_T represents the temperature difference mask loss, and its calculation formula is: loss_T=(1-T label )log(1-T)+T labellog(T), loss_cls represents category loss, loss_reg represents target position regression loss, loss_cls uses cross entropy loss function, loss_reg uses IOU loss, and after the total loss loss is back-propagated, the network weight parameters are updated.

[0043] The battlefield target detection method based on the fusion of visible light and infrared images provided in this embodiment can better cope with the complex ground battlefield environment, eliminate the adverse effects of various interference factors, have all-weather target detection capabilities, and can automatically align multi-spectral image features in space, perceive the temperature differences of different objects in infrared images, and further improve detection accuracy and efficiency by fusing features, thereby providing a reliable basis for command decisions in the battlefield.

[0044] Embodiment 3:

[0045] Based on Example 2, taking the size of the visible light image and the infrared image as (W, H, 3) as an example, the battlefield target detection method includes the following steps:

[0046] Step 1: feature extraction. The sizes of visible light images and infrared images are (W, H, 3). The dual-path feature extraction network extracts the features of visible light images and infrared images respectively, and obtains feature maps of different scales at three levels. The visible light image feature extraction network obtains visible light feature maps F_V1, F_V2, and F_V3, with scales of (W / 8, H / 8, C), (W / 16, H / 16, C), and (W / 32, H / 32, C), respectively. The infrared image feature extraction network obtains infrared feature maps F_T1, F_T2, and F_T3, with scales of (W / 8, H / 8, C), (W / 16, H / 16, C), and (W / 32, H / 32, C), respectively.

[0047] Step 2: feature alignment. The visible light feature map and infrared feature map of the same level obtained by the dual-path feature extraction network are processed by the feature alignment module to obtain an aligned infrared feature map based on the visible light feature map. A total of three levels of aligned infrared feature maps F_A_T1, F_A_T2 and F_A_T3 are obtained, with scales of (W / 8, H / 8, C), (W / 16, H / 16, C) and (W / 32, H / 32, C) respectively.

[0048] Step 3: Temperature difference mask prediction: After aligning the infrared feature maps F_A_T1, F_A_T2 and F_A_T3 through the temperature difference perception module, the predicted values ​​of the temperature difference mask T1, T2 and T3 are obtained, and the scales are (W / 8, H / 8, 1), (W / 16, H / 16, 1) and (W / 32, H / 32, 1) respectively.

[0049] Step 4: feature fusion; after aligning the predicted values ​​of the infrared feature map, visible light feature map and temperature difference mask, the three-level fusion feature maps FUS1, FUS2 and FUS3 are obtained after the feature fusion module, with scales of (W / 8, H / 8, 1), (W / 16, H / 16, 1) and (W / 32, H / 32, 1) respectively; the fusion process refers to Figure 4 As shown in the figure, first, the visible light feature map and the aligned infrared feature map are fused by adding corresponding pixels to obtain a preliminary fused feature map, and then the preliminary fused feature map and the temperature difference mask prediction value are feature weighted by multiplying the corresponding spatial position features to obtain the final fused feature maps FUS1, FUS2 and FUS3.

[0050] Step 5: Target category prediction and target position prediction; the three-level fusion feature maps FUS1, FUS2 and FUS3 are obtained after passing through the neck network and the task head network, and the category prediction cls_predict and the target position prediction reg_predict are post-processed to obtain the final prediction result.

[0051] It should be noted that the above detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present application belongs.

[0052] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments described in the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0053] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.

[0054] In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0055] For ease of description, spatially relative terms, such as "above", "above", "on the upper surface of", "above", etc., may be used herein to describe the spatial positional relationship between a device or feature and other devices or features as shown in the figure. It should be understood that spatially relative terms are intended to include different orientations of the device in use or operation in addition to the orientation described in the figure. For example, if the device in the accompanying drawings is inverted, the device described as "above other devices or structures" or "above other devices or structures" will be positioned as "below other devices or structures" or "below other devices or structures". Thus, the exemplary term "above" may include both "above" and "below". The device may also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and the spatially relative descriptions used herein are interpreted accordingly.

[0056] In the above detailed description, reference is made to the accompanying drawings, which form a part of this document. In the accompanying drawings, similar symbols typically identify similar components unless the context indicates otherwise. The illustrated embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be used, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein.

[0057] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A battlefield target detection method based on visible light and infrared image fusion, characterized in that: The steps include: Step 1: feature extraction: extract the features of the visible light image and obtain a visible light feature map, extract the features of the infrared image and obtain an infrared feature map through a dual-path feature extraction network; Step 2: feature alignment; The features of the visible light image and the infrared image are spatially aligned by a feature alignment module; the method is to perform channel splicing on the features of the visible light image and the infrared image to obtain a spliced ​​feature map; the spliced ​​feature map is passed through a convolution block to obtain an offset parameter offset of a feasible variable convolution, and the offset parameter offset represents the size of a feasible variable convolution kernel; the infrared feature map is passed through a feasible variable convolution to obtain an aligned infrared feature map; Step 3, prediction of temperature difference mask; The infrared feature map is convolved by the temperature difference perception module to obtain the temperature difference mask prediction value; wherein, the infrared feature map is convolved by the temperature difference perception module to obtain the temperature difference mask prediction value; during the convolution training process, the target area image and the target and background area images are extracted respectively, and the mean and variance of the regional pixels are calculated to obtain the mean u1 and variance σ1 of the target area image, the mean u2 and variance σ2 of the target and background area images, and the single Gaussian model N1(u1,σ1) of the target area and the single Gaussian model N2(u2,σ2) of the target and background area are constructed; the KL divergence D of the single Gaussian model of the target area and the single Gaussian model of the target and background area is calculated KL (N1(u1,σ1)||N2(u2,σ2)); calculate the temperature difference mask label value in the target area, and the temperature difference mask label value in the target area is calculated as follows: Among them, α is a hyperparameter, D KL ∈[0,+∞), T label ∈[0,1); Step 4: feature fusion: The visible light feature map and the infrared feature map are fused through the feature fusion module to generate a preliminary fused feature map, and then the preliminary fused feature map and the temperature difference mask prediction value are feature weighted by multiplying the corresponding spatial position features to obtain the final fused feature map; Step 5: Get the final prediction result. After the fusion feature map is processed by the neck network and the task head network, the target category prediction and target position prediction are obtained, and the final prediction result is obtained after the target bounding box regression processing.

2. The battlefield target detection method according to claim 1, characterized in that: In step 4, the calculation formula of the fusion feature map is F fus =(F T +F I )·e β*T , where F fus represents the fusion feature map, F T Represents the visible light feature map, F I represents the infrared feature map, T represents the temperature difference mask prediction value, and β is a hyperparameter used to control the degree of feature enhancement.

3. The battlefield target detection method according to claim 2, characterized in that: In step 5, the target bounding box regression processing method is as follows: provide a pair of visible light images and infrared images; use the visible light image as the main state and the infrared image as the auxiliary state, and use the annotation information of the visible light image as the label information; obtain the temperature difference mask prediction value, the category prediction cls_predict and the target space position prediction reg_predict, and calculate the total loss loss in combination with the label information. The calculation formula is loss=loss_T+loss_cls+loss_reg, where loss_T represents the temperature difference mask loss, and its calculation formula is: loss_T=(1-T label )log(1-T)+T label log(T), loss_cls represents category loss, loss_reg represents target position regression loss, loss_cls uses cross entropy loss function, loss_reg uses IOU loss function, and after the total loss loss is back-propagated, the network weight parameters are updated.

4. The battlefield target detection method according to claim 1, characterized in that: In step one, visible light image features are extracted by the visible light image feature extraction network in the dual-path feature extraction network, and visible light feature maps of different scales at three levels are obtained respectively. Infrared image features are extracted by the infrared image feature extraction network in the dual-path feature extraction network, and infrared feature maps of different scales at three levels are obtained respectively.

5. The battlefield target detection method according to claim 4, characterized in that: In step 2, the features of the visible light feature map and the infrared feature map at the same level of the dual-path feature extraction network are aligned through the feature alignment module to obtain three aligned infrared feature maps based on the visible light feature map; in step 3, the three aligned infrared feature maps are processed by the temperature difference perception module to obtain the predicted values ​​of the three temperature difference masks respectively; in step 4, the predicted values ​​of the aligned infrared feature map, the visible light feature map and the temperature difference mask are processed by the feature fusion module to obtain the preliminary fused feature map of the three levels, and the final fused feature map is obtained after combining the temperature difference mask processing.

Citation Information

Patent Citations

  • Equipment monitoring method, device and apparatus based on infrared and visible light image fusion

    CN110555819A

  • Face recognition system and method

    CN111860428A

  • Infrared image temperature estimation method and system based on full convolutional neural network

    CN113705788A