Method for generating pseudo-infrared image of confrontation type bird scene
By generating high-quality pseudo-infrared images through the Cycle-MUHGAN model, the problem of insufficient quality of bird image data in drone detection is solved, the detection accuracy and robustness are improved, and the data expansion effect and training stability of the model are enhanced.
Patent Information
- Application Number
- CN202510824366.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-19
AI Technical Summary
In existing drone detection systems, the quality of flying bird images captured by drone cameras is insufficient, resulting in poor small target detection results. In addition, it is difficult to obtain high-quality thermal infrared image datasets, which affects target detection accuracy.
A cyclic adversarial generative neural network model called Cycle-MUHGAN based on a state-space module is used to generate high-quality pseudo infrared images by training on the RGB-T dual-light dataset. This can make up for the lack of thermal infrared images in actual scenes and enhance the data augmentation effect of the model.
The generated pseudo-infrared images are of high quality and can effectively improve the accuracy and robustness of drone detection, enhance the model's ability to handle complex scenes, and improve the stability and convergence speed of the training process.
Smart Images

Figure CN120672893A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a method for generating a pseudo-infrared image of a confrontational flying bird scene. Background Art
[0002] Birds pose a significant threat to aviation, agriculture, and forestry. Addressing bird strikes near agricultural, forestry, and airports presents a significant challenge. With the advancement of technology and the maturation of the drone industry, drones are increasingly being used to repel birds, reducing the threat to aviation safety and protecting agricultural and sideline products from bird attacks. Furthermore, with the development of computer hardware and software, deep learning has become increasingly widely used in object detection, with increasing accuracy. Most deep learning-based object detection applications rely on relatively complete, high-quality datasets. Existing bird data required for drone detection primarily comes from video images captured by drone cameras. However, due to poor camera performance or inaccurate capture methods, the captured images often exhibit small pixels, high overlap, and a low model-to-height ratio. This makes drone bird detection prone to overlap and missed detections, resulting in poor detection and recognition of small bird targets. Thermal infrared images can complement the limitations of visible light images. The advantages of visible light images are that they are sensitive to lighting environments and contain rich detailed texture information. Thermal infrared images can represent the temperature (thermal radiation) information of objects through color. The two types of image information are complementary. In complex scenarios, the two images can be combined to improve the robustness of the algorithm.
[0003] However, due to the rarity of drones equipped with visible-light and thermal-infrared binocular cameras in real-world environments, capturing thermal infrared images in real-world scenarios is challenging, making it difficult to use thermal infrared images to assist algorithms in accurate target detection. Furthermore, the visible-light and thermal-infrared images captured by binocular cameras require strict temporal and spatial alignment, further complicating the acquisition of a complete, high-quality RGBT image dataset. Therefore, it is imperative to find a reliable method for generating thermal infrared images. Summary of the Invention
[0004] In response to the above-mentioned deficiencies in the prior art, the present invention provides a method for generating an adversarial pseudo-infrared image of a flying bird scene, which solves the problem of insufficient accuracy and poor visual effects of the generated pseudo-infrared images due to insufficient model performance.
[0005] In order to achieve the above objectives, the present invention adopts the following technical solution: a method for generating a pseudo infrared image of a countermeasure flying bird scene, comprising the following steps: Obtain RGB-T bi-light image data of birds in an adversarial flying bird scene and preprocess it; Divide the preprocessed bi-optical image data to obtain an RGB-T bi-optical dataset; The cyclic adversarial generative network model is trained using the RGB-T dual-light dataset, and the cyclic adversarial generative neural network model Cycle-MUHGAN based on the state-space module is obtained; The thermal infrared image generated by the visible light image is input into the trained cyclic adversarial generative neural network model Cycle-MUHGAN to generate a pseudo infrared image.
[0006] The beneficial effect of the present invention is as follows: the present invention designs a cyclic adversarial generative neural network model Cycle-MUHGAN based on a state-space module, the purpose of which is to generate high-quality pseudo infrared images to make up for the shortcomings of thermal infrared images in actual scenes and achieve the effect of data expansion.
[0007] Furthermore, the cyclic adversarial generative neural network model Cycle-MUHGAN includes: A first generator, configured to convert a visible light image into a thermal infrared image; a second generator, configured to convert the thermal infrared image into a visible light image; The first discriminator is used to determine whether the generated visible light image is close to the real visible light image.
[0008] The second discriminator is used to determine whether the generated thermal infrared image is close to the real thermal infrared image.
[0009] Furthermore, the first generator and the second generator each include a state-space-based jump-densely connected network sub-model MUHRNet, wherein the jump-densely connected network sub-model MUHRNet includes ten stages and five resolution streams, and the five resolution streams include 1 / 4 resolution, 1 / 8 resolution, 1 / 16 resolution, 1 / 32 resolution, and 1 / 64 resolution; Before the first stage, a state space module is added to obtain the feature information of the image; The first stage includes a single-branch high-resolution HR module composed of four bottleneck residual blocks; the feature map includes visible light images and thermal infrared images; The second stage, the third stage, and the fourth stage respectively include one dual-branch high-resolution HR module, five dual-branch high-resolution HR modules, and two dual-branch high-resolution HR modules, each of which includes four basic residual blocks; The fifth stage is the bottom layer of the Cycle-MUHGAN model, which includes two dual-branch high-resolution HR modules. Each dual-branch high-resolution HR module includes four basic residual blocks. Three residual blocks containing state-space modules are added after the dual-branch high-resolution HR module at the bottom layer of the Cycle-MUHGAN model. The sixth, seventh, and eighth stages each include a dual-branch high-resolution HR module, each of which includes four basic residual blocks. Before the eighth, seventh, and sixth stages, a skip fusion connection module is used to fuse the high-resolution low-level features output by the second, third, and fourth stages with the upsampled features of the seventh, sixth, and fifth stages, respectively. The ninth stage includes a single-branch high-resolution HR module, which consists of four residual blocks, where The first stage, the first half of the second stage, the first half of the eighth stage and the ninth stage are all 1 / 4 resolution streams; the second half of the second stage, the first half of the third stage, the first half of the seventh stage and the second half of the eighth stage are all 1 / 8 resolution streams; the second half of the third stage, the first half of the fourth stage, the first half of the sixth stage and the second half of the seventh stage are all 1 / 16 resolution streams; the second half of the fourth stage, the first half of the fifth stage and the second half of the sixth stage are all 1 / 32 resolution streams; the second half of the fifth stage is a 1 / 64 resolution stream.
[0010] The beneficial effect of the above further solution is that: through the above design, the present invention can maintain high-resolution features while enhancing the extraction of image context information and local details, and can increase the model's processing ability for complex structures.
[0011] Still further, the state space module includes a structured state space submodule; The expression of the structured state space submodule is as follows: in, and Respectively Moment and The system status at any moment, represents the state transition matrix, express The impact on the status express t The state variables at time , expresst Always respond to external input and system status, and Both represent the input and output mapping of the control state; The expression of the optimized attention mechanism used by the structured state space submodule is as follows: in, Represents the attention weight calculation process, Indicates the feature information being queried. represents the queried vector, Indicates the features obtained by the query, represents the state transition matrix, T Indicates transpose.
[0012] The beneficial effects of the above further solution are: the state space module enhances the model's ability to process complex data and improves the network's ability to integrate dependencies between long-distance data.
[0013] Furthermore, the downsampling part of the skip-densely connected network sub-model MUHRNet includes an attention module CBAM consisting of a channel attention sub-module and a spatial attention sub-module; The expression of the channel attention submodule is as follows: in, represents the channel attention weight, represents the input feature map, represents the Sigmoid function, represents a full connection operation, represents the global average pooling operation, Represents the global maximum pooling operation; The expression of the spatial attention submodule is as follows: in, represents the spatial attention weight, Represents a convolution operation.
[0014] The beneficial effect of the above further solution is that: through the above design, the present invention improves the feature representation ability and the interpretability of the network, and further improves the performance of the model.
[0015] Furthermore, the joint loss functions of the backbone networks of the first generator and the second generator are as follows: in, represents the joint loss function, and Represents Focal Loss loss function respectively and DiceLoss loss function The weight ratio of represents the predicted probability, represents the scaling factor, represents the direct hyperparameter, represents the cross entropy loss function, Represents pixels i The true label value of N represents the total number of samples, Represents pixels i The predicted value of Represents a constant.
[0016] The beneficial effect of the above further scheme is that the joint loss function adopted by the present invention can further promote the convergence of the model and improve the generalization ability of the model, while improving the model's understanding and adaptation to task collaboration when processing multiple tasks.
[0017] Furthermore, the iteration strategies of the first discriminator and the second discriminator are both as follows: The label value of the sample judged to be true is set to a random value in (0.9, 1.1), and the label value of the sample judged to be false is set to a random value in (0, 0.2), where the label represents the authenticity of the sample.
[0018] The number of iterations of the first discriminator and the second discriminator is changed from once per epoch to three times per epoch; an epoch represents the process in which the training data passes through the neural network once.
[0019] The beneficial effect of the above further solution is that the present invention accelerates the learning speed of the discriminator and increases the stability of the generator through the design of the above iterative strategy.
[0020] Furthermore, the loss function of the cyclic adversarial generative neural network model Cycle-MUHGAN is expressed as follows: in, Represents the loss function of the cyclic adversarial generative neural network model Cycle-MUHGAN, Represents the style mapping constraint loss function of the cyclic adversarial generative neural network model Cycle-MUHGAN, Indicates the double loop path interactive learning constraint loss function, and Represent the mapping loss function respectively and style loss function The weight ratio of k Represents a value on a feature set, j represents a feature set, represents the mapping loss of the corresponding feature weight in the neural network, Represents feature representation, represents the visible light image of the training input, represents the thermal infrared image of the training input, Indicates that the visible light image is input to the generator The thermal infrared image results generated in Indicates that the thermal infrared image is input to the generator The visible light image result generated in Indicates that The generated thermal infrared results are input to the generator The visible light image result generated in Indicates that The generated visible light result is input into the generator The thermal infrared image results generated in represents the style loss of the corresponding feature weight in the neural network, express j The weight of Represents the Gram matrix, which is used to suppress checkerboard artifacts in the image. Represents the output after Gram matrix processing, and Represent the content loss function respectively and marginal loss function The weight ratio of represents a thermal infrared image, represents a visible light image, represents a pseudo thermal infrared image, represents a pseudo visible light image, Represents the difference between the original thermal infrared image and the pseudo thermal infrared image, measured by the L1 norm loss, Represents the difference between the original visible light image and the pseudo visible light image, measured by the L1 norm loss, represents the Laplacian operator, and Represent the edge features corresponding to thermal infrared images and visible light images, respectively. and Represent the edge features corresponding to the pseudo thermal infrared image and pseudo visible light image, respectively. Indicates that Charbonnier Loss is used to measure the difference between the edge features of the original thermal infrared image and the edge features of the pseudo thermal infrared image. Indicates that Charbonnier Loss is used to measure the difference between the edge features of the original infrared image and the edge features of the pseudo thermal infrared image. Charbonnier Loss represents the loss function used in image processing tasks.
[0021] The beneficial effect of the above further solution is that: through the above design, the present invention can improve the quality of generated images and improve the stability of gradients during training, thereby reducing possible mode collapse during training. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Schematic diagram of the RGB-T bi-optical dataset in the present invention.
[0023] Figure 2 This figure is a schematic diagram of the cyclic adversarial generation process of the cyclic adversarial generation neural network model Cycle-MUHGAN of the present invention.
[0024] Figure 3 Schematic diagram of the training strategy of the cyclic adversarial generative neural network model Cycle-MUHGAN of the present invention.
[0025] Figure 4 Schematic diagram of the structure of the jump-dense link network model MUHRNet in the backbone network of generators A and B of the present invention.
[0026] Figure 5 Comparison diagram of the pseudo infrared image generated through ablation experiment and the original image.
[0027] Figure 6 This is a structural diagram of the two attention mechanism modules in the present invention.
[0028] Figure 7 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0029] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0030] Example like Figure 7 As shown, the present invention provides a method for generating a pseudo infrared image of a countermeasure flying bird scene, and its implementation method is as follows: Obtain RGB-T bi-light image data of birds in an adversarial flying bird scene and preprocess it; Divide the preprocessed bi-optical image data to obtain an RGB-T bi-optical dataset; The cyclic adversarial generative network model is trained using the RGB-T dual-light dataset, and the cyclic adversarial generative neural network model Cycle-MUHGAN based on the state-space module is obtained; The thermal infrared image generated by the visible light image is input into the trained cyclic adversarial generative neural network model Cycle-MUHGAN to generate a pseudo infrared image.
[0031] In this embodiment, a visible light image is input as a source image to be converted into a pseudo infrared image and is converted into a first generator to generate a corresponding pseudo infrared image (which has temperature distribution characteristics similar to those of a thermal infrared image).
[0032] In this embodiment, the cyclic adversarial generative neural network model Cycle-MUHGAN includes: A first generator, configured to convert a visible light image into a thermal infrared image; a second generator, configured to convert the thermal infrared image into a visible light image; The first discriminator is used to determine whether the generated visible light image is close to the real visible light image.
[0033] The second discriminator is used to determine whether the generated thermal infrared image is close to the real thermal infrared image.
[0034] In this embodiment, the first generator and the second generator both include a state-space-based jump-densely connected network sub-model MUHRNet, wherein the jump-densely connected network sub-model MUHRNet includes ten stages and five resolution streams, and the five resolution streams include 1 / 4 resolution, 1 / 8 resolution, 1 / 16 resolution, 1 / 32 resolution, and 1 / 64 resolution; Before the first stage, a state space module is added to obtain the feature information of the image; The first stage includes a single-branch high-resolution (HR) module consisting of four bottleneck residual blocks; the feature map includes visible light images and thermal infrared images; The second stage, the third stage, and the fourth stage respectively include one dual-branch high-resolution HR module, five dual-branch high-resolution HR modules, and two dual-branch high-resolution HR modules, each of which includes four basic residual blocks; The fifth stage is the bottom layer of the Cycle-MUHGAN model, which includes two dual-branch high-resolution HR modules. Each dual-branch high-resolution HR module includes four basic residual blocks. Three residual blocks containing state-space modules are added after the dual-branch high-resolution HR module at the bottom layer of the Cycle-MUHGAN model. The sixth, seventh, and eighth stages each include a dual-branch high-resolution HR module, each of which includes four basic residual blocks. Before the eighth, seventh, and sixth stages, a skip fusion connection module is used to fuse the high-resolution low-level features output by the second, third, and fourth stages with the upsampled features of the seventh, sixth, and fifth stages, respectively. The ninth stage includes a single-branch high-resolution HR module, which consists of four residual blocks, where The first stage, the first half of the second stage, the first half of the eighth stage and the ninth stage are all 1 / 4 resolution streams; the second half of the second stage, the first half of the third stage, the first half of the seventh stage and the second half of the eighth stage are all 1 / 8 resolution streams; the second half of the third stage, the first half of the fourth stage, the first half of the sixth stage and the second half of the seventh stage are all 1 / 16 resolution streams; the second half of the fourth stage, the first half of the fifth stage and the second half of the sixth stage are all 1 / 32 resolution streams; the second half of the fifth stage is a 1 / 64 resolution stream.
[0035] In this embodiment, the iteration strategy of the first discriminator and the second discriminator is as follows: The label value of the sample judged to be true is set to a random value in (0.9, 1.1), and the label value of the sample judged to be false is set to a random value in (0, 0.2), where the label represents the authenticity of the sample.
[0036] The number of iterations of the first discriminator and the second discriminator is changed from once per epoch to three times per epoch; an epoch represents the process in which the training data passes through the neural network once.
[0037] In this embodiment, the basic concept of the present invention is to use a specialized infrared pod to capture RGB-T bi-optical image data of individual birds and flocks of birds. Images with clear object boundaries and good temperature information are then selected from datasets such as VT821, VT1000, and VT5000. Image preprocessing and feature matching techniques are then used to generate an RGB-T bi-optical dataset with low parallax error and clear pixels. The corresponding recurrent adversarial generative network model is then trained based on the RGB-T bi-optical dataset. The resulting model (a recurrent adversarial generative neural network model based on a state-space module, Cycle-MUHGAN) can generate thermal infrared images from visible light images captured by a monocular camera as a supplement to the dataset, thereby assisting in improving target detection accuracy. Specifically, this includes the following: In this embodiment, a CMOS binocular infrared pod is used to capture RGB-T bird image data. The captured image data should meet the following requirements: 1. Ensure that the viewing angle difference of the captured dual-light image is relatively small to facilitate subsequent viewing angle correction and feature matching; 2. The visible light image pixels are clear and can accurately present the details in the scene; 3. The thermal infrared image can accurately represent the temperature information of the objects in the image (that is, living things have relatively low temperatures and appear orange; inanimate objects have relatively low temperatures and appear purple; the higher the temperature, the higher the color brightness or saturation of the corresponding area in the image); 3. The captured scene content should include both well-lit daytime and dark night scenes.
[0038] In this example, to ensure high-quality and information-rich bi-optical images captured by a CMOS binocular infrared pod, while also reducing noise interference and improving feature acquisition, a data preprocessing solution was employed, including image enhancement, image widening, and joint denoising. The final preprocessed dataset was then divided into training and validation sets at varying ratios (e.g., 8:2 or 7:3). The model training performance was found to be best achieved with a dataset split at a ratio of 9:1.
[0039] In this embodiment, the present invention designs a cyclic adversarial generative neural network model Cycle-MUHGAN based on a state-space module. This model improves the shortcomings of traditional cyclic adversarial generative networks in achieving pseudo-infrared image generation, optimizes the generation quality and feature retention results, and not only enhances the authenticity and structural information preservation capabilities of pseudo-infrared images, but also improves the stability and convergence speed of the network during training. The cyclic adversarial generative neural network model Cycle-MUHGAN based on the state-space module inherits the model framework of the traditional cyclic adversarial generative network: 1. The mutual conversion of images is realized through bidirectional cyclic generation (visible light-thermal infrared); 2. The essence of the work is unsupervised image-to-image conversion; 3. No paired training data is required (that is, no label file is required), only two data sets in different fields are required to achieve high-quality mapping; 4. It contains two generators: converting visible light images into thermal infrared images and converting thermal infrared images into visible light images; 5. It contains two discriminators: judging whether the generated visible light image is close to the real visible light image and judging whether the generated thermal infrared image is close to the real thermal infrared image; 6. It contains two losses: the cycle consistency loss that ensures that the image after bidirectional conversion (A→B→A or B→A→B) can be restored to the original input image as much as possible, and the adversarial loss between the generator and the discriminator through game optimization. The cyclic adversarial training flow chart of the cyclic adversarial generative neural network model Cycle-MUHGAN based on the state-space module is as follows: Figure 1 shown.
[0040] In this embodiment, in order to improve the conversion quality of visible light images to thermal infrared images, the present invention optimizes the adversarial generation network loss function, training strategy, generator backbone network and discriminator iteration strategy.
[0041] In terms of the adversarial generative network loss function, the present invention makes the following optimizations: the two losses used in the traditional cyclic adversarial generative network are replaced by style mapping constraint loss and dual-loop path interactive learning constraint loss. This improvement enhances the learning ability of the network, thereby more effectively generating high-quality pseudo infrared images.
[0042] In order to generate pseudo infrared images that can more clearly represent temperature information, the style mapping constraint loss is used (the following uses denoted) to control the cycle consistency of the cyclic adversarial generation network. It consists of two loss functions, the first one is the mapping loss (hereinafter used ), and the other is style loss (the following uses express).
[0043] is defined as follows: is defined as follows: is defined as follows: in, and Represent the mapping loss function respectively and style loss function The weight ratio can be adjusted by and The value of The numerical value of .
[0044] In order to strengthen the connection between the two branches, a dual-loop path interaction learning constraint loss is adopted (hereinafter Representation) is used to characterize the feature consistency between different domains, so that the input data can be restored to its original structure as much as possible after bidirectional path mapping. It consists of two loss functions. The first is content loss (hereinafter used ), the other is the edge loss (the following uses express).
[0045] is defined as follows: is defined as follows: is defined as follows: in, and Represent the content loss function respectively and marginal loss function The weight ratio can be adjusted by and The value of The numerical value of .
[0046] In summary, the loss function used in the cyclic adversarial generative network framework of the present invention is: in, Represents the loss function of the cyclic adversarial generative neural network model Cycle-MUHGAN, Represents the style mapping constraint loss function of the cyclic adversarial generative neural network model Cycle-MUHGAN, Indicates the double loop path interactive learning constraint loss function, and Represent the mapping loss function respectively and style loss function The weight ratio of represents the mapping loss of the corresponding feature weight in the neural network, Represents feature representation, represents the visible light image of the training input, represents the thermal infrared image of the training input, Indicates that the visible light image is input to the generator The thermal infrared image results generated in Indicates that the thermal infrared image is input to the generator The visible light image result generated in Indicates that The generated thermal infrared results are input to the generator The visible light image result generated in Indicates that The generated visible light result is input into the generator The thermal infrared image results generated in represents the style loss of the corresponding feature weight in the neural network, express j The weight of Represents the Gram matrix, which is used to suppress checkerboard artifacts in the image. represents the output after Gram matrix processing, and Represent the content loss function respectively and marginal loss function The weight ratio of represents a thermal infrared image, represents a visible light image, represents a pseudo thermal infrared image, represents a pseudo visible light image, Represents the difference between the original thermal infrared image and the pseudo thermal infrared image, measured by the L1 norm loss, Represents the difference between the original visible light image and the pseudo visible light image, measured by the L1 norm loss, represents the Laplacian operator, and Represent the edge features corresponding to the thermal infrared image and the visible light image, respectively. and Represent the edge features corresponding to the pseudo thermal infrared image and pseudo visible light image, respectively. Indicates that Charbonnier Loss is used to measure the difference between the edge features of the original thermal infrared image and the edge features of the pseudo thermal infrared image. Indicates that Charbonnier Loss is used to measure the difference between the edge features of the original infrared image and the edge features of the pseudo thermal infrared image. k Represents a value on a feature set, j represents a feature set, , the corresponding weight for Charbonnier Loss is a loss function commonly used in image processing tasks (especially in image denoising, image restoration, etc.). It improves the L1 and L2 loss functions, especially when processing noisy images, and can better cope with outliers and noise.
[0047] In this embodiment, in terms of the adversarial generative network training strategy, the training strategy of the present invention is as follows: Figure 2-Figure 3 shown.
[0048] In this embodiment, the present invention designs a state-space-based skip-dense link network model MUHRNet for the generator of the recurrent adversarial generative network. HRNet (High Resolution Net) performs well in capturing high-level semantic information and low-level spatial details, but it is prone to problems such as insufficient generalization and performance when dealing with complex scenes. Therefore, this model improves on the shortcomings of HRNet, using an encoding-decoding structure to extract image features while combining dense links with multi-resolution convolution in parallel. The specific optimization process is as follows: The high-resolution branches (i.e., the 1 / 4 and 1 / 8 resolution branches) in the original network, which consumed a lot of computational resources, were deleted to reduce computational overhead. Next, to enhance the semantic expression capability of the high-resolution output, four new stages were added after the lowest-resolution part (the lowest-resolution part refers to the fifth stage, and the four new stages are stages 6, 7, 8, and 9). By gradually upsampling the feature maps and fusing them with the feature maps from the downsampling stage, the feature representation effect is improved.
[0049] Low-level features contain rich texture and boundary information, but also carry a lot of noise and background, while high-level features have strong semantic expression capabilities and assist in target positioning and noise suppression. Therefore, it is very important to combine low-level and high-level features. However, simply concatenating the two makes it difficult to fully utilize the complementary information, and the useless information and noise in the low-level features may affect the network's pseudo-infrared image generation performance. To address this problem, a skip fusion connection module is introduced into the network to compensate for the lack of spatial detail in the high-level features. Three skip fusion connection modules are introduced before the eighth, seventh, and sixth stages, respectively, to fuse the high-resolution low-level features output by the second, third, and fourth stages with the upsampled features of the seventh, sixth, and fifth stages. The skip fusion connection uses a method similar to U-Net to enhance network connectivity and feature transfer.
[0050] like Figure 4 As shown in the figure, the main structure of the MUHRNet model consists of ten stages and five resolution streams, with resolutions of 1 / 4, 1 / 8, 1 / 16, 1 / 32, and 1 / 64. A state-space module is added before the first stage to obtain image feature information. The first stage consists of a single-branch high-resolution HR module consisting of four bottleneck residual blocks with a module width of 64. A 3×3 convolution is then used to resize the feature map width to C, forming a 1 / 4 resolution stream. Stages two through four respectively contain one dual-branch high-resolution HR module, five dual-branch high-resolution HR modules, and two dual-branch high-resolution HR modules, each branch consisting of four basic residual blocks. The fifth stage is the lowest layer of the MUHRNet model and consists of two dual-branch high-resolution HR modules, each of which is also composed of two of four basic residual blocks. Furthermore, because the lowest layer contains the richest high-dimensional semantic information, three consecutive residual modules containing state-space modules are added after the dual-branch high-resolution HR module to process this information. Next, stages 6 through 8 each contain a two-branch high-resolution HR module, with each branch consisting of four basic residual blocks. Finally, stage 9 contains a single-branch module consisting of four residual blocks. The convolution widths of the final five resolution streams are C, 2C, 4C, 8C, and 16C, respectively.
[0051] like Figure 6 As shown in the figure, in order to enhance the extraction of key features of pseudo infrared images and reduce the interference of environmental noise on the model, the CBAM attention module is introduced in the downsampling stage of the jump-dense link network model MUHRNet. The CBAM attention module consists of a channel attention submodule and a spatial attention submodule.
[0052] The channel attention module is used to identify the most important channel features for a given task in an image. During this process, the input feature map undergoes global average pooling and global max pooling, resulting in two description vectors (of size C × 1). These two description vectors are then fed into a fully connected layer, where their outputs are element-wise added. Finally, a sigmoid activation function is used to generate a channel attention weight. This attention weight is then multiplied by the original input feature channel to generate the enhanced channel features. The formula for the channel attention module is as follows: in, represents the channel attention weight, represents the input feature map, represents the Sigmoid function, represents a full connection operation, represents the global average pooling operation, Represents the global maximum pooling operation.
[0053] The spatial attention module is used to identify the most important spatial locations in an image. Unlike the channel attention module, the input feature map is only average pooled and max pooled in the channel direction, resulting in two single-channel description maps (of size H×W). These two feature maps are concatenated in the channel dimension and then processed using a convolution operation with a 3×3 convolution kernel. Finally, the spatial attention weight is obtained through the Sigmoid activation function. This spatial attention weight is multiplied by the original input feature channel to obtain the enhanced spatial features. The formula for the spatial attention module is as follows: in, represents the spatial attention weight, Represents a convolution operation.
[0054] In this embodiment, since the features of pseudo-infrared images are often spatially dispersed (such as the difference between heat source and background temperature), the state space module can capture long-range dependencies in the spatial domain, enabling the model to better simulate temperature distribution and small-scale hotspots.
[0055] The state-space module includes the structured state-space submodule, which is essentially a mathematical model for processing long-term data series. Compared to traditional neural networks, the state-space module uses state transition equations and output equations to represent changes in system state. The structured state-space submodule has two main equations, describing its state and output respectively: in, and Respectively Moment and The system status at any moment, represents the state transition matrix, express The impact on the status express t The state variables at time , express t Always respond to external input and system status, and Both represent the input and output mapping of the control state.
[0056] In addition, an optimized attention mechanism is also used in the state-space module. Unlike Transformr's self-attention mechanism, the optimized attention mechanism focuses more on structured masking strategies (i.e., allowing dependencies between some time steps to take effect), which improves the computational efficiency of the model. The formula for the structured state-space submodule is as follows: in, Represents the attention weight calculation process, Indicates the feature information being queried. represents the queried vector, Indicates the features obtained by the query, represents the state transition matrix, T Indicates transpose.
[0057] In this embodiment, in order to make the generated pseudo image higher in quality, the loss function used by the backbone network of the generator is optimized to a joint loss function (hereinafter referred to as In which, the joint loss function is composed of the Focal Loss loss function (hereinafter used Represented) and Dice Loss function (used below Indicates) composition.
[0058] is defined as follows: is defined as follows: in, represents the joint loss function, and Represents Focal Loss loss function respectively and DiceLoss loss function The weight ratio of represents the predicted probability, represents the scaling factor, represents the direct hyperparameter, represents the cross entropy loss function, Represents pixels i The true label value of N represents the total number of samples, Represents pixels i The predicted value of represents a constant, and Represents the weight ratio of the two losses respectively, which can be adjusted and The value can be used to change its proportion to indirectly adjust The numerical value of .
[0059] In this embodiment, during the training of the recurrent adversarial generative network, true values are generally set to 1 and false values to 0. When the discriminator determines that a sample is true, it outputs Label = 1; when the discriminator determines that a sample is false, it outputs Label = 0. To ensure smoother training, at the code level, the label value for samples judged as true is set to a random value in the range (0.9, 1.1), and the label value for samples judged as false is set to a random value in the range (0, 0.2).
[0060] In this example, the ratio of the number of training iterations for the discriminator to the number of training iterations for the generator in a traditional recurrent adversarial generative network is 1:1, meaning that both the discriminator and the generator are trained only once per epoch. However, training the discriminator multiple times can stimulate the training of the backbone network in the generator, thereby producing better renderings. Therefore, the number of training iterations for both discriminators was changed from once per epoch to three per epoch.
[0061] In this embodiment, three evaluation indicators, namely peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and visual information fidelity (VIF), are used for evaluation. The formulas are as follows: Here, MAX represents the maximum possible value of an image pixel; MSE represents the mean square error.
[0062] in, and Represents the image blocks x 、 y The mean of and Represents the image blocks x 、 yThe variance of Represents an image block x and y The covariance between and They all represent constants to avoid the situation where the denominator is 0.
[0063] in, represents the information in the generated image, Represents the information in the original image.
[0064] Use the above three evaluation indicators to verify the performance of the adversarial generative network on the validation set and save the best model.
[0065] like Figure 5 As shown, Figure 5 The pseudo infrared image generated by the ablation experiment is compared with the original image. Table 1 is the pseudo infrared image index table of the ablation experiment.
[0066] Table 1 The beneficial effects of the present invention are: 1) according to actual application requirements, the loss function of the adversarial generative network is optimized to improve the quality and stability of the generation effect; 2) in order to improve training efficiency and model convergence speed, the training strategy of the adversarial generative network is optimized to reduce the risk of model instability; 3) the generator backbone network of the adversarial generative network is the key model for generating pseudo-infrared images to more accurately capture data distribution; 4) in order to optimize the training smoothing effect, the discriminator iteration strategy is improved to ensure smooth and stable generation results.
Claims
1. A method for generating pseudo infrared images of adversarial flying bird scenes, characterized in that: The following steps are involved: Obtain RGB-T bi-light image data of birds in an adversarial flying bird scene and preprocess it; Divide the preprocessed bi-optical image data to obtain an RGB-T bi-optical dataset; The cyclic adversarial generative network model is trained using the RGB-T dual-light dataset, and the cyclic adversarial generative neural network model Cycle-MUHGAN based on the state-space module is obtained; The thermal infrared image generated by the visible light image is input into the trained cyclic adversarial generative neural network model Cycle-MUHGAN to generate a pseudo infrared image.
2. The method for generating pseudo infrared images of a countermeasure flying bird scene according to claim 1, characterized in that: The cyclic adversarial generative neural network model Cycle-MUHGAN includes: A first generator, configured to convert a visible light image into a thermal infrared image; a second generator, configured to convert the thermal infrared image into a visible light image; The first discriminator is used to determine whether the generated visible light image is close to the real visible light image. The second discriminator is used to determine whether the generated thermal infrared image is close to the real thermal infrared image.
3. The method for generating pseudo infrared images of a countermeasure flying bird scene according to claim 2, characterized in that: The first generator and the second generator both include a state-space-based jump-densely connected network sub-model MUHRNet, wherein the jump-densely connected network sub-model MUHRNet includes ten stages and five resolution streams, and the five resolution streams include 1 / 4 resolution, 1 / 8 resolution, 1 / 16 resolution, 1 / 32 resolution, and 1 / 64 resolution; Before the first stage, a state space module is added to obtain the feature information of the image; The first stage includes a single-branch high-resolution HR module composed of four bottleneck residual blocks; the feature map includes visible light images and thermal infrared images; The second stage, the third stage, and the fourth stage respectively include one dual-branch high-resolution HR module, five dual-branch high-resolution HR modules, and two dual-branch high-resolution HR modules, each of which includes four basic residual blocks; The fifth stage is the bottom layer of the Cycle-MUHGAN model, which includes two dual-branch high-resolution HR modules. Each dual-branch high-resolution HR module includes four basic residual blocks. Three residual blocks containing state-space modules are added after the dual-branch high-resolution HR module at the bottom layer of the Cycle-MUHGAN model. The sixth, seventh, and eighth stages each include a dual-branch high-resolution HR module, each of which includes four basic residual blocks. Before the eighth, seventh, and sixth stages, a skip fusion connection module is used to fuse the high-resolution low-level features output by the second, third, and fourth stages with the upsampled features of the seventh, sixth, and fifth stages, respectively. The ninth stage includes a single-branch high-resolution HR module, which consists of four residual blocks, where The first stage, the first half of the second stage, the first half of the eighth stage and the ninth stage are all 1 / 4 resolution streams; the second half of the second stage, the first half of the third stage, the first half of the seventh stage and the second half of the eighth stage are all 1 / 8 resolution streams; the second half of the third stage, the first half of the fourth stage, the first half of the sixth stage and the second half of the seventh stage are all 1 / 16 resolution streams; the second half of the fourth stage, the first half of the fifth stage and the second half of the sixth stage are all 1 / 32 resolution streams; the second half of the fifth stage is a 1 / 64 resolution stream.
4. The method for generating pseudo infrared images of a countermeasure flying bird scene according to claim 3, characterized in that: The state space module includes a structured state space submodule; The expression of the structured state space submodule is as follows: in, and Respectively Moment and The system status at the moment, represents the state transition matrix, express The impact on the status express t The state variables at time , express t Always respond to external input and system status, and Both represent the input and output mapping of the control state; The expression of the optimized attention mechanism used by the structured state space submodule is as follows: in, Represents the attention weight calculation process, Indicates the feature information being queried. represents the queried vector, Indicates the features obtained by the query, represents the state transition matrix, T Indicates transpose.
5. The method for generating pseudo infrared images of a countermeasure flying bird scene according to claim 3, characterized in that: The downsampling part of the jump-densely connected network sub-model MUHRNet includes an attention module CBAM consisting of a channel attention sub-module and a spatial attention sub-module; The expression of the channel attention submodule is as follows: in, represents the channel attention weight, represents the input feature map, represents the Sigmoid function, represents a full connection operation, represents the global average pooling operation, Represents the global maximum pooling operation; The expression of the spatial attention submodule is as follows: in, represents the spatial attention weight, Represents a convolution operation.
6. The method for generating pseudo infrared images of a countermeasure flying bird scene according to claim 3, characterized in that: The joint loss functions of the backbone networks of the first generator and the second generator are as follows: in, represents the joint loss function, and Represents Focal Loss loss function respectively and DiceLoss loss function The weight ratio of represents the predicted probability, represents the scaling factor, represents the direct hyperparameter, represents the cross entropy loss function, Represents pixels i The true label value of N represents the total number of samples, Represents pixels i The predicted value of Represents a constant.
7. The method for generating pseudo infrared images of a countermeasure flying bird scene according to claim 2, characterized in that: The iterative strategies of the first discriminator and the second discriminator are both as follows: The label value of the sample judged to be true is set to a random value in (0.9, 1.1), and the label value of the sample judged to be false is set to a random value in (0, 0.2), where the label represents the authenticity of the sample. The number of iterations of the first discriminator and the second discriminator is changed from once per epoch to three times per epoch; an epoch represents the process in which the training data passes through the neural network once.
8. The method for generating pseudo infrared images of a countermeasure flying bird scene according to claim 2, characterized in that: The loss function of the cyclic adversarial generative neural network model Cycle-MUHGAN is expressed as follows: in, Represents the loss function of the cyclic adversarial generative neural network model Cycle-MUHGAN, Represents the style mapping constraint loss function of the cyclic adversarial generative neural network model Cycle-MUHGAN, Indicates the double loop path interactive learning constraint loss function, and Represent the mapping loss function respectively and style loss function The weight ratio of k Represents a value on a feature set, j represents a feature set, represents the mapping loss of the corresponding feature weight in the neural network, Represents feature representation, represents the visible light image of the training input, represents the thermal infrared image of the training input, Indicates that the visible light image is input to the generator The thermal infrared image results generated in Indicates that the thermal infrared image is input to the generator The visible light image result generated in Indicates that The generated thermal infrared results are input to the generator The visible light image result generated in Indicates that The generated visible light result is input into the generator The thermal infrared image results generated in represents the style loss of the corresponding feature weight in the neural network, express j The weight of Represents the Gram matrix, which is used to suppress checkerboard artifacts in the image. represents the output after Gram matrix processing, and Represent the content loss function respectively and marginal loss function The weight ratio of represents a thermal infrared image, represents a visible light image, represents a pseudo thermal infrared image, represents a pseudo visible light image, Represents the difference between the original thermal infrared image and the pseudo thermal infrared image, measured by the L1 norm loss, Represents the difference between the original visible light image and the pseudo visible light image, measured by the L1 norm loss, represents the Laplacian operator, and Represent the edge features corresponding to thermal infrared images and visible light images, respectively. and Represent the edge features corresponding to the pseudo thermal infrared image and pseudo visible light image, respectively. Indicates that Charbonnier Loss is used to measure the difference between the edge features of the original thermal infrared image and the edge features of the pseudo thermal infrared image. Indicates that Charbonnier Loss is used to measure the difference between the edge features of the original infrared image and the edge features of the pseudo thermal infrared image. Charbonnier Loss represents the loss function used in image processing tasks.