Heterogeneous image searching method and device based on generative adversarial network and storage medium

Through the HR-GAN network based on the generative adversarial network, the visible light image is generated to generate infrared simulated images, which solves the problem of poor quality of visible light image in surface unexploded munitions detection in the whole day period, and realizes efficient heterologous image detection, meeting the requirement of detection accuracy higher than 85%.

CN120164196AActive Publication Date: 2025-06-17CHINESE PEOPLES LIBERATION ARMY ARMY ARTILLERY & AIR DEFENSE ACAD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510219398.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

When effectively detecting unexploded surface ammunition during the whole day, it is difficult for the prior art to effectively use visible light images and infrared images for heterologous image detection, especially when it is detected at night after day shooting, the visible light images are of poor quality and it is difficult to distinguish the target.

Method used

Using a method based on a generative adversarial network, the visible light image is generated by training the HR-GAN network, and the properties and quantity of unexploded surface ammunition are detected through comparative analysis. The HR-GAN network improves image generation quality by introducing high-resolution networks HRnet and Restomer modules, and alleviates the problems of unclear textures and missing structures.

Benefits of technology

The heterologous image detection is realized throughout the day, and the generated infrared simulation images are of high quality, with an average target detection rate of 88.89%, which can basically meet the requirements of detection accuracy higher than 85%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164196A_ABST
    Figure CN120164196A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogenous image searching method and device based on a generative adversarial network and a storage medium, and the method comprises the following steps: inputting a visible light image of a to-be-detected target area into a trained HR-GAN network, generating an infrared simulation image through the visible light image, and carrying out the recognition of the infrared simulation image through the HR-GAN network; comparing and analyzing the infrared image of the to-be-detected target area with the generated infrared simulation image so as to detect the property and quantity of the unexploded ammunition on the earth surface; the HR-GAN training step comprises the steps of introducing a high-resolution network HRnet into a Pix2pix architecture as a generative network, keeping high-resolution representation of an image in the network, and relieving the problems that the texture of the generated infrared image is not clear and the noise is relatively high; a Restomer module is introduced into the generative network, so that multi-scale local-global representation learning of the image is promoted, and the problem of structure deficiency of the generated infrared image is relieved. Target detection experiments show that the average target detection rate reaches 88.89%, and the requirement that the detection precision is higher than 85% can be basically met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of heterologous image detection, and particularly to a method, device and storage medium for searching heterologous images based on a generative adversarial network. Background Art

[0002] Unexploded ordnance on the ground generally includes landmines, grenades, shells, bombs and bullets of various calibers. For unexploded ordnance on the ground, ground equipment is generally used for flat-pushing detection or low-altitude detection by an unmanned aerial vehicle platform, and the detection process is time-consuming and laborious. Especially when live ammunition shooting is carried out during the day and unexploded ordnance is searched at night, that is, visible light images of the projectile impact area are collected before the shooting starts, but it is already dark after the shooting ends, and it is necessary to detect and eliminate the unexploded ordnance in the projectile impact area in time. At this time, the quality of the collected visible light images is not good and it is difficult to effectively distinguish, while the quality of the collected short-wave infrared images is better. How to effectively use visible light images and infrared images for heterologous image detection is an important way to solve all-day unexploded ordnance detection.

[0003] Heterologous images usually refer to images of different bands collected by different types of imaging detection devices. Due to the different imaging mechanisms of visible light images and infrared images, there are significant differences in their features. If the collected visible light images and infrared images are directly cross-matched, it is difficult to effectively detect the target information. At this time, if the visible light images are converted into infrared images, or the infrared images are converted into visible light images, so that the heterologous image fusion becomes homologous image matching, it is easier to achieve, and the method for detecting the difference points of homologous images is already mature. Summary of the Invention

[0004] A method, device and storage medium for searching heterologous images based on a generative adversarial network proposed by the present invention can at least solve one of the technical problems in the background art.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A method for searching heterologous images based on a generative adversarial network includes inputting a visible light image of a target area to be detected into a trained HR-GAN network, generating an infrared simulation image from the visible light image, and then comparing and analyzing the infrared image of the target area to be detected with the generated infrared simulation image, so as to detect the nature and quantity of unexploded ordnance on the ground;

[0007] Among them, the training steps of the HR-GAN network include introducing a high-resolution network HRnet as a generative network in the Pix2pix architecture to maintain the high-resolution representation of the image in the network and alleviate the problems of unclear texture and high noise in the generated infrared image;

[0008] Introduce the Restomer module into the generation network to promote multi-scale local-global representation learning of images and alleviate the problem of missing structures in the generated infrared images.

[0009] Furthermore, introduce the ECA model based on HRnet to form an improved ECA-HRnet network;

[0010] Among them, the input image size of ECA-HRnet is 256×192, starting with a stem composed of two 3×3 convolutions with a stride of 1; after passing through the stem, the resolution is reduced to 1 / 4 of the input image feature map resolution;

[0011] Its backbone network consists of four steps; in the first step, the network is composed of 4 residual units, and each residual unit is composed of ECA-Bottleneck, that is, the ECA module is embedded into the Bottleneck, and then a 3×3 convolutional kernel is used to reduce the width of the feature map resolution;

[0012] In the second step, a multi-resolution module is repeated respectively;

[0013] In the third step, four multi-resolution modules are repeated respectively;

[0014] In the fourth step, three multi-resolution modules are repeated respectively;

[0015] Among them, each multi-resolution module contains two parts, one is parallel multi-resolution convolution, and the other is multi-resolution fusion;

[0016] Branches with different resolutions are connected in parallel. Each branch contains four residual units, and each unit is composed of ECA-BasicBlock, that is, the ECA module is embedded into the BasicBlock; feature fusion is performed on branches with different resolutions to complete information exchange. The process from low resolution to high resolution is mainly achieved through bilinear upsampling, and the process from high resolution to low resolution uses one or more strided convolutions, that is, a convolutional layer with a 3×3 convolutional kernel and a stride of 2.

[0017] Furthermore, to obtain different local channel information, the mapping relationship between the channel C and the convolutional kernel size K is expressed as:

[0018] C = Φ(K) = 2 (γ×K-b) (1)

[0019] Therefore, the value of K will change adaptively with the number of channels, expressed as:

[0020]

[0021] Wherein, K is the size of the convolution kernel, representing the interaction between the current channel features and the information of K channels, C represents the number of channels, and b and γ are fixed values, which are set to 1 and 2 respectively.

[0022] Furthermore, the Restomer module is introduced into the generation network HRnet; first, the dimension is increased using 1×1 convolution, then the features are divided into three blocks using 3×3 grouped convolution, and finally, classical self-attention calculation is performed.

[0023] Furthermore, the loss function of the HR-GAN network is:

[0024]

[0025] Wherein, Drawing on the loss function of the conditional generative adversarial network, D represents the discriminator, G represents the generator, x is the input image, y is the corresponding real image, and z is a random variable.

[0026] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.

[0027] On yet another hand, the present invention also discloses a computer device including a memory and a processor, the memory storing a computer program, which when executed by the processor causes the processor to execute the steps of the above method.

[0028] As can be seen from the above technical solutions, the method for searching for heterologous images based on a generative adversarial network of the present invention optimizes and improves the existing generative adversarial network model in order to ensure full-time and all-weather uninterrupted detection and effectively solve the problem of detecting heterologous images of unexploded ordnance on the ground, thereby proposing the HR-GAN network model. The indicators of the infrared simulation images generated by this model are improved compared to other networks. Through target detection experiments, it shows that by using the method for detecting difference points of heterologous images of unexploded ordnance on the ground based on HR-GAN, the average target detection rate reaches 88.89%, which can basically meet the requirement of a detection accuracy higher than 85%. However, since noise is easily introduced during the process of generating infrared simulation images and some small target image information will be lost, problems such as missed detection and false detection occur when using infrared simulation images and actually collected infrared images for small target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flowchart of image conversion;

[0030] Figure 2 is a schematic diagram of the CycleGAN network;

[0031] Figure 3It is the flow chart of the InfraGAN algorithm;

[0032] Figure 4 It is the schematic diagram of the HR-GAN network architecture in the embodiment of the present invention;

[0033] Figure 5 It is the schematic diagram of the HRnet network architecture in the embodiment of the present invention;

[0034] Figure 6 It is the schematic diagram of the ECA-HRnet network architecture in the embodiment of the present invention;

[0035] Figure 7 It is the schematic diagram of the ECA network architecture in the embodiment of the present invention;

[0036] Figure 8 It is the schematic diagram of the ECA-BasicBlock network architecture in the embodiment of the present invention;

[0037] Figure 9 It is the schematic diagram of the ECA-Bottleneck network architecture in the embodiment of the present invention;

[0038] Figure 10 It is the schematic diagram of the Restomer network architecture in the embodiment of the present invention;

[0039] Figure 11 It is the curve graph of the change of the loss function in the embodiment of the present invention;

[0040] Figure 12 It is the image conversion result graph of four algorithms in the embodiment of the present invention;

[0041] Figure 13 It is the infrared image feature point extraction graph in the embodiment of the present invention;

[0042] Figure 14 It is the infrared image feature matching graph in the embodiment of the present invention;

[0043] Figure 15 It is the differential point detection result graph in the embodiment of the present invention;

[0044] Figure 16 It is the detection result graph of the heterogeneous images of unexploded ordnance on the ground in the embodiment of the present invention. Detailed implementation manners

[0045] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention.

[0046] Generative Adversarial Nets (GAN) is a deep learning model mainly composed of two neural networks that play against each other, namely the generator and the discriminator. The generator selects some random noise from a distribution and attempts to generate a distribution similar to the output from it, aiming to generate a forged distribution that is exactly similar to the real distribution. That is to say, the forged output images should look like real images. The discriminator is responsible for receiving real data or data samples generated by the generator and evaluating them. At the same time, the discriminator outputs a probability value. When this value is close to 1, it indicates that the data is real, and when it is close to 0, it indicates that the data is generated, which is used to judge the authenticity of the input data. The discriminator not only evaluates the authenticity of the generated data but also provides feedback to the generator to promote the generator to improve the quality of the generated data. The image conversion process is as Figure 1 shown.

[0047] Generative Adversarial Nets (GAN) was proposed by Goodfellow et al. in 2014. By comparison, it is found that the image quality generated by this method is better than that generated by other methods. Since then, many scholars have continued to explore along the path of generative adversarial in the field of image generation. At first, the generated images were all random, and it was difficult to control the image quality. To solve such problems, Mirza et al. proposed the CGAN network in 2015. This network uses label information in the evaluation process of the discriminator. After multiple rounds of the game between generation and discrimination, the images finally generated by the generator are directional, realizing the control of the content of the generated images for the first time. Subsequently, based on in-depth research on the CGAN network, Phillip et al. replaced the random noise input to the GAN network with images and used it to solve the image conversion problem, and proposed the Pix2pixGAN network in 2017, with good image conversion effects. However, this network also has limitations and is mainly applied to paired datasets. In the same year, to solve the image conversion problem of unpaired datasets, Park et al. proposed the CycleGAN network with cycle consistency, and its principle is as Figure 2 shown. In this network, the source domain image and the target domain image can be mutually converted, and the image obtained after one round of conversion is the same as the source domain image. In this process, two generators need to be constructed, one for converting from the source domain to the target domain, and the other for converting from the target domain to the source domain. By using the two generators for mutual conversion, the image conversion problem of unpaired datasets is effectively solved. Later, to solve the mode collapse problem, based on the previous research, the InfraGAN network was proposed by introducing the infra-network component. The generator and the discriminator of this network are trained together, and the input random noise is mapped into a latent space, and then the generator generates images from this latent space to improve the quality of the generated images. Its process is as Figure 3 shown.

[0048] HR-GAN algorithm

[0049] Aiming at the problems that the infrared images generated by the current infrared image generation algorithms are somewhat unclear in texture, lacking in structure, and having high noise, the embodiment of the present invention proposes an HR-GAN infrared image generation algorithm by improving on the Pix2pix model. Its overall architecture is as Figure 4 shown, to obtain the corresponding infrared equivalent image for a given input visible light image. The main contents are as follows:

[0050] (1) Introduce a high-resolution network (HRnet) as the generation network in the Pix2pix architecture to maintain the high-resolution representation of the image in the network and alleviate the problems of unclear texture and high noise in the generated infrared images;

[0051] (2) Introduce a Restomer module in the generation network to promote the learning of multi-scale local-global representations of the image and alleviate the problem of lacking structure in the generated infrared images.

[0052] Improvement of the HRnet network

[0053] Since the feature extraction network will lose the resolution of the image during the sampling process, to solve the problem of image resolution loss, a research team from the University of Science and Technology of China and Microsoft Research collaborated to propose the HRnet network, which can maintain the high resolution of the image throughout the learning process. Its model architecture is as Figure 5 shown.

[0054] Channel attention can improve the performance of convolutional neural networks to a certain extent. However, the designs of many current attention mechanisms are relatively complex in terms of the number of parameters and structure, etc., which increases the complexity of the model to a certain extent and makes the complex human pose detection network model even more complex. To balance the contradiction between complexity and model performance, the Efficient Channel Attention (ECA) model takes into account both the complexity and performance of the network and can perform local cross-channel interactions between channels without reducing the dimension. This efficient channel attention can effectively reduce the complexity of the model while ensuring that the network performance does not decrease. Therefore, this paper introduces the ECA model on the basis of HRnet to form an improved ECA-HRnet network. Its network architecture is as Figure 6As shown in the figure. The input image size of ECA-HRnet is 256×192, starting with a stem composed of two 3×3 convolutions with a stride of 1. After passing through the stem, the resolution is reduced to 1 / 4 of the input image feature map resolution. Its backbone network mainly consists of four steps. In the first step, the network is composed of 4 residual units, each residual unit is composed of ECA-Bottleneck (embedding the ECA module into the Bottleneck), and then a 3×3 convolution kernel is used to reduce the width of the feature map resolution. In the second (third, fourth) step, one (four, three) multi-resolution modules are repeated respectively. Among them, each multi-resolution module contains two parts, one is parallel multi-resolution convolution, and the other is multi-resolution fusion. Branches with different resolutions are connected in parallel. Each branch contains four residual units, and each unit is composed of ECA-BasicBlock (embedding the ECA module into the BasicBlock); feature fusion is performed on branches with different resolutions to complete information exchange. The process from low resolution to high resolution is mainly achieved through bilinear upsampling, and the process from high resolution to low resolution mainly uses one or more strided convolutions (convolution layer with a 3×3 convolution kernel and a stride of 2).

[0055] To obtain different local channel information, the mapping relationship between channel C and convolution kernel size K can be expressed as:

[0056] C = Φ(K) = 2 (γ×K-b) (1)

[0057] Therefore, the value of K will change adaptively with the number of channels and can be expressed as:

[0058]

[0059] In the formula, K is the convolution kernel size, representing the interaction between the current channel feature and the information of K channels. C represents the number of channels, and b and γ are fixed values, set to 1 and 2 respectively.

[0060] Assume that the input of the feature map is C×H×W. After passing through a GAP layer, the size of the feature map becomes C×1×1. Since the HRnet network is relatively deep and there are many intermediate layers, these intermediate layers determine the feature extraction and fusion capabilities of the network, and small convolution kernels can extract more features. In addition, in human pose detection, the connectivity of human joints is related to the nodes of the front and rear joints and has nothing to do with other joint nodes. Therefore, the convolution kernel K of the one-dimensional convolution is set to 3. The Sigmoid function is used to generate the corresponding channel weights, representing the importance of each channel feature, and the input features are multiplied and weighted to complete the re-calibration of the features. The structure of the ECA module is as Figure 7 shown.

[0061] Since the ECA module has a relatively simple structure and can be directly embedded into the existing network framework, the deep convolutional neural network embedded with the ECA module can be called ECA-Net. Bottleneck and BasicBlock are classic convolutional units commonly used in Resnet. ECA-BasicBlock is obtained by embedding the ECA module structure into the BasicBlock convolutional unit, and its structure is as shown in Figure 8 shown. ECA-Bottleneck is formed by embedding the ECA module structure into the Bottleneck structure, and its network architecture is as shown in Figure 9 shown.

[0062] Lightweight Transformer-Restomer

[0063] To obtain clear infrared images in real time, the Restomer module is introduced into the generation network HRnet. Since convolutional neural networks (CNNs) perform well in learning generalizable image priors from large-scale data, these models have been widely applied to image restoration tasks. It has been found that the Transformer model

[17] has good performance in speech and visual image processing. This model improves the disadvantages of the limited receptive field and inadaptability to input content of convolutional neural networks. At the same time, it increases the computational complexity, and the computational complexity shows a quadratic growth trend with the increase of spatial resolution. Therefore, it is difficult to use in high-resolution image processing. However, by adding an embedding design to the multi-head attention and feed-forward network, it can capture multiple long-range pixel interactions simultaneously and is applicable to large image processing. Different from the general Transformer model, Restomer is not the common patch-wise when performing token calculation in the self-attention model, but pixel-wise. First, use a 1×1 convolution to increase the dimension, then use a 3×3 grouped convolution to divide the features into three blocks, and finally perform classical self-attention calculation. The architecture of the Restomer module is as shown in Figure 10 shown.

[0064] Loss function

[0065] The loss functions of the generator and discriminator are the same as those of GAN. The purpose of the discriminator is to detect the simulated images generated by the generator with the highest probability. The goal of the generator is to generate images to mislead the discriminator so that the discriminator cannot correctly distinguish the true from the false. The loss function of the HR-GAN network is:

[0066]

[0067] In the formula, Drawing on the loss function of the conditional generative adversarial network, where D represents the discriminator, G represents the generator, x is the input image, y is the corresponding real image, and z is a random variable. For the discriminator, the maximum value of D(x,y) is taken, that is, the probability of recognizing the real image is the largest, while the minimum value of D(x,G(x,z)) is taken, that is, the probability of judging the generated simulated image as real when generating the simulated image is the smallest, that is, the maximum value of 1 - D(x,G(x,z)); for the generator ζ cGAN in terms of, the first term is a constant, and it is sufficient to minimize the second term. λζ L1 (G) is for the generator. Since the generator can not only mislead the discriminator but also be as realistic as possible and consistent with the content of the input image, an L1 loss function is added.

[0068] Algorithm performance test

[0069] Experimental purpose

[0070] To test the quality of the infrared simulation images generated by the HR-GAN network, comparative experiments are conducted using four network models, namely CycleGAN, Pix2pix, InfraGAN, and HR-GAN. By comparing and analyzing the quality of the infrared images generated by the four network models, it is shown that the HR-GAN network has significant advantages in converting visible light images into infrared images, laying a solid foundation for the next step of detecting the differences in heterogeneous images of surface unexploded ordnance.

[0071] Experimental steps

[0072] (1) Install the hardware platform

[0073] The experiment uses hardware such as an Intel Core i7-11700K@3.60GHz CPU and an NVIDIA RTX 3080 GPU. The experimental environment is python3.7, and the Pytorch learning framework is adopted.

[0074] (2) Train the network model

[0075] When training the network, first, the Adam algorithm is used to optimize the training network, where the two momentum parameters are preset to 0.5 and 0.999 respectively. Secondly, to ensure that the network can converge smoothly, the number of model training times is preset to 200 times. During the entire training process, for the first 100 times, the learning rate of the generative network is preset to 0.0002, and the learning rate of the adversarial network is preset to 0.000002. In the latter 100 times, the learning rates of the two networks both show a linear downward trend and finally both drop to 0. The change of its loss function is as Figure 11 shown.

[0076] (3) Calculate the evaluation index

[0077] Six objective evaluation metrics are adopted to evaluate the quality of the generated images, including Mean Squared Error (MSE), Mean Absolute Error (MAE), Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), Multi-scale Structural Similarity Index Measure (MS-SSIM), and Learned Perceptual Image Patch Similarity (LPIPS).

[0078] Experimental results

[0079] Experiments were conducted on converting visible light images of surface unexploded ordnance into infrared simulation images. The CycleGAN network, Pix2pix network, InfraGAN network, and the HR-GAN network proposed in this chapter were used for generation respectively, and infrared images of surface unexploded ordnance generated by 4 network models were obtained respectively (as Figure 12 shown), and the obtained image information was calculated, and the evaluation metrics are shown in Table 1.

[0080] Table 1 Comparison of infrared image simulation algorithm performance

[0081]

[0082] By comparing the results, it can be found that compared with the image generation method CycleGAN without registration, the image generation methods Pix2Pix, InfraGAN, and HR-GAN with registered images have achieved better results on the dataset. Among them, in terms of the peak signal-to-noise ratio (PSNR) index, the HR-GAN network has increased by at least 6.66% compared to other networks; in terms of the mean squared error (MSE) index, the HR-GAN network has decreased by at least 12.0%; in terms of the mean absolute error (MAE) index, the HR-GAN network also has good performance and is lower than other networks; in terms of the structural similarity (SSIM) index, the HR-GAN network has increased by at least 11.48% and reached a structural similarity of 95.78%; in terms of the multi-scale structural similarity (MS-SSIM) index, the HR-GAN network is still the best and reaches a multi-scale structural similarity of 96.9%; while in terms of the learned perceptual image patch similarity (LPIPS) index, the HR-GAN network has decreased by at least 57.03%. From the calculation results of the above six evaluation indicators, the quality of the infrared simulation images generated by the HR-GAN network is better than that of the CycleGAN network, Pix2pix network, and InfraGAN network, indicating that the infrared simulation images generated by the HR-GAN network are more similar to the original infrared images, and the generated infrared images have clearer structural information and texture information.

[0083] Detection of Difference Points in Heterogeneous Images Based on the HR-GAN Algorithm

[0084] Experimental Purpose

[0085] Through multiple experiments on detecting difference points in heterogeneous visible light images and infrared images, calculate the average detection rate to verify that the infrared simulation images generated by the generative adversarial network can be effectively applied to the detection of surface unexploded ordnance. Input the visible light image of the target area to be measured into the trained HR-GAN network, generate an infrared simulation image from the visible light image, and then use the infrared image of the target area to be measured and the generated infrared simulation image for comparative analysis to detect the nature and quantity of surface unexploded ordnance.

[0086] Experimental Steps

[0087] For the detection of difference points in heterogeneous images, first use the trained HR-GAN network to convert the previously collected visible light image (without unexploded ordnance) into an infrared simulation image, and then extract feature points from the infrared simulation image and the subsequently collected infrared image of the target area (with unexploded ordnance) respectively. The extraction results are as Figure 13 shown; then use the SURF algorithm to perform feature cross-matching on the infrared simulation image and the infrared image. The image matching results are as Figure 14 shown; after processing the pixel gray difference of the two images, a binary image can be obtained, and finally, the target information can be detected through threshold processing, as Figure 15 shown.

[0088] Experimental results

[0089] Through 45 heterologous image detection experiments, surface unexploded ordnance with a parachute drop structure was used for detection during the experiments. By statistically analyzing the ratio of the number of detected targets to the actual preset number of targets, the target detection results are as follows Figure 16 shown, and the final average target detection rate is 88.89%.

[0090] It was found through experiments that although the HR-GAN network can generate infrared simulation images of surface unexploded ordnance, and when comparing the generated infrared simulation images with the original images, the structural similarity reaches 95.78%, due to the relatively small pixel proportion of surface unexploded ordnance images under high-altitude imaging conditions, noise may be generated in the generated infrared simulation images. When cross-matching with the actual obtained infrared images with surface unexploded ordnance, the pixel points of surface unexploded ordnance may be offset, so the targets cannot be detected, resulting in an average target detection rate of 88.89%. At the same time, the generated infrared simulation images may lack some image features. After pixel difference processing of the original infrared images and the infrared simulation images, noise will appear on the difference map, so there will be false alarm misjudgment situations.

[0091] Conclusion

[0092] To ensure full-time and all-weather uninterrupted detection and effectively solve the problem of heterologous image detection of surface unexploded ordnance, the present invention optimizes and improves the existing generative adversarial network model, and thus proposes the HR-GAN network model. The indicators of the infrared simulation images generated by this model are improved compared with other networks. The target detection experiment shows that by using the method for detecting the difference points of heterologous images of surface unexploded ordnance based on HR-GAN, the average target detection rate reaches 88.89%, which can basically meet the requirement that the detection accuracy is higher than 85%. However, since noise is easily introduced during the process of generating infrared simulation images and some small target image information will be lost, problems such as missed detection and false detection occur when using the infrared simulation images and the actually collected infrared images for small target detection.

[0093] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.

[0094] On yet another aspect, the present invention also discloses a computer device, including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above method.

[0095] In another embodiment provided by the present application, a computer program product including instructions is further provided. When it runs on a computer, it causes the computer to execute any one of the above-mentioned cross-source image search methods based on a generative adversarial network.

[0096] It can be understood that the systems, devices, and storage media provided by the embodiments of the present invention correspond to the methods provided by the embodiments of the present invention. For the explanations, examples, and beneficial effects of the relevant content, reference can be made to the corresponding parts in the above methods.

[0097] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).

[0098] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0099] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0100] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A heterogeneous image search method based on a generative adversarial network, characterized in that: The following steps are included: The visible light image of the target area to be tested is input into the trained HR-GAN network, and the visible light image is used to generate an infrared simulation image. The infrared image of the target area to be tested is then compared and analyzed with the generated infrared simulation image to detect the nature and quantity of unexploded ordnance on the surface. The HR-GAN network training steps include introducing a high-resolution network HRnet as a generative network in the Pix2pix architecture to maintain the high-resolution representation of the image in the network and alleviate the problem of unclear texture and high noise in the generated infrared image. The Restomer module is introduced into the generative network to promote the learning of multi-scale local-global representation of images and alleviate the problem of missing structure in generating infrared images.

2. The heterogeneous image search method based on a generative adversarial network according to claim 1, characterized in that: The ECA model is introduced on the basis of HRnet to form an improved ECA-HRnet network; Among them, the ECA-HRnet input image size is 256×192, and it starts with a stem consisting of two 3×3 convolutions with a stride of 1; after the stem, the resolution is reduced to 1 / 4 of the input image feature map resolution; Its backbone network consists of 4 steps. In the first step, the network consists of 4 residual units, each of which consists of ECA-Bottleneck, that is, the ECA module is embedded in the Bottleneck, and then a convolution kernel 3×3 is used to reduce the width of the feature map resolution. In the second step, a multi-resolution module is repeated separately; In the third step, four multi-resolution modules are repeated one by one respectively; In the fourth step, three multi-resolution modules are repeated respectively; Each multi-resolution module contains two parts: one is parallel multi-resolution convolution and the other is multi-resolution fusion. Branches of different resolutions are connected in parallel, and each branch contains four residual units. Each unit is composed of ECA-BasicBlock, that is, the ECA module is embedded in BasicBlock; branches of different resolutions perform feature fusion to complete information exchange. The process from low resolution to high resolution is mainly achieved through bilinear upsampling, and the process from high resolution to low resolution adopts one or more strided convolutions, that is, convolution layers with a convolution kernel of 3×3 and a stride of 2.

3. The heterogeneous image search method based on generative adversarial network according to claim 2, characterized in that: In order to obtain different local channel information, the mapping relationship between channel C and convolution kernel size K is expressed as: C=Φ(K)=2 (γ×K-b) (1) Therefore, the K value will change adaptively with the number of channels, expressed as: Where K is the convolution kernel size, which represents the interaction between the current channel feature and K channel information, C represents the number of channels, and b and γ are fixed values, which are set to 1 and 2 respectively.

4. The heterogeneous image search method based on generative adversarial network according to claim 3, characterized in that: The Restomer module is introduced into the generative network HRnet; first, 1×1 convolution is used to increase the dimension, then 3×3 group convolution is used to divide the features into three blocks, and finally the classic self-attention calculation is performed.

5. The heterogeneous image search method based on generative adversarial network according to claim 4, characterized in that: The loss function of the HR-GAN network is: ζ cGAN (G,D)E x,y [log(x,y)]+E x [log(1-D(x,G(x)))](3) In the formula, The loss function of the conditional generative adversarial network is borrowed. D represents the discriminator, G represents the generator, x represents the input image, y represents the corresponding real image, and z represents the random variable.

6. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 5.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Visible light-infrared pedestrian re-identification method and system

    CN113887353A

  • Repairable casting defect image generation method based on HRNet improved CycleGAN

    CN117474837A

  • Unmanned aerial vehicle infrared image detection method, computer equipment and storage medium

    CN118570670A

  • Method for estimating depth of scene in image and computing device for implementation of the same

    WO2021096324A1