An intelligent detection and recognition method for weak and small targets

By constructing a conditional adversarial variational autoencoder and improving YOLOV3 model, combining denoising and sliding window iterative recognition technology, the accuracy and speed problems of weak target detection and recognition of infrared systems in complex backgrounds are solved, and high-precision and low false alarm rate detection effect is achieved.

CN114155411BActive Publication Date: 2025-06-10SICHUAN JIUZHOU ELECTRIC GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111501469.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-06-10
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

Existing infrared systems are difficult to accurately detect and identify weak targets in complex contexts, and the detection speed is slow and the false alarm rate is high.

Method used

A weak target intelligent detection and recognition method is adopted to extract semantic information of multiple types of images, a conditional adversarial variational autoencoder is constructed, and the YOLOV3 model is improved, combining the denoising method based on residual learning and sliding window iterative recognition technology to improve detection accuracy and speed and reduce false alarm rate.

Benefits of technology

It realizes high-precision weak object detection and recognition in complex backgrounds, improves detection speed, reduces false alarm rate, and improves the real-time performance and accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155411B_ABST
    Figure CN114155411B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for intelligent detection and recognition of small and weak targets, belonging to the technical field of image recognition, and solves the problems of inaccurate recognition results and slow recognition speed of small and weak infrared targets in the prior art. It includes: extracting semantic information of multiple types of images, and constructing a conditional adversarial variational autoencoder based on the semantic information; constructing an improved YOLOV3 model through a forward-shifted scale detection structure and an adjusted convolutional structure; performing denoising processing on the received infrared image to obtain a denoised image, and adding it to the noise-free image set; inputting the noise-free image set into the constructed conditional adversarial variational autoencoder to obtain an extended image set, and dividing the extended image set into a training set and a test set; training the improved YOLOV3 model based on the training set to obtain a target recognition model; inputting the test set into the target recognition model to obtain the recognition results of the targets in the infrared image, including the positions and types of the targets. It realizes accurate and rapid recognition of small and weak targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a method for intelligent detection and recognition of small targets. Background Art

[0002] With the increasing speed of offensive weapons such as aircraft and missiles, infrared optical systems are required to search for incoming targets at a longer distance to provide sufficient warning time for our countermeasure system. However, when the target is far away, the system receives a weak signal in a complex background. Not only is the background complex, the target signal is weak, the signal-to-noise ratio is low, the size is small, but also there is no shape and texture information, making it difficult to detect it from the background. As a result, how to achieve reliable detection and recognition of weak targets in complex backgrounds has become a hot topic widely studied at home and abroad.

[0003] In order to solve the problem of detecting and identifying weak targets in infrared systems, scholars have proposed two types of methods: detection first and then classification, and intelligent methods based on deep learning. The former can be mainly divided into detection methods based on single-frame spatial domain and time-domain sequence images. When the target signal-to-noise ratio is relatively low, it usually leads to a large number of false alarms and poor detection performance. It also has the disadvantages of many parameters that need to be adjusted manually and requires online adaptive calculation. With the development of artificial intelligence, intelligent detection methods based on deep learning have emerged, which have achieved good results in intelligent monitoring, face recognition, medical fields, etc. They can be mainly divided into two-stage detection models, such as FasterR-CNN, and single-stage models, such as YOLO, SSD and other detection and recognition algorithms. According to different scenarios and requirements of indicators such as timeliness and accuracy, corresponding methods can be selected and improved in a targeted manner.

[0004] However, traditional intelligent detection and recognition methods are mainly used for large-scale targets, and have poor detection and recognition performance for small targets. Specifically, there are the following shortcomings: First, for large-scale targets, the detection and recognition performance for small targets with target pixels less than 16×16 is poor; second, the infrared images used for training are noisy and insufficient in number, which affects the detection and recognition results; third, the detection and recognition speed cannot meet the real-time detection requirements of actual equipment; fourth, the target detection false alarm rate is relatively high. Summary of the invention

[0005] In view of the above analysis, an embodiment of the present invention aims to provide a small target intelligent detection and recognition method to solve the problems of inaccurate recognition results and slow recognition speed of existing infrared small target recognition methods.

[0006] On the one hand, an embodiment of the present invention provides a method for intelligent detection and recognition of small targets, comprising the following steps:

[0007] Extract semantic information of multiple categories of images, and build a conditional adversarial variational autoencoder based on the semantic information; build an improved YOLOV3 model by forward-shifting the scale detection structure and adjusting the convolution structure;

[0008] De-noising the received infrared image to obtain a denoised image, and add it to the noise-free image set;

[0009] Input the noise-free image set into the constructed conditional adversarial variational autoencoder to obtain an extended image set, and divide the extended image set into a training set and a test set;

[0010] Based on the training set, the improved YOLOV3 model is trained to obtain the target recognition model; the test set is input into the target recognition model to obtain the recognition results of the target in the infrared image, including the location and type of the target.

[0011] Based on the further improvement of the above method, the semantic information of multiple types of images is extracted, including:

[0012] Based on the public image dataset, a trained neural network is obtained; wherein the input of the trained neural network is the image in the public image dataset, and the output is the image feature;

[0013] From the public image dataset, randomly select multiple categories of background images, and select multiple images from each category as the public sample set;

[0014] Input each image in the public sample set into the trained neural network and obtain the image features output by the fully connected layer as the semantic information of each image;

[0015] According to the semantic information and image category of each image, the semantic information of each category of images is obtained by summing and averaging.

[0016] Based on the further improvement of the above method, the conditional adversarial variational autoencoder includes: an encoder module, a decoder module, and a discriminator module; wherein,

[0017] An encoder module, used for obtaining a feature map that obeys a normal distribution according to a feature map of an input image;

[0018] A decoder module is used to decode the feature map of the fused semantic information to generate a reconstructed image and a random image; the feature map of the normally distributed image, the semantic information of the image and the randomly generated semantic information are spliced ​​to obtain the feature map of the fused semantic information;

[0019] The discriminator module is used to discriminate the input image, the reconstructed image and the random image respectively, and obtain the discriminant result of whether the input image, the reconstructed image and the random image carry real information. Based on the further improvement of the above method, the loss function in the conditional adversarial variational autoencoder includes: variational autoencoder loss function and discriminator loss function; among them,

[0020] The variational autoencoder loss function includes: KL divergence, the variational lower bound of the marginal likelihood, and the discriminative loss function of the reconstructed image;

[0021] The discriminator loss function includes: the discriminative loss function of the input image and the discriminative loss function of the random image.

[0022] Based on the further improvement of the above method, an improved YOLOV3 model is constructed by moving the scale detection structure forward and adjusting the convolutional structure, including:

[0023] All three original scale detection structures in the YOLOV3 model are moved forward by one module. The three scale sizes are 26×26, 52×52, and 104×104, and three aspect ratios of 1:1, 1:2, and 2:1 are output respectively;

[0024] The feature map with 8 times downsampling output in the backbone network is upsampled by 2 times to match the size of the feature map with 4 times downsampling in the backbone network;

[0025] The 8 times downsampled feature map after 2 times upsampling processing is spliced and fused with the 4 times downsampled feature map, and classification and coordinate regression are performed on the obtained feature map with a size of 104×104;

[0026] The convolutional layer and the batch normalization layer are merged into one layer.

[0027] Based on the further improvement of the above method, it also includes: comparing the recognition result of the target with the test set, and identifying whether it exists in the corresponding image in the test set according to the type of the target in the recognition result of the target. If it does not exist, the target is a false alarm target, and a false alarm flag is set;

[0028] An image sequence is taken out from the recognition result of the target according to a fixed quantity, and it is identified whether the average false alarm rate of the image sequence is greater than the false alarm rate index; if it is greater, the real target in the current image sequence is identified through sliding window iteration, the corresponding false alarm flag is deleted, and the optimized image sequence is output; otherwise, the current image sequence is directly output, and the next batch of image sequences is taken out according to a fixed quantity for continuous recognition and output.

[0029] Based on the further improvement of the above method, the step size of the sliding window is 1, and the length of the sliding window is obtained according to the following formula:

[0030]

[0031] where S represents the infrared image frame rate of S frames per second, m represents the target speed of m meters per second, and [S / 2] represents rounding up the value of S / 2.

[0032] Based on the further improvement of the above method, the real targets in the current image sequence are identified through sliding window iteration, including:

[0033] Take the false alarm targets in the current image sequence as the targets to be identified and add them to the set of targets to be identified;

[0034] Set the sliding window length L, and extract the sliding window sequence according to the sliding window length L;

[0035] Identify whether the number of occurrences of each target to be identified in the sliding window sequence is greater than the corresponding threshold. If it is greater, the target to be identified is a real target, remove it from the set of targets to be identified, and delete the false alarm flag of the target to be identified in the current image sequence. Otherwise, enter the next sliding window iteration until the set of targets to be identified is empty or all sliding windows in the current image sequence are completed;

[0036] Output the optimized image sequence of the current image sequence, including:

[0037] Starting from the (L + 1)-th image of the current image sequence, if the false alarm target of the image is included in the set of targets to be identified, do not output the false alarm target.

[0038] Based on the further improvement of the above method, the received infrared image is denoised through a feed-forward convolutional neural network based on residual learning to obtain the denoised image, including:

[0039] The received infrared image is input into the input layer;

[0040] The middle layer adopts a residual structure, approximates the residual mapping as image noise for learning, does not adopt a pooling layer, and directly fills the boundary of the feature map using convolution; among them, the first layer adopts convolution and batch normalization, in the second layer to the penultimate layer, batch normalization is added between convolution and the ReLU activation function, and the last layer adopts convolution;

[0041] The output layer outputs the residual image;

[0042] Subtract the feature map of the infrared image from the feature map of the residual image to obtain the denoised image.

[0043] Based on the further improvement of the above method, approximating the residual mapping as image noise for learning, including:

[0044] Calculate the average error through the following formula to learn the training parameters; when the average error converges to the threshold, obtain the trained network model:

[0045]

[0046] Among them, N represents the image size, K represents the number of filters, b represents the expected residual image, a represents the estimated residual image, λ represents the regularization parameter, fk ·a represents the convolution of the estimated residual image a and the k-th filter f k and ρ k (·) represents the k-th adjustable penalty parameter in the network model, and p represents the p-th pixel.

[0047] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0048] 1. Provide a systematic solution. Through five aspects of infrared image reception, image denoising, dataset expansion, intelligent detection network optimization, and false alarm rate processing, automatically improve the quality of the training set, effectively expand the number of samples, improve the detection and recognition rate, and reduce the false alarm rate;

[0049] 2. A denoising method based on residual learning. This method does not learn a discriminant model with the prior of the display image, but regards image denoising as a simple discriminant learning problem, separates noise from the image through a feedforward convolutional neural network, realizes automatic image denoising, and solves the problem of low quality of the original image;

[0050] 3. Based on conditional variational auto-encoding technology, introduce the adversarial idea, add the features of the common background image, expand the dataset, overcome the limitations of the data acquisition channel, and provide a data basis for improving the generalization ability of the intelligent detection and recognition method for small and weak targets;

[0051] 4. Optimize and adjust the YOLOV3 model through feature forward movement, prediction forward movement, and inference acceleration, and solve the problem that traditional methods cannot well adapt to small and weak targets;

[0052] 5. Based on the target movement, flexibly set the sliding window size, and identify real targets through the detection of image sequences, solving the problem of high false alarm rate in the detection and recognition of small and weak targets in complex backgrounds.

[0053] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can be made obvious from the description, or understood by implementing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the content specifically pointed out in the description and the drawings. Description of the Drawings

[0054] The drawings are only for the purpose of showing specific embodiments, and are not considered as limiting the present invention. Throughout the drawings, the same reference signs represent the same components.

[0055] Figure 1 It is a flowchart of the intelligent detection and recognition method for small and weak targets in the embodiment of the present invention;

[0056] Figure 2 Schematic diagram of the conditional variational autoencoder framework in an embodiment of the present invention;

[0057] Figure 3 Schematic diagram of the improved YOLOV3 structure in an embodiment of the present invention;

[0058] Figure 4 Schematic diagram of the feedforward convolutional neural network based on residual learning in an embodiment of the present invention;

[0059] Figure 5 Schematic diagram of the sliding window and counting processing process in an embodiment of the present invention. Detailed implementation manners

[0060] The following will specifically describe the preferred embodiments of the present invention with reference to the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.

[0061] The infrared images in the present invention can be target images from different platforms and equipment equipped with different bands such as medium-long wave and medium wave, and different aperture infrared optical systems in different spatial domains such as air, land, and sea. An infrared image can include one or more targets. The infrared image information includes: image type, information acquisition time, image pixel information, the upper left coordinate position of each target in the image, the lower right coordinate position, and the target type.

[0062] A specific embodiment of the present invention discloses a method for intelligent detection and recognition of small and weak targets, as Figure 1 shown, including the following steps:

[0063] S11: Extract the semantic information of multiple types of images, and based on the semantic information, construct a conditional adversarial variational autoencoder;

[0064] Since the target information in the infrared image accounts for a small number of pixels in the entire image, in order to effectively utilize the limited target geometric attributes, such as: length, width, tilt angle, etc., semantics are further introduced on the basis of the conditional variational autoencoder (CVAE) to generate new samples that conform to specific semantic information.

[0065] Specifically, according to the public image dataset and the traditional neural network, extract the image semantic information, including:

[0066] ① Based on the public image dataset, obtain a trained traditional neural network; the input of the trained neural network is the image in the public image dataset, and the output is the image feature;

[0067] Exemplarily, a traditional CNN network is trained using the COCO dataset (Common Objects in Context, a dataset for image recognition provided by the Microsoft team).

[0068] ② Randomly select multiple categories of background images from the public image dataset, and select multiple images for each category as the public sample set;

[0069] It should be noted that: the number of background image categories selected from the public image dataset, and the number of images selected for each category, are greater than the number of categories and the number of images for each category in the noise-free image set obtained in step S11. Preferably, the number of categories is greater than 20, and the number of images for each category is greater than 100.

[0070] Exemplarily, select images such as borders, textures, and shapes in the public image dataset to form the public sample set.

[0071] ③ Input each image in the public sample set into the trained traditional neural network, and obtain the image features output by the fully connected layer as the semantic information of each image;

[0072] ④ According to the semantic information and image category of each image, sum and average to obtain the semantic information of each category of images. The formula is:

[0073]

[0074] where k represents the number of image categories, m k represents the number of images in the k-th category, represents the semantic information of the i-th image in the k-th category, and Fea k represents the semantic information of the k-th category of images.

[0075] Based on the extracted image semantic information, a conditional adversarial variational autoencoder is constructed. As Figure 2 shown, the conditional adversarial variational autoencoder includes: an encoder module, a decoder module, and a discriminator module; where,

[0076] The encoder module is used to obtain a feature map that follows a normal distribution N(μ, δ 2 ) according to the feature map of the input image; where μ represents the mean and δ 2 represents the covariance.

[0077] The decoder module is used to decode the feature map fused with semantic information to generate a reconstructed image and a random image; the feature map of the normal distribution, the semantic information of the image, and the randomly generated semantic information are concatenated to obtain the feature map fused with semantic information;

[0078] It should be noted that the randomly generated semantic information is randomly sampled from all possible latent distribution spaces of the input image.

[0079] The discriminator module is used to discriminate the input image, the reconstructed image, and the random image respectively, and obtain the discrimination results of whether the input image, the reconstructed image, and the random image carry real information.

[0080] It should be noted that the role of the discriminator is to discriminate the input image and the reconstructed image as carrying real images, and the random image as not carrying real information, while the encoder and the decoder make the random image with random information as real as possible. Therefore, the discriminator and the encoder and decoder are in confrontation with each other to promote the codec to generate more real images.

[0081] In the conditional adversarial variational autoencoder, the input image, the reconstructed image, the random image, and the related semantic information are input into the discriminator together. The discrimination loss of the discriminator for the original image and the generated image is defined as the discriminator loss, and the discrimination loss for the reconstructed image is still regarded as the loss in the Conditional AutoEncoder (CVAE), which is used to improve the situation where the CVAE only adjusts the network parameters depending on the reconstruction error, and at the same time enhance the training of the conditional adversarial variational autoencoder network. Therefore, the improved loss function is:

[0082] L CAVAE = L ICVAE + E x~p(x) [log(D(x))] + E h~p(h) [log(1 - D(G(h|y)))] Equation (2)

[0083] Among them, x represents the input image, y represents the semantic attribute of x, h represents the random image, and E x~p(x) [log(D(x))] is the discrimination loss of the input image, and E h~p(h) [log(1 - D(G(h|y)))] is the discrimination loss of the random image, and L ICVAE is the improved CVAE loss function, that is, the CVAE loss function that adds the discrimination loss of the reconstructed image on the basis of the original CVAE loss function L CVAE as follows:

[0084] L ICVAE = L CVAE + λE z~p(z) [log(D(G(z|y)))] Equation (3)

[0085]

[0086] Among them, z represents the latent variable of the input image x, φ represents the parameters of the encoder, θ represents the parameters of the decoder, and E z~p(z)[log(D(G(z|y)))] is the discriminative loss of the reconstructed image, λ is the balancing parameter, D KL (q φ (z|x,y)||p θ (z)) is the relative entropy or Kullback-Leiler (KL) divergence between the auxiliary distribution and the prior distribution, representing the variational lower bound of the marginal likelihood.

[0087] S12: Construct an improved YOLOV3 model by moving the scale detection structure forward and adjusting the convolutional structure;

[0088] It should be noted that the existing YOLOV3 model cannot well support the detection and recognition of small and weak targets. By moving the scale detection structure forward and adjusting the convolutional structure, an improved YOLOV3 model is constructed, as Figure 3 shown. Specifically, the improvement points include:

[0089] Move all three original scale detection structures in the YOLOV3 model forward by one module, that is, the three detection scales are 26×26, 52×52, and 104×104 respectively. The corresponding three prediction branches are set and adjusted based on the target ratio in the image, and output three aspect ratios of 1:1, 1:2, and 2:1 respectively to adapt to more resolutions and improve the detection and recognition ability of the algorithm for small and weak targets.

[0090] Perform 2x upsampling on the feature map with 8x downsampling output in the backbone network to match its size with the 4x downsampling feature map of the backbone network;

[0091] Concatenate and fuse the 8x downsampling feature map after 2x upsampling processing with the 4x downsampling feature map, and perform classification and coordinate regression on the obtained 104×104-sized feature map;

[0092] Merge the convolutional layer and the batch normalization layer into one layer to improve the inference speed.

[0093] By feature forward movement, prediction forward movement, and inference acceleration, the YOLOV3 model is optimized and adjusted, solving the problem that the traditional method cannot well adapt to small and weak targets.

[0094] It should be noted that S11 and S12 do not have an inevitable order in processing, and this embodiment does not limit their processing order.

[0095] S13: Denoise the received infrared image to obtain the denoised image and add it to the noise-free image set;

[0096] Preferably, the preprocessing of denoising the received infrared image is regarded as a discriminative learning problem. Through a feedforward convolutional neural network, the denoising method based on residual learning does not learn the prior information of the image and separates the noise from the image. The network architecture is as Figure 4 shown and includes:

[0097] The input layer is the received infrared image;

[0098] The middle layer adopts a residual structure, approximates the residual mapping as image noise for learning, improves the training speed, stability and accuracy; does not adopt a pooling layer to retain the original image features as much as possible; uses convolution to directly fill the boundaries of the feature map to ensure that each feature layer in the middle has the same size as the input image, and effectively removes boundary artifacts for subsequent processing; among them, the first layer adopts convolution and batch normalization, the batch normalization is added between the convolution and the ReLU activation function from the second layer to the penultimate layer, and the last layer adopts convolution;

[0099] The output layer is the residual image;

[0100] Subtract the feature map of the infrared image from the feature map of the residual image to obtain the denoised image.

[0101] During the training process of the network model, the training parameters are learned through the average error between the expected residual image and the estimated residual image from the noise input. When the average error converges to the threshold, the trained network model is obtained. The formula for calculating the average error is:

[0102]

[0103] where, N represents the image size, K represents the number of filters, b represents the expected residual image, a represents the estimated residual image, λ represents the regularization parameter, f k ·a represents the convolution of the estimated residual image a with the kth filter f k of, ρ k (·) represents the kth adjustable penalty parameter in the network model, and p represents the pth pixel.

[0104] To solve the problems of limited sample data quantity and sample imbalance, the denoised image is used for dataset expansion.

[0105] S14: Input the noise-free image set into the constructed conditional adversarial variational autoencoder to obtain an expanded image set, and divide the expanded image set into a training set and a test set;

[0106] It should be noted that the images in the noise-free image set are used as the sample set and input into the constructed conditional adversarial variational autoencoder. According to the loss function of the conditional adversarial variational autoencoder, the network parameters in the encoder, decoder, and discriminator are optimized respectively through the backpropagation algorithm, and the discriminant result of whether the generated image carries real information is output. The generated image carrying real information is used as the extended image. The extended image and the denoised image are compared pairwise to obtain the target type, and then the coordinate position of the target is manually marked to obtain the extended image set, which is divided into a training set and a test set according to a ratio of 8:2 or 7:3.

[0107] S15: Based on the training set, train the improved YOLOV3 model to obtain the target recognition model; input the test set into the target recognition model to obtain the recognition result of the target in the infrared image, including the position and type of the target.

[0108] Specifically, based on the training set, train the improved YOLOV3 model, compare the coordinate position and type of the target in the recognition result of the target with the coordinate position and type of the target obtained by manual marking, use the mAP (Mean Average Precision) value as the evaluation index. When the mAP value is less than the set threshold, adjust the network parameters and re-perform the network training process; when the network training reaches the set number of iterations and the mAP meets the standard, save the weight parameters of the trained network model to obtain the trained improved YOLOV3 model, that is, the target recognition model.

[0109] Then input the test set into the target recognition model to obtain the recognition result of the target, including the position and type of the target.

[0110] S16: Based on the recognition result of the target, through the sliding window image sequence detection, identify the real target and output the optimized image sequence.

[0111] Preferably, compare the recognition result of the target with the test set, and according to the type of the target in the recognition result of the target, identify whether it exists on the corresponding image in the test set. If not, the target is a false alarm target and a false alarm flag is set.

[0112] Since the false alarm rate of a single-frame image is generally high, especially for complex backgrounds, such as scenes where the target passes through clouds, between air and ground, etc., therefore, a sliding window mechanism is adopted, and the sliding window size is flexibly set based on the target movement to count and judge the false alarm targets for the continuous image sequence, identify the real target, and reduce the false alarm rate.

[0113] Specifically, an image sequence is taken from the recognition result of the target in a fixed quantity, and it is determined whether the average false alarm rate of the image sequence is greater than the false alarm rate index; if it is greater, the real target in the current image sequence is identified through sliding window iteration, the corresponding false alarm flag is deleted, and the optimized image sequence is output; otherwise, the current image sequence is directly output, and then the next batch of image sequences is taken in a fixed quantity for continuous recognition and output.

[0114] It should be noted that the fixed quantity is preset according to the moving speed of the target. Exemplarily, 100 consecutive images or 500 consecutive images are taken for the next detection. Moreover, considering that there may be one or more targets in an image, after calculating the false alarm rate of each target in the taken image sequence, the average is calculated as the average false alarm rate for comparison with the preset false alarm rate index. If it is greater than the false alarm rate index, it indicates that the false alarm rate is relatively high and further optimization is required. The real target in the current image sequence is identified through sliding window iteration, and the corresponding false alarm flag is deleted.

[0115] The step size of the sliding window is 1, and the length of the sliding window is obtained according to the following formula:

[0116]

[0117] where S represents the frame rate of the infrared image as S frames per second, m represents the target speed as m meters per second, and [S / 2] represents rounding up the value of S / 2.

[0118] As Figure 5 shown, the real target in the current image sequence is identified through sliding window iteration. Here, D represents the length of the current image sequence, and R represents the number of sliding window iterations. The specific process is as follows:

[0119] The false alarm targets in the current image sequence are used as the targets to be recognized and added to the set of targets to be recognized.

[0120] Set the sliding window length L according to formula (6), and take out the sliding window sequence according to the sliding window length L.

[0121] Identify whether the number of occurrences of each target to be recognized in the sliding window sequence is greater than the corresponding threshold. If it is greater, the target to be recognized is a real target, which is removed from the set of targets to be recognized, and the false alarm flag of the target to be recognized in the current image sequence is deleted. Otherwise, enter the next sliding window iteration until the set of targets to be recognized is empty or all sliding windows of the current image sequence are completed.

[0122] According to the sliding window result, the optimized image sequence of the current image sequence is output simultaneously, including:

[0123] Starting from the (L + 1)-th image of the current image sequence, if the false alarm target of the image is included in the set of targets to be recognized, the false alarm target is not output.

[0124] Exemplarily, assume that L is 60. Then, the output of the target in the 61st image depends on the set to be recognized in the sliding window sequence detection from 1 to 60, and the output of the target in the 62nd image depends on the set to be recognized in the sliding window sequence detection from 2 to 61.

[0125] By detecting the continuity of the images, false alarm targets can be converted into real targets, and possible false alarm targets can be not displayed according to the real-time sliding window detection results, thereby reducing the false alarm rate of the entire image sequence.

[0126] Optimally, compare and analyze the false alarm rate in the optimized image sequence with the false alarm rate in the recognition result of the target obtained in step S15, and optimize and adjust the model structure or parameters according to the difference result.

[0127] Furthermore, when performing real-time detection and recognition on weak infrared images, after obtaining the noise-free image through step S13, output the target recognition model in step S15 to obtain the recognition result of the target in the real-time infrared image, including the position and type of the target.

[0128] Compared with the prior art, a method for intelligent detection and recognition of small and weak targets provided in this embodiment provides a systematic solution. Through five aspects of infrared image reception, image denoising, dataset expansion, intelligent detection network optimization, and false alarm rate processing, it automatically improves the quality of the training set, effectively expands the number of samples, improves the detection and recognition rate, and reduces the false alarm rate; a denoising method based on residual learning, which does not learn a discriminant model with the prior of the display image, but regards image denoising as a simple discriminant learning problem, separates the noise from the image through a feedforward convolutional neural network, and realizes automatic image denoising to solve the problem of low quality of the original image; expands the dataset through conditional variational auto-encoding technology, overcomes the limitations of data acquisition channels, and provides a data basis for the improvement of the generalization ability of the method for intelligent detection and recognition of small and weak targets; optimizes and adjusts the YOLOV3 model through feature forward movement, prediction forward movement, and inference acceleration, and solves the problem that traditional methods cannot well adapt to small and weak targets; flexibly sets the sliding window size based on the target movement, and recognizes real targets through the detection of the image sequence, and solves the problem of high false alarm rate in the detection and recognition of small and weak targets in complex backgrounds.

[0129] Those skilled in the art can understand that all or part of the processes of implementing the method in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.

[0130] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for intelligent detection and recognition of small targets. It is characterized in that The steps include: Extract semantic information of multiple types of images, and construct a conditional adversarial variational autoencoder based on the semantic information; construct an improved YOLO V3 model by forward-shifting the scale detection structure and adjusting the convolution structure; the forward-shifted scale detection structure is to shift all three original scale detection structures in the YOLO V3 model forward by one module, and the three scale sizes are 26×26, 52×52 and 104×104 respectively; De-noising the received infrared image to obtain a denoised image, and add it to the noise-free image set; Inputting the noise-free image set into the constructed conditional adversarial variational autoencoder to obtain an extended image set, and dividing the extended image set into a training set and a test set; Based on the training set, the improved YOLO V3 model is trained to obtain a target recognition model; the test set is input into the target recognition model to obtain a recognition result of the target in the infrared image, including the position and type of the target; Comparing the recognition result of the target with the test set, and identifying whether the target type in the recognition result of the target exists in the corresponding image in the test set, if not, the target is a false alarm target, and a false alarm flag is set; A fixed number of image sequences are taken from the target recognition results to identify whether the average false alarm rate of the image sequence is greater than the false alarm rate index; if greater, the real target in the current image sequence is identified through sliding window iteration, the corresponding false alarm flag is deleted, and the optimized image sequence is output; Otherwise, directly output the current image sequence, and then take out the next batch of image sequences according to a fixed number to continue recognition and output; The step size of the sliding window is 1, and the length of the sliding window is obtained according to the following formula: Wherein, S means the infrared image frame rate is S frames / second, m means the target speed is m meters / second, and [S / 2] means rounding up the value of S / 2.

2. According to the method for intelligent detection and recognition of small targets according to claim 1, It is characterized in that The extracting semantic information of multiple types of images includes: Based on the public image dataset, a trained neural network is obtained; wherein the input of the trained neural network is the image in the public image dataset, and the output is the image feature; From the public image dataset, randomly select multiple categories of background images, and select multiple images from each category as the public sample set; Input each image in the public sample set into the trained neural network, and obtain the image features output by the fully connected layer as the semantic information of each image; According to the semantic information and image category of each image, the semantic information of each category of images is obtained by summing and averaging.

3. According to claim 2, the method for intelligent detection and recognition of small targets, It is characterized in that The conditional adversarial variational autoencoder includes: an encoder module, a decoder module, and a discriminator module; wherein, The encoder module is used to obtain a feature map that obeys a normal distribution according to the feature map of the input image; The decoder module is used to decode the feature map fused with semantic information to generate a reconstructed image and a random image; the feature map with a normal distribution, the semantic information of the image, and the randomly generated semantic information are concatenated to obtain the feature map fused with semantic information; The discriminator module is used to discriminate the input image, the reconstructed image, and the random image respectively, and obtain the discrimination results of whether the input image, the reconstructed image, and the random image carry real information.

4. The small target intelligent detection and recognition method according to claim 3, wherein, the loss function in the conditional adversarial variational autoencoder includes: a variational autoencoder loss function and a discriminator loss function; wherein, the variational autoencoder loss function includes: KL divergence, the variational lower bound of marginal likelihood, and the discrimination loss function of the reconstructed image; the discriminator loss function includes: the discrimination loss function of the input image and the discrimination loss function of the random image.

5. The small target intelligent detection and recognition method according to claim 1 or 4, wherein, the improved YOLO V3 model constructed by the forward-shifted scale detection structure and the adjusted convolutional structure includes: All three original scale detection structures in the YOLO V3 model are shifted forward by one module. The three scale sizes are 26×26, 52×52, and 104×104, and three aspect ratios of 1:1, 1:2, and 2:1 are output respectively; The feature map with 8 times downsampling output in the backbone network is upsampled by 2 times to match the size of the feature map with 4 times downsampling in the backbone network; The 8 times downsampled feature map after 2 times upsampling processing is concatenated and fused with the 4 times downsampled feature map, and classification and coordinate regression are performed on the obtained 104×104 sized feature map; The convolutional layer and the batch normalization layer are merged into one layer.

6. The small target intelligent detection and recognition method according to claim 1, wherein, identifying the real target in the current image sequence through sliding window iteration includes: Regarding the false alarm targets in the current image sequence as targets to be recognized and adding them to the set of targets to be recognized; Setting the sliding window length L, and taking out the sliding window sequence according to the sliding window length L; Identifying whether the number of occurrences of each target to be recognized in the sliding window sequence is greater than the corresponding threshold. If it is greater, the target to be recognized is a real target, removed from the set of targets to be recognized, and the false alarm flag of the target to be recognized in the current image sequence is deleted. Otherwise, enter the next sliding window iteration until the set of targets to be recognized is empty or all sliding windows in the current image sequence are completed; The output of the optimized image sequence of the current image sequence includes: Starting from the (L + 1)-th image of the current image sequence, if the false alarm target of the image is included in the set of targets to be recognized, the false alarm target is not output.

7. The small target intelligent detection and recognition method according to claim 1, wherein, denoising the received infrared image through a feedforward convolutional neural network based on residual learning to obtain a denoised image, including: The received infrared image is input into the input layer; The middle layer adopts a residual structure, approximates the residual mapping as image noise for learning, does not adopt a pooling layer, and directly fills the boundary of the feature map by convolution; among them, the first layer adopts convolution and batch normalization, and in the second layer to the penultimate layer, batch normalization is added between the convolution and the ReLU activation function, and the last layer adopts convolution; The output layer outputs a residual image; Subtract the feature map of the residual image from the feature map of the infrared image to obtain a denoised image.

8. The intelligent detection and recognition method for small and weak targets according to claim 7, characterized in that the approximating the residual mapping as image noise for learning includes: calculating the average error through the following formula to learn the training parameters; when the average error converges to a threshold, obtaining a trained network model: Where, N represents the image size, K represents the number of filters, b represents the desired residual image, a represents the estimated residual image, λ represents the regularization parameter, and f k ·a represents the convolution of the estimated residual image a with the k-th filter f k , and ρ k (·) represents the k-th adjustable penalty parameter in the network model, and p represents the p-th pixel.

Citation Information

Patent Citations

  • VAEGAN-based incomplete projection CT image reconstruction method

    CN109146988A

  • Vehicle logo recognition method based on improved YOLO-V3 model

    CN112200186A