A Passive Terahertz Image Target Detection Method Based on Filtering Enhanced Deep Learning
Enhanced samples are generated through multi-scale filtering and spatial geometric transformation, and deep learning training is combined with convolutional neural networks, which solves the problems of low signal-to-noise ratio of passive terahertz image samples and difficult target recognition, and achieves efficient recognition of different noise and rotational forms.
Patent Information
- Application Number
- CN202010684465.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-07-15
AI Technical Summary
Passive terahertz image samples have low signal-to-noise ratio and blurred images. Traditional deep learning algorithms are difficult to effectively identify image samples with severe noise, and fixed image filtering preprocessing cannot ensure that the output denoising samples are suitable for deep learning algorithms to obtain high-precision recognition effects.
Multi-scale filtering is used to remove sample noise, multi-scale filtering enhancement samples are generated through multi-directional spatial geometric transformation, features are extracted and deep learning training are performed in convolutional neural networks, and multi-channel feature prediction and results are fused on the denoised samples to obtain the final target detection result.
It greatly improves the recognition detection rate of terahertz targets with different noise levels and different steering patterns, can effectively filter out striped noise, while retaining image details, and has target rotation invariance, improving the robustness of deep learning to noise and rotation patterns.
Smart Images

Figure CN113850725B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of passive terahertz image target detection, and particularly relates to a passive terahertz image target detection method for filtering enhancement and deep learning. Background Art
[0002] Terahertz waves refer to electromagnetic waves with frequencies in the range of 0.1 - 10 THz. The terahertz wave band can cover the characteristic spectra of substances such as semiconductors, plasmas, organisms, and biological macromolecules, and has good penetrability. Secondly, terahertz has low energy and hardly affects the human body. Terahertz technology can be widely applied in fields such as radar, remote sensing, homeland security and anti-terrorism, high-security data communication and transmission, atmospheric and environmental monitoring, real-time biological information extraction, and medical diagnosis.
[0003] Passive terahertz imaging technology, compared with active terahertz imaging technology, does not require an active terahertz radiation source and completes imaging by passively receiving terahertz waves emitted by the human body itself. It has the advantages of low cost, safety without radiation, non-contact concealment, etc., and has a better application prospect in the field of security inspection. In traditional security inspection, the detection of targets is usually completed manually, which greatly tests the patience and perseverance of inspectors. With the booming development of deep learning, current algorithms can already perform effective target detection on optical images. However, in terahertz images, especially passive terahertz images, related methods still need to be improved. Currently, there are mainly the following types of problems:
[0004] 1. Since the passive terahertz scanning system does not use a radiation source, the terahertz wave power radiated by the target object itself is low, and various types of noise interference are introduced during the imaging process, resulting in a very low signal-to-noise ratio of passive terahertz image samples and very blurred images, which affects the effective recognition of image content;
[0005] 2. Traditional deep learning algorithms have not been improved for the problem of identifying noisy terahertz samples. Due to problems such as noise differences and target morphology differences in terahertz samples, it is difficult to achieve high-precision recognition of all sample targets through simple and fixed image filtering preprocessing.
[0006] Based on the above objective problems, the sample quality of passive detection is low, which has a great impact on manual annotation and model training. Through traditional deep learning methods, it is impossible to produce a significant recognition effect for image samples with severe noise, and using a fixed image filtering method for preprocessing still makes it difficult to ensure that the output denoised samples are all suitable for deep learning algorithms to obtain high-precision recognition effects.
[0007] Based on the above market demands and technical problems, there is an urgent need for a passive terahertz image target detection method based on multi-scale filtering enhancement. Based on deep learning algorithms, it is improved and innovated. Through multi-scale filtering, deep learning's enhanced learning for different noise levels is completed, and then through spatial geometric transformation, deep learning's enhanced learning for different steering forms is completed, thereby greatly improving the recognition and detection rate of terahertz targets with different noise levels and different steering forms. Summary of the Invention
[0008] The present invention discloses a passive terahertz image target detection method for filtering and enhancing deep learning. By multi-scale denoising and spatial geometric transformation, the deep learning sample set is enhanced and amplified to obtain samples with different denoising intensities and target forms. While greatly filtering out stripe noise, it better retains image details and has target rotation invariance, improving the robustness of deep learning for target recognition with different noise levels and different steering forms. To solve the above technical problems, the present invention aims to provide a passive terahertz image target detection method for filtering and enhancing deep learning, including the following steps:
[0009] Use multi-scale filtering to remove sample noise, use multi-directional spatial geometric transformation, and jointly generate multi-scale filtering enhanced samples;
[0010] Use a convolutional neural network to extract features, train model parameters, and perform deep learning training;
[0011] Perform multi-channel feature prediction on the denoised samples, and fuse the prediction results of multiple channels to obtain the final target detection result.
[0012] Preferably, when using multi-scale filtering to remove sample noise, where the neighborhood pixel value is f(k, l), and the pixel value of the output image is g(i, j), g(i, j) is a weighted value combination of the neighborhood pixel value f(k, l), where (i, j) and (k, l) represent the coordinates of pixel points, and w(i, j, k, l) is equal to the product of the spatial domain kernel ws and the range domain kernel wr, where σ s and σ r are respectively the filtering smoothing parameters of the spatial domain and the range domain. Different denoising threshold samples l x and the filtering smoothing parameters σ s of the spatial domain and the range domain and σ r The relationship is l x = l 0 ·w(σ s , σ r ), where σ s = x, σ r = 13x, and different x values are taken respectively, x ∈ (0, 2], to obtain multi-scale denoised samples and perform multi-scale filtering enhancement.
[0013] As a preferred method, multi-directional spatial geometric transformation is used to jointly generate multi-scale filter enhancement samples, including rotating and flipping the sample image. The original pixel value coordinates of the image after filtering enhancement are (x 0 ,y 0 ), the coordinates of the image center after rotation are (x 1 ,y 1 ), Where θ is the angle of counterclockwise rotation, and the coordinates of the image after flipping the y-axis are (x 2 ,y 2 ), (x 2 ,y 2 )=(2w-x 0 ,y 0 ), where w is the width of the image.
[0014] Preferably, the feature extraction is performed by a convolutional neural network, wherein the network is composed of N-1 convolutional layers and 1 fully connected layer, firstly a convolution kernel with 32 filters, then n groups of repeated residual units, each unit is composed of a separate convolutional layer and a group of repeated convolutional layers, and the repeated convolutional layers are repeated g respectively. 1 Times, g 2 times, ..., g n times; a single convolutional layer uses a convolution with a stride of 2 for downsampling. In each repeated convolutional layer, a 1x1 convolution operation is performed first, and then a 3x3 convolution operation is performed. The number of filters is first halved and then restored, for a total of N-1 layers.
[0015] Preferably, the training model parameters are subjected to deep learning training, including training and optimizing the parameter model through a loss function. Before training, the image is decomposed into S×S grids, each grid containing A pre-selected boxes and B predicted boxes. The loss function consists of a coordinate prediction loss function, a confidence loss function and a category loss function, as shown in the following formula:
[0016]
[0017] Where i and j represent the jth prediction box of the i-th grid. b x ,b y ,b w ,b h Respectively represent the center point coordinates and length and width of the directly predicted prediction box, g x ,g y ,g w ,g h Represents the center point coordinates and length and width of the real box, c x ,cy , a w , a h respectively represent the distance from the upper left corner of the current grid to the upper left corner of the image and the length and width of the anchor box, t x , t y , t w , t h are the parameters to be learned;
[0018] When the prediction box within the grid is responsible for predicting the ground truth box, Otherwise
[0019] When the prediction box is responsible for predicting the target within the grid, Otherwise
[0020] When a certain prediction box is not responsible for predicting the ground truth box in the corresponding grid but has an overlap ratio with the ground truth box greater than the set threshold, G ij = 0, otherwise G ij = 1;
[0021] Among them, the confidence of the j-th prediction box in the i-th grid Pr(object) represents the probability that the current prediction box has an object, represents the overlap ratio between the ground truth box and the prediction box, and are the class probability and the true probability when predicting the class c respectively.
[0022] Preferably, the multi-channel feature prediction of the denoised sample includes inputting the denoised sample l into a convolutional neural network for multi-channel downsampling and feature prediction at 2 n times, 2 n-1 times,... 2 n-M times, where M is the number of channels; performing upsampling on the feature map of 2 n-i times and feature fusion on the feature map of 2 n-i+1 times, i ∈ [1, M], to complete multi-channel simultaneous prediction.
[0023] The beneficial effects brought by adopting the above technical solutions are:
[0024] (1) Aiming at the problems of serious noise, blurred details, and target differences in passive terahertz images, by multi-scale bilateral filtering to expand the denoised samples of passive terahertz images at different denoising scales, it is possible to filter out noise and retain image details for samples of different qualities;
[0025] (2) Aiming at the problem of different rotation forms of the same target, performing spatial geometric transformation enhancement to expand the transformed samples of passive terahertz images at different turning forms, which is robust to the differences in target rotation forms;
[0026] (3) Input the augmented sample into the feature extraction network to extract features, learn and train the pre-trained weights, and use the trained model to predict the denoised sample, completing the deep learning of data augmentation for the recognition of terahertz sample targets, thereby greatly improving the recognition and detection rate of terahertz targets with different noise levels and different turning shapes. It can not only filter out severe stripe noise but also avoid the loss of image details caused by excessive denoising. Brief Description of the Drawings
[0027] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0028] Figure 2 It includes the original passive terahertz images (a1) and (a2) of Scene 1 and Scene 2, the multi-scale bilateral filtering images (b1) and (b2), (c1) and (c2), (d1) and (d2) of Scene 1 and Scene 2;
[0029] Figure 3 It includes the passive terahertz image (a1) after filtering and enhancement, and the images (a2) and (a3) after spatial geometric transformation enhancement;
[0030] Figure 4 It is a structural diagram of the feature extraction network of the present invention;
[0031] Figure 5 It is an explanatory diagram of the loss function parameters of the present invention;
[0032] Figure 6 It is a multi-scale prediction structural diagram of the present invention;
[0033] Figure 7 It includes the recognition effect diagrams (a1) before multi-scale filtering enhancement and (a2) after multi-scale filtering enhancement.
[0034] Figure 8 It includes the recognition effect diagrams (a1) before spatial geometric transformation enhancement and (a2) after spatial geometric transformation enhancement. Detailed Embodiment
[0035] The technical solution of the present invention will be described in detail below with reference to the drawings.
[0036] A passive terahertz image target recognition method based on multi-scale filtering enhancement deep learning, the processing flow of which is as Figure 1 shown. The specific steps are as follows:
[0037] (1) Use multi-scale filtering to remove sample noise and use multi-directional spatial geometric transformation to jointly generate multi-scale filtering enhanced samples.
[0038] (2) Extract features using a convolutional neural network, train the model parameters, and perform deep learning training.
[0039] (3) Perform multi-channel feature prediction on the denoised samples, and fuse the prediction results of multiple channels to obtain the final object detection result.
[0040] The specific process of step (1) is as follows:
[0041] (1-1) Perform bilateral filtering on the labeled passive terahertz image samples. The pixel value g(i,j) of the output image is a weighted value combination of the neighborhood pixel values f(k,l), where (i,j) and (k,l) represent the coordinates of pixel points, and w(i,j,k,l) is equal to the product of the spatial domain kernel ws and the range domain kernel wr, where σ s and σ r are the filtering and smoothing parameters of the spatial domain and the range domain respectively. The sample l x at different denoising thresholds s and the filtering and smoothing parameters σ r of the spatial domain and the range domain x The relationship is l 0 = l s ·w(σ r ),σ s = x, σ r = 13x. Different x values are taken to obtain multi-scale denoised samples for data augmentation. In the embodiment, x = 0.5, 1, 1.5 are taken respectively, and their denoising effects are as shown in Figure 2 (b1) and (b2), (c1) and (c2), (d1) and (d2) respectively. It can be seen that the noise is effectively filtered out.
[0042] (1-2) Spatial geometric transformation enhancement. Rotate and flip the sample image. The original pixel value coordinates of the filtered and enhanced image are (x 0 ,y 0 ), and the coordinates obtained after rotating the image center are (x 1 ,y 1 ), where θ is the counterclockwise rotation angle. The coordinates of the image flipped along the y-axis are (x 2 ,y 2 ), (x 2 ,y 2 ) = (2w - x 0 ,y 0 ), where w is the width of the image. In the embodiment, θ = 180° and w = 64 are taken respectively. The rotation and flipping effects are as shown in Figure 3 (a2) and (a3) respectively.
[0043] Furthermore, the specific process of step (2) is as follows:
[0044] (2-1) Convolutional neural network feature extraction. The convolutional neural network consists of N-1 convolutional layers and 1 fully connected layer. First, there is a convolution kernel with 32 filters, followed by n groups of repeated residual units. Each unit consists of a single convolutional layer and a group of repeated convolutional layers. The repeated convolutional layers are repeated g 1 Times, g 2 times, ..., g n times; a single convolution layer uses a convolution with a step size of 2 for downsampling. In each repeated convolution layer, a 1x1 convolution operation is performed first, and then a 3x3 convolution operation is performed. The number of filters is first halved and then restored, for a total of N-1 layers. In the embodiment, a darknet53 network is used, where N=53, n=5, g1=1, g2=2, g3=8, g4=8, and g5=4. Its structure is shown in the figure Figure 4 shown.
[0045] (2-2) The parameter model is trained and optimized by the loss function. Before training, the image is decomposed into S*S grids, each grid contains A pre-selected boxes and B predicted boxes. The loss function of Yolov3 consists of the coordinate prediction loss function, the confidence loss function and the category loss function, as shown in the following formula:
[0046]
[0047] Where i and j represent the jth prediction box of the i-th grid. b x ,b y ,b w ,b h Respectively represent the center point coordinates and length and width of the directly predicted prediction box, g x ,g y ,g w ,g h Represents the center point coordinates and length and width of the real box, c x ,c y ,a w ,a h Represents the distance from the upper left corner of the current grid to the upper left corner of the image and the length and width of the anchor box, t x ,t y ,t w ,t h are the parameters that need to be learned. When the prediction box in the grid is responsible for predicting the true box, otherwise When the prediction box is responsible for predicting the object within the grid, otherwise When a prediction box is not responsible for predicting the ground truth box in the corresponding grid but its overlap ratio with the ground truth box is greater than the set threshold, G ij = 0; otherwise, G ij = 1. The confidence of the j-th prediction box in the i-th grid Pr(object) represents the probability that the current prediction box has an object, represents the overlap ratio between the ground truth box and the prediction box. and are the class probability and the true probability when predicting class c, respectively. The explanations of some of its parameters are as Figure 5 shown.
[0048] Furthermore, the specific process of step (3) is as follows:
[0049] (3-1) Perform multi-channel feature prediction on the denoised samples, and fuse the prediction results of multiple channels to obtain the final object detection result. Performing multi-channel feature prediction on the denoised samples means inputting the denoised sample l into the convolutional neural network to perform multi-channel downsampling and feature prediction at 2 n times, 2 n-1 times,... 2 n-M times, where M is the number of channels; then perform upsampling on the feature map of 2 n-i times, and perform feature fusion with the feature map of 2 n-i+1 times, i ∈ [1, M], to complete multi-channel simultaneous prediction. Its structure diagram is as Figure 6 shown.
[0050] The final recognition effect is as Figure 7 shown, Figure 7 (a1) and Figure 7 (a2) are the recognition effects before and after data filtering and enhancement, respectively. The mean average precision before filtering is 87.45%, and it is improved to 92.16% after filtering. It can be clearly seen that in the traditional method, when the appropriate denoising scale is not clear, there are cases of missed detection and false detection in the recognition effect, and the recognition effect after denoising filtering data enhancement has a significant improvement. Figure 8 (a1) and Figure 8 (a2) are the recognition effects before and after spatial geometric transformation enhancement, respectively. The mean average precision after enhancement is 94.27%. It can be clearly seen that using spatial geometric transformation enhancement has better robustness and adaptability to the rotation of the target object.
[0051] The embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the present invention.
Claims
1. A passive terahertz image target detection method based on filtering enhanced deep learning, The following steps are involved: Multi-scale filtering is used to remove sample noise, and multi-directional spatial geometric transformation is used to jointly generate multi-scale filtering enhanced samples; Use convolutional neural networks to extract features, train model parameters, and perform deep learning training; Perform multi-channel feature prediction on the denoised samples, fuse the multi-channel prediction results, and obtain the final target detection results; The multi-scale filtering is used to remove sample noise, wherein the neighborhood pixel value is f(k, l), the pixel value of the output image is g(i, j), and g(i, j) is a weighted value combination of the neighborhood pixel value f(k, l). Where (i, j) and (k, l) represent the coordinates of the pixel points, w(i, j, k, l) is equal to the product of the spatial kernel ws and the range kernel wr. σ s and σ r are the filtering smoothing parameters in the spatial domain and the value domain, respectively, and the different denoising threshold samples l x and the filter smoothing parameter σ in the spatial domain and the range s With σ r The relationship is l x = l 0 ·w(σ s ,σ r ), where σ s =x,σ r =13x, take different x,x∈(0,2] respectively, obtain multi-scale denoising samples, and perform multi-scale filtering enhancement; Using multi-directional spatial geometric transformation, jointly generate multi-scale filtered and enhanced samples, including rotating and flipping the sample image. The original pixel value coordinates of the filtered and enhanced image are (x 0 , y 0 ). The coordinates obtained after rotating the image center are (x 1 , y 1 ), where θ is the counterclockwise rotation angle. The coordinates of the image after flipping along the y-axis are (x 2 , y 2 ). (x 2 , y 2 ) = (2w - x 0 , y 0 ), where w is the width of the image.
2. According to the passive terahertz image target detection method of filtering enhanced deep learning according to claim 1, Features: The feature extraction is performed using a convolutional neural network. The network consists of N-1 convolutional layers and 1 fully-connected layer. First, there is a convolutional kernel with 32 filters, followed by n groups of repeated residual units. Each unit is composed of a single convolutional layer and a group of repeatedly executed convolutional layers. The repeatedly executed convolutional layers are repeated g 1 times, g 2 times,..., g n times; the single convolutional layer uses convolution with a stride of 2 for downsampling. In each repeatedly executed convolutional layer, a 1x1 convolution operation is first performed, followed by a 3x3 convolution operation. The number of filters is first halved and then restored, for a total of N-1 layers.
3. According to the passive terahertz image target detection method of filtering enhanced deep learning in claim 1, Features: The training model parameters are trained for deep learning training, including training and optimizing the parameter model through a loss function. Before training, the image is decomposed into S×S grids, each grid containing A pre-selected boxes and B predicted boxes. The loss function consists of a coordinate prediction loss function, a confidence loss function and a category loss function, as shown in the following formula: where i and j represent the j-th prediction box in the i-th grid, b x , b y , b w , b h respectively represent the center point coordinates, length, and width of the directly predicted prediction box, g x , g y , g w , g h respectively represent the center point coordinates, length, and width of the ground truth box, c x , c y , a w , a h respectively represent the distance from the upper left corner of the current grid to the upper left corner of the image and the length and width of the anchor box, t x , t y , t w , t h are the parameters to be learned; When the predicted bounding box within the grid is responsible for predicting the ground truth bounding box, Otherwise When the prediction box is responsible for predicting the target within this grid, Otherwise When a prediction box is not responsible for predicting the ground truth box in the corresponding grid but has an overlap ratio with the ground truth box greater than the set threshold, G ij = 0, otherwise G ij = 1; Among them, the confidence of the j-th prediction box in the i-th grid Pr(object) represents the probability that there is an object in the current prediction box, represents the overlap rate between the ground truth box and the prediction box, and are the class probability and the true probability when predicting class c, respectively.
4. According to the passive terahertz image target detection method of filtering enhanced deep learning as described in claim 2, Features: Performing multi-channel feature prediction on the denoised samples includes inputting the denoised sample l into a convolutional neural network for multi-channel downsampling and feature prediction at 2 n times, 2 n-1 times, … 2 n-M times, where M is the number of channels; performing upsampling on the feature map of 2 n-i times and performing feature fusion with the feature map of 2 n-i+1 times, i ∈ [1, M], to complete multi-channel simultaneous prediction.