Target detection feature extraction enhancement method under severe conditions
By introducing directed differential convolution and convolution edge response modules into the target detection network, the contour information is used to extract the characteristics of constant illumination, and the performance degradation caused by domain offset in the target detection method under harsh conditions is solved, and efficient object detection under harsh conditions is achieved.
Patent Information
- Application Number
- CN202411902846.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-30
AI Technical Summary
Under harsh conditions, the object detection method has deteriorated detection performance due to the domain offset caused by the difference in contrast and saturation of the image, and the existing methods have failed to effectively solve this problem.
By constructing a neural network module containing directed differential convolution (DDCM) and convolution edge response (CBR) modules, the contour information is used to extract the characteristics of the illumination invariant, alleviate the domain offset problem, and detect it in the target detection network.
It effectively alleviates the problem of image domain offset under harsh conditions, improves the target detection performance, and has a small impact on detection speed, and does not destroy the real-time nature of the target detection network.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer image processing, and particularly to a method for enhancing target detection feature extraction under harsh conditions. Background Art
[0002] The target detection enhancement network is to improve the detection effect of ordinary target detection algorithms on targets under harsh conditions. It is usually placed in front of ordinary target detectors and is of great significance for unmanned driving, remote sensing, etc. Currently, there are various end-to-end target detection enhancement networks. Enhancing image features based on differentiable preprocessing modules: For example, for a thick fog environment, an atmospheric degradation model is used to reconstruct a fog-free image, so that the network only needs to extract the features of the fog-free image for detection. Based on classical operators: Traditional algorithms such as the Sobel operator and Gaussian pyramid are used to directly extract contour texture features independent of the environment, which can enhance the feature extraction ability of the network. For example, the Sobel operator is used to extract the object contour in a low-light environment to enhance the details of the object.
[0003] However, even with the use of these target detection enhancement networks, the contrast and saturation of images taken under harsh conditions still have a large difference from those of general images. The domain shift formed by this difference will cause the performance of target detection methods applied to traditional datasets to decline under harsh conditions. Existing methods often ignore the influence caused by the non-uniform change of contrast in harsh environments. Summary of the Invention
[0004] The present invention aims to provide a method for enhancing target detection feature extraction under harsh conditions, which can effectively alleviate the image domain shift problem under harsh conditions by means of contour information, and has little impact on the detection speed of the target detection network and does not destroy the real-time performance of the target detection network.
[0005] The technical solution of the present invention is as follows:
[0006] The method for enhancing target detection feature extraction under harsh conditions includes the following steps:
[0007] A. Construct a neural network module, where the neural network module includes a DDCM module and a CBR module;
[0008] B. Input the original image into the neural network module and divide it into two paths; the first path is first processed by the DDCM module, and after the processing result of the DDCM module is multiplied by a preset parameter α, the first path result is obtained; the second path is processed by the CBR module to obtain the second path processing result; the first path processing result and the second path processing result are concatenated by the Concat function, and the obtained result is processed by the CBR module, and then processed by the SUM module with the processing result of the DDCM module to obtain the output result;
[0009] C. Input the output result into the target detection convolutional neural network for detection.
[0010] The DDCM module mentioned above is a directed difference convolution, and the directed difference convolution formula is:
[0011]
[0012] where W is the weight of the convolution kernel, R is the rotation matrix, and p 0 represents the position of the current convolution in the grid set P, and p′ and p n are respectively the unknown and adjacent pixel grids in the same column in the grid set P.
[0013] The grid set P mentioned above is:
[0014] P = ((-1, -1); (-1, 0),..., (0, 1), (1, 1)) (2).
[0015] The derivation process of the directed difference convolution formula is as follows:
[0016] a. Based on the difference convolution formula, the difference convolution formula is as follows
[0017]
[0018] b. Use an anglelayer composed of a 7x7 CBR with an output channel of 16 and a 7X7 convolution with an output channel of one to convolve the original image to obtain an angle map; then use the tanh activation function to scale the angle map output by the angle layer to between [-pi / 2, pi / 2], so as to determine the rotation angle α of the directed difference convolution; after determining the angle α, use the rotation matrix R to determine the sampling points of the rotated convolution kernel, and the formula is as follows:
[0019]
[0020] c. Substitute the rotation matrix R into the standard convolution formula to obtain formula 4, and formula 4 is as follows:
[0021]
[0022] d. Use formula 4 to determine the receptive field direction and then rotate the image, and then use formula 3 to perform difference on the obtained rotated image, that is, obtain formula 1, that is, the directed difference convolution formula.
[0023] The preset parameter α = 1.2.
[0024] The target detection convolutional neural network is a general convolutional neural network.
[0025] The target detection convolutional neural network described above is a Yolov3 network, a Faster cnn network, or a CenterNet network.
[0026] The present invention designs a directed differential convolution. The directed differential convolution extracts local contrast-invariant contour features in an image with non-uniform contrast change through weight difference and self-adaptive direction change of the convolution kernel. By simulating the information processing mechanism of the primary visual cortex of the human brain, it can alleviate the performance degradation caused by domain shift of pictures by using illumination invariance. BPVENet can help the target detection algorithm adapt to different harsh conditions.
[0027] The method of the present invention can effectively alleviate the image domain shift problem under harsh conditions by means of contour information, and perform end-to-end training under different adverse conditions without a specific loss function. Since the neural network module provided in the method of the present invention is simple enough, it has little impact on the detection speed of the target detection network and does not destroy the real-time performance of the target detection network. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a schematic structural diagram of the neural network according to Embodiment 1 of the present invention;
[0029] Figure 2 is a visual comparison diagram of the target detection effect on the RTTS dataset in Embodiment 4;
[0030] Figure 3 is a visual comparison diagram of the target detection effect on the ExDark dataset in Embodiment 4. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The present invention will be specifically described below with reference to the drawings and embodiments.
[0032] Embodiment 1
[0033] The method for enhancing target detection feature extraction under harsh conditions described above includes the following steps:
[0034] A. Construct a neural network module, where the neural network module includes a DDCM module and a CBR module;
[0035] B. Input the original image into the neural network module and divide it into two paths; the first path is first processed by the DDCM module, and after the processing result of the DDCM module is multiplied by a preset parameter α, the first path result is obtained; the second path is processed by the CBR module to obtain the second path processing result; the first path processing result and the second path processing result are concatenated through the Concat function, and the obtained result is processed by the CBR module, and then processed by the SUM module with the processing result of the DDCM module to obtain the output result;
[0036] C. Input the output result into the target detection convolutional neural network for detection.
[0037] The DDCM module is a directed difference convolution, and the directed difference convolution formula is:
[0038]
[0039] Where W is the weight of the convolution kernel, R is the rotation matrix, and p 0 represents the position of the current convolution in the grid set P, and p' and p n are respectively the unknown and adjacent pixel grids in the same column in the grid set P.
[0040] The grid set P is:
[0041] P = ((-1, -1); (-1, 0),..., (0, 1), (1, 1)) (2).
[0042] The derivation process of the directed difference convolution formula is as follows:
[0043] a. Based on the difference convolution formula, the difference convolution formula is as follows
[0044]
[0045] b. Use an anglelayer composed of a 7x7 CBR with an output channel of 16 and a 7X7 convolution with an output channel of one to convolve the original image to obtain an angle map; then use the tanh activation function to scale the angle map output by the angle layer to between [-pi / 2, pi / 2], so as to determine the rotation angle α of the directed difference convolution; after determining the angle α, use the rotation matrix R to determine the sampling points of the rotated convolution kernel, and the formula is as follows:
[0046]
[0047] c. Substitute the rotation matrix R into the standard convolution formula to obtain formula 4, and formula 4 is as follows:
[0048]
[0049] d. Use formula 4 to determine the receptive field direction and then rotate the image, and then use formula 3 to perform difference on the obtained rotated image, that is, obtain formula 1, that is, the directed difference convolution formula.
[0050] The preset parameter α = 1.2.
[0051] Embodiment 2
[0052] Based on the method of Embodiment 1, the target detection convolutional neural network is the Faster cnn network.
[0053] Example 3
[0054] Based on the method of Example 1, the target detection convolutional neural network is the CenterNet network.
[0055] Example 4
[0056] For the performance evaluation of enhancing the enhancement effect of the network on the target detector, performance measurement criteria widely used in the field of target detection are adopted, and the specific evaluation is shown in Formulas (6) and (7).
[0057]
[0058] Among them, AP represents the integral of R (Recall) over P (Precision), the confidence threshold ranges from 0 to 1, mAP represents the average AP value of all classes in the dataset, and N is the number of classes. The higher the value of mAP, the stronger the detection performance of the model.
[0059] BPVENet: The neural network module of Example 1 is adopted, and the target detection convolutional neural network adopts the Yolov3 network.
[0060] Compare the detection results of BPVENet with the baseline target detector Yolov3, and the comparison results are shown in Table 1 below:
[0061] Table 1 summarizes the experimental data on the low-light dataset (ExDark) and the thick fog dataset (RTTS). Figure 2 and 3 Also respectively reflect the comparison results between BPVENet and Yolov3 network in this embodiment. It can be seen from Table 1 that the method of the embodiment has obvious improvement effects on the detection performance of the detector in both low-light and thick fog environments.
[0062] Table 1 Detection Comparison Results Table
[0063]
Claims
1. A method for enhancing feature extraction of target detection under harsh conditions, characterized in that: The following steps are involved: A. constructing a neural network module, wherein the neural network module includes a DDCM module and a CBR module; B. The original image is input into the neural network module and divided into two paths. The first path is processed by the DDCM module. The processing result of the DDCM module is multiplied by the preset parameter α to obtain the first path result. The second path is processed by the CBR module to obtain the second path processing result. The first processing result and the second processing result are concatenated by the Concat function, and the obtained result is processed by the CBR module, and then processed with the DDCM module processing result by the SUM module to obtain the output result; C. Input the output results into the target detection convolutional neural network for detection.
2. The method for enhancing feature extraction of target detection under harsh conditions as claimed in claim 1, characterized in that: The DDCM module is a directed differential convolution, and the directed differential convolution formula is: Among them, W is the weight of the convolution kernel, R is the rotation matrix, p0 represents the position of the current convolution in the grid set P, and p′ is n They are respectively the unknown pixel grids in the same column and adjacent to each other in the grid set P.
3. The method for enhancing feature extraction of target detection under harsh conditions as claimed in claim 2, characterized in that: The grid set P is: P=((-1,-1);(-1,0),...,(0,1),(1,1)) (2)。 4. The method for enhancing feature extraction of target detection under harsh conditions as claimed in claim 2, characterized in that: The derivation process of the directed differential convolution formula is: a. Based on the differential convolution formula, the differential convolution formula is as follows b. Convolve the original image with an angle layer consisting of a 7x7 CBR with an output channel of 16 and a 7X7 convolution with an output channel of 1 to obtain an angle map; then use the tanh activation function to scale the angle map output by the angle layer to between [-pi / 2,pi / 2] to determine the directed differential convolution rotation angle α; after determining the angle α, use the rotation matrix R to determine the sampling points of the rotated convolution kernel, the formula is as follows: c. Substitute the rotation matrix R into the standard convolution formula to obtain Formula 4, which is as follows: d. Use Formula 4 to determine the direction of the receptive field and then rotate the image. Then use Formula 3 to differentiate the resulting rotated image to obtain Formula 1, which is the directed differential convolution formula.
5. The method for target detection feature extraction and enhancement under harsh conditions as claimed in claim 1, characterized in that: The preset parameter α=1.
2.
6. The method for enhancing feature extraction of target detection under harsh conditions as claimed in claim 1, characterized in that: The target detection convolutional neural network is a general convolutional neural network.
7. The method for enhancing feature extraction of target detection under harsh conditions as claimed in claim 1, characterized in that: The target detection convolutional neural network is a Yolov3 network, a Faster cnn network, or a CenterNet network.