Target detection method and device based on image adaptive enhancement

By using image adaptive enhancement technology in the object detection system, using neural networks and filters to process images on foggy days, the problem of distortion and blurring of foggy days is solved, and the detection effect and robustness are improved.

CN120088577APending Publication Date: 2025-06-03XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510263617.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of image distortion and blurring in foggy scenes, affecting the performance of the target detection system.

Method used

The object detection method based on image adaptive enhancement is adopted to adaptively process images on foggy days through neural networks and multiple filters to enhance image quality and improve detection effect.

Benefits of technology

This method can perform adaptive processing according to the properties of the image itself, improve image quality, enhance detection effect, and make the model have strong robustness to both foggy and non-foggy images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088577A_ABST
    Figure CN120088577A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method and device based on image adaptive enhancement. The method comprises the following steps: acquiring a to-be-detected image; wherein the to-be-detected image is a foggy day image; performing down-sampling operation on the to-be-detected image to obtain a first image; processing the first image by adopting a trained neural network to obtain a plurality of parameters; assigning a plurality of filters to the plurality of parameters, and processing the to-be-detected image by using the filters to obtain an enhanced image; processing the enhanced image by using the trained image recognition network to obtain a classification result of the target in the to-be-detected image; wherein the trained image recognition network is obtained by taking data of a preset category as a training data set and training the initial image recognition network by adopting a loss function of the preset category. The target detection effect can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to a target detection method and a device based on image adaptive enhancement. Background Art

[0002] Object detection in foggy weather is a complex and challenging task in the field of computer vision, especially in the fields of autonomous driving, intelligent transportation systems (ITS), and security monitoring. Fog is formed by tiny water droplets or ice crystals suspended in the air, which scatter light and complicate the propagation path of light. This scattering effect leads to image degradation, which is manifested as reduced contrast, color deviation, and blurred edges of objects. Long-wavelength light, such as red light, can penetrate fog more easily than short-wavelength light. Therefore, photos taken in foggy conditions usually appear warm tones. In addition, in foggy conditions, visibility is significantly reduced, image contrast is reduced, colors are distorted, and details are blurred. These factors greatly affect the performance of vision-based object detection systems.

[0003] In order to overcome the above problems, there are mainly two solutions: foggy object detection based on traditional methods and foggy object detection based on deep learning. Among them, foggy object detection based on traditional methods mainly relies on image defogging, relies on physical models to estimate the atmospheric scattering effect, and attempts to restore clear images through defogging algorithms. However, the fog concentration in the actual environment is unpredictable and the lighting conditions are complex and changeable, making it difficult for traditional models based on fixed parameters to adapt to diverse scenarios. In addition, due to the different degrees of influence of fog on light of different wavelengths, the information obtained by relying solely on visible light cameras is limited, which also limits the effectiveness of traditional methods. In recent years, with the development of deep learning technology, advanced algorithms such as convolutional neural networks (CNNs) have been widely used in target detection. A variety of innovative solutions have been proposed for special situations in foggy days. For example, using adversarial generative networks (GANs) for image defogging preprocessing can generate fog-free images close to the actual situation, thereby improving the accuracy of subsequent target detection. At the same time, data enhancement technology has also been widely used. By simulating foggy environments of various concentrations, more robust detection models can be trained. In addition, some studies have also combined time series analysis to utilize the information association between consecutive frames. However, the above methods are difficult to solve problems such as image distortion and blur in foggy scenes.

[0004] Therefore, there is an urgent need to provide a target detection method and device thereof to improve the above-mentioned technical problems. Summary of the invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a target detection method and device based on image adaptive enhancement. The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0006] In a first aspect, the present invention provides an object detection method based on image adaptive enhancement, including:

[0007] Obtain an image to be detected; wherein, the image to be detected is a foggy image;

[0008] Perform a downsampling operation on the image to be detected to obtain a first image;

[0009] Process the first image using a trained neural network to obtain multiple parameters;

[0010] Assign the multiple parameters to multiple filters, and process the image to be detected using the filters to obtain an enhanced image;

[0011] Process the enhanced image using a trained image recognition network to obtain a classification result of the object in the image to be detected; wherein, the trained image recognition network is trained with data of a preset category as a training data set and using a loss function of a preset category to train an initial image recognition network.

[0012] In a second aspect, the present invention further provides an object detection device based on image adaptive enhancement, including:

[0013] A data acquisition module, configured to obtain an image to be detected; wherein, the image to be detected is a foggy image;

[0014] A first data processing module, configured to perform a downsampling operation on the image to be detected to obtain a first image;

[0015] A second data processing module, configured to process the first image using a trained neural network to obtain multiple parameters;

[0016] A data enhancement module, configured to assign the multiple parameters to multiple filters, and process the image to be detected using the filters to obtain an enhanced image;

[0017] A data recognition module, configured to process the enhanced image using a trained image recognition network to obtain a classification result of the object in the image to be detected; wherein, the trained image recognition network is trained with data of a preset category as a training data set and using a loss function of a preset category to train an initial image recognition network.

[0018] Advantages of the present invention:

[0019] A target detection method and device based on image adaptive enhancement provided by the present invention can process images whether in foggy days or non-foggy days according to the nature of the images themselves by setting an image adaptive module including a neural network and multiple filters, adaptively improve the image quality of the input images, enhance the detection effect, and thus make the model have strong robustness to foggy images and non-foggy images.

[0020] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flowchart of a target detection method based on image adaptive enhancement provided by an embodiment of the present invention;

[0022] Figure 2 is another flowchart of a target detection method based on image adaptive enhancement provided by an embodiment of the present invention;

[0023] Figure 3 is a schematic diagram of a simulation experiment comparison provided by an embodiment of the present invention;

[0024] Figure 4 is another schematic diagram of a simulation experiment comparison provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The present invention will be further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0026] Please refer to Figure 1 and Figure 2 , Figure 1 is a flowchart of a target detection method based on image adaptive enhancement provided by an embodiment of the present invention, Figure 2 is another flowchart of a target detection method based on image adaptive enhancement provided by an embodiment of the present invention. A target detection method based on image adaptive enhancement provided by the present invention includes:

[0027] S101. Obtain an image to be detected; wherein, the image to be detected is a foggy image.

[0028] Specifically, in this embodiment, the image to be detected is a real foggy image.

[0029] S102. Perform a downsampling operation on the image to be detected to obtain a first image.

[0030] Specifically, in this embodiment, a downsampling operation is performed on the image to be detected to reduce the data processing amount and build super noise.

[0031] S103. Process the first image using the trained neural network to obtain multiple parameters.

[0032] Specifically, in this embodiment, the image adaptation module includes a trained neural network and multiple cascaded filters.

[0033] Among them, the trained neural network includes five cascaded convolutional layers and one fully connected layer. The trained neural network is used to output the parameters of multiple cascaded filters. Processing the first image using the trained neural network to obtain multiple parameters includes:

[0034] Process the first image using five cascaded convolutional layers to extract the local features of the first image layer by layer and obtain a feature map.

[0035] Process the feature map using one fully connected layer to globally integrate the feature map and obtain multiple parameters.

[0036] In this embodiment, since the image filter is independent of the resolution, the input of the trained neural network is designed as a 256*256 image obtained by downsampling the original image, and then the hyperparameters of multiple cascaded filters are obtained. After the original image passes through multiple cascaded filters, the image quality will be improved.

[0037] S104. Assign multiple parameters to multiple filters, and process the image to be detected using the filters to obtain an enhanced image.

[0038] Specifically, in this embodiment, the filters include a defog filter (Defog), a white balance filter (WB), a gamma correction filter (Gamma), a tone filter (Tone), a contrast filter (Contrast), and a sharpening filter (Sharpen) cascaded in sequence. Processing the image to be detected using the filters to obtain an enhanced image includes:

[0039] Perform defogging processing on the image to be detected using the defog filter to remove the fog effect in the image to be detected.

[0040] Perform color correction on the defogged image using the white balance filter to eliminate the color cast caused by different light sources.

[0041] Perform brightness correction on the color-corrected image using the gamma correction filter to adjust the brightness and dark details of the image.

[0042] Perform tone correction on the brightness-corrected image using the tone filter to adjust the overall color style of the image.

[0043] Perform contrast correction on the tone-corrected image using the contrast filter to adjust the difference between the bright and dark parts of the image.

[0044] An edge information enhancement is performed on the contrast-corrected image using a sharpening filter to enhance the target edge contour, obtaining an enhanced image.

[0045] In this embodiment, the defogging filter can be regarded as the inverse process of fogging, and the expression of the defogging process of the defogging filter is:

[0046] J1(x) = (I1(x) - A(1 - t(x, ω))) / t(x, ω);

[0047] t(x, ω) = 1 - ω * ICA;

[0048] Wherein, J1(x) represents the image after defogging processing, I1(x) represents the image input to the defogging filter, A represents the global atmospheric light, t(x, ω) represents the medium projection function, ω represents one of multiple parameters, and ICA represents the point with the lowest pixel gray value of each channel of the image divided by A;

[0049] The expression of the color correction process of the white balance filter is:

[0050]

[0051] Wherein, P1 represents the image after color correction, Wr, Wg, and Wb represent three of multiple parameters, and represents the red, green, and blue values of the image input to the white balance filter;

[0052] The expression of the brightness correction process of the gamma correction filter is:

[0053]

[0054] Wherein, P2 represents the image after brightness correction, represents the image input to the gamma correction filter, and G represents one of multiple parameters;

[0055] The expression of the hue correction process of the hue filter is:

[0056]

[0057] Wherein, P4 represents the image after hue correction, tj ∈ [t1, t2,..., t7], and t1, t2,..., t7 represent some of multiple parameters, represents the image input to the hue filter;

[0058] The expression of the contrast correction process of the contrast filter is:

[0059]

[0060]

[0061] Among them, P3 represents the image after contrast correction, and represents the red, green, and blue values of the image input to the contrast filter, represents one of a plurality of parameters;

[0062] The expression for the process of enhancing the edge information of the sharpening filter is:

[0063] F(x,λ) = I(x) + λ(I(x) - Gau(x));

[0064] Among them, F(x,λ) represents the enhanced image, I2(x) represents the image input to the sharpening filter, λ represents one of a plurality of parameters, and Gau(·) represents the Gaussian filter.

[0065] S105. Use the trained image recognition network to process the enhanced image to obtain the classification result of the target in the image to be detected; among them, the trained image recognition network is trained with the data of the preset category as the training data set and the loss function of the preset category to train the initial image recognition network.

[0066] Specifically, in this embodiment, the trained image recognition network includes a trained backbone network, a trained neck network, and a trained detection head; using the trained image recognition network to process the enhanced image to obtain the classification result of the target in the image to be detected includes:

[0067] Use the trained backbone network to extract features from the enhanced image and increase the number of channels to obtain multi-scale features;

[0068] Use the trained neck network to perform feature fusion on the multi-scale features to obtain the fused multi-scale features;

[0069] Use the trained detection head to process the fused multi-scale features to obtain the classification results of the targets of different scales in the image to be detected.

[0070] In this embodiment, the trained backbone network includes multiple CBS modules, a spatial pyramid pooling module, and an EMA attention mechanism module; using the trained backbone network to extract features from the enhanced image and increase the number of channels to obtain multi-scale features includes:

[0071] Use multiple CBS modules to extract features from the enhanced image to obtain the first feature image;

[0072] The first feature image is fused using a spatial pyramid pooling module to improve the accuracy of object detection, and the fused first feature image is obtained.

[0073] The fused first feature image is weighted by channel or spatial information using an EMA attention mechanism module to dynamically adjust the feature weights, enhance important features and suppress noise, and multi-scale features are obtained.

[0074] In this embodiment, the training process of the trained image recognition network includes:

[0075] Obtain data of multiple preset categories.

[0076] The data of the preset categories are fogged to different degrees as the training dataset, and the true labels of the samples in the training dataset are obtained. Optionally, 8111 images including 5 types of objects, namely people, bicycles, cars, buses, and motorcycles, in the VOC2007 and VOC2012 datasets are used as samples in the training dataset.

[0077] The data of the preset categories are fogged to different degrees as a part of the test dataset, the data of the preset categories are used as another part of the test dataset, and the real foggy data are used as another part of the test dataset, and the true labels of the samples in the test dataset are obtained. Optionally, 2734 images including 5 types of objects, namely people, bicycles, cars, buses, and motorcycles, in the VOC2007 dataset are used as a part of the test dataset; fogging is performed to different degrees using a fogging expression as another part of the test dataset; the publicly available real foggy dataset is used as another part of the test dataset, with 4322 images, including 5 types of objects, namely people, bicycles, cars, buses, and motorcycles, and each category has more than 500 examples.

[0078] Part of the samples in the training dataset are input into the p-th image recognition network to be trained for training, and the prediction results output during the p-th training process are obtained.

[0079] According to the prediction results output during the p-th training process and the true labels of the samples for training the p-th image recognition network to be trained, the loss is calculated and used as the loss of the p-th training process.

[0080] Backpropagation is performed according to the loss of the p-th training process to update the network parameters of the p-th image recognition network to be trained, and the (p + 1)-th image recognition network to be trained is obtained; iterate in this way until the number of training times or the convergence degree meets the preset conditions, and the trained image recognition network is obtained.

[0081] The test data set is input into the trained image recognition network for processing to obtain a prediction result, and the prediction result is compared with the true label in the test data set to judge the performance of the trained image recognition network.

[0082] In this embodiment, the EMA attention mechanism reshapes some channels into the batch dimension and groups the channel dimension into multiple sub-features to retain the information of each channel and reduce the computational overhead. The EMA attention mechanism recalibrates the channel weights in each parallel branch by encoding global information and captures pixel-level relationships through cross-dimensional interactions, thereby helping the neural network improve its feature extraction ability.

[0083] The input of the EMA attention mechanism is a feature map with a shape of C*H*W, where C is the number of channels, H is the height, and W is the width. The channel dimension of the input feature map is grouped into multiple sub-features, and the number of channels of each sub-feature is C / g, where g is the number of groups. The EMA attention mechanism includes two parallel branches, one path is the 1*1 branch, and one path is the 3*3 branch. The 1*1 branch uses one-dimensional global average pooling operations to encode channel information in two spatial directions respectively. The 1*1 branch uses 1x1 convolution to model the channel information, capture global information and recalibrate the channel weights, and extracts features in the height and width directions respectively through adaptive average pooling. The 3*3 branch uses 3*3 convolution to capture local features and enhance the aggregation of multi-scale spatial structure information.

[0084] The output features of the two parallel branches are aggregated together through cross-dimensional interactions to capture pixel-level pairwise relationships. The EMA attention mechanism performs adaptive average pooling and Softmax operations on the features of the two branches respectively, then calculates the weights using matrix multiplication, and applies the weights to the original input feature map.

[0085] In this embodiment, the expression for performing fogging processing on data of a preset category to different degrees is:

[0086] I(x) = J(x)e-βd(x)+A(1 - e-βd(x));

[0087]

[0088] where I(x) represents the image generated by the fogging processing, J(x) represents the data of the preset category, ρ represents the Euclidean distance from the current pixel to the central pixel, row and col respectively represent the number of rows and columns of the image, i is a hyperparameter from 0 to 9, 10 different levels of fog can be added to the image, β = 0.01*i + 0.05, A represents a hyperparameter, and A = 0.5.

[0089] In this embodiment, SSIM is a structural measurement index that can measure the similarity between two images. SSIM measures the similarity between two images from three aspects: brightness, contrast, and structure. Its value ranges from -1 to 1, and the larger the value, the more similar the two images are.

[0090] The expression for calculating the loss is:

[0091] L = SSIM(x, y) + Lloss

[0092] SSIM(x, y) = l(x, y) * c(x, y) * s(x, y);

[0093]

[0094] Among them, l(x, y) represents a function for measuring the brightness of two images, c(x, y) represents a function for measuring the contrast of two images, σx represents the standard deviation of the gray levels of the pixels of one of the two estimated images, σy represents the standard deviation of the gray levels of the pixels of the other of the two estimated images, s(x, y) represents a function for measuring the contrast structure of two images, σxy represents the correlation coefficient, xn represents the gray level value of the pixels of one of the two images, yn represents the gray level value of the pixels of the other of the two images, μx represents the average gray level of the pixels of one of the two images, μy represents the average gray level of the pixels of the other of the two images, C1, C2, and C3 represent arbitrary constants, N represents the total number of pixel points in the corresponding image, n represents the pixel point index, and Lloss represents the loss function of the image recognition network.

[0095] Since the value of SSIM ranges from -1 to 1 and the larger its value, the more similar the two images are, the SSIM loss function is generally defined as 1 - SSIM.

[0096] It should be noted that during the training process, when loading images, there is a two-thirds probability of adding fog to the images (according to the above-mentioned fog-adding process with different degrees). It is necessary to normalize and perform image enhancement (including rotation, translation, and scaling, etc.) on the samples in the training dataset and the test dataset to enhance the generalization ability of the model. The Adam optimizer is used as the optimizer, with the initial learning rate set to 0.0001 and the momentum set to 0.937, which helps to accelerate convergence, and the preset number of iterations is 80 rounds.

[0097] In summary, a target detection method based on image adaptive enhancement provided by the present invention has the following beneficial effects:

[0098] First, with the help of the image adaptive enhancement module, the present invention adaptively enhances the input image to improve the detection accuracy. Traditional object detection models generally adopt general image enhancement techniques. The main function of these techniques is to make the model more robust and prevent overfitting. However, the main problem faced by object detection in foggy images is underfitting. Therefore, the role of general image enhancement techniques is limited. The image adaptive enhancement module includes a deep neural network and multiple image filters, enabling it to process images whether they are foggy or non-foggy according to the nature of the image itself, adaptively improve the image quality of the input image, enhance the detection effect, and thus make the model have strong robustness to both foggy and non-foggy images.

[0099] Second, adding the SSIM loss to the original YOLOv5 loss function in the present invention can greatly increase the robustness of the model for image detection. When the SSIM loss function is not used, although the image adaptive module can generally improve the image quality, it will also cause the image to be distorted in some details. After adding the SSIM loss function, this defect is basically solved, enabling the image adaptive enhancement module to better improve the quality of the input image.

[0100] Third, the present invention uses the EMA attention mechanism, which has significant advantages in detecting relatively blurred targets. Aiming at the problem that it is difficult to detect and recognize blurred images, the EMA attention mechanism can re-calibrate and weight the weights of each channel by encoding global information, retain the more important information in the feature layer, remove redundant information, and thus perform feature extraction more effectively and improve the detection accuracy.

[0101] In an optional embodiment of the present invention, the effect of the object detection method based on image adaptive enhancement provided in the above embodiment is verified through a simulation experiment, specifically:

[0102] 1. Simulation conditions

[0103] The hardware platform for the simulation experiment of the present invention is: the processor is an Intel i7 10700F CPU with a main frequency of 2.90GHz and 16GB of memory.

[0104] The software platform for the simulation experiment of the present invention is: Windows 10 operating system, python3.8, PyTorch1.13.0.

[0105] 2. Simulation content and result analysis

[0106] The simulation experiment of the present invention uses the present invention and an existing technology (the original YOLOV5 network) to detect the input foggy images respectively.

[0107] The effect of the present invention will be further described below with reference to the simulation diagram.

[0108] Please refer to Figure 3 and Figure 4 , Figure 3 which is a schematic diagram for comparison of simulation experiments provided by an embodiment of the present invention, Figure 4 and which is another schematic diagram for comparison of simulation experiments provided by an embodiment of the present invention, Figure 3 The image in the upper left is an image of the daily dataset used for detection in the simulation experiment of the present invention, Figure 3 the image in the upper right is an image of the constructed dataset used for detection, Figure 3 The image in the lower left is the image of the present invention without using the SSIM loss function and with the input being Figure 3 the image enhanced by the image adaptive enhancement module in the upper right. Figure 3 The image in the lower right is the image of the present invention using the SSIM loss function and with the input being Figure 3 the image enhanced by the image adaptive enhancement module in the upper right.

[0109] It can be seen from Figure 3 that the image adaptive enhancement module can effectively enhance the image quality of foggy images. However, when not combined with the SSIM loss function, some details of the image will become blurred, while when combined with the SSIM function, its ability to improve the image quality is greatly enhanced.

[0110] Figure 4 The image in the upper left is an image of the real foggy dataset used for detection in the simulation experiment of the present invention, Figure 4 the image in the upper right is the improved quality image after passing through the image enhancement module of the improved YOLOV5 network in the simulation experiment of the present invention, Figure 4 the image in the lower left is the detection result image using the original YOLOV5 network in the simulation experiment of the present invention, Figure 4 and the image in the lower right is the detection result image using the improved YOLOV5 network in the simulation experiment of the present invention.

[0111] It can be seen from Figure 4 that the detection accuracy of the original YOLOV5 network for foggy images is relatively low. Compared with the original image, the image quality of the image after image adaptive enhancement is better, and the subsequent detection effect is also better.

[0112] The detection results of the two networks are evaluated using mAP50 respectively. Using the previous formula, mAP50 is calculated, and all the calculation results are plotted in Table 1.

[0113] mAP is one of the commonly used performance indicators in the object detection task and is used to evaluate the detection performance of the model on multiple categories.

[0114]

[0115] Calculating the Precision-Recall Curve: For each class, the precision-recall curve is obtained by calculating precision and recall at different confidence thresholds. For each class, the average precision is obtained by calculating the integral (or the area between discrete points) under the precision-recall curve. Interpolation methods such as 11-point interpolation or smoother interpolation methods can be used.

[0116] Taking the average of the average precisions of all classes gives mAP, and its expression is:

[0117]

[0118] where N is the number of classes, and APi is the average precision of the i-th class.

[0119] Combined with Table 1, it can be seen that the mAP values of the present invention are all higher than those of the original YOLOV5 network, proving that the present invention can obtain higher detection accuracies for foggy and conventional images. Combined with Table 2, it can be seen that adding an image adaptive module to YOLOv5 can effectively improve the detection accuracies of the three datasets; adding the SSIM loss function can improve the detection accuracies of the first and second datasets on the premise of hardly losing the RTTS detection accuracy, indicating that the model has improved the ability to detect conventional images; while adding the EMA attention mechanism improves the feature extraction ability of the model, and the detection accuracies of the three datasets are all improved.

[0120] Table 1 Quantitative analysis table of detection results on the public dataset in the simulation experiment

[0121]

[0122] Table 2 Ablation experiment

[0123]

[0124]

[0125] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant are intended to cover non-exclusive inclusion, so that an article or device comprising a series of elements includes not only those elements but also other elements not expressly listed. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the article or device comprising the element. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The orientation or positional relationship indicated by "above", "below", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation on the present invention.

[0126] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0127] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A target detection method based on image adaptive enhancement, characterized in that: include: Acquire an image to be detected; wherein the image to be detected is a foggy image; Performing a downsampling operation on the image to be detected to obtain a first image; Processing the first image using a trained neural network to obtain multiple parameters; Assigning the multiple parameters to multiple filters, and using the filters to process the image to be detected to obtain an enhanced image; The enhanced image is processed using a trained image recognition network to obtain a classification result of the target in the image to be detected; wherein the trained image recognition network uses data of a preset category as a training data set and adopts a loss function of a preset category to train the initial image recognition network.

2. The target detection method based on image adaptive enhancement according to claim 1, characterized in that: The trained neural network includes five cascaded convolutional layers and one fully connected layer; the trained neural network is used to process the first image to obtain multiple parameters, including: Processing the first image using a cascade of five convolutional layers, extracting local features of the first image layer by layer, and obtaining a feature map; The feature map is processed by using the one fully connected layer, and the feature map is globally integrated to obtain multiple parameters.

3. The target detection method based on image adaptive enhancement according to claim 1, characterized in that: The filter includes a defogging filter, a white balance filter, a gamma correction filter, a hue filter, a contrast filter and a sharpening filter which are cascaded in sequence; the filter is used to process the image to be detected to obtain an enhanced image, including: Using the defogging filter to perform defogging processing on the image to be detected, so as to remove the fog effect in the image to be detected; The white balance filter is used to perform color correction on the dehazed image to eliminate color cast caused by different light sources; Using the gamma correction filter to perform brightness correction on the color-corrected image, adjusting the brightness and dark details of the image; Using the tone filter to perform tone correction on the brightness-corrected image to adjust the overall color style of the image; Using the contrast filter to perform contrast correction on the tone-corrected image to adjust the difference between the bright and dark parts of the image; The sharpening filter is used to enhance edge information of the contrast-corrected image, thereby enhancing the target edge contour and obtaining the enhanced image.

4. The target detection method based on image adaptive enhancement according to claim 3 is characterized in that: The expression of the defogging process of the defogging filter is: J1(x)=(I1(x)-A(1-t(x,ω))) / t(x,ω); Wherein, J1(x) represents the image after dehazing, I1(x) represents the image input to the dehazing filter, A represents the global atmospheric light, t(x,ω) represents the medium projection function, and ω represents one of the multiple parameters; The expression of the color correction process of the white balance filter is: Wherein, P1 represents a color-corrected image, Wr, Wg, and Wb represent three of the multiple parameters, and Represents the red, green and blue values ​​of the image input to the white balance filter; The expression of the brightness correction process of the gamma correction filter is: Among them, P2 represents the image after brightness correction, represents the image input to the gamma correction filter, G represents one of a plurality of parameters; The expression of the tone correction process of the tone filter is: Where P4 represents the image after tone correction, tj∈[t1,t2,…,t7], t1,t2,…,t7 represents part of multiple parameters, represents the image input to the hue filter; The expression of the contrast correction process of the contrast filter is: Among them, P3 represents the image after contrast correction, and represents the red, green and blue values ​​of the image input to the contrast filter, Represents one of multiple parameters; The expression of the edge information enhancement process of the sharpening filter is: F(x,λ)=I(x)+λ(I(x)-Gau(x)); Wherein, F(x,λ) represents the enhanced image, I2(x) represents the image input to the sharpening filter, λ represents one of a plurality of parameters, and Gau(·) represents a Gaussian filter.

5. The target detection method based on image adaptive enhancement according to claim 1, characterized in that: The trained image recognition network includes a trained backbone network, a trained neck network and a trained detection head; The step of using a trained image recognition network to process the enhanced image to obtain a classification result of the target in the image to be detected includes: Using the trained backbone network to extract features from the enhanced image, and increasing the number of channels to obtain multi-scale features; Using the trained neck network to perform feature fusion on the multi-scale features to obtain fused multi-scale features; The trained detection head is used to process the fused multi-scale features to obtain classification results of objects of different scales in the image to be detected.

6. The target detection method based on image adaptive enhancement according to claim 5, characterized in that: The trained backbone network includes multiple CBS modules, spatial pyramid pooling modules and EMA attention mechanism modules; The method of using the trained backbone network to extract features from the enhanced image and increasing the number of channels to obtain multi-scale features includes: Using the multiple CBS modules to extract features from the enhanced image to obtain a first feature image; The spatial pyramid pooling module is used to fuse the first feature image to improve the accuracy of target detection and obtain a fused first feature image; The EMA attention mechanism module is used to weight the channel or spatial information of the fused first feature image, dynamically adjust the feature weights, enhance important features and suppress noise, and obtain multi-scale features.

7. The target detection method based on image adaptive enhancement according to claim 1, characterized in that: The training process of the trained image recognition network includes: Acquire data of a plurality of preset categories; Perform different degrees of fogging on the data of the preset categories as training data sets, and obtain the true labels of the samples in the training data sets; Performing fogging processing of different degrees on the data of the preset category as a part of the test data set, taking the data of the preset category as another part of the test data set, taking real foggy day data as another part of the test data set, and obtaining the real labels of the samples in the test data set; Inputting some samples in the training data set into the image recognition network to be trained for the pth time for training, and obtaining the prediction result output during the pth training process; According to the prediction results output during the p-th training process and the true labels of the samples of the image recognition network to be trained for the p-th time, the loss is calculated and used as the loss of the p-th training process; Back propagation is performed according to the loss of the p-th training process to update the network parameters of the image recognition network to be trained for the p-th time, and the image recognition network to be trained for the p+1-th time is obtained; and the trained image recognition network is obtained by iterating in this way until the number of training times or the degree of convergence meets the preset conditions; The test data set is input into the trained image recognition network for processing to obtain a prediction result, and the prediction result is compared with the real label in the test data set to judge the performance of the trained image recognition network.

8. The target detection method based on image adaptive enhancement according to claim 7, characterized in that: The expressions for performing different degrees of fogging processing on the preset category of data are as follows: I ( x ) =J ( x ) e-βd ( x ) +A(1-e-βd ( x ) ); Among them, I(x) represents the image generated by fogging, J(x) represents the data of the preset category, ρ represents the Euclidean distance from the current pixel to the center pixel, row and col represent the number of rows and columns of the image respectively, i represents a hyperparameter from 0 to 9, β = 0.01*i+0.05, and A represents a hyperparameter.

9. The target detection method based on image adaptive enhancement according to claim 7, characterized in that: The expression for calculating the loss is: L = SSIM(x,y) + Lloss SSIM(x,y)=l(x,y)*c(x,y)*s(x,y); Among them, l(x,y) represents a function that measures the brightness of two images, c(x,y) represents a function that measures the contrast of two images, σx represents the estimated grayscale standard deviation of the pixels of one of the two images, σy represents the estimated grayscale standard deviation of the pixels of the other of the two images, s(x,y) represents a function that measures the contrast structure of the two images, σxy represents the correlation coefficient, xn represents the grayscale value of the pixels of one of the two images, yn represents the grayscale value of the pixels of the other of the two images, μx represents the average grayscale value of the pixels of one of the two images, μy represents the average grayscale value of the pixels of the other of the two images, C1, C2 and C3 represent arbitrary constants, N represents the total number of pixels in the corresponding image, n represents the pixel index, and Lloss represents the loss function of the image recognition network.

10. A target detection device based on image adaptive enhancement, characterized in that: include: A data acquisition module, used to acquire an image to be detected; wherein the image to be detected is a foggy image; A data processing module 1 is used to perform a downsampling operation on the image to be detected to obtain a first image; A second data processing module is used to process the first image using a trained neural network to obtain multiple parameters; A data enhancement module, used for assigning the multiple parameters to multiple filters, and using the filters to process the image to be detected to obtain an enhanced image; A data recognition module is used to process the enhanced image using a trained image recognition network to obtain a classification result of the target in the image to be detected; wherein the trained image recognition network uses data of a preset category as a training data set and adopts a loss function of a preset category to train the initial image recognition network.