Lightweight target detection method and device based on image defogging and electronic equipment
The foggy day images are processed through the depth separation convolution and feature fusion model, which solves the accuracy and real-time problems of traffic target detection in foggy environments, and realizes the efficient operation of lightweight target detection.
Patent Information
- Application Number
- CN202510483392.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
AI Technical Summary
The existing traffic target detection methods have low accuracy, complex calculations and are not real-time in foggy environments, making it difficult to meet the application needs of resource-constrained equipment.
The lightweight object detection method based on image defogging is adopted, and the images are defogged and detected by the depth-separable convolutional sub-model, feature extraction sub-model and feature fusion sub-model. The AOD network and the CSPDarknet-53 model are used for feature extraction and fusion.
It improves the accuracy and robustness of target detection, reduces the calculation complexity and parameter quantity, enhances the efficiency and real-timeness of target detection, and is suitable for traffic monitoring in complex environments.
Smart Images

Figure CN119992112A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of target detection, and in particular, relates to a lightweight target detection method based on image defogging. Background Art
[0002] Traffic monitoring, as an important part of modern traffic management, plays a vital role in ensuring traffic safety and smooth operation. Traffic target detection, as a core technology, can accurately identify traffic targets such as vehicles, pedestrians, and traffic signs, and provide important data support for traffic flow control, illegal behavior determination, and accident prevention. However, bad weather, especially foggy scenes, poses a huge challenge to traffic target detection. Tiny particles in the fog scatter and absorb light, resulting in a significant decrease in the quality of traffic images, which is manifested in problems such as reduced image contrast, color deviation, and blurred details.
[0003] Traditional traffic target detection methods are difficult to maintain high accuracy in foggy environments. For example, vehicle outlines are difficult to clearly define, traffic sign information is difficult to effectively read, and pedestrian features are difficult to accurately capture. Traditional deep learning-based models are insufficient in the ability to extract features from foggy images, with high computational costs and long processing times, making it difficult to meet real-time requirements. In addition, traditional image defogging algorithms are computationally complex and have limited adaptability. The combination of deep learning defogging and detection models is not efficient enough, the training process is cumbersome, and there are many parameters, which is not conducive to deployment and application in actual traffic monitoring equipment with limited resources.
[0004] In the related technologies, the target detection algorithm in foggy scenes still has certain limitations. It is difficult to fully reflect the image features in complex environments, and there are problems such as low accuracy and complex calculations. Summary of the invention
[0005] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a lightweight object detection method based on image defogging, which improves the accuracy and efficiency of object detection.
[0006] In a first aspect, the present application provides a lightweight object detection method based on image defogging, the method comprising: Acquire an original image, where the original image is acquired based on acquisition equipment; Inputting the original image into a target detection model to obtain a target image, wherein the target detection model is trained based on a sample image, and the target detection model includes a depth separable convolution sub-model, a feature extraction sub-model, and a feature fusion sub-model; Perform detection based on the target image to obtain a target detection result; Among them, the depth-separable convolution sub-model includes a first depth-separable convolution layer, a second depth-separable convolution layer, a third depth-separable convolution layer, a fourth depth-separable convolution layer, a fifth depth-separable convolution layer, a first feature fusion layer, a second feature fusion layer, a third feature fusion layer and a fourth feature fusion layer, the feature extraction sub-model includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer, and the feature fusion sub-model includes a fifth feature fusion layer, a sixth feature fusion layer, a seventh feature fusion layer, an eighth feature fusion layer, a first upsampling layer, a second upsampling layer, a first downsampling layer and a second downsampling layer.
[0007] According to one embodiment of the present application, inputting the original image into a target detection model to obtain a target image includes: Inputting the original image into the depth-separable convolution sub-model to obtain a first image, where the first image is a defogged image; Inputting the first image into the feature extraction sub-model to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map; The first feature map, the second feature map, the third feature map and the fourth feature map are input into the feature fusion sub-model to obtain the target image.
[0008] According to one embodiment of the present application, inputting the original image into the depth-separable convolution sub-model to obtain a first image includes: The original image is input into the first depth-separable convolutional layer to obtain the feature Figure 1 ; The characteristics Figure 1 Input into the second depth separable convolutional layer to obtain the feature Figure 2 ; The characteristics Figure 1 and the characteristics Figure 2 Input to the first feature fusion layer to obtain the first feature tensor; The first feature tensor is input into the third depth-separable convolutional layer to obtain the feature Figure 3 ; The characteristics Figure 2 and the characteristics Figure 3 Input to the second feature fusion layer to obtain the second feature tensor; The second feature tensor is input into the fourth depth-separable convolutional layer to obtain the feature Figure 4 ; The characteristics Figure 1 , the characteristics Figure 2 , the characteristics Figure 3 and the characteristics Figure 4Input to the third feature fusion layer to obtain the third feature tensor; The third feature tensor is input into the fifth depth-separable convolutional layer to obtain the feature Figure 5 ; The characteristics Figure 5 and the original image are input into the fourth feature fusion layer to obtain the first image.
[0009] According to one embodiment of the present application, the depthwise separable convolution layer includes a depthwise convolution layer and a pointwise convolution layer.
[0010] According to one embodiment of the present application, the expression of the depth convolution layer is as follows:
[0011] Among them, n represents the number of samples in the batch, c represents the channel number, h represents the position number in the height direction, w represents the position number in the width direction, and K d represents the depth convolution kernel, S d represents the step size, i represents the index of the depth convolution kernel in the height direction, P d represents padding, j represents the index of the depth convolution kernel in the width direction, It means that the elements at the corresponding positions are selected from the input feature map X according to the step size and padding to participate in the convolution calculation. is the element of the deep convolution output feature map at the corresponding position, Represents the weight value of the depth convolution kernel at the corresponding index.
[0012] According to one embodiment of the present application, the expression of the point-by-point convolution layer is as follows:
[0013] in, is the element value of the point-by-point convolution output feature map at the corresponding position, n represents the sequence number of the sample in the batch, c represents the channel sequence number, m traverses all channels of the depth convolution output, h represents the position sequence number in the height direction, and w represents the position sequence number in the width direction. Represents the point-by-point convolution kernel The weight value from input channel m to output channel c in, is the deep convolution output feature map, The value of the element at the corresponding index (sample n, channel m, height h, width w).
[0014] According to an embodiment of the present application, inputting the first feature map, the second feature map, the third feature map, and the fourth feature map into the feature fusion sub-model to obtain the target image includes: The fourth feature map is input into the first upsampling layer and the third feature map is input into the fifth feature fusion layer to obtain a fusion feature map. Figure 3 ; The fusion feature Figure 3 After inputting the second upsampling layer, the second feature map is input into the sixth feature fusion layer to obtain the fused feature Figure 2 ; The fusion feature Figure 2 After inputting the first downsampling layer and the fusion feature Figure 3 Input to the seventh feature fusion layer to obtain the re-fusion feature Figure 3 ; The reintegration feature Figure 3 After being input into the second downsampling layer, the fourth feature map is input into the eighth feature fusion layer to obtain the target image.
[0015] According to one embodiment of the present application, the training process of the target detection model includes: Construct a preset target detection model, wherein the preset target detection model is a CSPDarknet-53 model; Obtaining a data set based on the sample image; Based on the first loss function, the preset target detection model is trained according to the data set to obtain the target detection model.
[0016] In a second aspect, the present application provides a lightweight object detection device based on image defogging, the device comprising: An acquisition module is used to acquire an original image, where the original image is acquired based on acquisition equipment; A processing module, used for inputting the original image into a target detection model to obtain a target image, wherein the target detection model is obtained based on sample image training, and the target detection model includes a depth separable convolution sub-model, a feature extraction sub-model and a feature fusion sub-model; A detection module, used to perform detection based on the target image to obtain a target detection result; Among them, the depth-separable convolution sub-model includes a first depth-separable convolution layer, a second depth-separable convolution layer, a third depth-separable convolution layer, a fourth depth-separable convolution layer, a fifth depth-separable convolution layer, a first feature fusion layer, a second feature fusion layer, a third feature fusion layer and a fourth feature fusion layer, the feature extraction sub-model includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer, and the feature fusion sub-model includes a fifth feature fusion layer, a sixth feature fusion layer, a seventh feature fusion layer, an eighth feature fusion layer, a first upsampling layer, a second upsampling layer, a first downsampling layer and a second downsampling layer.
[0017] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the lightweight target detection method based on image dehazing as described in the first aspect above is implemented.
[0018] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the lightweight target detection method based on image dehazing as described in the first aspect above.
[0019] In a fifth aspect, the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the lightweight target detection method based on image dehazing as described in the first aspect.
[0020] In a sixth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the lightweight target detection method based on image defogging as described in the first aspect above.
[0021] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application.
[0022] The lightweight object detection method based on image defogging provided by the present invention has the following beneficial effects compared with the prior art: (1) The present invention obtains the original image and inputs it into the target detection model for processing. It combines the deep separable convolution sub-model, feature extraction sub-model and feature fusion sub-model to perform defogging on the image before detection, thereby improving the image quality. By extracting important feature information from the original image and fusing it to obtain the target image, the target image is detected and the target in the image can be identified and located, thereby improving the accuracy and robustness of target detection. Through the collaborative work of the deep separable convolution sub-model, feature extraction sub-model and feature fusion sub-model, the computational complexity and parameter quantity of the model are reduced, the efficiency and real-time performance of target detection are enhanced, the effective combination of image defogging and target detection is achieved, and the applicability and reliability of the target detection system in complex environments are improved.
[0023] (2) The present invention can effectively defog the original image by inputting it into a deep separable convolution sub-model for defogging, thereby obtaining a clear first image. This enables the model to focus more on the extraction and utilization of key features when processing the defogged image, thereby enhancing the ability to recognize and locate traffic targets (vehicles, pedestrians, etc.) of different scales in foggy environments, effectively improving the accuracy of target detection, and providing more reliable data support for functions such as violation monitoring and traffic flow statistics in traffic management, thereby improving the defect of poor performance in traffic target detection under severe weather conditions.
[0024] (3) The present invention can effectively defog the image and extract rich image feature information by inputting the original image into a deep separable convolution sub-model and performing multi-level feature extraction and fusion processing. By processing multiple deep separable convolution layers, the detailed features of the image are gradually extracted, and the multi-level information is integrated through the feature fusion layer to obtain a clearer first image. The performance of the traffic target detection task after image defogging and the efficiency of the traffic monitoring system in severe weather are improved, meeting the application requirements in resource-constrained scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which: Figure 1 This is one of the flow charts of the lightweight target detection method based on image defogging provided in the embodiment of the present application; Figure 2 This is the second flow chart of the lightweight target detection method based on image defogging provided in the embodiment of the present application; Figure 3 This is the third flow chart of the lightweight target detection method based on image defogging provided in the embodiment of the present application; Figure 4 This is the fourth flow chart of the lightweight object detection method based on image defogging provided in the embodiment of the present application; Figure 5 It is a structural schematic diagram of a lightweight object detection device based on image defogging provided in an embodiment of the present application; Figure 6 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0027] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0028] In the following, in combination with the accompanying drawings, the lightweight target detection method based on image defogging, the lightweight target detection device based on image defogging, the electronic device and the readable storage medium provided in the embodiments of the present application are described in detail through specific embodiments and their application scenarios.
[0029] Among them, the lightweight target detection method based on image defogging can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.
[0030] The terminal includes, but is not limited to, a portable communication device such as a mobile phone or a tablet computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad). It should also be understood that in some embodiments, the terminal may not be a portable communication device, but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad).
[0031] In the following various embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse and a joystick.
[0032] The lightweight target detection method based on image defogging provided in the embodiment of the present application may be executed by an electronic device or a functional module or functional entity in the electronic device that can implement the lightweight target detection method based on image defogging. The electronic devices mentioned in the embodiment of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, wearable devices, etc. The lightweight target detection method based on image defogging provided in the embodiment of the present application is described below using an electronic device as an example of the execution subject.
[0033] Figure 1 This is one of the flow charts of the lightweight target detection method based on image defogging provided in the embodiment of the present application, such as Figure 1 As shown, the lightweight target detection method based on image defogging includes: step 110, step 120 and step 130.
[0034] Step 110: Acquire an original image, where the original image is acquired based on acquisition equipment; Taking the traffic monitoring system as an example, the acquisition device captures images or videos on the traffic road in real time and transmits them to the electronic device. The electronic device uses the acquired image as the original image, or processes the video frame by frame and extracts each frame as the original image for analysis.
[0035] Optionally, the acquisition device may be an infrared camera, a thermal imaging camera, a pan-tilt camera, or the like.
[0036] Step 120: input the original image into a target detection model to obtain a target image, wherein the target detection model is trained based on a sample image, and the target detection model includes a depthwise separable convolution sub-model, a feature extraction sub-model, and a feature fusion sub-model; Furthermore, the original image is input into the target detection model for detection to obtain the target image. The target detection model is trained based on the sample image.
[0037] It is easy to understand that the target detection model includes a depth-separable convolution sub-model, a feature extraction sub-model and a feature fusion sub-model. The depth-separable convolution sub-model includes a first depth-separable convolution layer, a second depth-separable convolution layer, a third depth-separable convolution layer, a fourth depth-separable convolution layer, a fifth depth-separable convolution layer, a first feature fusion layer, a second feature fusion layer, a third feature fusion layer and a fourth feature fusion layer. The feature extraction sub-model includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer. The feature fusion sub-model includes a fifth feature fusion layer, a sixth feature fusion layer, a seventh feature fusion layer, an eighth feature fusion layer, a first up-sampling layer, a second up-sampling layer, a first down-sampling layer and a second down-sampling layer.
[0038] It should be noted that the target detection model can be an AOD-CSPDarknet-53 (Attention-Oriented Detection-Cross-Stage Partial Darknet-53) model, wherein the deep separable convolution sub-model is an AOD network (Attention-Oriented Detection Network), which automatically focuses on important parts of the image through an attention mechanism, thereby improving detection accuracy and reducing interference from irrelevant areas. The feature extraction sub-model and the feature fusion sub-model are CSPDarknet-53 models, which are a deep convolutional neural network based on the CSPNet (Cross-Stage Partial Networks) architecture, including an input layer, a convolutional layer, a CSP residual module, and feature fusion.
[0039] It is easy to understand that the target detection model is trained based on several collected sample images. During the training process of the target detection model, the traditional model depth adjustment often adopts a simple method of proportionally deleting or copying the network layer. This application adopts an adaptive depth adjustment algorithm. On the basis of maintaining the CSPDarknet-53 framework, according to the importance and information redundancy of different feature extraction stages, in the shallow part of the network, the network layers are selectively merged and streamlined. By designing a layer merging rule based on feature similarity measurement, adjacent convolutional layers with similar feature extraction effects are merged, which reduces the network depth while avoiding excessive information loss.
[0040] In one embodiment, let the original network layer sequence be , where m is the number of original network layers, and the cosine similarity of the feature outputs of adjacent layers is calculated , when the similarity exceeds the set threshold θd, the two layers are merged into a new convolutional layer lnew, whose parameters are obtained by weighted average of the parameters of the original two layers, and the new network layer sequence becomes , where k is the logarithm of the number of merged layers, which realizes the adaptive adjustment of depth, reduces the amount of calculation, and improves the operation efficiency of the network. Let D be the original network depth, D' be the adjusted network depth, and k be the logarithm of the number of merged layers, then , the formula for calculating the parameters of the new layer when the layers are merged is as follows:
[0041]
[0042] Among them, li is the merged i-th layer, l i+1 is the merged i+1th layer, Sim is the cosine similarity calculation, W i is the weight of the i-th layer, b i is the bias of the i-th layer, W i+1 is the weight of the i+1th layer, b i is the bias of the i-th layer, W new is the weight parameter of the new layer, b new is the bias for the new layer.
[0043] Step 130: Perform detection based on the target image to obtain a target detection result; Finally, the target image is detected by sliding a fixed-size window on the target image, gradually scanning each position to determine whether the window contains the target and obtain the target detection result.
[0044] According to the lightweight target detection method based on image defogging provided by the embodiment of the present application, by acquiring the original image and inputting it into the target detection model for processing, combined with the deep separable convolution sub-model, feature extraction sub-model and feature fusion sub-model, the image can be defogged before detection, thereby improving the image quality, and the target image is obtained by extracting important feature information from the original image and fusing it, and the target image is detected, and the target image can be identified and located, thereby improving the accuracy and robustness of target detection. Through the collaborative work of the deep separable convolution sub-model, feature extraction sub-model and feature fusion sub-model, the computational complexity and parameter amount of the model are reduced, the efficiency and real-time performance of target detection are enhanced, the effective combination of image defogging and target detection is achieved, and the applicability and reliability of the target detection system in complex environments are improved.
[0045] In some embodiments, inputting the original image into a target detection model to obtain a target image includes: Inputting the original image into the depth-separable convolution sub-model to obtain a first image, where the first image is a defogged image; Inputting the first image into the feature extraction sub-model to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map; The first feature map, the second feature map, the third feature map and the fourth feature map are input into the feature fusion sub-model to obtain the target image.
[0046] Figure 2 This is a second flow chart of a lightweight target detection method based on image defogging provided in an embodiment of the present application, such as Figure 2As shown, the original image is input into the depth separable convolution sub-model to obtain a first image, which is the dehazed image; the first image is input into the feature extraction sub-model to obtain a first feature map, a second feature map, a third feature map and a fourth feature map; the first feature map, the second feature map, the third feature map and the fourth feature map are input into the feature fusion sub-model to obtain a target image.
[0047] In one embodiment, the depth separable convolution sub-model is an AOD network, and the feature extraction sub-model and the feature fusion sub-model are CSPDarknet-53 models. On the basis of the detection network CSPDarknet-53, the AOD network is migrated, and the first convolution layer of the detection network CSPDarknet-53 is migrated and replaced with the AOD network to form a neural network capable of defogging and detection functions. Specifically, the first convolution layer of the detection network CSPDarknet-53 is replaced with the AOD network, so that the input image can first pass through the defogging process of the AOD network at the initial stage of entering the neural network. The AOD network uses its specific algorithm architecture and parameter settings to effectively estimate fog-related parameters, and then restores foggy images to restore clearer, more detailed and more accurate color images. After the image is defogged by the AOD network, feature extraction and feature fusion are performed along the subsequent network layer structure of the detection network CSPDarknet-53 to obtain the target image.
[0048] In this embodiment, by inputting the original image into the deep separable convolution sub-model for defogging, the original image can be effectively defogged to obtain a clear first image. This allows the model to focus more on the extraction and utilization of key features when processing the defogging image, enhances the recognition and positioning capabilities of traffic targets (vehicles, pedestrians, etc.) of different scales in foggy environments, effectively improves the accuracy of target detection, provides more reliable data support for functions such as illegal behavior monitoring and traffic flow statistics in traffic management, and improves the defect of poor performance of traffic target detection under severe weather conditions.
[0049] In some embodiments, inputting the original image into the depthwise separable convolution sub-model to obtain a first image comprises: The original image is input into the first depth-separable convolutional layer to obtain the feature Figure 1 ; The characteristics Figure 1 Input into the second depth separable convolutional layer to obtain the feature Figure 2 ; The characteristics Figure 1 and the characteristics Figure 2 Input to the first feature fusion layer to obtain the first feature tensor; The first feature tensor is input into the third depth-separable convolutional layer to obtain the feature Figure 3 ; The characteristics Figure 2 and the characteristics Figure 3 Input to the second feature fusion layer to obtain the second feature tensor; The second feature tensor is input into the fourth depth-separable convolutional layer to obtain the feature Figure 4 ; The characteristics Figure 1 , the characteristics Figure 2 , the characteristics Figure 3 and the characteristics Figure 4 Input to the third feature fusion layer to obtain the third feature tensor; The third feature tensor is input into the fifth depth-separable convolutional layer to obtain the feature Figure 5 ; The characteristics Figure 5 and the original image are input into the fourth feature fusion layer to obtain the first image.
[0050] Figure 3 This is a flow chart of the lightweight target detection method based on image defogging provided in the embodiment of the present application. Figure 3 As shown, the original image x is input into the first depth separable convolutional layer for feature extraction, and the feature Figure 1 x1; Input the feature map x1 into the second depth separable convolutional layer for feature extraction, and obtain the feature Figure 2 x2; input the feature map x1 and the feature map x2 into the first feature fusion layer to obtain the first feature tensor Concat1. The calculation formula of the first feature tensor is as follows:
[0051] Among them, x1 represents the feature Figure 1 , x2 represents the feature Figure 2 , Concat1 represents the first feature tensor.
[0052] Input the first feature tensor Concat1 into the third depth-separable convolutional layer to obtain the feature Figure 3 x3; the feature Figure 2 x2 and features Figure 3 x3 is input to the second feature fusion layer to obtain the second feature tensor Concat2. The calculation formula of the second feature tensor Concat2 is as follows:
[0053] Among them, x2 represents the feature Figure 2 , x3 represents the feature Figure 3, Concat2 represents the second feature tensor.
[0054] The second feature tensor is input into the fourth depth-separable convolutional layer to obtain the feature Figure 4 x4; the feature Figure 1 x1. Features Figure 2 x2. Features Figure 3 x3 and features Figure 4 x4 is input to the third feature fusion layer to obtain the third feature tensor Concat3. The calculation formula of the third feature tensor Concat3 is as follows:
[0055] Among them, x1 represents the feature Figure 1 , x2 represents the feature Figure 2 , x3 represents the feature Figure 3 , x4 represents the feature Figure 4 , Concat3 represents the third feature tensor.
[0056] Finally, the third feature tensor Concat3 is input into the fifth depth-separable convolutional layer to obtain the feature Figure 5 x5; the feature Figure 5 x5 and the original image x are input to the fourth feature fusion layer to obtain the first image Xout. The calculation formula of the first image is as follows:
[0057] Among them, Xout represents the first image, ReLU represents the activation function, and x5 represents the feature Figure 5 , x represents the original image.
[0058] In this embodiment, by inputting the original image into the deep separable convolution sub-model and performing multi-level feature extraction and fusion processing, the image can be effectively dehazed and rich image feature information can be extracted. By processing multiple deep separable convolution layers, the detailed features of the image are gradually extracted, and the multi-level information is integrated through the feature fusion layer to obtain a clearer first image, which improves the performance of traffic target detection tasks after image dehazing and the effectiveness of the traffic monitoring system in severe weather, and meets the application requirements in resource-constrained scenarios.
[0059] In some embodiments, the depthwise separable convolution layer includes a depthwise convolution layer and a pointwise convolution layer.
[0060] It should be noted that, taking the AOD network as an example with the depthwise separable convolution sub-model, on the basis of the AOD network, the convolution of the original AOD network is replaced by the depthwise separable convolution. The depthwise separable convolution is a combined convolution, including depthwise convolution and pointwise convolution. The depthwise convolution can independently perform convolution on each input channel with a small amount of calculation. The pointwise convolution mixes the information of different channels through 1×1 convolution to enhance the feature representation capability. In the forward propagation method forward, the calculation order of the forward propagation is adjusted, with depthwise convolution first and then pointwise convolution.
[0061] In this embodiment, by combining the deep convolution layer and the point-by-point convolution layer, the deep convolution reduces the number of convolution kernels by performing independent convolution on each input channel, thereby effectively reducing the computational complexity; the point-by-point convolution enhances the feature expression capability and improves the learning effect of the model; the deep separable convolution can significantly reduce the amount of computation and parameters while maintaining the model performance, thereby improving the accuracy.
[0062] In some embodiments, the expression of the depth convolution layer is as follows:
[0063] Among them, n represents the number of samples in the batch, c represents the channel number, h represents the position number in the height direction, w represents the position number in the width direction, and K d represents the depth convolution kernel, S d represents the step size, i represents the index of the depth convolution kernel in the height direction, P d represents padding, j represents the index of the depth convolution kernel in the width direction, It means that the elements at the corresponding positions are selected from the input feature map X according to the step size and padding to participate in the convolution calculation. is the element of the deep convolution output feature map at the corresponding position, Represents the weight value of the depth convolution kernel at the corresponding index.
[0064] It is easy to understand that the deep convolution performs convolution operations on each channel of the input feature map separately. In one embodiment, for a feature map with an input channel number of C_in, the deep convolution will use C_in convolution kernels, and the size of each convolution kernel is set to k×k (k is the set convolution kernel side length), which corresponds to a channel of the input feature map for convolution, and the information will not be mixed across channels during the convolution process. The number of channels of the feature map output after convolution of each channel remains single-channel, but the spatial size (width, height) of the feature map changes according to the convolution kernel size, step size, padding and other parameters. The expression of deep convolution is as follows:
[0065] Among them, n represents the number of samples in the batch, c represents the channel number, h represents the position number in the height direction, w represents the position number in the width direction, and K d represents the depth convolution kernel, S d represents the step size, i represents the index of the depth convolution kernel in the height direction, P d represents padding, j represents the index of the depth convolution kernel in the width direction, It means that the elements at the corresponding positions are selected from the input feature map X according to the step size and padding to participate in the convolution calculation. is the element of the deep convolution output feature map at the corresponding position, Represents the weight value of the depth convolution kernel at the corresponding index.
[0066] In this embodiment, the overall computational complexity can be reduced by using depthwise convolution, thereby achieving the purpose of being lightweight, which is suitable for the case where the number of input channels is large.
[0067] In some embodiments, the expression of the point-by-point convolutional layer is as follows:
[0068] in, is the element value of the point-by-point convolution output feature map at the corresponding position, n represents the sequence number of the sample in the batch, c represents the channel sequence number, m traverses all channels of the depth convolution output, h represents the position sequence number in the height direction, and w represents the position sequence number in the width direction. Represents the point-by-point convolution kernel The weight value from input channel m to output channel c in, is the deep convolution output feature map, The value of the element at the corresponding index (sample n, channel m, height h, width w).
[0069] It is easy to understand that point-by-point convolution uses a 1×1 convolution kernel to fuse the features of each channel output by the deep convolution according to the forward propagation forward, and change the number of channels to the number of output channels of the desired channel. In one embodiment, after the deep convolution, if the number of channels of the output feature map is C_mid (the C_mid parameter here depends on the deep convolution layer), it is converted into the desired output channel number C_out through the point-by-point convolution 1×1 convolution kernel. The expression of point-by-point convolution is as follows:
[0070] in, is the element value of the point-by-point convolution output feature map at the corresponding position, n represents the sequence number of the sample in the batch, c represents the channel sequence number, m traverses all channels of the depth convolution output, h represents the position sequence number in the height direction, and w represents the position sequence number in the width direction. Represents the point-by-point convolution kernel The weight value from input channel m to output channel c in, is the deep convolution output feature map, The value of the element at the corresponding index (sample n, channel m, height h, width w).
[0071] In this embodiment, by multiplying and summing the elements of each channel at the position of the deep convolution output with the corresponding convolution kernel weight, the channel information of the deep convolution output is fused through the point-by-point convolution kernel to obtain the output value of the point-by-point convolution at this position, thereby reducing the amount of calculation and being able to process information between channels and spatial dimensions at the same time.
[0072] In some embodiments, the step of inputting the first feature map, the second feature map, the third feature map, and the fourth feature map into the feature fusion sub-model to obtain the target image includes: The fourth feature map is input into the first upsampling layer and the third feature map is input into the fifth feature fusion layer to obtain a fusion feature map. Figure 3 ; The fusion feature Figure 3 After inputting the second upsampling layer, the second feature map is input into the sixth feature fusion layer to obtain the fused feature Figure 2 ; The fusion feature Figure 2 After inputting the first downsampling layer and the fusion feature Figure 3 Input to the seventh feature fusion layer to obtain the re-fusion feature Figure 3 ; The reintegration feature Figure 3 After being input into the second downsampling layer, the fourth feature map is input into the eighth feature fusion layer to obtain the target image.
[0073] Figure 4 This is a flow chart of the lightweight target detection method based on image defogging provided in the embodiment of the present application. Figure 4 As shown, in one embodiment, the output of the AOD network is input into CSPDarknet-53, and then through 4 layers of feature extraction, down-sampling by 2 times, 4 times, 8 times and 16 times respectively, to form 4 multi-scale features of different sizes, and obtain the first feature map P1, the second feature map P2, the third feature map P3 and the fourth feature map P4.
[0074] Since the fourth feature map P4 has been downsampled 16 times, the information it contains has strong semantics and global features. In traffic scenes, it can quickly determine the approximate category and rough location information of large objects (such as large trucks, buildings, etc.). Since it has been downsampled by a large multiple, the pixel resolution is low, but this highly abstract feature is very effective in eliminating some obvious background interference and determining the approximate range of the target in the early stage of the network. The third feature map P3 is generated by 8 times downsampling. While retaining certain detail information, it also has richer semantic information. Compared with feature map 4, it retains certain global semantic information while beginning to contain more local detail information, and can more accurately describe the shape and position of medium-sized targets (such as ordinary cars, pedestrian groups, etc.).
[0075] The fourth feature map is input into the first upsampling layer and the third feature map is input into the fifth feature fusion layer to obtain the fused feature Figure 3 RP3, can combine the advantages of both, so that the fusion feature Figure 3 It can not only obtain richer semantic information to facilitate the classification and judgment of the target, but also supplement some detailed information to improve the perception accuracy of the target position and contour, and fuse features Figure 3 RP3 can supplement higher-level semantic information on the basis of the original third feature map P3, enhance the depth of the overall semantic understanding of the target, thereby improving the accuracy of classification and positioning of medium-scale targets, and integrating features. Figure 3 The calculation formula of RP3 is as follows:
[0076] Among them, RP3 represents the fusion feature Figure 3 , Concat represents feature fusion, Upsample represents upsampling operation, represents the weight of the fourth feature map, represents the weight of the third feature map, P3 represents the third feature map, and P4 represents the fourth feature map.
[0077] The second feature map P2 is generated by 4 times downsampling, contains richer detail information, and is more accurate in expressing the features of small-scale targets (such as traffic signs, individual pedestrians, etc.). Its detail information is richer than the previous two, and is consistent with the fusion feature map P2. Figure 3 RP3 fusion can further refine the target feature representation and enhance the detection capability of small and medium-sized targets. Figure 3 RP3 is input into the second upsampling layer and the second feature map P2 is input into the sixth feature fusion layer to obtain the fusion feature Figure 2 RP2, which makes the fusion feature Figure 2 RP2 can combine the rich details and fusion features of the second feature map P2 Figure 3The semantic information in RP3 further improves the detection capability of small-scale targets, and is more accurate in capturing the edge and texture features of targets. Figure 2 The calculation formula of RP2 is as follows:
[0078] Among them, RP2 represents the fusion feature Figure 2 , Concat represents feature fusion, Upsample represents upsampling operation, represents the weight of the second feature map, Represents fusion features Figure 2 The weight of P2 represents the second feature map, and RP3 represents the fusion feature Figure 3 .
[0079] In order to integrate the information of different levels of features at different resolutions, in traffic scenes, there are situations such as vehicle occlusion, etc., multi-resolution feature fusion can better handle these complex situations, enable the network to understand the target features from multiple angles, enhance the robustness of feature expression, and fuse features. Figure 2 After inputting the first downsampling layer and fusion features Figure 3 Input to the seventh feature fusion layer to obtain the re-fusion feature Figure 3 RRP3, re-integration signature Figure 3 The calculation formula of RRP3 is as follows:
[0080] Among them, RP3 represents the fusion feature Figure 3 , Concat means feature fusion, Down means downsampling operation, Represents fusion features Figure 2 The weight of Represents fusion features Figure 3 The weight of RP2 represents the fusion feature Figure 2 , RRP3 indicates reintegration signature Figure 3 .
[0081] In order to make the target image integrate high-level semantic features and middle-level features after multiple fusion optimizations, while maintaining a strong semantic understanding of large-scale targets, the detailed information is supplemented to the greatest extent possible, providing a comprehensive and accurate feature basis for the final target detection output, so that large, medium and small-scale targets can be detected more effectively.
[0082] Re-integrate features Figure 3 After inputting the second downsampling layer and the fourth feature map, it is input into the eighth feature fusion layer to obtain the target image RP4. The calculation formula of the target image RP4 is as follows:
[0083] Among them, RP3 represents the fusion feature Figure 3 , Concat means feature fusion, Down means downsampling operation, represents the weight of the fourth feature map, Represents fusion features Figure 3 , RP4 represents the target image, and P4 represents the fourth feature map.
[0084] In this embodiment, a lightweight feature fusion is formed by modifying the feature fusion method of the model; the network layers are reasonably streamlined to reduce invalid calculations and redundant information transmission, and the large-, medium- and small-scale target detection accuracy is improved through multi-scale feature extraction and fusion. At the same time, the cosine similarity of the feature outputs of adjacent layers is calculated to accurately determine which layers can be merged, thereby reducing the network depth and effectively reducing the computational complexity of the model, so that the model can run more efficiently. It is suitable for application scenarios such as traffic monitoring that have high real-time requirements and limited resources, and improves the overall computing speed while ensuring a certain detection accuracy.
[0085] In some embodiments, the training process of the target detection model includes: Construct a preset target detection model, wherein the preset target detection model is a CSPDarknet-53 model; Obtaining a data set based on the sample image; Based on the first loss function, the preset target detection model is trained according to the data set to obtain the target detection model.
[0086] It is easy to understand that by constructing a preset target detection model, the preset target detection model is a CSPDarknet-53 model, a data set is obtained based on a sample image, and based on a first loss function, the preset target detection model is trained according to the data set to obtain a target detection model. The expression of the first loss function is as follows:
[0087] Among them, p is the scalar of the predicted classification score vector, for the positive samples in training, q is the intersection-over-union ratio between the generated predicted box and the true box, for the negative samples in training, the training target q of all categories is 0, α is the loss weight of the image foreground and background, and γ is the weight of each sample.
[0088] In this embodiment, by constructing a preset target detection model, generating a data set based on sample images, and training the model in combination with the first loss function, the performance and accuracy of the target detection model can be improved, the learning ability of the model is optimized, and the computational complexity is reduced while ensuring high accuracy.
[0089] The lightweight target detection method based on image defogging provided in the embodiment of the present application can be executed by a lightweight target detection device based on image defogging. In the embodiment of the present application, the lightweight target detection method based on image defogging is executed by a lightweight target detection device based on image defogging as an example to illustrate the lightweight target detection device based on image defogging provided in the embodiment of the present application.
[0090] The present application also provides a lightweight target detection device based on image defogging. Figure 5 As shown, the lightweight object detection device based on image defogging includes: an acquisition module 510, a processing module 520 and a detection module 530.
[0091] An acquisition module 510 is used to acquire an original image, where the original image is acquired based on acquisition equipment; A processing module 520 is used to input the original image into a target detection model to obtain a target image, wherein the target detection model is obtained based on sample image training, and the target detection model includes a depth separable convolution sub-model, a feature extraction sub-model and a feature fusion sub-model; A detection module 530 is used to perform detection based on the target image to obtain a target detection result; Among them, the depth-separable convolution sub-model includes a first depth-separable convolution layer, a second depth-separable convolution layer, a third depth-separable convolution layer, a fourth depth-separable convolution layer, a fifth depth-separable convolution layer, a first feature fusion layer, a second feature fusion layer, a third feature fusion layer and a fourth feature fusion layer, the feature extraction sub-model includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer, and the feature fusion sub-model includes a fifth feature fusion layer, a sixth feature fusion layer, a seventh feature fusion layer, an eighth feature fusion layer, a first upsampling layer, a second upsampling layer, a first downsampling layer and a second downsampling layer.
[0092] According to the lightweight target detection device based on image defogging provided by the embodiment of the present application, by acquiring the original image and inputting it into the target detection model for processing, combined with the deep separable convolution sub-model, feature extraction sub-model and feature fusion sub-model, the image can be defogged before detection, thereby improving the image quality, and the target image is obtained by extracting important feature information from the original image and fusing it, and the target image is detected, and the target image can be identified and located, thereby improving the accuracy and robustness of target detection. Through the collaborative work of the deep separable convolution sub-model, feature extraction sub-model and feature fusion sub-model, the computational complexity and parameter amount of the model are reduced, the efficiency and real-time performance of target detection are enhanced, the effective combination of image defogging and target detection is achieved, and the applicability and reliability of the target detection system in complex environments are improved.
[0093] The lightweight object detection device based on image defogging provided in the embodiment of the present application can achieve Figures 1 to 4 To avoid repetition, the various processes implemented in the embodiment of the lightweight target detection method based on image defogging are not described here.
[0094] In some embodiments, Figure 6 As shown, an embodiment of the present application also provides an electronic device 600, including a processor 601, a memory 602, and a computer program stored in the memory 602 and executable on the processor 601. When the program is executed by the processor 601, each process of the above-mentioned lightweight target detection method embodiment based on image defogging is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0095] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0096] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned lightweight target detection method embodiment based on image defogging are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0097] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0098] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned lightweight object detection method based on image defogging.
[0099] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0100] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned lightweight target detection method embodiment based on image defogging, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0101] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0102] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0103] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the lightweight target detection method based on image dehazing of each embodiment of the present application.
[0104] In the description of this application, "first feature" or "second feature" may include one or more of the features.
[0105] In the description of the present application, “plurality” means two or more.
[0106] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
[0107] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0108] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.
Claims
1. A lightweight target detection method based on image defogging, characterized in that: The method comprises: Acquire an original image, where the original image is acquired based on acquisition equipment; Inputting the original image into a target detection model to obtain a target image, wherein the target detection model is trained based on a sample image, and the target detection model includes a depth separable convolution sub-model, a feature extraction sub-model, and a feature fusion sub-model; Perform detection based on the target image to obtain a target detection result; Among them, the depth-separable convolution sub-model includes a first depth-separable convolution layer, a second depth-separable convolution layer, a third depth-separable convolution layer, a fourth depth-separable convolution layer, a fifth depth-separable convolution layer, a first feature fusion layer, a second feature fusion layer, a third feature fusion layer and a fourth feature fusion layer, the feature extraction sub-model includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer, and the feature fusion sub-model includes a fifth feature fusion layer, a sixth feature fusion layer, a seventh feature fusion layer, an eighth feature fusion layer, a first upsampling layer, a second upsampling layer, a first downsampling layer and a second downsampling layer.
2. The lightweight target detection method based on image defogging according to claim 1, characterized in that: The step of inputting the original image into the target detection model to obtain the target image comprises: Inputting the original image into the depth-separable convolution sub-model to obtain a first image, where the first image is a defogged image; Inputting the first image into the feature extraction sub-model to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map; The first feature map, the second feature map, the third feature map and the fourth feature map are input into the feature fusion sub-model to obtain the target image.
3. The lightweight target detection method based on image defogging according to claim 2, characterized in that: Inputting the original image into the depth-separable convolution sub-model to obtain a first image includes: Inputting the original image into a first depth-separable convolutional layer to obtain a feature map 1; Inputting the feature map 1 into a second depth-wise separable convolutional layer to obtain a feature map 2; Inputting the feature map 1 and the feature map 2 into a first feature fusion layer to obtain a first feature tensor; Inputting the first feature tensor into a third depth-wise separable convolutional layer to obtain a feature map 3; Inputting the feature map 2 and the feature map 3 into a second feature fusion layer to obtain a second feature tensor; Inputting the second feature tensor into a fourth depth-wise separable convolutional layer to obtain a feature map 4; Inputting the feature map 1, the feature map 2, the feature map 3 and the feature map 4 into a third feature fusion layer to obtain a third feature tensor; Inputting the third feature tensor into a fifth depth-wise separable convolutional layer to obtain a feature map 5; The feature map five and the original image are input into a fourth feature fusion layer to obtain the first image.
4. The lightweight target detection method based on image defogging according to claim 3 is characterized in that: The depth-wise separable convolution layer includes a depth-wise convolution layer and a point-wise convolution layer.
5. The lightweight target detection method based on image defogging according to claim 4, characterized in that: The expression of the depth convolution layer is as follows: ; Among them, n represents the number of samples in the batch, c represents the channel number, h represents the position number in the height direction, w represents the position number in the width direction, and K d represents the depth convolution kernel, S d represents the step size, i represents the index of the depth convolution kernel in the height direction, P d represents padding, j represents the index of the depth convolution kernel in the width direction, It means that the elements at the corresponding positions are selected from the input feature map X according to the step size and padding to participate in the convolution calculation. is the element of the deep convolution output feature map at the corresponding position, Represents the weight value of the depth convolution kernel at the corresponding index.
6. The lightweight target detection method based on image defogging according to claim 4, characterized in that: The expression of the point-by-point convolutional layer is as follows: ; in, is the element value of the point-by-point convolution output feature map at the corresponding position, n represents the sequence number of the sample in the batch, c represents the channel sequence number, m traverses all channels of the depth convolution output, h represents the position sequence number in the height direction, and w represents the position sequence number in the width direction. Represents the point-by-point convolution kernel The weight value from input channel m to output channel c in, is the deep convolution output feature map, The value of the element at the corresponding index (sample n, channel m, height h, width w).
7. The lightweight target detection method based on image defogging according to claim 2, characterized in that: The step of inputting the first feature map, the second feature map, the third feature map, and the fourth feature map into the feature fusion sub-model to obtain the target image includes: The fourth feature map is input into the first upsampling layer and the third feature map is input into the fifth feature fusion layer to obtain a fused feature map three; The fused feature map 3 is input into the second upsampling layer and the second feature map is input into the sixth feature fusion layer to obtain a fused feature map 2; Input the fused feature map 2 to the first downsampling layer and the fused feature map 3 to the seventh feature fusion layer to obtain a re-fused feature map 3; The re-fused feature map three is input into the second down-sampling layer and the fourth feature map is input into the eighth feature fusion layer to obtain the target image.
8. The lightweight target detection method based on image defogging according to claim 1, characterized in that: The training process of the target detection model includes: Construct a preset target detection model, wherein the preset target detection model is a CSPDarknet-53 model; Obtaining a data set based on the sample image; Based on the first loss function, the preset target detection model is trained according to the data set to obtain the target detection model.
9. A lightweight target detection device based on image defogging, implemented by the lightweight target detection method based on image defogging according to any one of claims 1 to 8, characterized in that: The device comprises: An acquisition module is used to acquire an original image, where the original image is acquired based on acquisition equipment; A processing module, used for inputting the original image into a target detection model to obtain a target image, wherein the target detection model is obtained based on sample image training, and the target detection model includes a depth separable convolution sub-model, a feature extraction sub-model and a feature fusion sub-model; A detection module, used to perform detection based on the target image to obtain a target detection result; Among them, the depth-separable convolution sub-model includes a first depth-separable convolution layer, a second depth-separable convolution layer, a third depth-separable convolution layer, a fourth depth-separable convolution layer, a fifth depth-separable convolution layer, a first feature fusion layer, a second feature fusion layer, a third feature fusion layer and a fourth feature fusion layer, the feature extraction sub-model includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer, and the feature fusion sub-model includes a fifth feature fusion layer, a sixth feature fusion layer, a seventh feature fusion layer, an eighth feature fusion layer, a first upsampling layer, a second upsampling layer, a first downsampling layer and a second downsampling layer.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the lightweight target detection method based on image defogging as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Image classification method and system based on densely connected MobileNets model
CN110489584A
Deep learning small target detection method and device based on cascade fusion and attention mechanism
CN112801158A
Image defogging method of depth separable convolutional neural network based on ResNet
CN116309165A
Foggy vehicle detection method and system based on deep learning
CN116597153A
Haze scene driving vision enhancement and target detection method
CN117974497A