A method and system for detecting small targets in low-light environments
Through the improved Zero-DCE network and simAM attention mechanism, small targets in dark light environments are detected, which solves the problem of difficult sample acquisition and poor model performance, and achieves efficient low-light small target detection.
Patent Information
- Application Number
- CN202510112801.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing dark light enhancement algorithm or dark light detection algorithm is trained based on a large amount of dark light data, and the sample acquisition is difficult, resulting in unsatisfactory performance in dark light small object detection scenarios.
The improved Zero-DCE network is used to dark-light enhance the original image, and a simAM attention mechanism is added to the backbone feature extraction module of YOLOv11. The enhanced image is input into the improved YOLOv11 network to obtain small object detection results.
There is no need to enhance the training data set in dark light, reduce the difficulty of data acquisition, avoid model overfitting, improve model training and inference speed, and reduce the requirements for device computing resources. Each pixel of the image has its own brightness enhancement curve, making the image smoother after enhancement.
Smart Images

Figure CN119559086B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection in images, and specifically to a small object detection method and system in a low-light environment. Background Art
[0002] Currently, traditional image enhancement techniques, such as histogram enhancement or RETINEX-based methods, although they can improve the brightness and contrast of low-light images to a certain extent, often amplify noise at the same time, or lose some key detail information while enhancing the target, and cannot effectively solve the contradiction between low light and small object detection.
[0003] Object detection algorithms based on deep learning can achieve high accuracy and recall rates under normal lighting conditions. For example, the PE-yolo method adds a low-frequency enhancement filter on the basis of the Laplacian pyramid and combines yolov3 to achieve end-to-end detection. While the low-frequency enhancement filter enhances the low-frequency detail features, it also introduces the enhancement of noise, and in the process of continuously sampling and extracting features by the backbone network Darknet-53 of yolov3, as the number of layers deepens, the detail features of small objects are prone to gradually be lost, and the performance in the low-light small object detection scenario is still not satisfactory. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by the present invention is that existing low-light enhancement algorithms or low-light detection algorithms are trained based on a large amount of low-light data, and it is difficult to obtain samples.
[0006] To solve the above technical problem, the present invention provides the following technical solution: A small object detection method in a low-light environment, including:
[0007] Collecting original image information;
[0008] Using the improved Zero-DCE network to enhance the low light of the original image;
[0009] Adding a simAM attention mechanism to the backbone feature extraction module of yolov11, and inputting the enhanced image into the improved YOLOv11 network to obtain the small object detection result;
[0010] The improved Zero-DCE network includes adding a low-pass enhancement filter to the basic Zero-DCE network to enhance the object texture features and filter the noise in camera imaging.
[0011] As a preferred solution of the small target detection method in the low-light environment of the present invention, wherein: the low-light enhancement includes enhancing the low-light image by using Zero-DCE, which consists of multiple curve estimation modules;
[0012] The curve estimation modules are connected in series, and each module is responsible for learning a curve transformation of the input image to enhance the lighting effect of the image;
[0013] For each of the curve estimation modules, it is assumed that by learning a set of pixel-level curves, the lighting of the image is enhanced;
[0014] For each pixel in the input low-light image, the goal of the network is to learn a curve so that after the pixel value is transformed by the curve, the enhanced pixel value is obtained;
[0015] The curve is a brightness enhancement curve, which is automatically learned by the network according to the input low-light image.
[0016] As a preferred solution of the small target detection method in the low-light environment of the present invention, wherein: the formula of the brightness enhancement curve is:
[0017]
[0018] Wherein, represents a pixel value of the original image I at index ( ), represents the brightness enhancement curve parameter, represents the pixel value of the image after the -th order enhancement; b is an unknown number, representing an equation of any order.
[0019] As a preferred solution of the small target detection method in the low-light environment of the present invention, wherein: the basic Zero-DCE network includes 6 hidden layers and 1 output layer;
[0020] Among them, 6 hidden layers use symmetric skip connections similar to U-Net;
[0021] All layers are ordinary 3x3 equal-length convolutional layers, and the output layer uses tanh as the activation function;
[0022] The improved Zero-DCE network includes using the output of the 3rd layer of Zero-DCE and the result of the low-frequency processing of the output of the 3rd layer as the input of the 4th layer; using the output of the 4th layer and the result of the low-frequency processing of the output of the 2nd layer as the input of the 5th layer; using the output of the 5th layer and the result of the low-frequency processing of the output of the 1st layer as the input of the 6th layer.
[0023] As a preferred solution of the low-light environment small target detection method described in the present invention, wherein: the steps of the filtering operation include: setting the improved Zero-DCE intermediate output feature map , where represents the pixel coordinates of the feature map, and i represents the hidden layer index of Zero-DCE;
[0024] Perform Fourier transform on the feature map and convert it to the frequency domain:
[0025]
[0026] where, W represents the width of the feature map, H represents the height of the feature map, represents the coordinates in the frequency domain, represents the intensity of the frequency component in the image;
[0027] Use the Butterworth low-pass filter BLPF to enhance the low-frequency information. The transfer formula of BLPF is:
[0028]
[0029] where, represents the feature point to the center of the feature spectrum distance; represents the cut-off frequency, which is used to control the separation degree of high and low frequency features. When designing the network structure, use the adaptive pooling network instead of ; n represents the filter order, which controls the attenuation degree of high-frequency features;
[0030] Multiply the frequency domain value by the low-pass filter to achieve low-frequency information enhancement. The low-frequency filtering formula is:
[0031]
[0032] where, represents the filtered spectrum;
[0033] Moderately amplify the low-frequency coefficients:
[0034]
[0035] where, k is the enhancement coefficient, , represents the amplified filtering result;
[0036] Finally, perform inverse Fourier transform to obtain the feature map after low-frequency enhancement filtering:
[0037]
[0038] Among them, represents the feature map obtained after low-frequency filtering of the feature map.
[0039] As a preferred solution of the small target detection method in low-light environment described in the present invention, wherein: the simAM attention mechanism is set in the feature map, and the feature at each pixel position is used as a neuron;
[0040] For a given neuron, its output , and the outputs of the surrounding neurons are , and the energy is measured by calculating the linear combination difference between the neuron and the surrounding neurons :
[0041]
[0042] Among them, and represent the weight coefficients for balancing the difference and the degree of association in the energy calculation, and are the means of the neuron outputs and respectively;
[0043] Convert the value of the calculated energy function into an attention weight, and use the Sigmod activation function as the conversion function. For the output of a given neuron, it is expressed as:
[0044] .
[0045] As a preferred solution of the small target detection method in low-light environment described in the present invention, wherein: the improved YOLOv11 network further includes using a low-light enhancement model to enhance the training set images, and then inputting the enhanced images into the improved YOLOv11 network for training. By continuously adjusting the network parameters, the network can accurately detect small targets in low light.
[0046] A small target detection system in low-light environment adopting the method as described in the present invention, characterized in that:
[0047] An acquisition unit for acquiring original image information;
[0048] An enhancement unit for enhancing the original image with the improved Zero-DCE network;
[0049] The detection unit adds the simAM attention mechanism to the backbone feature extraction module of yolov11, and inputs the enhanced image into the improved YOLOv11 network to obtain the small target detection result; the improved Zero-DCE network includes adding a low-pass enhancement filter to the basic Zero-DCE network to enhance the object texture features and filter the noise in camera imaging.
[0050] A computer device includes: a memory and a processor; the memory stores a computer program, wherein: when the processor executes the computer program, the steps of the method described in any one of the present inventions are implemented.
[0051] A computer-readable storage medium stores a computer program thereon, wherein: when the computer program is executed by a processor, the steps of the method described in any one of the present inventions are implemented.
[0052] The beneficial effects of the present invention: The small target detection method in a low-light environment provided by the present invention does not require a low-light enhancement training data set, reduces the difficulty of data acquisition, and avoids the occurrence of overfitting of the low-light enhancement model during the training process. The model training speed and inference speed are fast, and the computing power resources of the device are required low. Each pixel of the image has its own brightness enhancement curve, which is obtained during training, making the enhanced image smoother. Description of the Drawings
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 It is the overall flowchart of a small target detection method in a low-light environment provided by the first embodiment of the present invention;
[0055] Figure 2 It is the flowchart of the curve enhancement model in a small target detection method in a low-light environment provided by the first embodiment of the present invention;
[0056] Figure 3 It is the yolov11 detection architecture diagram of a small target detection method in a low-light environment provided by the first embodiment of the present invention;
[0057] Figure 4 It is the Zero-DCE architecture diagram of the conventional curve enhancement model in a small target detection method in a low-light environment provided by the first embodiment of the present invention;
[0058] Figure 5 Flowchart of the pixel attention mechanism for a small target detection method in low-light environment provided by the first embodiment of the present invention;
[0059] Figure 6 Original low-light image of a small target detection method in low-light environment provided by the second embodiment of the present invention;
[0060] Figure 7 Pixel distribution of each channel of the original low-light data of a small target detection method in low-light environment provided by the second embodiment of the present invention;
[0061] Figure 8 Image enhanced by Zero-DCE in a small target detection method in low-light environment provided by the second embodiment of the present invention;
[0062] Figure 9 Pixel distribution of each channel of the image enhanced by the Zero-DCE network in a small target detection method in low-light environment provided by the second embodiment of the present invention;
[0063] Figure 10 Image enhanced by the model in a small target detection method in low-light environment provided by the second embodiment of the present invention;
[0064] Figure 11 Pixel distribution of each channel of the image enhanced by the model in a small target detection method in low-light environment provided by the second embodiment of the present invention. Detailed implementation manners
[0065] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] Embodiment 1, referring to Figures 1-5 , which is an embodiment of the present invention, provides a small target detection method in a low-light environment, including:
[0067] S1: Collect original image information.
[0068] S2: Perform low-light enhancement on the original image.
[0069] It should be noted that in a dark environment, the overall brightness of the captured images is generally low, and low-brightness pixels account for a significant proportion. In essence, dark light enhancement aims to increase the pixel value of the foreground of the image, and effectively increase the brightness by increasing the contrast between the foreground and the background, thereby improving the visual effect of the image.
[0070] Because different objects have different reflective properties under different lighting conditions, the pixel values they present in the image show a diverse distribution state. If a linear method is used to enhance the lighting effect of an image, not only will the highly reflective areas be prone to overexposure, but the noise points in the image will also be amplified to the same extent, which will undoubtedly have a serious negative impact on image quality.
[0071] In view of this, a curve model is constructed that can enhance low-brightness pixels to a large extent while maintaining the original state of high-brightness pixels as much as possible, so as to achieve the goal of accurate and effective dark light enhancement. While significantly improving image brightness and contrast, it avoids adverse phenomena such as exposure and noise amplification caused by excessive enhancement, ensuring that the enhanced image can clearly and truly restore the scene information in a dark light environment.
[0072] Zero-DCE (Zero-Reference Deep Curve Estimation) is a deep learning method for low-light image enhancement (such as Figure 4 ), mainly composed of multiple curve estimation modules. It is assumed that the illumination of an image can be enhanced by learning a set of pixel-level curves. For each pixel in the input low-light image, the goal of the network is to learn a suitable curve so that the enhanced pixel value can be obtained after the pixel value is transformed through this curve. These curves are not pre-defined, but are automatically learned by the network based on the input low-light image. These modules are connected in series, and each module is responsible for learning a curve transformation of the input image to gradually enhance the illumination effect of the image.
[0073] Assume that the brightness enhancement curve formula is:
[0074]
[0075] in, Indicates that the original image I index is ( ), represents the brightness enhancement curve parameters, Indicates The pixel value of the image after nth-order enhancement; b is an unknown number representing an equation of any order. Its purpose is to effectively enhance low-light images without the need for a reference image (i.e., without the need for paired low-light and normal-light images for training).
[0076] S3: Use the improved Zero-DCE network to adjust the brightness of the image after low-light enhancement, such as Figure 4 .
[0077] Based on the DCE-Network, a low-pass enhancement filter is added to enhance the texture features of objects and filter the noise in camera imaging. The basic Zero-DCE network is very simple, with a total of 7 layers (6 hidden layers and 1 output layer, and 6 hidden layers use symmetric skip connections similar to U-Net). All layers are ordinary 3x3 equal-length (stride = 1) convolutional layers, and the output layer uses tanh as the activation function, as Figure 4 shown.
[0078] Since the low-frequency components contain the most image semantic information, to enrich the semantic information in the reconstructed image, the present invention adds low-pass filtering processing to the outputs of the first, second, and third layers of Zero-DCE to ensure the enhancement of low-frequency features. Among them, the improved Zero-DCE network includes using the output of the third layer of Zero-DCE and the result of the output of the third layer after low-frequency processing as the input of the fourth layer; using the output of the fourth layer and the result of the output of the second layer after low-frequency processing as the input of the fifth layer; using the output of the fifth layer and the result of the output of the first layer after low-frequency processing as the input of the sixth layer.
[0079] Let the intermediate output feature map of the improved Zero-DCE be , where represents the pixel coordinates of the feature map, and i represents the hidden layer index of Zero-DCE; the operation steps are as follows:
[0080] Perform Fourier transform on the feature map to convert it to the frequency domain:
[0081]
[0082] Among them, W is the width of the feature map, H is the height of the feature map, is the coordinate in the frequency domain, and different correspond to different frequencies, and the magnitude of the amplitude reflects the intensity of this frequency component in the image.
[0083] Use the Butterworth low-pass filter BLPF to enhance the low-frequency information. The transfer formula of BLPF:
[0084]
[0085] Among them, represents the distance from the feature point to the center of the feature spectrum ; represents the cut-off frequency, which is used to control the separation degree of high-frequency and low-frequency features. When designing the network structure, an adaptive pooling network is used instead of ; n is the filter order, which controls the attenuation degree of high-frequency features.
[0086] Finally, by multiplying the frequency domain value with the low-pass filter, the enhancement of low-frequency information is realized. The low-pass filter formula is:
[0087]
[0088] where is to multiply the original spectrum with the transfer function to obtain the filtered spectrum. This new spectrum is the result after being processed by the Butterworth low-pass filter, which realizes the retention and enhancement of low-frequency components (because the low-frequency region ), and at the same time attenuates high-frequency components (because the high-frequency region and decreases as the frequency increases).
[0089] Realize the retention and enhancement of low-frequency components, and at the same time attenuate high-frequency noise. To further highlight the low-frequency enhancement effect, the low-frequency coefficients are moderately amplified:
[0090]
[0091] where k is the enhancement coefficient, , represents the amplified filtered result.
[0092] Finally, the inverse Fourier transform is performed to obtain the feature map after low-frequency enhancement filtering:
[0093]
[0094] where represents the feature map obtained after low-frequency filtering, where represents the pixel coordinates of the feature map.
[0095] S4: Add the simAM attention mechanism to the backbone feature extraction module of yolov11, input the enhanced image into the improved YOLOv11 network, and obtain the small target detection result.
[0096] The core of the simAM attention mechanism lies in measuring the importance of each neuron through an energy function. For a given neuron, its output and the outputs of the surrounding neurons are (where the surrounding neurons can be defined according to a certain neighborhood range, such as other neurons within a rectangular area centered on this neuron etc.). The energy of this neuron is measured by calculating the linear combination difference between it and these surrounding neurons :
[0097]
[0098] where and represent the weight coefficients used to balance the proportion of the difference and the degree of correlation in the energy calculation. and are the means of the neuron outputs and respectively. This calculation method of the linear combination difference aims to capture the uniqueness and coordination of neurons in the local area. For example, if the linear combination difference between a neuron and its surrounding neurons is large, it means that its performance in this local area is relatively "prominent", and it may carry some key feature information in the image, such as representing the mutation of the object edge in the image. At this time, its energy value will be relatively high; on the contrary, if its linear relationship with the surrounding neurons is relatively coordinated and the change trend is similar, its energy value will be low.
[0099] According to the value of the calculated energy function, it is necessary to further convert it into an attention weight so that it can be directly applied to the feature map to adjust the attention degree to different neurons (that is, the features at different pixel positions). This conversion function plays a role of normalization and mapping, converting the relatively abstract and numerically uncertain quantity of the energy value into an attention weight, making it fall within a reasonable range (usually between 0 and 1). The present invention uses the Sigmod activation function as the conversion function:
[0100]
[0101] When has a very large value (i.e., high energy), will approach 0, indicating that the attention weight corresponding to this neuron is very low and will be relatively weakened in subsequent feature weighting and other operations; while when has a very small value (low energy), Approaching 1 means that the attention weight of this neuron is high, and the corresponding feature will be focused on in subsequent processing, enabling the network to adaptively focus on the feature information represented by neurons that are more important in the local area and more in line with the overall feature consistency, thereby enhancing the entire network's ability to extract and utilize image features.
[0102] Add the simAM attention mechanism to the backbone feature extraction module of yolov11, use the low-light enhancement model to enhance the training set images, and then input the enhanced images into the improved YOLOv11 network for training. By continuously adjusting the network parameters, the network can accurately detect small low-light targets. As Figure 2 、 Figure 3 shown, the design steps of the curve enhancement model integrating low-frequency enhancement and the main steps of integrating the attention module based on yolov11 are as follows:
[0103] (1) Build a transfer network at the outputs of the 2nd and 3rd hidden layers of Zero-DCE.
[0104] (2) Add a channel separation layer to separate the three RGB channels, and perform steps (3)-(6) on each channel:
[0105] (3) Add a convolution for filtering and use Gaussian padding.
[0106] (4) Add adaptive pooling and use mean padding.
[0107] (5) Add upsampling.
[0108] (6) Add concatenation to concatenate the features processed for RGB.
[0109] (7) Concatenate the output of the 2nd process and the output of the 4th layer and give it to the 6th layer.
[0110] (8) Concatenate the output of the 3rd layer process and the original output and give it to the 5th layer.
[0111] (9) Train the curve enhancement model.
[0112] (10) Modify the yolov11 backbone structure and add the simAM network after c3k2.
[0113] (11) Modify the loss and add the NWD loss.
[0114] Among them, a low-frequency filtering enhancer is added to the adjacent layer of Zero-DCE to improve the detailed texture features. The yolov11 network architecture with a pixel attention mechanism is added. An attention mechanism is added after the C3k2 module of the backbone of yolo to improve the pixel accuracy of the segmented targets at different scales.
[0115] On the other hand, this embodiment also provides a small target detection system for low-light environments, which includes:
[0116] An acquisition unit that acquires the original image information.
[0117] An enhancement unit that uses the improved Zero-DCE network to perform low-light enhancement on the original image.
[0118] A detection unit that adds a simAM attention mechanism to the backbone feature extraction module of yolov11, and inputs the enhanced image into the improved YOLOv11 network to obtain the small target detection result. The improved Zero-DCE network includes adding a low-pass enhancement filter to the basic Zero-DCE network to enhance the object texture features and filter the noise in camera imaging.
[0119] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0120] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite ordered listing of executable instructions for implementing logical functions, and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
[0121] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0122] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0123] Example 2, referring to Figures 6-11 , which is an embodiment of the present invention, provides a small target detection method in a low-light environment. To verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0124] Figure 6 : Original low-light data. Figure 7 : Pixel distribution of the original low-light data. Figure 8 : Zero-DCE enhanced data. Figure 9 : Pixel distribution of the Zero-DCE enhanced data. Figure 10:The data after model enhancement provided by the present invention. Figure 11 :The pixel distribution of the data after model enhancement provided by the present invention.
[0125] From Figure 7 、 Figure 9 、 Figure 11 It can be seen that from the characteristics of the image itself, there is a non-linear correspondence between the illumination intensity of the image and the pixel value. Specifically, for the area with a lower pixel value, since it corresponds to a lower illumination intensity, the magnification of this area should be increased to enhance the brightness of the dark area and improve the overall visibility of the image; while for the area with a higher pixel value, considering that it already corresponds to a relatively high illumination intensity, in order to prevent overexposure due to excessive enhancement, the magnification of this area needs to be reduced. Thus, on the basis of effectively enhancing the overall illumination effect of the dark image, it is ensured that the brightness of each area of the image is within a reasonable range, avoiding problems such as detail loss and poor visual effects caused by overexposure.
[0126] From the effect after image enhancement, the dark-light enhancement method provided by the present invention can effectively improve the illumination of the detailed texture part in the image and eliminate the overexposure problem in the area with a large illumination intensity in the original data.
[0127] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for detecting small targets in a dark environment, characterized in that: include: Collecting original image information; The improved Zero-DCE network is used to enhance the dark light of the original image; Add the simAM attention mechanism to the backbone feature extraction module of YOLOv11, input the enhanced image into the improved YOLOv11 network, and obtain the small target detection result; The improved Zero-DCE network includes, in the basic Zero-DCE network, adding a low-pass enhancement filter to enhance the texture features of the object and filter the noise in the camera imaging, and outputting a feature map; The basic Zero-DCE network includes 6 hidden layers and 1 output layer; Six of the hidden layers use symmetric skip connections similar to U-Net; All layers are ordinary 3x3 equal-length convolutional layers, and the output layer uses tanh as the activation function; The improved Zero-DCE network specifically includes: using the output of the third layer of the basic Zero-DCE and the result of the low-frequency processing of the output of the third layer as the input of the fourth layer; using the output of the fourth layer and the result of the low-frequency processing of the output of the second layer as the input of the fifth layer; using the output of the fifth layer and the result of the low-frequency processing of the output of the first layer as the input of the sixth layer; The simAM attention mechanism is set in the feature map, and each pixel position feature is used as a neuron; For a given neuron, its output is δ, and the outputs of surrounding neurons are δ e , the energy E(δ) is measured by calculating the linear combination difference between a neuron and its surrounding neurons: Among them, θ and σ represent the weight coefficients used to balance the difference and correlation degree in energy calculation, and The neuron outputs δ and δ are e The mean of The calculated energy function value is converted into attention weight through Sigmod activation function. The attention weight acts on the YOLOv11 network and is expressed as:
2. The method for detecting small targets in a dark environment as claimed in claim 1, characterized in that: The dark light enhancement includes enhancing the low light image using the improved Zero-DCE, which is composed of a plurality of curve estimation modules; The curve estimation modules are connected in series, each module is responsible for learning a curve transformation of the input image to enhance the lighting effect of the image; For each of the curve estimation modules, it is assumed that the illumination of the image is enhanced by learning a set of pixel-level curves; For each pixel in the input low-light image, the goal of the network is to learn a curve so that the pixel value is transformed by the curve to obtain the enhanced pixel value; The curve is a brightness enhancement curve, which is automatically learned by the network based on the input low-light image.
3. The method for detecting small targets in a dark environment as claimed in claim 2, characterized in that: The brightness enhancement curve formula is: L i (I(x,y);α)=I(x,y)+αI(x,y)(1-I(x,y)) Among them, I(x,y) represents a pixel value of the original image I with index (x,y), α∈[0,1] represents the brightness enhancement curve parameter, and L i (...) represents the pixel value of the image after the i∈[1,b)th order enhancement; b is an unknown number, representing an equation of any order.
4. The method for detecting small targets in a dark environment as claimed in claim 3, characterized in that: The steps of the filtering operation include: assuming that the improved Zero-DCE intermediate output feature map F i (x,y), where (x,y) represents the pixel coordinates of the feature map and i represents the hidden layer index of Zero-DCE; For the feature map F i (x,y) is transformed into the frequency domain by Fourier transform: Among them, W represents the width of the feature map, H represents the height of the feature map, (u, v) represents the coordinates in the frequency domain, and F i (u,v) represents the intensity of the frequency component in the image; Use Butterworth low-pass filter BLPF to enhance low-frequency information. The transfer formula of BLPF is: Among them, D(u,v) represents the distance from the feature point (u,v) to the center of the feature spectrum (W / 2,H / 2); D0 represents the cutoff frequency, which is used to control the degree of separation of high-frequency and low-frequency features. When designing the network structure, an adaptive pooling network is used instead of D0; n represents the filter order, which controls the degree of attenuation of high-frequency features. By multiplying the frequency domain value with the low-pass filter, low-frequency information enhancement is achieved. The low-frequency filter formula is: G i (u,v)=H(u,v)×F i (u,v) Among them, G i (u,v) represents the spectrum after filtering; Moderately amplify the low-frequency coefficients: G i ′=(1+k)G i (u,v) Where k is the enhancement coefficient, D(u,v)<D0, G i ′ represents the amplified filtering result; Finally, the inverse Fourier transform obtains the feature map after low-frequency enhancement filtering: Among them, g i (x,y) represents the feature map F i (x,y) is filtered through low frequency to get the feature map.
5. The method for detecting small targets in a dark environment as claimed in claim 4, characterized in that: The improved YOLOv11 network also includes: after using the dark light enhancement model to enhance the training set images, the enhanced images are input into the improved YOLOv11 network for training, and the network parameters are continuously adjusted so that the network can accurately detect small targets in dark light.
6. A small target detection system in a dark light environment using the method according to any one of claims 1 to 5, characterized in that: A collection unit, collecting original image information; The enhancement unit uses the improved Zero-DCE network to perform dark light enhancement on the original image; Detection unit, add simAM attention mechanism to the backbone feature extraction module of yolov11, input the enhanced image into the improved YOLOv11 network, and obtain the small target detection result; The improved Zero-DCE network includes, in the basic Zero-DCE network, adding a low-pass enhancement filter to enhance the texture features of the object and filter the noise in the camera imaging.
7. A computer device comprising: Memory and processor; The memory stores a computer program, wherein the processor implements the steps of any one of the methods of claims 1-5 when executing the computer program.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Road target fusion sensing method and system under weak light condition
CN115830567A
Dark light image enhancement method based on improved multi-scale Retinex algorithm
CN116612033A
Flame and smoke lightweight detection and monitoring early warning method and device and storage medium
CN118379679A