Low-illumination target detection network model and low-illumination target detection method

By adopting an end-to-end joint optimization framework in low-illumination object detection technology, the low-illumination enhancement module is deeply coupled with the target detection network, solving the problem of target deviation and error accumulation effects in the prior art, and achieving efficient low-illumination object detection performance.

CN120070912APending Publication Date: 2025-05-30NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510137938.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing low-illumination object detection technology has target deviation and error accumulation effects in image enhancement and object detection tasks, resulting in noise amplification and artifact generation, and reducing the feature discrimination ability of the detection network.

Method used

The end-to-end joint optimization framework is adopted to deeply couple the low-illumination enhancement module with the object detection network, and the two-way information interaction at the feature level is achieved through a differentiable architecture, and the image enhancement loss function and the object detection loss function are empowered and fused, and the weight coefficient is dynamically adjusted to adapt to task uncertainty.

Benefits of technology

It effectively improves the performance indicators of low-illumination object detection, improves feature separability and boundary positioning accuracy, reduces calculation complexity and delay, and ensures the efficient and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070912A_ABST
    Figure CN120070912A_ABST
Patent Text Reader

Abstract

The invention discloses a low-illumination target detection network model and a low-illumination target detection method.The model structure comprises an enhanced network, a backbone network, an improved neck network and a C2FDy module, the backbone network, the improved neck network and the C2FDy module are connected with the enhanced network, and the backbone network, the improved neck network and the C2FDy module are subjected to lightweight improvement based on a YOLOv8 model; embedded equipment can be deployed and an end-to-end detection network can be formed; and the enhancement network uses a gate function to control dynamic adjustment of brightness, details and feature weights of the enhanced low-illumination image so as to realize two-way information interaction of feature levels in the model structure. According to the method, the detection performance in a low-illumination scene can be improved, the detection requirements of a complex scene under the low-illumination condition can be met, efficient operation on embedded equipment is guaranteed through model lightweight design, the calculation time is short, the efficiency is high, the resource consumption is low, quick response can be achieved, and bidirectional information interaction of feature levels is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision processing, and particularly relates to a low-light target detection network model, a low-light target detection method, a device, and a storage medium. Background Art

[0002] The importance of low-light target detection technology in industrial applications has become increasingly prominent. Especially in scenarios such as night monitoring, intelligent security systems, and automated production, such as warehouses, production lines, ports, and logistics centers, the lighting conditions are often limited. Traditional image processing techniques are difficult to effectively identify and locate targets because images in low-light environments usually have problems such as noise, blurring, and insufficient contrast, which makes the target features unclear and seriously affects the effectiveness and security of automated and intelligent systems.

[0003] Different from general general target detection methods, existing low-light target detection generally uses low-light image enhancement algorithms for image data processing. These methods face two main problems:

[0004] (1) In computer vision tasks in low-light environments, traditional solutions usually regard image enhancement and target detection as separate serial tasks;

[0005] Although the low-light enhancement algorithms (such as RetinexNet, Zero-DCE, etc.) used in existing traditional solutions perform excellently in image quality evaluation metrics such as PSNR and SSIM, experimental studies have shown that simply optimizing the image quality at the pixel level cannot effectively improve the performance metrics (such as mAP) of downstream target detection tasks. This task decoupling paradigm has two key limitations: First, there is a significant target deviation between the objective function optimized by the enhancement network (such as L1 / L2 reconstruction loss) and the core requirements of the detection task (such as feature separability, boundary localization accuracy); Second, the error accumulation effect in the cascade system will lead to noise amplification and artifact generation during the enhancement process, which instead reduces the feature discrimination ability of the detection network.

[0006] To break through the above bottleneck, the present application proposes an end-to-end joint optimization framework, deeply couples the low-light enhancement module and the target detection network through a differentiable architecture, realizes two-way information interaction at the feature level, fuses the image enhancement loss function and the target detection loss function with weights, and adaptively adjusts the dynamic weight coefficient through task uncertainty.

[0007] (2) Although existing low-light enhancement algorithms (such as KinD, MBLLEN, etc.) have made significant progress in image quality restoration, their complex multi-stage processing architectures (usually containing dozens of convolutional modules and recursive structures) lead to a significant increase in computational latency. On the one hand, the excessive consumption of hardware resources limits the multi-task parallel processing ability. On the other hand, the uncontrollability of processing latency may result in the loss of key frames, directly affecting the temporal coherence of object detection.

[0008] To overcome this challenge, the low-light image enhancement module of the present invention adopts a lightweight design. This module only performs enhancement through 3 cascaded 3×3 residual modules. This design not only ensures efficient image quality restoration, but also the increase in computational volume and the number of parameters is very limited, significantly reducing the computational complexity and latency, and maintaining extremely high operating efficiency.

[0009] In summary, the present application specifically proposes a low-light object detection method to solve the above technical problems. Summary of the Invention

[0010] The main purpose of the present invention is to provide a low-light object detection method to solve the technical problems proposed in the background art.

[0011] The present invention adopts the following technical solutions to solve the above technical problems:

[0012] A low-light object detection network model, the model structure includes an enhancement network and a backbone network, an improved neck network, and a C2F_Dy module connected to the enhancement network. The backbone network, the improved neck network, and the C2F_Dy module are lightweight improved based on the YOLOv8 model, and can be deployed on embedded devices and form an end-to-end detection network;

[0013] The enhancement network uses a gate function to control the dynamic adjustment of the brightness, details, and feature weights of the enhanced low-light image to achieve two-way information interaction at the feature level in the model structure.

[0014] Preferably, the enhancement network includes a light estimation stage and an image enhancement stage, wherein:

[0015] In the light estimation stage, a convolutional layer with a convolution kernel size of 3×3, a convolutional layer with a convolution kernel size of 5×5, and a ReLU activation function are sequentially used to capture image information at different scales;

[0016] In the image enhancement stage, it is used to transfer the image details extracted by the light estimation module to the multi-layer convolution stacking operation. Each layer of convolution uses a 3×3 convolution kernel and combines an activation function for non-linear feature enhancement. Through the residual connection mechanism and the multi-layer convolution stacking operation, the feature expression ability is strengthened.

[0017] Preferably, the performance of the enhancement network is dynamically balanced by the weighting coefficients α and β to constrain the network output quality. For the dynamic adjustment of the feature weights, we have:

[0018] L_enhance = α × Fidelity Loss + β × Smooth Loss

[0019] Where L_enhance represents the enhancement loss function, Fidelity Loss represents the fidelity loss function, and the mean square error is used to calculate the difference between the input image and the enhanced image to measure the similarity between the enhanced image and the original image. Smooth Loss represents the smooth loss function, which effectively suppresses noise and unnatural boundaries by calculating gradients in different directions and applying weights. We have:

[0020]

[0021] Where, is the predicted value, representing the output of the model, y i is the true value, representing the input of the model, and N is the total number of pixels in the image. is the p-norm of the gradient between neighboring pixels. is the pixel value at position (i, j) in the predicted image, w ij is the weight calculated based on pixel differences, which encourages smaller difference values between adjacent pixels, and the weight w ij is calculated through the gradient difference between adjacent pixels in the image. The calculation formula is:

[0022]

[0023] Where ΔI ij is the gradient difference between adjacent pixels in the image, and σ is the smooth loss coefficient used to adjust the sensitivity of the weight.

[0024] Preferably, the low-light target detection network model is jointly optimized and trained using the total target detection loss function. During the process of jointly optimizing and training using the total target detection loss function:

[0025] The low-light image enhancement loss function L_enhance is combined with the classification loss function L_c1s, the bounding box regression loss function L_box, and the object confidence loss function L_obj in the target detection network to form the total loss function L_TOTAL, and the hyperparameter λ is introduced to guide the training optimization. We have:

[0026] L_TOTAL = λL_enhance + L_c1s + L_box + L_obj

[0027] Among them, the classification loss function \(L_{c1s}\) and the object confidence loss function \(L_{obj}\) are both binary cross-entropy losses for the conventional N objects. The bounding box regression loss function \(L_{box}\) is composed of the conventional IOU loss and DFL loss in the YOLOv8 system framework. The hyperparameter \(\lambda\) is a dynamic hyperparameter obtained by the Adam optimizer at a learning rate of 0.001.

[0028] Preferably, the backbone network is obtained by stacking 14 inverted residual modules, generating channel weights through global average pooling, dynamically weighting the specified features, and combining residual connections to avoid gradient vanishing in deep networks.

[0029] Preferably, the inverted residual module realizes channel expansion and compression through 1×1 pointwise convolution, and combines 3×3 depthwise separable convolution for efficient spatial feature extraction;

[0030] When the convolution stride of the inverted residual module is 1 and the number of input channels is the same as the number of output channels, the residual connection is enabled to enhance the training effect.

[0031] Preferably, the improved neck network adopts a cross-scale feature fusion module, using the output feature maps at different stages of the backbone network, specifically including:

[0032] a1. Perform channel compression and feature mapping through 1×1 convolution to ensure that features at each scale are fused within the same semantic space;

[0033] a2. Adopt layer-by-layer upsampling and concatenation operations to fuse the context information of high-level features and the semantic details of middle-level features, combined with the upsampling mechanism to reduce the computational overhead and ensure the transmission of spatial information;

[0034] a3. Introduce repeated convolution blocks to extract context information to enhance the interaction and complement of features at different scales;

[0035] a4. Perform feature reconstruction through element-wise addition and perform global feature mapping on the fused features to output an efficient fused feature representation.

[0036] Preferably, the C2F_Dy module divides the feature map into multiple subgroups through channel splitting and concatenation operations, introduces a Bottleneck structure on each subgroup for feature compression and extraction, and then realizes cross-group feature aggregation through the concatenation operation.

[0037] A low-light object detection method includes:

[0038] S1. Select and expand the low-light object detection dataset;

[0039] S2. Set up the low-light object detection network model as described above and train it;

[0040] S3. The trained low-light object detection network model inputs the detection image data for low-light object detection operations.

[0041] Preferably, the specific operation process in the S1 step includes:

[0042] S11. Select low-light images containing specified items in the ExDark dataset as the data source;

[0043] S12. Extract the images and labels of the specified categories in the existing VOC dataset according to the category index of the ExDark dataset, remove the redundant category label information and remap the category index, perform simulated low-light operations on the dataset, increase the number of dataset samples, and generate a VOC dataset consistent with the ExDark categories;

[0044] S13. The filtered VOC dataset and the low-light image dataset are evenly divided into a training set, a test set, and a validation set and the resolutions are unified to ensure data consistency.

[0045] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.

[0046] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, the memory storing a computer program, which when executed by the processor causes the processor to execute the steps of the above method.

[0047] As can be seen from the above technical solutions, the present invention provides a low-light object detection method and an object detection network model. Compared with the prior art, the present invention has the following advantages:

[0048] 1. In the object detection network model of the present invention, the inverted residual module realizes significant optimization in terms of the number of parameters and the amount of computation through the combination of expansion and compression operations, while maintaining the efficient feature extraction and multi-scale information fusion capabilities, and is suitable for application in lightweight object detection networks, especially on mobile devices and resource-constrained devices, significantly improving the detection performance of different-scale objects, and at the same time taking into account the efficiency and adaptability of the low-light object detection network model.

[0049] 2. In the object detection network model of the present invention, the improved neck network can fully combine detailed features and global context through lightweight design and an efficient information integration mechanism, effectively improving the performance of the network in multi-scale object detection tasks, especially showing excellent robustness and generalization in low-light scenarios.

[0050] 3. In the object detection network model of the present invention, the C2F_Dy module significantly enhances the feature expression ability of the model by splitting and splicing the channel dimension, without damaging the object information, while reducing the computational complexity, optimizing the efficiency of information flow, and achieving the balance between model lightweight and high performance.

[0051] 4. In the object detection network model of the present invention, the C2F_Dy module processes the input feature map and the output feature map in groups, and each group of convolutional kernels only acts on the corresponding input feature group, thereby realizing sparse connection and parallel processing. This not only greatly reduces the number of model parameters and computational complexity, but also retains the representation ability for complex features, and effectively improves the cross-scale feature expression and detection performance of the model.

[0052] 5. The low-light object detection network model framework of the present invention is improved based on the yolov8 model framework, which can improve the detection performance in low-light scenarios. It can not only meet the detection requirements of complex scenarios under low-light conditions, but also ensure efficient operation on embedded devices through model lightweight design, with fast calculation time, high efficiency and low resource consumption, and can respond quickly.

[0053] 6. The method of the present invention generates low-illumination visual characteristics similar to the ExDark dataset by applying different degrees of low-light image simulation operations to the image, which can provide diverse data samples, thereby enhancing the robustness and generalization of the low-illumination object detection network model in low-light environments.

[0054] 7. The method of the present invention realizes comprehensive feature analysis and dynamic enhancement control of the input image by setting up a light estimation module, and gradually improves the quality of low-light images through subsequent multi-layer convolution operations. Thus, it not only realizes multi-scale feature extraction, but also dynamically constrains the enhancement degree through a gating function mechanism to ensure the naturalness and robustness of the enhancement effect.

[0055] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Of course, any product implementing the present invention does not necessarily need to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The accompanying drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0057] Figure 1 is a schematic diagram of the overall structure of the low-illumination object detection network model of the present invention;

[0058] Figure 2 Schematic diagram of the low-light image enhancement network structure of the present invention;

[0059] Figure 3 Schematic diagram of the residual block network structure of the backbone network of the present invention;

[0060] Figure 4 Schematic diagram of the backbone network structure of the present invention;

[0061] Figure 5 Schematic diagram of the neck network structure of the present invention;

[0062] Figure 6 Schematic diagram of the C2F_Dy module network structure of the present invention;

[0063] Figure 7 Schematic diagram of the low-light target detection process of the present invention. Detailed implementation manners

[0064] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0065] In the embodiment, refer in detail to Figures 1 to 7 .

[0066] As Figure 1 shown, an object detection method using a low-light object detection network model is proposed in the embodiment of the present invention. The specific implementation process includes the following operation steps:

[0067] S1. Select and expand the low-light object detection data set. The specific operation process includes:

[0068] Select the low-light images containing the specified items in the ExDark data set as the main data source. This data set contains 12 types of objects, a total of 7363 low-light images, of which 4800 are used for training, 1563 are used for testing, and 1000 are used for verification. All images are uniformly adjusted to 640×640 pixels to ensure data consistency.

[0069] Select the low-light image data set ExDark, use the VOC image data set with the same detection categories, and perform the low-light simulation operation to generate a VOC data set with low-light visual characteristics.

[0070] Filter and extract the images and labels of the corresponding categories from the VOC dataset according to the category index of the ExDark dataset, remove the redundant category label information and the data of redundant categories.

[0071] Subsequently, remap the category index to ensure consistency between the two datasets. In the processed VOC dataset, 4,535 images are used for training, 1,065 for testing, and 986 for validation. All images are uniformly adjusted to 640×640 pixels to ensure data consistency.

[0072] Perform low-light image simulation operations based on the VOC dataset and the low-light image dataset to generate a VOC dataset with low-light visual characteristics, including:

[0073] Data content Training set Validation set Test set Remarks ExDark 4800 1000 1563 Low - illumination dataset Normal - illumination VOC 4535 986 1065 Rich generalization ability Low - light VOC 4535 986 1065 Simulated low - illumination environment

[0074] Among them, the low-light image simulation operation specifically includes image degradation, adding Gaussian noise and salt-and-pepper noise, adjusting brightness and contrast, etc., to generate low-illumination visual characteristics similar to the ExDark dataset, providing more diverse data samples for the training of the model, and enhancing the robustness and generalization of the model in low-light environments.

[0075] S2. Design the structure of the low-light image enhancement network model based on the low-illumination object detection dataset. As shown in Figure 2 , an enhancement network is composed of a low-light image enhancement module and a loss function to process the dataset data and construct an end-to-end low-illumination object detection network model;

[0076] Based on the Retinex theory, this enhancement network divides the illumination estimation and image quality optimization of low-light images into multiple processing stages to gradually optimize the illumination distribution. According to the theory, an image can be decomposed into a reflection component and an illumination component, and their relationship expression is as follows:

[0077] y = Z⊙X

[0078] Among them, y represents the input low-illumination image, Z represents the illumination estimation distribution, X represents the clear target image, and ⊙ represents pixel multiplication.

[0079] In the specific implementation process, the overall architecture of the low-light image enhancement module takes the illumination estimation module as the core, uses a gate function to control the dynamic adjustment settings of enhancing the brightness, details, and feature weights of low-light images. The illumination estimation module solves the illumination estimation distribution problem by introducing the parameter network H θ and a step-by-step residual learning mechanism. The iterative formula of this module is as follows:

[0080]

[0081] where xt Denotes the light estimation distribution result at the t-th stage. When t = 0, it represents the input low-light image y, H θ (x t ) represents the residual learning result of the parametric network for the current light estimation distribution result x t , and u t represents the residual term output by the network, which is used as the light compensation amount to further optimize the light estimation distribution result, x t+1 Denotes the light estimation distribution result at the (t + 1)-th stage. Through multiple iterations, the light estimation distribution result x t will be gradually optimized at each stage, and on the basis of the previous stage result, further improve the brightness, contrast and detail quality of the image.

[0082] This step-by-step optimization method avoids the difficulty of one-time estimation and helps the network converge more stably to the ideal light estimation distribution result.

[0083] At this time, the working process of the low-light image enhancement network is divided into two main stages: the light estimation stage and the image enhancement stage, where:

[0084] (a) In the light estimation stage, the input image undergoes multi-branch feature extraction by the light estimation module. The module performs multi-scale feature extraction on the input image through parallel 3×3 and 5×5 convolutional kernels. The former focuses on capturing local detail features, and the latter focuses on overall light features;

[0085] At this time, the light optimization part not only analyzes the brightness, contrast and light distribution of the input image, but also uses the gate function mechanism to dynamically adjust the weights of the extracted features, so as to provide guidance and constraints for the subsequent image enhancement stage;

[0086] (b) In the image enhancement stage, it is used to transfer the image details extracted by the light estimation module to the multi-layer convolution stacking operation to gradually optimize the brightness, contrast and detail quality of the input image;

[0087] At this time, each layer of convolution uses a 3×3 convolutional kernel and combines an activation function for non-linear feature enhancement. At the same time, through the residual connection mechanism and the multi-layer convolution stacking operation, the residual connection mechanism ensures that the original characteristics of the input image are retained, and the features of the multi-layer convolution stacking are fused through the splicing operation to gradually strengthen the expression ability of local and global features. A three-time residual structure is adopted, and the input and output weight ratios of each scale residual are dynamically adjusted through the gate function in the light estimation module.

[0088] It should also be noted that the gate function mechanism in the light estimation module plays a role of dynamic regulation in the whole image enhancement process. According to the light distribution and feature importance of the input image, the gate function dynamically adjusts the degree of feature enhancement, so as to achieve precise control of the enhancement intensity. This design enables image enhancement to not only adaptively adjust according to specific low-light characteristics, but also effectively avoid image distortion caused by over-enhancement.

[0089] Moreover, the finally enhanced feature map generates the final enhancement result through the output convolutional layer, and the Sigmoid activation function is used to limit the output pixel values within the range of [0, 1], ensuring that the brightness and detail distribution of the image conform to the natural visual law. At the same time, the output image generates a final enhancement result with stronger contrast, higher brightness and richer details by adding it to the original input image.

[0090] Generally speaking, this module realizes comprehensive feature analysis and dynamic enhancement control of the input image through the part of light optimization, and gradually improves the quality of low-light images by combining subsequent multi-layer convolutional operations. In addition, this part not only realizes multi-scale feature extraction, but also dynamically constrains the enhancement degree through the gate function mechanism, ensuring the naturalness and robustness of the enhancement effect.

[0091] Furthermore, the loss function of the low-light image enhancement module includes Fidelity Loss and Smooth Loss. By combining these two losses with weights, the training process can be accelerated and the image enhancement effect can be improved. Therefore, the enhancement loss function is set to dynamically balance the performance of the enhancement network through the weighting coefficients α and β to constrain the output quality of the network, so that the training process can be effectively accelerated and the enhancement effect can be guaranteed in terms of quality. There is:

[0092] L_enhance = α × Fidelity Loss + β × Smooth Loss

[0093] Where:

[0094] L_enhance represents the enhancement loss function;

[0095] Fidelity Loss represents the fidelity loss function, which uses the mean square error to calculate the difference between the input image and the enhanced image to measure the similarity between the enhanced image and the original image, ensuring that the quality of the enhanced image does not deviate too much from the original image;

[0096] Smooth Loss represents the smooth loss function, which calculates the gradients in different directions and gives weights to effectively suppress noise and unnatural boundaries, so as to achieve the protection and enhancement of image details.

[0097] For the fidelity loss function and the smooth loss function, there is further:

[0098]

[0099] Among them, is the predicted value, expressed as the output of the model, y i is the true value, expressed as the input of the model, and N is the total number of pixels in the image. is the p-norm of the gradient between neighboring pixels. is the pixel value at position (i, j) in the predicted image, w ij is the weight calculated based on pixel differences, which encourages smaller difference values between adjacent pixels.

[0100] In addition, in actual calculations, the smooth loss first converts the image to the YCbCr color space, focuses on the Y channel, and is used to calculate the pixel gradient of the enhanced image. By smoothing the image boundary, the noise and abrupt boundaries in the enhanced image are reduced, thereby ensuring that the image has a natural transition effect. At this time, larger pixel differences will be given smaller weights, while smaller pixel differences will obtain larger weights. Therefore, the weight w ij is calculated through the gradient difference between adjacent pixels in the image, and the calculation formula is:

[0101]

[0102] Among them, ΔI ij is the gradient difference between adjacent pixels in the image, and σ is the smooth loss coefficient, which is used to adjust the sensitivity of the weight.

[0103] And in this way, the details in the image are effectively protected, and the noise and unnatural boundaries are effectively suppressed.

[0104] S3. Improve the object detection network model, including a backbone network connected to a low-light image enhancement module, and improve the neck network for lightweight improvement, so that it can be deployed on embedded devices and form an end-to-end detection network.

[0105] Based on the improvement of the YOLOv8 model, as Figure 7 shown, applied to the low-light object detection method mentioned in the above embodiments, including a backbone network connected to an enhancement network, an improved neck network, and a C2F_Dy module. Among them, the backbone network, the improved neck network, and the C2F_Dy module are lightweight improved around yolov8, and can be deployed on embedded devices and form an end-to-end detection network. There is:

[0106] (1) In the improved general object detection network model, the core module of the lightweight object detection backbone network is the inverted residual module. Its design achieves a good balance between computational efficiency and feature extraction ability. It realizes channel expansion and compression through 1×1 pointwise convolution, reduces the computational overhead while increasing the feature dimension, and combines 3×3 depthwise separable convolution for efficient spatial feature extraction, significantly reducing the computational complexity;

[0107] In addition, the module also integrates a channel attention mechanism and residual connections, further enhancing the network's expressive power and stability while effectively extracting key features.

[0108] Specifically, the backbone network is obtained by stacking 14 inverted residual modules. It generates channel weights through global average pooling, dynamically weights the specified features, and combines residual connections to avoid the vanishing gradient in deep networks. At the same time, the hierarchical structure is clear. The network takes the enhanced image as input. Through the combination of expansion and compression operations, the inverted residual module has achieved significant optimization in terms of the number of parameters and computational volume, while maintaining the ability of efficient feature extraction and multi-scale information fusion. It is suitable for application in lightweight object detection networks, especially on mobile devices and resource-constrained devices, significantly improving the detection performance of different-scale objects while taking into account the efficiency and adaptability of the model.

[0109] Furthermore, the processing flow of the inverted residual module refers to the following table and Figure 4 as shown, specifically including:

[0110] L1. The input of the inverted residual module first undergoes channel expansion through a 1×1 pointwise convolution, and is combined with batch normalization and activation functions for non-linear transformation, thus significantly enhancing the feature expression ability;

[0111] L2. Subsequently, the module uses a 3×3 depthwise separable convolution to further extract spatial features while significantly reducing the computational volume. This process also combines batch normalization and activation functions to ensure computational efficiency while retaining the necessary feature integrity;

[0112] At this time, in order to further enhance the feature expression ability, the inverted residual module integrates a channel attention mechanism. It generates global feature expressions between channels through global average pooling, and generates weight factors through two fully connected layers, respectively combined with ReLU and Sigmoid activation functions, and multiplies them with the input features channel by channel to achieve dynamic weighting of key channels, thus effectively enhancing the feature expression ability. At this time, a residual block is formed as Figure 3 as shown;

[0113] L3. Finally, another 1×1 pointwise convolution is used to achieve channel compression and is combined with batch normalization processing.

[0114]

[0115]

[0116] At this time, it should be particularly noted that when the convolution stride in the reverse residual module is 1 and the number of input channels is the same as the number of output channels, the residual connection is enabled between this layer and the previous layer. This design avoids the problem of gradient disappearance in the training of deep networks, effectively improves the expression ability of the network, and ensures the stability and training effect of the network.

[0117] (2) The improved neck network adopts a cross-scale feature fusion module, which utilizes the output feature maps at different stages of the backbone network, such as Figure 5 as shown, specifically including:

[0118] a1. Channel compression and feature mapping are performed through 1×1 convolution to ensure that features at each scale are fused within the same semantic space;

[0119] a2. Adopt layer-by-layer upsampling and splicing operations to fuse the context information of high-level features and the semantic details of middle-level features, combined with the upsampling mechanism (such as bilinear interpolation), to reduce the computational overhead and ensure the transmission of spatial information;

[0120] a3. Introduce repeated convolution blocks to extract context information to enhance the interaction and complement of features at different scales;

[0121] a4. Feature reconstruction is performed through element-wise addition, and global feature mapping is performed on the fused features to output an efficient fused feature representation.

[0122] At this time, the cross-scale feature fusion module can fully combine detailed features and global context through lightweight design and efficient information integration mechanism, effectively improving the performance of the network in multi-scale object detection tasks, especially showing excellent robustness and generalization in low-light scenarios and small object detection.

[0123] (3) The C2F_Dy module divides the feature map into multiple subgroups through channel splitting and splicing operations, and introduces a Bottleneck structure on each subgroup for feature compression and extraction. Subsequently, cross-group feature aggregation is achieved through splicing operations, thereby realizing feature extraction and fusion, significantly improving the expression ability and computational efficiency of the model. Refer to Figure 6 , specifically including:

[0124] b1. The input feature map undergoes a convolution operation, and the number of channels is expanded to 2 times c1;

[0125] b2. The feature map is divided into two groups in the channel dimension through channel splitting, and the size of each group of feature maps is (1, c1, 256, 256), (1, c1, 256, 256), (1, c1, 256, 256);

[0126] b3. Introduce a Bottleneck neck structure on each subgroup to further extract and compress features, and maintain the spatial dimension of the output features through N times of processing;

[0127] b4. The feature map after being processed by the Bottleneck neck structure is concatenated with the original segmentation feature map in the channel dimension to achieve cross-group feature aggregation. The size of the concatenated feature map is (1, (2 + N) * 64, 256, 256), (1, (2 + N) * 64, 256, 256), (1, (2 + N) * 64, 256, 256);

[0128] b5. Compress the number of channels to c2 through the second convolution operation, and output the efficiently fused feature map.

[0129] At this time, the C2F_Dy module significantly enhances the feature expression ability of the model without destroying the target information by splitting and concatenating the channel dimension, while reducing the computational complexity, optimizing the efficiency of information flow, and achieving the balance between model lightweight and high performance.

[0130] At the same time, the C2F_Dy module combines the advantages of 3×3 convolution kernels and 1×1 convolution kernels, which are responsible for capturing rich spatial information and efficient channel information fusion respectively, and significantly reduces the number of parameters and computational complexity through group convolution technology, optimizing the efficiency of information flow;

[0131] Among them, the 3×3 convolution kernel effectively captures rich spatial information and improves the spatial perception ability of features, while the 1×1 convolution kernel realizes efficient channel information fusion through low-computation-cost channel interaction and feature integration;

[0132] On this basis, the module introduces group convolution technology (Group Convolution), which processes the input feature map and output feature map in groups, and each group of convolution kernels only acts on the corresponding input feature group, so as to achieve sparse connection and parallel processing.

[0133] Therefore, the system framework structure of the C2F_Dy module not only greatly reduces the number of parameters and computational complexity of the model, but also retains the ability to represent complex features, effectively improves the cross-scale feature expression and detection performance of the model. In addition, the channel splitting and recombination mechanism promotes the information flow between features, provides the model with stronger feature fusion and expression ability, and finally achieves the balance between model lightweight and high performance.

[0134] In summary, the low-light object detection network model framework is improved based on the YOLOv8 model framework, which can enhance the detection performance in low-light scenarios. It can not only meet the detection requirements of complex scenarios under low-light conditions, but also ensure efficient operation on embedded devices through model lightweight design. It has the advantages of fast calculation time, high efficiency, and low resource consumption, can respond quickly, provides strong technical support for industrial intelligent applications in multiple fields, and can achieve two-way information interaction at the feature level in the model structure.

[0135] At the same time, as shown in the following table, compared with the existing YOLOV8 model, the structure of this application is more lightweight, can perform rapid operation and response on the basis of greatly reducing the number of data layers and parameters, and is more convenient to use.

[0136] Model name Number of layers Number of parameters (in millions) Computational volume (GFLOPs) YOLOV8n 225 315 8.9 Low - illumination object detection model n 353 200 5.6 YOLOV8s 225 1116 28.8 Low - illumination object detection model s 353 346 12.1

[0137] S4. Use the low-light object detection network model and perform joint tuning training using the total object detection loss function.

[0138] Specifically, in the process of joint tuning training, by combining the low-light image enhancement module and the object detection network, a comprehensive optimization loss function design is used, and the cooperation and dynamic feedback mechanism between modules are realized through end-to-end training. The core of this design is to combine the low-light image enhancement loss function L_enhance with the classification loss function L_c1s, the bounding box regression loss function L_box, and the object confidence loss function L_obj in the object detection network to form the total loss function L_TOTAL, and introduce the hyperparameter λ to guide the training to adjust the weight of the image enhancement module loss term.

[0139] At this time

[0140] So there is:

[0141] L_TOTAL = λL_enhance + L_c1s + L_box + L_obj

[0142] Among them, the classification loss function L_c1s and the object confidence loss function L_obj are both binary cross-entropy losses for N conventional objects, and the bounding box regression loss function L_box is composed of the conventional IOU loss and DFL loss in the YOLOv8 system framework;

[0143] In addition, for the rapid convergence of training, the dynamic hyperparameter Adam optimizer is used, and the initial parameter values are as follows:

[0144] Learning rate (lr): 0.001;

[0145] β1 (decay rate of the first moment estimate): 0.9;

[0146] β2 (Decay rate of second moment estimation): 0.999;

[0147] ε (Small constant to prevent division by zero): 1e-7.

[0148] This method realizes the cooperation and dynamic feedback mechanism between modules through end-to-end training by combining a low-light image enhancement module and a target detection network, using a comprehensively optimized loss function design;

[0149] Moreover, it should be further noted that the image enhancement module aims to optimize image brightness, contrast, and detail quality, while the target detection module reversely guides the image enhancement module to dynamically adjust its strategy and intensity by improving classification, regression, and confidence performance. Through this collaborative feedback mechanism, the network can achieve dual optimization of enhancement effects and detection performance. In addition, by introducing a dynamic hyperparameter adjustment mechanism, the training efficiency and effect can be further improved.

[0150] For example, in the initial stage of training, the weight of λ can be dynamically increased to prioritize optimizing the image enhancement module to quickly improve the quality of the input image; while in the later stage of training, λ is gradually decreased, and the focus of optimization is shifted to the detection module, making the model pay more attention to detection performance.

[0151] In summary, this low-light target detection network model is based on end-to-end low-light target detection of the image enhancement module, fully combining the design characteristics of low-light image enhancement and lightweight target detection networks. The overall architecture can not only meet the detection requirements of complex scenes under low-light conditions but also ensure efficient operation on embedded devices through model lightweight design, providing strong technical support for industrial intelligent applications in multiple fields.

[0152] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor, causes the processor to execute the steps of the above method.

[0153] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above method.

[0154] In yet another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when run on a computer, causes the computer to execute any of the low-light target detection methods in the above embodiments.

[0155] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention, and the explanations, examples, and beneficial effects of related content can refer to the corresponding parts in the above method.

[0156] An embodiment of the present application further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus.

[0157] The memory is used to store a computer program.

[0158] When the processor is used to execute the program stored on the memory, the above-mentioned low-light target detection method is implemented.

[0159] The communication bus mentioned in the above electronic device may be a peripheral component interconnect standard (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0160] The communication interface is used for communication between the above electronic device and other devices.

[0161] The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0162] The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0163] It should also be noted that the electronic device further includes a terminal device, and the terminal device may also be referred to as a terminal, a user equipment, a mobile station, a mobile terminal, etc. The terminal device may be a mobile phone, a smart TV, a wearable device, a tablet computer, a computer with wireless transceiver function, a virtual reality terminal device, an augmented reality terminal device, a wireless terminal in industrial control, a wireless terminal in unmanned driving, a wireless terminal in remote surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device.

[0164] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid-state drive SSD), etc.

[0165] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

[0166] In addition, it should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.

[0167] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the meaning of "and / or" appearing throughout the text includes three parallel scenarios. Taking "A and / or B" as an example, it includes scenario A, scenario B, or the scenario where A and B are satisfied simultaneously. In addition, in the embodiments of the present invention, "a plurality" means two or more. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

Claims

1. A low-light target detection network model, the model structure includes an enhanced network and a backbone network connected to the enhanced network, an improved neck network and a C2F_Dy module, characterized in that: The backbone network, improved neck network and C2F_Dy module are lightweight and improved based on the YOLOv8 model; The enhancement network uses a gate function to control the dynamic adjustment of the brightness, details and feature weights of the enhanced low-light image to achieve two-way information interaction at the feature level in the model structure.

2. The low-light target detection network model according to claim 1, characterized in that: The enhancement network includes an illumination estimation stage and an image enhancement stage, wherein: In the illumination estimation stage, a convolutional layer with a convolution kernel size of 3×3, a convolutional layer with a convolution kernel size of 5×5, and a ReLU activation function are used in sequence to capture image information of different scales; The image enhancement stage is used to pass the image details extracted by the illumination estimation module to the multi-layer convolution stacking operation. Each layer of convolution uses a 3×3 convolution kernel and combines the activation function for nonlinear feature enhancement. The feature expression capability is enhanced through the residual connection mechanism and multi-layer convolution stacking operations.

3. The low-light target detection network model according to claim 2, characterized in that: In the enhanced network, the performance of the enhanced network is dynamically balanced by weight coefficients α and β to constrain the network output quality and the dynamic adjustment of feature weights, as follows: L_enhance=α×Fidelity Loss+β×Smooth Loss Among them, L_enhance is represented as the enhancement loss function, Fidelity Loss is represented as the fidelity loss function, and the mean square error is used to calculate the difference between the input image and the enhanced image to measure the similarity between the enhanced image and the original image. SmoothLoss is represented as the smooth loss function, which calculates the gradients in different directions and gives weights to effectively suppress noise and unnatural boundaries. There are: in, is the predicted value, expressed as the output of the model, y i is the true value, which is represented as the input of the model, N is the total number of pixels in the image, is the p-norm of the gradient between pixels in the neighborhood, is the pixel value at position (i, j) in the predicted image, w ij is a weight calculated based on pixel differences, which encourages adjacent pixels to have smaller differences, and the weight w ij The gradient difference between adjacent pixels in the image is calculated using the following formula: Among them, ΔI ij is the gradient difference between adjacent pixels in the image, and σ is the smoothing loss coefficient, which is used to adjust the sensitivity of the weight.

4. The low-light target detection network model according to claim 3, characterized in that: The low-light target detection network model uses the target detection total loss function for joint tuning training, wherein during the joint tuning training using the target detection total loss function: The low-light image enhancement loss function L_enhance is combined with the classification loss function L_c1s, the bounding box regression loss function L_box and the target confidence loss function L_obj in the target detection network to form the total loss function L_TOTAL, and the hyperparameter λ is introduced to guide the training and tuning, as follows: L_TOTAL=λL_enhance+L_c1s+L_box+L_obj Among them, the classification loss function L_c1s and the target confidence loss function L_obj are both binary cross entropy losses of conventional N targets. The bounding box regression loss function L_box is composed of the conventional IOU loss and DFL loss of the YOLOv8 system framework. The hyperparameter λ is a dynamic hyperparameter obtained by the Adam optimizer at a learning rate of 0.

001.

5. The low-light target detection network model according to claim 1, characterized in that: The backbone network is obtained by stacking 14 reverse residual modules, generating channel weights through global average pooling, dynamically weighting the specified features, and combining residual connections to avoid gradient disappearance in deep networks.

6. The low-light target detection network model according to claim 5, characterized in that: The reverse residual module implements channel expansion and compression through 1×1 point-by-point convolution, combined with 3×3 depth-wise separable convolution for efficient spatial feature extraction; When the convolution stride of the reverse residual module is 1 and the number of input channels is the same as the number of output channels, the residual connection is enabled to enhance the training effect.

7. The low-light target detection network model according to claim 1, characterized in that: The improved neck network adopts a cross-scale feature fusion module and uses the output feature maps of different stages of the backbone network, specifically including: a1. Perform channel compression and feature mapping through 1×1 convolution to ensure that features of different scales are integrated in the same semantic space; a2. Adopt layer-by-layer upsampling and concatenation operations to fuse the contextual information of high-level features with the semantic details of mid-level features, combined with upsampling mechanisms to reduce computational overhead and ensure the transmission of spatial information; a3. Introduce repeated convolution blocks to extract contextual information to enhance the interaction and complementation of features at different scales; a4. Reconstruct features by element-by-element addition, perform global feature mapping on the fused features, and output an efficient fused feature representation.

8. The low-light target detection network model according to claim 1, characterized in that: The C2F_Dy module divides the feature map into multiple subgroups through channel segmentation and splicing operations, introduces a Bottleneck structure on each subgroup for feature compression and extraction, and then realizes cross-group feature aggregation through splicing operations.

9. A low-light target detection method, characterized in that: include: S1. Select and expand low-light target detection dataset; S2. Setting and training a low-light target detection network model as described in any one of claims 1 to 8 above; S3. The trained low-light target detection network model inputs the detection image data to perform low-light target detection operations.

10. The low-light target detection method according to claim 9, characterized in that: The specific operation process in step S1 includes: S11. Select low-light images containing specified objects in the ExDark dataset as data sources; S12. Extract the images and labels of the specified categories in the existing VOC dataset according to the category index of the ExDark dataset, remove the redundant category label information and remap the category index, simulate the low-light operation on the dataset, increase the number of dataset samples, and generate a VOC dataset consistent with the ExDark category; S13. The screened VOC dataset and low-light image dataset are divided into training set, test set and validation set, and the resolution is unified to ensure data consistency.