A method for constructing a lightweight target detection network adaptive to dark light scenes

CN119091347BActive Publication Date: 2026-09-25SHANGHAI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411136957.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-09-25
Estimated Expiration
2044-08-19

AI Technical Summary

Technical Problem

由于光线不足,图像中的目标往往难以清晰地显示出来,导致目标边缘模糊、颜色纹理信息丢失等问题,目标检测的准确性会受到严重影响,从而影响整个视觉系统的性能

Benefits of technology

[0035]1.本发明在引入特征增强模块时,采用三条路径并行输入的方式,以融合不同尺度的特征信息,可以获得鲁棒性更强、语义信息更丰富的特征,经过特征增强网络模块之后所获得的特征图更容易融合;并且在信息提取网络中,设有残差连接,保证重要信息不丢失的同时,扩充了特征图的空间信息,在编码与解码的过程中,丰富了图像的色彩和纹理信息。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091347B_ABST
    Figure CN119091347B_ABST
Patent Text Reader

Abstract

A method for constructing a lightweight target detection network suitable for dark light scenes, comprising the steps of constructing a training set, a validation set and a test set, training and validating the model, and the model training process comprising the following steps: S1. The original image F is sequentially subjected to two convolution layers to eliminate redundant information, to obtain feature maps F1 and F2; S2. The original image F, the feature map F1 and the feature map F2 are input into a feature enhancement network in parallel for semantic information coding to obtain output feature maps O, O1 and O2 on three paths; S3. The output feature map O2 is subjected to a layer of up-sampling network and is spliced with the output feature map O1 in the channel dimension, and then is subjected to once more up-sampling to adjust the channel dimension, to obtain a spliced convolution feature map; the output feature map O and the spliced convolution feature map are added element by element to obtain an intermediate image after image enhancement by decoding; and S4. MobileViT is adopted to extract features from the intermediate image, the extracted feature information of different scales is fused and then is input into a detection head, and a detection result is predicted by the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of target detection, specifically relating to a method for constructing a lightweight target detection network adapted to low-light scenes, a method for target detection using the constructed network, and corresponding electronic products. Background Technology

[0002] In recent years, the development of deep learning has enabled object detection tasks to achieve good results on most benchmark datasets. However, most existing detection networks are studied under normal conditions. In real-world environments, there are often many harsh environmental conditions, such as nighttime and exposure issues. Therefore, object detection under low-light conditions has been a long-standing and continuously important challenge in the field of computer vision. Due to insufficient light, objects in images are often difficult to display clearly, leading to problems such as blurred object edges and loss of color and texture information. This severely affects the accuracy of object detection, thereby impacting the performance of the entire vision system.

[0003] Studies have shown that by appropriately enhancing images and recovering more latent information about the original blurred targets based on environmental conditions, target detection models can adapt to different low-light conditions, which is significant for improving the performance and stability of computer vision systems. Furthermore, during deployment, lightweighting the model reduces the burden on terminal devices, allowing it to better adapt to resource-constrained environments such as embedded devices and mobile devices. Therefore, this invention provides a method for target detection in low-light scenes using feature enhancement networks and lightweight feature extraction networks, enabling accurate detection of the target's specific location information even under low-light conditions. Summary of the Invention

[0004] To address the problems and shortcomings of existing technologies, this invention provides a lightweight method for constructing a target detection network in low-light scenarios.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] In a first aspect, the present invention provides a method for constructing a lightweight target detection network in low-light scenes, comprising the following steps:

[0007] (1) Collect single-view video in low-light scenes, divide the image sequence, and after labeling the images, divide them into training set, validation set and test set;

[0008] (2) The training set is used for model training. By continuously optimizing the overall network loss function and adjusting the weights of the trained model, a model that can perform object detection in low-light scenes is obtained after training is completed.

[0009] (3) The model obtained in step (2) is evaluated using a validation set and a test set. The validation set is used to evaluate the parameter indicators of the model, and the test set is used to verify the detection performance of the model.

[0010] The specific process of model training in step (2) is as follows:

[0011] S1. The original image F is passed through two convolutional layers to eliminate redundant information, resulting in feature map F1 and feature map F2; the original image is represented by the data size F∈R. H×W×3 The feature maps have data sizes of 1 / 4 and 1 / 8 of the original image, respectively, and are represented as follows: and

[0012] S2. Input the original image F, feature map F1, and feature map F2 into the feature enhancement network module in parallel to obtain the output feature maps O, O1, and O2 on the three paths, denoted as O∈R. H×W×32 ,

[0013] S3. The output feature map O2 is passed through an upsampling network and concatenated with the output feature map O1 along the channel dimension. Then, it is upsampled again to adjust the channel dimension, resulting in a concatenated convolutional feature map. The output feature map O and the concatenated convolutional feature map are added element by element, and the image is decoded to obtain the enhanced intermediate image.

[0014] S4. The lightweight feature extraction backbone network MobileViT is used to extract features from the intermediate image. The extracted feature information at different scales is fused at different levels to obtain a multi-layer fused feature map. The multi-layer fused feature map is input into the detection head, and the network predicts the detection result.

[0015] Preferably, the method for dividing the training set, validation set, and test set in step (1) is as follows:

[0016] a. Acquire raw image sequences: The video was acquired in a low-light environment and divided into 10 image sequences;

[0017] b. Create image labels: Manually label the objects in the image using the labelImg tool, and then export the image label file;

[0018] c. Divide the collected image sequences into training, validation, and test sets according to a 7:1:2 ratio.

[0019] Furthermore, the scene data collected in step a is acquired by the ZED next-generation binocular camera, and the corrected left camera video is obtained through the camera software package.

[0020] Preferably, in step S1, redundant information is eliminated through two convolutional layers, and the kernel size of each convolutional layer is 3, expressed by the formula:

[0021] F1 = Conv(F1), S = 4,

[0022] F2 = Conv(F1), S = 2,

[0023] In the formula, Conv represents the convolution operation, and S represents the stride of the convolution.

[0024] Preferably, the feature enhancement network described in S2 consists of six CBR modules with residual connections, wherein the first and sixth layers have residual connections, the second and fifth layers have residual connections, and the third and fourth layers have residual connections. Each CBR module performs Cov, BN, and ReLU operations in sequence. The kernel size of the Cov convolution operation is 3, and the stride is 1. The size of the feature map is not changed during the encoding process, and the number of channels is increased from 3 to 32, integrating the extracted information into the channel dimension.

[0025] Preferably, the process described in S3 is expressed by the following mathematical formula:

[0026]

[0027] In the formula, the upsample operation represents upsampling, the concat operation represents concatenation along the channel dimension, and the add operation represents element-wise addition of feature maps. This represents the feature map of the final output.

[0028] Preferably, the specific method for feature extraction in S4 is as follows: each region of the intermediate image is encoded into three vectors, namely query vector Q, key vector K and value vector V; the similarity between Q of each region and K of the other regions is calculated; and the effective information is extracted by weighted summation according to the weights.

[0029] Preferably, the overall network loss function in step (2) is defined as the reconstruction loss L1 and the detection loss L2. d The reconstruction loss is mainly used to constrain the information loss of the input image after it passes through the feature enhancement network; the detection loss is mainly used to constrain the detection loss of the object detection network, including the object loss L. obj Category loss L cls and regression loss L loc This can be expressed as a formula:

[0030] L=λ1L1+λ d L d =λ1L1+λ obj L obj +λcls L cls +λ loc L loc

[0031] Wherein, λ1, λ obj , λ cls and λ loc These represent the weighting coefficients for each part of the loss.

[0032] Secondly, a target detection method is provided, the steps of which include inputting the image to be detected into the target detection network constructed in the first aspect for target detection.

[0033] Thirdly, an electronic product is provided, including a processor and a memory communicatively connected to the processor; the memory is used to store instructions; the processor is used to invoke the instructions stored in the memory to execute the target detection method described in the second aspect.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] 1. When introducing the feature enhancement module, this invention adopts a three-path parallel input method to fuse feature information at different scales, which can obtain features with stronger robustness and richer semantic information. The feature map obtained after passing through the feature enhancement network module is easier to fuse. Furthermore, residual connections are set in the information extraction network to ensure that important information is not lost while expanding the spatial information of the feature map, enriching the color and texture information of the image during the encoding and decoding process.

[0036] 2. This invention obtains an intermediate image after preliminary decoding during the decoding stage. During the training process, the effect of the preliminary enhancement can be monitored more intuitively, solving the problem of achieving accurate and stable detection when the semantic information of the image is not obvious.

[0037] 3. This invention employs the lightweight feature extraction backbone network MobileViT, which not only has excellent feature extraction capabilities but also allows the model to better adapt to resource-constrained environments such as embedded devices and mobile devices.

[0038] 4. The detection method of this invention is mainly designed for stable and accurate target detection in poor external environments and insufficient light. It utilizes a feature enhancement network to improve the overall training speed of the model and employs the lightweight feature extraction backbone network MobileViT, which not only has good feature extraction capabilities but also allows the model to better adapt to resource-constrained environments such as embedded devices and mobile devices. This enables the network to be applied to autonomous robots equipped with vision systems, allowing them to better cope with low-light scenes, ensure the integrity of image information, and improve the accuracy and robustness of robot detection capabilities. Attached Figure Description

[0039] Figure 1 This is an overall framework diagram of a lightweight target detection network adapted to low-light scenes in a specific embodiment of the present invention;

[0040] Figure 2 This is a schematic diagram of the feature enhancement network module in a specific embodiment of the present invention;

[0041] Figure 3 This is a schematic diagram of the decoding module in a specific embodiment of the present invention. Detailed Implementation

[0042] The present invention will now be described in detail with reference to specific embodiments.

[0043] A lightweight object detection network adapted to low-light scenes, the overall framework of which is as follows: Figure 1 As shown, the construction method includes the following steps:

[0044] (1) Single-view video of low-light scenes is acquired using a stereo camera. Based on the video acquired by the left camera, an image sequence dataset is constructed and divided into a training set, a validation set, and a test set, as follows:

[0045] a. Acquire raw image sequences: Equip the ZED next-generation binocular camera on the vision platform, initialize the device, ensure that the scene data to be acquired is in a low-light environment, obtain the corrected left camera video through the camera program package, and divide the acquired video into 10 image sequences.

[0046] b. Create image labels: Manually label the objects in the images using the labelImg tool, export the image label files, and ensure that the dataset provided for network training is of high quality.

[0047] c. Divide the dataset into training, validation, and test sets: Divide the collected image sequences into training, validation, and test sets in a 7:1:2 ratio.

[0048] (2) The model is trained using a training set. By continuously optimizing the overall network loss function and adjusting the weights of the trained model, a model capable of object detection in low-light scenes is obtained after training. Specifically:

[0049] S1. The original image is represented by a data size F∈R. H×W×3 The original image is passed through two convolutional layers sequentially to eliminate redundant information: the first convolutional layer has a kernel size of 3 and a stride of 4, resulting in feature map F1. Feature map F1 is reduced to 1 / 4 of its original size while the channel dimension remains unchanged. Feature map F1 is then passed through a second convolutional layer with a kernel size of 3 and a stride of 2, resulting in feature map F2. Feature map F2 is reduced to 1 / 8 of its original size while maintaining the same channel dimensions. The above process can be expressed by the following formula:

[0050] F1 = Conv(F), S = 4,

[0051] F2 = Conv(F1), S = 2,

[0052] In the formula, Conv represents the convolution operation, and S represents the stride of the convolution.

[0053] S2. Input the original image F, feature map F1, and feature map F2 into the feature enhancement network module in parallel to obtain the output feature maps O, O1, and O2 on the three paths, denoted as O∈R. H×W×32 ,

[0054] Feature enhancement network modules such as Figure 2 As shown, the system consists of six CBR modules with residual connections, which encode semantic information and extract important information. Each CBR module sequentially performs Cov, BN, and ReLU operations. The Cov convolution operation has a kernel size of 3 and a stride of 1, and the size of the feature map is not changed during the encoding process. The number of channels is increased from 3 to 32, integrating the extracted information into the channel dimension. In the feature enhancement network module, residual connections are set between the first and sixth layers, the second and fifth layers, and the third and fourth layers to ensure that important information is not lost.

[0055] S3. The output feature map O2 is concatenated with the output feature map O1 through an upsampling network layer, adjusting the channel dimensions once more to obtain a concatenated convolutional feature map. The output feature map O2 and the concatenated convolutional feature map are then added element-wise, and the resulting image is processed by a decoding module to obtain the enhanced intermediate image. The decoding module is as follows: Figure 3 As shown, this process can be expressed mathematically as follows:

[0056]

[0057] In the formula, the upsample operation represents upsampling, the concat operation represents concatenation along the channel dimension, and the add operation represents element-wise addition of feature maps. This represents the feature map of the final output.

[0058] S4. A lightweight feature extraction backbone network, MobileViT, is used to extract features from the intermediate image. Each region of the intermediate image is encoded into three vectors: a query vector Q, a key vector K, and a value vector V. The similarity between Q of each feature block and K of the other feature blocks is calculated, and weighted summation is performed according to weights to extract effective information. The extracted feature information at different scales is fused at different levels (channel-dimensional concatenation) to obtain a multi-layer fused feature map. The multi-layer fused feature map is input into the detection head, and finally the network predicts the detection result (location and category information).

[0059] The overall network loss function is optimized based on the prediction results, and the weights of the trained model are adjusted. The overall network loss function is defined as the reconstruction loss L1 and the detection loss L2. d The reconstruction loss is mainly used to constrain the information loss of the input image after it passes through the feature enhancement network; the detection loss is mainly used to constrain the detection loss of the object detection network, including the object loss L. obj Category loss L cls and regression loss L loc This can be expressed as a formula:

[0060] L=λ1L1+λ d L d =λ1L1+λ obj L obj +λ cls L cls +λ loc L loc

[0061] Wherein, λ1, λ obj , λ cls and λ loc These represent the weighting coefficients for each part of the loss.

[0062] (3) The model obtained in step 2 is evaluated using a validation set and a test set. The validation set is used to evaluate the parameter indicators of the model, and the test set is used to verify the detection performance of the model.

[0063] The foregoing has described preferred embodiments of the present invention. However, it should be understood that the invention is not limited to the content disclosed herein. Any non-substantial improvements made using the inventive concept and technical solution, or any application of the inventive concept and technical solution to other situations, are within the protection scope of the present invention.

Claims

1. A method for constructing a lightweight target detection network adapted to low-light scenes, characterized in that, Includes the following steps: (1) Collect single-view video in low-light scenes, divide the image sequence, and after labeling the images, divide them into training set, validation set and test set; (2) The model is trained using a training set. By continuously optimizing the overall network loss function and adjusting the weights of the trained model, a model capable of object detection in low-light scenes is obtained after training. The overall network loss function is defined as the reconstruction loss L1 and the detection loss L2. d The reconstruction loss is mainly used to constrain the information loss of the input image after it passes through the feature enhancement network; the detection loss is mainly used to constrain the detection loss of the object detection network, including the object loss L. obj Category loss L cls and regression loss L loc This can be expressed as a formula: in, , , and These represent the weighting coefficients for each part of the loss; (3) The model obtained in step (2) is evaluated using the validation set and the test set, wherein the validation set is used to evaluate the parameter indicators of the model, and the test set is used to verify the detection performance of the model; The specific process of model training in step (2) is as follows: S1. The original image F is passed through two convolutional layers sequentially to eliminate redundant information, resulting in feature map F1 and feature map F2; the original image is represented by a data size of [data size not specified in the original text]. The feature map data sizes are respectively those of the original image. 1 / 4 and 1 / 8 are represented as and ; S2. Input the original image F, feature map F1, and feature map F2 into the feature enhancement network module in parallel to obtain three paths. The output feature maps O, O1, and O2 are represented as follows: , , The feature enhancement network consists of six CBR modules with residual connections. The first and sixth layers have residual connections, the second and fifth layers have residual connections, and the third and fourth layers have residual connections. Each CBR module performs Cov, BN, and ReLU operations in sequence. The kernel size of the Cov convolution operation is 3, and the stride is 1. The size of the feature map is not changed during the encoding process, but the number of channels is increased from 3 to 32, and the extracted information is integrated into the channel dimension. S3. The output feature map O2 is concatenated with the output feature map O1 through an upsampling network layer, and then upsampled again to adjust the channel dimensions, resulting in a concatenated convolutional feature map. The output feature map O and the concatenated convolutional feature map are added element-wise, and the image is decoded to obtain the enhanced intermediate image. The process of adding the output feature map O and the concatenated convolutional feature map element-wise can be expressed by the following mathematical formula: In the formula, the upsample operation represents upsampling, the concat operation represents concatenation along the channel dimension, and the add operation represents element-wise addition of feature maps. This represents the feature map of the final output; S4. A lightweight feature extraction backbone network, MobileViT, is used to extract features from the intermediate image. The extracted feature information at different scales is fused at different levels to obtain a multi-layer fused feature map. The multi-layer fused feature map is input into the detection head, and the network predicts the detection result. The specific method of feature extraction is as follows: each region of the intermediate image is encoded into three vectors, namely query vector Q, key vector K, and value vector V. The similarity between Q of each region and K of the other regions is calculated, and the weighted summation is performed according to the weights to extract effective information.

2. The method for constructing a lightweight target detection network adapted to low-light scenes as described in claim 1, characterized in that, The method for dividing the training set, validation set, and test set in step (1) is as follows: a. Acquire the original image sequence: The video was acquired in a low-light environment, and the acquired video sequence was divided into 10 image sequences; b. Create image labels: Use the labelImg tool to manually label the categories of objects in the image and export the image label file; c. Divide the collected image sequences into training, validation, and test sets according to a 7:1:2 ratio.

3. The method for constructing a lightweight target detection network adapted to low-light scenes as described in claim 1, characterized in that, S1 describes the process of eliminating redundant information through two convolutional layers. The kernel size of each convolutional layer is 3, which can be expressed by the formula: In the formula, Conv represents the convolution operation, and S represents the stride of the convolution.

4. A target detection method, comprising the steps of inputting the image to be detected into the target detection network constructed according to any one of claims 1-3 for target detection.

5. An electronic product, comprising a processor and a memory communicatively connected to the processor; the memory being used to store instructions; the processor being used to invoke the instructions stored in the memory to execute the target detection method as described in claim 4.