Low-illumination image target detection method and device, electronic equipment and medium
Through the low-illumination end-to-end object detection model, image enhancement and object detection are integrated into a unified network structure, solving the problems of low-illumination image object detection efficiency and error accumulation in the prior art, and achieving more efficient and robust object detection effects.
Patent Information
- Application Number
- CN202510118858.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing low-illumination image object detection methods have problems such as inefficiency, accumulation of errors and insufficient information sharing in harsh environments, resulting in low target detection efficiency and robustness.
The low-illumination end-to-end object detection model is adopted to integrate image enhancement and object detection into a unified network structure, and process it through prompt encoder, trained image enhancement network and object detector to reduce information loss and error transmission.
It improves the efficiency and robustness of the object detection, reduces information loss and error transmission, and enables more accurate object detection in low-illumination environments.
Smart Images

Figure CN120047737A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and more particularly to a low-light image target detection method, device, electronic device and medium. Background Art
[0002] Low-light image target detection is an important research field in computer vision, mainly facing challenges in target recognition in low-light environments, such as insufficient light, noise interference and dynamic range limitations. Although traditional target detection methods have achieved gratifying results on general datasets, it is still challenging to accurately locate objects in low-light images captured in harsh environments. In the target detection task in low-light environments, it is usually necessary to cope with challenges brought by complex lighting changes and image quality degradation, and at the same time solve problems such as difficult target recognition caused by insufficient light. At present, research in this field mainly focuses on how to effectively enhance image quality, extract target features, and achieve high-precision target localization in complex low-light scenes. Although low-light target detection technology has made certain progress, existing methods usually use two-stage algorithms, that is, first use an enhancement algorithm to enhance image quality, and then use a detector to perform target localization, but there are still problems such as low efficiency, error accumulation and insufficient information sharing in actual applications. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a low-light image target detection method, device, electronic device and medium, so as to reduce information loss and error transmission, and improve the efficiency and robustness of target detection.
[0004] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0005] In a first aspect, the present invention provides a low-light image target detection method, including: obtaining a low-light image to be detected in a target scene; inputting the low-light image to be detected into a pre-trained low-light end-to-end target detection model to obtain an enhanced image and a target detection result of the low-light image to be detected; wherein, the low-light end-to-end target detection model includes: a prompt encoder, a trained image enhancement network and a target detector; the prompt encoder outputs prompt coding information of the low-light image to be detected to the trained image enhancement network, the trained image enhancement network enhances the low-light image to be detected based on the prompt coding information, outputs the enhanced image to the target detector, and the target detector performs target detection on the enhanced image and outputs the target detection result.
[0006] Optionally, input the low-light image to be detected into a pre-trained end-to-end low-light object detection model to obtain the enhanced image and object detection result of the low-light image to be detected, including: downsampling the low-light image to be detected to obtain a sampled image; segmenting the sampled image using adaptive threshold segmentation to obtain a binary light intensity mask image, and inputting the binary light intensity mask image into a prompt encoder to obtain prompt encoding information; inputting the sampled image and the prompt encoding information into a trained image enhancement network to obtain an adjustment curve corresponding to the low-light image to be detected; wherein, the adjustment curve includes: an adjustment curve for the R channel, an adjustment curve for the G channel, and an adjustment curve for the B channel; based on the adjustment curve corresponding to the low-light image to be detected and the low-light image to be detected, obtain the enhanced image of the low-light image to be detected; input the enhanced image of the low-light image to be detected into an object detector to obtain the object detection result.
[0007] Optionally, inputting the sampled image and the prompt encoding information into a trained image enhancement network to obtain an adjustment curve corresponding to the low-light image to be detected, including: inputting the sampled image and the prompt encoding information into a trained image enhancement network to obtain an adjustment factor matrix corresponding to the sampled image; upsampling the adjustment factor matrix to obtain a restored adjustment factor matrix; wherein, the adjustment factor matrix includes: an adjustment factor matrix for the R channel, an adjustment factor matrix for the G channel, and an adjustment factor matrix for the B channel; based on the restored adjustment factor matrix, determine the adjustment curve corresponding to the low-light image to be detected.
[0008] Optionally, based on the adjustment curve corresponding to the low-light image to be detected and the low-light image to be detected, obtain the enhanced image of the low-light image to be detected, including: multiplying the adjustment curve corresponding to each channel by the low-light image to be detected respectively to obtain a mapping matrix corresponding to each channel; adding the mapping matrix and the low-light image to be detected to obtain the enhanced image of the low-light image to be detected.
[0009] Optionally, before obtaining the low-light image to be detected in the target scene, it further includes: obtaining a training set; wherein, the training set includes: first low-light images in different scenes and a second low-light image in the target scene; pre-training the image enhancement network based on the first low-light images to obtain a trained image enhancement network; loading the trained image enhancement network into the end-to-end low-light object detection model and freezing the model parameters of the trained image enhancement network; training the end-to-end low-light object detection model based on the second low-light image to obtain a trained end-to-end low-light object detection model.
[0010] Optionally, pre-train the image enhancement network based on the first low-light image to obtain a trained image enhancement network, including: inputting the first low-light image into the image enhancement network to obtain an adjustment factor matrix corresponding to each first low-light image; performing adaptive learning on the adjustment factor matrix based on a preset loss function until a preset number of iterations is reached to obtain a trained image enhancement network; wherein, the loss function includes: a spatial consistency loss function, an exposure suppression loss function, a color constancy loss function, and an illumination smoothness loss function.
[0011] Optionally, train the low-light end-to-end object detection model based on the second low-light image to obtain a trained low-light end-to-end object detection model, including: determining the prompt coding information of each second low-light image based on the prompt encoder; determining an adjustment curve corresponding to each second low-light image based on the trained image enhancement network and the prompt coding information; obtaining an enhanced image of each second low-light image based on the adjustment curve corresponding to each second low-light image and the second low-light image; determining the object detection result of each second low-light image based on the enhanced image of each second low-light image and the object detector; calculating the joint loss of the trained image enhancement network and the object detector based on the enhanced image of each second low-light image and the object detection result, and optimizing the model parameters of the low-light end-to-end object detection model based on the joint loss to obtain a trained low-light end-to-end object detection model.
[0012] In a second aspect, the present invention provides a low-light image object detection device, including: an image acquisition module for acquiring a to-be-detected low-light image in a target scene; an object detection module for inputting the to-be-detected low-light image into a pre-trained low-light end-to-end object detection model to obtain an enhanced image and an object detection result of the to-be-detected low-light image; wherein, the low-light end-to-end object detection model includes: a prompt encoder, a trained image enhancement network, and an object detector; the prompt encoder outputs the prompt coding information of the to-be-detected low-light image to the trained image enhancement network, the trained image enhancement network enhances the to-be-detected low-light image based on the prompt coding information, outputs the enhanced image to the object detector, and the object detector performs object detection on the enhanced image and outputs the object detection result.
[0013] In a third aspect, the present invention provides an electronic device, including a processor and a memory, the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the steps of any method provided in the first aspect above.
[0014] Fourthly, the present invention provides a computer-readable storage medium with a computer program stored thereon. When the computer program is run by a processor, it executes the steps of the method provided in any one of the above first aspects.
[0015] The present invention brings the following beneficial effects:
[0016] The above low-light image target detection method, device, electronic device and medium provided by the present invention first obtain the low-light image to be detected in the target scene; then input the low-light image to be detected into a pre-trained low-light end-to-end target detection model to obtain the enhanced image and target detection result of the low-light image to be detected; wherein, the low-light end-to-end target detection model includes: a prompt encoder, a trained image enhancement network and a target detector. The above method uses a low-light end-to-end target detection model for target detection, integrating the entire process from the input image to the final detection result into a unified network structure, eliminating the conversion and intermediate processing between multiple independent steps in the traditional method, making the information transfer inside the model more direct, reducing information loss and error propagation, and at the same time improving the target detection efficiency and robustness.
[0017] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are realized and obtained by the structures specifically pointed out in the specification, claims and drawings.
[0018] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a flowchart of a low-light image target detection method provided by an embodiment of the present invention;
[0021] Figure 2 It is a network structure diagram of a low-light end-to-end target detection model provided by an embodiment of the present invention;
[0022] Figure 3 It is a comparison diagram of an original multi-stage method and an end-to-end model provided by an embodiment of the present invention;
[0023] Figure 4 It is a schematic structural diagram of a low-illumination image target detection device provided by an embodiment of the present invention;
[0024] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] To facilitate better understanding of the present application by those skilled in the art, the technical terms involved in the present application will be briefly introduced below.
[0027] 1. Prompt Learning: It is a strategy that uses external prompts or input information to guide the model to learn. It usually involves introducing specific information or formatted prompts (such as text, images, or other context information) into the model to help the model better understand the task requirements or data characteristics. This method can improve the performance of the model in multiple fields such as large language models and computer vision.
[0028] In the field of computer vision, the prompt can be descriptive text, such as "Identify the dog in the image" or "Mark all the people in the picture". This form of prompt can help the model clarify the task objective; it can be a label prompt, such as "Category: Cat" or "Object: Tree", directly indicating the object to be detected or classified by the model; it can be a visual prompt, such as a certain preprocessing result of the image, a mask image, showing the area that needs attention.
[0029] 2. End-to-End Learning: It is a deep learning method, which means that the entire process from input to output is automatically completed by a single model or system without complex process design and processing steps in the middle. It has a unified model structure and directly learns from the input data (such as image pixel values) to the output (such as detection results). All intermediate steps, including feature extraction and representation learning, are learned by the network itself during the training process. It can directly learn from a large amount of data and often be able to capture the complex non-linear relationships in the task, achieving better performance than traditional methods. Compared with non-end-to-end models with multi-stage processing, it has the following advantages:
[0030] 1) Automated feature learning: It can automatically extract and optimize features from data, reducing human intervention and dependence on domain knowledge, and is more suitable for complex, high-dimensional, and unstructured data (such as images, speech, natural language processing, etc.). In contrast, each stage of the non-end-to-end model is disjointed and does not achieve the goal of global automatic learning.
[0031] 2) Global optimization: From input to output, through the joint training of the entire model, all parameters of the model are optimized simultaneously, ensuring that all parts of the model can work better together. In contrast, non-end-to-end models often face problems of matching and tuning between modules.
[0032] 3) Advantages of transfer learning: It is easier to transfer between multi-tasks or different domains, and can quickly adapt to new tasks through means such as fine-tuning to achieve task adaptation in different domains. In contrast, non-end-to-end models rely on domain knowledge and can only be fine-tuned within a single domain to achieve optimality.
[0033] 3. Low light image enhancement: In the field of image processing, it refers to processing images taken under low light conditions through specific algorithms and technologies to improve the visibility and quality of the images. Since images taken in low light environments usually have problems such as noise, insufficient contrast, and missing details, enhancement techniques need to be adopted to improve the visual effects of these images. The methods of low light image enhancement mainly include histogram equalization, color enhancement, and deep learning methods. Finally, low light images with missing details and color distortion are enhanced into high-quality visual effect images with high color saturation and prominent details.
[0034] After introducing the technical terms involved in this application, the technical solutions provided by the embodiments of this application will be described below.
[0035] Currently, in the field of low light object detection, existing methods usually use two-stage algorithms, that is, first use an enhancement algorithm to enhance the image quality, and then use a detector for object localization. However, in practical applications, the following problems are still faced:
[0036] 1) Inconsistency of targets and data: The inconsistency of targets means that image enhancement focuses on human vision, while object detection focuses on machine vision; the inconsistency of data means that existing methods cannot share the same data samples in the training stages of image enhancement and object detection tasks. There is a lack of coordination between different tasks and no sharing and reference between each other.
[0037] 2) Separation of steps and error propagation: Image quality enhancement and object detection are carried out separately, resulting in the need for independent computing resources for each process. This separate processing not only increases the processing time but also introduces more error accumulation, making the overall inference efficiency lower than that of end-to-end models.
[0038] 3) Lack of information sharing and collaborative optimization: Image enhancement and object detection are independent modules and cannot share information with each other. During the enhancement process, features useful for detection may be ignored, and the detector cannot fully utilize the intermediate information of image enhancement. The overall system cannot perform global optimization, which limits the performance ceiling.
[0039] Based on this, a low-light image object detection method, device, electronic device, and medium provided by the embodiments of the present invention can reduce information loss and error transmission, and improve the efficiency and robustness of object detection.
[0040] For ease of understanding of this embodiment, first, a low-light image object detection method disclosed in the embodiments of the present invention will be introduced in detail. This method can be executed by an electronic device, such as a smartphone, a computer, a tablet, etc. Refer to Figure 1 The flowchart of a low-light image object detection method shown in the figure schematically shows that the method mainly includes the following steps S101 to step S102:
[0041] Step S101: Obtain a low-light image to be detected in a target scene.
[0042] In one implementation, the target scene can be an urban street, an intersection traffic signal, an indoor environment, a natural landscape, etc. The low-light image to be detected can be an image of the target scene captured by a device such as a mobile phone or a camera in a low-light environment. The low-light image to be detected includes detection targets, such as: people, vehicles, buildings, and so on.
[0043] Step S102: Input the low-light image to be detected into a pre-trained low-light end-to-end object detection model to obtain an enhanced image of the low-light image to be detected and an object detection result.
[0044] Among them, the low-light end-to-end object detection model includes: a prompt encoder, a trained image enhancement network, and an object detector; the prompt encoder outputs prompt coding information of the low-light image to be detected to the trained image enhancement network, and the trained image enhancement network enhances the low-light image to be detected based on the prompt coding information and outputs the enhanced image to the object detector. The object detector performs object detection on the enhanced image and outputs an object detection result.
[0045] In one embodiment, a pre-trained low-light end-to-end object detection model can be used to perform object detection on the acquired low-light image to be detected, and an enhanced image and an object detection result of the low-light image to be detected are obtained. Among them, the low-light end-to-end object detection model includes: a prompt encoder, a trained image enhancement network, and an object detector. In the embodiment of the present invention, the context light intensity information of the low-light image to be detected can be used as prompt information, and the prompt encoder is used to perform feature representation on the prompt information and integrate it into the trained image enhancement network to obtain an enhanced image. Finally, the enhanced image is input into the object detector to obtain the object detection result.
[0046] The above-mentioned low-light image object detection method provided by the present invention uses a low-light end-to-end object detection model for object detection, integrates the entire process from the input image to the final detection result into a unified network structure, eliminates the conversion and intermediate processing between multiple independent steps in the traditional method, makes the transmission of information inside the model more direct, reduces information loss and error transmission, and improves the object detection efficiency and robustness at the same time.
[0047] In one embodiment, referring to Figure 2 As shown in the network structure diagram of a low-light end-to-end object detection model, for the aforementioned step S102, that is, when the low-light image to be detected is input into the pre-trained low-light end-to-end object detection model to obtain the enhanced image and the object detection result of the low-light image to be detected, the following methods can be used, including but not limited to:
[0048] First, downsample the low-light image to be detected to obtain a sampled image.
[0049] In a specific implementation, the low-light image to be detected can be downsampled by 16 times to obtain a sampled image.
[0050] Then, use adaptive threshold segmentation to segment the sampled image to obtain a binary light intensity mask image, and input the binary light intensity mask image into the prompt encoder to obtain prompt coding information;
[0051] In a specific implementation, adaptive threshold segmentation can be used to segment the sampled image to obtain a binary light intensity mask image, and then the binary light intensity mask image is input into the prompt encoder. In the embodiment of the present invention, the network structure of the prompt encoder consists of multiple linear layers. Specifically, a linear layer is used to perform linear mapping on the prompt information to increase the feature dimension, and then a non-linear activation function is superimposed to learn the non-linear process to improve the diversity of expression ability. Multiple linear layers are parallel at the output stage to output multiple prompt coding information.
[0052] Next, input the sampled image and the prompt encoding information into the trained image enhancement network to obtain the adjustment curve corresponding to the low-light image to be detected; wherein, the adjustment curve includes: the adjustment curve of the R channel, the adjustment curve of the G channel, and the adjustment curve of the B channel.
[0053] In specific implementation, first input the sampled image and the prompt encoding information into the trained image enhancement network to obtain the adjustment factor matrix corresponding to the sampled image; then perform upsampling on the adjustment factor matrix to obtain the restored adjustment factor matrix; finally, based on the restored adjustment factor matrix, determine the adjustment curve corresponding to the low-light image to be detected.
[0054] In the embodiment of the present invention, the image enhancement network is pre-trained to obtain a trained image enhancement network, and then the trained image enhancement network is loaded into the low-light end-to-end object detection model, and the model parameters of the trained image enhancement network are frozen. The image enhancement network is composed of several simple convolutional layers, is a UNet-like network architecture, the core components are depthwise separable convolution layers (Depthwise Separable Convolution, DSConv) and non-linear activation functions, and skip connections are used for feature map reuse to enhance the utilization efficiency of feature information in different stages. In the output layer part of the network, there are 3 output channels, corresponding to the RGB 3-channel image of the image respectively, and an adjustment curve will be generated for each channel.
[0055] In practical applications, input the sampled image and the prompt encoding information into the trained image enhancement network to obtain the adjustment factor matrix corresponding to the sampled image (including: the adjustment factor matrix of the R channel, the adjustment factor matrix of the G channel, and the adjustment factor matrix of the B channel), wherein the values of the adjustment factor matrix are in the range of 0 to 1, corresponding to the pixel value range of 0 to 255 of the image. In the output part of the image enhancement network, the adjustment factor matrix is upsampled and restored to the scale of the low-light image to be detected, thereby reducing the computational amount of the image enhancement network, and further reducing the overall burden of the low-light end-to-end object detection model. The prompt encoding information is incorporated in the feature decoding stage. After the feature fusion of each skip connection, the prompt encoding information is added element-wise to the fused feature to ensure that the prompt information can effectively guide the image enhancement network to optimize and enhance the important regions.
[0056] After that, based on the adjustment curve corresponding to the low-light image to be detected and the low-light image to be detected, obtain the enhanced image of the low-light image to be detected.
[0057] In specific implementation, the mapping matrices of different channels can be calculated from the mapping relationship between the input low-light image to be detected and the adjustment curve, and the mapping matrix and the original low-light image are added together to obtain the final enhanced image. Specifically, first, the adjustment curve corresponding to each channel is multiplied by the low-light image to be detected respectively to obtain the mapping matrix corresponding to each channel; then, the mapping matrix and the low-light image to be detected are added together to obtain the enhanced image of the low-light image to be detected.
[0058] Finally, the enhanced image of the low-light image to be detected is input into the target detector to obtain the target detection result.
[0059] In specific implementation, the input of the target detector is the enhanced high-quality image (i.e., the enhanced image), and the target detector can be any type of detector, such as: YOLO, DETR, FasterR-CNN, CenterNet, etc. This flexibility enables this method to optimize performance for different scenarios. For example, YOLO is used in real-time monitoring environments to achieve fast response; DETR is selected in occasions where high-precision positioning is required to improve accuracy. Specifically, the target detector outputs the visualized result after target detection through coordinate decoding and class prediction.
[0060] The above low-light image target detection method provided by the embodiments of the present invention uses adaptive threshold segmentation to use the context light intensity information of the low-light image as the prompt information, performs feature representation through the prompt encoder, integrates it into the unsupervised image enhancement network, learns the adaptive adjustment curve, obtains the mapping matrix through the adjustment curve and the original image, obtains the enhanced image through the mapping matrix, and sends the enhanced image into the target detector for target detection, improving the detection effect under extreme lighting conditions. In addition, the prompt information is input into the image enhancement network through the lightweight prompt encoder, and its number of parameters is much smaller than that of traditional target detection models, will not increase too many model parameters, and can be flexibly configured and seamlessly integrated into the existing target detection framework, with good model scalability and real-time performance.
[0061] The embodiments of the present invention also provide a training process for a low-light end-to-end target detection model, which mainly includes the following steps 1 to 4:
[0062] Step 1: Obtain the training set; wherein, the training set includes: the first low-light images in different scenarios and the second low-light images in the target scenario.
[0063] In specific implementation, the first low-light images under large-scale low-light conditions can be collected for the image enhancement network. The first low-light images can be images taken in low-light environments or images in public datasets. It is necessary to ensure that the collected images cover various scenarios, such as urban streets, indoor environments, natural landscapes, etc., to enhance the generalization ability of the model. In addition, the second low-light images of a specific application scenario (i.e., the target scenario, which is also the scenario where the trained low-light end-to-end object detection model is applied) are collected as the low-light object detection dataset. The collected second low-light images are screened and filtered, and the object detection frames are labeled, and the corresponding format conversion of the annotation files is performed.
[0064] Step 2: Pre-train the image enhancement network based on the first low-light images to obtain a trained image enhancement network.
[0065] Before performing the low-light object detection training task, it is necessary to pre-train the image enhancement network. In the embodiments of the present invention, in the image enhancement task, the image enhancement task is converted into a task of using a CNN network to estimate specific curves of an image, that is, a lightweight CNN network is trained to estimate pixel values and high-order curves to adjust the dynamic range of a given image. At the same time, an unsupervised loss calculation method is adopted to implicitly enhance the image quality and drive the network to learn. Since it is an unsupervised network, no paired images are required during the training process, avoiding the trouble caused by data loss.
[0066] In specific implementation, when pre-training the image enhancement network based on the first low-light images, first input the first low-light images into the image enhancement network to obtain an adjustment factor matrix corresponding to each first low-light image; then perform adaptive learning on the adjustment factor matrix based on a preset loss function until a preset number of iterations is reached to obtain a trained image enhancement network; where the loss function includes: a spatial consistency loss function, an exposure suppression loss function, a color constancy loss function, and an illumination smoothness loss function.
[0067] See Figure 2As shown, the image enhancement network consists of several simple convolutional layers and is a UNet-like network architecture. The core components are depthwise separable convolutional layers and non-linear activation functions, and skip connections are used to reuse feature maps to enhance the utilization efficiency of feature information at different stages. In the output layer part of the network, there are 3 output channels, corresponding to the RGB 3-channel image of the image respectively. For each channel, an adjustment factor matrix M will be generated, and each value in the matrix corresponds to a pixel point in the image. The Tanh function is used to limit the values of the adjustment factor matrix M within the range of -1 to 1. The enhancement network will adaptively learn the adjustment factor matrix M according to the loss function to enhance the low-light image, and obtain the adjustment curve corresponding to the adjustment factor matrix M through continuous iteration, so as to obtain the trained enhancement network. The formula of the adjustment curve is as follows:
[0068] AC(I(α); M(α)) = Iter N (I(α) + M(α)(I(α) 2 -I(α)))
[0069] Among them, I represents the input first low-light image, α represents the coordinates of the image, AC(I(α); M(α)) represents the output of the adjustment curve, and Iter N represents iterating N times.
[0070] Due to the difference of the adjustment factor matrix M, the dynamic adjustment range of the adjustment curve is also different, Figure 2 shows the trend change of the adjustment curve for different values of the adjustment factor matrix M, and the adjusted values are within the range of 0 to 1, corresponding to the pixel value range of 0 to 255 of the image.
[0071] In the embodiment of the present invention, for the design of the unsupervised loss function, 4 types of losses are used to train the image enhancement network: spatial consistency loss L spa , exposure suppression loss L exp , color constancy loss L col , illumination smoothness loss L ill .
[0072]
[0073] Among them, in L spa , K is the number of local regions, Ω(i) are 4 neighboring regions centered on this region, Y and I respectively represent the average intensities of the enhanced image and the input image. By retaining the differences between adjacent regions of the input image and the enhanced image, the spatial consistency of the enhanced image is promoted; in L exp , M represents the number of non-overlapping local regions of size 16×16, and the average intensity value of the local region in the enhanced image is represented as Y, which can suppress the cases of under-exposure and over-exposure; in L col , Jp Represents the average intensity of the p-channel. The channel pair is denoted as (p, q), following the gray world color constancy assumption, i.e., the color in each channel is gray on average across the entire image, which is used to correct potential color deviations in the enhanced image; L ill where N is the number of iterations, and represent the horizontal and vertical directions, and M represents the adjustment factor matrix, which can maintain the monotonic relationship between adjacent pixels.
[0074] During the training process, online data augmentation techniques can be used to improve the generalization ability and robustness of the model. To ensure the diversity of the training data, in the embodiments of the present invention, the probability factors and expansion coefficients of different data augmentation methods are restricted within a certain range. Table 1 shows the probability factors and expansion coefficients corresponding to different augmentation techniques. After the image enhancement network is trained, it can be used in specific application scenarios without repeated training.
[0075] Table 1 Probability factors and expansion coefficients of different data augmentation methods
[0076]
[0077] Step 3: Load the trained image enhancement network into the low-light end-to-end object detection model, and freeze the model parameters of the trained image enhancement network.
[0078] In specific implementation, load the weights of the above-trained image enhancement network model into the low-light end-to-end object detection model, and freeze the training parameters of the image enhancement network.
[0079] Step 4: Train the low-light end-to-end object detection model based on the second low-light image to obtain a trained low-light end-to-end object detection model.
[0080] In specific implementation, in a specific application scenario (such as: when the trained model is used for detecting people or vehicles on the road, low-light images of the street are used for model training), train the low-light end-to-end object detection model using the second low-light images collected in the specific scenario. Prompt the encoder to input the light intensity mask information (i.e., the prompt coding information) as the prompt information into the image enhancement network to guide the image enhancement network to generate high-quality images. The high-quality images (i.e., enhanced images) are input into the detector in the specific scenario, and the joint loss of the dual task is used to jointly guide the collaborative learning of the low-light end-to-end object detection model.
[0081] Specifically, when training the low-light end-to-end object detection model based on the second low-light image, the following methods can be adopted, including but not limited to:
[0082] First, determine the prompt encoding information of each second low-illumination image based on the prompt encoder.
[0083] In specific implementation, the input of the prompt encoder is a binary mask image obtained by adaptive threshold segmentation, which emphasizes the dark regions. The image enhancement network can better capture the foreground and background of the entire image. The adaptive threshold segmentation can be dynamically adjusted according to the local brightness characteristics of the image, effectively adapting to images under different lighting conditions, enabling useful prompt information to be extracted in various low-illumination situations. Specifically, first convert each second low-illumination image into a binary mask image through adaptive threshold segmentation, and then process the binary mask image through the prompt encoder to output the prompt encoding information of each second low-illumination image.
[0084] Then, determine the adjustment curve corresponding to each second low-illumination image based on the trained image enhancement network and the prompt encoding information.
[0085] Next, based on the adjustment curve corresponding to each second low-illumination image and the second low-illumination image, obtain the enhanced image of each second low-illumination image.
[0086] After that, determine the object detection result of each second low-illumination image based on the enhanced image of each second low-illumination image and the object detector.
[0087] In specific implementation, input the prompt encoding information as prompt information into the trained image enhancement network to obtain the adjustment curve corresponding to each second low-illumination image; then calculate the mapping matrix of different channels according to the mapping relationship between the second low-illumination image and the adjustment curve, add the mapping matrix and the second low-illumination image to obtain the enhanced image of each second low-illumination image; furthermore, input the enhanced image of each second low-illumination image into the object detector to obtain the object detection result of each second low-illumination image.
[0088] Finally, based on the enhanced image of each second low-illumination image and the object detection result, calculate the joint loss of the trained image enhancement network and the object detector, and optimize the model parameters of the low-illumination end-to-end object detection model based on the joint loss to obtain the trained low-illumination end-to-end object detection model.
[0089] In specific implementation, the model is optimized through joint loss calculation. The joint loss calculation mainly comes from the calculation loss of the image enhancement network and the calculation loss of the object detector. In the training stage, the sum of the two losses jointly guides the learning of the two tasks, forming an end-to-end unified framework, which not only improves the image quality, but also helps the object detector to more accurately identify and locate the object, as well as improve the overall collaborative learning effect and dynamic optimization ability to cope with various complex scenarios.
[0090] In the embodiments of the present invention, image enhancement and object detection are integrated into a unified end-to-end learning model. The context light intensity information is incorporated into the unsupervised image enhancement network through prompt learning to guide the generation of high-quality visual images. The enhanced image is then fed into the object detector, and the detection loss of the object detector and the enhancement loss of the image enhancement network form a joint loss, jointly guiding the learning of the image enhancement network and the object detector to form a low-light end-to-end object detection model. The low-light end-to-end object detection model can effectively integrate image enhancement and object detection, and through real-time feedback and learning mechanisms, it can achieve dynamic adjustment to ensure that the dual tasks of image enhancement and object detection work in coordination under a common goal, significantly improving the object detection efficiency in low-light environments. See Figure 3 As shown, compared with the prior art, the present invention has the following beneficial effects:
[0091] 1) Introduce prompt learning in low-light object detection, use the light intensity information of low-light images as key prompts, and combine adaptive threshold segmentation technology to extract lighting features, providing additional semantic information for the image enhancement network and a new idea for object detection in low-light environments.
[0092] 2) By integrating the entire process from the input image to the final detection result into an end-to-end training method that integrates the input image to the detection result into a single network structure, eliminating the conversion between multiple independent steps and improving the efficiency of information transfer. Through the joint training mechanism, image enhancement and object detection share data and computing resources, enhancing the training efficiency and detection robustness.
[0093] 3) Adopt an unsupervised image enhancement method, which is not restricted by training sample pairs. This method can be seamlessly integrated with various object detection models (such as YOLO, DETR, Faster R-CNN, CenterNet), and has good scalability. During the image enhancement process, it has lightweight characteristics, ensuring high inference speed when processing low-light images and not imposing a burden on the original object detection model.
[0094] For the low-light image object detection method provided in the foregoing embodiments, the embodiments of the present invention also provide a low-light image object detection device. See Figure 4 As shown in the structural schematic diagram of a low-light image object detection device, it shows that the device mainly includes the following parts:
[0095] An image acquisition module 401, configured to acquire a low-light image to be detected in a target scene;
[0096] The target detection module 402 is configured to input the low-light image to be detected into a pre-trained end-to-end low-light target detection model to obtain an enhanced image and a target detection result of the low-light image to be detected; wherein, the end-to-end low-light target detection model includes: a prompt encoder, a trained image enhancement network, and a target detector; the prompt encoder outputs the prompt coding information of the low-light image to be detected to the trained image enhancement network, the trained image enhancement network enhances the low-light image to be detected based on the prompt coding information, and outputs the enhanced image to the target detector, and the target detector performs target detection on the enhanced image and outputs the target detection result.
[0097] The above-mentioned low-light image target detection device provided by the present invention uses an end-to-end low-light target detection model for target detection, integrates the entire process from the input image to the final detection result into a unified network structure, eliminates the conversion and intermediate processing between multiple independent steps in the traditional method, makes the transmission of information in the model more direct, reduces information loss and error transmission, and improves the target detection efficiency and robustness at the same time.
[0098] In one implementation, the above-mentioned target detection module 402 is specifically configured to: downsample the low-light image to be detected to obtain a sampled image; segment the sampled image using adaptive threshold segmentation to obtain a binary light intensity mask image, and input the binary light intensity mask image into the prompt encoder to obtain prompt coding information; input the sampled image and the prompt coding information into the trained image enhancement network to obtain an adjustment curve corresponding to the low-light image to be detected; wherein, the adjustment curve includes: an adjustment curve for the R channel, an adjustment curve for the G channel, and an adjustment curve for the B channel; based on the adjustment curve corresponding to the low-light image to be detected and the low-light image to be detected, obtain an enhanced image of the low-light image to be detected; input the enhanced image of the low-light image to be detected into the target detector to obtain a target detection result.
[0099] In one implementation, the above-mentioned target detection module 402 is specifically configured to: input the sampled image and the prompt coding information into the trained image enhancement network to obtain an adjustment factor matrix corresponding to the sampled image; upsample the adjustment factor matrix to obtain a restored adjustment factor matrix; wherein, the adjustment factor matrix includes: an adjustment factor matrix for the R channel, an adjustment factor matrix for the G channel, and an adjustment factor matrix for the B channel; based on the restored adjustment factor matrix, determine an adjustment curve corresponding to the low-light image to be detected.
[0100] In one implementation, the above-mentioned target detection module 402 is specifically configured to: multiply the adjustment curve corresponding to each channel by the low-light image to be detected respectively to obtain a mapping matrix corresponding to each channel; add the mapping matrix and the low-light image to be detected to obtain an enhanced image of the low-light image to be detected.
[0101] In one embodiment, the above device further includes a model training module, configured to: obtain a training set; wherein, the training set includes: first low-light images in different scenarios and second low-light images in a target scenario; pre-train an image enhancement network based on the first low-light images to obtain a trained image enhancement network; load the trained image enhancement network into a low-light end-to-end object detection model, and freeze the model parameters of the trained image enhancement network; train the low-light end-to-end object detection model based on the second low-light images to obtain a trained low-light end-to-end object detection model.
[0102] In one embodiment, the above model training module is specifically configured to: input the first low-light images into the image enhancement network to obtain an adjustment factor matrix corresponding to each first low-light image; perform adaptive learning on the adjustment factor matrix based on a preset loss function until a preset number of iterations is reached to obtain a trained image enhancement network; wherein, the loss function includes: a spatial consistency loss function, an exposure suppression loss function, a color constancy loss function, and an illumination smoothness loss function.
[0103] In one embodiment, the above model training module is specifically configured to: determine prompt coding information for each second low-light image based on a prompt encoder; determine an adjustment curve corresponding to each second low-light image based on the trained image enhancement network and the prompt coding information; obtain an enhanced image of each second low-light image based on the adjustment curve corresponding to each second low-light image and the second low-light image; determine an object detection result of each second low-light image based on the enhanced image of each second low-light image and an object detector; calculate a joint loss of the trained image enhancement network and the object detector based on the enhanced image of each second low-light image and the object detection result, and optimize the model parameters of the low-light end-to-end object detection model based on the joint loss to obtain a trained low-light end-to-end object detection model.
[0104] It should be noted that the device provided in the embodiments of the present invention has the same implementation principle and the same technical effects as those in the foregoing method embodiments. For a brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding contents in the foregoing method embodiments.
[0105] The embodiments of the present invention further provide an electronic device. Specifically, the electronic device includes a processor and a storage device; a computer program is stored on the storage device, and when the computer program is run by the processor, it executes the method according to any one of the above embodiments.
[0106] Figure 5A schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 100 includes: a processor 50, a memory 51, a bus 52, and a communication interface 53. The processor 50, the communication interface 53, and the memory 51 are connected through the bus 52. The processor 50 is configured to execute an executable module stored in the memory 51, such as a computer program.
[0107] Among them, the memory 51 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 53 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element. The Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0108] The bus 52 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 5 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0109] Among them, the memory 51 is used to store a program. After receiving an execution instruction, the processor 50 executes the program. The method executed by the device defined by the flow process disclosed in any embodiment of the foregoing embodiments of the present invention can be applied to the processor 50 or implemented by the processor 50.
[0110] The processor 50 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 50 or the instructions in the form of software. The above-mentioned processor 50 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 51, and the processor 50 reads the information in the memory 51 and combines its hardware to complete the steps of the above method.
[0111] The computer program product of the readable storage medium provided by the embodiments of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For the specific implementation, reference can be made to the foregoing method embodiments, which will not be elaborated here.
[0112] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0113] Finally, it should be noted that the above-mentioned embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A low-light image target detection method, characterized in that: include: Acquire the low-light image to be detected in the target scene; The low-illuminance image to be detected is input into a pre-trained low-illuminance end-to-end target detection model to obtain an enhanced image and a target detection result of the low-illuminance image to be detected; wherein the low-illuminance end-to-end target detection model includes: a prompt encoder, a trained image enhancement network and a target detector; the prompt encoder outputs the prompt coding information of the low-illuminance image to be detected to the trained image enhancement network, the trained image enhancement network enhances the low-illuminance image to be detected based on the prompt coding information, and outputs the enhanced image to the target detector, the target detector performs target detection on the enhanced image, and outputs the target detection result.
2. The method according to claim 1, characterized in that Inputting the low-light image to be detected into a pre-trained low-light end-to-end target detection model to obtain an enhanced image and target detection result of the low-light image to be detected, including: Downsampling the low-illumination image to be detected to obtain a sampled image; Adopting adaptive threshold segmentation to segment the sampled image to obtain a binary light intensity mask image, and inputting the binary light intensity mask image into the prompt encoder to obtain prompt encoding information; Inputting the sampled image and the prompt coding information into the trained image enhancement network to obtain an adjustment curve corresponding to the low-light image to be detected; wherein the adjustment curve includes: an adjustment curve of the R channel, an adjustment curve of the G channel, and an adjustment curve of the B channel; Obtaining an enhanced image of the low-illuminance image to be detected based on the adjustment curve corresponding to the low-illuminance image to be detected and the low-illuminance image to be detected; The enhanced image of the low-illuminance image to be detected is input into the target detector to obtain a target detection result.
3. The method according to claim 2, characterized in that Inputting the sampled image and the prompt coding information into the trained image enhancement network to obtain an adjustment curve corresponding to the low-light image to be detected, including: Inputting the sampled image and the hint coding information into the trained image enhancement network to obtain an adjustment factor matrix corresponding to the sampled image; Upsampling the adjustment factor matrix to obtain a restored adjustment factor matrix; wherein the adjustment factor matrix includes: an adjustment factor matrix of the R channel, an adjustment factor matrix of the G channel, and an adjustment factor matrix of the B channel; Based on the restored adjustment factor matrix, an adjustment curve corresponding to the low-illuminance image to be detected is determined.
4. The method according to claim 2, characterized in that: Based on the adjustment curve corresponding to the low-illuminance image to be detected and the low-illuminance image to be detected, obtaining an enhanced image of the low-illuminance image to be detected, including: Multiplying the adjustment curve corresponding to each channel by the low-light image to be detected to obtain a mapping matrix corresponding to each channel; The mapping matrix and the low-illuminance image to be detected are added to obtain an enhanced image of the low-illuminance image to be detected.
5. The method according to claim 1, characterized in that Before obtaining the low-light image to be detected in the target scene, the following steps are also included: Acquire a training set; wherein the training set includes: a first low-light image in different scenes and a second low-light image in a target scene; Pre-training an image enhancement network based on the first low-illumination image to obtain a trained image enhancement network; Loading the trained image enhancement network into a low-light end-to-end object detection model, and freezing model parameters of the trained image enhancement network; The low-illumination end-to-end object detection model is trained based on the second low-illumination image to obtain a trained low-illumination end-to-end object detection model.
6. The method according to claim 5, characterized in that Pre-training an image enhancement network based on the first low-illumination image to obtain a trained image enhancement network includes: Inputting the first low-light image into an image enhancement network to obtain an adjustment factor matrix corresponding to each of the first low-light images; The adjustment factor matrix is adaptively learned based on a preset loss function until a preset number of iterations is reached to obtain a trained image enhancement network; wherein the loss function includes: a spatial consistency loss function, an exposure suppression loss function, a color constancy loss function and an illumination smoothness loss function.
7. The method according to claim 5, characterized in that The low-illumination end-to-end object detection model is trained based on the second low-illumination image to obtain a trained low-illumination end-to-end object detection model, including: determining, based on the prompt encoder, prompt coding information of each of the second low-light images; Determining an adjustment curve corresponding to each of the second low-light images based on the trained image enhancement network and the prompt encoding information; Obtaining an enhanced image of each second low-illuminance image based on the adjustment curve corresponding to each second low-illuminance image and the second low-illuminance image; determining an object detection result of each of the second low-illuminance images based on the enhanced image of each of the second low-illuminance images and the object detector; Based on the enhanced image and the target detection result of each of the second low-light images, the joint loss of the trained image enhancement network and the target detector is calculated, and the model parameters of the low-light end-to-end target detection model are optimized based on the joint loss to obtain a trained low-light end-to-end target detection model.
8. A low-light image target detection device, characterized in that: include: An image acquisition module is used to acquire a low-light image to be detected in a target scene; The target detection module is used to input the low-illuminance image to be detected into a pre-trained low-illuminance end-to-end target detection model to obtain an enhanced image and a target detection result of the low-illuminance image to be detected; wherein the low-illuminance end-to-end target detection model includes: a prompt encoder, a trained image enhancement network and a target detector; the prompt encoder outputs the prompt coding information of the low-illuminance image to be detected to the trained image enhancement network, the trained image enhancement network enhances the low-illuminance image to be detected based on the prompt coding information, and outputs the enhanced image to the target detector, the target detector performs target detection on the enhanced image, and outputs the target detection result.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are performed.
Citation Information
Cited By
Low-illumination image target detection method and device based on color channel conversion enhancement
CN120912915A
A low-illumination image target detection method and device based on color channel transformation enhancement
CN120912915B
Target detection method and device, electronic equipment and computer program product
CN120912937A