Low-illumination image object detection method and device, equipment and storage medium

The pre-trained object recognition model filters the illumination constant information, which solves the robustness of object detection in low-illumination environments and achieves high-precision image object recognition.

CN120147665APending Publication Date: 2025-06-13SUZHOU GAIDE PHOTOELECTRIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510324422.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing low-illumination object detection algorithms rely on complex image decomposition processes and lack robustness in unknown scenes, making it difficult to accurately identify targets in low-light environments.

Method used

Through the pre-trained first object recognition model, the illumination constant information is filtered out and used for image object recognition in low-light environments, avoiding the complex image decomposition process. The model includes at least a backbone network, a KM-H prior decoder and a detection head.

Benefits of technology

It significantly improves the image object recognition ability in low-light environments, improves robustness, and is suitable for high-precision image object detection scenarios under low-light conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147665A_ABST
    Figure CN120147665A_ABST
Patent Text Reader

Abstract

The invention discloses a low-illumination image object detection method and device, equipment and a storage medium. The method comprises the steps that a first image to be subjected to image object recognition is obtained, and the first image is image information obtained through shooting under the condition that the illumination intensity is lower than preset illumination intensity; a target image object in the first image is determined according to a first object recognition model obtained through pre-training and the first image, and the first object recognition model at least comprises a backbone network, a KM-H prior decoder and a detection head. According to the technical scheme of the embodiment of the invention, the illumination invariant information is selectively screened out through the first object recognition model, the recognition capability of the image object in a low-illumination environment can be remarkably improved, a complex image decomposition process is not needed, and the robustness in practical application is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and particularly to a method, device, equipment and storage medium for detecting objects in low-light images. Background Art

[0002] In recent years, general object detection technology has made remarkable progress in the field of computer vision and is widely used in many scenarios such as autonomous driving, intelligent monitoring, medical image analysis, etc. However, although the performance of existing algorithms under standard lighting conditions has approached or even exceeded the human level, their performance in low-light environments still has obvious deficiencies. Object detection tasks in low-light environments face many challenges, such as image quality degradation, noise increase, and loss of key details. These problems seriously affect the detection accuracy and robustness, making it difficult for traditional feature extraction-based detection methods to accurately identify objects.

[0003] Currently, object detection methods under low-light conditions can be divided into image enhancement methods and low-light adaptive fine-tuning methods. Image enhancement methods improve the image quality, providing clearer input data for the object detector, thereby improving the detection performance; the low-light adaptive fine-tuning method is to further fine-tune the detector trained under normal lighting conditions on low-light images to make it better adapt to complex lighting environments.

[0004] Although the above two solutions can alleviate the low-light object detection problem to a certain extent, they both highly rely on a large number of low-light images collected from the real world. However, compared with well-lit images, the proportion of low-light images in existing public datasets is quite low, and the acquisition scenarios are also relatively single. At the same time, due to the low visibility of objects in low-light images, it is much more difficult to label the object bounding boxes on low-illuminance images. The difficulties in collecting and labeling low-light data have hindered the development of low-illuminance object detection.

[0005] To overcome the limitations brought by the lack of low-light data, Attila L, Sourav G, Michael M, Jan C et al. proposed a method based on zero-shot light and dark domain adaptation, that is, trying to train an object detector only using well-lit images and requiring the model to have strong generalization ability when migrating the target domain to a low-light environment. To achieve the above goal, we must be able to obtain "domain-invariant representations" from the source domain that can span the target domain, that is, the model should learn illumination-invariant information on well-lit images.

[0006] However, existing low-light object detection algorithms have the following disadvantages: The method based on zero-shot light and dark domain adaptation relies on image decomposition, which is challenging. Existing methods rely on manually formulated strategies or paired training data with different brightness levels, and these methods lack robustness when facing unknown scenarios. Summary of the Invention

[0007] The present invention provides a method, device, equipment and storage medium for detecting objects in low-light images. By selectively screening out illumination-invariant information through a first object recognition model, the recognition ability of image objects in a low-light environment can be significantly improved, without a complex image decomposition process, and the robustness in practical applications is enhanced.

[0008] According to one aspect of the present invention, a method for detecting objects in low-light images is provided. The method includes:

[0009] Obtain a first image to be recognized for image objects, where the first image is image information obtained by shooting under a light intensity lower than a preset light intensity;

[0010] Determine a target image object in the first image according to a first object recognition model obtained by pre-training and the first image, where the first object recognition model at least includes a backbone network, a KM-H prior decoder, and a detection head.

[0011] According to another aspect of the present invention, a device for detecting objects in low-light images is provided. The device includes:

[0012] A first image acquisition module, configured to obtain a first image to be recognized for image objects, where the first image is image information obtained by shooting under a light intensity lower than a preset light intensity;

[0013] A target object determination module, configured to determine a target image object in the first image according to a first object recognition model obtained by pre-training and the first image, where the first object recognition model at least includes a backbone network, a KM-H prior decoder, and a detection head.

[0014] According to another aspect of the present invention, an electronic device is provided, and the electronic device includes:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; where

[0017] The memory stores a computer program executable by the at least one processor. When executed by the at least one processor, the computer program enables the at least one processor to perform the low-light image object detection method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the low-light image object detection method according to any embodiment of the present invention when executed.

[0019] The technical solution of the embodiment of the present invention performs image object recognition processing on a first image captured under a light intensity lower than a preset light intensity through a first object recognition model obtained by pre-training, solves the problem that the light and dark domain adaptive target detection method overly relies on a complex image decomposition process, and through the pre-trained first object recognition model, achieves high robustness to light changes, and is particularly suitable for high-precision image object detection scenarios under low-light conditions.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 is a flowchart of a low-light image object detection method provided in Embodiment 1 of the present invention;

[0023] Figure 2 is a flowchart of a low-light image object detection method provided in Embodiment 2 of the present invention;

[0024] Figure 3 is a structural diagram of a low-light image object detection device provided in Embodiment 3 of the present invention;

[0025] Figure 4 is a schematic structural diagram of an electronic device for implementing the low-light image object detection method of the embodiment of the present invention. Detailed Embodiments

[0026] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] Embodiment 1

[0029] Figure 1 FIG. 10 is a flowchart of a low-light image object detection method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of identifying and confirming objects in an image, especially applicable to the situation of object recognition in an image taken under low light. This method can be executed by a low-light image object detection device, which can be implemented in the form of hardware and / or software, and the low-light image object detection device can be configured in an electronic device. As Figure 1 shown, the method includes:

[0030] S101. Obtain a first image to be used for image object recognition.

[0031] It should be noted that the first image can be image information obtained by shooting under a light intensity lower than a preset light intensity. The preset light intensity is not specifically limited in the present invention. That is to say, the technical solution of the embodiment of the present invention can be applicable to the first image taken in a low-light environment. However, it should be noted that the technical solution of the embodiment of the present invention can also be applicable to the first image taken in a high-light environment. Specifically, the technical solution of the embodiment of the present invention can be applicable to the first image taken under any light intensity greater than 0.1 lux.

[0032] Specifically, in the intelligent security scenario, a first image is captured by a security camera device; in the autonomous driving scenario, a first image is captured by an in-vehicle camera device; in the UAV aerial photography scenario, a first image is captured by an aerial photography device; in the medical surgery scenario, a first image is captured by an abdominal cavity imaging device. In different application scenarios, an image under any light intensity greater than 0.1 lux in that scenario is obtained and used as the first image.

[0033] S102. Determine the target image object in the first image according to the first object recognition model obtained through pre-training and the first image.

[0034] It should be noted that the first object recognition model is an object detection model equipped with a KM-H prior decoder. The first object recognition model is an object detection model obtained by training with the pseudo ground truth of the KM-H prior generated by the prior extraction sub-network and training samples. The first object recognition model is similar to a teacher-student model. The prior extraction sub-network serves as the "teacher" to generate pseudo labels, and the KM-H prior decoder serves as the "student" to learn the prior representation adapted to the target task.

[0035] Among them, the first object recognition model includes at least a backbone network, a KM-H prior decoder, and a detection head.

[0036] The backbone network is a core component in a deep learning model. In the technical solution of the embodiment of the present invention, the VGG16 pre-trained model in Pytorch can be used as the backbone network. It is responsible for extracting multi-level feature representations from input data (such as images, texts, etc.). It usually serves as the "backbone" of the model and provides basic features for subsequent tasks (such as classification, detection, segmentation, etc.).

[0037] The KM-H prior decoder is a component used to generate high-quality images in image generation or image reconstruction tasks. It can capture the key features of the data, optimize the image generation process by introducing prior knowledge, and guide the decoder to generate high-quality and more realistic images that conform to the real data distribution.

[0038] The detection head is a key component in an object detection model. It is responsible for completing the tasks of target localization and classification based on the features extracted by the backbone network. It is usually located at the end of the model and directly outputs the detection results (such as bounding boxes and categories). The detection head can efficiently detect targets in images and has a wide range of applications in the field of computer vision.

[0039] Specifically, the first image is input into a first object recognition model for object recognition, and according to the output of the first object recognition model, the target image object in the first image is determined. It should be noted that the output result of the first object recognition model depends on the type of training samples. In different application scenarios, after being trained with corresponding sample data, the first object recognition model can accurately recognize the corresponding image object in the first image, and is particularly applicable to the first image with relatively low light intensity.

[0040] Exemplarily, determining the target image object in the first image according to the pre-trained first object recognition model and the first image includes: based on the backbone network in the first object recognition model, performing feature extraction processing on the first image to obtain a shallow feature image; based on the KM-H prior decoder in the first object recognition model, determining the illumination-invariant information in the shallow feature image; based on the detection head in the first object recognition model, determining the target image object in the first image according to the illumination-invariant information.

[0041] It should be noted that illumination changes (such as overexposure, shadows, low light) will distort the color, contrast, and texture in the image, resulting in a shift in the feature map distribution. Shallow features (such as edges and corners) are sensitive to illumination, and stable illumination-invariant information needs to be extracted through specific processing. The illumination invariance of the illumination-invariant information aims to enable the model to maintain a stable feature extraction ability under different illumination conditions. Especially in the object detection task, it directly affects the recognition of small objects, boundary localization, and robustness in complex environments.

[0042] Specifically, after the first image is input into the first object recognition model, the first image is subjected to feature extraction processing through the backbone network to obtain feature images of each layer. The shallow feature image is separately retained as an intermediate result and sent into the KM-H prior decoder to determine the illumination-invariant information. The detection head in the first object recognition model identifies the target image object in the first image according to the illumination-invariant information.

[0043] The technical solution of the embodiment of the present invention uses the pre-trained first object recognition model to perform image object recognition processing on the first image obtained by shooting under a preset light intensity lower than the preset value, solving the problem that the bright-dark domain adaptive object detection method overly relies on a complex image decomposition process. The pre-trained first object recognition model achieves high robustness to illumination changes and is particularly applicable to the high-precision image object detection scenario under low light conditions.

[0044] Embodiment 2

[0045] Figure 2The flowchart of a low-light image object detection method provided in the second embodiment of the present invention. On the basis of the above embodiments, the training process of the first object recognition model is refined. As Figure 2 shown, the method includes:

[0046] S201. Obtain first sample images under different lighting conditions.

[0047] Among them, the first sample images can be sample images determined in advance for the to-be-used scenarios. The first sample images can include indoor sample images, outdoor sample images, daytime sample images, nighttime sample images, etc., which can provide diverse training samples for the model.

[0048] S202. Extract the illumination-invariant information of the first sample images through a preset object recognition model, and determine the target training error during the extraction process.

[0049] Among them, the preset object recognition model can refer to an object detection model including an initialized KM-H prior decoder.

[0050] Specifically, input the first sample images into the preset object recognition model for the object recognition process, and extract the illumination-invariant information of the first sample images. At the same time, during the extraction process, determine the target training error of the preset object recognition model during the training process.

[0051] Exemplarily, the extracting the illumination-invariant information of the first sample images through a preset object recognition model includes:

[0052] Darken the first sample images to obtain second sample images after darkening; input the first sample images and the second sample images into the backbone network in the preset object recognition model for feature extraction processing to obtain a first feature image corresponding to the first sample images and a second feature image corresponding to the second sample images; input the first feature image and the second feature image into the KM-H prior decoder in the preset object recognition model to determine a first illumination-invariant information corresponding to the first feature image and a second illumination-invariant information corresponding to the second feature image.

[0053] Among them, the second sample image may refer to the first sample image after being darkened. Exemplarily, the darkening process can be processed by Dark ISP to simulate the light degradation in the real world and synthesize an artificial dark image, so as to obtain the second sample image for model training. The first feature image may refer to the shallow feature image of the first sample image, and the second feature image may refer to the shallow feature image of the second sample image. The first light invariant information may refer to the light invariant information of the first feature image. The second light invariant information may refer to the light invariant information of the second feature image.

[0054] Specifically, the first sample image with good lighting and its corresponding second sample image of synthesized low light are combined into an image pair and input into a preset object recognition model. Feature extraction is performed through the backbone network. Here, an enhancement strategy for light invariant features is designed to guide the feature extraction network to learn light invariant representations. According to the output of the preset object recognition model, the first feature image and the second feature image are obtained.

[0055] The first feature image and the second feature image are respectively input into the preset object recognition model to extract the first light invariant information corresponding to the first feature image and the second light invariant information corresponding to the second feature image through the KM-H prior decoder.

[0056] Exemplarily, the determination of the target training error in the extraction process includes: determining a first training error according to the first feature image and the second feature image; determining a second training error according to the first light invariant information and the second light invariant information; determining a third training error according to the first light invariant information, the second light invariant information, and the pre-determined pseudo-label light invariant information; and determining the target training error according to the first training error, the second training error, the third training error, and the detection loss error of the detection head.

[0057] Specifically, the determination method of the first training error is as follows:

[0058]

[0059] Among them, represents the KL divergence, F l and F d are respectively the first feature image and the second feature image.

[0060] The determination method of the second training error is as follows:

[0061] L cst =MSE(P l ,P d )+(1 - SSIM(P l,P d ));

[0062] Among them, MSE represents the mean square error, and SSIM represents the structural similarity index. P l and P d respectively represent the first illumination-invariant information and the second illumination-invariant information.

[0063] The third training error can be determined by the third training error formula, calculated based on the first illumination-invariant information, the first pseudo-label information, the second illumination-invariant information, and the second pseudo-label information. The detection loss error of the detection head can be determined according to the loss algorithm of the actual object detection model. Then, the sum of the first training error, the second training error, the third training error, and the detection loss error of the detection head is determined as the target training error.

[0064] Exemplarily, determining the third training error according to the first illumination-invariant information, the second illumination-invariant information, and the pre-determined pseudo-label illumination-invariant information includes: determining a first component error according to the first illumination-invariant information and the first pseudo-label information; determining a second component error according to the second illumination-invariant information and the second pseudo-label information; determining the third training error according to the first component error and the second component error.

[0065] It should be noted that the pseudo-label illumination-invariant information includes the first pseudo-label information corresponding to the first sample image and the second pseudo-label information corresponding to the second sample image. The pseudo-label illumination-invariant information provides reliable prior knowledge for the KM-H prior decoder, and finally realizes high-precision object recognition under complex illumination.

[0066] Among them, the first pseudo-label information and the second pseudo-label information can be obtained through a pre-trained pseudo-label generator. Exemplarily, the first sample image and the second sample image can be respectively input into the pseudo-label generator, so that the first pseudo-label information and the second pseudo-label information can be respectively obtained.

[0067] The pseudo-label generator is trained through low-light enhancement tasks (such as low-light image restoration), and its core objective is to learn how to extract robust features independent of illumination from low-light images. It should be noted that although the pseudo-label generator can generate illumination-invariant information that the KM-H prior decoder ultimately aims to obtain, this illumination-invariant information does not fully fit the object detection task. The ultimate goal of the technical solution of the present invention is to enable the KM-H prior decoder to serve the object detection task. Therefore, it is necessary to retrain the pseudo-label generator on the object detection task and fine-tune the model parameters to make it better adapt to the object detection task. Directly using the pre-trained pseudo-label generator may lead to a performance decline due to the mismatch between its prior representation and the object detection task. Therefore, the present invention chooses to only use the pseudo-label generator to generate pseudo-labels to guide the training of the KM-H prior decoder, rather than directly replacing the position of the KM-H prior decoder. In this way, the model can be retrained and fine-tuned on the object detection task to ensure the adaptability of the KM-H prior decoder to the object detection task, thereby improving the overall performance of the model.

[0068] It should be noted that both the first component error and the second component error can be obtained in the following manner:

[0069]

[0070] where MAE represents the mean absolute error and SSIM represents the structural similarity index; in L prior When it is the first component error, P can refer to the first illumination-invariant information, and can refer to the first pseudo-label information; in L prior When it is the second component error, P can refer to the second illumination-invariant information, and can refer to the second pseudo-label information.

[0071] Determine the sum of the first component error and the second component error as the third training error.

[0072] S203. Backpropagate the target training error into the preset object recognition model and adjust the network parameters of the preset object recognition model.

[0073] Specifically, backpropagate the target training error into the preset object recognition model so that the preset object recognition model adjusts the network parameters of the backbone network, the KM-H prior decoder, and the detection head respectively, realizes optimizing the network to continuously update and optimize the parameters, thereby guiding the feature extraction network to learn the illumination-invariant representation, and still maintaining the accuracy and precision of the entire detection network when migrating to the low-illumination environment in the target domain.

[0074] S204. When the preset convergence condition is met, it is determined that the training of the preset object recognition model is completed, and the first object recognition model is obtained.

[0075] Among them, the preset convergence condition may include reaching the preset number of training iterations or the target training error being less than the preset training error. Specifically, when the preset convergence condition is met, it indicates that the training of the preset object recognition model is completed. At this time, the preset object recognition model with the training completed can be determined as the first object recognition model.

[0076] Exemplarily, after obtaining the first object recognition model, it further includes: obtaining a first verification image and the first verification object corresponding to the first verification image; inputting the first verification image into the first object recognition model for object recognition to obtain an object recognition result, and determining the recognition error of the first object recognition model according to the object recognition result and the first verification object; in the case where the recognition error is greater than the preset error, continuing to train the first object recognition model based on the first sample image.

[0077] Among them, the first verification image can be a test image of the first object recognition model to verify the adaptability of the first object recognition model under different lighting conditions. The first verification image is specifically designed for the target detection task in a low-light environment. Since the first sample image has extensive data coverage and rich lighting variations, and can cover all target categories in the first verification image at the same time, it is very suitable for verifying the generalization ability of the model under more complex lighting conditions. The first verification object can refer to the object to be recognized in the first verification image. The first verification object can be obtained through manual annotation.

[0078] Specifically, input each first verification image into the first object recognition model for object recognition to obtain the object recognition result corresponding to each first verification image. By comparing the object recognition result with the first verification object, the recognition error of the first object recognition model can be determined. When the recognition error is greater than the preset error, it indicates that the recognition error of the first object recognition model is large and the recognition accuracy is poor. Continue to train the first object recognition model using the first sample image. When the recognition error is less than the preset error, it indicates that the recognition error of the first object recognition model is small and it can be applied in practice.

[0079] S205. Obtain a first image for which image object recognition is to be performed.

[0080] S206. Determine the target image object in the first image according to the first object recognition model obtained by pre-training and the first image.

[0081] In the technical solution of the embodiment of the present invention, starting from the Kubelka-Munk theory and HSV hue information, a learnable illumination-invariant prior (KM-H prior) is innovatively proposed. The light and dark domain consistency network (LDC-Net) constructed based on this prior extracts the KM-H prior representation from the shallow feature map through the decoder branch, and uses this to guide the feature extraction network to selectively filter out the illumination-invariant information for detector training. This method can significantly improve the generalization ability of the baseline model in low-light environments.

[0082] Embodiment III

[0083] Figure 3 FIG. is a schematic structural diagram of a low-light image object detection device provided in Embodiment III of the present invention. As Figure 3 shown, the device includes:

[0084] A first image acquisition module 301, configured to acquire a first image to be subjected to image object recognition, where the first image is image information captured under a light intensity lower than a preset light intensity;

[0085] A target object determination module 302, configured to determine a target image object in the first image according to a first object recognition model obtained through pre-training and the first image, where the first object recognition model at least includes a backbone network, a KM-H prior decoder, and a detection head.

[0086] In the technical solution of the embodiment of the present invention, through the first object recognition model obtained through pre-training, the first image captured under a light intensity lower than the preset light intensity is subjected to image object recognition processing, which solves the problem that the light and dark domain adaptive target detection method overly relies on a complex image decomposition process. The first object recognition model obtained through pre-training achieves high robustness to illumination changes and is particularly suitable for high-precision image object detection scenarios under low-light conditions.

[0087] Optionally, the target object determination module 302 is specifically configured to:

[0088] Based on the backbone network in the first object recognition model, perform feature extraction processing on the first image to obtain a shallow feature image;

[0089] Based on the KM-H prior decoder in the first object recognition model, determine the illumination-invariant information in the shallow feature image;

[0090] Based on the detection head in the first object recognition model, determine the target image object in the first image according to the illumination-invariant information.

[0091] Optionally, the device further includes a model training module. Wherein,

[0092] The model training module includes:

[0093] A sample image acquisition unit for acquiring first sample images under different lighting conditions, where the first sample images include at least indoor sample images, outdoor sample images, daytime sample images, and nighttime sample images;

[0094] A training error determination unit for extracting illumination-invariant information of the first sample images through a preset object recognition model and determining the target training error during the extraction process;

[0095] A network parameter adjustment module for backpropagating the target training error into the preset object recognition model to adjust the network parameters of the preset object recognition model;

[0096] An identification model training unit for determining that the training of the preset object recognition model is completed and obtaining a first object recognition model when a preset convergence condition is met.

[0097] Optionally, the training error determination unit is used for:

[0098] Darkening the first sample images to obtain second sample images after darkening processing;

[0099] Inputting the first sample images and the second sample images into the backbone network in the preset object recognition model respectively for feature extraction processing to obtain a first feature image corresponding to the first sample images and a second feature image corresponding to the second sample images;

[0100] Inputting the first feature image and the second feature image into the KM-H prior decoder in the preset object recognition model respectively to determine a first illumination-invariant information corresponding to the first feature image and a second illumination-invariant information corresponding to the second feature image.

[0101] Optionally, the training error determination unit further includes:

[0102] A first error determination subunit for determining a first training error according to the first feature image and the second feature image;

[0103] A second error determination subunit for determining a second training error according to the first illumination-invariant information and the second illumination-invariant information;

[0104] A third error determination subunit, configured to determine a third training error according to the first illumination invariant information, the second illumination invariant information, and pre-determined pseudo-label illumination invariant information, where the pseudo-label illumination invariant information includes first pseudo-label information corresponding to the first sample image and second pseudo-label information corresponding to the second sample image;

[0105] A fourth error determination subunit, configured to determine a target training error according to the first training error, the second training error, the third training error, and a detection loss error of a detection head.

[0106] Optionally, the third error determination subunit is specifically configured to:

[0107] Determine a first component error according to the first illumination invariant information and the first pseudo-label information;

[0108] Determine a second component error according to the second illumination invariant information and the second pseudo-label information;

[0109] Determine a third training error according to the first component error and the second component error.

[0110] Optionally, the model training module is further configured to:

[0111] After obtaining a first object recognition model, obtain a first verification image and a first verification object corresponding to the first verification image;

[0112] Input the first verification image into the first object recognition model for object recognition, obtain an object recognition result, and determine an identification error of the first object recognition model according to the object recognition result and the first verification object;

[0113] In a case where the identification error is greater than a preset error, continue to train the first object recognition model based on the first sample image.

[0114] The low-light image object detection device provided by the embodiments of the present invention can execute the low-light image object detection method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0115] Embodiment 4

[0116] Figure 4FIG. 0 shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0117] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0118] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0119] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the low-light image object detection method.

[0120] In some embodiments, the low-light image object detection method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the low-light image object detection method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the low-light image object detection method by any other suitable means (e.g., by means of firmware).

[0121] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0122] The computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0123] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0125] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0126] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0127] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0128] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting objects in low-light images, characterized in that: include: Acquire a first image to be subjected to image object recognition, wherein the first image is image information obtained by shooting under a light intensity lower than a preset light intensity; According to a first object recognition model obtained by pre-training and the first image, a target image object in the first image is determined, wherein the first object recognition model at least includes a backbone network, a KM-H prior decoder and a detection head.

2. The method according to claim 1, characterized in that The step of determining a target image object in the first image based on a pre-trained first object recognition model and the first image includes: Based on the backbone network in the first object recognition model, performing feature extraction processing on the first image to obtain a shallow feature image; Determining illumination invariant information in the shallow feature image based on a KM-H prior decoder in the first object recognition model; Based on the detection head in the first object recognition model, a target image object in the first image is determined according to the illumination invariant information.

3. The method according to claim 1, characterized in that The training process of the first object recognition model includes: Acquire first sample images under different lighting conditions, wherein the first sample images at least include indoor sample images, outdoor sample images, daytime sample images, and nighttime sample images; Extracting illumination invariant information of the first sample image by using a preset object recognition model, and determining a target training error during the extraction process; Back-propagating the target training error to the preset object recognition model to adjust the network parameters of the preset object recognition model; When the preset convergence condition is met, it is determined that the training of the preset object recognition model is completed, and a first object recognition model is obtained.

4. The method according to claim 3, characterized in that The extracting illumination invariant information of the first sample image by using a preset object recognition model includes: Darkening the first sample image to obtain a darkened second sample image; Inputting the first sample image and the second sample image into the backbone network in the preset object recognition model for feature extraction processing, respectively, to obtain a first feature image corresponding to the first sample image and a second feature image corresponding to the second sample image; The first feature image and the second feature image are respectively input into the KM-H priori decoder in the preset object recognition model to determine first illumination invariant information corresponding to the first feature image and second illumination invariant information corresponding to the second feature image.

5. The method according to claim 4, characterized in that The determining of the target training error in the extraction process includes: Determining a first training error according to the first feature image and the second feature image; Determining a second training error according to the first illumination invariant information and the second illumination invariant information; Determining a third training error according to the first illumination invariant information and the second illumination invariant information, and predetermined pseudo-label illumination invariant information, wherein the pseudo-label illumination invariant information includes first pseudo-label information corresponding to the first sample image and second pseudo-label information corresponding to the second sample image; A target training error is determined according to the first training error, the second training error, the third training error, and a detection loss error of a detection head.

6. The method according to claim 5, characterized in that The determining a third training error according to the first illumination invariant information, the second illumination invariant information, and predetermined pseudo-label illumination invariant information includes: determining a first component error according to the first illumination invariant information and the first pseudo-label information; determining a second component error according to the second illumination invariant information and the second pseudo label information; A third training error is determined based on the first component error and the second component error.

7. The method according to claim 3, characterized in that After obtaining the first object recognition model, the method further includes: Acquire a first verification image and a first verification object corresponding to the first verification image; Inputting the first verification image into the first object recognition model for object recognition to obtain an object recognition result, and determining a recognition error of the first object recognition model based on the object recognition result and the first verification object; When the recognition error is greater than a preset error, the first object recognition model is continuously trained based on the first sample image.

8. A low-light image object detection device, characterized in that: include: A first image acquisition module, used to acquire a first image to be subjected to image object recognition, wherein the first image is image information obtained by shooting under a light intensity lower than a preset light intensity; A target object determination module is used to determine the target image object in the first image based on a pre-trained first object recognition model and the first image, wherein the first object recognition model at least includes a backbone network, a KM-H prior decoder and a detection head.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the low-light image object detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the low-light image object detection method according to any one of claims 1 to 7 when executed.