3D-LUT-based target detection method, device, electronic device, system, and storage medium
The target detection method that combines 3D-LUT and YOLOv8 algorithms solves the problem of poor image enhancement and target detection in low-light environments, achieves high-precision target detection and image quality enhancement, and adapts to various lighting conditions.
Patent Information
- Application Number
- CN202410269108.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-11
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-03-11
AI Technical Summary
Existing image enhancement and target detection technologies perform poorly in low-light environments and cannot effectively handle the problems of detail loss and noise amplification. In particular, deep learning-based target detection algorithms perform poorly in low-light conditions.
A 3D-LUT-based target detection method is adopted. The trained 3D-LUT model and convolutional neural network CNN are combined with the interpolation method to perform image enhancement processing on dark and bright light sample images, and the YOLOv8 algorithm is used for target detection to achieve image feature capture and target detection under low illumination conditions.
It improves the accuracy and reliability of target detection, enhances image quality, reduces noise, can adapt to dynamically changing lighting conditions, and improves target detection effects in low-light and complex environments.
Smart Images

Figure CN118135251B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a 3D-LUT-based target detection method, device, electronic device, system, and storage medium. Background Art
[0002] Image enhancement and object detection in low-light environments have always been challenging in the field of image processing. While traditional image enhancement techniques, such as histogram equalization and gamma correction, can improve image brightness and contrast to a certain extent, they are generally unable to effectively address detail loss and noise amplification in low-light environments. Furthermore, existing image enhancement techniques employ a global processing approach and are unable to optimize local features within an image.
[0003] Target detection technologies in related technologies, especially deep learning-based target detection networks, such as the YOLO series of neural networks, have achieved remarkable results in environments with good lighting conditions, that is, they have good target recognition effects for target detection in bright-light photos. However, these algorithm models are rarely trained on image data collected in low-light environments, resulting in loss of details and increase in noise in image data under low-light conditions. Therefore, they cannot cope with target detection in low-light or uneven lighting environments, and the target detection effect is poor.
[0004] At present, no effective solution has been proposed to the problem that target detection technology in related technologies cannot cope with low-light and complex environments and the target detection effect is poor. Summary of the Invention
[0005] The embodiments of the present application provide a 3D-LUT-based target detection method, device, electronic device, system and storage medium to at least solve the problem in the related art that target detection technology cannot cope with low-light complex environments and the target detection effect is poor.
[0006] In a first aspect, an embodiment of the present application provides a target detection method based on 3D-LUT, comprising: obtaining a first image frame to be detected in a video frame of acquired video data, wherein the first image frame includes a target image taken under dark light conditions; using a trained first 3D-LUT model and a preset interpolation method, performing brightness and contrast enhancement processing on the first image frame to obtain an enhanced image frame, wherein the first 3D-LUT model uses a currently generated second 3D-LUT model to perform image enhancement verification on a paired dark light sample image and a bright light sample image, and guides the currently generated second 3D-LUT model to perform image enhancement verification on the paired dark light sample image and the ... A convolutional neural network (CNN) trained by iterative optimization is performed, and the interpolation method includes one of the following: trilinear interpolation, nonlinear interpolation, and neighboring order interpolation; based on a target detection model, target detection is performed in the enhanced image frame to obtain a detection result, wherein the target detection model is based on the YOLOv8 algorithm and is a neural network model trained according to a second image frame and a measured target corresponding to the second image frame, the second image frame is generated by enhancing a preset sample image frame by using the trained first 3D-LUT model, and the sample image frame includes a dark light sample image frame and a bright light sample image frame mixed in a preset ratio.
[0007] In a second aspect, an embodiment of the present application provides a target detection device based on a 3D-LUT, comprising: an acquisition module for acquiring a first image frame to be detected in a video frame of acquired video data, wherein the first image frame includes a target image taken under dark light conditions; an enhancement module for using a trained first 3D-LUT model and a preset interpolation method to perform brightness and contrast enhancement processing on the first image frame to obtain an enhanced image frame, wherein the first 3D-LUT model uses a currently generated second 3D-LUT model to perform image enhancement verification on a paired dark light sample image and a bright light sample image, and guides the currently generated second 3D-LUT model to enhance the brightness and contrast of the first image frame. A convolutional neural network CNN trained by iterative optimization of a UT model, wherein the interpolation method includes one of the following: trilinear interpolation, nonlinear interpolation, and adjacent order interpolation; an identification module is used to perform target detection in the enhanced image frame based on the target detection model to obtain a detection result, wherein the target detection model is based on the YOLOv8 algorithm and is a neural network model trained according to the second image frame and the measured target corresponding to the second image frame, and the second image frame is generated by enhancing the preset sample image frame by the trained first 3D-LUT model, and the sample image frame includes a dark light sample image frame and a bright light sample image frame mixed in a preset ratio.
[0008] In a third aspect, an embodiment of the present application provides a target recognition system, comprising: an acquisition module, a transmission device, and a server device; wherein, the acquisition module is connected to the server device via the transmission device; the acquisition module is used to acquire video data; the transmission device is used to transmit the video data to the server device; the server device is used to execute the 3D-LUT-based target detection method described in the first aspect.
[0009] In a fourth aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the 3D-LUT-based target detection method described in the first aspect.
[0010] In a fifth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored. When the program is executed by a processor, the target detection method based on 3D-LUT as described in the first aspect above is implemented.
[0011] Compared with the related art, the target detection method, device, electronic device, system and storage medium based on 3D-LUT provided by the embodiments of the present application obtain a first image frame to be detected in the video frame of the collected video data; use the trained first 3D-LUT model and the preset interpolation method to perform brightness and contrast enhancement processing on the first image frame to obtain an enhanced image frame, wherein the first 3D-LUT model is a convolutional neural network CNN trained by using the currently generated second 3D-LUT model to perform image enhancement verification on the paired dark light sample image and the bright light sample image, and guide the currently generated second 3D-LUT model to perform iterative optimization; based on the target detection model, target detection is performed in the enhanced image frame to obtain a detection result, wherein the target detection model is based on the YOLOv8 algorithm, and according to the second image frame and the first The neural network model is trained by the measured target corresponding to the second image frame. The second image frame is generated by enhancing the preset sample image frame through the trained first 3D-LUT model. The sample image frame includes a dark light sample image frame and a bright light sample image frame mixed in a preset ratio. The video image frame is enhanced through the paired trained 3D-LUT model in combination with the set interpolation method to achieve accurate pixel color mapping, effectively enhance image quality, reduce noise while retaining important details, and process the training image of the target detection model through the 3D-LUT model in combination with the set interpolation method to improve the target detection accuracy and reliability, and can effectively capture low-light image features, improve the adaptability to dynamically changing lighting conditions, and solve the problem that the target detection technology in the related technology cannot cope with weak light and complex environments and the target detection effect is poor.
[0012] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0014] Figure 1 3D-LUT-based target detection method of the embodiment of the present application is a hardware structure block diagram of the terminal;
[0015] Figure 2 is a flowchart of a 3D-LUT-based target detection method according to an embodiment of the present application;
[0016] Figure 3 3D-LUT is a block diagram of a target detection device according to an embodiment of the present application. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for ordinary technicians in the field related to the contents disclosed in the present application, some changes such as design, manufacturing or production based on the technical contents disclosed in the present application are only conventional technical means and should not be understood as the contents disclosed in the present application being insufficient.
[0018] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments unless there is a conflict.
[0019] Unless otherwise defined, technical or scientific terms used in this application shall have the ordinary meaning as understood by persons of ordinary skill in the art to which this application belongs. The use of "a," "an," "an," "the," and similar expressions in this application does not denote a limitation of quantity and may refer to either the singular or the plural. The terms "comprise," "include," "have," and any variations thereof, as used in this application, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device comprising a series of steps or modules (units) is not limited to the listed steps or units but may also include steps or units not listed, or may include other steps or units inherent to the process, method, product, or device. As used in this application, "multiple steps" means two or more steps. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" may mean: A exists alone, A and B exist simultaneously, or B exists alone. The terms "first," "second," and "third," etc., as used in this application, simply distinguish similar objects and do not imply a specific ordering of the objects.
[0020] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. Taking running on a terminal as an example, Figure 1 3D-LUT-based target detection method of the embodiment of the present application is a hardware structure diagram of the terminal. Figure 1 As shown, the terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0021] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the target detection method based on 3D-LUT in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0022] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of terminal 10. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0023] This embodiment provides a 3D-LUT-based target detection method running on the above terminal. Figure 2 : is a flow chart of a target detection method based on 3D-LUT according to an embodiment of the present application. Figure 2 As shown, the process includes the following steps:
[0024] Step S201 : obtaining a first image frame to be detected from the video frames of the collected video data, wherein the first image frame includes a target image captured under dark light conditions.
[0025] In this embodiment, video data is acquired through a deployed acquisition device (e.g., a camera), and the video data exists in the form of video image frames one by one, each video image frame corresponding to an image, and the video image frames acquired by the acquisition device include images shot under low-light or dark-light conditions. It can be understood that, in this embodiment, the target image frames to be processed are mainly video image frames corresponding to images shot under dark-light conditions (dark-light images), and by enhancing the brightness and contrast of the dark-light images, it is ensured that accurate target detection can be performed on all video frames in the video data. Of course, corresponding processing is also performed on images shot under bright-light conditions (bright-light images) in the video data. However, because the image quality of the bright-light images themselves can meet the requirements for accurate target detection, the magnitude of the enhancement processing performed on the bright-light images is smaller than that on the dark-light images, thereby saving computing resources for processing image data.
[0026] Step S202: Using the trained first 3D-LUT model and a preset interpolation method, brightness and contrast enhancement processing is performed on the first image frame to obtain an enhanced image frame, wherein the first 3D-LUT model is a convolutional neural network CNN trained by using the currently generated second 3D-LUT model to perform image enhancement verification on paired dark light sample images and bright light sample images, and guiding the currently generated second 3D-LUT model to perform iterative optimization. The interpolation method includes one of the following: trilinear interpolation, nonlinear interpolation, and neighboring order interpolation.
[0027] In this embodiment, the 3D-LUT model is a method of enhancing the pixel values in an image using a three-dimensional lookup table. The 3D-LUT can map and process all color information, including non-existent colors and color gamuts that cannot be reached by existing images. The 3D-LUT is an ideal choice for processing complex color transformations. It should be understood that the 3D-LUT can affect hue, saturation, and brightness in a fully three-dimensional color space. The 3D-LUT is composed of three RGB 1D-LUTs, that is, the color values of the three input RGB color channels are mapped according to the three lookup tables of the 3D-LUT to obtain the converted color. For example, if a pixel value in the two images is 80, and the trained 3D-LUT contains a mapping group f(80)=180, then after enhancement processing by the 3D-LUT model, the pixel value 80 is updated to the pixel value 180. It should also be understood that the size of the 3D-LUT represents the size of the pixel mapping group it contains, that is, each pixel value is 80. The number of pixel mapping groups for each color channel (corresponding to the color levels on the three primary color channels R (red), G (green), and B (blue)) is determined. For example, if the 3D-LUT is set to 8×8×8, it means that there are 8 pixel mapping groups on each primary color channel of the 3D-LUT (the corresponding point in the three-dimensional coordinates is 26144); when the size of the 3D-LUT is set to 64×64×64, there are 64 pixel mapping groups on each primary color channel (the corresponding point in the three-dimensional coordinates is 26144).
[0028] In this embodiment, the training of the first 3D-LUT model does not use input samples, and the dark-light image is enhanced by the continuously generated 3D-LUT model, and the loss calculation of the enhanced image and the bright-light image paired with the corresponding dark-light image is performed (the corresponding pixel offset calculation is performed), and then the enhancement effect of the corresponding 3D-LUT model on the dark-light image is determined, so as to guide the generation method of the 3D-LUT model, and achieve an effect that is infinitely close to artificial brightening through training; for example: a convolutional neural network CNN is used to output a 3D-LUT, at this time, the target pixel in the paired dark-light image and bright-light image is A, and the pixel value of A in the dark-light image is 50, and the pixel value of A in the bright-light image is 230. After the dark-light image is enhanced by the 3D-LUT, the pixel value of A is enhanced to 190, and the enhanced The loss is calculated for the dark-light image and the original paired bright-light image. For the target pixel A, after being processed by the 3D-LUT, its pixel value is still different from the expected target by (230-190=40). 40 is used as the offset to guide the CNN to generate a corresponding new 3D-LUT, so that after the newly generated 3D-LUT processes the paired dark-light image and bright-light image, the pixel value of the target pixel A in the enhanced dark-light image is infinitely close to or equal to 230. At this time, the corresponding 3D-LUT is the first 3D-LUT model obtained by training. It should be noted that in this embodiment, the units (or dimensions) corresponding to the pixel values involved are clear to those skilled in the art, for example: 1 pixel = 0.035278 cm, and the relevant units are commonly used units used in image processing calculations.
[0029] In this embodiment, the training of the 3D-LUT differs from the training of the traditional 3D-LUT in that it adopts a method for obtaining pixel mapping groups without manually setting them. Traditional 3D-LUT training relies on manually setting pixel mapping groups for image enhancement, which has the drawbacks of being overly dependent on human intervention and prone to errors in the set pixel mapping groups.
[0030] In this embodiment, a 3D-LUT combined with an interpolation method is used to enhance a dark light image. That is, when the 3D-LUT cannot enhance some pixels, the corresponding pixels are enhanced by the interpolation method. For example, the pixel value of the target pixel B in the first image frame is 60, and there is no mapping corresponding to the pixel value 60 in all pixel mapping groups of the 3D-LUT. At this time, an interpolation method is used to insert a pixel value with a matching enhancement operation effect. For example, a pixel value of 240 is inserted as the corresponding pixel value of the target pixel B after image enhancement. This avoids the situation where the enhanced image quality is not high due to the pixel mapping group not covering all pixels within the color gamut, thereby accelerating the image enhancement speed and enabling real-time update of target detection.
[0031] Step S203: Perform target detection in the enhanced image frame based on the target detection model to obtain a detection result, wherein the target detection model is based on the YOLOv8 algorithm and is a neural network model trained according to the second image frame and the measured target corresponding to the second image frame. The second image frame is generated by enhancing the preset sample image frame through the trained first 3D-LUT model, and the sample image frame includes a dark light sample image frame and a bright light sample image frame mixed in a preset ratio.
[0032] In this embodiment, in order to overcome the limitation that it is only applicable to dark light environments and not suitable for bright light environments, a 3D-LUT is used to enhance the sample images required by the target detection model, and in the training process of the target detection model, a mixture of dark light and bright light environment data sets is used to enable the target detection model to achieve accurate and reliable target detection for images under low illumination conditions and standard brightness environments, thereby improving the accuracy and reliability of target detection. At the same time, in this embodiment, the training of the target detection model is clear and understandable to those skilled in the art, but the training samples used in the training of the target detection model in the embodiment of the present application are processed by the corresponding 3D-LUT model.
[0033] It should be understood that, in this embodiment, the target detection result is to detect the target to be monitored from the video data, such as a person, a vehicle, or a moving object. In this embodiment, there is no specific limitation on the target to be detected.
[0034] Through the above steps S201 to S203, a first image frame to be detected is obtained in the video frame of the collected video data; the first image frame is subjected to brightness and contrast enhancement processing using the trained first 3D-LUT model and the preset interpolation method to obtain an enhanced image frame, wherein the first 3D-LUT model is a convolutional neural network CNN trained by using the currently generated second 3D-LUT model to perform image enhancement verification on the paired dark light sample image and the bright light sample image, and guiding the currently generated second 3D-LUT model to perform iterative optimization; based on the target detection model, target detection is performed in the enhanced image frame to obtain a detection result, wherein the target detection model is based on the YOLOv8 algorithm and is trained according to the second image frame and the measured target corresponding to the second image frame. Neural network model, the second image frame is generated by enhancing the preset sample image frame through the trained first 3D-LUT model. The sample image frame includes a dark light sample image frame and a bright light sample image frame mixed in a preset ratio. The video image frame is enhanced through the paired trained 3D-LUT model in combination with the set interpolation method to achieve accurate pixel color mapping, effectively enhance image quality, and reduce noise while retaining important details. The training image of the target detection model is processed through the 3D-LUT model in combination with the set interpolation method to improve the target detection accuracy and reliability, and can effectively capture low-light image features, improve the adaptability to dynamically changing lighting conditions, and solve the problem that the target detection technology in the related technology cannot cope with weak light and complex environments and the target detection effect is poor.
[0035] It should be noted that the embodiment of the present application proposes an image enhancement method that combines a convolutional neural network (CNN) and a 3D-LUT (three-dimensional lookup table) model to improve image enhancement and target detection in low-light environments. The 3D-LUT model obtained through paired training can specifically adjust the brightness and contrast of the image while maintaining image details and reducing noise, so as to improve image processing and target detection in low-light environments. The target detection method of the embodiment of the present application has a wide range of applications, including but not limited to night monitoring, unmanned vehicles, wildlife observation in low light, and other scenarios.
[0036] It should be further explained that the target detection method of the embodiment of the present application also has the following beneficial effects: First, it enhances image quality and detail preservation. The embodiment of the present application adopts a self-designed CNN and three-dimensional lookup table (3D-LUT) model to achieve the purpose of maintaining image details and reducing noise while enhancing image brightness. Compared with traditional image enhancement technology, this method shows higher efficiency in dealing with the problem of detail loss in images under low light conditions, and provides clearer and more natural image quality; second, it optimizes target detection performance. The embodiment of the present application has made special optimizations for target detection in low-light environments. In order to overcome the limitation that it is only applicable to dark light environments and not to bright light environments, the "3D-LUT + The YOLOv8 algorithm uses a mixture of dark and bright environment data sets during training, so that the target detection model can accurately and reliably detect targets for images in low-light conditions as well as standard brightness environments, thereby improving the accuracy and reliability of target detection. Third, it has strong adaptability to complex environments. The target detection method of the embodiment of the present application is not limited to processing specific types of images or specific lighting conditions, but automatically adapts to various scenes and lighting conditions through the algorithm model, realizing image point operations that do not rely on direct manual settings, but rely on the algorithm for adaptive processing, thereby significantly increasing its flexibility and effectiveness in practical applications.
[0037] In some embodiments, the training of the first 3D-LUT model includes the following steps:
[0038] Step 21: Obtain a current 3D-LUT, input the dark-light sample image into the current 3D-LUT, and obtain a bright-light enhanced sample image corresponding to the dark-light sample image, wherein each color channel of the current 3D-LUT includes a preset number of pixel mapping groups, and the current 3D-LUT includes one of the following output by the CNN network: an initial 3D-LTU, or a 3D-LUT corresponding to a completed iterative optimization.
[0039] In some optional implementations, the sizes of the current 3D-LUT and the first 3D-LUT are set to 16x16x16, and correspondingly, each color channel of the current 3D-LUT includes 16 pixel mapping groups.
[0040] Step 22 : Calculate the pixel offset between the dark light sample image and the bright light sample image based on the bright light sample image and the bright light enhanced sample image paired with the dark light sample image, wherein the pixel offset is used to represent the degree of correction of the pixel values in the pixel mapping group.
[0041] Step 23: Iteratively optimize the current 3D-LUT according to the pixel offset, and use the iteratively optimized 3D-LUT to process the corresponding dark and light sample images until the calculated pixel offset is less than a preset pixel offset threshold, thereby generating a first 3D-LUT model.
[0042] In this embodiment, a CNN network is designed that can output a 16x16x16 current 3D-LUT, and the current 3D-LUT is used to map dark-light images to bright-light images. The CNN network of the embodiment of the present application uses paired dark-light images and bright-light images for network training. The dark-light images are used as image enhancement verification inputs, and the bright-light images are used to calculate the enhancement effect loss, so as to guide the CNN network optimization through the calculated loss to obtain the target 3D-LUT. In this embodiment, the CNN is trained by a paired training method, using one-to-one corresponding pairs of low-light and bright-light images to ensure that the training data covers scenes under various lighting conditions. The 3D-LUT model generated by the training process is used to achieve accurate pixel color mapping, effectively enhance image quality, while retaining important details and reducing noise.
[0043] By obtaining the current 3D-LUT in the above steps, the dark-light sample image is input into the current 3D-LUT to obtain a bright-light enhanced sample image corresponding to the dark-light sample image; based on the bright-light sample image and the bright-light enhanced sample image paired with the dark-light sample image, the pixel offset between the dark-light sample image and the bright-light sample image is calculated; according to the pixel offset, the current 3D-LUT is iteratively optimized, and the corresponding dark-light sample image is processed by the 3D-LUT that has completed the iterative optimization until the calculated pixel offset is less than a preset pixel offset threshold, thereby generating a first 3D-LUT model, thereby achieving training of the first 3D-LUT. Unlike traditional 3D-LUT training that requires manual setting of a pixel mapping group for image enhancement, this method does not rely excessively on human intervention, achieves accurate pixel color mapping, effectively enhances image quality, and retains important details and reduces noise.
[0044] In some embodiments, the pixel offset between the dark light sample image and the bright light sample image is calculated based on the bright light sample image and the bright light enhanced sample image paired with the dark light sample image, which is achieved by the following steps:
[0045] Step 31, respectively obtaining a first pixel set and a second pixel set corresponding to the bright light sample image and the bright light enhanced sample image;
[0046] Step 32: Calculate the mean square error of the pixel values corresponding to the first pixel set and the second pixel set to obtain a pixel offset.
[0047] By obtaining the first pixel set and the second pixel set corresponding to the bright light sample image and the bright light enhanced sample image respectively in the above steps; calculating the mean square error of the pixel values corresponding to the first pixel set and the second pixel set to obtain the pixel offset, it is achieved to guide the optimization of the CNN model and generate a more accurate 3D-LUT
[0048] In some embodiments, processing the first image frame using the trained first 3D-LUT model and a preset interpolation method to obtain an enhanced image frame includes the following steps:
[0049] Step 41: traverse all third pixel values of the first image frame using the first 3D-LUT model.
[0050] In step 42 , it is determined whether there is a fourth pixel value mapped to the third pixel value in all pixel mapping groups corresponding to the first 3D-LUT model, wherein the fourth pixel value is used to represent the pixel value corresponding to the first image frame after highlight enhancement.
[0051] Step 43 : When it is determined that there is a fourth pixel value mapped to the third pixel value, the corresponding third pixel value is updated to the corresponding fourth pixel value to obtain an enhanced image frame.
[0052] In this embodiment, a search is performed based on the third pixel value to determine whether there is a corresponding mapping group. For example, if the third pixel value is a, and there is a mapping group f(a)=b in all pixel mapping groups corresponding to the first 3D-LUT model, then b is the fourth pixel value.
[0053] In step 44, when it is determined that there is no fourth pixel value mapped to the third pixel value, an interpolation method is used to generate a fifth pixel value corresponding to the third pixel value, and the fifth pixel value is used as the pixel value corresponding to the bright light enhancement of the first image frame to obtain an enhanced image frame.
[0054] In this embodiment, when searching for whether there is a corresponding mapping group based on the third pixel value, and determining that there is no corresponding mapping group, for example: the third pixel value is a, and there is no mapping group f(a) in all pixel mapping groups corresponding to the first 3D-LUT model, a set interpolation method is used to insert a matching pixel value as the pixel value corresponding to the third pixel value after the corresponding first image frame completes image enhancement, thereby achieving enhancement of the first image frame.
[0055] It should be noted that in order to reduce computer memory storage pressure, the 3D-LUT adopts a 16×16×16 format instead of 256×256×253 or 180×180×180. At the same time, although common interpolation methods include nonlinear interpolation, adjacent-order interpolation, and trilinear interpolation, nonlinear interpolation consumes a long time, and adjacent-order interpolation may result in color blur and uneven color. Trilinear interpolation has a stable effect and can greatly shorten the image enhancement process. Therefore, in this embodiment, trilinear interpolation is adopted.
[0056] It should be understood that in this embodiment, the bright light image is also enhanced, but because the pixel values corresponding to the bright light image itself are relatively large, for example, the pixel value of pixel c in the bright light image is 230, and f(230)=250 exists in the mapping group, the pixel mapping group of the 3D-LUT basically does not change the original pixel value, or the change is very small. At the same time, the corresponding rules can only be obtained through deep learning training.
[0057] This embodiment also provides a target detection device based on 3D-LUT, which is used to implement the above embodiments and preferred implementations, and will not be repeated here. As used below, the terms "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, it is also possible and conceivable to implement it in hardware, or a combination of software and hardware.
[0058] Figure 3 is a structural block diagram of a target detection device based on 3D-LUT according to an embodiment of the present application, such as Figure 3 As shown, the device includes an acquisition module 31, an enhancement module 32 and an identification module 33, wherein:
[0059] An acquisition module 31 is configured to acquire a first image frame to be detected from the video frames of the collected video data, wherein the first image frame includes a target image captured under dark light conditions;
[0060] The enhancement module 32 is coupled to the acquisition module 31 and configured to perform brightness and contrast enhancement processing on the first image frame using a trained first 3D-LUT model and a preset interpolation method to obtain an enhanced image frame, wherein the first 3D-LUT model is a convolutional neural network (CNN) trained by performing image enhancement verification on paired dark light sample images and bright light sample images using a currently generated second 3D-LUT model, and guiding the currently generated second 3D-LUT model to perform iterative optimization, and the interpolation method includes one of the following: trilinear interpolation, nonlinear interpolation, and neighboring order interpolation;
[0061] The recognition module 33 is coupled to the enhancement module 32 and is used to perform target detection in the enhanced image frame based on the target detection model to obtain a detection result, wherein the target detection model is based on the YOLOv8 algorithm and is a neural network model trained according to the second image frame and the measured target corresponding to the second image frame. The second image frame is generated by enhancing the preset sample image frame through the trained first 3D-LUT model, and the sample image frame includes a dark light sample image frame and a bright light sample image frame mixed in a preset ratio.
[0062] Through the target detection device based on 3D-LUT of the embodiment of the present application, a first image frame to be detected is obtained in the video frame of the collected video data; the first image frame is subjected to brightness and contrast enhancement processing by using the trained first 3D-LUT model and the preset interpolation method to obtain an enhanced image frame, wherein the first 3D-LUT model is a convolutional neural network CNN trained by using the currently generated second 3D-LUT model to perform image enhancement verification on the paired dark light sample image and the bright light sample image, and guiding the currently generated second 3D-LUT model to perform iterative optimization; based on the target detection model, target detection is performed in the enhanced image frame to obtain a detection result, wherein the target detection model is based on the YOLOv8 algorithm, and is based on the second image frame and the measured target corresponding to the second image frame. The trained neural network model, the second image frame is generated by enhancing the preset sample image frame through the trained first 3D-LUT model, and the sample image frame includes a dark light sample image frame and a bright light sample image frame mixed in a preset ratio. The video image frame is enhanced through the paired trained 3D-LUT model in combination with the set interpolation method to achieve accurate pixel color mapping, effectively enhance image quality, and reduce noise while retaining important details. The training image of the target detection model is processed through the 3D-LUT model in combination with the set interpolation method to improve the target detection accuracy and reliability, and can effectively capture low-light image features, improve the adaptability to dynamically changing lighting conditions, and solve the problem that the target detection technology in the related technology cannot cope with weak light and complex environments and the target detection effect is poor.
[0063] In some embodiments, the device is further configured to obtain a current 3D-LUT, input a dark-light sample image into the current 3D-LUT, and obtain a bright-light enhanced sample image corresponding to the dark-light sample image, wherein each color channel of the current 3D-LUT includes a preset number of pixel mapping groups, and the current 3D-LUT includes one of the following output by the CNN network: an initial 3D-LTU, a 3D-LUT corresponding to a result of one iterative optimization; based on a bright-light sample image and a bright-light enhanced sample image paired with the dark-light sample image, a pixel offset between the dark-light sample image and the bright-light sample image is calculated, wherein the pixel offset is used to characterize a degree of correction to pixel values in the pixel mapping group; iteratively optimize the current 3D-LUT based on the pixel offset, and use the iteratively optimized 3D-LUT to process the corresponding dark-light sample image until the calculated pixel offset is less than a preset pixel offset threshold, thereby generating a first 3D-LUT model.
[0064] In some embodiments, the device is further configured to respectively obtain a first pixel set and a second pixel set corresponding to the bright light sample image and the bright light enhanced sample image; and calculate a mean square error of pixel values corresponding to the first pixel set and the second pixel set to obtain a pixel offset.
[0065] In some embodiments, the enhancement module 32 further includes:
[0066] The first processing unit is configured to traverse all third pixel values of the first image frame by using a first 3D-LUT model.
[0067] a first determination unit, coupled to the first processing unit, configured to determine whether, in all pixel mapping groups corresponding to the first 3D-LUT model, there exists a fourth pixel value mapped to the third pixel value, wherein the fourth pixel value is used to represent a pixel value corresponding to the first image frame after highlight enhancement is performed;
[0068] The first enhancement unit is coupled to the first judgment unit and is used to update the corresponding third pixel value to the corresponding fourth pixel value to obtain an enhanced image frame when the first judgment unit determines that there is a fourth pixel value mapped to the third pixel value.
[0069] In some embodiments, the enhancement module 32 is also used to generate a fifth pixel value corresponding to the third pixel value by using an interpolation method when the first judgment unit determines that there is no fourth pixel value mapped to the third pixel value, and use the fifth pixel value as the pixel value corresponding to the first image frame after bright light enhancement, to obtain an enhanced image frame.
[0070] In some embodiments, each color channel of the current 3D-LUT includes 16 pixel mapping groups, and / or the interpolation method is trilinear interpolation.
[0071] This embodiment also provides a target recognition system, including: an acquisition module, a transmission device and a server device; wherein the acquisition module is connected to the server device via the transmission device; wherein the acquisition module is connected to the server device via the transmission device; the acquisition module is used to acquire video data; the transmission device is used to transmit the video data to the server device; the server device is used to execute the steps of any of the above methods.
[0072] This embodiment further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0073] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0074] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0075] S1. Obtain a first image frame to be detected from video frames of collected video data, wherein the first image frame includes a target image captured under dark light conditions.
[0076] S2, using the trained first 3D-LUT model and a preset interpolation method, performing brightness and contrast enhancement processing on the first image frame to obtain an enhanced image frame, wherein the first 3D-LUT model uses the currently generated second 3D-LUT model to perform image enhancement verification on paired dark light sample images and bright light sample images, and guides the currently generated second 3D-LUT model to iteratively optimize the trained convolutional neural network CNN, and the interpolation method includes one of the following: trilinear interpolation, nonlinear interpolation, and adjacent order interpolation.
[0077] S3. Based on the target detection model, target detection is performed in the enhanced image frame to obtain a detection result, wherein the target detection model is based on the YOLOv8 algorithm and is a neural network model trained according to the second image frame and the measured target corresponding to the second image frame. The second image frame is generated by enhancing a preset sample image frame using the trained first 3D-LUT model. The sample image frame includes a dark light sample image frame and a bright light sample image frame mixed in a preset ratio.
[0078] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be repeated here.
[0079] In addition, in conjunction with the 3D-LUT-based object detection method in the above embodiments, embodiments of the present application may provide a storage medium for implementation. The storage medium stores a computer program; when the computer program is executed by a processor, it implements any of the 3D-LUT-based object detection methods in the above embodiments.
[0080] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0081] The above embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A target detection method based on three-dimensional color gamut conversion 3D-LUT, characterized in that: include: Acquire a first image frame to be detected from the video frames of the collected video data, wherein the first image frame includes a target image captured under dark light conditions; performing brightness and contrast enhancement processing on the first image frame using a trained first 3D-LUT model and a preset interpolation method to obtain an enhanced image frame, wherein the first 3D-LUT model is a convolutional neural network (CNN) trained by using the current 3D-LUT to perform image enhancement verification on paired dark-light sample images and bright-light sample images and guiding iterative optimization of the current 3D-LUT; Based on the target detection model, target detection is performed in the enhanced image frame to obtain a detection result, wherein the target detection model is a neural network model trained based on the YOLOv8 algorithm and according to the second image frame and the measured target corresponding to the second image frame, and the second image frame is generated by enhancing a preset sample image frame using the trained first 3D-LUT model, and the preset sample image frame includes a dark light sample image frame and a bright light sample image frame mixed in a preset ratio; wherein the training of the first 3D-LUT model includes: Obtaining a current 3D-LUT, inputting the dark-light sample image into the current 3D-LUT, and obtaining a bright-light enhanced sample image corresponding to the dark-light sample image, wherein each color channel of the current 3D-LUT includes a preset number of pixel mapping groups, and the current 3D-LUT includes one of the following output by the CNN network: an initial 3D-LUT, or a 3D-LUT corresponding to a result of one iterative optimization; Calculating a pixel offset between a bright light enhanced sample image corresponding to the dark light sample image and the bright light sample image; Iteratively optimizing the current 3D-LUT according to the pixel offset to generate the first 3D-LUT model; The first image frame is processed using the trained first 3D-LUT model and a preset interpolation method to obtain an enhanced image frame, including: Using the first 3D-LUT model, traverse all third pixel values of the first image frame; Determining whether there is a fourth pixel value mapped to the third pixel value in all pixel mapping groups corresponding to the first 3D-LUT model, wherein the fourth pixel value is used to represent a pixel value corresponding to the first image frame after highlight enhancement; When it is determined that there is a fourth pixel value mapped to the third pixel value, updating the corresponding third pixel value to the corresponding fourth pixel value to obtain the enhanced image frame; When it is determined that there is no fourth pixel value mapped to the third pixel value, an interpolation method is used to generate a fifth pixel value corresponding to the third pixel value, and the fifth pixel value is used as the pixel value corresponding to the first image frame after bright light enhancement to obtain the enhanced image frame.
2. The method according to claim 1, characterized in that Calculating a pixel offset between a bright light enhanced sample image corresponding to the dark light sample image and the bright light sample image includes: respectively acquiring a first pixel set and a second pixel set corresponding to the bright light sample image and the bright light enhanced sample image; The pixel offset is obtained by calculating a mean square error of pixel values corresponding to the first pixel set and the second pixel set.
3. The method according to claim 1, characterized in that Each color channel of the current 3D-LUT includes 16 pixel mapping groups, and / or the interpolation method is a trilinear interpolation method.
4. A 3D-LUT-based target detection device, used to execute the target detection method based on three-dimensional color gamut conversion 3D-LUT according to claim 1, characterized in that: include: An acquisition module, configured to acquire a first image frame to be detected from the video frames of the acquired video data, wherein the first image frame includes a target image captured under dark light conditions; an enhancement module, configured to perform brightness and contrast enhancement processing on the first image frame using a trained first 3D-LUT model and a preset interpolation method to obtain an enhanced image frame, wherein the first 3D-LUT model is a convolutional neural network (CNN) trained by using the current 3D-LUT to perform image enhancement verification on paired dark light sample images and bright light sample images, and guiding the current 3D-LUT to perform iterative optimization; An identification module is configured to perform target detection in the enhanced image frame based on a target detection model to obtain a detection result, wherein the target detection model is based on the YOLOv8 algorithm and is a neural network model trained based on a second image frame and a measured target corresponding to the second image frame, the second image frame is generated by enhancing a preset sample image frame using the trained first 3D-LUT model, the preset sample image frame including a dark light sample image frame and a bright light sample image frame mixed in a preset ratio.
5. A target recognition system, characterized in that: include: An acquisition module, a transmission device, and a server device; wherein the acquisition module is connected to the server device via the transmission device; The acquisition module is used to acquire video data; The transmission device is used to transmit the video data to the server device; The server device is used to execute the target detection method based on three-dimensional color gamut conversion 3D-LUT according to any one of claims 1 to 3.
6. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the steps of the object detection method based on three-dimensional color gamut conversion 3D-LUT according to any one of claims 1 to 3.
7. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target detection method based on three-dimensional color gamut conversion 3D-LUT according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer readable storage medium
CN113888437A
Target detection method, apparatus and device for continuous images, and storage medium
US20210319565A1