Key point locking method and device based on dark light enhancement and target detection

Through the combination of Darklighter enhanced model and yolov8 series models, the detection difficulties of infrared imaging technology under low illumination conditions are solved, efficient key point locking of infrared images is achieved, and the accuracy of target recognition and system adaptability are improved.

CN120387943APending Publication Date: 2025-07-29ARMY ENG UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510544382.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing infrared imaging technology is difficult to detect non-heating objects under low illumination conditions, with low signal-to-noise ratio and poor image effects, resulting in low target recognition accuracy.

Method used

The Darklighter enhancement model is used to enhance infrared images, and the target object segmentation and the yolov8-seg model are combined to perform key point locking. The image quality is optimized through multi-layer convolutional layers and loss functions to achieve accurate key points locking.

Benefits of technology

It significantly improves the accuracy and reliability of target recognition in dark light environments, enhances the adaptability and flexibility of the system, and can work stably in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387943A_ABST
    Figure CN120387943A_ABST
Patent Text Reader

Abstract

The invention discloses a key point locking method and device based on dark light enhancement and target detection, and belongs to the technical field of image detection, and the locking method comprises the steps: obtaining an infrared image of a dark light environment; performing enhancement processing on the infrared image through a pre-trained Darklight enhancement model based on a Retinex theory; inputting the infrared image after enhancement processing into a pre-trained yolov8-seg model to carry out target object segmentation; stroke processing is carried out on the target object segmentation result, key points are marked, and a training sample is generated; the yolov8-pose model is trained through the training sample, and a trained yolov8-pose model is obtained; and deploying the Darklight enhancement model, the yolov8-seg model and the yolov8-pose model at the same time, so as to lock the key points on the infrared image. The detection locking effect is accurate and efficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and particularly to a key point locking method and device based on low-light enhancement and target detection. Background Art

[0002] In the current complex and changeable night environment, especially under low illuminance conditions, accurately identifying targets such as vehicles is crucial for timely avoidance. This ability not only relates to the order of night public transportation but also directly affects people's life and property safety. To address this challenge, infrared enhanced image technology has become one of the more widely used solutions in this field.

[0003] However, despite the unique advantages of infrared imaging technology, there are still some significant problems and limitations in practical applications:

[0004] 1. Unable to detect non-heating objects: For obstacles that do not generate heat or have good heat dissipation performance, traditional infrared devices are difficult to effectively detect their presence.

[0005] 2. Signal-to-noise ratio problem: When the environmental light is extremely weak, even for high-health parts such as the human body, the collected signal may be overwhelmed by background noise, further reducing the recognition accuracy.

[0006] 3. Poor visual effect: Compared with the clear images under visible light, infrared images are usually darker and have a single color, posing a significant challenge to observers. Summary of the Invention

[0007] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a key point locking method and device based on low-light enhancement and target detection, so as to solve the technical problem that the existing infrared imaging technology has certain limitations, resulting in difficulties in subsequent image detection and low accuracy.

[0008] To achieve the above object, the present invention is implemented by the following technical solutions:

[0009] In the first aspect, the present invention provides a key point locking method based on low-light enhancement and target detection, including:

[0010] Obtain an infrared image of a low-light environment;

[0011] Perform enhancement processing on the infrared image through a pre-trained Darklighter enhancement model based on the Retinex theory;

[0012] Input the enhanced infrared image into a pre-trained yolov8-seg model for target object segmentation;

[0013] Perform stroke processing on the segmentation result of the target object and mark key points to generate training samples;

[0014] Train the yolov8-pose model with the training samples to obtain a trained yolov8-pose model;

[0015] Deploy the Darklighter enhancement model, the yolov8-seg model, and the yolov8-pose model simultaneously to achieve key point locking on infrared images.

[0016] Optionally, the Darklighter enhancement model gradually strips the light and noise in the input image through multiple iterations, and its expression is:

[0017]

[0018] In the formula, is the observed image obtained by the decomposition in the th round of iteration, is the noise map and light map in the th round of iteration;

[0019] When taking 0, is the input image, When taking , is the output image, is the maximum number of iterations; the noise map and the light map are estimated and obtained by the trained ME-Net network according to the input image .

[0020] Optionally, the ME-Net network is sequentially connected with multiple convolutional layers, and the number of output channels of the multiple convolutional layers gradually increases. The output tensor dimension of the last convolutional layer is , being the height and width of the input image;

[0021] Take the feature maps of the first channels as the light map , ;

[0022] Take the feature maps of the last channels as the noise map , .

[0023] Optionally, the loss function for training the ME-Net network is:

[0024]

[0025] In the formula, are all weight parameters, is the color intensity loss, is the central attention brightness loss, is the illumination regularization loss, is the semantic fidelity loss, is the noise estimation loss;

[0026] Color intensity loss is:

[0027]

[0028] In the formula, is the image of the th channel in the RGB channels of the output image of the Darklighter enhancement model;

[0029] Central attention brightness loss is:

[0030]

[0031] In the formula, is the weight matrix, is the average brightness matrix, is the target brightness, is the all-ones matrix; the output image of the Darklighter enhancement model is evenly divided into P non-overlapping blocks, the average brightness y of each block is calculated and the average brightness matrix is generated, the weight of each block is calculated and the weight matrix is generated, is the position of the block;

[0032] Illumination regularization loss is:

[0033]

[0034] In the formula, is the first-order gradient operator;

[0035] Noise estimation loss is:

[0036]

[0037] Semantic fidelity loss is:

[0038]

[0039] In the formula, It is the intermediate layer feature extraction operation of the VGG-16 network.

[0040] In a second aspect, the present invention provides a key point locking device based on low-light enhancement and target detection, including:

[0041] An image acquisition module, configured to acquire an infrared image of a low-light environment;

[0042] An image enhancement module, configured to perform enhancement processing on the infrared image through a pre-trained Darklighter enhancement model based on the Retinex theory;

[0043] An image segmentation module, configured to input the enhanced infrared image into a pre-trained yolov8-seg model for target object segmentation;

[0044] A stroke annotation module, configured to perform stroke processing on the target object segmentation result and annotate key points to generate training samples;

[0045] A model training module, configured to train a yolov8-pose model through the training samples to obtain a trained yolov8-pose model;

[0046] A model deployment module, configured to deploy the Darklighter enhancement model, the yolov8-seg model, and the yolov8-pose model simultaneously to achieve key point locking on the infrared image.

[0047] In a third aspect, the present invention provides an electronic device, including a processor and a storage medium;

[0048] The storage medium is used to store instructions;

[0049] The processor is used to operate according to the instructions to execute the steps of the above method.

[0050] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.

[0051] In a fifth aspect, the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.

[0052] Compared with the prior art, the beneficial effects achieved by the present invention:

[0053] The present invention provides a key point locking method and device based on low-light enhancement and target detection. In the core link of target recognition in low-light environments, the cutting-edge YOLO target detection model is organically integrated with the Darklighter enhancement model. The Darklighter enhancement model uses a deep neural network to analyze and optimize the brightness, contrast, and detail information in the image, thereby significantly improving the image quality of infrared images in low-light environments. The YOLO target detection model includes the yolov8-seg model and the yolov8-pose model. The yolov8-seg model can achieve target object segmentation, and the yolov8-pose model can achieve key point positioning. Through the application of model fusion, not only the accuracy and reliability of target recognition are improved, but also the adaptability and flexibility of the system are enhanced, enabling it to work stably in various complex environments. Brief Description of the Drawings

[0054] Figure 1 is a schematic flowchart of the key point locking method based on low-light enhancement and target detection provided by an embodiment of the present invention;

[0055] Figure 2 is a schematic structural diagram of the key point locking device based on low-light enhancement and target detection provided by an embodiment of the present invention. Detailed Embodiments

[0056] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.

[0057] Embodiment 1:

[0058] First, an embodiment of the present invention provides a key point locking method based on low-light enhancement and target detection, including the following steps:

[0059] Step S1, obtain an infrared image of a low-light environment.

[0060] Various sources of videos or images can be aggregated as needed, covering video clips and image files with different resolutions, formats, and shooting environments. They carry rich and diverse information and can provide sufficient data support for subsequent model training.

[0061] Step S2, perform enhancement processing on the infrared image through a pre-trained Darklighter enhancement model based on the Retinex theory.

[0062] The Retinex theory is an image enhancement model that simulates the color constancy of the human visual system. Its core idea is to achieve dynamic range compression, detail enhancement, and color restoration by decomposing the illumination and reflection components of the image.

[0063] In this embodiment, a Darklighter enhancement model is constructed based on the Retinex theory. The Darklighter enhancement model gradually strips the illumination and noise in the input image through multiple iterations, and its expression is:

[0064]

[0065] In the formula, is the observed image obtained by the -th round of iterative decomposition, is the noise map and illumination map of the -th round of iteration;

[0066] When taking 0, is the input image, When taking , is the output image, is the maximum number of iterations; the noise map and the illumination map are estimated and obtained by the trained ME-Net network according to the input image .

[0067] Among them, the ME-Net network is successively connected with multiple convolutional layers. The number of output channels of the multiple convolutional layers gradually increases, and the output tensor dimension of the last convolutional layer is , being the height and width of the input image;

[0068] Taking the feature maps of the first channels as the illumination map , respectively;

[0069] Taking the feature maps of the last channels as the noise map , respectively.

[0070] The ME-Net network adopts unsupervised training, and the loss function for training is:

[0071]

[0072] In the formula, are all weight parameters, is the color intensity loss, is the central attention brightness loss, is the illumination regularization loss, is the semantic fidelity loss, is the noise estimation loss;

[0073] Color intensity loss is:

[0074]

[0075] In the formula, is the th channel image in the RGB channels of the output image of the Darklighter enhancement model;

[0076] Central attention brightness loss is:

[0077]

[0078] In the formula, is the weight matrix, is the average brightness matrix, is the target brightness, is the all-ones matrix; The output image of the Darklighter enhancement model is evenly divided into P non-overlapping blocks, the average brightness y of each block is calculated and the average brightness matrix is generated, the weight of each block is calculated and the weight matrix is generated, is the position of the block; For blocks distributed in a 16*16 pattern, (1,1) represents the block in the first row and first column;

[0079] Illumination regularization loss is:

[0080]

[0081] In the formula, is the first-order gradient operator;

[0082] Noise estimation loss is:

[0083]

[0084] Semantic fidelity loss is:

[0085]

[0086] In the formula, is the intermediate layer feature extraction operation of the VGG-16 network.

[0087] By combining multiple loss functions, it is ensured that the enhanced image is both suitable for visual perception and retains semantic information.

[0088] Step S3: Input the enhanced infrared image into the pre-trained yolov8-seg model for object segmentation.

[0089] The YOLOv8-seg model is the first model in the YOLO series to integrate object detection and instance segmentation functions, achieving efficient and accurate instance segmentation capabilities through multi-dimensional optimization. In the embodiments of the present invention, various objects in videos or images are deeply segmented and accurately detected through the yolov8-seg model. Whether it is pedestrians and vehicles in street scenes or potential suspicious objects on night roads, they can be accurately identified, clearly outlining the regions of objects of interest and providing key object pointers for subsequent analysis and decision-making.

[0090] Step S4: Perform edge tracing on the object segmentation results and label key points to generate training samples.

[0091] Perform delicate and clear edge tracing on the object segmentation results, making the objects more prominent in the picture. Key points will also be accurately labeled according to professional knowledge and preset rules, such as the key vital points of the human body, providing key guidance for the subsequent training of the yolov8-pose model.

[0092] Step S5: Train the yolov8-pose model with the training samples to obtain a trained yolov8-pose model.

[0093] The YOLOv8-Pose model is an efficient human pose estimation model developed based on the YOLOv8 framework, combining object detection and key point localization capabilities, taking into account real-time performance and accuracy. Through training, the yolov8-pose model is continuously adjusted and optimized to ensure that the system can always accurately identify and process various objects in complex and changing real-world environments and cope with various challenges.

[0094] Step S6: Deploy the Darklighter enhancement model, yolov8-seg model, and yolov8-pose model simultaneously to achieve key point locking on infrared images.

[0095] Embodiment 2:

[0096] As Figure 2 shown, the embodiments of the present invention provide a key point locking device based on low-light enhancement and object detection, including:

[0097] An image acquisition module configured to acquire infrared images of low-light environments;

[0098] An image enhancement module configured to enhance the infrared image through the pre-trained Darklighter enhancement model based on the Retinex theory;

[0099] An image segmentation module, configured to input the enhanced infrared image into a pre-trained yolov8-seg model for object segmentation;

[0100] A stroke annotation module, configured to perform stroke processing on the object segmentation result and annotate key points to generate training samples;

[0101] A model training module, configured to train the yolov8-pose model with the training samples to obtain a trained yolov8-pose model;

[0102] A model deployment module, configured to deploy the Darklighter enhancement model, the yolov8-seg model, and the yolov8-pose model simultaneously to achieve key point locking on the infrared image.

[0103] Example 3:

[0104] Based on the key point locking method provided in Example 1, an embodiment of the present invention provides an electronic device, including a processor and a storage medium;

[0105] The storage medium is used to store instructions;

[0106] The processor is used to operate according to the instructions to execute the steps of the above method.

[0107] Example 4:

[0108] Based on the key point locking method provided in Example 1, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.

[0109] Example 5:

[0110] Based on the key point locking method provided in Example 1, an embodiment of the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.

[0111] A computer program product, such as an intelligent helmet, serves as an information presentation terminal directly facing the user. The helmet is closely connected to the RK3588 chip through a high-speed transmission line. The Darklighter enhancement model, yolov8-seg model, and yolov8-pose model obtained in Example 1 are simultaneously deployed and burned onto the RK3588 chip. Meanwhile, the high-resolution and high-brightness display screen equipped inside the intelligent helmet can present the results processed by the chip to the user in the most intuitive visual form. Whether in complex environments such as high-speed movement or strong light irradiation, it can ensure that the user clearly and immediately obtains key information, providing strong support for their action decisions, and is a capable information assistant for the user in front-line scenarios.

[0112] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0113] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0114] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps of the functions specified in one block or a plurality of blocks.

[0116] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A key point locking method based on low-light enhancement and target detection, characterized in that Including: Obtain an infrared image of a low-light environment; Enhance the infrared image through a pre-trained Darklighter enhancement model based on the Retinex theory; Input the enhanced infrared image into a pre-trained yolov8-seg model for object segmentation; Perform stroke processing and key point annotation on the object segmentation result to generate a training sample; Train the yolov8-pose model with the training sample to obtain a trained yolov8-pose model; Deploy the Darklighter enhancement model, the yolov8-seg model, and the yolov8-pose model simultaneously to achieve key point locking on the infrared image.

2. The key point locking method based on low-light enhancement and target detection according to claim 1, characterized in that, The Darklighter enhancement model gradually strips the light and noise in the input image through multiple iterations, and its expression is: In the formula, is the observed image obtained by the -th round of iterative decomposition, is the noise map and illumination map of the -th round of iteration; When taking 0, it is the input image, When taking , it is the output image, is the maximum number of iterations; the noise map and the illumination map are estimated and obtained by the trained ME-Net network according to the input image .

3. The key point locking method based on low-light enhancement and target detection according to claim 2, characterized in that The ME-Net network is successively connected with multiple convolutional layers, the number of output channels of the multiple convolutional layers gradually increases, and the output tensor dimension of the last convolutional layer is , being the height and width of the input image; Take the first feature maps of the channels as the illumination maps respectively , ; After taking The feature maps of the channels are used as noise maps respectively .

4. The key point locking method based on low-light enhancement and target detection according to claim 3, characterized in that, The loss function for training the ME-Net network is: In the formula, are all weight parameters, is the color intensity loss, is the central attention luminance loss, is the illumination regularization loss, is the semantic fidelity loss, is the noise estimation loss; Color intensity loss is as follows: In the formula, is the th channel image in the RGB channels of the output image of the Darklighter enhancement model; Central concern: luminance loss is In the formula, is the weight matrix, is the average luminance matrix, is the target luminance, is the all-ones matrix; the output image of the Darklighter enhancement model is evenly divided into P non-overlapping blocks, the average luminance y of each block is calculated and the average luminance matrix is generated, the weight of each block is calculated and the weight matrix is generated, is the position of the block; Illumination regularization loss is as follows: In the formula, is a gradient operator; Noise estimation loss is as follows: Semantic fidelity loss is as follows: In the formula, is the intermediate layer feature extraction operation of the VGG-16 network.

5. A key point locking device based on low-light enhancement and target detection, characterized in that, Including: An image acquisition module configured to obtain an infrared image of a low-light environment; An image enhancement module configured to enhance the infrared image through a pre-trained Darklighter enhancement model based on the Retinex theory; An image segmentation module configured to input the enhanced infrared image into a pre-trained yolov8-seg model for object segmentation; A stroke annotation module configured to perform stroke processing and key point annotation on the object segmentation result to generate a training sample; A model training module configured to train the yolov8-pose model with the training sample to obtain a trained yolov8-pose model; A model deployment module configured to deploy the Darklighter enhancement model, the yolov8-seg model, and the yolov8-pose model simultaneously to achieve key point locking on the infrared image.

6. An electronic device, characterized in that, Including a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1-4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1-4.

8. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, it implements the steps of the method according to any one of claims 1-4.