Personnel identification method and system for low-light environment

By using light change detection models and deep learning technologies for image enhancement and personnel recognition in low-light environments, the technical challenges of personnel recognition in low-light environments are solved, and the monitoring effect with high accuracy and robustness is achieved.

CN120048001APending Publication Date: 2025-05-27夏浩洎
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510132653.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In low-light environments, personnel identification and tracking face technical challenges such as declining image quality, increasing noise, changes in light and personnel occlusion, making it difficult for traditional algorithms to accurately separate and identify individuals.

Method used

A person recognition method for low-light environments is proposed, including obtaining and pre-processing of surveillance video image sequences, constructing a light change detection model to detect light change parameters, dynamically adjusting the parameters of the image enhancement algorithm, and performing key points detection and individual separation and recognition of images based on human posture algorithm and attention mechanism.

Benefits of technology

Through adaptive histogram equalization and deep learning technology, the accuracy and robustness of personnel monitoring in low-light environments are significantly improved, and intelligent monitoring in complex scenarios can be effectively handled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048001A_ABST
    Figure CN120048001A_ABST
Patent Text Reader

Abstract

The invention discloses a personnel identification method and system for a low-light environment. The method comprises the following steps: obtaining a monitoring video image sequence in the low-light environment and carrying out the preprocessing of the monitoring video image sequence; constructing an illumination change detection model, and detecting an illumination change parameter of the preprocessed monitoring video image based on the illumination change detection model; adjusting parameters of an image enhancement algorithm based on the illumination change parameters, and processing the preprocessed monitoring video image based on the adjusted image enhancement algorithm to obtain an enhanced image; key point detection and skeleton extraction are carried out on the enhanced image based on a human body posture algorithm, an attention mechanism and context information are introduced, and individuals in the enhanced image are separated and recognized. According to the method, the accuracy and robustness of personnel monitoring in a low-light environment are remarkably improved, and an effective solution is provided for intelligent monitoring in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and particularly relates to a method and system for personnel recognition in low-light environments. Background Art

[0002] Performing personnel recognition and tracking in low-light environments faces many technical challenges. First, insufficient lighting leads to a decline in image quality and an increase in noise, making it difficult to capture the outlines and detailed features of personnel, which brings difficulties to feature extraction and analysis. Second, the dynamic change of lighting will cause fluctuations in image brightness and contrast, increasing the uncertainty of personnel matching during the tracking process. Moreover, when people gather, there are often complex occlusion and interleaving phenomena, and the mutual interference between people makes it difficult for traditional object detection and segmentation algorithms to accurately separate individuals.

[0003] The above problems are intertwined, bringing huge challenges to low-light personnel recognition technology. Therefore, there is an urgent need to propose a method and system for personnel recognition in low-light environments. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a method and system for personnel recognition in low-light environments to solve the problems existing in the above prior art.

[0005] To achieve the above object, the present invention provides a method for personnel recognition in low-light environments, including the following steps:

[0006] Obtain a sequence of surveillance video images in a low-light environment and perform preprocessing;

[0007] Construct a lighting change detection model, and detect the lighting change parameters of the preprocessed surveillance video images based on the lighting change detection model;

[0008] Adjust the parameters of the image enhancement algorithm based on the lighting change parameters, and process the preprocessed surveillance video images based on the adjusted image enhancement algorithm to obtain enhanced images;

[0009] Perform key point detection and skeleton extraction on the enhanced images based on the human pose algorithm, and introduce an attention mechanism and context information to separate and recognize individuals in the enhanced images.

[0010] Optionally, the process of obtaining a sequence of surveillance video images in a low-light environment and performing preprocessing to obtain an enhanced sequence of surveillance video images includes:

[0011] Divide each frame image in the monitored video image sequence into blocks, and perform histogram equalization processing and smoothing filtering processing on each block; then extract the contour information in the image through an edge detection algorithm, and extract the personnel area according to the gradient direction and magnitude of the contour pixel points; perform binary processing and detail enhancement processing on the extracted personnel area, and fuse the processed personnel area with the original image to obtain the preprocessed monitored video image.

[0012] Optionally, the process of constructing a light change detection model and detecting the light change parameters of the preprocessed monitored video image based on the light change detection model includes:

[0013] Extract the brightness feature and contrast feature of the preprocessed monitored video image, and based on the brightness feature and contrast feature, judge whether the current image has a light change according to the pre-constructed light change detection model; if a light change is detected, extract the light change area in the image and estimate the light change parameters.

[0014] Optionally, the process of adjusting the parameters of the image enhancement algorithm based on the light change parameters and processing the preprocessed monitored video image based on the adjusted image enhancement algorithm to obtain the enhanced image includes:

[0015] Based on the light change parameters, dynamically adjust the parameters of the image enhancement algorithm, enhance the brightness, local contrast and texture of the image respectively based on the adjusted parameters of the image enhancement algorithm, and then fuse the images after brightness adjustment, local contrast enhancement and texture enhancement to obtain the final enhanced image;

[0016] Among them, the parameters of the image enhancement algorithm include a brightness adjustment coefficient and a contrast adjustment coefficient.

[0017] Optionally, the process of separating and identifying individuals in the enhanced image includes the processing of images without personnel occlusion and the processing of images with personnel occlusion;

[0018] Among them, the process of processing images without personnel occlusion includes: performing key point detection on the human body area in the enhanced image based on the human body pose algorithm to obtain the key point coordinates representing the positions of human joints, connecting the key point coordinates to form a human body skeleton, and extracting the human body skeleton feature vector to represent the human body pose information; then extracting the human body pose features for each frame image in the monitored video image sequence, constructing the personnel movement trajectory in combination with time information, dividing the personnel movement trajectory into different individuals through the trajectory clustering algorithm to realize the separation of personnel; extracting the appearance features of each individual, and comprehensively using the pose information, movement trajectory and appearance features of the individual, and adopting the method of multi-feature fusion to perform identity recognition on the separated individuals to obtain the identity information of each person and complete the recognition.

[0019] Optionally, the process of processing an image with people occluding each other includes:

[0020] Use a convolutional neural network to extract features from an image with people occluding each other to obtain image features, and then weight the image features through an attention mechanism; use a recurrent neural network to model the image with people occluding each other to capture the context information between people; fuse the attention-weighted image features and the context information learned by the recurrent neural network to form a fused feature representation. Based on the fused feature representation, use a region proposal network to generate candidate target regions, and determine whether the candidate target regions contain the target person. For the regions where the target person is detected, further use a fully convolutional network to perform pixel-level segmentation on it, and post-process the detected target person according to the segmentation result to obtain the final person separation result.

[0021] The present invention also proposes a personnel recognition system for low-light environments, which is used to implement the personnel recognition method for low-light environments, including: an image acquisition module, an image enhancement module, and an individual recognition module;

[0022] The image sequence acquisition module is used to acquire and preprocess the surveillance video image sequence in a low-light environment;

[0023] The image enhancement module is used to build a light change detection model, detect the light change parameters of the preprocessed surveillance video image based on the light change detection model; and is used to adjust the parameters of the image enhancement algorithm based on the light change parameters, and process the preprocessed surveillance video image based on the adjusted image enhancement algorithm to obtain an enhanced image;

[0024] The individual recognition module is used to perform key point detection and skeleton extraction on the enhanced image based on the human pose algorithm, and introduce an attention mechanism and context information to separate and recognize individuals in the enhanced image.

[0025] The present invention also provides an electronic device, including: a memory and a processor; the memory is used to store a program; the processor is used to execute the program to implement each step of the personnel recognition method for low-light environments.

[0026] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each step of the personnel recognition method for low-light environments is implemented.

[0027] Compared with the prior art, the present invention has the following advantages and technical effects:

[0028] In view of the problems of the decline in the quality of surveillance video images, the increase in noise, the occlusion of people gathering, and background interference under low-light conditions, the present invention adopts adaptive histogram equalization to enhance the image contrast, constructs a light change detection model to dynamically adjust the image enhancement parameters, and introduces deep learning human pose estimation and attention mechanism to improve the individual recognition ability. This method significantly improves the accuracy and robustness of personnel surveillance in low-light environments, and provides an effective solution for intelligent surveillance in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings that form a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0030] Figure 1 is the overall flowchart of the method for identifying people in a low-light environment according to an embodiment of the present invention;

[0031] Figure 2 is the schematic diagram of the specific process of separating and identifying individuals according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0033] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0034] Embodiment 1

[0035] As Figure 1 shown, this embodiment provides a method for identifying people in a low-light environment, including the following steps:

[0036] Obtain the surveillance video image sequence in the low-light environment and perform preprocessing;

[0037] Construct a light change detection model, and detect the light change parameters of the preprocessed surveillance video image based on the light change detection model;

[0038] Adjust the parameters of the image enhancement algorithm based on the light change parameters, and process the preprocessed surveillance video image based on the adjusted image enhancement algorithm to obtain an enhanced image;

[0039] Perform key point detection and skeleton extraction on the enhanced image based on the human pose algorithm, and introduce the attention mechanism and context information to separate and identify individuals in the enhanced image.

[0040] As a specific embodiment, to obtain the monitoring video image sequence in the low-light environment, for the problems of image quality degradation and increased noise, the process of preprocessing the image using the adaptive histogram equalization algorithm to enhance the image contrast and highlight the personnel contour and detail features includes:

[0041] Obtain the monitoring video image sequence in the low-light environment. For the problems of image quality degradation and increased noise, the following steps are taken for processing: By dividing the image into blocks, calculate the histogram distribution of each block, adaptively determine the equalization parameters of each block according to the histogram distribution, perform histogram equalization processing on the blocks to enhance the contrast of local regions. Perform smoothing filtering on the histogram-equalized image to suppress noise and make the image smoother. Extract the contour information in the image through the edge detection algorithm, and judge whether the contour is a personnel contour according to the gradient direction and magnitude of the contour pixel points, and extract the personnel area. Perform binarization processing on the extracted personnel area, and adaptively determine the binarization threshold according to the gray value distribution of the pixel points inside the personnel contour to highlight the detail features of the personnel area. Perform further detail enhancement on the binarized personnel area through morphological processing to extract the detail features of the personnel, such as clothing texture, facial features, etc. Merge the processed personnel area with the original image to obtain the enhanced monitoring video image, with improved image quality and clearer personnel contour and detail features. By performing the above processing on consecutive multiple frames of monitoring video images, a video image sequence with improved quality is obtained, providing high-quality image data for subsequent video analysis applications.

[0042] As an implementable approach, to enhance the quality of surveillance video images in low-light environments, the method of adaptive block histogram equalization can be adopted. First, the image is divided into blocks of 16x16 pixels, and the grayscale histogram distribution of each block is calculated. According to the uniformity of the histogram distribution, the equalization parameters of each block are adaptively determined. For example, the threshold for contrast stretching can be set to twice the gray level with the largest number of pixels in the histogram. Then, histogram equalization is performed on each block to enhance the contrast of the local area. To suppress the noise in the equalized image, a 5x5 Gaussian filter can be used to smooth the image. Next, the Canny operator is used to detect the edges of the smoothed image and extract the contour information in the image. According to the gradient direction and magnitude of the contour pixel points, it can be determined whether the contour is a human contour. For example, a human contour usually has a continuous gradient direction and a large gradient magnitude. The extracted human area is binarized, and the binarization threshold is adaptively determined according to the gray value distribution of the pixel points inside the human contour. For example, the threshold can be set to the mean gray value within the human area. Morphological processing, such as erosion and dilation, is performed on the binarized human area to extract the detailed features of the person, such as clothing texture and facial features. Finally, the processed human area is weighted and fused with the original image to obtain the enhanced surveillance video image. By performing the above processing on consecutive frames of surveillance video images, a sequence of video images with improved quality can be obtained, providing high-quality image data for subsequent video analysis applications.

[0043] As a specific embodiment, the process of ensuring the stability of image quality by constructing a light change detection model, estimating the light change parameters in real time, and dynamically adjusting the parameters of the image enhancement algorithm according to the parameters for the problem of image brightness and contrast fluctuations caused by dynamic light changes includes:

[0044] Obtain the image data to be processed and extract the brightness and contrast features of the image; according to the pre-constructed light change detection model, determine whether the current image has undergone a light change; if a light change is detected, extract the light change area in the image and estimate the light change parameters; according to the estimated light change parameters, dynamically adjust the parameters of the image enhancement algorithm, including the brightness adjustment coefficient and the contrast adjustment coefficient; adopt an adaptive threshold segmentation algorithm to extract the target area in the image and perform local contrast enhancement; extract the texture features of the image through the histogram of oriented gradients algorithm and perform texture enhancement processing; fuse the images after brightness adjustment, contrast enhancement, and texture enhancement to obtain the final enhanced image with unchanged light.

[0045] As an implementable approach, it is determined whether there is a change in illumination in the current image according to a pre-constructed illumination change detection model. This model uses the support vector machine algorithm. By extracting features such as the color histogram and gradient histogram of the image, the illumination change detection model is trained. By inputting the features of the current image into the model for prediction, it is determined that there is an illumination change in the current image. Then, the illumination change region in the image is extracted. The region growing algorithm is adopted. Taking the pixels with obvious illumination change as the seed points, the region is continuously expanded until the gray value difference of the pixels in the region is less than the threshold, and the illumination change region is obtained. According to the gray value distribution of the pixels in the illumination change region, the illumination change parameters are estimated to obtain the brightness adjustment coefficient and the contrast adjustment coefficient. According to the estimated illumination change parameters, the parameters of the image enhancement algorithm are dynamically adjusted to adjust the brightness and contrast of the image, so that the brightness and contrast distribution of the image is more balanced. Then, the adaptive threshold segmentation algorithm is adopted. According to the gray value distribution of the local region of the image, the segmentation threshold is adaptively determined, the target region in the image is extracted, and the local contrast of the target region is enhanced. Finally, through the histogram of oriented gradients algorithm, the texture features of the image are extracted, the gradient direction distribution of the pixels in the local region of the image is calculated to obtain the histogram of oriented gradients, and the image is processed for texture enhancement according to the histogram distribution to make the texture details clearer. The images after brightness adjustment, contrast enhancement, and texture enhancement are weighted and fused to obtain the final enhanced image with invariant illumination. The visual quality of the image is significantly improved, the target region is clearer, and the texture details are richer.

[0046] Further, the process of adopting the adaptive threshold segmentation algorithm to extract the target region in the image and perform local contrast enhancement includes:

[0047] According to the acquired image, an adaptive threshold segmentation algorithm is adopted. By dynamically adjusting the threshold, the image is segmented into a target region and a background region, and the image of the segmented target region is obtained. For the image of the segmented target region, through the local histogram equalization algorithm, the pixel gray value distribution within the target region is adjusted to enhance the local contrast, and the image of the target region with enhanced contrast is obtained. According to the image of the target region with enhanced contrast, the edge contour features of this region are extracted. Through the contour feature matching algorithm, the extracted target contour is matched with the preset target template contour. If the matching degree exceeds the preset threshold, it is determined that this target region is the target to be extracted. For the extracted target region, through morphological processing, the edge contour of the target region is smoothed and corrected to eliminate noise points and breakpoints, and the corrected target edge contour is obtained. According to the corrected target edge contour, geometric feature parameters such as the size, position, and direction of this target region are extracted, and the extracted geometric feature parameters are used as the basis for subsequent target tracking and recognition. The convolutional neural network algorithm is adopted to extract features and classify and recognize the image of the extracted target region. Through multi-layer convolution and pooling operations, deep features such as the texture and shape of the target region are extracted, and the extracted features are sent to the fully connected layer for classification decision to obtain the class label to which this target region belongs. The geometric feature parameters and class label of the target region are fused to construct a comprehensive feature vector of the target, and this feature vector is stored in the feature database for subsequent target retrieval and comparison.

[0048] Further, the process of extracting the texture features of the image and performing texture enhancement processing through the histogram of oriented gradients algorithm includes:

[0049] According to the acquired image to be processed, the histogram of oriented gradients algorithm is adopted to extract the texture features of the image. For the extracted texture features, through feature analysis, the statistical histogram of the texture features is determined. According to the statistical histogram of the texture features, the distribution of the texture features is judged. If the texture features are unevenly distributed, texture enhancement processing is performed. When performing texture enhancement processing, a texture synthesis algorithm is adopted to generate a texture pattern similar to the original texture features. The generated texture pattern is superimposed on the original image to obtain the image with enhanced texture. Through the image quality assessment algorithm, the quality of the image with enhanced texture is evaluated to judge the effect of texture enhancement. If the texture enhancement effect is not ideal, return to the texture synthesis step to regenerate the texture pattern until the texture enhancement effect meets the preset threshold.

[0050] The process of separating and recognizing individuals in the enhanced image includes the processing of images without human mutual occlusion and the processing of images with human mutual occlusion;

[0051] Among them, the process of processing images without human mutual occlusion includes:

[0052] Such asFigure 2 As shown, a video image sequence containing a personnel gathering scene is obtained and each frame of the image is processed. A pre-trained human pose estimation model is used to detect key points in the human region of the image through a convolutional neural network, and the key point coordinates representing the positions of human joints are obtained. The human skeleton is formed by connecting the key point coordinates, and the skeleton feature vector is extracted to represent the human pose information. The human pose features are extracted from each frame of the video sequence, and the personnel movement trajectory is constructed by combining the time information. The movement trajectory is analyzed, and the trajectory is divided into different individuals through a trajectory clustering algorithm to achieve the separation of personnel. The appearance features of each individual are extracted, including clothing color, texture, etc., to form an appearance feature vector. By integrating the pose information, movement trajectory and appearance features of the individual, a multi-feature fusion method is used to train an individual recognition model to identify the identity of the separated individual and obtain the identity information of each person.

[0053] As an implementable method, first, a video image sequence containing a personnel gathering scene is obtained and each frame of the image is processed. The pre-trained human pose estimation model OpenPose is used to detect key points in the human region of the image through a convolutional neural network, and 18 key point coordinates representing the positions of human joints are obtained. The human skeleton is formed by connecting the key point coordinates, and a 36-dimensional skeleton feature vector is extracted to represent the human pose information. The human pose features are extracted from each frame of the video sequence, and the personnel movement trajectory is constructed by combining the time information with a time interval of 5 seconds for the trajectory points. The movement trajectory is analyzed, and the trajectory is divided into different individuals through the DBSCAN trajectory clustering algorithm with the clustering parameter eps set to 5 meters and minPts set to 10 to achieve the separation of personnel. The appearance features of each individual are extracted, including clothing color histogram, SIFT texture features, etc., to form a 128-dimensional appearance feature vector. By integrating the pose information, movement trajectory and appearance features of the individual, a multi-feature fusion method is used to train an individual recognition model using the support vector machine SVM to identify the identity of the separated individual and obtain the identity information of each person. The performance of the model is evaluated on the test set, and the accuracy rate reaches 95%, demonstrating the effectiveness of this method for individual recognition in a personnel gathering scene.

[0054] The process of processing images with personnel occlusion includes:

[0055] Regarding the problem of difficult target detection and segmentation caused by personnel interference, an attention mechanism and context information are introduced. By learning the interaction patterns between personnel and prior knowledge of the scene, the discrimination ability of the algorithm for individuals is enhanced, and the accuracy of detection and segmentation is improved.

[0056] According to the input image data containing mutual occlusion of people, a convolutional neural network is used to extract features from the image to obtain a deep feature representation of the image. The extracted image features are weighted through an attention mechanism to highlight the significant features of individual people, weaken the mutual interference between people, and enhance the discrimination ability of the algorithm for individuals. A recurrent neural network is used to model the image feature sequence to capture the interaction patterns and context information between people and learn the prior knowledge of the scene. The attention-weighted image features and the context information learned by the recurrent neural network are fused to form the final feature representation. Based on the fused feature representation, a region proposal network is used to generate candidate target regions, and a classifier is used to determine whether the candidate regions contain the target people. For the detected target people regions, a fully convolutional network is further used to perform pixel-level segmentation on them to obtain fine-grained people contours and semantic masks. Post-processing is performed on the detected target people according to the segmentation results to eliminate overlapping and redundant detection boxes, and the final people detection and segmentation results are obtained.

[0057] As an implementable approach, when processing image data containing people occlusion, a pre-trained convolutional neural network model, such as VGG-16 or ResNet-50, is first used to extract features from the input image. Through the stacking of convolutional layers and pooling layers, deep feature representations of the image are extracted, and the feature dimension can reach 512 dimensions or 1024 dimensions. Then, an attention mechanism is introduced to weight the extracted image features by learning attention weights. Specifically, the self-attention mechanism is used to generate an attention weight matrix by calculating the similarity between features, and the original features are weighted and summed to obtain an enhanced feature representation. The attention mechanism can highlight the significant features of individual people and weaken the mutual interference between people. Next, a long short-term memory network (LSTM) is used to model the image feature sequence to capture the interaction patterns and context information between people. The image features are input into the LSTM in chronological order to learn the prior knowledge of the scene, and the hidden state dimension of the LSTM can be set to 256 dimensions. The attention-weighted image features and the context information learned by the LSTM are concatenated to form the final feature representation, and the dimension can reach 768 dimensions. Based on the fused feature representation, a region proposal network (RPN) is used to generate candidate target regions. The RPN generates a large number of anchor boxes on the feature map by sliding a window and uses a binary classifier to determine whether the anchor box contains the target person, generating high-quality candidate regions. For the selected candidate regions, a fully convolutional network (FCN) is further used for pixel-level semantic segmentation. The FCN obtains fine-grained person contours and semantic masks through per-pixel classification, and the number of pixel categories in the segmentation result can be set to 2, representing people and the background. Finally, post-processing is performed on the detected target people according to the segmentation result. Through the non-maximum suppression (NMS) algorithm, with an intersection over union (IoU) threshold of 5, overlapping and redundant detection boxes are eliminated to obtain the final person detection and segmentation results. The entire algorithm process realizes the accurate detection and segmentation of people in complex scenes through the organic combination of deep learning models, providing a good foundation for subsequent person tracking and behavior analysis.

[0058] In this embodiment, aiming at problems such as the degradation of the quality of surveillance video images, the increase in noise, the occlusion of people gathering, and background interference under low-light conditions, adaptive histogram equalization is used to enhance the image contrast, a light change detection model is constructed to dynamically adjust the image enhancement parameters, and the human pose estimation and attention mechanism of deep learning are introduced to improve the individual recognition ability. This method significantly improves the accuracy and robustness of people surveillance in low-light environments, providing an effective solution for intelligent surveillance in complex scenarios.

[0059] On the other hand, based on the same inventive concept as the above embodiments, this embodiment also provides a personnel recognition system for low-light environments, which is used to implement the personnel recognition method for low-light environments. The usage method of this system can be mutually referred to with the recognition method provided in the above embodiments. The personnel recognition system includes: an image acquisition module, an image enhancement module, and an individual recognition module;

[0060] The image sequence acquisition module is used to acquire the surveillance video image sequence in a low-light environment and perform preprocessing;

[0061] The image enhancement module is used to construct a light change detection model, detect the light change parameters of the preprocessed surveillance video image based on the light change detection model; and is used to adjust the parameters of the image enhancement algorithm based on the light change parameters, and process the preprocessed surveillance video image based on the adjusted image enhancement algorithm to obtain an enhanced image;

[0062] The individual recognition module is used to perform key point detection and skeleton extraction on the enhanced image based on the human body pose algorithm, and introduce an attention mechanism and context information to separate and recognize individuals in the enhanced image.

[0063] Embodiment 2

[0064] The present invention also provides an electronic device, including: a memory and a processor; the memory is used to store a program; the processor is used to execute the program to implement each step of the personnel recognition method for low-light environments.

[0065] Embodiment 3

[0066] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each step of the personnel recognition method for low-light environments is implemented.

[0067] The above is only the preferred specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for identifying people in a low-light environment, characterized in that: The following steps are involved: Acquire surveillance video image sequences in low-light environments and perform preprocessing; Constructing an illumination change detection model, and detecting illumination change parameters of the preprocessed surveillance video image based on the illumination change detection model; Adjusting the parameters of the image enhancement algorithm based on the illumination change parameters, and processing the preprocessed surveillance video image based on the adjusted image enhancement algorithm to obtain an enhanced image; Based on the human body posture algorithm, key point detection and skeleton extraction are performed on the enhanced image, and an attention mechanism and context information are introduced to separate and identify individuals in the enhanced image.

2. The method for identifying people in a low-light environment according to claim 1, characterized in that: The process of obtaining a surveillance video image sequence in a low-light environment and preprocessing it to obtain an enhanced surveillance video image sequence includes: Each frame image in the surveillance video image sequence is divided into blocks, and each block is subjected to histogram equalization processing and smoothing filtering processing; then, the contour information in the image is extracted by an edge detection algorithm, and the personnel area is extracted according to the gradient direction and size of the contour pixel points; the extracted personnel area is subjected to binarization processing and detail enhancement processing, and the processed personnel area is fused with the original image to obtain a preprocessed surveillance video image.

3. The method for identifying people in a low-light environment according to claim 1, characterized in that: The illumination change detection model is constructed, and the process of detecting illumination change parameters of the preprocessed surveillance video image based on the illumination change detection model includes: The brightness features and contrast features of the preprocessed surveillance video image are extracted, and based on the brightness features and contrast features, a pre-built illumination change detection model is used to determine whether illumination changes have occurred in the current image; if illumination changes are detected, the illumination change area in the image is extracted and the illumination change parameters are estimated.

4. The method for identifying people in a low-light environment according to claim 3, characterized in that: The process of adjusting the parameters of the image enhancement algorithm based on the illumination change parameters, and processing the preprocessed surveillance video image based on the adjusted image enhancement algorithm to obtain an enhanced image includes: Based on the illumination change parameters, dynamically adjusting the parameters of the image enhancement algorithm, enhancing the brightness, local contrast and texture of the image based on the adjusted parameters of the image enhancement algorithm, and then fusing the images after brightness adjustment, local contrast enhancement and texture enhancement to obtain a final enhanced image; The parameters of the image enhancement algorithm include a brightness adjustment coefficient and a contrast adjustment coefficient.

5. The method for identifying people in a low-light environment according to claim 1, characterized in that: The process of separating and identifying individuals in the enhanced image includes processing images without mutual occlusion of individuals and processing images with mutual occlusion of individuals; Among them, the process of processing images without mutual occlusion of people includes: performing key point detection on the human body area in the enhanced image based on the human body posture algorithm, obtaining the key point coordinates representing the position of the human body joints, connecting the key point coordinates to form a human skeleton, and extracting the human skeleton feature vector for representing the human body posture information; then extracting the human body posture features for each frame image in the monitoring video image sequence, combining the time information to construct the movement trajectory of the person, and dividing the movement trajectory of the person into different individuals through the trajectory clustering algorithm to achieve the separation of the people; extracting the appearance features of each individual, combining the individual's posture information, movement trajectory and appearance features, and using a multi-feature fusion method to identify the separated individuals, obtain the identity information of each person, and complete the identification.

6. The method for identifying people in a low-light environment according to claim 5, characterized in that: The process of processing images where people occlude each other includes: A convolutional neural network is used to extract features from images where people occlude each other to obtain image features, and then the image features are weighted through an attention mechanism. A recurrent neural network is used to model images where people occlude each other to capture contextual information between people. The image features weighted by attention and the contextual information learned by the recurrent neural network are fused to form a fused feature representation. Based on the fused feature representation, a region proposal network is used to generate candidate target regions, and it is determined whether the candidate target region contains the target person. For the region where the target person is detected, a full convolutional network is further used to perform pixel-level segmentation on it. The detected target person is post-processed according to the segmentation result to obtain the final person separation result.

7. A person recognition system for low-light environments, characterized in that: Used to implement the method for identifying people in a low-light environment as described in any one of claims 1 to 6, comprising: an image acquisition module, an image enhancement module and an individual identification module; The image sequence acquisition module is used to acquire and pre-process surveillance video image sequences in a low-light environment; The image enhancement module is used to construct an illumination change detection model, and detect illumination change parameters of the preprocessed surveillance video image based on the illumination change detection model; and is used to adjust parameters of the image enhancement algorithm based on the illumination change parameters, and process the preprocessed surveillance video image based on the adjusted image enhancement algorithm to obtain an enhanced image; The individual recognition module is used to perform key point detection and skeleton extraction on the enhanced image based on a human posture algorithm, and introduce an attention mechanism and context information to separate and identify individuals in the enhanced image.

8. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the method for identifying people in a low-light environment as described in any one of claims 1 to 6.

9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the method for identifying a person in a low-light environment as described in any one of claims 1 to 6 is implemented.