Pedestrian gait recognition method under dense shielding condition
By adopting the main and auxiliary dual-channel network with DHPP structure and CoordConv method in the gait recognition method, combining data augmentation and network cropping, the problems of low gait recognition accuracy and high feature dimensions under dense occlusion conditions are solved, and the gait recognition effect with high robustness and low deployment cost is achieved.
Patent Information
- Application Number
- CN202311726381.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-17
AI Technical Summary
The existing gait recognition methods have low accuracy and high feature dimensions under dense occlusion conditions, resulting in high deployment costs and difficult to achieve large-scale promotion.
The main and auxiliary dual-channel gait recognition network based on DHPP structure and CoordConv method is adopted to reduce feature dimensions and improve robustness through data augmentation and network cropping.
Under dense occlusion conditions, the robustness of gait recognition is significantly improved, the feature dimension is reduced, the deployment cost is saved, and it is suitable for large-scale promotion.
Smart Images

Figure CN120164249A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image detection, and particularly relates to a pedestrian gait recognition method under dense occlusion conditions. Background Art
[0002] As a long-distance non-cooperative authentication method, gait recognition methods have important applications in the fields of security, criminal investigation, social security, etc. Compared with face recognition solutions, gait recognition methods have lower requirements for target distance and resolution, and do not require clear frontal images to be taken. Compared with re-identification solutions, gait recognition methods are more robust to changes in target clothing and accessories, and are more suitable for non-cooperative scenarios and long-term target search. The research on gait recognition algorithms has received increasing attention in recent years and has become a key technology for the new generation of identity authentication.
[0003] At present, the mainstream gait recognition methods mainly rely on the target contour sequence as the input to identify the target identity. In addition, in recent years, there have been some research results on gait recognition that use dense 3D representations as the input. However, due to its dependence on 3D pose recovery, it is relatively difficult to be applied in practice. Currently, the gait recognition algorithms based on contour sequences are mainly divided into three categories. The first type of method converts the gait sequence images into a single gait energy map for processing, and the subsequent processing methods are similar to those of face recognition, person re-identification, etc. The network structure of this type of method is relatively simple, and the computational complexity is more controllable. However, this type of method ignores the important temporal information of gait images, and at the same time, the synthesis of gait energy maps makes some spatial detail information blurred, making it difficult for the method to recognize similar targets and generally difficult to achieve high accuracy. The second type of method relies on temporal convolutional methods to process the continuous sequence images of gait contours, emphasizing the continuous changes of gait features between several consecutive frames, and has been well developed due to its good short-term recognition effect. However, this type of method is most vulnerable to interference from factors such as walking frequency and sampling frequency, and its performance will be greatly reduced in actual application scenarios. The third type of method, represented by the Gaitset method, is a gait recognition algorithm that does not rely on the sequential constraint of gait images, has high flexibility, and has obvious advantages for applications in actual scenarios. However, this type of method still has its limitations. First, like the other two types of gait recognition methods, the input of this type of method depends on the contour extraction of pedestrian targets. Whether using instance segmentation methods or moving target detection methods, the contours extracted in the case of severe pedestrian crowd occlusion will always be affected by occlusion. The contours obtained by combining moving target detection with pedestrian detection methods will have redundant responses, while the target contours obtained by instance segmentation methods will be incomplete due to occlusion. This redundancy and incompleteness will greatly reduce the accuracy of gait recognition methods. Second, the robustness of this type of method in interference scenarios such as carrying backpacks and wearing coats still needs to be improved. Finally, the features extracted by this type of method are as high as 15,872 dimensions, and it is very difficult to perform large-scale and fast matching searches for such high-dimensional features in the case of large-scale and long-term gait matching problems in actual use. Summary of the Invention
[0004] In view of the above problems or improvement requirements of the prior art, the present invention provides a pedestrian gait recognition method under dense occlusion conditions. While ensuring the accuracy of the algorithm, it reduces the output feature dimension of the algorithm, saves deployment costs, and is suitable for large-scale promotion and use.
[0005] The technical solution adopted by the present invention to achieve the above object is as follows:
[0006] A pedestrian gait recognition method under dense occlusion conditions, comprising the following steps:
[0007] 1) Preprocess the pedestrian gait profile training images to obtain a multi-pedestrian overlapping and occluding image dataset;
[0008] 2) Use a data augmentation method based on random binary dilation to perform secondary augmentation on the training data;
[0009] 3) Construct a main and auxiliary dual-channel gait recognition network based on the DHPP structure and the CoordConv method, and use the augmented data to jointly train the main and auxiliary dual-channel gait recognition network;
[0010] 4) Prune the trained main and auxiliary dual-channel gait recognition network, only retain the main channel part, completely remove the auxiliary channel, obtain a pedestrian gait recognition network under dense occlusion conditions, and perform pedestrian gait recognition through the pedestrian gait recognition network.
[0011] The step 1) includes the following steps:
[0012] 1.1) For each target image to be trained, randomly select a random contour image from other ID targets at random angles as interference, merge the contours with the image to be trained to obtain a mixed image;
[0013] 1.2) Calculate the target contour after being interfered in the mixed image;
[0014] 1.3) Extract the target contour after being interfered through the original bounding box position to obtain a multi-pedestrian overlapping and occluding image;
[0015] 1.4) Repeat steps 1.1) to 1.3) to construct a multi-pedestrian overlapping and occluding image database as the image dataset to be trained.
[0016] The step 1.2) is specifically:
[0017] Let w and h be the width and height of the original target bounding box in the mixed image, then the center coordinate range of the interference image is [0:w, -h / 2:3h / 2], and the mixed image is cropped according to the position and size of the original target image as the target contour image.
[0018] The step 2) includes the following steps:
[0019] 2.1) Randomly extract a partial area on each original image in the multi-pedestrian overlapping and occluding image dataset, copy and map it to a zero matrix of the same size as the original image to obtain an image of the area to be dilated;
[0020] 2.2) Perform binary dilation operation on the image of the area to be dilated using dilation kernels of random sizes and shapes;
[0021] 2.3) Perform a logical OR operation on the dilated image and the original image to obtain the finally output augmented image.
[0022] The shape of the expansion kernel includes an ellipse and a rectangle, and the size of the expansion kernel includes (3,3), (5,5), (7,7), (9,9), (3,7), (7,3).
[0023] The said step 3) includes the following steps:
[0024] 3.1) Construct a main-channel feature extraction network based on the CoorConv method;
[0025] 3.2) Construct an auxiliary channel for the main-channel feature extraction network;
[0026] 3.3) Add the output features of the two channels bit by bit to obtain the final output features, and complete the construction of the main-auxiliary dual-channel gait recognition network;
[0027] 3.4) Use the triplet loss function to train the main-auxiliary dual-channel gait recognition network.
[0028] The main-channel feature extraction network extracts features from the image sequence data of the same ID and the same angle through a weight-sharing CoordConv convolutional network, performs feature fusion on the finally output feature map using the global pooling method, and constructs a DHPP structure to extract the features in the spatial dimension in blocks.
[0029] The auxiliary channel fuses the main-channel features by means of global pooling after each downsampling of the main-channel feature extraction network, and uses the DHPP structure to extract the block features of the fused features in the spatial dimension.
[0030] The DHPP structure horizontally divides the feature map to obtain different sub-regions, adds an attention structure to each of the average pooling and max pooling branches respectively, and adaptively weights and screens the features selected for each sub-region.
[0031] A pedestrian gait recognition system under dense occlusion conditions includes:
[0032] A data preprocessing module, which is used to preprocess the pedestrian gait contour training images to obtain a multi-pedestrian overlapping occlusion image dataset;
[0033] A data augmentation module, which is used to perform secondary augmentation on the training data using a data augmentation method based on random binary dilation;
[0034] A main-auxiliary dual-channel gait recognition network construction module, which is used to construct a main-auxiliary dual-channel gait recognition network based on the DHPP structure and the CoorConv method, and jointly train the main-auxiliary dual-channel gait recognition network using the augmented data;
[0035] The pedestrian gait recognition module is used to prune the trained main - auxiliary dual - channel gait recognition network, only retain the main channel part, completely remove the auxiliary channel, obtain the pedestrian gait recognition network under dense occlusion conditions, and perform pedestrian gait recognition through the pedestrian gait recognition network.
[0036] The present invention has the following beneficial effects and advantages:
[0037] 1. The present invention realizes pedestrian gait recognition under dense occlusion conditions through an advanced algorithm, and uses the main - auxiliary dual - channel gait recognition network based on the DHPP structure and the CoordConv method to greatly improve the robustness of the gait recognition algorithm for dense occlusion scenarios.
[0038] 2. The present invention uses the training - joint - test pruning method to simplify the network structure, reduce the algorithm output feature dimension while ensuring the algorithm accuracy, save the deployment cost, and is suitable for large - area promotion and use. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is the flowchart of pedestrian gait recognition under dense occlusion conditions of the present invention;
[0040] Figure 2 is the schematic diagram of the DHPP structure constructed in step three;
[0041] Figure 3 is the schematic diagram of the gait recognition network structure constructed in step three. DETAILED DESCRIPTION OF THE INVENTION
[0042] The following further describes the present invention in detail with reference to the drawings and embodiments.
[0043] As Figure 1 shown, the present invention provides a method for pedestrian gait recognition under dense occlusion conditions. The specific steps of the method are as follows:
[0044] Step 1: First, for each target image to be trained, randomly select a random contour image from a random other ID target at a random angle as interference for contour merging. The center coordinate range of the interference image is [0:w, -h / 2:3h / 2], where w and h are the width and height of the original target bounding box. The obtained mixed image can well simulate the extraction result of the target contour in a dense occlusion crowd. When the center ordinate of the occluded target is in [-h / 2:h / 2], the contour shape is similar to the target extraction result from the monitoring perspective with other pedestrians interfering behind the target. When the center ordinate of the occluded target is in [h / 2:3h / 2], the contour shape is similar to the target extraction result from the monitoring perspective when the target is occluded by other pedestrians in front. Finally, re - extract the interfered target contour through the original bounding box position to obtain the training image dataset.
[0045] Step 2: Based on the training dataset obtained in Step 1, randomly select some regions of the image on each piece of data to be trained, extract them, and copy and map them into a zero matrix of the same size as the image to obtain the image of the region to be dilated. Then, perform binary dilation operation on the image of the region to be dilated using dilation kernels of random sizes and shapes. The shapes of the dilation kernels include ellipse and rectangle, and the sizes include (3,3), (5,5), (7,7), (9,9), (3,7), and (7,3). Perform logical OR operation on the dilated image and the original image to obtain the finally output augmented image.
[0046] Step 3: As Figure 2 and Figure 3 shown, construct a main and auxiliary dual-channel gait recognition network based on the DHPP structure and the CoordConv method. The CoordConv method can provide absolute position information in the network as feature supplement to make up for the lack of absolute position information caused by the translational invariance of the convolutional structure. This problem is particularly serious in binary image processing. Through this method, the influence of the cropped multi-scale horizontal pyramid on the accuracy can be reduced to a certain extent. First, construct the main-channel feature extraction network with the help of the CoordConv method. The input of the main channel comes from the training sample library. For each batch of samples, 16 randomly angled gait contour image sets of 8 IDs are selected for training, and 30 disordered images are randomly extracted from each image set with the image resolution of 64*64. After input into the main channel, the data of each image sequence with the same ID and the same angle are subjected to feature extraction through the CoordConv convolutional network with weight sharing. Feature fusion is performed on the finally output feature map using the global pooling method, and the DHPP structure is constructed to perform block extraction on the features in the spatial dimension. The auxiliary channel, also known as the multi-level global channel, performs global fusion on the multi-level multi-scale features extracted from the pedestrian contours at different moments in the sequence images of the main channel. Its input comes from the CoordConv network of the main channel. After each downsampling of the main channel, the features of the main channel are fused and extracted with the help of global pooling, and an independent feature extraction network is constructed to further abstract the features. Similar to the main channel, the feature output part also uses the DHPP structure to perform block feature extraction in the spatial dimension. The DHPP structure first performs horizontal segmentation on the feature map into 16 parts. On this basis, attention structures are added to the average pooling and max pooling branches respectively to adaptively weight and screen the features selected for each sub-region, and this structure removes the problem that the spatial position expression of the gait network features is relatively loose. Finally, the output features of the two channels are added bit by bit to obtain the final output features. The network is trained using the triplet loss function, where the margin is set to 0.2, as shown in the following formula. The learning rate is set to 0.0001.
[0047]
[0048] Step 4: Prune the trained network, only retain the main channel part, and completely remove the auxiliary channels. The final underlying structure of the network is a feature extraction structure formed by multiple CoordConv networks with shared weights, and the top-level structure is a 1 / 16 spatial horizontal segmentation feature description structure formed by a single DHPP. The features output by the DHPP are directly connected to form a 4096-dimensional gait feature description. Finally, deploy the pruned gait recognition network on the server side, extract and analyze the features of the real-time acquired pedestrian contour sequence, output the feature description, and compare the similarity with the pedestrian gait feature data in the database to achieve the gait recognition function.
Claims
1. A pedestrian gait recognition method under dense occlusion conditions, characterized in that, It includes the following steps: 1) Preprocess the pedestrian gait contour training images to obtain a multi-pedestrian overlapping and occluding image dataset; 2) Use a data augmentation method based on random binary dilation to perform secondary augmentation on the training data; 3) Construct a main and auxiliary dual-channel gait recognition network based on the DHPP structure and the CoordConv method, and use the augmented data to jointly train the main and auxiliary dual-channel gait recognition network; 4) Prune the trained main and auxiliary dual-channel gait recognition network, only retain the main channel part, completely remove the auxiliary channel, obtain a pedestrian gait recognition network under dense occlusion conditions, and perform pedestrian gait recognition through the pedestrian gait recognition network.
2. The pedestrian gait recognition method under dense occlusion conditions according to claim 1, characterized in that, The step 1) includes the following steps: 1.1) For each target image to be trained, randomly select random contour images from other ID targets at random angles as interference, merge the contours with the image to be trained to obtain a mixed image; 1.2) Calculate the target contour after being interfered in the mixed image; 1.3) Extract the target contour after being interfered through the original bounding box position to obtain a multi-pedestrian overlapping and occluding image; 1.4) Repeat steps 1.1) to 1.3) to construct a multi-pedestrian overlapping and occluding image database as the image dataset to be trained.
3. The pedestrian gait recognition method under dense occlusion conditions according to claim 2, characterized in that, The step 1.2) is specifically: Let w and h be the width and height of the original target bounding box in the mixed image, then the center coordinate range of the interference image is [0:w, -h / 2:3h / 2], and the mixed image is cropped according to the position and size of the original target image as the target contour image.
4. The pedestrian gait recognition method under dense occlusion conditions according to claim 1, characterized in that, The step 2) includes the following steps: 2.1) Randomly extract partial regions on each original image in the multi-pedestrian overlapping and occluding image dataset, copy and map them into a zero matrix of the same size as the original image to obtain an image of the region to be dilated; 2.2) Perform binary dilation operation on the image of the region to be dilated using dilation kernels of random sizes and shapes; 2.3) Perform a logical OR operation on the dilated image and the original image to obtain the finally output augmented image.
5. The pedestrian gait recognition method under dense occlusion conditions according to claim 4, characterized in that, The shapes of the dilation kernels include ellipse and rectangle, and the sizes of the dilation kernels include (3,3), (5,5), (7,7), (9,9), (3,7), (7,3).
6. The pedestrian gait recognition method under dense occlusion conditions according to claim 1, characterized in that, The step 3) includes the following steps: 3.1) Construct a main channel feature extraction network based on the CoorConv method; 3.2) Construct an auxiliary channel for the main channel feature extraction network; 3.3) Add the output features of the two channels bit by bit to obtain the final output feature, and complete the construction of the main and auxiliary dual-channel gait recognition network; 3.4) Use a triplet loss function to train the main and auxiliary dual-channel gait recognition network.
7. The pedestrian gait recognition method under dense occlusion conditions according to claim 6, characterized in that, The main channel feature extraction network extracts features from the image sequence data of each same ID and same angle through a CoordConv convolutional network with weight sharing, performs feature fusion on the finally output feature map using the global pooling method, and constructs a DHPP structure to perform block extraction on the features in the spatial dimension.
8. The pedestrian gait recognition method under dense occlusion conditions according to claim 6, characterized in that, After each downsampling of the main channel feature extraction network, the auxiliary channel fuses the main channel features by means of global pooling, and uses the DHPP structure to extract the block features in the spatial dimension of the fused features.
9. The pedestrian gait recognition method under dense occlusion conditions according to claim 7 or 8, characterized in that, The DHPP structure horizontally divides the feature map to obtain different sub-regions, and adds attention structures to the average pooling and max pooling branches respectively to adaptively weight and screen the selected features of each sub-region.
10. A pedestrian gait recognition system under dense occlusion conditions, characterized in that, It includes: A data preprocessing module for preprocessing pedestrian gait contour training images to obtain a multi-pedestrian overlapping occlusion image dataset; A data augmentation module for using a data augmentation method based on random binary dilation to perform secondary augmentation on the training data; A main-auxiliary dual-channel gait recognition network construction module for constructing a main-auxiliary dual-channel gait recognition network based on the DHPP structure and the CoordConv method, and jointly training the main-auxiliary dual-channel gait recognition network with the augmented data; A pedestrian gait recognition module for pruning the trained main-auxiliary dual-channel gait recognition network, only retaining the main channel part, completely removing the auxiliary channel, obtaining a pedestrian gait recognition network under dense occlusion conditions, and performing pedestrian gait recognition through the pedestrian gait recognition network.