Semantic guidance feature extension method for sheltered pedestrian re-identification
By introducing a semantic guided feature extension method in pedestrian recognition technology, the LFSE module is used to fuse important local features and their adjacent features, and the problem of insufficient pedestrian recognition feature extraction under occlusion situations is solved, and the recognition accuracy of the model is improved.
Patent Information
- Application Number
- CN202510070041.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN119992593A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence computer vision technology and image retrieval fields, and in particular to a semantically guided feature expansion method for re-identification of occluded pedestrians. Background Art
[0002] Person Re-identification (Person Re-ID) refers to retrieving pedestrian information from different cameras to determine whether images captured by different cameras contain the same pedestrian. This technology is widely used in video surveillance, identity authentication, smart city management and other fields, and is one of the important technical means of social security management. In reality, pedestrian images under surveillance show high complexity and high variability. Traditional pedestrian re-identification methods that mainly rely on manual work are difficult to identify in a timely and efficient manner. With the rapid development of computer vision technology, deep learning models have been widely used in pedestrian re-identification, have made breakthrough progress, and have become the main research direction of pedestrian re-identification.
[0003] At present, pedestrian re-identification technology still faces many challenges due to the existence of occlusion. Since the surveillance cameras in real scenes such as stations, airports and shopping malls are fixed, the pedestrian images from the surveillance cameras are easily blocked to varying degrees by some obstacles, such as plants, cars, other pedestrians, etc. The appearance of occlusion causes the appearance characteristics of the same pedestrian identity in different environments to have significant differences, which seriously limits the accuracy of pedestrian re-identification. Especially for scenes with extreme performance requirements such as criminal investigation, it is usually necessary to use severely occluded images as input to query the whereabouts of suspects, which makes it almost impossible for traditional pedestrian re-identification models to extract enough distinguishable features, causing the model to completely lose its recognition ability under occlusion.
[0004] In order to better deal with the problem of insufficient extraction of discriminable features of pedestrians in pedestrian re-identification under occlusion, some studies focus on mining more pedestrian identity information. Yang et al. (2023) proposed to suppress the influence of interference information by discretizing posture information into visibility labels of body parts, and obtain more posture information of visible areas as additional semantic clues to improve the model's recognition ability under different occlusion conditions. In order to extract more semantic information that is not paid attention to in the pre-trained model, Zhang et al. (2023) proposed a complementary network, which uses two branches to extract potential discriminant information in the background that is useful for pedestrian recognition and salient areas in the global range. The two branches complement each other to achieve the extraction of additional semantic information such as backpacks and handbags. Through this design, the network can extract richer semantic features and improve the recognition accuracy of the model.
[0005] Other studies focus on obtaining pedestrian identity features in visible areas. Gao et al. (2024) proposed a teacher-student decoder (TSD) framework that combines the Transformer decoder and human body analysis information to better locate and aggregate the features of pedestrian body parts. Dong et al. (2024) proposed a framework called multi-view information fusion and propagation (MVI2P), which generates a comprehensive pedestrian representation by integrating feature maps of multi-view images, and designs positioning and quantization modules to selectively integrate discernible pedestrian identity features. However, when the highly recognizable features of pedestrians are occluded, the similarity of different pedestrian samples may increase. Existing methods lack effective strategies to specifically mine local features in visible areas to enhance the expressiveness of pedestrian samples, which limits the recognition accuracy of the model. Summary of the invention
[0006] Purpose of the invention: To address the problem that the similarity of different pedestrians increases due to the loss of some identifiable features of pedestrians caused by occlusion, the present invention proposes a semantic-guided feature expansion method for re-identification of occluded pedestrians. This method uses a ResNet-50 backbone network and initializes the backbone network using the LUPerson pre-trained model; a local feature semantic expansion module (LFSE) is proposed to extract more discriminative semantic features of pedestrians and improve the expressiveness of pedestrian features. Specifically, LFSE selects important unobstructed local features of pedestrians and their neighboring features for fusion, introduces more discernible detail information to enhance the expressiveness of local features, and improves the adaptability of the model in occluded scenarios.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] A semantically guided feature expansion method for occluded pedestrian re-identification comprises the following steps:
[0009] (1) Data preprocessing: Market-1501 and Occluded-DukeMTMC large-scale pedestrian datasets are selected as experimental datasets, and the datasets are divided into three parts: training set, test set, and query set. Data preprocessing is performed using data normalization, random horizontal flipping, random erasing, random cropping, and other techniques;
[0010] (2) Model construction: The backbone network is constructed based on the residual network ResNet-50, the backbone network is initialized using the pre-trained model, and S LFSE is applied to expand the feature semantics to construct a semantic-guided feature expansion method;
[0011] (3) Model training: Use the sampler to sample a predefined number of images, package them into a batch and input them into the model. Use the defined and initialized neural network to extract the global and local features of the pedestrians in the image. Use the loss function to calculate the loss value of the extracted feature vector and back propagate to update the model parameters. Iterate continuously to minimize the value of the loss function and obtain the optimized pedestrian re-identification model.
[0012] (4) Model testing: Use the trained model to initialize a custom neural network, extract and fuse global and local features to represent the appearance feature vector of the pedestrian, calculate the Euclidean distance between the feature vectors of the pedestrian to be queried and the pedestrians in the query library, and sort them by similarity to find the pedestrian most similar to the pedestrian to be queried.
[0013] The Market-1501 dataset selected in step (1) is specifically divided as follows:
[0014]
[0015] The dataset Occluded-DukeMTMC-reID selected in step (1) is specifically divided as follows:
[0016]
[0017] Among them, in the data preprocessing in step (1):
[0018] (1-1) The first step of data preprocessing is to resize the input image to 256×128;
[0019] (1-2) Perform data normalization, scaling the data to between 0 and 1 or scaling the data to a standard normal distribution with a mean of 0 and a variance of 1 to avoid the negative impact of data size differences on model training;
[0020] (1-3) Perform a random horizontal flip operation to flip the image left and right with a pre-set probability to generate new images and increase the diversity of training data;
[0021] (1-4) Perform a random erasing operation to fill a rectangular area with a random value at a random position in the image to generate a new training image, forcing the model to focus on the features of different areas and improve the robustness of the model;
[0022] (1-5) Perform a random cropping operation to randomly crop an area from the original image to generate more training data. This method can increase the diversity of data and help improve the generalization performance of the model.
[0023] Among them, the construction and initialization of the backbone network in step (2) are:
[0024] (2-1) ResNet is composed of multiple convolutional layers, including multiple residual blocks. Each residual block includes a skip connection, which directly adds the input to the output, thereby retaining the original feature information. This design idea not only makes the network easier to optimize, but also alleviates the problems of gradient vanishing and gradient explosion, and improves the accuracy of the network. This method uses the ResNet-50 version, and sets the step size of the last convolutional layer to 1 to meet the needs of the pedestrian re-identification task. The construction of the branch network is constructed using the basic unit of ResNet. Each basic unit consists of a convolution filter, a batch normalization layer (BN) and a linear rectifier (ReLU). The convolution kernel size is divided into 1×1 and 3×3. We use the LFSE module to select and fuse important features and their neighboring local features to obtain more expressive important local features;
[0025] (2-2) The pre-trained model for initializing the backbone network comes from the LUPerson dataset. ResNet is pre-trained on LUPerson and can then be trained on other datasets for specific tasks through fine-tuning or transfer learning.
[0026] Wherein, in step (3):
[0027] (3-1) After the backbone network extracts features, the global features obtained by global average pooling are combined with the local features obtained by semantic expansion of the LFSE module to obtain pedestrian features representing identity information;
[0028] (3-2) Use the cross entropy loss function to optimize global features, and use the difficult sample sampling triple loss function and the center loss function to optimize the combined features. The cross entropy loss is as follows:
[0029]
[0030] Here, N is the number of samples, y i One-hot encoding of the pedestrian with identity i. It is the probability score of a pedestrian i in the feature vector G extracted by the model.
[0031] The Center Loss formula is as follows:
[0032]
[0033] in, represents the features of the kth class before the fully connected layer, represents the feature center of the kth category, and n is the size of a batch.
[0034] The formula for the triplet loss function of difficult sample sampling is as follows:
[0035]
[0036] Where f(x) is the trained model, represents the distance between positive sample pairs, Represents the distance between negative sample pairs. [·] + is the max(·,0) function, P is the number of pedestrian IDs in the sample space, K is the number of pictures of each pedestrian, is the anchor sample, is a positive sample, is a negative sample. m is a threshold used to set the relaxation margin between positive and negative sample pairs;
[0037] (3-3) In summary, the total loss of the semantic-guided feature expansion method is as follows:
[0038]
[0039] Among them, λ 1,2,3 is the weight coefficient of each loss function.
[0040] The beneficial effects of the present invention are as follows: the present invention constructs a semantically guided feature expansion method for re-identification of occluded pedestrians, which enhances the expression ability of local features by fusing important local features of pedestrians and their neighboring features, suppresses the problem of increased similarity of pedestrian samples caused by the loss of distinguishable features due to occlusion, and improves the recognition ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flowchart of pedestrian re-identification of the present invention;
[0042] Figure 2 It is the overall network model diagram of the present invention;
[0043] Figure 3 It is a detailed structural diagram of the local feature semantic expansion module of the present invention. DETAILED DESCRIPTION
[0044] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0045] The semantically guided feature expansion method for occluded pedestrian re-identification described in the present invention has an overall model training and recognition process as follows: Figure 1 The detailed structure of the semantic-guided feature expansion method is shown in Figure 2 As shown, the following steps are included:
[0046] (1) Dataset selection: The present invention selects the Market-1501 dataset and the Occlude-DukeMTMC-reID dataset for experiments from the pedestrian re-identification dataset. The Occlude-DukeMTMC-reID dataset is an occluded dataset derived from the DukeMTMC-reID dataset, which contains 15,618 training images, 17,661 gallery images, and 2,210 occluded query images. It is a very challenging dataset. The Market-1501 dataset is a dataset for pedestrian re-identification (Person Re-Identification). The dataset was released by Dr. Li et al. in 2015. The images in the dataset were collected from 6 cameras on the Tsinghua University campus, and a total of 1,501 pedestrians were annotated. Among them, 751 pedestrian annotations are used for the training set, and 750 pedestrian annotations are used for the test set. Each pedestrian image is labeled with a pedestrian ID and a camera ID. The pedestrian ID is used to distinguish different pedestrians, and the camera ID is used to identify which camera the image comes from.
[0047] (2) Training data preprocessing: The present invention scales the training data to 256×128, and then randomly flips, crops, erases, and blocks the images horizontally, where the preprocessing probability of each image is set to p=0.3. We then select the Adam optimizer as the optimization method for our model. In the first 10 rounds of training, we use the wormup learning rate to warm up, and then decay it in the 30th and 70th rounds, with the decay rate set to 0.1. The batch size is set to 64, and the total number of training epochs is set to 130.
[0048] (3) Backbone network model loading settings: The model is based on ResNet-50 and introduces the semantic-guided feature semantic expansion (LFSE) module proposed before and after. ResNet-50 is mainly divided into 5 stages: Layer0, Layer1, Layer2, Layer3 and Layer4. Layer0 includes a 7x7 convolution layer and a maximum pooling layer for preliminary feature extraction and dimensionality reduction of the input image. Layer1, Layer2, Layer3 and Layer4 are four groups of convolution layers containing residual blocks, with 3 residual blocks, 4 residual blocks, 6 residual blocks and 3 residual blocks respectively. Each residual block contains multiple convolution layers and identity mapping, which can help the network better learn the feature representation of the input image.
[0049] (4) Constructing a semantically guided local feature semantic expansion (LFSE) module:
[0050] (4-1) Constructing the Local Feature Semantic Expansion (LFSE) module: In order to address the problem that the similarity between different pedestrian samples may increase when the highly recognizable features of pedestrians are blocked, we proposed the Local Feature Semantic Expansion (LFSE), which forms more expressive local features by fusing important local features and neighboring features, thereby introducing more discernible information in important local features and enhancing the expressiveness of important local features.
[0051] The specific expansion process is as follows Figure 3 As shown, we use the convolutional network with spatial and channel attention to obtain the important area channels of interest to the model and the surrounding area channels corresponding to them in space. We represent the input original image as Y and use the formula:
[0052] F i =Cam(Conv(Y)), F′=SC(Conv(Y)),
[0053] Obtain important area channel F i The corresponding surrounding area channel F′={f′ n , n=1,...,N}, where the convolution function Conv(·) follows the same structure as Resnet50, SC(·), Cam(·) are used to extract the required channels. Next, in order to obtain an additional clue set F′ sorted by importance sort ={f″ n , n=1,...,N}, we first calculate the channel importance weights, where the result obtained by applying the global average pooling operation to the surrounding area channel F′ is used as the input and then passes through two fully connected layers and the activation function to obtain the weight matrix The formula is as follows.
[0054]
[0055] Next, these importance weights are used to sort the feature maps in F′, as follows:
[0056]
[0057] Where GAP(f′ n ) is the global average pooling effect, W1 and W2 are the learned parameter matrices, δ is the ReLu activation function, σ represents the Sigmoid activation function, and sort(·) represents the operation of sorting channels according to their weights.
[0058] Since the information represented by the elements in the sorted channel set is often subject to certain interference or duplication, we also need to select them and select features that have more semantic information and are more beneficial to the semantic expansion of important features. Here we select them based on the attention weight between each feature. Specifically, the feature channel set F′ sorted by importance sort As input, the output is the weight of the difference between each channel and other channels, which constitutes the weight set W Diff ={w n , n=1,...,N}, where each weight element is called a difference weight, and the formula is as follows:
[0059]
[0060] in represents the difference between channel n and other channels, diff(·) is the function to calculate the difference between channels, and w n is the selection weight of channel n, It is the comprehensive difference score obtained by averaging the difference between channel n and other channels, and the difference weight set W obtained through training Diff Finally, we can select the final fusion channel based on the difference between channels and their importance. The formula is as follows:
[0061]
[0062] Among them, α, β, A are all learnable hyperparameters, S(·) is the selection function, and F′ select It is the channel feature that is selected and fused as additional semantic information. Finally, the important area feature channel is fused with it to obtain the final semantic extension feature F exp :
[0063]
[0064] Through the proposed LFSE module, we can expand the semantic information of important local features according to additional clues, increase the discernible information of important local features, and improve its expressive ability of pedestrian identity.
[0065] (5) Training model
[0066] The total loss is calculated and back propagated to minimize the loss function, and the optimized semantic-guided feature expansion method is obtained. After 140 rounds of iterations in the following environment, the trained model is obtained.
[0067]
[0068] (6) Test model
[0069] (6-1) The test set images in the dataset are enlarged from 64×128 to 128×384 and then normalized. The trained network model is loaded into the system and the local and global features are fused.
[0070] (6-2) In order to ensure the fairness and credibility of the experimental comparison, we use the Cumulative Matching Characteristics (CMC) curve and mean average precision (mAP) as evaluation indicators to measure the recognition accuracy and mean average precision, and follow the evaluation settings provided by the existing occluded pedestrian re-identification methods. The experimental settings include training and testing using a specific deep learning model or algorithm, and evaluating on the test set. The Rank-k accuracy A can be expressed as:
[0071]
[0072] Here p is probe, q is query, N q is the number of images, f CMC As follows:
[0073]
[0074] The CMC curve is plotted with k as the horizontal axis and Rank-k accuracy A as the vertical axis. In actual use, the more representative Rank-1, Rank-5 and Rank-10 accuracy are usually selected to replace the CMC curve, among which Rank-1 accuracy is the most important indicator. The mean average precision (mAP) is another important evaluation indicator. The CMC curve usually only cares about the ranking of the top positive samples in the retrieval library, while mAP is determined by the ranking results of all positive samples in the gallery, so it can usually reflect the performance of the model more robustly. Calculating AP requires the following three steps: Precision: For a probe image q in the query j , a series of sorted results of the gallery are returned. Consider the first n query results. Assume that the number of pedestrian IDs in the first n results that are the same as the probe image is c(n,q j ):
[0075]
[0076] Average Precision: For the C search images of the query, Precision records the precision of all N positive samples under an identity category and calculates their average precision AP, as shown in the following formula:
[0077]
[0078] Here TP represents the number of correctly predicted positive samples, and FP represents the number of incorrectly predicted negative samples. n , Precision n are the number of images of the identity category and the precision of the nth sample
[0079] Mean Average Precision (mAP): The average of the Average Precision of all search images, that is:
[0080]
[0081] The mean average precision (mAP) measures the average recognition ability of the model in different categories. Generally speaking, the more images an image dataset contains, the more the mAP value obtained by the model on this dataset can reflect its generalization ability. At the same time, if a dataset contains a variety of different difficult scenes, the model needs to be able to adapt to these changes to obtain better detection results, so that the mAP value trained on a difficult dataset can better reflect the recognition ability of the model in difficult scenes. Therefore, the CMC curve and mAP indicators are usually used together to evaluate the accuracy of the pedestrian re-identification model.
[0082] Finally, the network model was optimized using the above method, and the trained model was tested using the above evaluation indicators. The experiments were conducted on the Market-1501 and Occlude-DukeMTMC-reID datasets. The experiments proved that the semantic-guided feature expansion method effectively improved the model recognition results, and made significant progress in both mAP and Rank-1 evaluation directions.
[0083] The above embodiments have introduced in detail the specific implementation methods of the semantically guided feature expansion method for occluded pedestrian re-identification proposed by the present invention. The introduction of the above embodiments is only used to help understand the proposed method and core ideas of the present invention. According to the ideas of the present invention, there may be some differences in the specific implementation methods. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A semantically guided feature expansion method for occluded pedestrian re-identification, characterized in that: Includes steps: (1) Data preprocessing: Select the data set required for the experiment from the person re-ID data set, divide the data set and perform data preprocessing; (2) Constructing the main network structure: Building a backbone network based on the residual network ResNet-50 and initializing the backbone network using the pre-trained model; then performing local feature semantic expansion, selecting important features and their neighboring features for fusion to generate more discriminative local features, and improving the model's recognition ability under occlusion; (3) Model training: First, the batch-sampled pedestrian samples are preprocessed and input into a predefined model to extract the high-order features of pedestrians. Then, the global features extracted by the backbone network are concatenated with the local features extracted by the local feature semantic expansion module and input into the joint loss function to calculate the loss and perform back propagation, update the model parameters, and continuously iterate to minimize the value of the loss function to form an optimized pedestrian re-identification model. (4) Model testing: The proposed network model is used as a pedestrian feature extractor after training. Then, pedestrian features with individual identity identification information are used to represent pedestrians. The Euclidean distance between the feature vectors of the target pedestrian and the pedestrians in the test dataset is calculated. The pedestrians in the dataset are sorted according to their similarity with the target pedestrian. The pedestrian image most similar to the target pedestrian is obtained and the recognition accuracy of the model is calculated.
2. The semantically guided feature expansion method for occluded pedestrian re-identification according to claim 1, characterized in that: Step (1) scales the original image to a size of 256×128, and uses preprocessing data enhancement such as data normalization, random horizontal flipping, random erasing, random cropping, and random blocking techniques on the original pedestrian image dataset.
3. The semantically guided feature expansion method for occluded pedestrian re-identification according to claim 1, characterized in that: (a) The convolutional neural network defined in step (2) uses ResNet-50 as the backbone network. ResNet-50 has 4 residual blocks in each stage. The first stage contains 3 residual blocks, and the second, third, and fourth stages contain 4, 6, and 3 residual blocks respectively. (b) The semantic-guided feature expansion method in step (2) is to extract more discernible features of pedestrians in occluded scenes. The local feature semantic expansion module forms more discriminative local features by fusing important visible features and their neighboring features.
4. The semantically guided feature expansion method for occluded pedestrian re-identification according to claim 1, characterized in that: The joint loss function L constructed in step (3) is as follows: where λ 1,2,3 is the weight coefficient of each loss function, is the triplet loss function for difficult sample sampling, is the center loss function, It is the cross entropy loss function. By combining multiple loss functions, we can more comprehensively consider problems such as multi-objective and overfitting and improve the generalization performance of the model.
Citation Information
Patent Citations
Pedestrian re-identification method based on multi-layer fusion and alignment division
CN111881780A
Cross-border ethnic text sorting method and device fusing document theme features
CN115114400A
Pedestrian re-identification method based on residual multi-channel attention multi-feature fusion
CN115830531A
Transform-based semantic information enhanced behavior recognition method
CN116363555A
Effective and lightweight multi-scale pedestrian re-identification method
CN117078967A