Residual error expansion fusion method for pedestrian re-identification
By adopting the residual expansion fusion method in pedestrian re-identification technology, the loss of pedestrian feature information during the transmission process is solved and the recognition performance is improved.
Patent Information
- Application Number
- CN202510202576.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-24
AI Technical Summary
When existing pedestrian re-identification technology faces occlusion, complex scenarios, and other situations, feature information is easily lost, resulting in a degradation of recognition performance.
Residual expansion and fusion method is adopted to extract low and high-level features through the modified ResNet-50 backbone network, and a replica is generated using the weight sharing mechanism in the residual fusion module to expand and fusion in vertical and horizontal directions, retaining the key feature information of the original layer.
It effectively reduces the loss during feature transmission and improves the accuracy and recognition performance of pedestrian re-identification.
Smart Images

Figure CN120047972A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision technology and image retrieval. Specifically, it relates to a residual augmentation fusion method for pedestrian re-identification. Background Art
[0002] Pedestrian re-identification is a technology that aims to identify a target pedestrian in video sequences from existing possible sources and non-overlapping camera fields of view, and has a wide range of applications in many fields such as public safety and the construction of smart cities. In this context, pedestrian re-identification, as a key means of urban security monitoring, can assist urban managers and security monitoring personnel in tracking suspicious persons, and thus more effectively ensure social order stability and the lives and property safety of the public.
[0003] Previous pedestrian re-identification technologies mainly relied on object detection and tracking algorithms. However, when faced with occlusion and complex scenes, the performance of such technologies is often unsatisfactory, which has also to some extent promoted the birth of the cross-technology of deep learning and pedestrian re-identification.
[0004] Pedestrian re-identification technology not only has a profound impact on urban management and security protection, but also plays a certain role in the monitoring of pedestrian behavior. By analyzing pedestrian behavior, abnormal behaviors such as loitering and tailing can be discovered in a timely manner. Using pedestrian re-identification technology can effectively prevent criminal acts and ensure public safety. Therefore, this technology has attracted much attention in the fields of computer vision and artificial intelligence research, and has attracted in-depth research by many scholars.
[0005] With the successful application of deep learning in the field of pedestrian re-identification by researchers, this field has developed rapidly. However, in real situations, pedestrians face challenges such as cluttered backgrounds, partial occlusion, various postures, illumination changes, and low resolution, which greatly blur the features of the same person. Especially in the case of occlusion, different body parts are often occluded and become invisible, greatly increasing the challenge of pedestrian re-identification.
[0006] Currently, the commonly used method is to use a multi-scale feature fusion strategy to extract global features and local features. For example, Zhang et al. proposed a lightweight feature pyramid branch to extract features from different levels of the network and aggregate them into a bidirectional pyramid structure. However, most current methods do not consider that the features extracted by the low-level network are propagated through too many network layers, resulting in the loss of original information features and reducing the diversity of pedestrian features. Summary of the Invention
[0007] Objective of the Invention: As mentioned above, in the process of pedestrian feature extraction, more representative information is generally obtained by fusing multi-layer pedestrian features. However, in the process of transmitting low-level pedestrian features to high-level features, the loss of pedestrian representation information will occur, which will affect the recognition performance of the model. To alleviate this problem, the present invention proposes a residual expansion fusion method for pedestrian re-identification. The backbone network is constructed using a modified ResNet-50 to extract effective low-level and high-level features. Then, the residual fusion module uses a weight sharing mechanism to generate copies of the input feature vectors for each network layer using two different sizes of residual units. Subsequently, vertical and horizontal expansion and fusion operations are performed on these copies respectively to effectively retain the key feature information of the original layer and minimize the loss of important features during transmission.
[0008] The objective of the present invention can be achieved through the following technical solutions:
[0009] A residual expansion fusion method for pedestrian re-identification, comprising the following steps:
[0010] (1) The Market-1501 and CUHK03 pedestrian datasets are selected as the experimental datasets, and the datasets are divided into three parts: a training set, a test set, and a query set. Data preprocessing is performed using techniques such as data normalization, random horizontal flipping, random flipping, random erasing, and random rotation;
[0011] (2) The convolutional neural network used has ResNet-50 as the backbone network, and its structure includes 4 stages. Each stage uses skip connections to alleviate the problem of gradient disappearance, so that deeper networks can be trained. This design enables the network to efficiently extract multi-scale feature information while maintaining high computational efficiency, providing rich feature representations for subsequent tasks;
[0012] (3) The batch-sampled pedestrian samples are preprocessed and input into a pre-defined model to extract high-order features of pedestrians; then, the low-level and high-level features extracted by the backbone network are input into the residual expansion fusion module to reduce feature loss through expansion, so that the low-level and high-level features are fused to the same size, and are input into the loss function to calculate the loss and perform backpropagation, update the model parameters, and continuously iterate to minimize the value of the loss function, forming an optimized pedestrian re-identification model;
[0013] (4) After training, the proposed network model is used as a pedestrian feature extractor. Each pedestrian is represented by pedestrian features containing identity recognition information. By calculating the Euclidean distance between the target pedestrian and the feature vectors of each pedestrian in the test dataset, the pedestrians in the dataset are sorted according to their similarity to the target pedestrian. Finally, the pedestrian image most similar to the target pedestrian is selected, and the recognition accuracy of the model is calculated.
[0014] Among them, the dataset Market-1501 selected in step (1) is specifically divided as follows:
[0015]
[0016] Among them, the dataset CUHK03 selected in step (1) is specifically divided as follows:
[0017]
[0018] Among them, in the data preprocessing of step (1):
[0019] (1-1) The first step of data preprocessing is to resize the input image size;
[0020] (1-2) Perform data normalization, scale the pixel values of the image to between 0 and 1, or convert the pixel values to a standard normal distribution with a mean of 0 and a variance of 1. Normalization can eliminate the order-of-magnitude differences between input features, avoid adverse effects of data size on model parameter optimization, and improve the convergence speed and stability of the model;
[0021] (1-3) Perform a random horizontal flip operation, flip the image horizontally along the horizontal axis with a certain probability. This operation simulates the symmetry feature changes in the real scene, helps to enhance the diversity of data, and improves the model's recognition ability for targets in different directions. The flipping probability can be flexibly set according to task requirements;
[0022] (1-4) Perform a random erasing operation. Erase a rectangular area at a random position in the image and fill the area with random values. This operation forces the model to focus on other part features of the image during training instead of relying on certain fixed patterns, and helps to improve the model's robustness to occlusion, noise, etc.;
[0023] (1-5) Perform a random rotation operation. Randomly rotate the input image within a certain angle range, for example, generate a new rotated image between -30 degrees and 30 degrees. Random rotation can simulate the image orientation changes in the real scene and improve the model's adaptability to targets at different angles. The rotated image needs to be padded or cropped at the edges to match the target size;
[0024] (1-6) Perform a random flipping operation. Flip the image horizontally and vertically with a certain probability respectively. This operation can further increase the diversity of data and simulate the distribution of targets in different spatial directions in reality. The choice of flipping probability needs to be adjusted according to the data augmentation requirements of the task;
[0025] (1-7) Perform a random brightness adjustment operation. Randomly increase or decrease the brightness value of the input image to simulate the shooting effects under different lighting conditions. Brightness adjustment can enhance the robustness of the model to lighting changes and prevent the model from overfitting to specific lighting conditions during training. The adjusted brightness range should be controlled within a reasonable interval to maintain the recognizability of the image.
[0026] Among them, the construction and initialization of the backbone network and the branch network in step (2):
[0027] (2-1) ResNet consists of multiple convolutional layers, including multiple residual blocks. Each residual block includes a skip connection that directly adds the input to the output, thus retaining the original feature information. This design idea not only makes the network easier to optimize but also alleviates the problems of gradient vanishing and gradient explosion, improving the accuracy of the network. ResNet-50 is used in this method, and the stride of the last convolutional layer is set to 1 to meet the needs of the person re-identification task. The branch network is constructed using the basic units of ResNet, and each basic unit consists of a convolutional filter, a batch normalization layer (BN), and a rectified linear unit (ReLU);
[0028] (2-2) The pre-trained model for initializing the backbone network comes from the ImageNet dataset. ResNe50 has been pre-trained on ImageNet and can then be trained on other datasets for specific tasks through fine-tuning or transfer learning.
[0029] Among them, in step (3):
[0030] (3-1) After the backbone network extracts features, the low-level features and high-level features are input into the residual expansion and fusion module to reduce feature loss through expansion, so that the low-level and high-level features are fused to the same size;
[0031] (3-2) Use the triplet loss function to optimize the global features, and use the hard sample mining triplet loss function and the center loss function to optimize the combined features. Among them, the triplet loss is shown as follows:
[0032]
[0033] Among them, N tp represents the number of available triplets, δ is a hyperparameter used to represent the interval, and [·] + represents the hinge loss function;
[0034] Center Loss learns a center vector for each category, so that all sample points in the same category are as close as possible to this center, while the centers of different categories are as far away from each other as possible:
[0035]
[0036] Among them, it is worth noting that c i is the center of all sample Features of category y i ;
[0037] (3-3) In summary, the total loss of the residual augmentation fusion branch network for person re-identification is as follows:
[0038] L = AL t + BL c
[0039] where A and B are the weight coefficients of each loss function.
[0040] The beneficial effects of the present invention are as follows: The present invention constructs a residual augmentation fusion method for person re-identification, generates copies of the respective input feature vectors for each network layer using residual units, first expands the copies generated by the lower network layers in the vertical and horizontal directions, and then fuses them with the copies generated by the adjacent higher network layers at the same size. In order to reduce the loss of important features of the original layer during feature fusion and be able to match and identify pedestrians more accurately. Description of the Drawings
[0041] Figure 1 is the person recognition flowchart of the present invention;
[0042] Figure 2 is the overall network model diagram of the present invention;
[0043] Figure 3 is the residual augmentation fusion structure diagram of the present invention; Detailed Embodiment
[0044] The embodiments of the present invention will be described in detail below with reference to the drawings.
[0045] A residual augmentation fusion method for person re-identification according to the present invention, its feature extraction and recognition process is as Figure 1 shown, and the detailed structure of the residual augmentation fusion for person re-identification is as Figure 2 shown, including the following steps:
[0046] (1) Dataset selection: The CUHK 03 dataset selected from the pedestrian re-identification dataset for experiments in this invention is divided into two subsets, CUHK 03-Labeled and CUHK 03-Detected. Among them, the CUHK 03-Labeled dataset contains a training set consisting of 767 pedestrians and 7368 images, and a test set consisting of another 700 pedestrians, 1400 query images and 5328 gallery images. The CUHK 03-Detected dataset contains a training set of 767 pedestrians and 7365 images, and a test set of another 700 pedestrians, which contains 1400 query images and 5332 gallery images. Market-1501 is a dataset for pedestrian re-identification and is one of the largest and most realistic pedestrian re-identification task datasets collected on the campus of Tsinghua University. This dataset contains 32668 images depicting 1501 pedestrians and is divided into a training set and a test set, consisting of 12936 and 19732 images depicting 751 and 750 people respectively;
[0047] (2) Training data preprocessing: In this invention, the training data is scaled to 288x288, and then the images are subjected to data normalization, random horizontal flipping, random erasing, random rotation, random flipping, and random brightness adjustment. Finally, normalization is performed using the mean and standard deviation of ImageNet;
[0048] (3) Model loading: Use the ResNet-50 pre-trained model trained from the ImageNet dataset to initialize the backbone network.
[0049] (4) Construct a residual augmentation fusion module:
[0050] Specifically, first, a residual module is designed in the horizontal connection direction. This module consists of a CBR3x3 unit, a CBR1x1 unit, and a residual operation. The CBR unit is equivalent to a network layer and consists of a convolutional filter, a batch normalization layer (BN), and a rectified linear unit (ReLU). The input feature vector Xi of each layer passes through this module to generate a copy Yi that retains the discriminative features learned by this network layer. Among them, the weights in the two CBR units are shared, which helps the copy capture important regions in the image. As shown in the following formula.
[0051] CBR = ReLU(BN(C i (X i )))
[0052] Y i = RS(CBR 3x3 (X i ), CBR 1x1 (Xi ))
[0053] Then, utilize the characteristics of the feature maps extracted by ResNet50 in the upsampling direction: the resolution of the low-level feature maps is four times that of the adjacent high-level feature maps, and perform two-way expansion on the high-level feature maps, enabling the high-level features to be fused with the adjacent low-level features from top to bottom. The expansion method is shown in the following formula.
[0054] BE = (1 - α)(1 - β)I(x 1 ,y 1 ) + α(1 - β)I(x 2 ,y 1 ) + (1 - α)βI(x 1 ,y 2 ) + αβI(x 2 ,y 2 )
[0055]
[0056] Among them, I(x, y) is the target pixel point to be expanded, α and β are the ratios of the horizontal and vertical distances of the position I(x, y) from the four discrete points, (x 1 ,y 1 ), (x 1 ,y 2 ) and (x 2 ,y 1 ), (x 2 ,y 2 ) represent the four discrete points in the horizontal and vertical directions.
[0057] Finally, fuse the high-level feature maps with the adjacent low-level feature maps at the same size. The fusion method is shown in the following formula.
[0058]
[0059] Among them represents the fusion operation, g(·) represents global average pooling, and are the learnable parameters in the two fully connected layers, followed by the Sigmoid activation function σ(·). The BEF module sequentially fuses the first four feature maps to obtain the fused feature F f .
[0060] (5) Construction of the loss function.
[0061] (5-1) The present invention selects to introduce Triplet Loss and Center Loss to train the model.
[0062] Triplet Loss: Assume I, Ip , I n are three samples, Sample I, I p are of the same pedestrian identity, Sample I, I n are of different pedestrian identities. The triplet loss is to find a new feature space: the distance from Sample I to I p is much smaller than the distance from Sample I to I n , as shown in the following formula:
[0063]
[0064] where, N tp represents the number of available triplets, δ is a hyperparameter used to represent the interval, and [·] + represents the hinge loss function.
[0065] Center Loss: Learn a center vector for each category so that all sample points in the same category are as close as possible to this center, while the centers of different categories are as far apart as possible from each other. It is usually expressed as shown in the following formula:
[0066]
[0067] where, it should be noted that c i is the center of all sample Features of the category y i .
[0068] (5-2) Train the model
[0069] Calculate the total loss and perform backpropagation to minimize this loss function, that is, obtain the model samples with the smallest distance from the positive samples and the largest distance from the negative samples, and obtain the optimized person re-identification model. After 120 iterations in the following environment, the trained model is obtained.
[0070]
[0071] (6) Test the model
[0072] (6-1) After enlarging the images in the test set from the original size of 64×128 to 128×384, normalize the images to ensure the consistency of data distribution and the stability of the model. Subsequently, load the trained person re-identification model into the predefined network architecture, and extract discriminative pedestrian feature vectors by fusing global features and local features. Finally, use the currently recognized objective evaluation metrics - mean average precision (mAP) and first hit rate (Rank-1) to evaluate the model performance to measure the effect of the model in the person re-identification task.
[0073] (6-2) The Cumulative Matching Characteristic (CMC) curve is a commonly used metric in retrieval tasks, mainly used to calculate the hit probability of Rank-k. Specifically, for a probe image to be retrieved in the query set of the test set from the gallery to be retrieved, a series of sorted results in the gallery of the test set are returned. The sorting is based on similarity. The Rank-k accuracy A can be expressed as:
[0074]
[0075] where p is the probe, q is the query, and N q is the number of images, and f CMC is as follows:
[0076]
[0077] In practical applications, representative Rank-1, Rank-5, and Rank-10 accuracies are usually selected to replace the CMC curve. Among them, the Rank-1 accuracy is the most important metric. The CMC curve is plotted with k on the x-axis and the Rank-k accuracy A on the y-axis. The Rank-1 accuracy directly reflects the probability that the target object is the first in the retrieval results and is a key metric for measuring the performance of the retrieval system.
[0078] (6-3) Mean Average Precision (mAP) is another important evaluation metric. The CMC curve usually only cares about the ranking of the top positive samples in the gallery, while mAP is determined by the ranking results of all positive samples in the gallery. Therefore, it can usually more robustly reflect the performance of the model. Calculating AP requires the following three steps:
[0079] (6-3-1) Precision: For a probe image q in the query j , a series of sorted results in the gallery are returned. Considering the first n query results, assume the number of results with the same pedestrian ID as the probe image among the first n results is c(n, q j ):
[0080]
[0081] (6-3-2) Average Precision: For the probe image q of the query j , record the set {i 1 , i 2 ,...., i m} of the ranking results of all M positive samples. Calculate their average Precision as shown in the following formula:
[0082]
[0083] (6 - 3 - 3) Mean Average Precision (mAP): For all N q The mean of the Average Precision of all probe images, that is:
[0084]
[0085] The mAP metric is a key parameter for measuring the performance of a retrieval system. The higher its value, the better the performance of the retrieval system. When the performance of the retrieval system reaches its extreme, that is, it can accurately retrieve all positive samples and these samples rank first in the ranking, the value of the mAP metric will reach the maximum value of 1. As the number of samples in the gallery image library increases, the mAP metric can better reflect the stability and reliability of the retrieval system when facing complex data, that is, robustness. However, if the number of samples in the gallery image library is small, the reference value of the mAP metric will be greatly reduced. In view of this, in order to more comprehensively evaluate the accuracy of the person re - identification model, the CMC curve and the mAP metric are generally considered comprehensively.
[0086] Finally, the network model was optimized using the above - mentioned method, and the trained model was tested using the above - mentioned evaluation metrics. The experiment was carried out on the Market - 1501 dataset. The experiment proves that a residual expansion and fusion method for person re - identification effectively improves the model recognition results, and has made remarkable progress in both the mAP and Rank - 1 evaluation directions.
[0087] The above embodiments have introduced in detail the specific implementation manners of a residual expansion and fusion method for person re - identification proposed by the present invention. The introduction of the above embodiments is only used to help understand the proposed method and the core idea of the present invention. There may be some differences in the specific implementation manners according to the idea of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A residual expansion fusion method for pedestrian re-identification, characterized in that: Includes steps: (1) Data preprocessing: Select the data set required for the experiment from the person re-ID data set, divide the data, and perform preprocessing operations; (2) Constructing the main network structure: The backbone network is constructed based on ResNet-50 and initialized using the pre-trained model. Each network layer uses residual units to generate copies of their own feature vectors, and the copies generated by the lower network layers are expanded vertically and horizontally, and then fused with the copies generated by the adjacent higher-level network layers at the same size to reduce the loss of important features of the original layer during feature fusion. (3) Model training: First, the batch-sampled pedestrian samples are preprocessed and input into a predefined model to extract the high-order features of pedestrians. Then, the low-level features extracted by the backbone network and the high-level features are input into the residual expansion fusion module to reduce the feature loss through expansion, so that the high-level and low-level features are fused at the same size. The modules are input into the loss function to calculate the loss and perform back propagation, update the model parameters, and continuously iterate to minimize the value of the loss function to form an optimized pedestrian re-identification model. (4) Model testing: After training, the proposed network model is used as a pedestrian feature extractor, using pedestrian features containing identity information to represent each pedestrian. By calculating the Euclidean distance between the target pedestrian and the feature vectors of each pedestrian in the test dataset, the pedestrians in the dataset are sorted according to their similarity to the target pedestrian. Finally, the pedestrian image most similar to the target pedestrian is selected, and the recognition accuracy of the model is calculated.
2. The residual expansion and fusion method for person re-identification according to claim 1, characterized in that: Step (1) preprocesses the original pedestrian image dataset using data enhancement techniques such as data normalization, random horizontal flipping, random erasing, random rotation, random flipping, and random brightness adjustment.
3. The residual expansion and fusion method for pedestrian re-identification according to claim 1, characterized in that: (a) In step (2), the convolutional neural network defined uses ResNet-50 as the backbone network. Its structure contains four gradient mitigation stages. Each stage alleviates the gradient vanishing problem through jump connections, so that deeper networks can be trained. This design enables the network to effectively extract multi-level feature information while maintaining high computational efficiency, providing rich feature representation for subsequent tasks; (b) The residual expansion fusion module in step (2) is designed to alleviate the problem of information loss in the transmission of low-level features when fusion is performed. This paper uses a residual unit composed of two CBR units of different sizes to generate a copy of the input feature vector for each network layer through weight sharing, and minimizes the loss of important features of the original layer by retaining the copy.
4. The residual expansion and fusion method for pedestrian re-identification according to claim 1, characterized in that: The loss function L constructed in step (3) is as follows: <h2 style=";text-align:left;direction:ltr">L=AL<h2 style=";text-align:left;direction:ltr"> t <h2 style=";text-align:left;direction:ltr"> +BL<h2 style=";text-align:left;direction:ltr"> c Among them, N tp represents the number of available triplets, δ is a hyperparameter used to represent the interval, [·] + represents the hinge loss function, c i Is of category y i The center of all sample features, A and B are the weights of each loss function.
Citation Information
Patent Citations
Pedestrian re-identification method based on multi-attention joint multilevel features
CN112183468A
Spatial relation model pedestrian feature recognition method based on multi-region attention association
CN115188019A
Dressing pedestrian re-identification method based on image style migration and residual network
CN119479005A
Person re-identification method based on perspective-guided multi-adversarial attention
US20230267725A1
Person image re-identification method based on autonomous model structure evolution
WO2024093466A1