End-to-end pedestrian re-identification method based on index prediction

By adopting the encoder-decoder architecture and multi-label cross-entropy loss function in pedestrian re-identification technology, joint optimization of feature extraction, enhancement and sorting is achieved, solving the problems of inconsistent optimization goals and limited generalization capabilities in the existing technology, and significantly improving the performance and generalization capabilities of the model.

CN120126181AActive Publication Date: 2025-06-10NORTHWESTERN POLYTECHNICAL UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510285742.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing pedestrian re-identification technology has problems such as inconsistent optimization goals, limited error accumulation and generalization capabilities. Especially when dealing with multi-match images, the traditional cross-entropy loss function is not applicable.

Method used

The end-to-end pedestrian re-identification method based on index prediction is adopted, and the joint optimization of feature extraction, feature enhancement and sorting is achieved through the encoder-decoder architecture, and a customized multi-label cross-entropy loss function is designed to adapt to the situation of multiple matching images.

Benefits of technology

It significantly improves the performance and generalization ability of the pedestrian re-identification model, avoids error accumulation, and adapts to the multi-match image prediction problem in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126181A_ABST
    Figure CN120126181A_ABST
Patent Text Reader

Abstract

The invention discloses an end-to-end pedestrian re-identification method based on index prediction, which constructs a brand new ReID method, namely a StreamReID model, redefines a ReID task as an index prediction problem, and realizes joint optimization of feature extraction, feature enhancement and sorting steps by using an encoder-decoder architecture. And meanwhile, a customized multi-label cross entropy loss function is designed to adapt to the condition of multiple matching images, so that the problems in the prior art are effectively solved, and the performance and generalization ability of the ReID model are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an end-to-end pedestrian re-identification method based on index prediction. Background Art

[0002] Pedestrian re-identification (Person Re-Identification, ReID) is an important task in the field of computer vision, and its goal is to match images of the same pedestrian through a large-scale database. This technology has broad application prospects in the fields of public security, video surveillance, and intelligent transportation. With the wide deployment of surveillance cameras and the development of artificial intelligence technology, the ReID problem has received increasing attention.

[0003] Traditional methods usually decompose the ReID task into multiple key subtasks, including feature extraction, feature enhancement, and ranking. However, this method of decomposing the task into multiple independent subtasks has some problems. First, independently optimizing these subtasks may lead to inconsistent training and testing objectives. For example, the feature extraction stage may focus on learning discriminative features, but the ranking stage may be affected by the feature enhancement strategy, resulting in limited overall performance. Second, the feature enhancement and ranking stages usually rely on manually designed strategies, which may not be able to adapt to complex scene changes such as occlusion, illumination changes, or clothing changes, affecting the generalization and robustness of the model. Third, due to the lack of joint optimization between the various subtasks, errors in the previous subtask may be transmitted to the subsequent subtasks, resulting in a decline in overall performance.

[0004] There have also been some studies that attempt to optimize the ReID model as a whole. For example, some studies use an encoder-decoder architecture to handle the ReID task, but most of these methods regard the ranking problem as an independent step and do not jointly optimize it with the feature extraction and enhancement steps. In addition, when dealing with the situation where there are multiple matching images in the database, these methods usually face the problem of non-unique correct predictions, which makes the traditional cross-entropy loss function no longer applicable.

[0005] In the field of pedestrian re-identification, although existing studies have attempted to use an encoder-decoder architecture to handle the ReID task for overall optimization, most of these methods regard the ranking problem as an independent step and do not jointly optimize it with the feature extraction and enhancement steps. This separate processing method leads to problems such as inconsistent optimization objectives, error accumulation, and limited generalization ability. In addition, existing methods usually face the problem of non-unique correct predictions when dealing with the situation where there are multiple matching images in the database, which makes the traditional cross-entropy loss function no longer applicable, further restricting the optimization and performance improvement of the model. Summary of the Invention

[0006] To overcome the deficiencies of the prior art, the present invention provides an end-to-end pedestrian re-identification method based on index prediction, constructs a brand-new ReID method - the StreamReID model, redefines the ReID task as an index prediction problem, and uses an encoder-decoder architecture to jointly optimize the feature extraction, feature enhancement, and ranking steps. At the same time, a customized multi-label cross-entropy loss function is designed to adapt to the situation of multi-matching images, thereby effectively solving the problems existing in the prior art and significantly improving the performance and generalization ability of the ReID model.

[0007] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0008] Step 1: Hard sampling strategy;

[0009] Use the ViT model as the feature extractor;

[0010] Input the sample image into the ViT model to extract features;

[0011] Resample the extracted features for training the subsequent encoder-decoder module;

[0012] Step 2: Construct an encoder for realizing relational information interaction;

[0013] Step 2-1: Perform spatial average pooling on the image features extracted from the feature extractor to compress them into one-dimensional vectors; the pooled query features and database features are respectively represented as F = {f q , f d0 , f d1 ,..., f di ,...., f dn}, where f q represents the query feature, f di represents the database feature, and i is the index of the feature in the database; the database feature f di and the query feature where C is the channel dimension;

[0014] Step 2-2: Input the pooled features into the encoder module, which consists of 6 Transformer layers, and each layer contains a self-attention layer followed by a feed-forward layer; for the input feature F, the self-attention mechanism is expressed as:

[0015] Projection: Q = FC Q (F), K = FC K (F), V = FC V (F)

[0016]

[0017] Aggregation: F a = S`V(1)

[0018] where FC Q , FC K , FC V all represent fully connected layers, used to embed the input features into the query Q, key K, and value V vectors respectively; d represents the dimension of the feature vector; S is the calculated affinity metric, reflecting the relationship between the compared features; is the output feature;

[0019] Then send F a to the feed-forward layer;

[0020] Step 2-3: Repeat Step 2-2 six times to obtain the features enhanced by the encoder, denoted as F en ;

[0021] Step 3: Construct a decoder for index prediction;

[0022] Step 3-1: Aggregate database image features Construct a dynamic codebook through a set of learnable position embeddings L = {l 0 , l 1 ,..., l n , l End}, and this process is formally expressed as:

[0023]

[0024] where φ = {φ 0 , φ 1 ,..., φ n , φ End} is the constructed dynamic codebook, n is the number of images in the database, and φ End is the end marker;

[0025] Step 3-2: Take the query feature as the start marker, denoted as Then send it to the decoder module, which is composed of two layers of Transformer, and this process is written as:

[0026] D 0 = Decoder(H 0 )(3)

[0027] where is the output of the decoder module;

[0028] Step 3-3: Based on S 0The correlation between the last element in and each index feature in the dynamic codebook predicts the index position of the matching element, and this process is written as:

[0029] ψ 0 = softmax(η(S 0 )·φ) (4)

[0030] where η is a function for selecting the last element in the set S 0 ; the index of the maximum value in the vector ψ indicates the position of the currently predicted matching image;

[0031] Step 3-4: Assume the predicted position is j, and concatenate the j-th element in the dynamic codebook with H 0 to construct the input for the next round After that, input H 1 into the decoder, and repeatedly apply equations (3) and (4) to iteratively generate the index positions of all matching images;

[0032] Step 4: Loss function;

[0033] Adopt a multi-label cross-entropy loss function, written as:

[0034]

[0035] where ψ i is the i-th element in ψ; pos represents the indexes of all matching images in the database;

[0036] By using the above loss function, the decoder model can be trained to predict the indexes of any matching images in each round.

[0037] Preferably, step 1 is specifically:

[0038] Randomly select an image feature as a query and construct a database for it; specifically, use a memory bank to retain the features of the first 10,000 images, and select 100 images that are most similar to the query but have different identities from them to construct a database; next, randomly insert the image features with the same identity as the query in the current mini-batch into different positions in the database; the constructed database and the query feature are used as the input of the encoder-decoder model; finally, train the encoder-decoder model to predict the position of the matching image.

[0039] Preferably, when training the decoder model: to avoid the decoder from making repeated predictions, discard the tokens corresponding to the previous prediction indices from the codebook after each round. Once all the index tokens corresponding to the matching elements are discarded, the decoder is trained to predict the index of the end token; during the training phase, in the decoder module, arrange all the matching image index tokens in the dynamic codebook in ascending order and input them into the decoder in parallel. Apply the masked attention mechanism to prevent forward tokens from accessing subsequent tokens, thereby enabling parallel training of the decoder.

[0040] A computer program that causes a computer to execute the above-mentioned end-to-end person re-identification method.

[0041] An electronic device, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned end-to-end person re-identification method.

[0042] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned end-to-end person re-identification method is implemented.

[0043] A chip, comprising: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the above-mentioned end-to-end person re-identification method.

[0044] A computer program product, the computer program product includes a computer storage medium, the computer storage medium stores a computer program, the computer program includes instructions that can be executed by at least one processor, and when the instructions are executed by the at least one processor, the above-mentioned end-to-end person re-identification method is implemented.

[0045] The beneficial effects of the present invention are as follows:

[0046] The present invention introduces a hard sampling strategy in terms of input data, constructs more challenging training samples, and significantly improves the generalization ability of the model; in terms of the model structure, it adopts an encoder-decoder architecture, uses the self-attention mechanism of Transformer for feature enhancement, and redefines the sorting problem as an index prediction task to achieve the joint optimization of feature extraction, enhancement, and sorting; during the training process, a customized multi-label cross-entropy loss function is designed to adapt to the prediction problem of multiple matching images. At the same time, an end-to-end training method is adopted to avoid error accumulation and improve the overall performance. In addition, the present invention also optimizes the self-attention mechanism formula and the multi-label cross-entropy loss function formula to further enhance the feature discrimination and the optimization ability of the model. Description of the Drawings

[0047] Figure 1Schematic diagram of the StreamReID model of the present invention. Detailed implementation manners

[0048] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0049] The present invention proposes a new method of integrating the ReID task into an end-to-end framework that can uniformly optimize the feature extraction, feature enhancement, and ranking steps, aiming to improve the performance of the ReID model through a unified optimization objective.

[0050] The present invention proposes a brand-new ReID method - the StreamReID model. By redefining the ReID task as an index prediction problem and using an encoder-decoder architecture to jointly optimize the feature extraction, feature enhancement, and ranking steps, and designing a customized multi-label cross-entropy loss function to adapt to the case of multi-matching images, the problems existing in the prior art are effectively solved, and the performance and generalization ability of the ReID model are significantly improved. Its structure is as Figure 1 shown. The main work of the present invention includes three components: a feature extractor, a Transformer encoder module, and a Transformer decoder module.

[0051] 1. Hard sampling strategy;

[0052] In each mini-batch, the present invention extracts 128 images, including 32 pedestrian instances, with 4 images for each instance. These images are then fed into the ViT model to extract basic features. After that, the extracted features are resampled to train the subsequent encoder-decoder module. In this step, a random image feature is selected as the query, and a challenging database is constructed for it. Specifically, a memory bank is used to retain the features of the first 10,000 images, and 100 images that are most similar to the query but have different identities are selected from it to construct the database. Next, the image features with the same identity as the query in the current mini-batch are randomly inserted into different positions in the database. The constructed database and the query feature are used as the input of the encoder-decoder model. Finally, the encoder-decoder model is trained to predict the matching image positions.

[0053] 2. Encoder: realizing relationship information interaction;

[0054] In the StreamReID model, the self-attention mechanism of the encoder module is used to capture and exchange the relationship information between the query image feature and the database image feature, thereby enhancing the distinctiveness of the features. Different from some existing works that use graph neural networks based on manual strategies (such as adjacency matrices) to model relationship information, the adopted Transformer architecture does not require such strategies, thus enabling more effective optimization.

[0055] Actually, first, spatial average pooling is performed on the image features extracted from the feature extractor (ViT) to compress them into one-dimensional vectors. The pooled query features and database features are respectively represented as F = {f q , f d0 , f d1 ,..., f dn}, where f q represents the query feature, while f di represents the database feature, and i is the index of the feature in the database. The database feature f di and the query feature where C is the channel dimension. These pooled features are then fed into an encoder module, which consists of 6 Transformer layers, each containing a self-attention layer followed by a feed-forward layer. The self-attention mechanism within the Transformer layer can capture and exchange the relationship information between the compared features. For the input feature F, the self-attention mechanism can be expressed as:

[0056] Projection: Q = FC Q (F), K = FC K (F), V = FC V (F)

[0057]

[0058] Aggregation: F a = S`V

[0059] where FC Q / K / V represents the fully connected layer used to embed the input features into query (Q), key (K), and value (V) vectors; S is the calculated affinity metric reflecting the relationship between the compared features; is the output feature. Thereafter, F a is fed into a feed-forward layer. By repeating the above operations 6 times, the features enhanced by the encoder are obtained, denoted as F en .

[0060] 3. Decoder: Realize index prediction;

[0061] In the model of the present invention, a decoder module is used to redefine the sorting problem as a prediction task for the index positions of matching images in the database. This method can achieve the integration and joint optimization of the sorting process. In this step, the enhanced query features and database features F enIt is input into the decoder model to train it to generatively predict the index position of the database image with the same identity as the query. To establish the association between the database image and its index, a dynamic codebook is first constructed for the decoder based on the database features.

[0062] Specifically, given the enhanced feature F en , first aggregate the database image features through a set of learnable position embeddings L = {l 0 , l 1 ,..., l n , l End} to construct the dynamic codebook. This process can be formally expressed as:

[0063]

[0064] where φ = {φ 0 , φ 1 ,..., φ n , φ End} is the constructed dynamic codebook, n is the number of images in the database, and φ End is an end marker. Subsequently, to initialize the input of the decoder module, the query feature is used as the start marker, denoted as and then it is fed into the decoder module, which consists of two layers of Transformer. This process can be written as:

[0065] S 0 = Decoder(H 0 ) (3)

[0066] where is the output of the decoder module. Then, based on the correlation between the last element in S 0 and each index feature in the dynamic codebook, the index position of the matching element is predicted. This process can be written as:

[0067] ψ 0 = softmax(η(S 0 )·φ) (4)

[0068] where η is a function for selecting the last element in the set S 0 ; the index of the maximum value in the vector ψ indicates the position of the currently predicted matching image. Assuming the predicted position is j, then the j-th element in the codebook is concatenated with H 0 to construct the input for the next round After that, H 1 is input into the decoder, and equations (3) and (4) are repeatedly applied to iteratively generate the index positions of all matching images.

[0069] During the training process, it was noticed that at each step, there were multiple database images with the same identity as the query image, resulting in multiple potential correct prediction results. This made it impossible for the standard cross-entropy loss function (which assumes a unique prediction result) to effectively train the model. To address this issue, a customized multi-label cross-entropy loss function was proposed, which can be written as:

[0070]

[0071] where ψ i is the i-th element in ψ; pos represents the indices of all matching images in the database. By using this loss function, the decoder model can be trained to predict the indices of any matching image in each round. To avoid the decoder from making repeated predictions, the tokens corresponding to the previously predicted indices are discarded from the codebook after each round. Once all the index tokens corresponding to the matching elements are discarded, the decoder is trained to predict the index of the end token. Additionally, during the training phase, in the decoder module, all the matching image index tokens in the dynamic codebook are sorted in ascending order and fed into the decoder in parallel. The masked attention mechanism is applied to prevent the forward tokens from accessing the subsequent tokens, thus enabling the parallel training of the decoder.

[0072] Examples:

[0073] 1. Dataset selection;

[0074] In terms of dataset setting, the present invention first uses three widely used person re-identification datasets, namely Market-1501, DukeMTMC-reID, and MSMT17. Additionally, an occluded person re-identification dataset Occluded Duke is incorporated to evaluate the robustness of the model of the present invention to occlusion situations. Moreover, the present invention also uses two clothing change datasets PRCC and LTCC to evaluate the performance of the model under different dressing conditions. This comprehensive evaluation can better understand the robustness of the method of the present invention in different scenarios.

[0075] 2. Implementation detail setting;

[0076] The present invention adopts the visual encoder of the CLIP model as the feature extractor and trains the model for 60 epochs with a batch size of 128. During the training and testing processes, the images of each person are resized to the size of (256×128) to ensure consistency and facilitate direct comparison with previous works. The learning rate is set to 3.5×10 -4. During training, random horizontal flipping is performed with a probability of 0.5, and data augmentation is performed using random erasing. To evaluate the trained model, the present invention adopts the mainstream CMC and mAP metrics. All experiments are carried out on a single NVIDIA RTX 4090 GPU, which provides sufficient computing power to handle large-scale training and evaluation processes, ensuring that the model converges efficiently within a reasonable time.

[0077] 3. Training and inference steps;

[0078] In the training phase, the present invention first feeds the input image into the feature extractor to extract basic features. In this step, a metric loss such as triplet loss is adopted to train the feature extractor. The loss function used for training the extractor is expressed as Subsequently, the extracted features are resampled, and the resampled features are fed into the encoder model. Then, an identity classification loss is applied Enhanced features are generated through the encoder module, guiding it to capture effective relationship information and enhance the discriminability of the extracted features. Then, the enhanced features are input into the decoder module and trained using MCE loss. The overall loss function of StreamReID can be written as:

[0079]

[0080] In the testing phase, to ensure inference efficiency, the present invention uses the feature extractor of the StreamReID model to extract basic features for query images and database images. Based on their cosine similarity, 100 most difficult-to-distinguish candidate objects are initially selected for each query, and then these candidate objects are input into the StreamReID model together with the query image for prediction. In the decoder module, the query feature serves as the initial token, and the previously predicted index tokens are gradually removed from the codebook in each round to avoid repeated predictions. In this step, the end token is discarded, and the decoder is required to continuously predict until the entire list of candidate objects is traversed, so as to construct the final ranking list according to the prediction order.

[0081] 4. Implementation environment;

[0082] The present invention is implemented using a Sugon-W580-G20 server and a Linux Ubuntu 16.04.4 LTS operating system. One NVIDIA GeForce GTX 3090Ti graphics processor is used for training, and the NVIDIA CUDA 10.2 platform is used to accelerate training. In terms of software configuration, Python 3.8.13 (GCC 7.5.0) and PyTorch 1.11.0, as well as dependency libraries such as Numpy 1.19.2 and Pillow 9.0.2 are used.

Claims

1. An end-to-end person re-identification method based on index prediction, characterized in that: The steps include: Step 1: Hard sampling strategy; The ViT model is used as the feature extractor; Input the sample image into the ViT model to extract features; Resample the extracted features for training subsequent encoder-decoder modules; Step 2: Construct an encoder to realize the interaction of relational information; Step 2-1: Perform spatial average pooling on the image features extracted from the feature extractor and compress them into a one-dimensional vector; the query features and database features after pooling are represented as F = {f q ,f d0 ,f d1 ,...,f di ,....,f dn }, where f q represents the query feature, f di Represents a database feature, i is the index of the feature in the database; Database features di and query features Where C is the channel dimension; Step 2-2: Input the pooled feature into the encoder module, which consists of 6 Transformer layers, each of which contains a self-attention layer followed by a feed-forward layer; for the input feature F, the self-attention mechanism is expressed as: Projection:Q=FC Q (F),K=FC K (F),V=FC V (F) Aggregation:F a =S`V (1) Among them FC Q , FC K , FC V Both represent fully connected layers, which are used to embed input features into query Q, key K, and value V vectors respectively; d represents the dimension of the feature vector; S is the calculated affinity measure, which reflects the relationship between the compared features; is the output feature; Then F a is fed into the feed-forward layer; Step 2-3: Repeat step 2-2 six times to obtain the features enhanced by the encoder, denoted as F en ; Step 3: Build a decoder to implement index prediction; Step 3-1: Aggregate database image features Through a set of learnable position embeddings L = {l0,l1,...,l n ,l End }Construct a dynamic codebook. This process is formally expressed as: Where φ={φ0,φ1,...,φ n ,φ End } is the constructed dynamic codebook, n is the number of images in the database, and φ End is the end marker; Step 3-2: Use the query feature as the starting tag, denoted as Then it is sent to the decoder module, which consists of two layers of Transformer. This process is written as: S 0 =Decoder(H 0 ) (3) in is the output of the decoder module; Step 3-3: Based on S 0 The correlation between the last element in and each index feature in the dynamic codebook predicts the index position of the matching element. This process is written as: ψ 0 =softmax(η(S 0 )·φ) (4) Where η is a function S used to select the last element in a set 0 ; The index of the maximum value in the vector ψ indicates the location of the current predicted matching image; Step 3-4: Assume that the predicted position is j, and compare the jth element in the dynamic codebook with H 0 Connect them together to build the next round of input Afterwards, H 1 Input the decoder and repeatedly apply equations (3) and (4) to iteratively generate the index positions of all matching images; Step 4: Loss function; Using the multi-label cross entropy loss function, written as: where ψ i is the i-th element in ψ; pos represents the index of all matching images in the database; By using the above loss function, the decoder model can be trained to predict the index of any matching image at each round.

2. The end-to-end person re-identification method based on index prediction according to claim 1, characterized in that: The step 1 is specifically as follows: Randomly select an image feature as the query and build a database for it; specifically, use a memory bank to retain the features of the first 10,000 images, and select the 100 images that are most similar to the query but have different identities to build the database; next, randomly insert the image features with the same identity as the query in the current mini-batch into different positions in the database; the constructed database is used as the input of the encoder-decoder model together with the query features; finally, the encoder-decoder model is trained to predict the matching image positions.

3. The end-to-end person re-identification method based on index prediction according to claim 2, characterized in that: When training the decoder model: to avoid repeated predictions by the decoder, the tokens corresponding to the previously predicted index are discarded from the codebook after each round, and once all index tokens corresponding to matching elements are discarded, the decoder is trained to predict the index of the end token; in the training phase, in the decoder module, all matching image index tokens in the dynamic codebook are arranged in ascending order and input into the decoder in parallel; A masked attention mechanism is applied to prevent the forward token from accessing the subsequent token, thus enabling parallel training of the decoder.

4. A computer program, characterized in that The computer program enables a computer to execute the method according to any one of claims 1 to 3.

5. An electronic device, characterized in that: include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method as claimed in any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.

7. A chip, characterized in that: include: A processor, configured to call and run a computer program from a memory, so that a device equipped with the chip executes a method as claimed in any one of claims 1 to 3.

8. A computer program product, characterized in that The computer program product comprises a computer storage medium storing a computer program, wherein the computer program comprises instructions executable by at least one processor, and when the instructions are executed by the at least one processor, the method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Pedestrian re-identification method based on contrast language image pre-training model CLIP

    CN115393902A

  • Pedestrian re-identification method based on relation attention transformer

    CN116012771A

  • Pedestrian re-identification method

    CN116824469A

  • Pedestrian search method and device, equipment and storage medium

    CN119007243A

  • Athlete style recognition system and method

    US20200394413A1