Video pedestrian re-identification method based on space-time domain feature extraction and adaptive aggregation

By constructing a video pedestrian recognition model based on space-time domain feature extraction and adaptive aggregation, and using attention mechanism to perform feature aggregation in the air and time domains, the problem of difficult to extract discriminant and robust video-level features in the prior art is solved, and higher video pedestrian recognition accuracy and stability are achieved.

CN120071389AActive Publication Date: 2025-05-30CHINA ACAD OF CIVIL AVIATION SCI & TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510075234.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-30
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract discriminant and robust video-level features in video pedestrian recognition, especially in capturing cross-frame dependencies and handling frame occlusion problems.

Method used

Using a method based on space-time domain feature extraction and adaptive aggregation, a video pedestrian re-identification model consisting of a backbone network, a spatial feature aggregation module (SFA) and a timing feature aggregation module (TFA) is constructed. This model uses attention mechanism to perform feature aggregation in the airspace and time domains to generate discriminant and robust video-level features.

Benefits of technology

It significantly improves the discriminant and robustness of video-level representations, effectively deals with challenges such as frame occlusion and small differences between classes, and improves the accuracy and stability of video pedestrian re-identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071389A_ABST
    Figure CN120071389A_ABST
Patent Text Reader

Abstract

The invention discloses a video pedestrian re-identification method based on space-time domain feature extraction and adaptive aggregation, and the method comprises the following steps: constructing a video pedestrian re-identification model; extracting an initial frame-level feature map through the backbone network; acquiring local features and global features by a spatial feature aggregation module, and extracting frame-level features fusing the local features and the global features through a feature aggregation module; the time domain feature aggregation module is used for aggregating frame-level features based on an attention mechanism so as to extract video-level features used for video re-identification, and through the process, an end-to-end video pedestrian re-identification process is realized. Compared with the prior art, the method has the advantages that two core modules, namely the spatial feature aggregation module and the time sequence feature aggregation module, are designed, video-level features with discrimination and robustness can be generated, challenges such as frame shielding and small inter-class difference can be effectively handled, and discrimination and robustness of video-level representation can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video pedestrian re-identification in computer vision, and particularly to a video pedestrian re-identification method based on spatio-temporal feature extraction and adaptive aggregation. Background Art

[0002] Video-based pedestrian re-identification is a crucial and challenging task in the field of computer vision, which aims to associate similar pedestrians from a set of non-overlapping camera views. Since pedestrian re-identification has the characteristics of being not restricted by recognition distance and environment, it can be widely used in social security fields such as video surveillance, intelligent security, and criminal investigation, and has received extensive attention from researchers. Currently, image-based pedestrian re-identification methods mainly focus on extracting appearance features of single frames. Since the feature information that a single frame can provide is limited, it is difficult to capture discriminative pedestrian appearance features only relying on a single frame. At this time, a set of continuous multi-frame video segments with temporal complementarity is particularly important, which can not only provide rich pedestrian appearance information, but also reveal the relationship between motion, perspective change and identity appearance change, thereby alleviating the limitations of image pedestrian re-identification and providing more representative and discriminative video-level features.

[0003] Currently, video-based pedestrian re-identification research mainly focuses on extracting discriminative features of pedestrians. For example, temporal modeling methods capture temporal consistency features by modeling the correlation between frames, and feature aggregation methods extract effective spatio-temporal features to generate robust video-level representations. However, extracting discriminative appearance features is still a challenging task. Temporal feature modeling work has limited ability to capture long-distance dependencies and is difficult to effectively model cross-frame dependencies. For example, 3D convolution captures dynamic information by expanding the temporal dimension, and its computational complexity limits the ability to model cross-frame dependencies; while the optical flow method relies on accurate motion estimation and is easily affected by frame occlusion problems, resulting in unsatisfactory dynamic feature extraction effects. Feature aggregation methods focus on extracting spatial detail features and can improve the representativeness of video-level features to a certain extent. Due to the lack of the ability to model inter-frame changes and inter-frame interactions, these methods are difficult to capture static local details, thus limiting the discriminability and robustness of video-level representations. Therefore, there are still challenges in extracting discriminative and robust video-level representations with existing methods. Summary of the Invention

[0004] The purpose of the present invention is to provide a video pedestrian re-identification method based on spatio-temporal feature extraction and adaptive aggregation, which solves the above problems.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is: a video pedestrian re-identification method based on spatio-temporal feature extraction and adaptive aggregation, and the method steps are as follows:

[0006] Step 1: Construct a video person re-identification model, which consists of three main modules: a backbone network, a spatial feature aggregation module SFA, and a temporal feature aggregation module TFA;

[0007] Step 2: Given an input video clip, extract initial frame-level feature maps through the backbone network, including a shallow feature map S t and a deep feature map D t ; The backbone network is based on the ResNeSt50 neural network model for classifying and extracting feature images, obtaining a shallow feature map S t and a deep feature map D t ;

[0008] Step 3: Input the shallow feature map S t and the deep feature map D t into the spatial feature aggregation module SFA respectively. The local and global feature extraction modules in SFA are used to obtain local features and global feature g t , and extract frame-level feature Z that fuses local and global features through a feature aggregation module based on the attention mechanism SA ;

[0009] Step 4: Using the frame-level feature Z SA as the input, the temporal feature aggregation module is based on the attention mechanism to aggregate frame-level features to extract video-level features z TA for video re-identification. Through this process, an end-to-end video person re-identification process is achieved.

[0010] Preferably, in Step 1, the spatial feature aggregation module SFA includes a local feature extraction module LFE, a global feature extraction module GFE, and a feature fusion module AFA based on the attention mechanism,

[0011] The local feature extraction module LFE detects local significant regions through the attention mechanism and extracts local feature descriptions with geometric characteristics from the shallow features extracted by the backbone network;

[0012] The global feature extraction module GFE extracts global feature descriptions with geometric characteristics from the deep features extracted by the backbone network through generalized mean pooling;

[0013] The feature fusion module AFA is based on the attention mechanism, uses a self-attention module to fuse local and global features, and obtains frame-level feature descriptions with local and global semantics.

[0014] Preferably, the method for the local feature extraction module LFE to extract local feature descriptions with geometric characteristics is as follows:

[0015] Denote the shallow feature of the t-th frame image extracted by the backbone network as Then the local feature description based on saliency region detection Can be expressed as:

[0016]

[0017] In the formula, Represents the attention matrix for local saliency region detection, and * represents the element-wise product operation by channel. Represents the max pooling operation with a kernel of m×m;

[0018] The above attention matrix A s Can be formulated as:

[0019] A s = Softplus(W 2 (ReLU(W 1 S t )))#(2)

[0020] In the formula, W 1 And W 2 Represent the linear mapping weights, and ReLU(·) and Softplus(·) represent the ReLU and Softplus non-linear activation functions respectively.

[0021] Preferably, the method for the global feature extraction module GFE to extract global feature descriptions with geometric characteristics is as follows:

[0022] Denote the deep feature of the t-th frame image extracted by the backbone network as Then the global feature description extracted based on generalized mean pooling is Where g k Can be expressed as:

[0023]

[0024] In the formula, Represents the i-th element of the k-th channel in the feature map D t , N = H 2 ×W 2 ; p > 0 is a learnable hyperparameter, initialized to 3.0.

[0025] Preferably, the method for the feature fusion module AFA to fuse local features and global features and obtain frame-level feature descriptions with local and global semantics is as follows:

[0026] First, the global feature g t and the local feature are transformed into feature vectors with the same dimension;

[0027] Then, the fusion of the local feature and the global feature is achieved based on the multi-head self-attention mechanism;

[0028] Finally, based on the multi-head self-attention mechanism, the frame-level feature description z 0 .

[0029] Preferably, the method by which the feature fusion module AFA transforms the global feature g t and the local feature into feature vectors with the same dimension is as follows:

[0030] For the global feature g t , the transformed feature vector x g =W g g t , where represents a learnable weight matrix;

[0031] Similarly, for the local feature its transformed feature description is:

[0032]

[0033] In the formula, represents a learnable weight matrix, represents unfolding along the spatial dimension into a matrix with dimension C 1 ×N 1 , where

[0034] Preferably, the method by which the feature fusion module AFA achieves the fusion of the local feature and the global feature based on the multi-head self-attention mechanism is as follows,

[0035] Denote where x fuse is an added learnable feature fusion token vector; taking X as the input, the output of the self-attention mechanism can be calculated as:

[0036]

[0037] In the formula, (·) T represents the matrix transpose operation, and are respectively the input X passing through the learnable linear matrix W q , Wk and W v The projection result after

[0038] Preferably, the feature fusion module AFA is based on the multi-head attention mechanism and outputs the frame-level feature description z after fusing local features and global features 0 The method is as follows:

[0039] Denote Z i as the output of the i-th self-attention, then the output of the multi-head self-attention is:

[0040] Z MultiHead =[Z 1 , Z 2 , …, Z h W O #(6)

[0041] In the formula, h represents the number of multi-head attentions, and W O represents the learnable weight. Based on the above multi-head self-attention mechanism for spatial domain feature fusion, then is the output state of x fuse , that is, the frame-level spatial domain fusion feature description.

[0042] Preferably, the temporal feature aggregation module TFA performs temporal feature aggregation. Temporal feature aggregation is to aggregate the frame-level feature descriptions of the video sequence into video-level feature representations based on the attention mechanism. The specific method is as follows

[0043] For the video sequence The frame-level feature description after its spatial domain aggregation is denoted as where represents the frame-level spatial domain fusion feature of the t-th frame I t . Then the temporal feature aggregation feature vector z TA is:

[0044]

[0045] In the formula, is the output result of the frame-level feature description Z SA after multi-head self-attention. The specific calculation process is the same as in formulas (5)-(6); represents average pooling processing.

[0046] Preferably, in step one, the spatial domain loss function and the temporal domain loss function are designed to supervise frame-level feature extraction and video-level feature extraction respectively. The spatial domain loss function and the temporal domain loss function are jointly minimized for model optimization. The design and optimization methods of the spatial domain loss function and the temporal domain loss function are as follows

[0047] Given training data where x n represents video data with label y n the model optimization process can be formulated as the following minimization problem:

[0048]

[0049] In the formula, W k and b k represent the weights and bias terms of the k-th layer, represents the predicted label of x n ;

[0050] where the overall loss function can be expressed as:

[0051]

[0052] In the formula, and represent the spatial domain loss function and the temporal domain loss function respectively, and λ space is a hyperparameter;

[0053] where the spatial domain loss function is composed of a local loss function and a global loss function, and is used to supervise the fusion of local features and global features. The formula is as follows,

[0054]

[0055] In the formula, and represent the local loss function and the global loss function respectively, and have the same form, and are both composed of the cross-entropy loss and the triplet loss ;

[0056] where the temporal domain loss function is used to supervise video-level feature extraction. The formula is as follows,

[0057]

[0058] In the formula, λ cont is a hyperparameter, represents the contrastive loss, and it can be expressed as:

[0059]

[0060] In the formula, z a , z p , z nrespectively represent the feature vectors of the target sample, positive sample, and negative sample; P(a) and N(a) respectively represent the a index sets of all positive samples and all negative samples, and d(·) represents the cosine distance function.

[0061] Compared with the prior art, the advantages of the present invention are as follows: The invention designs two core modules: a spatial feature aggregation module and a temporal feature aggregation module. The former provides rich frame-level features for the temporal aggregation module based on a frame-level feature extraction strategy that integrates local and global features. The latter aggregates the frame-level features based on an attention mechanism to generate discriminative and robust video-level features. Therefore, the spatial feature aggregation module and the temporal feature aggregation module proposed by this model can effectively address challenges such as frame occlusion and small inter-class differences, and can significantly improve the discriminability and robustness of video-level representations. Brief Description of the Drawings

[0062] Figure 1 is the structural schematic diagram of the video person re-identification method of the present invention;

[0063] Figure 2 is the structural diagram of the spatial feature aggregation module of the present invention;

[0064] Figure 3 is the structural diagram of the local feature extraction module of the present invention;

[0065] Figure 4 is the structural diagram of the temporal feature aggregation module of the present invention. Detailed Embodiments

[0066] The following will further describe the present invention. A video person re-identification method based on spatio-temporal feature extraction and adaptive aggregation is as Figures 1 to 4 follows, and the method steps are as follows:

[0067] Step 1: Construct a video person re-identification model, which is composed of three main modules: a backbone network, a spatial feature aggregation module SFA (Spatial Feature Aggregation, SFA), and a temporal feature aggregation module TFA (Temporal Feature Aggregation, TFA);

[0068] Step 2: Given an input video clip, extract initial frame-level feature maps through the backbone network, including a shallow feature map S t and a deep feature map D t ; The backbone network is based on the ResNeSt50 neural network model for the classification and extraction of feature images, obtaining a shallow feature map S t and a deep feature map D t, compared with traditional technologies, as the backbone network, ResNeSt50 can perceive context information in a larger range, thereby enhancing the ability to understand complex scenes;

[0069] Step three, input the shallow feature map S t and the deep feature map D t into the spatial feature aggregation module SFA respectively. The local and global feature extraction modules in SFA are used to obtain local features and the global feature g t respectively, and the frame-level feature Z that fuses local features and global features is extracted through the attention-based feature aggregation module (AFA) SA ;

[0070] Step four, using the frame-level feature Z SA as the input, the temporal feature aggregation module based on the attention mechanism is used to aggregate the frame-level features to extract the video-level feature z for video re-identification TA . Through this process, an end-to-end video pedestrian re-identification process is realized.

[0071] In order to extract rich frame-level feature descriptions, the present invention proposes a spatial domain feature fusion strategy that integrates local features and global features based on the attention mechanism. As Figure 2 shown, the specific implementation manner of the spatial feature aggregation module SFA is as follows:

[0072] The spatial feature aggregation module SFA includes a local feature extraction module (Local Feature Extraction, LFE), a global feature extraction module (Global Feature Extraction, GFE), and an attention-based feature fusion module (AFA), which are introduced separately below:

[0073] The local feature extraction module (LFE) detects local significant regions through the attention mechanism, and extracts local feature descriptions with geometric characteristics from the shallow features extracted by the backbone network. The module designed in the present invention can effectively capture the significant regions in the image. Even under the conditions of frame occlusion or complex background, it can still accurately focus on the local significant regions of pedestrians, thereby significantly enhancing the expression ability and robustness of features. As Figure 3 shown, the specific method is as follows:

[0074] Denote the shallow feature of the t-th frame image extracted by the backbone network as Then the local feature description based on significant region detection can be expressed as:

[0075]

[0076] In the formula, represents the attention matrix for local saliency region detection, and * represents the element-wise product operation by channel, represents the max pooling operation with a kernel of m×m;

[0077] The above attention matrix A s can be formulated as:

[0078] A s = Softplus(W 2 (ReLU(W 1 S t ))) #(2)

[0079] In the formula, W 1 and W 2 represent the linear mapping weights, and ReLU(·) and Softplus(·) represent the ReLU and Softplus non-linear activation functions respectively.

[0080] The local features extracted based on the above method have the characteristic of describing the geometric structure of the spatial saliency region.

[0081] The global feature extraction module (GFE) extracts a global feature description with geometric characteristics from the deep features extracted by the backbone network through Generalized-Mean pooling (GeM). The module designed in the present invention can effectively extract a comprehensive and discriminative global feature representation by weighted aggregation of each important pixel point in the feature map, thereby further enriching the semantic information of the global feature description. The specific method is as follows:

[0082] Denote the deep features of the t-th frame image extracted by the backbone network as Then the global feature description extracted based on generalized mean pooling is where g k can be expressed as:

[0083]

[0084] In the formula, represents the i-th element of the k-th channel in the feature map D t , N = H 2 ×W 2 ; p>0 is a learnable hyperparameter, initialized to 3.0.

[0085] The global feature description extracted based on the above method has a general description of the overall spatial information.

[0086] The Attention-based Feature Aggregation module (AFA) utilizes the self-attention module to fuse local and global features, in order to obtain frame-level feature descriptions with both local and global semantics. The frame-level feature descriptions generated by the AFA module through feature aggregation are highly discriminative, robust, and semantically consistent, providing rich frame-level feature representations for the temporal feature aggregation module. The specific method is as follows:

[0087] First, the global feature g t and the local feature are transformed into feature vectors with the same dimension, and the method is as follows:

[0088] For the global feature g t , the transformed feature vector x g = W g g t , where represents the learnable weight matrix;

[0089] Similarly, for the local feature its transformed feature description is:

[0090]

[0091] In the formula, represents the learnable weight matrix, represents unfolding along the spatial dimension into a matrix with dimension C 1 × N 1 , where

[0092] Then, the fusion of local and global features is achieved based on the multi-head self-attention mechanism, and the method is as follows:

[0093] Denote where x fuse is the added learnable feature fusion token vector; taking X as the input, the calculation process of the output of the self-attention mechanism can be expressed as:

[0094]

[0095] In the formula, (·) T represents the matrix transpose operation, and are the projection results of the input X after passing through the learnable linear matrices W q , W k and W v respectively.

[0096] Finally, based on the above self-attention mechanism, output z 0 is the frame-level feature description after fusing local features and global features.

[0097] Based on the above attention mechanism feature fusion mechanism, output the frame-level feature description z after fusing local features and global features 0 The method is as follows:

[0098] Denote Z i as the output of the i-th self-attention, then the output of the multi-head self-attention is:

[0099] Z MultiHead = [Z 1 , Z 2 , …, Z h W O #(6)

[0100] In the formula, h represents the number of multi-head attentions, and W O represents the learnable weight. Based on the spatial domain feature fusion of the above multi-head self-attention mechanism, then is the output state of x fuse , which is the frame-level spatial domain fusion feature description.

[0101] The temporal feature aggregation module TFA of the present invention performs temporal feature aggregation. Temporal feature aggregation is to aggregate the frame-level feature descriptions of the video sequence into a video-level feature representation based on the attention mechanism. The module set by the present invention effectively mines more discriminative clues by comprehensively exploring the global correlation between frame-level features and assigning appropriate weights to frames of different importance levels, significantly improving the accuracy and robustness of the video feature representation. As Figure 4 shown, the specific implementation method is as follows:

[0102] For the video sequence its frame-level feature description after spatial domain aggregation is denoted as where represents the frame-level spatial domain fusion feature of the t-th frame I t . Then, the temporal feature aggregation feature vector z TA based on the attention mechanism is:

[0103]

[0104] In the formula, is the output result of the frame-level feature description Z SA after multi-head self-attention. The specific calculation process is the same as that of formulas (5)-(6) and will not be elaborated here; represents average pooling processing.

[0105] Based on the above-mentioned attention mechanism for temporal feature aggregation, video-level feature descriptions for re-identification are extracted.

[0106] The present invention designs a spatial loss function and a temporal loss function to supervise frame-level feature extraction and video-level feature extraction respectively, and jointly minimizes the spatial loss function and the temporal loss function to optimize the model.

[0107] The design and optimization methods of the spatial loss function and the temporal loss function are as follows.

[0108] Given training data where x n represents video data with label y n the model optimization process can be formulated as the following minimization problem:

[0109]

[0110] In the formula, W k and b k represent the weight and bias terms of the k-th layer. represents the predicted label of x n ;

[0111] The overall loss function can be expressed as:

[0112]

[0113] In the formula, and represent the spatial loss function and the temporal loss function respectively, and λ space is a hyperparameter;

[0114] The spatial loss function is composed of a local loss function and a global loss function, and is used to supervise the fusion of local features and global features. The formula is as follows.

[0115]

[0116] In the formula, and represent the local loss function and the global loss function respectively. and have the same form, and are both composed of cross-entropy loss and triplet loss ;

[0117] The temporal loss function is used to supervise video-level feature extraction. The formula is as follows.

[0118]

[0119] In the formula, λ cont is a hyperparameter, denotes the contrastive loss, which can be expressed as:

[0120]

[0121] In the formula, z a , z p , z n respectively represent the feature vectors of the target sample, the positive sample, and the negative sample; P(a) and N(a) respectively represent the index sets of all positive samples and all negative samples of z a , and d(·) represents the cosine distance function. The optimization of the model is achieved by jointly minimizing the spatial domain loss function and the temporal domain loss function.

[0122] In summary, the jointly optimized loss function can simultaneously focus on the important features in the spatial and temporal dimensions, effectively guide the model to generate more discriminative and representative video-level features during the feature extraction process, and significantly improve the target retrieval ability of the model in complex scenarios.

[0123] To better present the effectiveness of the video pedestrian re-identification method based on spatio-temporal feature extraction and adaptive aggregation of the present invention, the present invention is also verified through comparative experiments and ablation experiments.

[0124] The experiments in this embodiment use four commonly used benchmark datasets and evaluation metrics for performance evaluation. The specific information is as follows:

[0125] The MARS dataset is the first large-scale video pedestrian re-identification dataset, which consists of 20,715 video sequences collected from 6 non-overlapping regions of 1,261 pedestrians. Each video sequence contains an average of 59 frames; among them, 625 pedestrian data are used for the training set, and 636 pedestrian data are used for the test set.

[0126] The DukeMTMC-VideoReID (abbreviated as DukeV) dataset is collected by cameras in 8 non-overlapping regions on the Duke University campus, containing 4,832 video sequences of 1,812 pedestrians. Each video sequence contains an average of 168 frames; among them, 702 pedestrian data are used for the training set, and 702 pedestrian data are used for the test set.

[0127] The LS-VID dataset is collected by cameras in 15 non-overlapping regions, containing 3,772 pedestrians and a total of 14,943 video sequences. Each video sequence contains an average of 200 frames. Among them, the training set contains 842 pedestrians, and the test set contains 2,730 pedestrians.

[0128] The iLIDS-VID dataset is collected by cameras in two non-overlapping regions and consists of 600 video sequences of 300 pedestrians, with each video sequence containing an average of 71 frames. Among them, the data of 150 pedestrians are used for the training set, and the data of the remaining pedestrians are used for the test set. This data contains more occlusion samples and is more challenging.

[0129] The evaluation metrics include the Cumulated Matching Characteristics (CMC) curve from Rank1 to Rank5 and the Mean Average Precision (mAP).

[0130] In the present invention, video clips are generated from the above benchmark dataset through sparse temporal sampling, and each frame within the video clips is data-augmented by random flipping and random erasing with a probability of 0.5. The size of each frame is set to 256×128. During the model training stage, Adam is used for model optimization, and the specific training parameters are set as follows: the initial learning rate and weight decay are 3×10 -4 and 5×10 -4 respectively. After every 40 training epochs, the learning rate decays to one-tenth of the original; the batch size is 32, which includes 8 pedestrians, with 4 video clips for each pedestrian and each video clip containing 8 frames.

[0131] In the method evaluation stage, in order to analyze the performance advantages of the method proposed in the present invention, it is compared with existing CNN-based methods and Transformer-based methods. The specific experimental results are shown in Table 1.

[0132] Table 1 Experimental comparison results between the method proposed in the present invention and existing methods

[0133]

[0134] From the analysis of the experimental results in Table 1, it can be seen that the performance of the method proposed in the present invention on the four benchmark datasets is higher than that of the comparative methods. Specifically, on the large-scale MARS and LS-VID datasets, the method proposed in the present invention achieves the best mAP, which are 89.1% and 85.6% respectively, significantly superior to the comparative methods. On the DukeV dataset and the iLIDS-VID dataset, the method proposed in the present invention achieves comparable performance to the comparative methods; for example, on the DukeV dataset, the mAP of the method proposed in the present invention is 97.4%, which is basically equivalent to the highest mAP of 97.6% of the comparative methods.

[0135] On the other hand, in order to analyze the overall impact of each module in the method proposed in the present invention on the model, the proposed method is divided into 3 parts, and ablation experiments are conducted on two key modules except the backbone network.

[0136] Table 2 Ablation experiment results of the key modules of the method proposed in this invention

[0137]

[0138] Table 2 shows the ablation experiments of two key modules in the proposed method. From the experimental results in Table 2, it can be seen that compared with the baseline method in the first row of Table 2, after introducing TFA, the mAP is increased by 0.3%; after introducing SFA, the mAP is increased by 1.8%; after introducing TFA and SFA simultaneously, the mAP and Rank1 are increased by 1.7% and 2.2% respectively. At the same time, it can be seen that compared with the baseline method, the TFA and SFA modules increase the number of parameters and computational complexity, and the computational complexity increases by 0.34 GFLOPs. The above experimental results verify the effectiveness of the spatial domain feature aggregation of the SFA module and the temporal domain feature aggregation of the TFA module.

[0139] Table 3 Ablation experiment results of each sub-module in the spatial domain feature aggregation module

[0140]

[0141] Table 3 shows the ablation experiment results of each sub-module of the spatial domain feature aggregation. From the experimental results in Table 3, it can be seen that on the basis of the backbone network extracting shallow features and deep features, by introducing local feature and global feature aggregation based on the attention mechanism, its mAP is increased by 1.2% compared with the baseline method. The above experiment shows that the effective local feature and global feature aggregation can be realized based on the attention mechanism. In addition, when introducing the local feature extraction module (LFE) and the global feature extraction module (GFE), the performance of the model is further improved. This also verifies that LFE can capture the significance in local features and GFE can extract more comprehensive global features. The improvement of these performances further shows that SFA is beneficial to aggregating local and global features to provide richer frame-level features.

[0142] Table 4 Ablation experiment results of the temporal domain feature aggregation module

[0143]

[0144] Table 4 shows the influence of different temporal domain feature aggregation methods on the performance of the proposed method. As shown in Table 4 of the experimental results, TFA has achieved significant improvements compared with other methods in the two key indicators of mAP and Rank1. This strongly confirms the effectiveness of TFA in temporal domain feature aggregation and demonstrates the superior performance of TFA for video feature aggregation.

[0145] The above has introduced in detail a video pedestrian re-identification method based on spatio-temporal feature extraction and adaptive aggregation provided by the present invention. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. It is possible to make changes and improvements to the present invention without exceeding the concept and scope defined by the appended claims. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A video pedestrian re-identification method based on spatial-temporal feature extraction and adaptive aggregation, characterized by: The method steps are as follows: Step 1: construct a video pedestrian re-identification model, which consists of three main modules: a backbone network, a spatial feature aggregation module SFA, and a temporal feature aggregation module TFA; Step 2: Given an input video clip, the backbone network is used to extract the initial frame-level feature map, including the shallow feature map S t and deep feature map D t ; Step 3: transform the shallow feature map S t and deep feature map D t They are input into the spatial feature aggregation module SFA respectively, and the local features are obtained by the local and global feature extraction modules in SFA respectively. and the global feature g t , and extract the frame-level feature Z that integrates local features and global features through the feature aggregation module based on the attention mechanism SA ; Step 4: Using frame-level features Z SA As input, the temporal feature aggregation module is used to aggregate frame-level features based on the attention mechanism to extract video-level features z for video re-identification. TA ,Through this process, an end-to-end video pedestrian re-identification process is realized.

2. The video pedestrian re-identification method based on spatial-temporal feature extraction and adaptive aggregation according to claim 1 is characterized by: In step 1, the spatial feature aggregation module SFA includes a local feature extraction module LFE, a global feature extraction module GFE and a feature fusion module AFA based on an attention mechanism. The local feature extraction module LFE detects local salient areas through an attention mechanism and extracts local feature descriptions with geometric characteristics from shallow features extracted by the backbone network; The global feature extraction module GFE extracts global feature descriptions with geometric characteristics from deep features extracted by the backbone network through generalized mean pooling; The feature fusion module AFA is based on the attention mechanism and utilizes the self-attention module to fuse local features and global features, and obtains a frame-level feature description with local and global semantics.

3. The video pedestrian re-identification method based on spatial-temporal feature extraction and adaptive aggregation according to claim 2 is characterized by: The method for extracting local feature descriptions with geometric characteristics by the local feature extraction module LFE is as follows: The shallow features of the t-th frame image extracted by the backbone network are Then the local feature description based on salient region detection It can be expressed as: In the formula, represents the attention matrix for local salient region detection, * represents the element-wise product operation by channel, Represents a maximum pooling operation with a kernel of m×m; The above attention matrix A s It can be formulated as: A s =Softplus(W2(ReLU(W1S t )))#(2) Where W1 and W2 represent linear mapping weights, ReLU(·) and Softplus(·) represent ReLU and Softplus nonlinear activation functions, respectively.

4. The video pedestrian re-identification method based on spatial-temporal feature extraction and adaptive aggregation according to claim 2 is characterized by: The method for extracting the global feature description with geometric characteristics by the global feature extraction module GFE is as follows: The deep features of the t-th frame image extracted by the backbone network are The global feature description based on generalized mean pooling is where g k It can be expressed as: In the formula, Represents the feature map D t The i-th element of the k-th channel in , N = H2 × W2; p>0 is a learnable hyperparameter, initialized to 3.

0.

5. The video pedestrian re-identification method based on spatial-temporal feature extraction and adaptive aggregation according to claim 2 is characterized by: The method in which the feature fusion module AFA fuses local features and global features and obtains a frame-level feature description with local and global semantics is as follows: First, the global feature g t and local features Convert to feature vector of the same dimension; Then, the fusion of local features and global features is realized based on the multi-head self-attention mechanism; Finally, based on the multi-head self-attention mechanism, the frame-level feature description z0 is output after fusing local features and global features.

6. The video pedestrian re-identification method based on spatial-temporal feature extraction and adaptive aggregation according to claim 5 is characterized by: The feature fusion module AFA combines the global feature g t and local features The method of converting to a feature vector of the same dimension is as follows: For the global feature g t , the transformed feature vector x g =W g g t ,in represents the learnable weight matrix; Similarly, for local features Its transformed feature description for: In the formula, represents the learnable weight matrix, Indicates that along the space dimension Expand into a matrix of dimension C1×N1, where 7. The video pedestrian re-identification method based on spatial-temporal feature extraction and adaptive aggregation according to claim 6 is characterized by: The method of the feature fusion module AFA to realize the fusion of local features and global features based on the multi-head self-attention mechanism is as follows: remember where x fuse The token vector is added to the learnable feature fusion; with X as input, the output of the self-attention mechanism is The calculation process can be expressed as: In the formula, (·) T represents the matrix transpose operation, as well as The input X is the learnable linear matrix W. q , W k and W v The projection result after .

8. The video pedestrian re-identification method based on spatial-temporal feature extraction and adaptive aggregation according to claim 7 is characterized by: The feature fusion module AFA is based on a multi-head attention mechanism, and the method of outputting the frame-level feature description z0 after fusing local features and global features is as follows: Remember Z i is the i-th self-attention output, then the multi-head self-attention output for: WITH MultiHead =[From 1 ,WITH 2 ,…,WITH h ]IN O #(6) In the formula, h represents the number of multi-head attention, W O represents the learnable weights. Based on the spatial feature fusion of the multi-head self-attention mechanism, For x fuse The output state of is the frame-level spatial fusion feature description.

9. The video pedestrian re-identification method based on spatial-temporal feature extraction and adaptive aggregation according to claim 1, characterized in that: The temporal feature aggregation module TFA performs temporal feature aggregation. Temporal feature aggregation is based on the attention mechanism to aggregate the frame-level feature description of the video sequence into a video-level feature representation. The specific method is as follows: For video sequences The frame-level feature description after spatial aggregation is recorded as in Indicates the tth frame I t The frame-level spatial fusion features are then aggregated into feature vector z based on the temporal features of the attention mechanism. TA for: In the formula, is the frame-level feature description Z SA The output result of multi-head self-attention is calculated in the same way as equations (5)-(6); Represents average pooling processing.

10. The video pedestrian re-identification method based on spatial-temporal feature extraction and adaptive aggregation according to claim 1, characterized in that: In step 1, a spatial loss function and a temporal loss function are designed to supervise frame-level feature extraction and video-level feature extraction respectively, and the model is optimized by jointly minimizing the spatial loss function and the temporal loss function. The design and optimization methods of the spatial loss function and the temporal loss function are as follows: Given training data where x n Indicates that the label is y n video data, the model optimization process can be formulated as the following minimization problem: Where W k and b k represents the weights and biases of the kth layer, Represents x n The predicted label of Among them, the overall loss function It can be expressed as: In the formula, and Respectively represent the spatial domain loss function and the temporal domain loss function, λ space is a hyperparameter; Among them, the spatial domain loss function It is composed of a local loss function and a global loss function, which is used to supervise the fusion of local features and global features. The formula is as follows: In the formula, and Represent the local loss function and the global loss function respectively, and The form is consistent, both are derived from the cross entropy loss and triplet loss constitute; Among them, the time domain loss function For supervised video-level feature extraction, the formula is as follows: In the formula, λ cont is a hyperparameter, represents the contrast loss, which can be expressed as: In the formula, z a ,z p ,z n Represent the feature vectors of target samples, positive samples and negative samples respectively; P(a) and N(a) represent z a The index set of all positive samples and the index set of all negative samples, d(·) represents the cosine distance function.

Citation Information

Patent Citations

  • Face orientation aggregation method and device

    CN109598213A

  • Pedestrian re-identification method based on residual multi-channel attention multi-feature fusion

    CN115830531A

  • Compact fiber structures for snapshot spectral and volumetric oct imaging

    WO2024059556A2