Shielding pedestrian re-identification method based on shielding simulation and feature fusion

By building a feature coupling network based on occlusion simulation and Token constraints, the problem of insufficient robustness of the occlusion pedestrian re-identification method in the prior art in complex scenarios is solved, and efficient feature fusion and robustness improvement is achieved without the need for additional auxiliary models.

CN119992597AActive Publication Date: 2025-05-13XINJIANG UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510108769.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing occlusion pedestrian re-identification methods are inadequate in robustness when dealing with complex occlusion scenarios, and the introduction of additional auxiliary models increases model complexity and computing resource requirements.

Method used

Using the method based on occlusion simulation and feature fusion, a feature coupling network based on occlusion simulation and Token constraints is constructed. The network can handle occlusion scenarios through occlusion simulation strategies, and the distinction and robustness of feature representations are improved through the local-global feature coupling module and the Token orthogonal embedding module.

Benefits of technology

It effectively improves the robustness of the model in various occlusion scenarios, reduces the adverse impact of occlusion or non-target areas on the network, and does not require additional auxiliary models, improving computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992597A_ABST
    Figure CN119992597A_ABST
Patent Text Reader

Abstract

The invention discloses a sheltered pedestrian re-identification method based on sheltering simulation and feature fusion, which belongs to the technical field of computer vision, needs to construct a feature coupling network based on sheltering simulation and Token constraint, and comprises the following specific steps: S1, integrating a pedestrian image data set; s2, designing a blocking simulation strategy based on block mixing; s3, establishing a backbone network; s4, designing a blocking simulation strategy based on block mixing; s5, establishing a local-global feature coupling module; s6, designing a loss function; according to the method based on Transform, the dependency relationship of long-distance information can be effectively processed, and meanwhile, the robustness of the model to deal with various shielding scenes can be improved without the help of an additional auxiliary model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an occluded pedestrian re-identification method based on occlusion simulation and feature fusion. Background Art

[0002] In recent years, many researchers have been studying occluded person re-identification methods. Most existing methods address these challenges by locating visible body parts and aligning them. These methods can be roughly divided into two categories: 1) segmentation-based methods and 2) external cue-based methods. Segmentation-based methods segment pedestrian images or feature maps into blocks and extract local features from them for comparison, thereby aligning feature blocks. Although this alleviates the impact of occlusion to a certain extent, it will cause the loss of semantic information of adjacent parts and destroy the orderliness of the feature space. External cue-based methods usually use auxiliary models of pose estimation and semantic segmentation to indicate unoccluded human body parts, both of which help to eliminate occluded features. However, the introduction of additional network structures increases the complexity of the model and requires more training time and computing resources. In addition, the effectiveness of external cue-based methods cannot be guaranteed in complex occlusion scenarios due to the accuracy of additional models. Summary of the invention

[0003] The present invention aims to provide an occluded pedestrian re-identification method based on occlusion simulation and feature fusion to solve the problems raised in the above background technology.

[0004] In order to achieve the above object, the present invention provides the following technical solutions:

[0005] A method for re-identifying occluded pedestrians based on occlusion simulation and feature fusion requires building a feature coupling network based on occlusion simulation and Token constraints. The specific steps include:

[0006] S1: Integrate the pedestrian image dataset and divide it into training set, validation set and test set for use in different stages. At the same time, crop the dataset photos to 256*128;

[0007] S2: Design an occlusion simulation strategy based on block blending, select two different training samples from a batch of original input images to fuse into a new training sample, simulate more diverse occlusion situations in the real world, and enhance the network's ability to perceive the target person under various occlusion situations;

[0008] S3: Establish the backbone network, input the new training image into ViT, and use the ViT-B-16 pre-trained weight feature extraction backbone to extract global and local features;

[0009] S4: Design an occlusion simulation strategy based on block mixing, passing the class tokens of the last Transformer layer through the Token Orthogonal Embedding Module to obtain a discriminative class token representation, ensuring that it maintains a representation difference with the patch tokens in the feature embedding space;

[0010] S5: Establish a local-global feature coupling module to reshape the local classification features and local patch features and feed them into the multi-scale Transformer feature integration module to obtain more fine-grained local enhancement features and complement them with the global representation, which greatly reduces the adverse effects of occlusion or non-target areas on the network;

[0011] S6: Design a loss function and use it to evaluate the model prediction results.

[0012] Preferably, the occlusion simulation strategy based on block mixing proposed in step S2 specifically implements the following steps:

[0013] S21: Get a batch of original images from the dataset Where B represents the number of pedestrian categories in a batch. In order to better utilize occlusion enhancement to solve the occlusion problem, two different training samples (x1, y1) and (x2, y2) are selected to merge into a new training sample.

[0014] S22: Considering the color difference in the changing environment, the data is locally grayed out so that the model can adapt to the significant changes caused by local color loss and reduce the resulting deviation; a maximum number of attempts is set to select a random aspect ratio and then determine the random rectangular area R g Height H r and width W r ;

[0015] S23: Different from the traditional method of directly converting the selected area RGB color space into grayscale color space, a random angle θ∈[0, 360°) is selected to convert the selected area RGB color space into grayscale color space. g The geometric center (W r / 2,H r / 2) is the rotation axis, for the rectangular area R g To rotate, the rotation transformation matrix M(θ) is defined as:

[0016]

[0017] in:

[0018] t x =(1-cos(θ))·W r / 2+sin(θ)·H r / 2

[0019] t y =(1-cos(θ))·H r / 2-sin(θ)·W r / 2

[0020] The new image after rotation transformation It can be obtained by the following transformation:

[0021] R′ g (u, v) = R g (M -1 (θ)·(u, v, 1) T )

[0022] In the above formula, M -1 (θ) is the inverse matrix of M(θ), which is used to map from the transformed space back to the original space, (u, v) is the coordinate point in the transformed image, and then choose to replace the grayscale image to all color channels or only to a randomly selected color channel, which provides more variations for model training.

[0023] S24: Due to the vertical symmetry of human features, vertical segmentation may damage key body parts, so only horizontal occlusion simulation is considered, the image is divided into N blocks of equal size along the horizontal axis, the 2D coordinates of each block are calculated, and stored in the coordinate list Then assign an index to each block and generate an index list I l =[I1, I2, ..., I N ]; At this time, the equally divided image can be expressed as In the real world, occlusions are diverse and irregular, and appear randomly at any position in space. The divided image blocks are sampled according to N / 2 to obtain the number of sampling blocks n, and the coordinate list C l The index of n blocks is randomly selected in , and the formula is as follows:

[0024] I′ l =sample(range(len(C l )), n)

[0025] S25: Mix the images in equal proportions according to the indexes, so as to effectively cover different positions in the space. The specific form is as follows:

[0026]

[0027] Since the generated occluded image contains labels from different identities, providing only one identity label at the network input stage may disrupt the recognition ability of the model; therefore, for labels y1 and y2, a combination ratio λ is set to combine them; here λ is expressed as n / N, and the formula is as follows:

[0028]

[0029] So as to obtain new training samples because It is obtained by segmenting and merging two different images, but each image needs to be optimized by IDLoss and TripletLoss, so the loss function is redefined:

[0030]

[0031] Preferably, the step S3 proposes to construct a backbone network, and the specific steps include:

[0032] S31: Given an input image Where H, W, and C represent the height, width, and number of channels of the image respectively;

[0033] S32: Use a sliding window to divide the image X into N (h×w) Where N can be expressed as:

[0034]

[0035] In the above formula, S and P represent the step size of the sliding window and the size of the image block respectively; since vit requires a sequence as input, each patch is embedded into a D-dimensional space through a linear projection function F(·);

[0036] S33: Add an additional learnable class token x cls is appended to the input sequence to capture the semantic information of pedestrians and is used as the global feature representation f of the encoder. gb ; A learnable position embedding is added to the patch embedding before feeding into the transformer block and camera embedded They are used to preserve the position information and camera-specific information of the image respectively. The embedding process of the input sequence can be expressed as:

[0037]

[0038] In the above formula, λ is a hyperparameter used to balance the camera embedding weights;

[0039] S34: Input Embedding Will be processed by L transformer layers, the final output of the encoder Divided into two parts (global features and local features): and In order to learn more distinctive pedestrian features, the local features are divided into K groups to form K groups of local patch features The size of each group is (N / / K)×D; set a local class token in the local feature The local class token and each group of local patch features are concatenated and fed into a shared transformer layer to learn local classification features for K groups.

[0040] Preferably, in step S4, an occlusion simulation strategy based on block mixing is proposed, and the specific steps include:

[0041] S41: Denote the class tokens of the last Transformer layer as ycls, where K is the number of class tokens; before this, in order to suppress overfitting while reducing the co-adaptability between features and stabilize the gradient, Dropout and BatchNormalization are introduced to preprocess the input features; then L2 normalization is performed along the feature dimension to ensure that the feature embedding has a stable scale before entering the matrix operation. The above process can be expressed as:

[0042]

[0043] In the above formula, ||·|| F represents the Frobenius norm, ∈ is a number designed to prevent division by zero;

[0044] S42: Based on prior knowledge, orthogonal constraints can achieve inter-class separation while enhancing intra-class compactness, which helps the model improve the ability to distinguish similar classes. In order to maintain the difference between class tokens and patch tokens, the class tokens that have been normalized by L2 are The orthogonality constraint is imposed as:

[0045]

[0046] In the above formula, I i,j ∈K×K is the identity matrix; HuberLoss is used to constrain the similarity between features to make them close to orthogonal;

[0047] S43: In real scenes, occlusion has great uncertainty. While imposing constraints on class tokens, it may cause the network to learn overly complex and highly differentiated feature representations. Therefore, regularization constraints are introduced to control the size of model weights to help the network learn more sparse feature representations and avoid excessive reliance on certain features due to excessive weights. The formula is as follows:

[0048]

[0049] In the above formula, Hong represents the weight decay coefficient used to control the strength of regularization; finally, the orthogonal constraint and the regularization constraint work together, as described below:

[0050]

[0051] Preferably, in step S5, it is proposed to construct a local-global feature coupling module, and the specific steps include:

[0052] S51: Local classification feature f l =[f l 1 , f l 2 , ..., f l K ] and local patch features They are concatenated separately and then reshaped to obtain the characteristic shape and

[0053] S52: Specific analysis is performed on the local classification feature branch. The same process is applicable to the other branch. First, the multi-dimensional attention mechanism of ODconv is used to mine detail features from multi-view and multi-scale images. Different from acting directly on the feature map, a learnable dynamic weight w1 is initialized in the dynamic convolution branch to increase the weight of the main line and make the dynamic convolution parts complement each other. The specific process is as follows:

[0054]

[0055] S53: Utilize the MHSA mechanism and hierarchical feature extraction of the Transformer encoder to aggregate more nuances. Prior to this, a token of a specified size was generated based on the feature map of a branch, as follows:

[0056]

[0057] In the above formula, the token Represents global information of local patch features;

[0058] S54: The two branch outputs after ODconv are expressed as and Then with Cascaded and fed into the Transformer encoder for self-attention calculation. The specific calculation formula is described as follows:

[0059]

[0060] In the above formula, D represents the number of channels, represents matrix multiplication; q, k, and v represent the matrix multiplication by and The query matrix, key matrix and value matrix obtained by concatenation are and w q , w k and w v Represent different weight matrices respectively; subsequently, the output A is fed into the FFN network to enhance the embedded representation capability; the final output of the local classification feature branch is expressed as: Similarly, the output of the other branch is: Next, add the outputs of the two branches to get the integrated features.

[0061] S55: Further extract and enhance the integrated features, using a feature extractor consisting of a convolutional layer and a pooling layer to obtain local enhanced features The formula is as follows:

[0062] f k =σ(Avg(G(F out )))

[0063] S55: Define a learnable weight matrix Where C represents the number of categories, the weight matrix and f k Perform Hadamard product to force the model to give priority to categories that contribute to local features; introduce global features f g Matmul multiplies the above weights to obtain weights with higher classification accuracy The specific process is described as follows:

[0064]

[0065] S56: Through the above operations, the local features and global representations are coupled at multiple scales to enhance the classification accuracy. Softmaxloss is usually used for optimization. Considering that it has a lot of room for optimization in terms of expanding the inter-class distance between heterogeneous samples and reducing the intra-class distance between homogeneous samples, an additional angle margin is added to the original Softmaxloss for optimization. The formula is as follows:

[0066]

[0067] Preferably, the step S6 proposes to construct a loss function, and the specific steps include:

[0068] S61: Re-ID is considered an image classification task, each ID is a class, so IDloss is used to train the model:

[0069]

[0070] In the above formula, B represents the number of images in batch size, p(y i |x i ) represents x i Identified as y i The predicted probability of

[0071] S62: In order to reduce the intra-class distance and increase the inter-class distance, triplet loss is used for optimization:

[0072]

[0073] In the above formula, f a is the anchor point sample; f p is a positive sample; f n is a negative sample, ||·||2 represents L2-norm;

[0074] S63: Finally, the loss function can be expressed as follows:

[0075]

[0076] In the above formula, α and β are the equilibrium and Hyperparameters of .

[0077] Compared with the prior art, the present invention has the following beneficial effects:

[0078] The Transformer-based method of the present invention can effectively process the dependency of long-distance information without the need for additional auxiliary models, and can improve the robustness of the model in dealing with various occlusion scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 It is a schematic diagram of the overall framework;

[0080] Figure 2 A schematic diagram of an occlusion simulation strategy based on block blending;

[0081] Figure 3 This is a schematic diagram of the architecture of the multi-scale Transformer feature integration module;

[0082] Figure 4 This is a flow chart of an occluded pedestrian re-identification method based on occlusion simulation and feature fusion. DETAILED DESCRIPTION

[0083] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments:

[0084] A pedestrian re-identification method based on occlusion simulation and feature fusion requires the construction of a feature coupling network based on occlusion simulation and Token constraints, such as Figure 1 The figure shows the overall architecture of the proposed feature coupling network based on occlusion simulation and Token constraint, which consists of a block-mixing-based occlusion simulation strategy, a local-global feature coupling module, and a Token orthogonal embedding module. The occluded samples generated from the block-mixing-based occlusion simulation strategy are fed into the pre-trained ViT as the backbone for feature extraction. In the last Transformer layer, the class labels are passed through the Token orthogonal embedding module to obtain more discriminative feature representations. Then, the local-global feature coupling module couples local features and global representations at multiple scales for comprehensive feature learning.

[0085] The specific implementation steps are as follows: Figure 4 shown):

[0086] Step 1: Integrate the pedestrian image dataset and divide it into training set, validation set and test set for use in different stages. At the same time, crop the dataset photos to 256*128;

[0087] Step 2: Select two different training samples from a batch of original input images to fuse into a new training sample, simulating more diverse occlusion situations in the real world and enhancing the network's ability to perceive the target person under various occlusion situations;

[0088] Step 3: Input the new training image into the ViT (using ViT-B-16 pre-trained weights) feature extraction backbone to extract global and local features;

[0089] Step 4: Pass the class tokens of the last Transformer layer through the Token Orthogonal Embedding Module to obtain a discriminative class token representation, ensuring that it maintains representation differences with the patch tokens in the feature embedding space;

[0090] Step 5: The local classification features and local patch features are reshaped and fed into the multi-scale Transformer feature integration module to obtain more fine-grained local enhancement features and complement each other with the global representation, which greatly reduces the adverse effects of occlusion or non-target areas on the network;

[0091] Step 6: Evaluate the model prediction results through the loss function.

[0092] Furthermore, an occlusion simulation strategy based on block blending is constructed;

[0093] The proposed occlusion simulation in step 2 is as follows Figure 2 As shown in Figure 1, most of the current occlusion enhancement strategies use random cropping and random erasing to form an occlusion scene. However, they can only cover some common occlusion scenes and ignore the influence of the external environment, thus failing to effectively solve the problem of occlusion diversity. In order to solve the above problems, the present invention proposes an occlusion simulation strategy based on block mixing. First, a batch of original images are obtained from the dataset. Where B represents the number of pedestrian categories in a batch. In order to better utilize occlusion enhancement to solve the occlusion problem, two different training samples (x1, y1) and (x2, y2) are selected to merge into a new training sample. Prior to this, considering the color difference in the changing environment, the data was locally grayed out so that the model can better adapt to the significant changes caused by local color loss and reduce the resulting deviation. The random aspect ratio is selected by setting a maximum number of attempts to determine the random rectangular area R. g Height H r and width W r After this selection, unlike the traditional method of directly converting the selected region RGB color space into grayscale color space, a random angle θ∈[0,360°) is selected to convert the selected region RGB color space into grayscale color space. g The geometric center (W r / 2,H r / 2) is the rotation axis, for the rectangular area R g To rotate, the rotation transformation matrix M(θ) is defined as:

[0094]

[0095] in:

[0096] t x =(1-cos(θ))·Wr / 2+sin(θ)·H r / 2

[0097] t y =(1-cos(θ))·H r / 2-sin(θ)·W r / 2

[0098] The new image after rotation transformation It can be obtained by the following transformation:

[0099] R′ g (u, v) = R g (M -1 (θ)·(u, v, 1) T )

[0100] In the above formula, M -1 (θ) is the inverse matrix of M(θ), which is used to map from the transformed space back to the original space, (u, v) is the coordinate point in the transformed image, and then choose to replace the grayscale image to all color channels or only to a randomly selected color channel, thereby providing more variations for model training.

[0101] Furthermore, due to the vertical symmetry of human features, vertical segmentation may damage key body parts (such as the head), so only horizontal occlusion simulation is considered, the image is split into N blocks of equal size along the horizontal axis, the 2D coordinates of each block are calculated, and stored in the coordinate list Then assign an index to each block and generate an index list I l =[I1, I2, ..., I N ]. At this time, the equally divided image can be expressed as In the real world, occlusion is diverse and irregular, and may appear anywhere in space. The segmented image blocks are sampled according to N / 2 to obtain the number of sampling blocks n, and the coordinate list C l The index of n blocks is randomly selected in , and the formula is as follows:

[0102] I′ l =sample(range(len(C l )), n)

[0103] Furthermore, the images are mixed in equal proportions according to the indexes, thereby effectively covering different positions in the space.

[0104] The specific form is as follows:

[0105]

[0106] Since the generated occluded image contains labels from different identities, providing only one identity label at the network input stage may disrupt the recognition ability of the model. Therefore, for labels y1 and y2, a combination ratio λ is set to combine them. Here λ is expressed as n / N, and the formula is as follows:

[0107]

[0108] So as to obtain new training samples because It is obtained by segmenting and merging two different images, but each image needs to be optimized by IDLoss and TripletLoss, so the loss function is redefined:

[0109]

[0110] Further, build a backbone network:

[0111] The ViT feature extraction backbone mentioned in step 3 is described as follows: Given an input image Where H, W and C represent the height, width and number of channels of the image respectively; the image X is divided into N (h×w) Where N can be expressed as:

[0112]

[0113] In the above formula, S and P represent the step size of the sliding window and the size of the image patch, respectively. Since vit requires a sequence as input, each patch is embedded into a D-dimensional space through a linear projection function F(·). In addition, an additional learnable class token x cls is appended to the input sequence to capture the semantic information of pedestrians and is used as the global feature representation f of the encoder. gb ; A learnable position embedding is added to the patch embedding before feeding into the transformer block and camera embedded They are used to preserve the position information and camera-specific information of the image, respectively.

[0114] The embedding process of the input sequence can be expressed as:

[0115]

[0116] In the above formula, λ is a hyperparameter used to balance the camera embedding weights.

[0117] Furthermore, the input embedding It will be processed by L transformer layers. The final output of the encoder is Divided into two parts (global features and local features): and In order to learn more distinctive pedestrian features, the local features are divided into K groups to form K groups of local patch features The size of each group is (N / / K) × D. Set a local class token in the local feature The local class token and each group of local patch features are then concatenated and fed into a shared transformer layer to learn local classification features f for K groups. l =[f l 1 , f l 2 , ..., f l K ].

[0118] Furthermore, an occlusion simulation strategy based on block blending is constructed:

[0119] The Token orthogonal embedding module proposed in step 4 is as follows Figure 1 Specifically, the class tokens of the last Transformer layer are represented as ycls, where K is the number of class tokens; before this, in order to suppress overfitting while reducing the co-adaptability between features and stabilize the gradient, Dropout and BatchNormalization are introduced to preprocess the input features; then L2 normalization is performed along the feature dimension to ensure that the feature embedding has a stable scale before entering the matrix operation. The above process can be expressed as:

[0120]

[0121] In the above formula, ||·|| F represents the Frobenius norm, and ∈ is a number designed to prevent division by zero.

[0122] Based on prior knowledge, orthogonal constraints can achieve inter-class separation while enhancing intra-class compactness, which helps the model improve the ability to distinguish similar classes. Therefore, in order to maintain the difference between class tokens and patch tokens, the class tokens that have been normalized by L2 are The orthogonality constraint is imposed as:

[0123]

[0124] In the above formula, I i,j ∈K×K is the identity matrix; HuberLoss is used to constrain the similarity between features to make them close to orthogonal.

[0125] In real scenes, occlusion has great uncertainty. While imposing constraints on class tokens, it may cause the network to learn overly complex and highly differentiated feature representations. Therefore, regularization constraints are introduced to control the size of model weights to help the network learn more sparse feature representations and avoid excessive reliance on certain features due to excessive weights. The formula is as follows:

[0126]

[0127] In the above formula, Hong represents the weight decay coefficient used to control the strength of regularization; finally, the orthogonal constraint and the regularization constraint work together, as described below:

[0128]

[0129] Furthermore, a local-global feature coupling module is constructed:

[0130] In order to complement local features and global representation, the present invention designs a local-global feature coupling module, such as Figure 1 As shown. The multi-scale Transformer feature integration module mentioned in step 5 is as follows Figure 3 As shown. The local classification feature f l =[f l 1 , f l 2 , ..., f l K ] and local patch features They are concatenated separately and then reshaped to obtain the characteristic shape and The local classification feature branch is analyzed in detail, and the other branch is also applicable to the same process. It first uses the multi-dimensional attention mechanism of ODconv to mine detail features from multi-view and multi-scale images. Different from directly acting on the feature map, a learnable dynamic weight w1 is initialized in the dynamic convolution branch, thereby increasing the weight of the main line and allowing the dynamic convolution parts to complement each other.

[0131] The specific process is as follows:

[0132]

[0133] Furthermore, the MHSA mechanism and hierarchical feature extraction of the Transformer encoder are used to aggregate more subtle differences. Prior to this, a token of a specified size is generated based on the feature map of a branch, as follows:

[0134]

[0135] In the above formula, the token Represents the global information of local patch features. The two branch outputs after ODconv are expressed as and Then with Cascaded and fed into the Transformer encoder for self-attention calculation. The specific calculation formula is described as follows:

[0136]

[0137] In the above formula, D represents the number of channels, represents matrix multiplication. q, k and v represent the matrix multiplication by and The query matrix, key matrix and value matrix obtained by concatenation are and w q , w k and w v Represent different weight matrices respectively. Subsequently, the output A is fed into the FFN network to enhance the embedded representation capability. The final output of the local classification feature branch is expressed as: Similarly, the output of the other branch is: Next, add the outputs of the two branches to get the integrated features.

[0138] In order to further extract and enhance the integrated features, a feature extractor consisting of convolutional layers and pooling layers is used to obtain local enhanced features. The formula is as follows:

[0139] f k =σ(Avg(G(F out )))

[0140] Furthermore, we define a learnable weight matrix Where C represents the number of categories. and f k Performing Hadamard product forces the model to prioritize categories that contribute to local features. Next, introduce the global feature f g Matmul multiplies the above weights to obtain weights with higher classification accuracy

[0141] The specific process is described as follows:

[0142]

[0143] Through the above operation, the local features and the global representation are coupled at multiple scales to enhance the classification accuracy. Softmaxloss is usually used for optimization. The present invention considers that there is a lot of room for optimization in terms of expanding the inter-class distance between heterogeneous samples and reducing the intra-class distance between homogeneous samples. An additional angle margin is added to the original Softmaxloss for optimization. The formula is as follows:

[0144]

[0145] Furthermore, we construct the loss function:

[0146] Re-ID is considered an image classification task, each ID is a class, so IDloss is used to train the model:

[0147]

[0148] In the above formula, B represents the number of images in batch size, p(y i |x i ) represents x i Identified as y i In order to reduce the intra-class distance and increase the inter-class distance, the triplet loss is used for optimization:

[0149]

[0150] In the above formula, f a is the anchor point sample; f p is a positive sample; f n is a negative sample, and ||·||2 represents L2-norm.

[0151] Finally, the loss function of the present invention can be expressed by the following formula:

[0152]

[0153] In the above formula, α and β are the equilibrium and Hyperparameters of .

[0154] The above is only an embodiment of the present invention, and the common knowledge such as the known specific technical solutions and / or characteristics in the solution is not described in detail here. It should be pointed out that for those skilled in the art, without departing from the technical solution of the present invention, several modifications and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A method for re-identifying occluded pedestrians based on occlusion simulation and feature fusion, characterized by: Construct a feature coupling network based on occlusion simulation and Token constraints. The specific steps include: S1: Integrate the pedestrian image dataset and divide it into training set, validation set and test set for use in different stages. At the same time, crop the dataset photos to 256*128; S2: Design an occlusion simulation strategy based on block blending, select two different training samples from a batch of original input images to fuse into a new training sample, simulate more diverse occlusion situations in the real world, and enhance the network's ability to perceive the target person under various occlusion situations; S3: Establish the backbone network, input the new training image into ViT, and use the ViT-B-16 pre-trained weight feature extraction backbone to extract global and local features; S4: Design an occlusion simulation strategy based on block mixing, passing the class tokens of the last Transformer layer through the Token Orthogonal Embedding Module to obtain a discriminative class token representation, ensuring that it maintains a representation difference with the patch tokens in the feature embedding space; S5: Establish a local-global feature coupling module to reshape the local classification features and local patch features and feed them into the multi-scale Transformer feature integration module to obtain more fine-grained local enhancement features and complement them with the global representation, which greatly reduces the adverse effects of occlusion or non-target areas on the network; S6: Design a loss function and use it to evaluate the model prediction results.

2. The method for re-identifying an occluded pedestrian based on occlusion simulation and feature fusion according to claim 1, characterized in that: The occlusion simulation strategy based on block mixing proposed in step S2 is specifically implemented by the following steps: S21: Get a batch of original images from the dataset Where B represents the number of pedestrian categories in a batch. In order to better utilize occlusion enhancement to solve the occlusion problem, two different training samples (x1, y1) and (x2, y2) are selected to merge into a new training sample. S22: Considering the color difference in the changing environment, the data is locally grayed out so that the model can adapt to the significant changes caused by local color loss and reduce the resulting deviation; a maximum number of attempts is set to select a random aspect ratio and then determine the random rectangular area R g Height H r and width W r ; S23: Different from the traditional method of directly converting the selected area RGB color space into grayscale color space, a random angle θ∈[0,360°) is selected to convert the selected area RGB color space into grayscale color space. g The geometric center (W r / 2,H r / 2) is the rotation axis, for the rectangular area R g To rotate, the rotation transformation matrix M(θ) is defined as: in: t x =(1-cos(θ)) W r / 2+sin(θ)·H r / 2 t y =(1-cos(θ))·H r / 2-sin(θ)·W r 2 The new image after rotation transformation It can be obtained by the following transformation: R′ g (u,v)=R g (M -1 (θ)·(u,v,1) T ) In the above formula, M -1 (θ) is the inverse matrix of M(θ), which is used to map from the transformed space back to the original space, (u,v) is the coordinate point in the transformed image, and then choose to replace the grayscale image to all color channels or only to a randomly selected color channel, which provides more variations for model training. S24: Due to the vertical symmetry of human features, vertical segmentation may damage key body parts, so only horizontal occlusion simulation is considered, the image is divided into N blocks of equal size along the horizontal axis, the 2D coordinates of each block are calculated, and stored in the coordinate list Then assign an index to each block and generate an index list I l =[I1,I2,…,I N ]; At this time, the equally divided image can be expressed as In the real world, occlusions are diverse and irregular, and appear randomly at any position in space. The divided image blocks are sampled according to N / 2 to obtain the number of sampling blocks n, and the coordinate list C l The index of n blocks is randomly selected in , and the formula is as follows: I′ l =sample(range(len(C l )),n) S25: Mix the images in equal proportions according to the indexes, so as to effectively cover different positions in the space. The specific form is as follows: Since the generated occluded image contains labels from different identities, providing only one identity label at the network input stage may disrupt the recognition ability of the model; therefore, for labels y1 and y2, a combination ratio λ is set to combine them; here λ is expressed as n / N, and the formula is as follows: So as to obtain new training samples because It is obtained by segmenting and merging two different images, but each image needs to be optimized by IDLoss and TripletLoss, so the loss function is redefined:

3. The method for re-identifying an occluded pedestrian based on occlusion simulation and feature fusion according to claim 1, characterized in that: The step S3 proposes to construct a backbone network, and the specific steps include: S31: Given an input image Where H, W, and C represent the height, width, and number of channels of the image respectively; S32: Use a sliding window to divide the image X into N (h×w) patches Where N can be expressed as: In the above formula, S and P represent the step size of the sliding window and the size of the image block respectively; since vit requires a sequence as input, each patch is embedded into a D-dimensional space through a linear projection function F(·); S33: Add an additional learnable class token x cls is appended to the input sequence to capture the semantic information of pedestrians and is used as the global feature representation f of the encoder. gb ; A learnable position embedding is added to the patch embedding before feeding into the transformer block and camera embedded They are used to preserve the position information and camera-specific information of the image respectively. The embedding process of the input sequence can be expressed as: In the above formula, λ is a hyperparameter used to balance the camera embedding weights; S34: The input embedding z0 will be processed by L transformer layers, and the final output of the encoder is Divided into two parts (global features and local features): and In order to learn more distinctive pedestrian features, the local features are divided into K groups to form K groups of local patch features The size of each group is (N / / K)×D; set a local class token in the local feature The local class token and each group of local patch features are concatenated and fed into a shared transformer layer to learn local classification features for K groups.

4. The method for re-identifying an occluded pedestrian based on occlusion simulation and feature fusion according to claim 1, characterized in that: In step S4, an occlusion simulation strategy based on block mixing is proposed, and the specific steps include: S41: Denote the class token of the last Transformer layer as y cls ,in K is the number of class tokens; before this, in order to suppress overfitting while reducing the co-adaptability between features and stabilize the gradient, Dropout and BatchNormalization are introduced to preprocess the input features; then L2 normalization is performed along the feature dimension to ensure that the feature embedding has a stable scale before entering the matrix operation. The above process can be expressed as: In the above formula, ∥·∥ F represents the Frobenius norm, ∈ is a number designed to prevent division by zero; S42: Based on prior knowledge, orthogonal constraints can achieve inter-class separation while enhancing intra-class compactness, which helps the model improve the ability to distinguish similar classes. In order to maintain the difference between class tokens and patch tokens, the class tokens that have been normalized by L2 are The orthogonality constraint is imposed as: In the above formula, I i,j ∈K×K is the identity matrix; HuberLoss is used to constrain the similarity between features to make them close to orthogonal; S43: In real scenes, occlusion has great uncertainty. While imposing constraints on class tokens, it may cause the network to learn overly complex and highly differentiated feature representations. Therefore, regularization constraints are introduced to control the size of model weights to help the network learn more sparse feature representations and avoid excessive reliance on certain features due to excessive weights. The formula is as follows: In the above formula, μ represents the weight decay coefficient used to control the strength of regularization; the final orthogonal constraint and regularization constraint work together, as follows:

5. The method for re-identifying an occluded pedestrian based on occlusion simulation and feature fusion according to claim 1, characterized in that: The step S5 proposes to construct a local-global feature coupling module, and the specific steps include: S51: Local classification features and local patch features They are concatenated separately and then reshaped to obtain the characteristic shape and S52: Specific analysis is performed on the local classification feature branch. The same process is applicable to the other branch. First, the multi-dimensional attention mechanism of ODconv is used to mine detail features from multi-view and multi-scale images. Different from acting directly on the feature map, a learnable dynamic weight w1 is initialized in the dynamic convolution branch to increase the weight of the main line and make the dynamic convolution parts complement each other. The specific process is as follows: S53: Utilize the MHSA mechanism and hierarchical feature extraction of the Transformer encoder to aggregate more nuances. Prior to this, a token of a specified size was generated based on the feature map of a branch, as follows: In the above formula, the token Represents global information of local patch features; S54: The two branch outputs after ODconv are expressed as and Then with Cascaded and fed into the Transformer encoder for self-attention calculation. The specific calculation formula is described as follows: In the above formula, D represents the number of channels, represents matrix multiplication; q, k, and v represent the matrix multiplication by and The query matrix, key matrix and value matrix obtained by concatenation are and w q , w k and w v Represent different weight matrices respectively; subsequently, the output A is fed into the FFN network to enhance the embedded representation capability; the final output of the local classification feature branch is expressed as: Similarly, the output of the other branch is: Next, add the outputs of the two branches to get the integrated features. S55: Further extract and enhance the integrated features, using a feature extractor consisting of a convolutional layer and a pooling layer to obtain local enhanced features The formula is as follows: f k =σ(Avg(G(F out ))) S55: Define a learnable weight matrix Where C represents the number of categories, the weight matrix and f k Perform Hadamard product to force the model to give priority to categories that contribute to local features; introduce global features f g Matmul multiplies the above weights to obtain weights with higher classification accuracy The specific process is described as follows: S56: Through the above operations, the local features and global representations are coupled at multiple scales to enhance the classification accuracy. Softmaxloss is usually used for optimization. Considering that it has a lot of room for optimization in terms of expanding the inter-class distance between heterogeneous samples and reducing the intra-class distance between homogeneous samples, an additional angle margin is added to the original Softmaxloss for optimization. The formula is as follows:

6. The method for re-identifying an occluded pedestrian based on occlusion simulation and feature fusion according to claim 1, characterized in that: The step S6 proposes to construct a loss function, and the specific steps include: S61: Re-ID is considered an image classification task, each ID is a class, so IDloss is used to train the model: In the above formula, B represents the number of images in batch size, p(y i |x i ) represents x i Identified as y i The predicted probability of S62: In order to reduce the intra-class distance and increase the inter-class distance, triplet loss is used for optimization: In the above formula, f a is the anchor point sample; f p is a positive sample; f n is a negative sample, ||·||2 represents L2-norm; S63: Finally, the loss function can be expressed as follows: In the above formula, α and β are the equilibrium and Hyperparameters of .

Citation Information

Patent Citations

  • Shielding image recognition model training method and device, equipment and medium

    CN116311106A

  • Speech emotion recognition method based on time-frequency feature separation type transformer cross fusion architecture

    CN117746908A

  • ViT distillation training method, system and device and readable storage medium

    CN119129701A

  • Information processing apparatus, method for controlling the same, and non-transitory computer-readable storage medium

    US20230077498A1