Gating adaptive image-text feature fusion-based reloading pedestrian re-identification algorithm
By introducing a gated adaptive graphic feature fusion mechanism in the pedestrian re-identification algorithm, combining image and text description information, the problem of insufficient recognition of traditional methods in complex scenarios is solved, and accurate recognition under different clothing conditions is achieved.
Patent Information
- Application Number
- CN202510087927.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-27
AI Technical Summary
The traditional pedestrian recognition method performs poorly when facing complex scenes such as target characters changing clothes and background changes, and relies too much on single-modal image matching, lacking text or other information supplementation.
A dress-changing pedestrian re-identification algorithm based on gated adaptive graphic and text feature fusion is proposed. By combining image and text description information, clothing change information from text description is introduced, and accurate pedestrian recognition under different clothing conditions is achieved through the modal fusion, dynamic weighting and correction mechanism of image features.
Overcoming the limitations of traditional methods in complex scenarios, achieving accurate pedestrian recognition under different clothing conditions, and improving the application ability of the model in real scenarios.
Smart Images

Figure CN120047892A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of pedestrian re-identification, and particularly to an algorithm for re-identifying dressed-changing pedestrians. Background Art
[0002] Pedestrian re-identification has always played a crucial role in the field of computer vision. Especially in video surveillance and security applications, it has become an indispensable technology. The core goal of pedestrian re-identification is to accurately identify and retrieve images or video clips of the same person by analyzing perspectives from different cameras, especially under constantly changing lighting, poses, and background conditions. This task not only involves image retrieval but also needs to handle complex scene changes and perspective differences of different cameras. Therefore, the research on pedestrian re-identification has always been a hot topic in the field of computer vision.
[0003] Traditional pedestrian re-identification methods mainly match pedestrians by extracting appearance features of pedestrians, such as clothing color, style, texture, etc. This means that when a pedestrian changes clothes, the performance of the model in re-identifying the pedestrian will decrease significantly. In addition, traditional pedestrian re-identification methods rely too much on a single modality, that is, image-to-image matching, lacking the supplement of text or other information, which limits the ability of the model to handle complex scenes and results in limitations in real-world applications. Therefore, some researchers have also incorporated multi-modal information into the pedestrian re-identification task, such as combining text descriptions with image features to improve performance. Text-based pedestrian re-identification solves the cross-modal matching problem that traditional pedestrian re-identification methods cannot effectively handle by combining visual and text features.
[0004] The dressed-changing pedestrian re-identification algorithm retrieves the target image of the same person through the clothing changes in the text description. By fusing the biological information of the pedestrian in the source image and the information describing the clothing changes in the text description, the target person wearing the specified clothing is retrieved in the database. The dressed-changing pedestrian re-identification algorithm utilizes the information of clothing changes from the text description and combines the modal fusion of image features. When there is a reference image of the target person and the clothing style and changes of the target person, this method can be applied to a wider range of practical scenarios. For example, when the police are tracking a criminal suspect, the police can quickly find the target suspect from a large amount of surveillance information based on the reference picture of the criminal suspect and the information of clothing changes, avoiding more serious losses, which is of great help in improving the case-handling efficiency of law enforcement officers such as the police and narrowing down the scope of the target suspect. Summary of the Invention
[0005] The present invention proposes a clothing-changing pedestrian re-identification algorithm based on gated adaptive image-text feature fusion, aiming to overcome the limitations of traditional pedestrian re-identification methods in the face of complex scenarios such as target person's clothing change and background change by combining image and text description information. By introducing clothing change information from text descriptions and modal fusion of image features, accurate pedestrian recognition under different clothing conditions is achieved through a dynamic weighting and correction mechanism.
[0006] The technical solution of the present invention is as follows:
[0007] The present invention provides a clothing-changing pedestrian re-identification algorithm based on gated adaptive image-text feature fusion, and the algorithm includes the following steps:
[0008] Step 1: In order to use the CLIP large model as an image encoder, first adjust the smaller side of the source image to the CLIP input size input_dim, and then perform central cropping to generate a square image reference_image of input_dim×input_dim;
[0009] reference_image = {input_dim×input_dim}
[0010] Step 2: In order to use the CLIP large model as a text encoder, convert the input text description captions into a digital ID sequence that the model can understand, and specify the maximum length of the input text as 77 tokens to obtain the tokenized text input text_input;
[0011] text_input = tokenize(captions)
[0012] Step 3: In order to use the CLIP large model as an image encoder, adjust the smaller side of the target image to the CLIP input size input_dim, and then perform central cropping to generate a square image target_image of input_dim×input_dim;
[0013] target_image = {input_dim×input_dim}
[0014] Step 4: Input a batch of preprocessed source images reference_image into the image encoder of the CLIP large model to obtain a source image feature vector reference_features of B×D, where B is the batch size and D is the dimension of the feature vector;
[0015] reference_features = {B×D}
[0016] Step 5: Input a batch of tokenized texts into the text encoder of the CLIP large model to obtain a text feature vector text_features of B×D, where B is the batch size and D is the dimension of the feature vector;
[0017] text_features = {B×D}
[0018] Step 6: Input a batch of preprocessed target images target_image into the image encoder of the CLIP large model to obtain a target image feature vector target_features of B×D, where B is the batch size and D is the dimension of the feature vector;
[0019] target_features = {B×D}
[0020] Step 7: Concatenate the source image feature vector reference_features and the text feature vector text_features extracted by the CLIP large model encoder. The concatenation operation occurs in the last dimension, i.e., the feature dimension. By combining the feature information of the image and the text, a merged feature vector raw_com_feat that integrates image visual information and text semantic information is obtained. Its shape is B×(2×D), where B is the batch size and D is the dimension of the feature vector. This merged feature vector can represent both the image content and the corresponding text description simultaneously, providing a richer feature input for subsequent multimodal tasks. At the same time, add a multi-head attention mechanism Multi-Head Attention module to further explore the relationship between the image and text features. Multi-head attention can focus on the mutual correlation between features in different subspaces, improving the fineness of the fusion. Its input is all raw_com_feat;
[0021] raw_com_feat = Concat(reference_features, text_features)
[0022] attn_output
[0023] = Multi_Head_Attention(raw_com_feat, raw_com_feat, raw_com_feat)
[0024] raw_com_feat = raw_com_feat + attn_output
[0025] raw_com_feat = {B×(2×D)}
[0026] Step 8: Use the image feature attention module att_I and the text feature attention module att_T to adjust the image features and text features respectively, obtaining the reference_features and text_features after attention adjustment;
[0027] reference_features = reference_features × att_I(raw_combined_features)
[0028] text_features = text_features × att_I(raw_combined_features)
[0029] Step 9: Concatenate the reference_features and text_features after attention adjustment along the last dimension to obtain a new combined feature new_combined_features. The dimension after concatenation is {B×(2×D)}, where B is the batch size and D is the dimension of the feature vector;
[0030] new_combined_features = Concat(reference_features, text_features)
[0031] new_combined_features = {B×(2×D)}
[0032] Step 10: Process the concatenated new_combined_features through the weight calculation dynamic_scalar module to output a dynamic scalar dynamic. This scalar represents the relative importance of the image features and text features, ranging from [0, 1]. Through weighted operation, the synthetic feature com_features is finally generated. Here, dynamic is the weight of the image features, and 1 - dynamic is the weight of the text features;
[0033] dynamic = dynamic_scalar(new_combined_features)
[0034] Step 11: Weight-average reference_features and text_features according to the weights of the dynamic scalars obtained in Step 10, and fuse them through the Gated Adaptive Feature Fusion module, hereinafter abbreviated as GAFF in the following formula, to generate the final synthetic feature com_feat, whose dimension is {B×D}, where B is the batch size and D is the dimension of the feature vector;
[0035] com_feat = GAFF(dynamic * image_features, (1 - dynamic) * text_features)
[0036] com_feat = {B×D}
[0037] Step 12: In the training stage, the final synthetic feature com_feat of the network is denoted as Feat s and the target image feature target_features is denoted as Feat tgt for Loss calculation. In the formula, N represents the batch size, κ represents the cosine similarity calculation, and λ is a scaling factor used to adjust the logits range. The λ parameter is set to 100 to enhance the dynamic range of the logits. In the prediction stage, the distance between com_feat and target_features is calculated to determine whether the prediction is correct;
[0038]
[0039] Furthermore, for the image feature attention module att_I and the text feature attention module att_T in Step 8, the specific method for processing each layer of data is as follows:
[0040] Step 8.1: The inputs of the image feature attention module att_I and the text feature attention module att_T are the concatenation of the image feature and the text feature. Assume that the dimension of the image feature is image_embed_dim and the dimension of the text feature is text_embed_dim. Then the input dimension after concatenation is image_embed_dim + text_embed_dim;
[0041] Step 8.2: Use the fully connected layer as the first layer of the image feature attention module att_I and the text feature attention module att_T. Its input is raw_combined_features with a dimension of D+D, and the output dimension is D. The role of this layer is to map the concatenated image and text features to a new space with the same dimension as the image features. This process is equivalent to fusing the two modalities (image and text) and laying the foundation for feature weighting in the next step through the transformation of this layer;
[0042] att_I_fc_layer_output = Linear(raw_combined_features)
[0043] att_T_fc_layer_output = Linear(raw_combined_features)
[0044] att_I_FC_layer_output = {B×D}
[0045] att_T_FC_layer_output = {B×D}
[0046] Step 8.3: The output after passing through the fully connected layer will undergo a non-linear transformation through a ReLU activation function. The role of the ReLU activation function is to introduce non-linearity, enabling the network to better fit complex mapping relationships;
[0047] att_I_FC_ReLU_output = ReLU(raw_combined_features)
[0048] att_T_FC_ReLU_output = ReLU(raw_combined_features)
[0049] Step 8.4: The output activated by ReLU will pass through a Dropout layer. Dropout is a technique to prevent overfitting. It randomly discards a certain proportion of neurons to help the model improve its generalization ability;
[0050] att_I_FC_Dropout_output = Dropout(att_I-FC_ReLU_output)
[0051] att_T_FC_Dropout_output = Dropout(att_T_FC_ReLU_output)
[0052] Step 8.5: The output enters the second fully connected layer, whose input dimension is D and output dimension is also D. The role of this layer is to further process the features transformed by the previous layers, ensure the consistency of the feature dimensions, and prepare data for the subsequent Sigmoid activation;
[0053] att_I_FC_layer_output = Linear(att_I_FC_Dropout_output)
[0054] att_T_FC_layer_output = Linear(att_T_FC_Dropout_output)
[0055] att_I_FC_layer_output = {B×D}
[0056] att_T_FC_layer_output = {B×D}
[0057] Step 8.6: The output after passing through the second fully connected layer will go through a Sigmoid activation function. The role of the Sigmoid function is to compress the output value into the range [0, 1]. In this way, we obtain a numerical value representing the weighted coefficient of the image features, and this coefficient is used to adjust the importance of the image features in the subsequent steps;
[0058] atten_I_output = Sigmoid(att_I_FC_layer_output)
[0059] att_T_output = Sigmoid(att_T_FC_layer_output)
[0060] Furthermore, in the weight calculation dynamic_scalar module in Step 10, the specific method for processing each layer of data is as follows:
[0061] Step 10.1: The weight calculation dynamic_scalar module is a sequence composed of multiple neural network layers. The first layer is a fully connected layer, with the input being new_combined_features, the input dimension being the sum of the image feature dimension and the text feature dimension, that is, D + D, and the output dimension being D;
[0062] Output 1 = Linear(new_combined_features)
[0063] Output 1 = {B×D}
[0064] Step 10.2: Use the ReLU activation function as the weight to calculate the second layer of the dynamic_scalar module. ReLU is a common activation function used to introduce non-linearity. It processes the input data, turning negative values to zero and leaving positive values unchanged;
[0065] Output 2 = ReLU(Output 1 )
[0066] Output 2 = {B × D}
[0067] Step 10.3: Use the Dropout regularization technique as the weight to calculate the third layer of the dynamic_scalar module. During training, it randomly "drops out" the outputs of some neurons to prevent overfitting. The role of this layer is to randomly discard the intermediate outputs of the neural network to improve the generalization ability of the model. It also does not change the dimension of the data, and the output dimension is still D;
[0068] Output 3 = Dropout(Output 2 )
[0069] Output 3 = {B × D}
[0070] Step 10.4: Use another fully connected layer as the weight to calculate the fourth layer of the dynamic_scalar module. This layer maps the previous output (dimension D) to a 1-dimensional output to obtain a scalar value. The role of this layer is to compress the data dimension through the fully connected layer, from D to 1;
[0071] Output 4 = Linear(Output 3 )
[0072] Output 4 = {B × 1}
[0073] Step 10.5: Use the Sigmoid activation function as the weight to calculate the fifth layer of the dynamic_scalar module. The Sigmoid activation function compresses the input value into the range [0, 1], usually used to calculate probabilities or weighting coefficients. The role of this layer is to map the scalar value obtained in the previous layer to a value between 0 and 1 through the Sigmoid function, representing a dynamic weighting coefficient, and its dimension is B × 1;
[0074] dynamic = Sigmoid(Output 4 )
[0075] dynamic = {B×1}
[0076] Further, in the Gated Adaptive Feature Fusion module in step 11, the method of data processing for each layer is specifically as follows:
[0077] Step 11.1: Take batch normalization, ReLU activation function, and linear layer as the first part of the Gated Adaptive Feature Fusion module. This part is used for the preliminary fusion of text feature T and image feature V. Its inputs are text feature T and image feature V, and the output is the preliminary fusion feature fusion;
[0078] fusion = Linear(ReLU(BN(T, V)))
[0079] where Linear is the linear layer and BN is batch normalization;
[0080] Step 11.2: Reduce the input feature dimension from D to D / 2 through a fully connected layer, then perform batch normalization to standardize the data. Subsequently, apply the ReLU activation function to introduce non-linearity. Next, another fully connected layer keeps the feature dimension at D / 2, and batch normalization and activation function processing are performed again. Finally, restore the feature dimension to the original D through a fully connected layer to form the final output. By gradually reducing the feature dimension, adding non-linear activation, and then expanding back to the original dimension, redundant information is reduced, and finally the expressive ability of feature representation is improved. The input is the preliminary fusion feature fusion obtained in step 11.1;
[0081] e = Linear(ReLU(BN(Linear(ReLU(BN(Linear(fusion)))))))
[0082] where Linear is the linear layer and BN is batch normalization;
[0083] Step 11.3: Standardize the data through batch normalization, then apply the specified ReLU activation function to introduce non-linear transformation. Next, use another fully connected layer to keep the feature dimension at D, and map the output to the range [0, 1] through the Sigmoid function. Its input is the preliminary fusion feature fusion obtained in step 11.1;
[0084] g = Sigmoid(Linear(ReLU(BN(fusion))))
[0085] where Linear is the linear layer and BN is batch normalization;
[0086] Step 11.4: Combine the image feature V and e through g to obtain the fused feature com_feat of the image and text;
[0087] com_feat = (V * g) + (e * (1 - g))
[0088] Advantages of the present invention:
[0089] The advantages of the present invention compared with the prior art are as follows: The present invention proposes a clothing-changing pedestrian re-identification algorithm based on gated adaptive image-text feature fusion, aiming to overcome the limitations of traditional pedestrian re-identification methods in the face of complex scenarios such as target person's clothing change and background change by combining image and text description information. By introducing clothing change information from text descriptions and modal fusion of image features, accurate pedestrian recognition under different clothing conditions is achieved through a dynamic weighting and gated correction mechanism.
[0090] Other features and advantages of the present invention will be described in detail in the following specific implementation part. Brief Description of the Drawings
[0091] By describing the exemplary embodiments of the present invention in more detail with reference to the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more obvious. Among them, in the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0092] Figure 1 It is a flowchart of a clothing-changing pedestrian re-identification algorithm based on gated adaptive image-text feature fusion.
[0093] Figure 2 It is the backbone network diagram of the image feature attention module att_I and the text feature attention module att_T.
[0094] Figure 3 It is the network diagram of the weight calculation dynamic_scalar module.
[0095] Figure 4 It is the network diagram of the gated adaptive feature fusion GatedAdaptiveFeatureFusion module. Detailed Description of the Preferred Embodiments
[0096] The following will describe the preferred embodiments of the present invention in more detail with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein.
[0097] As Figure 1As shown in the figure, the present invention provides a clothing-changing pedestrian re-identification algorithm based on gated adaptive graphic and text feature fusion. The method includes the following steps:
[0098] Step 1: In order to use the CLIP large model as an image encoder, first adjust the smaller side of the source image to the CLIP input size input_dim, and then perform central cropping to generate a square image reference_image of input_dim×input_dim;
[0099] reference_image = {input_dim×input_dim}
[0100] Step 2: In order to use the CLIP large model as a text encoder, convert the input text description captions into a digital ID sequence that the model can understand, and specify the maximum length of the input text as 77 tokens to obtain the tokenized text input text_input;
[0101] text_input = tokenize(captions)
[0102] Step 3: In order to use the CLIP large model as an image encoder, adjust the smaller side of the target image to the CLIP input size input_dim, and then perform central cropping to generate a square image target_image of input_dim×input_dim;
[0103] target_image = {input_dim×input_dim}
[0104] Step 4: Input a batch of preprocessed source images reference_image into the image encoder of the CLIP large model to obtain source image feature vectors reference_features of B×D, where B is the batch size and D is the dimension of the feature vector;
[0105] reference_features = {B×D}
[0106] Step 5: Input a batch of tokenized text inputs text_input into the text encoder of the CLIP large model to obtain text feature vectors text_features of B×D, where B is the batch size and D is the dimension of the feature vector;
[0107] text_features = {B×D}
[0108] Step 6: Input a batch of preprocessed target images target_image into the image encoder of the CLIP large model to obtain a target image feature vector target_features of B×D, where B is the batch size and D is the dimension of the feature vector;
[0109] target_features = {B×D}
[0110] Step 7: Concatenate the source image feature vector reference_features and the text feature vector text_features extracted by the CLIP large model encoder. The concatenation operation occurs in the last dimension, i.e., the feature dimension. By combining the feature information of the image and the text, a merged feature vector raw_com_feat that integrates image visual information and text semantic information is obtained. Its shape is B×(2×D), where B is the batch size and D is the dimension of the feature vector. This merged feature vector can represent both the image content and the corresponding text description simultaneously, providing a richer feature input for subsequent multimodal tasks. At the same time, add a multi-head attention mechanism Multi-Head Attention module to further explore the relationship between the image and text features. Multi-head attention can focus on the mutual correlation between features in different subspaces, improving the fineness of the fusion. Its input is all raw_com_feat;
[0111] raw_com_feat = Concat(reference_features, text_features)
[0112] attn_output
[0113] = Multi_Head_Attention(raw_com_feat, raw_com_feat, raw_com_feat)
[0114] raw_com_feat = raw_com_feat + attn_output
[0115] raw_com_feat = {B×(2×D)}
[0116] Step 8: Use the image feature attention module att_I and the text feature attention module att_T to adjust the image features and text features respectively to obtain reference_features and text_features after attention adjustment;
[0117] referencefeatures = reference_features × att_I(raw_combined_features)
[0118] text_features = text_features × att_I(raw_combined_features)
[0119] Step 9: Concatenate the reference_features and text_features after attention adjustment along the last dimension to obtain a new combined feature new_combined_features. The dimension after concatenation is {B × (2 × D)}, where B is the batch size and D is the dimension of the feature vector;
[0120] new_combined_features = Concat(reference_features, text_features)
[0121] new_combined_features = {B × (2 × D)}
[0122] Step 10: Process the concatenated new_combined_features through the dynamic_scalar module for weight calculation, and output a dynamic scalar dynamic. This scalar represents the relative importance of the image features and text features, ranging from [0, 1]. Through weighted operation, the final synthetic feature com_features is generated. Here, dynamic is the weight of the image features, and 1 - dynamic is the weight of the text features;
[0123] dynamic = dynamic_scalar(new_combined_features)
[0124] Step 11: Weight-average the reference_features and text_features according to the weights of the dynamic scalar obtained in Step 10, and fuse them through the GatedAdaptiveFeatureFusion module, abbreviated as GAFF in the following formula, to generate the final synthetic feature com_feat, whose dimension is {B × D}, where B is the batch size and D is the dimension of the feature vector;
[0125] com_feat = GAFF(dynamic * image_features, (1 - dynamic) * text_features)
[0126] com_feat = {B×D}
[0127] Step 12: In the training stage, the final synthetic feature com_feat of the network is denoted as Feat s and the target image feature target_features is denoted as Feat tgt For the calculation of Loss, in the formula, N represents the batch size, κ represents the cosine similarity calculation, λ is a scaling factor used to adjust the logits range, the λ parameter is set to 100 to enhance the dynamic range of the logits, and in the prediction stage, it is judged whether the prediction is correct by calculating the distance between com_feat and target_features;
[0128]
[0129] Furthermore, for the image feature attention module att_I and the text feature attention module att_T in step 8, the specific method for processing each layer of data is as follows:
[0130] Step 8.1: The inputs of the image feature attention module att_I and the text feature attention module att_T are the concatenation of the image feature and the text feature. Assuming the dimension of the image feature is image_embed_dim and the dimension of the text feature is text_embed_dim, then the dimension of the concatenated input is image_embed_dim + text_embed_dim;
[0131] Step 8.2: Use the fully connected layer as the first layer of the image feature attention module att_I and the text feature attention module att_T. Its input is raw_combined_features with a dimension of D + D, and the output dimension is D. The role of this layer is to map the concatenated image and text features to a new space with the same dimension as the image feature. This process is equivalent to fusing two modalities (image and text) and laying the foundation for feature weighting in the next step through the transformation of this layer;
[0132] att_I_fc_layer_output = Linear(raw_combined_features)
[0133] att_T_fc_layer_output = Linear(raw_combined_features)
[0134] att_I_FC_layer_output = {B×D}
[0135] att_T_FC_layer_output = {B×D}
[0136] Step 8.3: The output after passing through the fully connected layer will undergo a non - linear transformation through a ReLU activation function. The role of the ReLU activation function is to introduce non - linearity, enabling the network to better fit complex mapping relationships;
[0137] att_I_FC_ReLU_output = ReLU(raw_combined_features)
[0138] att_T_FC_ReLU_output = ReLU(raw_combined_features)
[0139] Step 8.4: The output activated by ReLU will pass through a Dropout layer. Dropout is a technique to prevent overfitting. It randomly discards a certain proportion of neurons to help the model improve its generalization ability;
[0140] att_I_FC_Dropout_output = Dropout(att_I_FC_ReLU_output)
[0141] att_T_FC_Dropout_output = Dropout(att_T_FC_ReLU_output)
[0142] Step 8.5: The output enters the second fully connected layer. Its input dimension is D, and the output dimension is also D. The role of this layer is to further process the features transformed by the previous layers, ensure the consistency of the feature dimensions, and prepare data for the subsequent Sigmoid activation;
[0143] att_I_FC_layer_output = Linear(att_I_FC_Dropout_output)
[0144] att_T_FC_layer_output = Linear(att_T_FC_Dropout_output)
[0145] att_I_FC_layer_output = {B×D}
[0146] att_T_FC_layer_output = {B×D}
[0147] Step 8.6: The output passing through the second fully-connected layer will go through a Sigmoid activation function. The role of the Sigmoid function is to compress the output value into the range [0, 1]. In this way, we obtain a numerical value representing the weighted coefficient of the image features, and this coefficient is used to adjust the importance of the image features in subsequent steps;
[0148] atten_I_output = Sigmoid(att_I - FC_layer_output)
[0149] att_T_output = Sigmoid(att_T_FC_layer_output)
[0150] Furthermore, in the weight calculation dynamic_scalar module in Step 10, the specific method for processing each layer of data is as follows:
[0151] Step 10.1: The weight calculation dynamic_scalar module is a sequence composed of multiple neural network layers. The first layer is a fully-connected layer. The input is new_combined_features, and the input dimension is the sum of the image feature dimension and the text feature dimension, that is, D + D, and the output dimension is D;
[0152] Output 1 = Linear(new_combined_features)
[0153] Output 1 = {B × D}
[0154] Step 10.2: Use the ReLU activation function as the second layer of the weight calculation dynamic_scalar module. ReLU is a common activation function used to introduce non-linearity. It processes the input data, turning negative values to zero and keeping positive values unchanged;
[0155] Output 2 = ReLU(Output 1 )
[0156] Output 2 = {B × D}
[0157] Step 10.3: Use the Dropout regularization technique as the third layer of the weight calculation dynamic_scalar module. During training, randomly "discard" the outputs of some neurons to prevent overfitting. The role of this layer is to randomly discard the intermediate outputs of the neural network to improve the generalization ability of the model. It also does not change the data dimension, and the output dimension is still D;
[0158] Output 3 = Dropout(Output 2 )
[0159] Output 3 = {B×D}
[0160] Step 10.4: Use another fully connected layer as the fourth layer of the dynamic_scalar module. This layer maps the previous output (with dimension D) to a 1D output, obtaining a scalar value. The role of this layer is to compress the data dimension through the fully connected layer, from D to 1;
[0161] Output 4 = Linear(Output 3 )
[0162] Output 4 = {B×1}
[0163] Step 10.5: Use the Sigmoid activation function as the fifth layer of the dynamic_scalar module. The Sigmoid activation function compresses the input value into the range [0, 1], usually used to calculate probabilities or weighting coefficients. The role of this layer is to map the scalar value obtained in the previous layer to a value between 0 and 1 through the Sigmoid function, representing a dynamic weighting coefficient, and its dimension is B×1;
[0164] dynamic = Sigmoid(Output 4 )
[0165] dynamic = {B×1}
[0166] Furthermore, in the GatedAdaptiveFeatureFusion module in Step 11, the specific method for data processing in each layer is as follows:
[0167] Step 11.1: Use batch normalization, ReLU activation function, and linear layer as the first part of the GatedAdaptiveFeatureFusion module. This part is used for the preliminary fusion of text feature T and image feature V. Its input is text feature T and image feature V, and the output is the preliminary fusion feature fusion;
[0168] fusion = Linear(ReLU(BN(T, V)))
[0169] where Linear is the linear layer and BN is batch normalization;
[0170] Step 11.2: Reduce the input feature dimension from D to D / 2 through a fully connected layer, then perform batch normalization to standardize the data. Subsequently, apply the ReLU activation function to introduce non-linearity. Next, another fully connected layer keeps the feature dimension at D / 2, and batch normalization and activation function processing are performed again. Finally, a fully connected layer restores the feature dimension to the original D to form the final output. By gradually reducing the feature dimension, adding non-linear activation, and then expanding back to the original dimension, redundant information is reduced, and ultimately the expressive ability of the feature representation is improved. The input is the preliminary fusion feature fusion obtained in Step 11.1;
[0171] e = Linear(ReLU(BN(Linear(ReLU(BN(Linear(fusion)))))))
[0172] where Linear is the linear layer and BN is batch normalization;
[0173] Step 11.3: Standardize the data through batch normalization. Then, apply the specified ReLU activation function to introduce non-linear transformation. Next, use another fully connected layer to maintain the feature dimension at D, and map the output to the range [0, 1] through the Sigmoid function. The input is the preliminary fusion feature fusion obtained in Step 11.1;
[0174] g = Sigmoid(Linear(ReLU(BN(fusion))))
[0175] where Linear is the linear layer and BN is batch normalization;
[0176] Step 11.4: Combine the image feature V and e through g to obtain the fused feature com_feat of the image and text;
[0177] com_feat = (V * g) + (e * (1 - g))
Claims
1. A pedestrian re-identification algorithm based on gated adaptive image and text feature fusion, characterized in that: The method comprises the following steps: Step 1. In order to use the CLIP large model as an image encoder, first adjust the smaller side of the source image to the CLIP input size input_dim, and then perform center cropping to generate a square image reference_image of input_dim×input_dim; reference_image={input_dim×input_dim} Step 2: In order to use the CLIP large model as a text encoder, convert the input text description captions into a sequence of numeric IDs that the model can understand, and specify the maximum length of the input text as 77 tokens to obtain the tokenized text input text_input; teXt_input=tokenize(captions) Step 3. In order to use the CLIP large model as an image encoder, the smaller side of the target image is adjusted to the CLIP input size input_dim, and then the center is cropped to generate a square image target_image of input_dim×input_dim; target_image={input_dim×input_dim} Step 4: Input a batch of preprocessed source images reference_image into the image encoder of the CLIP large model to obtain a B×D source image feature vector reference_features, where B is the batch size and D is the dimension of the feature vector; reference_features={B×D} Step 5: Input a batch of tokenized text input text_input into the text encoder of the CLIP large model to obtain a B×D text feature vector text_features, where B is the batch size and D is the dimension of the feature vector; text_features = {B×D} Step 6: Input a batch of preprocessed target images target_image into the image encoder of the CLIP large model to obtain B×D target image feature vectors target_features, where B is the batch size and D is the dimension of the feature vector; target_features = {B×D} Step 7: Concatenate the source image feature vector reference_features and the text feature vector teXt_features extracted by the CLIP large model encoder. The concatenation operation occurs in the last dimension, i.e., the feature dimension. By combining the feature information of the image and the text, a merged feature vector raw_com_feat that combines the image visual information and the text semantic information is obtained. Its shape is B×(2×D), where B is the batch size and D is the dimension of the feature vector. This merged feature vector can simultaneously represent the image content and the corresponding text description, providing richer feature input for subsequent multimodal tasks. At the same time, a multi-head attention mechanism Multi-Head Attention module is added to further explore the relationship between image and text features. Multi-head attention can focus on the correlation between features in different subspaces and improve the fineness of fusion. Its input is raw_com_feat. raw_com_feat=Concat(reference_features,text_features) raw_com_feat =Multi_Head_Attention(raw_com_feat,raw_com_feat,raw_com_feat) raw_com_feat = {B×(2×D)} Step 8: Use the image feature attention module att_I and the text feature attention module att_T to adjust the image features and text features respectively, and obtain reference_features and text_features after attention adjustment; reference_features=reference_features×att_I(raw_combined_features) text_features=text_features×att_I(raw_combined_features) Step 9: Concatenate the reference_features and text_features after attention adjustment along the last dimension to obtain a new merged feature new_combined_features. The concatenated dimension is {B×(2×D)}, where B is the batch size and D is the dimension of the feature vector. new_combined_features=Concat(reference_features,text_features) new_combined_features={B×(2×D)} Step 10: Process the concatenated new_combined_features through the weight calculation dynamic_scalar module, and output a dynamic scalar dynamic, which represents the relative importance of image features and text features, ranging from [0, 1]. Through weighted operations, the synthetic features com_features are finally generated. Here, dynamic is the weight of the image feature, and 1-dynamic is the weight of the text feature. dynamic=dynamic_scalar(new_combined_features) Step 11: perform weighted average of reference_features and text_features according to the weight of the dynamic scalar obtained in step 10, and fuse them through the gated adaptive feature fusion GatedAdaptiveFeatureFusion module, referred to as GAFF in the following formula, to generate the final synthetic feature com_feat, whose dimension is {B×D}, where B is the batch size and D is the dimension of the feature vector; com_feat=GAFF(dynamic*image_features, (1-dynamic)*text_features) com_feat = {B×D} Step 12: The final synthetic feature com_feat of the network during the training phase is recorded as Feat s And the target image features target_features are recorded as Feat tgt Used for Loss calculation. In the formula, N represents the batch size, κ represents the cosine similarity calculation, and λ is the scaling factor used to adjust the range of logits. The λ parameter is set to 100 to enhance the dynamic range of logits. In the prediction stage, the distance between com_feat and target_features is calculated to determine whether the prediction is correct.
2. The pedestrian re-identification algorithm based on gated adaptive image-text feature fusion according to claim 1 is characterized in that The network process of the image feature attention module att_I and the text feature attention module att_T in step 8 is as follows: Step 8.1, the input of the image feature attention module att_I and the text feature attention module att_T is the concatenation of the image feature and the text feature. Assuming that the dimension of the image feature is image_embed_dim and the dimension of the text feature is text_embed_dim, then the concatenated input dimension is image_embed_dim+text_embed_dim; Step 8.2: Use the fully connected layer as the first layer of the image feature attention module att_I and the text feature attention module att_T. Its input is raw_combined_features, with a dimension of D+D and an output dimension of D. The function of this layer is to map the spliced image and text features to a new space with the same dimension as the image features. This process is equivalent to fusing the two modalities (image and text) together, and the transformation of this layer lays the foundation for the next step of feature weighting. att_I_fc_layer_output=Linear(raw_combined_features) att_T_fc_layer_output=Linear(raw_combined_features) att_I_FC_layer_output={B×D} att_T_FC_layer_output={B×D} Step 8.3: The output after the fully connected layer will be transformed nonlinearly through a ReLU activation function. The function of the ReLU activation function is to introduce nonlinearity so that the network can better fit complex mapping relationships. att_I_FC_ReLU_output=ReLU(raw_combined_features) att_T_FC_ReLU_output=ReLU(raw_combined_features) Step 8.4: The output of ReLU activation will pass through a Dropout layer. Dropout is a technique to prevent overfitting. It randomly discards a certain proportion of neurons to help the model improve its generalization ability. att_I_FC_Dropout_output=Dropout(att_I_FC_ReLU_output) att_T_FC_Dropout_output=Dropout(att_T_FC_ReLU_output) Step 8.5, the output enters the second fully connected layer, whose input dimension is D and output dimension is also D. The function of this layer is to further process the features transformed by the previous layer, ensure the consistency of feature dimensions, and prepare data for subsequent Sigmoid activation; att_I_FC_layer_output=Linear(att_I_FC_Dropout_output) att_T_FC_layer_output=Linear(att_T_FC_Dropout_output) att_I_FC_layer_output={B×D} att_T_FC_layer_output={B×D} Step 8.6: The output of the second fully connected layer passes through a Sigmoid activation function. The Sigmoid function compresses the output value into the range of [0, 1]. In this way, we get a value representing the weighted coefficient of the image feature. This coefficient is used to adjust the importance of the image feature in subsequent steps. att_I_output=Sigmoid(att_I_FC_layer_output) att_T_output=Sigmoid(att_T_FC_layer_output) 3. The pedestrian re-identification algorithm based on gated adaptive image-text feature fusion according to claim 1 is characterized in that The network process of the weight calculation dynamic_scalar module in step 10 is as follows: Step 10.
1. Weight calculation The dynamic_scalar module is a sequence of multiple neural network layers. The first layer is a fully connected layer. The input is new_combined_features. The input dimension is the sum of the image feature dimension and the text feature dimension, that is, D+D. The output dimension is D. Output1=Linear(new_combined_features) Output1=(B×D} Step 10.2: Use the ReLU activation function as the weight calculation for the second layer of the dynamic_scalar module. ReLU is a common activation function used to introduce nonlinearity. It processes the input data, converting negative values to zero and keeping positive values unchanged. Output2 = ReLU (Output1) Output2=(B×D} Step 10.3: Use the Dropout regularization technique as the third layer of the weight calculation dynamic_scalar module. During the training process, the output of some neurons is randomly discarded to prevent overfitting. This layer randomly discards the intermediate outputs of the neural network to improve the generalization ability of the model. It also does not change the dimension of the data. The dimension of the output is still D. Output3=Dropout(Output2) Output3={B×D} Step 10.4, use another fully connected layer as the fourth layer of the weight calculation dynamic_scalar module. This layer maps the previous output dimension D to a 1-dimensional output to obtain a scalar value. The function of this layer is to compress the data dimension from D to 1 through the fully connected layer; Output4=Linear(Output3) Output4={B×1} Step 10.
5. Use the Sigmoid activation function as the weight calculation in the fifth layer of the dynamic_scalar module. The Sigmoid activation function compresses the input value into the range of [0, 1] and is usually used to calculate probabilities or weighting coefficients. The function of this layer is to map the scalar value obtained in the previous layer to a value between 0 and 1 through the Sigmoid function, indicating a dynamic weighting coefficient. Its dimension is B×1. dynamic=Sigmoid(Output4) dynamic={B×1} 4. The pedestrian re-identification algorithm based on gated adaptive image-text feature fusion according to claim 1 is characterized in that The network process of the gated adaptive feature fusion GatedAdaptiveFeatureFusion module in step 11 is as follows: Step 11.1, batch normalization, ReLU activation function, and linear layer are used as the first part of the gated adaptive feature fusion GatedAdaptiveFeatureFusion module. This part is used for the preliminary fusion of text features T and image features V. Its input is text features T and image features V, and its output is preliminary fusion feature fusion; fusion=Linear(ReLU(BN(T,V))) Among them, Linear is the linear layer, and BN is batch normalization; Step 11.2: Reduce the input feature dimension from D to D / 2 through a fully connected layer, then perform batch normalization to standardize the data, then apply the ReLU activation function to introduce nonlinearity, then another fully connected layer keeps the feature dimension at D / 2, and performs batch normalization and activation function processing again, and finally, restore the feature dimension to the original D through a fully connected layer to form the final output. By gradually reducing the feature dimension, adding nonlinear activation, and expanding back to the original dimension, redundant information is reduced, and the expressive power of the feature representation is finally improved. The input is the preliminary fusion feature fusion obtained in step 11.1; e=Linear(ReLU(BN(Linear(ReLU(BN(Linear(fusion))))))) Among them, Linear is the linear layer, and BN is batch normalization; Step 11.3, standardize the data through batch normalization, then apply the specified ReLU activation function to introduce nonlinear transformation, then use a fully connected layer to maintain the feature dimension at D, and map the output to the range of [0, 1] through the Sigmoid function, whose input is the preliminary fusion feature fusion obtained in step 11.1; g=Sigmoid(Linear(ReLU(BN(fusion)))) Among them, Linear is the linear layer, and BN is batch normalization; Step 11.4, combine the image features V and e through g to obtain the fusion feature com_feat of the image and text; com_feat=(V*g)+(e*(1-g))
Citation Information
Cited By
Visual feature-based character consensus method and system
CN121259265A