A Transformer-based Cross-modal Person Re-identification Method

By adopting a Transformer-based method in cross-modal pedestrian recognition, data augmentation and grayscale transformation of images and performing feature fusion, the problem of matching failure in the existing methods is solved, and the matching accuracy and feature capture ability are improved.

CN116863501BActive Publication Date: 2025-06-17BEIJING TRIONIS PETRO-TECH DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310711644.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2025-06-17
Estimated Expiration
2043-06-15

AI Technical Summary

Technical Problem

The existing cross-modal pedestrian recognition method has the problem of matching failure when processing images of different modalities, and lacks efficient and convenient processing methods.

Method used

The cross-modal pedestrian re-identification method based on Transformer is adopted. By obtaining the cross-modal pedestrian re-identification data set, the image is enhanced and grayscale transformed, and respectively sent to the modal unique linear integration module, and feature fusion is performed through the shared Transformer encoding layer.

Benefits of technology

The matching accuracy of cross-modal image pairs is improved, the model structure is simplified, and the capture ability of modal-independent features is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116863501B_ABST
    Figure CN116863501B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-modal pedestrian re-identification method based on Transformer, which relates to the field of artificial intelligence technology. The present invention at least includes the following steps: S1: Obtain a cross-modal pedestrian re-identification data set and perform data augmentation on the two-modal pedestrian images. The model structure used in the present invention is simple and does not have complex component designs. Moreover, the model does not cross-modally generate model features that have not been extracted. Instead, it directly collects more of its own features as much as possible, and then fuses the features of different modalities and extracts the feature information of those common modality-independent aspects. The idea is simpler and the matching accuracy is higher; the grayscale data augmentation strategy used in the present invention is a plug-and-play augmentation strategy, which can greatly improve the model's ability to capture modality-independent features without changing the original model architecture, and significantly improve the matching accuracy of cross-modal image pairs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and specifically to a cross-modal pedestrian re-identification method based on Transformer. Background Art

[0002] Pedestrian re-identification aims to retrieve a specific pedestrian among multiple non-overlapping cameras. Most existing studies are based on a set of strong assumptions, that is, the pedestrians in the retrieved images all contain a complete human torso and are all single-modal visible light (RGB) images. However, in the real world, to improve the poor shooting quality of RGB cameras under low light conditions, infrared (IR) mode cameras are introduced, and data of two shooting modalities will be obtained during shooting. Traditional models face the problem of failure when matching images of different modalities. The above scenarios are common and difficult to avoid in reality, which makes pedestrian re-identification in the real world face many challenges. Studying methods to deal with the above situations has far-reaching significance in both scientific research and industrial fields.

[0003] Cross-modal based pedestrian re-identification has received extensive attention from all walks of life in recent years. Existing advanced cross-modal pedestrian re-identification methods are based on the modal shared feature learning route. This type of method first extracts their corresponding features from data of different modalities, and then uses the idea of feature mapping or modal disentanglement to learn modal shared features. Although the research route of this type of method is clear;

[0004] However, existing models are generally complex, and the design ideas are difficult to understand. There is a lack of an efficient and convenient method to process cross-modal pedestrian image data. Among many technical routes, the strategy based on modal shared feature learning is a type that has been studied the most extensively and has a relatively good matching effect. According to the summary, there are few studies on cross-modal pedestrian re-identification tasks based on pure ViT methods.

[0005] Therefore, a method with a moderate structural complexity and capable of improving the matching accuracy of cross-modal image pairs is needed. Summary of the Invention

[0006] The purpose of the present invention is to provide a cross-modal pedestrian re-identification method based on Transformer.

[0007] To achieve the above purpose, the present invention provides the following technical solution: A cross-modal pedestrian re-identification method based on Transformer, at least including the following steps:

[0008] S1: Obtain a cross-modal pedestrian re-identification dataset, and perform data augmentation on pedestrian images of two modalities;

[0009] S2: Enhance the grayscale of the original image to be input into the RGB image branch;

[0010] S3: Split the dataset images into two parts, the RGB modality and the IR modality, according to the modality. Then, send the two divided parts and the part after grayscale enhancement into the modality-specific linear integration module respectively to obtain the corresponding sequences.

[0011] S4: Send the sequences obtained in S3 into the S specific Transformer encoder modules of the corresponding modality respectively.

[0012] S5: Combine the image sets after passing through the two specific branches, send them into the L - S shared Transformer encoding layers, and calculate the overall loss according to the corresponding loss function.

[0013] Preferably, the pedestrian images in the cross-modal pedestrian re-identification dataset obtained in S1 all contain the complete human torso.

[0014] Preferably, the data augmentation in S1 at least includes the following steps:

[0015] For the pedestrian images of the two modalities, at least use one of horizontal flipping, edge padding, grayscale transformation, uniform cropping, random cropping, random cropping and stitching, random slight angle rotation, random position transformation, random noise addition, random erasing, and image sharpening for data augmentation.

[0016] Convert to a tensor, normalize it, and then standardize it.

[0017] Preferably, the grayscale enhancement in S2 at least includes the following steps:

[0018] Perform grayscale transformation on the image that will originally be input to the RGB image branch, and set its corresponding label to be the same as the original RGB image.

[0019] Input the original RGB image and the grayscale-transformed image into the branch specific to the RGB modality together.

[0020] The specific formula is:

[0021] Gray = R×0.299 + G×0.587 + B×0.114.

[0022] Preferably, for each batch of images input to the model in S3, which contains pedestrian pictures from two modalities, each modality accounts for half of the total number of images in this batch. Input the divided IR images into the IR modality-specific linear integration module, and send the RGB images and the grayscale-enhanced images into the RGB modality-specific linear integration module respectively.

[0023] For the modality-specific linear integration module, given an input image \(x\in\mathbb{R}\) H×W×C , where \(H\), \(W\), and \(C\) respectively refer to the height, width, and number of channels of the image;

[0024] Use the overlapping sampling strategy to process the input image \(x\) to obtain better local neighborhood representation ability;

[0025] Set the sampling step size to \(s\) and the side length of the sampling block size to \(p\), then the input image \(x\) is divided into \(N\) fixed-size sub-blocks \([x i |i = 1, 2, …, N]\), and the calculation formula for \(N\) is:

[0026]

[0027] where represents taking the lower boundary of the corresponding result, \(X H and \(X W respectively refer to the number of sub-blocks in the height and width axes;

[0028] When \(s\) is less than \(p\), the effect of overlapping sampling is obtained, and when \(s\) is smaller, the overlapping sampling area is larger;

[0029] Adopt the same settings as Vision Transformer. After linearly mapping the \(N\) sub-blocks, a classification identifier is added before the first sub-block to capture global information, and then a learnable position encoding integration \(E p is added to each sub-block to maintain spatial information. The final output is:

[0030] \(Z_0=[X CLS ,x 1 E,x 2 E,…,x N E]+E P

[0031] where \(X CLS \in\mathbb{R} 1×D represents the classification identifier, \(E\in\mathbb{R} P×P×C×D represents the transformation matrix for linearly transforming the sampled \(x i \((i\in[1,N],i\in\mathbb{Z})\) sub-blocks, and \(E P \in\mathbb{R} (N+1)×D represents the position encoding integration.

[0032] Preferably, for the modality-specific branches in S3 and S4, the network structures are exactly the same but with non-shared weight configurations. Each branch contains a module for linearly transforming the image sub-blocks and \(S\) identical encoding modules for encoding the sub-blocks. The specific change process formula:

[0033] T m = P(I m )

[0034]

[0035] where P(·) refers to the linear transformation operation, and E(·) refers to the encoding operation of the Transformer;

[0036] Combine the image sets after passing through the two specific branches and and send them into S5 to obtain L - S shared Transformer encoding layers.

[0037] Preferably, each layer of the Transformer encoding layer in S4 and S5 is composed of a multi - head self - attention mechanism and a multi - layer perceptron model. Apply hierarchical regularization before each multi - head self - attention mechanism module and multi - layer perceptron module, and apply a residual connection between the above two modules;

[0038] The Transformer encoding layer consists of two stages. The formula for the data input from the previous layer passing through the multi - head self - attention mechanism is:

[0039] Z i ' = Z i-1 + MSA(LN(Z i-1 )) l ∈ 1, 2, …, L

[0040] The output formula of the i - th Transformer encoding layer is:

[0041] Z i = Z i '+ MLP(LN(Z i ')) l ∈ 1, 2, …, L

[0042] In S5, complete the screening and fusion of the modality - specific features mined separately in the modality - specific branches before, select the common modality - independent features as the final pedestrian feature representation. Among them, both the batch hard sampling triplet loss and the cross - entropy loss are used for the first classification identifier of the last Transformer encoding layer. Before using the cross - entropy loss, the batch regularization bottleneck strategy is also used to synchronize the convergence of the two different losses;

[0043] For the similarity calculation between two images, use one of the cosine distance and the Euclidean distance as the similarity value of the two images;

[0044] For two images x1, x2, their similarity value is denoted as

[0045] In the triplet loss, the network f receives three images, namely the anchor (x a ), the positive sample (x p ), and the negative sample (x n ). Among them, the positive sample and the anchor form a positive sample pair, and the negative sample and the anchor form a negative sample pair. Then the triplet loss formula is as follows:

[0046]

[0047] where α refers to the margin threshold parameter, which is used to control the intensity of optimization;

[0048] The cross-entropy loss and the total loss function of S5 are respectively expressed as:

[0049]

[0050] L i = L ID + L tri

[0051] where p(i) refers to the probability value that the input image x belongs to the i-th (i ∈ [1, N], i ∈ Z) pedestrian identity, and q(i) refers to the true label of the image. If the true label of the image x is i, then q(i) = 1; otherwise, q(i) = 0.

[0052] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0053] 1. The model structure used in the present invention is simple and does not have complex component designs. Moreover, the model does not cross-modally generate model features that have not been extracted. Instead, it directly collects more of its own features as much as possible, and then fuses the features of different modalities and extracts the feature information of those common modality-independent aspects. The idea is simpler and the matching accuracy is higher;

[0054] 2. The gray-scale data augmentation strategy used in the present invention is a plug-and-play augmentation strategy, which can significantly improve the model's ability to capture modality-independent features without changing the original model architecture, and significantly improve the matching accuracy of cross-modal image pairs. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1It is a diagram of a cross-modal pedestrian re-identification benchmark model based on Transformer according to the present invention;

[0057] Figure 2 It is a diagram of a cross-modal pedestrian re-identification model based on Transformer using grayscale image data augmentation according to the present invention. Specific implementation manner

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0059] Please refer to Figure 1 - Figure 2 , a cross-modal pedestrian re-identification method based on Transformer, at least including the following steps:

[0060] S1: Obtain a cross-modal pedestrian re-identification data set, and perform data augmentation on the pedestrian images of the two modalities;

[0061] The pedestrian images in the cross-modal pedestrian re-identification data set obtained in S1 all contain the complete human torso;

[0062] The data augmentation in S1 at least includes the following steps:

[0063] For the pedestrian images of the two modalities, at least one of horizontal flipping, edge padding, grayscale transformation, unified cropping, random cropping, random cropping and stitching, random slight angle rotation, random position transformation, random noise addition, random erasing, and image sharpening is used for data augmentation, and after being converted into tensors and normalized, standardization is performed;

[0064] The following data augmentation methods are specifically used:

[0065] Pad the image edges, and set the padding value to 10;

[0066] Perform central cropping on the image, and set H×W to 288×144;

[0067] Perform random horizontal flipping on the image;

[0068] Convert it into a tensor, normalize it, and then perform standardization, and the data format H×W×C will be changed to C×H×W;

[0069] Finally, perform random erasing, set the probability value to 0.5, and set the maximum number of erased pixels to 1.

[0070] S2: Perform grayscale enhancement on the original image to be input into the RGB image branch;

[0071] The grayscale enhancement in S2 includes at least the following steps:

[0072] Perform grayscale transformation on the image originally to be input to the RGB image branch, and set its corresponding label to be the same as the original RGB image;

[0073] Input the original RGB image and the grayscale-transformed image into the branch unique to the RGB modality together; Briefly, converting an RGB image to a grayscale image is to perform a specific set of superimpositions on the values of the three channels in each pixel. The specific formula is:

[0074] Gray = R×0.299 + G×0.587 + B×0.114

[0075] S3: Split the dataset images into two parts, the RGB modality and the IR modality, according to the modality, and then send the two divided parts and the part after grayscale enhancement into the modality-specific linear integration module respectively to obtain the corresponding sequences;

[0076] For each batch of images input to the model in S3, which contains pedestrian pictures from two modalities, and each modality accounts for half of the total number of images in this batch. Input the divided IR images into the IR-modality-specific linear integration module, and send the RGB images and the grayscale-enhanced images into the RGB-modality-specific linear integration module respectively;

[0077] For the modality-specific linear integration module, given an input image x ∈ R H×W×C , where H, W, and C respectively refer to the height, width, and number of channels of the image;

[0078] Use the overlapping sampling strategy to process the input image x to obtain better local neighborhood representation ability;

[0079] Set the sampling step size to s and the side length of the sampling block to p, then the input image x is divided into N fixed-size sub-blocks [x i |i = 1, 2, …, N], and the calculation formula for N is:

[0080]

[0081] where represents taking the lower boundary of the corresponding result, X H and X W respectively refer to the number of sub-blocks in the height and width axes;

[0082] When s is less than p, the effect of overlapping sampling is obtained, and when s is smaller, the overlapping sampling area is larger, but more computing resources will also be consumed;

[0083] Using the same settings as the Vision Transformer, after linearly mapping N patches, the model adds a classification token before the first patch to capture global information and then adds a learnable position embedding E to each patch p to maintain spatial information. The final output is:

[0084] Z0 = [X CLS , x 1 E, x 2 E, …, x N E] + E P

[0085] where X CLS ∈R 1×D represents the classification token, and E ∈ R P×P×C×D denotes the transformation matrix for linear transformation on the sampled x i (i ∈ [1, N], i ∈ Z) patches, and E P ∈R (N+1)×D represents the position embedding

[0086] S4: Feed the sequences obtained in S3 into S specific Transformer encoder modules of the corresponding modalities respectively;

[0087] In the modality-specific branches of S3 and S4, the network structures are exactly the same, but the weight configurations are not shared. Each branch contains a module for linearly transforming the image patches and S identical encoding modules for encoding the patches. The specific transformation process formula:

[0088] T m = P(I m )

[0089]

[0090] where P(·) refers to the linear transformation operation and E(·) refers to the encoding operation of the Transformer;

[0091] Combine the image sets after passing through the two specific branches and and feed them into S5 to obtain L - S shared Transformer encoding layers

[0092] S5: Combine the image sets after passing through the two specific branches and feed them into L - S shared Transformer encoding layers, and calculate the overall loss according to the corresponding loss function.

[0093] Each layer of the Transformer encoding layer in S4 and S5 consists of a multi-head self-attention mechanism and a multi-layer perception (MLP) model. Layer normalization (Layer-Norm) is applied before each multi-head self-attention mechanism module and MLP module, and a residual connection is applied between the above two modules;

[0094] The Transformer encoding layer consists of two stages. The formula for the data input from the previous layer passing through the multi-head self-attention mechanism is:

[0095] Z i ′ = Z i-1 + MSA(LN(Z i-1 )) l ∈ 1, 2, …, L

[0096] The output formula of the i-th Transformer encoding layer is:

[0097] Z i = Z i ′ + MLP(LN(Z i ′)) l ∈ 1, 2, …, L

[0098] In S5, the screening and fusion of the modality-specific features previously mined in the modality-specific branches are completed, and the common modality-independent features are selected as the final pedestrian feature representation. Among them, the batch hard sampling triplet loss and cross-entropy loss are both used for the first classification identifier of the last Transformer encoding layer. Before using the cross-entropy loss, the batch normalization bottleneck strategy is also used to synchronize the convergence of the two different losses;

[0099] For the similarity calculation between two images, one of the cosine distance and Euclidean distance is used as the similarity value of the two images;

[0100] For two images x1, x2, their similarity value is denoted as

[0101] In the triplet loss, the network f receives three images, which are the anchor (x a ), the positive sample (x p ), and the negative sample (x n ). Among them, the positive sample and the anchor form a positive sample pair, and the negative sample and the anchor form a negative sample pair. Then the triplet loss formula is:

[0102]

[0103] where α refers to the boundary threshold parameter, which is used to control the intensity of optimization;

[0104] The cross-entropy loss and the total loss function of S5 are respectively expressed as:

[0105]

[0106] L i = L ID + L tri

[0107] Where p(i) refers to the probability value that the input image x belongs to the i-th (i ∈ [1, N], i ∈ Z) pedestrian identity, and q(i) refers to the true label of the image. If the true label of the image x is i, then q(i) = 1, otherwise q(i) = 0.

[0108] In the present invention, the sampling step of ViT is set to 12, the number of images of each modality in each batch is 32 (from 8 pedestrian identities, 4 images for each identity), and a total of 80 rounds of iterative training are performed. The learning rate of the final fully-connected classification layer of the model is 10 times that of other modules. The learning rate of the fully-connected classification layer is the base learning rate, and the base learning rate is set to 0.1. The learning rate adopts a preheating strategy, the number of linear preheating rounds is 10 rounds, the initial preheating learning rate is 0.1 times the base learning rate, and the learning rate is adjusted to 0.1 times the original at the 20th and 50th rounds of iteration respectively. The optimizer adopts stochastic gradient descent, the momentum value is set to 0.9, and the weight decay is set to 5e-4. In the model of this chapter, the number S of specific Transformer encoders in the modality-specific branch is set to 0. The model is built using Pytorch and the training and inference processes are completed.

[0109] According to a cross-modal pedestrian re-identification method based on Transformer of the present invention, the model structure used is simple and there is no complex component design. The model of the present invention does not cross-modally generate model features that have not been extracted from each other, but directly collects more of its own features as much as possible, and then fuses the features of different modalities and extracts the feature information of those common modality-independent aspects. The idea is simpler and the matching accuracy is higher. The grayscale data augmentation strategy used in the present invention is a plug-and-play augmentation strategy, which can greatly improve the model's ability to capture modality-independent features without changing the original model architecture, and significantly improve the matching accuracy of cross-modal image pairs.

[0110] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A cross-modal person re-identification method based on Transformer, characterized in that: At least include the following steps: S1: Obtain a cross-modal pedestrian re-identification dataset, and perform data augmentation on pedestrian images of two modalities; S2: Enhance the grayscale of the image that will originally be input into the RGB image branch; The grayscale enhancement in S2 at least includes the following steps: Perform grayscale transformation on the image that will originally be input into the RGB image branch, and set its corresponding label to be the same as the original RGB image; Jointly input the original RGB image and the grayscale-transformed image into the branch unique to the RGB modality; The specific formula is: Gray = R × 0.299 + G × 0.587 + B × 0.114 S3: Split the dataset images into two parts, the RGB modality and the IR modality, according to the modality, and then send the two divided parts and the part after grayscale enhancement into the modality-specific linear integration modules respectively to obtain corresponding sequences; For each batch of images input into the model in S3, it contains pedestrian pictures from two modalities, and each modality accounts for half of the total number of images in this batch. Input the divided IR images into the IR modality-specific linear integration module, and send the RGB images and the grayscale-enhanced images into the RGB modality-specific linear integration module respectively; For the modality-specific linear integration module, given an input image \(x\in A\) H×W×C , where \(H\), \(W\), and \(C\) respectively refer to the height, width, and number of channels of the image; Use the strategy of overlapping sampling to process the input image x to obtain better local proximity representation ability; Set the sampling step size as s and the side length of the sampling block size as p. Then the input image x is divided into N blocks of fixed size [x i |i = 1, 2, …, N]. The calculation formula for N is: Among them represents taking the lower boundary of the corresponding result, X H and X W respectively refer to the number of blocks in the height and width axes; When s is less than p, the effect of overlapping sampling is obtained, and when s is smaller, the overlapping sampling area is larger; With the same settings as the Vision Transformer, after linearly mapping N patches, the model adds a classification identifier before the first patch to capture global information and then adds a learnable positional encoding for each patch to integrate E P to maintain spatial information. The final output is: Z0 = [X CLS , x 1 E, x 2 E, …, x N E] + E P Wherein X CLS ∈ A 1×D represents a classification identifier, E ∈ A p×p×C×D indicates the linear transformation matrix performed on the sampled x i , i ∈ [1, N], i ∈ Z, on the block, E P ∈ A (N+1)×D represents the positional encoding integration; S4: Send the sequences obtained in S3 into S specific Transformer encoder modules of the corresponding modality respectively; S5: Combine the image sets after passing through the two specific branches, send them into L - S shared Transformer encoding layers, and calculate the overall loss according to the corresponding loss function, where L is the total number of Transformer encoding layers in the entire model.

2. The cross-modal person re-identification method based on Transformer according to claim 1, characterized in that: The pedestrian images in the cross-modal pedestrian re-identification dataset obtained in S1 all contain the complete human torso.

3. The cross-modal person re-identification method based on Transformer according to claim 1, characterized in that: The data augmentation in S1 at least includes the following steps: For pedestrian images of two modalities, perform data augmentation using at least one of horizontal flipping, edge padding, grayscale transformation, uniform cropping, random cropping, random cropping and stitching, random slight angle rotation, random position transformation, random noise addition, random erasing, and image sharpening; Convert to a tensor, perform normalization, and then perform standardization.

4. The cross-modal person re-identification method based on Transformer according to claim 1, characterized in that: The modality-specific branches in S3 and S4 have exactly the same network structure but do not share weight configurations. Each branch contains a module for performing linear transformation on image blocks and S identical encoding modules for encoding the blocks. The specific transformation process formula: T m = P(I m ) Where P(·) refers to the linear transformation operation, and E(·) refers to the encoding operation of the Transformer; The set of images after passing through two specific branches and are combined and sent into S5 to obtain L-S shared Transformer encoding layers.

5. The cross-modal person re-identification method based on Transformer according to claim 1, characterized in that: Each layer of the Transformer encoding layer in S4 and S5 is composed of a multi-head self-attention mechanism and a multi-layer perceptron model. Layer normalization is applied before each multi-head self-attention mechanism module and multi-layer perceptron module, and a residual connection is applied between the above two modules; The Transformer encoding layer consists of two stages. The formula for the data input from the previous layer passing through the multi-head self-attention mechanism is: Z′ c = Z c-1 + MSA(LN(Z c-1 )) c ∈ 1, 2, …, L The output formula of the i-th Transformer encoding layer is: Z c = Z' c + MLP(LN(Z' c )) c ∈ 1, 2, …, L In S5, the screening and fusion of the modality-specific features previously mined in the modality-specific branches are completed, and the common modality-independent features are selected as the final pedestrian feature representation. Among them, the batch hard sampling triplet loss and cross-entropy loss are used for the first classification identifier of the last Transformer encoding layer. Before using the cross-entropy loss, the batch normalization bottleneck strategy is also used to synchronize the convergence of the two different losses; For the similarity calculation between two images, one of the cosine distance and Euclidean distance is used as the similarity value of the two images; For two images x1 and x2, their similarity value is denoted as In the triplet loss, the network f receives three images, which are the anchor (x a ), the positive sample (x p ), and the negative sample (x n ). Among them, the positive sample and the anchor form a positive sample pair, and the negative sample and the anchor form a negative sample pair. Then the triplet loss formula is as follows: Among them, α refers to the boundary threshold parameter, which is used to control the intensity of optimization; The cross-entropy loss and the total loss function in S5 are respectively expressed as: L = L ID + L tri Among them, p(v) refers to the probability value that the input image x belongs to the v-th, v ∈ [1, N], v ∈ Z, pedestrian identity, and q(v) refers to the true label of the image. If the true label of the image x is v, then q(v) = 1, otherwise q(v) = 0.

Citation Information

Patent Citations

  • Diurethanes containing sulphur, process for their preparation and their use as herbicides

    EP0000030A1

  • Near infrared-visible light cross-modal double-current pedestrian re-identification method and system

    CN114220124A

  • Pedestrian re-identification method, related equipment and readable storage medium

    CN114445859A