A pedestrian re-identification method based on graph convolution network multi-source domain fusion

CN117935330BActive Publication Date: 2026-10-09BEIJING INST OF TECH ZHUHAI CAMPUS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311787209.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2026-10-09
Estimated Expiration
2043-12-22

AI Technical Summary

Technical Problem

因此,使用多源域适应会比单源域适应有更好的效果,而目前还缺少多源域适应的行人重识别方法

Benefits of technology

[0053] The computer-executable instructions are executed by the processor to implement any of the steps of the pedestrian re-identification method based on graph convolutional network multi-source domain fusion described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117935330B_ABST
    Figure CN117935330B_ABST
Patent Text Reader

Abstract

The application provides a pedestrian re-identification method based on graph convolution network multi-source domain fusion, which comprises the following steps: selecting a data set and dividing it into a target domain data set and a source domain data set; constructing a feature extraction model based on ViT; performing multi-source domain fusion based on a graph convolution network; setting a classifier for each source domain; calculating the distance between all features in the target domain and all class centers, and assigning the pseudo label of the distance to the nearest class center to the class to which the nearest class center belongs; dividing the extracted features into two branches to calculate a loss function; using multiple loss functions to train a pedestrian re-identification network model; inputting image data into the trained model; and outputting a pedestrian re-identification result. The application can effectively realize the deep fusion of features in different domains and achieve the purpose of high-precision unsupervised domain adaptation pedestrian re-identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and artificial intelligence, specifically to a pedestrian re-identification method based on multi-source domain fusion of graph convolutional networks. Background Technology

[0002] Pedestrian re-identification is a cross-camera retrieval technology. Because it can identify and track pedestrians appearing at different locations and times across multiple cameras, it has enormous application potential in smart cities, smart surveillance, and autonomous driving. However, in real-world scenarios, the high cost of annotation often hinders the application of pedestrian re-identification. Therefore, unsupervised adaptive pedestrian re-identification, which does not require a new round of data annotation in the target scene, has greater research value and application potential. Unsupervised adaptive pedestrian re-identification assumes that multiple labeled datasets from different scenes have already been collected (called the source domain). When facing a new scene, i.e., the target domain, it does not need to annotate the target domain data. It uses the source domain data as a foundation, combined with unlabeled target domain data for learning, and performs retrieval in the target domain, resulting in better performance and applicability.

[0003] Current domain adaptation methods focus on using a single source domain for transfer learning. However, relying solely on a single source domain leads to performance dependence on the similarity between the selected source domain and the target domain. In other words, if the selected source domain is unsuitable, the network's performance will significantly degrade. This dependence on source domain selection greatly hinders the model's usability in real-world scenarios. In real-world scenarios, data often comes from multiple different scenes. Therefore, multi-source domain adaptation would be more effective than single-source domain adaptation, but currently, there is a lack of multi-source domain adaptation methods for person re-identification. Summary of the Invention

[0004] To address the challenge of domain distribution differences in multi-source domain adaptation, this invention aims to provide a pedestrian re-identification method based on multi-source domain fusion using graph convolutional networks. This method, based on the domain fusion structure of graph convolutional networks, effectively achieves deep fusion of features from different domains and realizes high-precision unsupervised domain-adaptive pedestrian re-identification.

[0005] The present invention achieves the above objectives through the following technical solutions:

[0006] A pedestrian re-identification method based on multi-source domain fusion using graph convolutional networks, comprising the following steps:

[0007] The datasets are selected and divided into target domain datasets and source domain datasets;

[0008] Construct a feature extraction model based on ViT;

[0009] Multi-source domain fusion based on graph convolutional networks includes: maintaining a domain center node for each domain and generating source domain center nodes; constructing a first-layer directed graph convolutional network; constructing a second-layer graph convolutional network; representing edges as high-dimensional features to obtain high-dimensional edge features; and updating the obtained nodes based on the obtained high-dimensional edge features through a message passing mechanism.

[0010] Classifiers are set up for different source domains. The pseudo-labels are assigned to the category to which the nearest category center belongs by calculating the distance between all features in the target domain and all category centers.

[0011] The extracted features are divided into two branches to calculate the loss function. The pedestrian re-identification network model is trained using multiple loss functions. The image data is input into the trained model, and the pedestrian re-identification result is output.

[0012] According to the present invention, a pedestrian re-identification method based on graph convolutional network multi-source domain fusion is provided, wherein maintaining a domain center node for each domain includes:

[0013] The weighted average of all nodes in a given domain within a batch is calculated using the following formula:

[0014]

[0015] Among them, agt d Q is the center node of the d-th domain, and Q represents the number of samples belonging to the d-th domain in a mini-batch. From the sample The weight values ​​learned in the process, and through Normalization is performed to ensure that the sum of the weights is 1, and f(·) is a fully connected layer;

[0016] The center node of the domain d is obtained by the weighted average.

[0017] According to the present invention, a pedestrian re-identification method based on graph convolutional network multi-source domain fusion uses a moving exponential average to maintain a robust domain center node, as shown in the following formula:

[0018]

[0019] Where t represents the t-th iteration, and α represents the size of the updated parameters;

[0020] By updating the sliding exponential average, a domain center node covering the entire domain d is obtained.

[0021] According to the present invention, a pedestrian re-identification method based on graph convolutional network multi-source domain fusion is provided, wherein generating source domain center nodes includes:

[0022] A weighted average of the domain center nodes of all source domains is taken to obtain a source domain center node that can cover all source domains. The specific formula is as follows:

[0023]

[0024] Where D represents the number of source domains, a d =f′(agt) d ) is a learned weight value, obtained through Normalization is performed to ensure that the sum of the weights is 1, and f′(·) is a fully connected layer;

[0025] By weighting all the source domain center nodes, we obtain the source domain center node agt representing the entire source domain. source .

[0026] According to the present invention, a pedestrian re-identification method based on multi-source domain fusion of graph convolutional networks is provided, wherein the construction of the first layer of directed graph convolutional network includes:

[0027] In the directed graph convolutional network, the center node of the source domain points to the center node of the target domain, and the center node of the target domain points to the center nodes of the two source domains respectively.

[0028] According to the present invention, a person re-identification method based on graph convolutional network multi-source domain fusion includes constructing a second-layer graph convolutional network, wherein:

[0029] First, a one-way connection is established between each node and the central node of the domain, and a two-way connection is established between the central nodes of the domain; wherein, a single node indirectly receives information from other domains through the connection with the central node, making the domains increasingly closer.

[0030] According to the present invention, a pedestrian re-identification method based on graph convolutional network multi-source domain fusion is provided, wherein obtaining high-dimensional edge features includes:

[0031] Representing an edge as a high-dimensional feature: if nodes are adjacent and the directed edge points to each other, then there is an edge feature, which is a vector of length . This feature is generated by learning through a multilayer perceptron, as shown in the following formula:

[0032] e i,j =MLP([v i ,v j ])

[0033] Among them, [v i ,v j ] represents the node features v of length D. i and vj Concatenate them into a new vector of length 2D;

[0034] The high-dimensional side feature e of length P is output through a two-layer MLP. i,j , used to represent v j Point to v i The depth-related information between the two.

[0035] According to the present invention, a pedestrian re-identification method based on graph convolutional network multi-source domain fusion is provided, wherein updating the obtained nodes through a message passing mechanism includes:

[0036] Based on the obtained high-dimensional edge features, the node g(v) after receiving information from neighboring nodes is obtained through the message passing mechanism g(·). i The updated node features are obtained through residual connections. The specific formula is as follows:

[0037]

[0038]

[0039] in, It is node v i The set of neighboring nodes, i.e., v k It is v i The neighboring nodes have v j Point to v i edge feature e i,k .,symbol Represents broadcast multiplication. It is a learnable weight used to fuse features from P channels.

[0040] According to the present invention, a pedestrian re-identification method based on graph convolutional network multi-source domain fusion is provided, wherein classifiers are set for different source domains respectively, including:

[0041] Set up separate classifiers for different source domains, with their output dimensions corresponding to the pedestrian identities in the source domain datasets;

[0042] A dedicated classifier is also set up for the target domain. Since the target domain is unlabeled, pseudo-labels need to be generated through clustering. Therefore, the output dimension is aligned with the number of pseudo-labels. Furthermore, the target domain is clustered during the transfer training phase to generate pseudo-labels to assist training. The pseudo-label generation method used is expressed by the following formula.

[0043]

[0044]

[0045]

[0046] Among them, X t It is the set of all samples in the target domain. F represents the feature extractor, i.e., the base model used, and C is the classifier. The target domain features are input into the classifier and the class scores belonging to different categories are obtained. Then, the features are grouped according to the highest class scores. This represents the set of all samples belonging to the k-th class; then, the mean of the set is calculated to obtain the center μ of each class. k Finally, the distances between all features and the different category centers are calculated, and the pseudo-labels are assigned according to these distances. Assign it to the category of the nearest category center.

[0047] According to the present invention, a pedestrian re-identification method based on graph convolutional network multi-source domain fusion is provided, wherein the extracted features are divided into two branches to calculate the loss function, including:

[0048] The extracted features are divided into two branches to calculate the loss function. One branch directly calculates the cross-entropy loss after passing through the BN Neck, while the other branch is input into the domain fusion mechanism and used to calculate the triplet loss after fusion.

[0049] Therefore, compared with the existing technology, the present invention uses graph convolutional networks to reduce differences in multiple source domains, establishes a multi-source domain fusion pedestrian re-identification method based on graph convolutional networks, and uses multiple source domains for domain adaptation research, which can expand the number of samples, enhance the training effect, improve the generalization ability, and achieve higher accuracy pedestrian re-identification.

[0050] The present invention also provides an electronic device, comprising:

[0051] Memory, which stores computer-executable instructions;

[0052] The processor is configured to run computer-executable instructions.

[0053] The computer-executable instructions are executed by the processor to implement any of the steps of the pedestrian re-identification method based on graph convolutional network multi-source domain fusion described above.

[0054] The present invention also provides a storage medium storing a computer program, which, when executed by a processor, is used to implement the steps of any of the above-described pedestrian re-identification methods based on graph convolutional network multi-source domain fusion.

[0055] Therefore, the present invention also provides an electronic device and a storage medium for a pedestrian re-identification method based on graph convolutional network multi-source domain fusion, comprising: one or more memories and one or more processors. The memories are used to store program code and intermediate data generated during program execution, storage of model output results, and storage of the model and model parameters; the processors are used for the processor resources occupied by the code execution and the multiple processor resources occupied during model training.

[0056] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. Attached Figure Description

[0057] Figure 1 This is a flowchart of an embodiment of a pedestrian re-identification method based on graph convolutional network multi-source domain fusion according to the present invention.

[0058] Figure 2 This is a schematic diagram of the graph convolutional structure in an embodiment of the pedestrian re-identification method based on multi-source domain fusion of graph convolutional networks according to the present invention.

[0059] Figure 3 This is a schematic diagram of an embodiment of a pedestrian re-identification method based on graph convolutional network multi-source domain fusion according to the present invention. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0061] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0062] See Figures 1 to 3 This invention provides a pedestrian re-identification method based on multi-source domain fusion of graph convolutional networks, the method comprising the following steps:

[0063] Step S1: Select the dataset, which is divided into target domain dataset and source domain dataset; in this embodiment, the Market-1501 public dataset is used as the target domain; the DukeMTMC, CUHK03 and MSMT17 public datasets are used as the source domains.

[0064] Step S2: Construct a feature extraction model based on ViT (Vision Transformer).

[0065] Step S3, performing multi-source domain fusion based on graph convolutional networks, includes: maintaining a domain center node for each domain and generating source domain center nodes; constructing a first-layer directed graph convolutional network; constructing a second-layer graph convolutional network; representing edges as high-dimensional features to obtain high-dimensional edge features; and updating the obtained nodes based on the obtained high-dimensional edge features through a message passing mechanism.

[0066] Step S4: Classifiers are set up for different source domains. The distance between all features in the target domain and all class centers is calculated, and the pseudo-label is assigned to the class to which the nearest class center belongs according to the distance.

[0067] Step S5: Divide the extracted features into two branches and calculate the loss function. Train the pedestrian re-identification network model using multiple loss functions, input image data into the trained model, and output the pedestrian re-identification result.

[0068] like Figure 2 As shown, a domain center node is maintained for each domain to serve as the representative of that domain, which specifically includes:

[0069] The weighted average of all nodes in a given domain within a batch is calculated using the following formula:

[0070]

[0071] Among them, agt d Q is the center node of the d-th domain, and Q represents the number of samples belonging to the d-th domain in a mini-batch. From the sample The weight values ​​learned in the process, and through Normalization is performed to ensure that the sum of the weights is 1, and f(·) is a fully connected layer;

[0072] The center node of the domain d is obtained by the above weighted average.

[0073] In this embodiment, a robust domain center node is maintained using a moving exponential average, as shown in the following formula:

[0074]

[0075] Where t represents the t-th iteration, and α represents the size of the updated parameter, the sliding exponential average update can obtain a domain center node that covers the entire domain d.

[0076] A weighted average of the domain center nodes of all source domains is taken to obtain a source domain center node that can cover all source domains. The specific formula is as follows:

[0077]

[0078] Where D represents the number of source domains, a d =f′(agt) d ) is a learned weight value, which will also be... Normalization is performed to ensure that the sum of the weights is 1, and f′(·) is a fully connected layer; by weighting all the source domain center nodes, the source domain center node agt representing the entire source domain is obtained. source .

[0079] In this embodiment, the first layer of the directed graph convolutional network is constructed, and the specific connection method is as follows: Figure 2 GCN (1) As shown, it includes:

[0080] In directed graph convolutional networks, the central nodes of the source domain point to the central nodes of the target domain, while the central nodes of the target domain point to the central nodes of the two source domains. This connection method allows the central nodes of the target domain to receive information from the source domain as a whole, and also allows the two different central nodes of the source domains to receive information from the target domain, thus achieving the goal of moving the source and target domains towards each other. This approach prioritizes the reality that the distribution differences between the target and source domains are relatively large, while the distribution differences within the source domains are relatively small. Therefore, the source domain is first treated as a whole, ignoring the merging of elements within the source domain, and prioritizing the merging of the target and source domains—that is, target domain priority. The first layer of the graph convolutional network can shorten the distance between the source and target domains.

[0081] In this embodiment, the second layer of the graph convolutional network is constructed, such as... Figure 2 GCN in (2) As shown, it specifically includes:

[0082] First, a one-way connection is established between each node and the central node of that domain, and a two-way connection is established between the domain central nodes. Through this connection method, the domain central nodes can receive information from each other, achieving domain fusion. Individual nodes can also indirectly receive information from other domains through connections with the central nodes, ultimately bringing the domains closer together.

[0083] In this embodiment, the high-dimensional edge features are obtained, specifically including:

[0084] Representing an edge as a high-dimensional feature: if nodes are adjacent and the directed edge points to each other, then there is an edge feature, which is a vector of length . This feature is generated by learning through a multilayer perceptron, as shown in the following formula:

[0085] e i,j =MLP([v i ,v j ])

[0086] Among them, [v i ,v j ] represents the node features v of length D. i and v j The vectors are concatenated to form a new vector of length 2D; then, through a two-layer MLP, a high-dimensional side feature e of length P is output. i,j , used to represent v j Point to v i The depth-related information between the two.

[0087] In this embodiment, updating the obtained nodes through a message passing mechanism includes:

[0088] Based on the obtained high-dimensional edge features, with node v i For example, the following message passing and node update methods can be implemented, with the specific formulas as follows:

[0089]

[0090]

[0091] in, It is node v i The set of neighboring nodes, i.e., v k It is v i The neighboring nodes have v j Point to v i edge feature e i,k .,symbol Represents broadcast multiplication. This is a learnable weight used to fuse features from P channels. Through the message passing mechanism g(·), the node g(v) after receiving information from its neighboring nodes is obtained. i The updated node features are obtained through residual connections. Through high-dimensional edge features e i,k The message passing mechanism can aggregate information from P different channels, making node updates more diversified and enabling the discovery of more correlations. In the two-layer graph convolutional network constructed in this embodiment, the above-described update method is used to update nodes.

[0092] In this embodiment, classifiers are set for different source domains, including:

[0093] Since unsupervised multi-source domain adaptation requires training with multiple different datasets that do not share a class space (i.e., the total number of pedestrians differs depending on their identity), this embodiment sets up three classifiers for the three different source domains, with their output dimensions corresponding to the pedestrian identities in the source domain datasets. A dedicated classifier is also set up for the target domain; since the target domain is unlabeled and requires pseudo-labels to be generated through clustering, its output dimension is aligned with the number of pseudo-labels.

[0094] In addition, a new small structure called the BN Neck (Browser Normalization) layer is added between the base model and the fully connected layers. This ensures that the features are constrained to a hypersphere, thus avoiding the network optimizing the length of the features and instead directly optimizing their position on the hypersphere, resulting in smoother gradient flow. The specific model structure is as follows: Figure 3 As shown, features before the BN Neck are used to calculate the Triplet Loss, while features after the BNNeck and FC layers are used to calculate the cross-entropy loss.

[0095] For unsupervised multi-source domain adaptive person re-identification, since the target domain is unlabeled, clustering of the target domain is required during the transfer training phase to generate pseudo-labels to assist training. The pseudo-label generation method is expressed by the following formula:

[0096]

[0097]

[0098]

[0099] Among them, X t It is the set of all samples in the target domain. F represents the feature extractor, i.e., the base model used, and C is the classifier. The target domain features are input into the classifier and the class scores belonging to different categories are obtained. Then, the features are grouped according to the highest class scores. This represents the set of all samples belonging to the k-th class; then, the mean of the set is calculated to obtain the center μ of each class. k Finally, the distances between all features and the different category centers are calculated, and the pseudo-labels are assigned according to these distances. Assign it to the category of the nearest category center.

[0100] In this embodiment, the extracted features are divided into two branches to calculate the loss function, including:

[0101] The extracted features are divided into two branches to calculate the loss function. One branch directly calculates the Cross Entropy Loss after passing through a Batch Normalization (BN) neck. The other branch is input into the domain fusion mechanism and used to calculate the Triplet Loss after fusion. This branch-based loss calculation is because during domain fusion, each node's features receive information from other domains and nodes, inevitably being influenced by the category information of other nodes. This can severely interfere with the classifier, hindering the network's learning process. Therefore, the fused features only calculate the Triplet Loss to optimize the distribution of these features in the latent space.

[0102] Cross-entropy loss and triplet loss The cross-entropy loss uses Label Smoothing to smooth the labels and process noise in the data. The specific formula is shown below:

[0103]

[0104] Where, p i Let be the output value of the classifier, i.e., the predicted value of the sample as being of class i, and y be the true label of the sample. It can be seen that, unlike the traditional cross-entropy loss which directly assigns 0 to all values ​​except the true label, label smoothing assigns them relatively small values, effectively preventing overfitting; where ∈ takes the value 0.1 in this experiment.

[0105] This embodiment uses Batch Hard Triplet Loss, which selects the positive samples with the largest distance and the negative samples with the smallest distance from a mini-batch. This ensures that the model focuses more on samples that are difficult to distinguish. Compared to the standard Triplet Loss, Batch Hard Triplet Loss places greater emphasis on distinguishing different samples, thereby improving the model's robustness and accuracy. The specific formula is shown below:

[0106]

[0107] Where D(·) is the distance metric function, which is the conventional Euclidean distance in this embodiment, and m is short for margin, which is a unique hyperparameter related to Triplet Loss, used to control the distance between positive and negative samples. In the experiments of this embodiment, it is set to 0.1.

[0108] After applying labeled smoothed cross-entropy loss and Batch Hard Triplet Loss, the model's total loss function is as follows:

[0109]

[0110] Here, α is a hyperparameter used to balance the two loss functions, and it is set to 0.9 in the relevant experiments of this embodiment.

[0111] In summary, this embodiment uses graph convolutional networks for multi-source domain difference reduction, establishes a multi-source domain fusion pedestrian re-identification method based on graph convolutional networks, and uses multiple source domains for domain adaptation research, which can expand the number of samples, enhance training effect, improve generalization ability, and achieve higher accuracy pedestrian re-identification.

[0112] In one embodiment, an electronic device is provided, which may be a server. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the electronic device provides computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the electronic device stores data. The network interface of the electronic device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a pedestrian re-identification method based on graph convolutional network multi-source domain fusion.

[0113] Those skilled in the art will understand that the electronic device structure shown in this embodiment is only a partial structure related to the solution of this application and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than shown in this embodiment, or combine certain components, or have different component arrangements.

[0114] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0115] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0116] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0117] Therefore, this embodiment also provides an electronic device and storage medium for a pedestrian re-identification method based on graph convolutional network multi-source domain fusion, which includes: one or more memories and one or more processors. The memories are used to store program code and intermediate data generated during program execution, storage of model output results, and storage of the model and model parameters; the processors are used for the processor resources occupied by the code execution and the multiple processor resources occupied when training the model.

[0118] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0119] The above embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention shall fall within the scope of protection claimed by the present invention.

Claims

1. A pedestrian re-identification method based on multi-source domain fusion using graph convolutional networks, characterized in that, The method includes the following steps: The datasets are selected and divided into target domain datasets and source domain datasets; Construct a feature extraction model based on ViT; Multi-source domain fusion based on graph convolutional networks includes: maintaining a domain center node for each domain and generating source domain center nodes; constructing a first-layer directed graph convolutional network; constructing a second-layer graph convolutional network; representing edges as high-dimensional features to obtain high-dimensional edge features; and updating the obtained nodes based on the obtained high-dimensional edge features through a message passing mechanism. Classifiers are set up for different source domains. The pseudo-labels are assigned to the category to which the nearest category center belongs by calculating the distance between all features in the target domain and all category centers. The extracted features are divided into two branches to calculate the loss function. The pedestrian re-identification network model is trained using multiple loss functions. Image data is input into the trained model, and pedestrian re-identification results are output. The process of obtaining high-dimensional edge features includes: Representing an edge as a high-dimensional feature: if nodes are adjacent and the directed edge points to each other, then there is an edge feature, which is a vector of length . This feature is generated by learning through a multilayer perceptron, as shown in the following formula: in This represents the node features of length D. and spliced ​​together to form a length of A new vector; A two-layer MLP is used to output high-dimensional side features of length P. , used to represent point to The depth-related information between the two; The step of updating the obtained nodes through a message passing mechanism includes: Based on the obtained high-dimensional edge features, through a message passing mechanism The node obtains the information received from its neighboring nodes. The updated node features are obtained through residual connections. The specific formula is as follows: in, It is a node The set of neighboring nodes, i.e. yes The neighboring nodes have point to edge features ,symbol Represents broadcast multiplication. This is a learnable weight used for fusion. Features of each channel.

2. The method according to claim 1, characterized in that, Maintaining a domain center node for each domain includes: The weighted average of all nodes in a given domain within a batch is calculated using the following formula: in, It is the first The central node of each field This represents the first in a small batch. The number of samples in each field; From the sample The weight values ​​learned in the process, and through To perform normalization to ensure the sum of the weights is , It is a fully connected layer; The domain is obtained through the weighted average. The central node.

3. The method according to claim 2, characterized in that: A robust domain center node is maintained using a moving exponential average, as shown in the following formula: in, Representing the Round iteration, and Represents the size of the updated parameter; By updating the sliding exponent average, a coverage area is obtained. The global domain center node.

4. The method according to claim 3, characterized in that, The generated source domain center node includes: A weighted average of the domain center nodes of all source domains is taken to obtain a source domain center node that can cover all source domains. The specific formula is as follows: in, Represents the number of source domains. It is a learned weight value, obtained through... Normalization is performed to ensure that the sum of the weights is , It is a fully connected layer; By weighting all the source domain center nodes, we obtain the source domain center node representing the entire source domain. .

5. The method according to claim 1, characterized in that, The construction of the first layer of the directed graph convolutional network includes: In the directed graph convolutional network, the center node of the source domain points to the center node of the target domain, and the center node of the target domain points to the center nodes of the two source domains respectively.

6. The method according to claim 5, characterized in that, The construction of the second-layer graph convolutional network includes: First, a one-way connection is established between each node and the central node of the domain, and a two-way connection is established between the central nodes of the domain; wherein, a single node indirectly receives information from other domains through the connection with the central node, making the domains increasingly closer.

7. The method according to claim 1, characterized in that, The method involves setting classifiers for different source domains, including: Set up separate classifiers for different source domains, with their output dimensions corresponding to the pedestrian identities in the source domain datasets; A dedicated classifier is also set up for the target domain. Since the target domain is unlabeled, pseudo-labels need to be generated through clustering. Therefore, the output dimension is aligned with the number of pseudo-labels. Furthermore, the target domain is clustered during the transfer training phase to generate pseudo-labels to assist training. The pseudo-label generation method used is expressed by the following formula. in, It is the set of all samples in the target domain. This represents the feature extractor, which is the base model used. It's a classifier. You input the target domain features into the classifier and get class scores for different categories. Then, you group the features according to the highest class score. That means it belongs to the first The set of all samples for each class; then the mean of the set is calculated to obtain the center of each class. Finally, the distances between all features and the different category centers are calculated, and the pseudo-labels are assigned according to these distances. Assign it to the category of the nearest category center.

8. The method according to claim 1, characterized in that, The step of dividing the extracted features into two branches to calculate the loss function includes: The extracted features are divided into two branches to calculate the loss function. One branch directly calculates the cross-entropy loss after passing through the BN Neck, while the other branch is input into the domain fusion mechanism and used to calculate the triplet loss after fusion.