Domain Generalization Person Re-identification Method for Adaptive Modeling of Domain Features

Through the adaptive modeling of domain features, a pedestrian re-identification network with cross-domain embedded blocks and feature fusion modules is built, which solves the cross-domain recognition problem, realizes efficient pedestrian re-identification in different scenarios, and improves the model's adaptability and accuracy.

CN115661923BActive Publication Date: 2025-07-04ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211281152.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-19
Publication Date
2025-07-04
Estimated Expiration
2042-10-19

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively solve the cross-domain identification problem in pedestrian re-identification. Traditional methods lead to insufficient feature modeling capabilities or high computational storage costs, and cannot adapt to deployment requirements in different scenarios.

Method used

Adaptive modeling domain features is adopted, and a pedestrian recognition network including backbone deep neural network, cross-domain embedding block, sub-feature embedding network, domain sample adaptive sub-feature combination module and static domain universal feature extraction module is used to perform cross-domain feature expression and dynamic and static features fusion, and meta-learning technology is used to improve model adaptability.

Benefits of technology

It realizes robustness and generalization in different domain scenarios, can train models in a fixed set of data and deploy them directly, without additional data adjustments, and improves the accuracy and efficiency of pedestrian re-identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661923B_ABST
    Figure CN115661923B_ABST
Patent Text Reader

Abstract

The present invention discloses a domain generalization pedestrian re-identification method for adaptively modeling domain features, which is used to eliminate the conflicts of different domains and perform good feature expressions for pedestrian modeling in the case of training with multiple different domain data for retrieval and matching tasks. The method specifically includes the following steps: establishing a backbone deep neural network for extracting color images; establishing a sub-feature embedding network; establishing a sub-feature combination module for domain sample adaptation; establishing a static domain general feature extraction module; establishing a fusion module for domain sample adaptive features and domain general features; training a prediction model based on the foregoing model structure to obtain a final trained neural network model. The present invention is applicable to pedestrian re-identification under multi-domain data training and testing, and has better effects and robustness in the face of various complex situations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a domain generalization pedestrian re-identification method for adaptively modeling domain features. Background Art

[0002] Pedestrian re-identification is a problem that frequently faces cross-domain identification. To solve the cross-domain problem, traditional machine learning methods will jointly train data from multiple domains. However, there are many unique features within different domains. Such features are very helpful for modeling samples within their own domains, but are useless or even harmful for modeling samples from other domains. For such domain conflict problems, traditional methods will choose to directly discard or separately maintain the features for each domain, which will lead to poor feature modeling capabilities or extremely high computational storage cost overheads, and cannot meet the requirements of pedestrian re-identification models deployed in different scenarios. Summary of the Invention

[0003] In view of the above problems, the present invention provides a domain generalization pedestrian re-identification method for adaptively modeling domain features.

[0004] The specific technical solution adopted by the present invention is as follows:

[0005] A domain generalization pedestrian re-identification method for adaptively modeling domain features, comprising the following steps:

[0006] S1. Obtain a multi-source domain dataset for training a pedestrian re-identification network;

[0007] S2. Construct a pedestrian re-identification network, where the pedestrian re-identification network includes a backbone deep neural network, a plurality of cross-domain embedding blocks embedded in the backbone deep neural network, and a classification network connected to the backbone deep neural network;

[0008] The backbone deep neural network adopts a ResNet-50-IBN model, which is used to extract basic features from the input domain samples for subsequent cross-domain feature representation; in each stage of the ResNet-50-IBN model except the first stage, the last feature expression bottleneck block is replaced by a cross-domain embedding block;

[0009] The cross - domain embedding block includes a sub - feature embedding network, a domain - sample - adaptive sub - feature combination module, a static domain - general feature extraction module, and a static - dynamic feature fusion module. The block input of the cross - domain embedding block is the output of the feature expression bottleneck block cascaded at the front end of the current cross - domain embedding block. There are multiple sub - feature embedding networks, which are respectively used to extract domain - independent sub - features from the block input, and construct a common sub - feature embedding space for all the domain - independent sub - features extracted by the sub - feature embedding networks. In the domain - sample - adaptive sub - feature combination module, first, a domain - aware adapter is used to generate the weights of each domain - independent sub - feature in the common sub - feature embedding space adaptively according to the current domain sample, and then all the domain - independent sub - features in the common sub - feature embedding space are weighted and aggregated to obtain dynamic domain - sample - adaptive features. The static domain - general feature extraction module is used to extract static domain - general features from the block input. The static - dynamic feature fusion module is used to fuse the static domain - general features and the dynamic domain - sample - adaptive features and perform fine - tuning. The finally obtained features are used as the block output of the cross - domain embedding block.

[0010] The classification network is used to perform classification according to the final output of the backbone deep neural network.

[0011] S3. Train the constructed person re - identification network based on the multi - source domain dataset, and use the finally trained person re - identification network to perform person re - identification on unknown target - domain image data.

[0012] Preferably, in S1, the multi - source domain dataset contains domain - sample data from different source domains, and each domain sample is an RGB color image I train .

[0013] Preferably, in S2, the ResNet - 50 - IBN model pre - trained on ImageNet is used as the backbone deep neural network, and the last feature expression bottleneck block in each stage except the first stage in the ResNet - 50 - IBN model is removed. For each frame of color image I train , its features output at each stage with a cross - domain embedding block inserted in the network are obtained by inputting it into the backbone deep neural network. The features output at the l - th stage are input into the cross - domain embedding block connected at the end of the l - th stage, and after being processed by the cross - domain embedding block, they are input into the next cascaded module.

[0014] Preferably, in S2, the processing steps within the cross - domain embedding block include:

[0015] S21. Based on the features output by the front - end module of the current l - th cross - domain embedding block N domain - independent sub - features are obtained respectively through N sub - feature embedding networks, forming the common sub - feature embedding space of the current domain samples; among them, the domain - independent sub - feature extracted by the nth sub - feature embedding network is

[0016]

[0017] In the formula: represents the nth sub - feature embedding network in the lth cross - domain embedding block, which is implemented by a feature expression bottleneck block;

[0018] S22. The sub - feature combination module adaptive to domain samples predicts the combination weights w of all domain - independent sub - features in the common sub - feature embedding space in S21 l :

[0019]

[0020] In the formula: w l is an N - dimensional vector, and MLP() represents the multi - layer perceptron used to predict weights;

[0021] Based on the combination weights w l weighted aggregation is performed on all domain - independent sub - features in the common sub - feature embedding space in S21 to obtain a sample - adaptive dynamic combined feature expression:

[0022]

[0023] In the formula: represents the dynamic domain sample - adaptive feature in the lth cross - domain embedding block, w l,n represents the nth weight value in the combination weights w l where n = 1, 2, …, N;

[0024] S23. According to the static domain - general feature extraction module, static domain - general features are additionally extracted from the feature :

[0025]

[0026] In the formula: MLP S () represents the feature expression bottleneck block in the static domain - general feature extraction module;

[0027] S24. The dynamic domain sample - adaptive feature and the static domain - general feature are feature - fused by the dynamic - static feature fusion module. The dynamic - static feature fusion module includes a numerical normalization module, a feature weighted fusion module, and a feature recombination and fine - tuning module, and finally the feature

[0028]

[0029]

[0030] In the formula: Norm S ( ) and Norm D ( ) are respectively the normalization functions in the numerical normalization module; a l is the fusion weight in the feature weighted fusion module and is a learnable parameter; Sigmoid( ) and MLP( ) respectively represent the activation function and the multi-layer perceptron in the feature recombination fine-tuning module.

[0031] Preferably, in S3, the specific method for training the constructed person re-identification network is as follows:

[0032] S31. For each training sample in the multi-source domain dataset, input the feature train of the color image I predicted in S24 into the pooling layer, and input the pooling result into the classification network to obtain the predicted classification result and calculate the cross-entropy loss function L using ce and the in-domain category result Y manually labeled, the triplet loss function L triplet , the center clustering loss function L center and the domain-aware clustering loss function L cluster for four indicators;

[0033] S33. For each training sample, calculate the final loss function as:

[0034] L total = L ce + L triplet + L center + L cluster

[0035] and use the SGD optimization method and the backpropagation algorithm to train the entire person re-identification network under the loss function L total until the network converges.

[0036] Preferably, the domain-aware clustering loss function L cluster is calculated based on the combined weights predicted in S22. First, obtain the global combined weight W = [w 1 , w 2 , …, w L , calculate the in-domain clustering loss L intra and the inter-domain clustering loss L inter based on the combined weights, and then sum the two to obtain L cluster :

[0037]

[0038]

[0039] L cluster = L intra + L inter

[0040] Where: L is the total number of stages where cross - domain embedding blocks are inserted in the ResNet - 50 - IBN model, K represents the number of domains, M represents the number of instances within a domain, m1 and m2 are distance - constraint hyperparameters, W i,j represents the global combination weight of the j - th sample in the i - th domain, and respectively represent the average global combination weights of the i - th and j - th domains.

[0041] Preferably, there are five stages in the backbone deep neural network, and cross - domain embedding blocks are inserted at the ends of the second to fifth stages respectively.

[0042] Preferably, within the cross - domain embedding block, the number N of sub - feature embedding networks is 4.

[0043] Preferably, in the domain - aware clustering loss function, the distance - constraint hyperparameter m1 is 0.1 and m2 is 0.3.

[0044] Preferably, in S3, for unknown target - domain image data, it is input into the person re - identification network after removing the classification network to obtain a feature representation in the pooling layer, and the feature representation is used for retrieval in the database to achieve person re - identification.

[0045] This method is based on a deep neural network, models the domain - specific features and domain - shared feature spaces in RGB images, explores a domain - adaptive sub - feature combination method, adaptively extracts feature representations sensitive to the sample domain, and uses meta - learning techniques to improve the model's expression ability when dealing with unknown domains, and can better meet the requirements of person re - identification models in different domain scenarios. Compared with the methods in the prior art, the present invention has the following beneficial effects:

[0046] First, the present invention proposes a method for decomposing general or specific sub - features of different domains, so as to be able to construct a common sub - feature space independent of different domains.

[0047] Second, the present invention constructs a domain - aware sub - feature combiner by using the method of generating meta - parameters. By designing this part, the adaptability of the model to samples in different domains can be greatly improved, domain - adaptive feature representations can be extracted, and a more generalizable person re - identification model can be realized.

[0048] In the pedestrian re-identification task, this method can effectively improve the robustness to scene changes and has good application value. For example, the model can be trained on a fixed set of data and directly deployed in different scenarios without the need to collect data in a specified area again for model adjustment. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a schematic diagram of the basic steps of the method of the present invention;

[0050] Figure 2 is a schematic diagram of the structure of the adaptive feature extraction pedestrian re-identification network of the present invention;

[0051] Figure 3 is a partial experimental effect diagram shown in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0052] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0053] On the contrary, the present invention covers any alternatives, modifications, equivalent methods and solutions made within the spirit and scope of the present invention as defined by the claims. Further, in order to enable the public to better understand the present invention, in the following detailed description of the present invention, some specific details are described in detail. Those skilled in the art can fully understand the present invention without the description of these details.

[0054] Reference Figure 1 , in a preferred embodiment of the present invention, a method for pedestrian re-identification using adaptive domain feature extraction is provided, which is used for the process of conflict-free modeling of color images in different domains given multiple domains of color images. The method specifically includes the following steps:

[0055] S1. Obtain a multi-source domain dataset for training a pedestrian re-identification network.

[0056] In the embodiment of the present invention, in the above S1, the multi-source domain dataset contains domain sample data from different source domains, and each domain sample is an RGB color image I train .

[0057] S2. Construct a pedestrian re-identification network, which includes a backbone deep neural network, several cross-domain embedding blocks (CODE-Block) embedded in the backbone deep neural network, and a classification network connected to the backbone deep neural network.

[0058] The overall network structure of the pedestrian re-identification network is as Figure 2As shown in the figure, the backbone deep neural network adopts the ResNet-50-IBN model, which is used to extract basic features from the input domain samples for subsequent cross-domain feature representation; in each stage of the ResNet-50-IBN model except the first stage, the last representational bottleneck block is replaced by a cross-domain embedding block;

[0059] The cross-domain embedding block includes a sub-feature embedding network, a domain-sample adaptive sub-feature combination module, a static domain-general feature extraction module, and a static-dynamic feature fusion module. The block input of the cross-domain embedding block is the output of the representational bottleneck block cascaded at the front end of the current cross-domain embedding block; there are multiple sub-feature embedding networks, which are respectively used to extract domain-agnostic sub-features from the block input, and the domain-agnostic sub-features extracted by all sub-feature embedding networks are constructed into a common sub-feature embedding space; in the domain-sample adaptive sub-feature combination module, first, the domain-aware adapter is used to adaptively generate the weights of each domain-agnostic sub-feature in the common sub-feature embedding space according to the current domain sample, and then all domain-agnostic sub-features in the common sub-feature embedding space are weighted and aggregated to obtain dynamic domain-sample adaptive features; the static domain-general feature extraction module is used to extract static domain-general features from the block input; the static-dynamic feature fusion module is used to fuse the static domain-general features and the dynamic domain-sample adaptive features and perform fine-tuning, and the finally obtained features are used as the block output of the cross-domain embedding block;

[0060] The classification network is used to classify according to the final output of the backbone deep neural network.

[0061] In the embodiment of the present invention, in the above S2, the ResNet-50-IBN model pre-trained on ImageNet is used as the backbone deep neural network. The last representational bottleneck block in each stage of the ResNet-50-IBN model except the first stage is removed to ensure the simplicity of the network, and the position of the removed representational bottleneck block is replaced by a cross-domain embedding block. The backbone deep neural network refers to the ResNet-50-IBN model and has a total of five stages. The first stage is convolutional and pooling operations, and there is no need to insert a representational bottleneck block. Therefore, in this embodiment, cross-domain embedding blocks (CODE1-4 above) are inserted at the end of the second to fifth stages respectively. Figure 2 in the above

[0062] For each frame of color image I train , input it into the backbone deep neural network to obtain the features output by each stage in the network where cross-domain embedding blocks are inserted. The features output by the l-th stage It is input into the cross - domain embedding block connected at the end of the l - th stage, and after being processed by the cross - domain embedding block, it is input into the next cascaded module.

[0063] In the embodiment of the present invention, in the above S2, the processing steps in the cross - domain embedding block include:

[0064] S21. Based on the features output by the front - end module of the current l - th cross - domain embedding block N domain - independent sub - features are respectively obtained through N sub - feature embedding networks, constituting the common sub - feature embedding space of the current domain sample; among them, the domain - independent sub - feature extracted by the n - th sub - feature embedding network is

[0065]

[0066] In the formula: represents the n - th sub - feature embedding network in the l - th cross - domain embedding block, which is implemented by a feature expression bottleneck block.

[0067] In the embodiment of the present invention, within the above - mentioned cross - domain embedding block, the number N of sub - feature embedding networks is taken as 4.

[0068] S22. According to the sub - feature combination module adaptive to the domain sample, predict the combination weight w of all domain - independent sub - features in the common sub - feature embedding space in S21 l :

[0069]

[0070] In the formula: w l is an N - dimensional vector, and MLP() represents a multi - layer perceptron for predicting weights;

[0071] Based on the combination weight w l weighted aggregation is performed on all domain - independent sub - features in the common sub - feature embedding space in S21 to obtain a sample - adaptive dynamic combined feature expression:

[0072]

[0073] In the formula: represents the dynamic domain - sample - adaptive feature in the l - th cross - domain embedding block, w l,n represents the combination weight w l the n - th weight value in, n = 1, 2,..., N.

[0074] S23. According to the static domain - general feature extraction module, additionally extract static domain - general features from the feature :

[0075]

[0076] Wherein: MLP S () represents the feature expression bottleneck block in the static domain general feature extraction module.

[0077] S24. Use the dynamic and static feature fusion module to adaptively extract features from the dynamic domain samples and the static domain general features for feature fusion. The dynamic and static feature fusion module includes a numerical normalization module, a feature weighted fusion module, and a feature recombination fine-tuning module, and finally obtains the features for matching

[0078]

[0079]

[0080] Wherein: Norm s () and Norm D () are respectively the normalization functions in the numerical normalization module; α l is the fusion weight in the feature weighted fusion module, which is a learnable parameter; Sigmoid() and MLP() respectively represent the activation function and the multi-layer perceptron in the feature recombination fine-tuning module.

[0081] In addition, it should be noted that in the framework of the person re-identification network, the classifier data in the classification network needs to be determined according to the actual task situation, and there can be one or more. Figure 2 Three classifiers are shown in. The classification network is mainly used to assist the backbone network in training, and the classification network needs to be removed in the test stage or the actual application stage. Therefore, for the unknown target domain image data, input it into the person re-identification network after removing the classification network to obtain the feature expression in the pooling layer, and use this feature expression to retrieve in the library to achieve person re-identification. In actual retrieval, the feature expression of the target domain image data can be calculated for similarity with the feature expressions of the existing pedestrians in the library, and matching recommendations can be made according to the similarity.

[0082] S3. Train the constructed person re-identification network based on the multi-source domain dataset, and use the finally trained person re-identification network to perform person re-identification on the unknown target domain image data.

[0083] In the embodiment of the present invention, in the above S3, the specific method for training the constructed person re-identification network is as follows:

[0084] S31. For each training sample in the multi-source domain dataset, input the features train of the color image I predicted in S24 into the pooling layer, and then input the pooling result into the classification network to obtain the predicted classification result And use to calculate the cross-entropy loss function L with the manually labeled in-domain class results Y ce , triple loss function L triplet , center clustering loss function L center and domain-aware clustering loss function L cluster , four metrics;

[0085] In the embodiments of the present invention, the above domain-aware clustering loss function L cluster is calculated based on the combined weights predicted in S22. First, obtain the global combined weights W = [w 1 , w 2 , …, w L , calculate the in-domain clustering loss L intra and the inter-domain clustering loss L inter based on the combined weights, and then sum the two to obtain L cluster :

[0086]

[0087]

[0088] L cluster = L intra + L inter

[0089] where: L is the total number of stages where the cross-domain embedding block is inserted in the ResNet-50-IBN model, K represents the number of domains, M represents the number of instances within the domain, m1 and m2 are distance constraint hyperparameters, W i,j represents the global combined weight of the jth sample in the ith domain, and represent the average global combined weights of the ith and jth domains, respectively.

[0090] In the embodiments of the present invention, in the above domain-aware clustering loss function, the distance constraint hyperparameter m1 takes 0.1 and m2 takes 0.3.

[0091] In the present invention, the metrics for measuring classification accuracy and feature metric differences are the cross-entropy loss and the triple loss, respectively. The specific calculation methods of both belong to the prior art and will not be elaborated here.

[0092] S33. For each training sample, calculate the final loss function as:

[0093] L totar = L ce + L triplet + L center + L cluster

[0094] And use SGD optimization method and back propagation algorithm to optimize the loss function L total The entire person re-identification network is trained until the network converges.

[0095] The person re-identification network that converges after the above training can be used to perform person re-identification feature expression on actual RGB color images of different domains. When applied, it is only necessary to input the RGB color image to be characterized into the person re-identification network, and the corresponding feature expression can be output. By performing similar matching with the feature expressions of known pedestrians in the library, person re-identification can be achieved. The methods described in S1 to S3 above are applied to a specific embodiment below so that those skilled in the art can better understand the effect of the present invention.

[0096] Example

[0097] The implementation method of this embodiment is as described in S1 to S3 above, and the specific steps will not be elaborated in detail. The following only shows its effect based on case data. The present invention is implemented on four data sets with true value annotations, namely:

[0098] Market1501 dataset: This dataset contains 12936 images for training and 19281 images for testing.

[0099] CUHK03 dataset: This dataset contains 7368 images for training and 6728 images for testing.

[0100] MSMT17 dataset: This dataset contains 30,248 images for training and 96,193 images for testing.

[0101] CUHKSYSU dataset: This dataset contains 34,574 images for training.

[0102] In this example, three of the four data sets are selected for training based on the above method, and the others are used as test sets. The experiment is repeated three times and the average value is taken.

[0103] like Figure 3 As shown, the red triangle points in the figure are query points, and the green pentagon points are reference points in the library. The features obtained by the method of the present invention can make the query points more similar to the corresponding reference points.

[0104] The detection accuracy of the test results in this embodiment is shown in Table 1 below. The prediction accuracy of various methods is mainly compared using two indicators: average Rank1 and mAP. Among them, the average Rank1 indicator is used to measure the correct rate of the result with the highest prediction confidence, and the larger the value, the more accurate the prediction result; the mAP indicator is the average accuracy of all predictions, and the larger the value, the more comprehensive and accurate the prediction result. As shown in Table 1, compared with other methods, the method of the present invention (denoted as Our network) has obvious advantages in both the average Rank1 and mAP indicators.

[0105] Table 1

[0106]

[0107] In the above embodiment, the cross-domain adaptive person re-identification method of the present invention first models a domain-independent common feature space, and then combines domain-adaptive feature expressions according to a domain-aware weight generator, directly differentiates different domains for supervised training, and obtains a better person re-identification model.

[0108] Through the above technical solutions, the embodiment of the present invention develops a domain generalization person re-identification method for adaptively modeling domain features based on deep learning technology. The present invention can model domain-adaptive features for samples in different domains and can better adapt to the person re-identification task in complex scenarios of different domains.

[0109] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. An adaptive domain feature modeling-based domain generalization pedestrian re-identification method, characterized in that It includes the following steps: S1. Obtain a multi-source domain dataset for training a person re-identification network; S2. Construct a person re-identification network, which includes a backbone deep neural network, several cross-domain embedding blocks embedded in the backbone deep neural network, and a classification network connected after the backbone deep neural network; The backbone deep neural network adopts a ResNet-50-IBN model, which is used to extract basic features from the input domain samples for subsequent cross-domain feature expression; in each stage of the ResNet-50-IBN model except the first stage, the last feature expression bottleneck block is replaced by a cross-domain embedding block; The cross-domain embedding block includes a sub-feature embedding network, a domain-sample adaptive sub-feature combination module, a static domain-general feature extraction module, and a static and dynamic feature fusion module. The block input of the cross-domain embedding block is the output of the feature expression bottleneck block cascaded in front of the current cross-domain embedding block; there are multiple sub-feature embedding networks, which are respectively used to extract domain-agnostic sub-features from the block input, and the domain-agnostic sub-features extracted by all sub-feature embedding networks are constructed into a common sub-feature embedding space; in the domain-sample adaptive sub-feature combination module, first use a domain-aware adapter to adaptively generate the weights of each domain-agnostic sub-feature in the common sub-feature embedding space according to the current domain sample, and then perform weighted aggregation on all domain-agnostic sub-features in the common sub-feature embedding space to obtain a dynamic domain-sample adaptive feature; The static domain-general feature extraction module is used to extract static domain-general features from the block input; the static and dynamic feature fusion module is used to fuse the static domain-general features and the dynamic domain-sample adaptive features and perform fine-tuning, and the finally obtained features are used as the block output of the cross-domain embedding block; The classification network is used to classify according to the final output of the backbone deep neural network; S3. Based on the multi-source domain dataset, train the constructed person re-identification network, and after removing the classification network from the finally trained person re-identification network, use it to perform person re-identification on unknown target domain image data.

2. The domain generalization pedestrian re-identification method for adaptive domain feature modeling according to claim 1, characterized in that In the above S1, the multi-source domain dataset contains domain sample data from different source domains, and each domain sample is an RGB color image I train .

3. The domain generalization pedestrian re-identification method for adaptive domain feature modeling according to claim 1, wherein In S2, the pre-trained ResNet-50-IBN model on ImageNet is used as the backbone deep neural network, and the last feature expression bottleneck block in each stage except the first stage in the ResNet-50-IBN model is removed; for each frame of color image I train , it is input into the backbone deep neural network to obtain the features output by each stage in the network with cross-domain embedding blocks inserted. The features output by the l-th stage are input into the cross-domain embedding block connected at the end of the l-th stage. After being processed by the cross-domain embedding block, they are then input into the next cascaded module.

4. The domain generalization pedestrian re-identification method for adaptive domain feature modeling according to claim 3, wherein In S2, the processing steps in the cross-domain embedding block include: S21. Based on the features output by the front-end module of the current l-th cross-domain embedding block N domain-agnostic sub-features are respectively obtained through N sub-feature embedding networks to form the common sub-feature embedding space of the current domain samples; among them, the domain-agnostic sub-feature extracted by the n-th sub-feature embedding network is In the formula: represents the nth sub-feature embedding network in the lth cross-domain embedding block, which is implemented by a feature expression bottleneck block; S22. Predict the combination weights w of all domain-agnostic sub-features in the common sub-feature embedding space in S21 according to the domain-sample adaptive sub-feature combination module l : where: w l is an N-dimensional vector, and MLP() represents a multi-layer perceptron for predicting weights; Based on the combined weight w l Perform weighted aggregation on all domain-agnostic sub-features in the common sub-feature embedding space described in S21 to obtain a sample-adaptive dynamic combined feature representation: In the formula: represents the dynamic domain sample adaptive feature in the l-th cross-domain embedding block, and w l,n represents the combined weight w l The n-th weight value in, where n = 1, 2,..., N; S23. According to the static domain general feature extraction module, additionally extract the static domain general features from the features : Where: MLP S () represents the feature expression bottleneck block in the static domain general feature extraction module; S24. Use the dynamic and static feature fusion module to adaptively extract features from dynamic domain samples and general features of the static domain for feature fusion. The dynamic and static feature fusion module includes a numerical normalization module, a feature weighted fusion module, and a feature recombination and fine-tuning module, and finally obtains the features for matching Where: Norm S () and Norm D () are the normalization functions in the numerical normalization module; α l is the fusion weight in the feature weighted fusion module and is a learnable parameter; Sigmoid() and MLP() respectively represent the activation function and multi-layer perceptron in the feature recombination fine-tuning module.

5. The method for constructing a pedestrian re-identification network using the multi-source domain dataset according to claim 4, wherein In S3, the specific method for training the constructed person re-identification network is as follows: S31. For each training sample in the multi-source domain dataset, input the features of the color image I predicted in S24 into the pooling layer, and then input the pooling result into the classification network to obtain the predicted classification result train and use the pooling result to calculate the cross-entropy loss function L using and the intra-domain category result Y manually labeled and calculate the cross-entropy loss function L ce , the triplet loss function L triplet , the center clustering loss function L center and the domain-aware clustering loss function L cluster for four metrics; S33. For each training sample, calculate the final loss function as: L total = L ce + L triplet + L center + L cluster and use the SGD optimization method and the backpropagation algorithm to train the entire person re-identification network under the loss function L total until the network converges.

6. The method for constructing a pedestrian re-identification network using the multi-source domain dataset according to claim 5, wherein The domain-aware clustering loss function L cluster is calculated based on the combined weights predicted in S22. First, obtain the global combined weight W = [w 1 , w 2 , …, w L . Calculate the intra-domain clustering loss L intra and the inter-domain clustering loss L inter based on the combined weights, and then sum the two to obtain L cluster : L cluster = L intra + L inter Where: L is the total number of stages in the ResNet-50-IBN model where the cross-domain embedding block is inserted, K represents the number of domains, M represents the number of instances within a domain, m1 and m2 are distance constraint hyperparameters, W i,j represents the global combined weight of the j-th sample in the i-th domain, and represent the average global combined weights of the i-th and j-th domains, respectively.

7. The method for constructing a pedestrian re-identification network using the multi-source domain dataset according to claim 1, characterized in that, There are five stages in the backbone deep neural network, and cross-domain embedding blocks are inserted at the ends of the second to fifth stages respectively.

8. The method for constructing a person re-identification network using the multi-source domain dataset according to claim 1, wherein Inside the cross-domain embedding block, the number N of sub-feature embedding networks is taken as 4.

9. The method for constructing a pedestrian re-identification network using the multi-source domain dataset according to claim 6, wherein In the domain-aware clustering loss function, the distance constraint hyperparameter m1 is taken as 0.1 and m2 is taken as 0.

3.

10. The method for constructing a pedestrian re-identification network using the multi-source domain dataset according to claim 1, characterized in that, In S3, for unknown target domain image data, input it into the person re-identification network after removing the classification network to obtain a feature expression in the pooling layer, and use this feature expression to perform retrieval in the library to achieve person re-identification.

Citation Information

Patent Citations

  • Cross-modal retrieval method for sketch retrieval three-dimensional model based on spatiotemporal feature information

    CN112085072A

  • Modeling and learning character traits and medical condition based on 3D facial features

    US20180190377A1