Cross-modal Person Re-identification Method Based on Multi-agent Similarity Aggregation
Through the multi-agent similarity aggregation method, multiple agents are established for each category and corresponding aggregation mechanism is designed, which solves the accuracy of cross-modal pedestrian re-identification under day and night lighting differences, improves the accuracy of cross-modal pedestrian re-identification, and is suitable for smart video surveillance systems in smart cities, smart transportation and smart security.
Patent Information
- Application Number
- CN202211386276.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-11-07
AI Technical Summary
The existing cross-modal pedestrian re-identification method is insufficient in accuracy under day and night illumination differences, and the single-agent Softmax loss function cannot effectively handle intra-class differences in cross-modal images, resulting in limited practical capabilities of the video surveillance system.
Multi-agent similarity aggregation method is adopted to establish multiple learnable agents for each category, and a corresponding aggregation mechanism is designed to optimize similarity aggregation through mean, attention mechanism or difficulty mining strategies to improve the accuracy of cross-modal pedestrian re-identification.
Through multi-agent similarity aggregation, it can better characterize the in-class differences of cross-modal images, improve the accuracy of cross-modal pedestrian re-identification, adapt to modal changes, and apply to smart video surveillance systems in smart cities, smart transportation and smart security.
Smart Images

Figure CN115620343B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of machine vision and intelligent video monitoring, and particularly relates to a cross-modal pedestrian re-identification method based on multi-agent similarity aggregation. Background Art
[0002] Due to the large difference in illumination between day and night, video surveillance cameras work in the visible light mode during the day and in the infrared mode at night. It is difficult to distinguish targets in different imaging modes, making it extremely difficult to re-identify day-time target images and night-time target images. Existing research pays insufficient attention to cross-modal pedestrian re-identification, and most adopt the single-agent Softmax loss function, which is the same as that for single-modal pedestrian re-identification. The single-agent Softmax loss function does not pay enough attention to the intra-class differences of cross-modal images, restricting the accuracy of cross-modal pedestrian re-identification, and subsequently limiting the practical ability of cross-modal pedestrian re-identification algorithms in video surveillance systems. Summary of the Invention
[0003] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a cross-modal pedestrian re-identification method based on multi-agent similarity aggregation, which establishes multiple learnable agents for each category and uses an aggregation mechanism to re-aggregate the multi-agent similarities, so as to better describe the intra-class variations and similarity learning of cross-modal images, thereby enhancing the adaptability of the cross-modal pedestrian re-identification model to modal changes and further improving the accuracy of cross-modal pedestrian re-identification.
[0004] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0005] A cross-modal pedestrian re-identification method based on multi-agent similarity aggregation constructs a loss function by using multi-agent similarity aggregation, and the steps are as follows:
[0006] Step 1.1: Construct a loss function by using a multi-agent similarity aggregation mechanism, as shown in formula (1),
[0007]
[0008] where M represents the number of training samples, C represents the number of categories, represents the multi-agent aggregation similarity between the i-th sample x i and the c i -th category, c i ∈[1, 2, 3,..., C] represents the category label of x i , ]>represents the multi-agent aggregation similarity between the sample x i and the j-th category.
[0009] Step 1.2: Use the aggregation mechanism to aggregate the similarities between the sample and multiple proxies of each category to obtain the sample x i The multi-agent aggregation similarity with each category, as shown in formula (2),
[0010]
[0011] where, represents the i-th sample x i The multi-agent aggregation similarity with the j-th category, Aggregate represents the aggregation function, represents the i-th sample x i The similarity between the i-th sample x and the k-th proxy of the j-th category, K represents the number of proxies in each category.
[0012] Step 1.3: The similarity between the i-th sample x i and the k-th proxy of the j-th category is calculated as shown in formula (3),
[0013]
[0014] where, represents multiple k-th learnable proxies of the j-th category, T is the transpose operation, represents the d-dimensional feature vector extracted from x i by the neural network Net, θ is the weight parameter of the neural network Net.
[0015] Step 1.4: The similarity aggregation method is aggregated by one of the following three methods:
[0016] i) As Figure 1 shown, the similarity aggregation method uses mean aggregation, as shown in formula (4),
[0017]
[0018] ii) As Figure 2 shown, the similarity aggregation method uses attention mechanism aggregation, as shown in formula (5),
[0019]
[0020] where, represents the aggregation weight, calculated by the attention sub-network ANet, that is represents the d-dimensional feature vector extracted from x i by the neural network Net, φ is the weight parameter of ANet.
[0021] iii) As Figure 3As shown, the similarity aggregation method adopts a hard mining strategy for aggregation, as shown in Equation (6).
[0022]
[0023] Among them, c i ∈[1,2,3,...,C] represents the class label of sample x i . Among the proxies of the same class as x i (i.e., j = c i ), the minimum similarity is mined as the hard similarity between x i and the j-th class for optimization. Among the proxies of different classes from x i (i.e., j ≠ c i ), the maximum similarity is mined as the hard similarity between x i and the j-th class for optimization.
[0024] After adopting the above scheme, the present invention assigns multiple learnable proxies to each class, obtains multi-proxy similarities, and designs a multi-proxy similarity aggregation mechanism to achieve cross-modal pedestrian re-identification. On the one hand, the present invention learns multiple proxies for each class, which can better characterize the significant intra-class differences caused by data cross-modality; on the other hand, the present invention designs an aggregation mechanism to learn the optimal multi-proxy similarity aggregation method, improving the accuracy of cross-modal pedestrian re-identification. Therefore, the present invention can be widely applied to intelligent video surveillance systems in smart cities, smart transportation, and smart security. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 FIG. is a schematic diagram of the cross-modal pedestrian re-identification method based on multi-proxy similarity aggregation of the present invention, where the multi-proxy similarity aggregation method adopts mean aggregation.
[0026] Figure 2 FIG. is a schematic diagram of the cross-modal pedestrian re-identification method based on multi-proxy similarity aggregation of the present invention, where the multi-proxy similarity aggregation method adopts attention mechanism aggregation.
[0027] Figure 3 FIG. is a schematic diagram of the cross-modal pedestrian re-identification method based on multi-proxy similarity aggregation of the present invention, where the multi-proxy similarity aggregation method adopts hard mining strategy aggregation. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In this embodiment, for the cross-modal pedestrian re-identification method based on multi-proxy similarity aggregation, its training stage is as follows:
[0029] 1.1) Obtain visible light and infrared training set images of pedestrians, where each image is equipped with an identity identifier.
[0030] 1.2) Construct a dual-branch deep learning network Net. In this embodiment, as shown in Figure 1 , Figure 2 and Figure 3 , the dual-branch deep learning network consists of ResNet50, a Generalized-mean Pooling (GeP) layer, and a Fully Connected (FC) layer. In this embodiment, the specific dual-branch construction method is that the Stem, residual groups Layer1 and Layer2 of ResNet50 do not share parameters, while Layer3 and Layer4 share parameters.
[0031] 1.3) Train the constructed dual-branch deep learning network using a multi-agent similarity aggregation loss function:
[0032] 1.3.1) Adopt a multi-agent similarity aggregation mechanism to construct a loss function, as shown in formula (1).
[0033]
[0034] Among them, M represents the number of training samples, C represents the number of categories, represents the multi-agent aggregation similarity between the i-th sample x i and the c i -th category, c i ∈ [1, 2, 3,..., C] represents the category label of x i ; represents the multi-agent aggregation similarity between the sample x i and the j-th category.
[0035] 1.3.2) Use the aggregation mechanism to aggregate the similarities between the sample and multiple agents of each category, and obtain the multi-agent aggregation similarity between the sample x i and each category, as shown in formula (2).
[0036]
[0037] Among them, represents the multi-agent aggregation similarity between the i-th sample x i and the j-th category, Aggregate represents the aggregation function, represents the similarity between the i-th sample x i and the k-th agent of the j-th category, and K represents the number of agents for each category.
[0038] 1.3.3) The similarity i between the i-th sample x and the k-th agent of the j-th category is calculated as shown in formula (3).
[0039]
[0040] Among them, represents multiple k-th learnable agents of the j-th category, T is the transpose operation, represents the d-dimensional feature vector extracted from x i through the neural network Net, and θ is the weight parameter of the neural network Net.
[0041] 1.3.4) The similarity aggregation method can be aggregated in the following three ways:
[0042] i) As Figure 1 shown, the similarity aggregation method can adopt mean aggregation, as shown in formula (4),
[0043]
[0044] ii) As Figure 2 shown, the similarity aggregation method can adopt attention mechanism aggregation, as shown in formula (5),
[0045]
[0046] Among them, represents the aggregation weight, calculated by the attention sub-network ANet, that is Among them, represents the d-dimensional feature vector extracted from x i through the neural network Net, and φ is the weight parameter of ANet.
[0047] iii) As Figure 3 shown, the similarity aggregation method can adopt hard mining strategy aggregation, as shown in formula (6),
[0048]
[0049] Among them, c i ∈[1, 2, 3,..., C] represents the class label of the sample x i . Among the agents of the same class as x i (i.e., j = c i ), the minimum similarity is mined as the hard similarity between x i and the j-th category for optimization. Among the agents of different classes from x i (i.e., j ≠ c i ), the maximum similarity is mined as the hard similarity between x i and the j-th category for optimization.
[0050] 1.4) Minimize the loss function (formula (1)) using the gradient descent method to complete the training of the cross-modal pedestrian re-identification model and obtain the cross-modal pedestrian re-identification method model.
[0051] In this embodiment, for the cross-modal pedestrian re-identification method based on multi-agent similarity aggregation, its test phase is as follows: Use the cross-modal pedestrian re-identification method model obtained in the training phase to extract features from the query image and the registration image set, and obtain the features of the query image and the registration image (i.e., Figures 1 to 3 the output of the FC layer in). Based on the extracted features, calculate the distances between the query image and the registration images respectively, and sort them in ascending order according to the distances. Select the registration images with the top rankings as the registration images similar to the query image, which are used as the recognition results of the cross-modal pedestrian re-identification model.
[0052] The key of the present invention lies in fully realizing the feature representation of each class through the multi-agent method, which can better describe the intra-class variation of cross-modal images, and designing three multi-agent similarity aggregation methods, which can better perform similarity learning, thereby enhancing the adaptability of the cross-modal pedestrian re-identification model to modal changes, and further improving the accuracy of cross-modal pedestrian re-identification. Therefore, the present invention can be widely applied to intelligent video surveillance systems in smart cities, intelligent transportation, and intelligent security.
[0053] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.
Claims
1. A cross-modal pedestrian re-identification method based on multi-agent similarity aggregation, characterized in that: Adopt a multi-agent similarity aggregation mechanism to construct a loss function: Step 1.1: Adopt a multi-agent similarity aggregation mechanism to construct a loss function, as shown in formula (1). Among them, M represents the number of training samples, and C represents the number of categories. represents the multi-agent aggregation similarity between the i-th sample x i and the c i -th category, where c i ∈ [1, 2, 3,..., C] represents the category label of x i , represents the multi-agent aggregation similarity between the sample x i and the j-th category; Step 1.2: Aggregate the similarities between the sample and multiple proxies of each category using an aggregation mechanism to obtain the sample x i The multi-agent aggregation similarity with each category is shown in Equation (2). Among them, represents the multi-agent aggregation similarity between the i-th sample x i and the j-th category, where Aggregate represents the aggregation function. represents the similarity between the i-th sample x i and the k-th agent of the j-th category, where K represents the number of agents in each category. Step 1.3: The similarity between the \(i\)-th sample \(x\) i and the \(k\)-th proxy of the \(j\)-th category is calculated as shown in Equation (3). Among them, represents the k-th learnable agent of multiple j-th categories, T is the transpose operation, represents the d-dimensional feature vector extracted from x i by the neural network Net, and θ is the weight parameter of the neural network Net.
2. The cross-modal pedestrian re-identification method based on multi-agent similarity aggregation according to claim 1, wherein: The similarity aggregation method adopts mean aggregation, as shown in formula (4). 。 3. The cross-modal pedestrian re-identification method based on multi-agent similarity aggregation according to claim 1, characterized in that: The similarity aggregation method adopts attention mechanism aggregation, as shown in formula (5). Among them, represents the aggregation weight, which is calculated by the attention sub-network ANet, that is represents the d-dimensional feature vector extracted from x i by the neural network Net, and φ is the weight parameter of ANet.
4. The cross-modal pedestrian re-identification method based on multi-agent similarity aggregation according to claim 1, characterized in that: The similarity aggregation method adopts hard mining strategy aggregation, as shown in formula (6). Among them, c i ∈ [1, 2, 3, ..., C] represents the class label of the sample x i . Among the surrogates of the same class as x i (i.e., j = c i ), the minimum similarity is mined as the difficult similarity between x i and the j-th class for optimization. Among the surrogates of different classes from x i (i.e., j ≠ c i ), the maximum similarity is mined as the difficult similarity between x i and the j-th class for optimization.
Citation Information
Patent Citations
Unsupervised cross-modal retrieval method based on attention mechanism enhancement
CN113971209A
Visible light infrared pedestrian re-identification method based on multi-modal relation aggregation
CN114511878A