Re-identification method and system based on local perception and multivariate mutual information

Through the re-identification method of local perception and multivariate mutual correlation information, using Transformer and high-order reconstruction models, combined with tensor reconstruction and local perception branches, the vehicle re-identification model is optimized, which solves the shortcomings of the existing model in feature capture and improves the recognition performance.

CN117274621BActive Publication Date: 2025-09-23CHANGZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311153922.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-07
Publication Date
2025-09-23
Estimated Expiration
2043-09-07

AI Technical Summary

Technical Problem

Existing vehicle re-identification models based on convolutional neural networks are insufficient in capturing long-term dependencies between features and high-order fine-grained feature relationships, especially when the appearance similarity of vehicles of the same model makes recognition difficult.

Method used

A re-identification method based on local perception and multi-dimensional mutual correlation information is adopted. By introducing the mutual correlation model of Transformer and high-order reconstruction, combined with multi-head self-supervision module and tensor reconstruction technology, the complementary relationship between multi-dimensional information of vehicle images is captured, and the local perception branch is used to extract fine-grained features. The model is optimized by combining triple loss and cross entropy loss.

Benefits of technology

The performance of vehicle re-identification is improved, the mAP and CMC scores of the model on the VeRi-776 and VehicleID datasets are improved, and the discriminability and robustness of vehicle features are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117274621B_ABST
    Figure CN117274621B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and in particular to a re-identification method and system based on local perception and multivariate correlation information. The method comprises collecting vehicle images and constructing a dataset; introducing a Transformer and a high-order reconstruction cross-correlation model into a multivariate correlation information extraction branch, fusing multivariate key information of Transformer-based features and high-order reconstruction features, and capturing complementary information between multivariate information through pixel-level correlation operations; utilizing a fine-grained representation branch of local perception to divide a feature map into multiple sub-regions to learn local perception features to locate important regions of the vehicle image; and optimizing the target feature maps of the two branches using an objective function constructed using a triplet loss and a cross-entropy loss to obtain a vehicle re-identification result. The present invention overcomes the problem that existing methods are insufficient in capturing high-order, fine-grained features and the relationships between features in an image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a re-identification method and system based on local perception and multivariate mutual correlation information. Background Art

[0002] Re-ID aims to identify a specific vehicle in a dataset captured by non-overlapping cameras, which plays a huge role in the development of intelligent transportation systems.

[0003] Although convolutional neural network-based models have achieved impressive performance in re-identification tasks, their Gaussian receptive fields are limited in capturing long-term dependencies between features. Furthermore, vehicles of the same model and manufacturer may have similar appearances; therefore, capturing as many high-level, fine-grained features and relationships between features as possible from an image is crucial. Summary of the Invention

[0004] In view of the shortcomings of existing methods, the present invention overcomes the problem that existing methods are insufficient in capturing high-order fine-grained features and the relationship between features in images.

[0005] The technical solution adopted by the present invention is: a re-identification method based on local perception and multivariate mutual correlation information includes the following steps:

[0006] Step 1: Collect vehicle images and build a dataset;

[0007] Step 2: Introduce the Transformer and high-order reconstruction cross-correlation models into the multivariate cross-correlation information extraction branch, fuse the multivariate key information of the Transformer-based features and the high-order reconstruction features, and capture the complementary information between the multivariate information through pixel-level correlation operations;

[0008] Furthermore, step 2 specifically includes:

[0009] Step 21: Extract feature maps from CNN Flatten into 2D patches C, H and W are height, width and channel respectively, N = H × W;

[0010] Step 22: Use the transformer module of the first branch to process the input and output of the multi-head self-supervision module through residual connections and normalization layers to obtain the feature vector M;

[0011] Step 23: Introduce the feedforward network FFN to generate the pixel context-aware feature map F'.

[0012] Furthermore, it also includes:

[0013] Step 24: Feature Map Perform tensor reconstruction;

[0014] Step 25: Use channel generator, width generator and height generator to generate feature map Through pooling, convolution and Sigmoid activation at different angles, we obtain contextual attention maps at three angles.

[0015] Step 26: The obtained context attention map is multiplied by the element level to obtain fine-grained context features, the formula is:

[0016] Y={y1,y2,...,y CHW} (6)

[0017]

[0018] Where T={t1,t2...t i …t CHW}, T is the input feature map, t i Represents the i-th feature of the feature map; A={a1,a2,...a i ,a CHW} is the context attention map, Y={y1,y2,...,y CHW} is a fine-grained context feature;

[0019] Step 27: Aggregate the contextual sub-attention graph using weighted average to obtain a high-order tensor F, which is formulated as:

[0020]

[0021] Among them, λ i ∈(0,1) is the normalization factor, A i For the sub-attention map.

[0022] Step 28: Reconstruct the high-order tensor Perform global average pooling to obtain the target feature map of high-order information;

[0023] Step 29: Use pixelation to fuse the feature vectors F' and The formula is:

[0024]

[0025] Among them, F' is the context-aware feature map.

[0026] Step 3: Use the local perception fine-grained representation branch to divide the feature map into multiple sub-regions to learn local perception features to locate important areas of the vehicle image;

[0027] Furthermore, step three specifically includes:

[0028] Step 31: Input the image into the ResNet50 network to obtain a three-dimensional feature map G; where G = [G1, G2, ... G p ]stp∈[1,P],G p is the feature of the p-th part, p is the number of blocks;

[0029] Step 32: G p Add pooling layer and convolution layer to each feature G p Convert to P feature vectors to generate feature H p .

[0030] Step 4: Optimize the target feature maps of the two branches using the objective function constructed by triple loss and cross entropy loss to obtain the vehicle re-identification result;

[0031] Furthermore, the objective function formula is:

[0032] L=L softmax +λL tri (11)

[0033] Where λ is the equilibrium parameter, L softmax and L tri They represent softmax loss and triplet loss respectively.

[0034] Furthermore, the cross entropy loss is defined as:

[0035]

[0036] Among them, N i is the number of images in the mini-batch, N id represents the number of classes in the entire training set, y is the true value of the input feature, and x[j] represents the output of the fully connected layer containing two branch features for the jth identity.

[0037] Furthermore, the formula for triplet loss is:

[0038]

[0039] Among them, N k represents the number of identifications in a mini-batch, x P ,x N and x A Represents the positive, negative and anchor images of the vehicle dataset, x i,j are the identity identification and vehicle index, and α is the margin parameter.

[0040] Furthermore, a re-identification system based on local perception and multivariate mutual correlation information includes: a memory for storing instructions executable by a processor; and a processor for executing the instructions to implement a re-identification method based on local perception and multivariate mutual correlation information.

[0041] Furthermore, a computer-readable medium storing computer program code implements a re-identification method based on local perception and multivariate mutual correlation information when the computer program code is executed by a processor.

[0042] Beneficial effects of the present invention:

[0043] 1. The vehicle re-identification method of the present invention is used to extract discriminative feature representations of different vehicles, thereby optimizing the performance of vehicle re-identification;

[0044] 2. Comparing the proposed model with the existing model, the proposed model outperforms the existing model in terms of mAP and CMC scores on the VeRi-776 and VehicleID datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a flowchart of the re-identification method based on local perception and multivariate mutual correlation information of the present invention. DETAILED DESCRIPTION

[0046] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner, and therefore only shows the components related to the present invention.

[0047] like Figure 1 As shown in FIG, the re-identification method based on local perception and multivariate mutual correlation information includes the following steps:

[0048] Step 1: Collect vehicle images and build a dataset;

[0049] A dual-branch network is used to extract discriminative feature representations of different vehicles, thereby optimizing the performance of vehicle re-identification; the model consists of two branches, namely the multivariate mutual correlation information extraction branch and the fine-grained representation branch based on local perception.

[0050] Step 2: Introduce the Transformer and high-order reconstruction cross-correlation models into the multivariate cross-correlation information extraction branch, fuse the multivariate key information of the Transformer-based features and the high-order reconstruction features, and capture the complementary information between the multivariate information through pixel-level correlation operations;

[0051] In order to drive the model to capture information from different areas of the target vehicle and retain more discriminative features, the Transformer module is applied to obtain the context information of the entire vehicle image; therefore, the two branches extract feature maps from the vehicle image through the CNN backbone network and are represented as Among them, C, H and W are height, width and channel respectively.

[0052] In order to extract transformer-based feature representation, we flatten the feature map T into 2D patches, represented as Where N = H × W is the length of the feature patch sequence;

[0053] Then, in the multi-head self-supervised module (MHAM), the query matrix of each head of F is Key-value matrix Sum Matrix is generated by linear projection; where d k is the dimension of query and key matrix, d v is the dimension of the value matrix; the weighted sum of the attention weights can be obtained through the multi-head self-supervision module; therefore, the scaled dot product self-attention applied to Q, K, V can be expressed as follows:

[0054]

[0055] in, is the normalization coefficient, and the Softmax function is applied to each row of the input feature; the attention score reflects the relationship between the query matrix and the key matrix.

[0056] Then, the transformer module processes the input and output of the multi-head self-supervision module through residual connections and normalization layers to obtain the mapped feature vector M, as follows:

[0057] M=LN(F+MHAM(F)) (2)

[0058] Where LN represents the layer norm.

[0059] According to the standard transformer structure, a feedforward network (FFN) is introduced to generate the final pixel context-aware feature map F', as follows:

[0060] Q=LN(M+FFN(M)) (3)

[0061] Among them, FFN represents the linear projection of M, which can be defined as:

[0062] FFN(M)=ReLU(MW1)W2 (4)

[0063] Where W1 and W2 are represented as the parameters of the linear projection module with ReLU activation function.

[0064] F'=Conv(Reshape(Q)) (5)

[0065] Here, F' represents the discriminative features obtained from the feature map T in the multivariate correlation information extraction branch after passing through the Transformer module; Conv represents convolution, and Equation 5 converts the one-dimensional features into a feature matrix, facilitating matrix operations with the features output by the local information module. Based on these considerations, the multivariate correlation information extraction branch extracts a transformer-based feature representation to integrate the global information of a given image, which enables the CNN to capture more discriminative features from a global perspective.

[0066] In order to explore the context information of vehicle images to obtain higher-order and more discriminative features, the present invention uses the multivariate mutual correlation information extraction branch to perform tensor reconstruction on the CNN output feature map T; because high-rank tensors can be represented as reconstructions of multiple first-order tensors, there is no need to model the high-order context information of the target under channel compression to explore its high-order correlation; the high-order reconstruction part is divided into two parts, namely the first-order tensor generation module and the high-order tensor reconstruction module; by decomposing the high-rank problem into low-rank problems and maintaining the three-dimensional representation of the original space. The tensor reconstruction part mainly reconstructs the feature map obtained by the backbone network. Features are generated from the three perspectives of channel, width and height to obtain context information fragments in three directions; the context features generated in the three directions are processed to obtain a sub-attention map representing a part of the three-dimensional context feature; this process is repeated to reconstruct other context fragments, and the sub-maps are activated and aggregated through element-level multiplication and weighted averaging to obtain fine-grained context features in the spatial and channel dimensions.

[0067] For the tensor generation module, the feature generators of the three angles in the present invention are channel generator, width generator and height generator respectively; each generator is composed of a Pool-Conv-Sigmoid sequence, and the three-dimensional tensor obtained by the backbone network is pooled, convolved and activated with Sigmoid at different angles to obtain context fragments of the three angles, namely, the context attention map A.

[0068] For the tensor reconstruction module, the obtained context attention map A is obtained by element-wise multiplication to obtain fine-grained context features, as follows:

[0069] Y={y1,y2,...,y CHW} (6)

[0070] Where A={a1,a2,...a i ,a CHW} is the context attention map, Y={y1,y2,...,y CHW} is a fine-grained context feature, and the formula is:

[0071]

[0072] Where T={t1,t2...t i …t CHW}, T is the input feature map, t i Represents the i-th feature of the feature map.

[0073] The high-level features of the context features are processed by using tensor reconstruction, and the three context fragments are transformed into and Synthesize a sub-attention map A1 with a rank of 1. This sub-attention map represents part of the 3D context feature. Then, follow the same process to reconstruct other context fragments and use weighted average to aggregate these sub-attention maps. The formula is:

[0074]

[0075] Among them, λ i ∈(0,1) is a learnable normalization factor, A i is a sub-attention map; although each sub-attention map represents low-rank contextual information, their combination becomes a high-rank tensor.

[0076] Finally, the reconstructed high-order tensor matrix F is subjected to global average pooling to obtain a target feature map with higher-order information.

[0077] The feature vector F passed through the transformer module is connected to the feature vector F of the tensor reconstruction branch and optimized using the minimized cross-entropy loss. By using a block network to extract feature maps with more detailed features, and a low-order tensor reconstruction network to extract feature maps with higher-order feature information, the two feature maps are connected for prediction, resulting in a more robust and accurate re-identification prediction result.

[0078] In order to extract more discriminative features of the vehicle and integrate the features of the first branch of the network, the present invention adopts a cross-correlation operation to enhance the key information of a specific area; specifically, a pixel-wise operation is used to fuse the feature map F and Among them, the feature fusion module reconstructs the branch features of the tensor Decompose it into a small kernel of H×W and calculate the correlation between the two features to get the correlation representation feature H of the first branch. The mathematical formula can be expressed as:

[0079]

[0080] Therefore, the multivariate mutual information extraction branch considers the long-term dependencies between global vehicle features and can effectively preserve the interaction information between different features. At the same time, the tensor reconstruction module effectively provides auxiliary information for global features by highlighting the information of key features, improving the performance of the ReID model.

[0081] Step 3: Use the local perception fine-grained representation branch to divide the feature map into multiple sub-regions to learn local perception features to locate important areas of the vehicle image;

[0082] For the fine-grained representation branch of local perception, the features extracted from the backbone network are divided into multiple parts, and fine-grained features are extracted from the vehicle image. Specifically, the input image is fed into the backbone network ResNet50 to obtain a three-dimensional feature map G. For the fine-grained representation branch of local perception, the feature G is divided into six parts, which can be composed as follows:

[0083] G=[G1,G2,...G p ]stp∈[1,P] (10)

[0084] Among them, G p is the feature of the p-th part, and p is the number of blocks.

[0085] Therefore, the fine-grained representation branch explores the fine-grained information in the vehicle image. It captures the discriminative information representation of each sub-region from multiple similar vehicles. After that, a traditional pooling layer and a convolution layer with a 1×1 kernel are added to convert the features of each sub-region into P feature vectors to generate the feature H. p .

[0086] Step 4: Optimize the target feature maps of the two branches using the objective function constructed by triple loss and cross entropy loss to obtain the vehicle re-identification result;

[0087] Softmax loss and triplet loss are used as objective functions to optimize the network model. The formula is:

[0088] L=L softmax +λL tri (11)

[0089] Where λ is the equilibrium parameter, L softmax and L tri They represent softmax loss and triplet loss respectively; the purpose of softmax loss is to calculate the predicted probability that x belongs to the i-th category, and the softmax loss is defined as:

[0090]

[0091] Among them, N iis the number of images in the mini-batch, N id represents the number of classes in the entire training set, y is the true value of the input feature, and x[j] represents the output of the fully connected layer containing two branch features for the jth identity.

[0092] The purpose of triplet loss is to ensure the intra-class consistency and inter-class separability of features extracted by parallel networks; define and The triplet loss can be adjusted as follows:

[0093]

[0094] Among them, N k represents the number of identifications in a mini-batch, x P ,x N and x A Represents the positive, negative and anchor images of the vehicle dataset, x i,j are the identity and vehicle index, and α is a margin parameter that controls the distance between positive and negative images.

[0095] experiment:

[0096] Comparative experimental results of the model of the present invention on two vehicle datasets, namely the VeRi-776 dataset and the VehicleID dataset; wherein mAP and CMC are used to evaluate the performance of the model of the present invention.

[0097] Table 1 Comparison of mAP (%) and CMC scores (%) of Rnak1 and 5 of VeRi-776

[0098]

[0099] VeRi-776 is a large-scale urban surveillance vehicle dataset collected from real-world scenarios. It contains over 50,000 images of 776 vehicle IDs, including identity annotations, image timestamps, camera geolocation, vehicle color, and vehicle type information. Each vehicle ID has different view angles, lighting, and occlusion information. 576 vehicle classes are used as the training set, and the remaining data is used as the test set. A subset of 1,678 images in the test set is used to generate the query set.

[0100] The VehicleID dataset contains a total of 26,267 vehicles and 221,763 images; the training set contains 110,178 images of 13,134 vehicles; four subsets containing 800, 1,600, 2,400, and 3,200 vehicles respectively were extracted from the original test data, and searches were performed at different scales to provide test results when the test set was 800.

[0101] Table 2 Comparison of mAP (%) and CMC scores (%) of Rnak1 and 5 of VehicleID

[0102]

[0103]

[0104] The present invention adopts a dual-branch network method to extract the target vehicle to enhance the complementarity of the two branches. Specifically, the present invention introduces a cross-correlation model of Transformer and high-order reconstruction in the first branch, which fuses the multivariate key information of Transformer-based features and high-order reconstruction features, and captures the complementary information between the multivariate information through pixel-level correlation operations. Among them, the Transformer features are used to explore the global feature information of the vehicle, and the high-order tensor reconstruction branch is mainly used to obtain the high-order information of the vehicle. The above features are fused through the pixel-level feature fusion module to capture the local long-distance dependency information of the vehicle. The second branch divides the feature map into multiple sub-regions to learn local perception features to locate important areas of the vehicle image; by combining the above two branches, robust discriminant features can be extracted to improve the vehicle re-identification performance.

[0105] With the above-described preferred embodiments of the present invention as a guide, and with reference to the above description, relevant personnel are fully capable of making various changes and modifications without departing from the technical scope of this invention. The technical scope of this invention is not limited to the contents of the specification and must be determined according to the scope of the claims.

Claims

1. A re-identification method based on local perception and multivariate mutual correlation information, characterized in that: The following steps are involved: Step 1: Collect vehicle images and build a dataset; Step 2: Introduce the Transformer and high-order reconstruction cross-correlation models into the multivariate cross-correlation information extraction branch, fuse the multivariate key information of the Transformer-based features and the high-order reconstruction features, and capture the complementary information between the multivariate information through pixel-level correlation operations; Step 2 specifically includes: Step 21: Extract feature maps from CNN Flatten into 2D patches ; and are height, width and channel respectively, ; Step 22: Use the Transformer module to process the input and output of the multi-head self-supervision module through residual connections and normalization layers to obtain the feature vector ; Step 23: Introduce the feedforward network FFN to generate pixel context-aware feature maps ; Step 24: Feature Map Perform tensor reconstruction; Step 25: Use channel generator, width generator and height generator to generate feature map Through pooling, convolution and Sigmoid activation at different angles, we obtain contextual attention maps at three angles. Step 26: The obtained context attention map is multiplied by the element level to obtain fine-grained context features, the formula is: in, T ={ t 1, t 2... t i … t CHW }, T is the input feature map; t i Represents the feature map i Features is the context attention map, is a fine-grained context feature; Step 27: Aggregate the contextual sub-attention graph using weighted average to obtain a high-order tensor , the formula is: in, is the normalization factor, is the sub-attention map; Step 28: Reconstruct the high-order tensor Perform global average pooling to obtain the target feature map of high-order information; Step 29: Use pixelation to fuse feature vectors and , the formula is: in, is the context-aware feature map; Step 3: Use the local perception fine-grained representation branch to divide the feature map into multiple sub-regions to learn local perception features to locate important areas of the vehicle image; Step 4: Optimize the target feature maps of the two branches using the objective function constructed by triple loss and cross entropy loss to obtain the vehicle re-identification result.

2. The re-identification method based on local perception and multivariate mutual correlation information according to claim 1, characterized in that Step three specifically includes: Step 31: Input the image into the ResNet50 network to obtain a three-dimensional feature map ;in , It is The characteristics of each part, is the number of blocks; Step 32: Add pooling layer and convolution layer to transform each feature Convert to feature vectors to generate features .

3. The re-identification method based on local perception and multivariate mutual correlation information according to claim 1, characterized in that The formula of the objective function is: in, is the equilibrium parameter, and Respectively represent softmax Loss and triplet loss.

4. The re-identification method based on local perception and multivariate mutual correlation information according to claim 1, characterized in that The cross entropy loss is defined as: in, is the number of images in the mini-batch, represents the number of classes in the entire training set, is the true value of the input feature, Represents the output of the fully connected layer containing two branch features, used for the j an identity.

5. The re-identification method based on local perception and multivariate mutual correlation information according to claim 1, characterized in that The formula for triplet loss is: in, represents the number of identifications in a mini-batch, and Positive, negative and anchor images representing the vehicle dataset, is identification and vehicle indexing, is the margin parameter.

6. A re-identification system based on local perception and multivariate mutual correlation information, characterized in that: include: a memory for storing instructions executable by the processor; A processor, configured to execute instructions to implement the re-identification method based on local perception and multivariate mutual correlation information as described in any one of claims 1 to 5.

7. A computer-readable medium storing computer program code, characterized in that When the computer program code is executed by a processor, the computer program code implements the re-identification method based on local perception and multivariate mutual correlation information according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Vehicle re-identification method based on double sub-networks

    CN114067143A