A Multi-View Vehicle Re-Identification Model and Method in a Vehicle-Road Collaboration Scenario

By establishing the V2I-CARLA simulation data set in vehicle-road collaboration scenarios and using feature pyramid attention module and circle loss function, the problem of difficulty in vehicle feature extraction under different camera perspectives is solved, and high-precision multi-view vehicle re-identification is achieved, improving vehicle matching effect and robustness.

CN114743172BActive Publication Date: 2025-07-29JIANGSU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210064408.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-20
Publication Date
2025-07-29
Estimated Expiration
2042-01-20

AI Technical Summary

Technical Problem

In vehicle-road collaboration scenarios, there are differences in perspectives and changes in position of vehicles under different camera perspectives, making it difficult to capture vehicle details, existing data sets are small in number and low in quality, vehicle re-identification accuracy is not high, and existing algorithms ignore detailed features, resulting in high misidentification rate.

Method used

Establish a V2I-CARLA simulation dataset, adopt the feature pyramid attention module and circle loss function, combine the IBN block and CBAM in the feature extraction module, and use the circle loss optimization network to improve the re-identification accuracy of multi-view vehicle.

Benefits of technology

In the vehicle-road collaboration scenario, high-precision multi-view vehicle re-identification is achieved, which improves vehicle matching effect, data set richness and feature extraction integrity, and improves the accuracy and robustness of vehicle re-identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114743172B_ABST
    Figure CN114743172B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-view vehicle re-identification model and method in a vehicle-road collaborative scenario, and establishes a simulation dataset V2I-CARLA based on the vehicle-road collaborative scenario. Aiming at problems such as large differences in vehicle target scales in visual information collected by fixed platforms and mobile platforms and inability to obtain detailed features, a Feature Pyramid Attention module (FPA) is introduced in the feature extraction module to obtain multi-scale features of vehicles and improve the integrity of capturing vehicle features. Aiming at the problem of insufficient discrimination of similar detailed features, it is proposed to apply Circle loss to the vehicle re-identification task to strengthen the discrimination power of the algorithm and improve the vehicle matching effect. The present invention achieves a high matching accuracy in the vehicle-road collaborative simulation scenario, and also has a good matching effect on the publicly available vehicle re-identification dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent connected vehicles, and particularly designs a multi-view vehicle re-identification model and method based on vehicle-road cooperation. Background Technique

[0002] Vehicle-road collaborative perception is an effective direction for the future development of intelligent vehicles towards higher-level autonomous driving. Based on vehicle-road cooperation, not only can the intelligent level of individual vehicles be improved, but also the intelligent development of roads can be realized. During the autonomous driving process based on vehicle-road cooperation, not only can the vehicle itself complete the perception of the surrounding environment, but also it can obtain in advance the road conditions and real-time dynamics of various traffic participants in its surrounding area or even a larger range, and make scientific and reasonable decision-making plans and controls based on this, thereby improving driving safety and the overall traffic efficiency. Compared with individual vehicle intelligence, vehicle-road collaborative autonomous driving has advantages such as over-the-horizon vision, strong environmental adaptability, strong prediction ability, and reasonable computing power configuration, providing strong support for safe, efficient, and low-cost autonomous driving.

[0003] Vehicle-road cooperation is mainly reflected in four aspects: communication, perception, decision-making, and control. Among them, the perception of vehicle-road cooperation includes perceiving various traffic participation objects such as road people and vehicles, weather conditions, traffic accidents, and traffic signals. Accurate road perception plays a crucial role in the subsequent decision-making and control tasks of vehicle-road cooperation. The main objects of the vehicle-road cooperation system are fixed platforms and mobile platforms. Through the identification and information interaction between the vehicle end and the road end, a high-reliability and over-the-horizon driving environment information is established. As the most important participant in the road traffic system, the multi-view re-identification of vehicle targets on the mobile platform and the roadside platform has become the core of vehicle-road collaborative perception and is a key step to achieve vehicle and even group intelligence.

[0004] The cross-camera image matching of vehicles is to solve the problem of re-matching the vehicle target when it disappears in one field of view and appears in another field of view. In the vehicle-road cooperation scenario, it is to re-match the target vehicle through the road target detection of the fixed platform after it disappears in the field of view of the on-vehicle platform. In recent years, with the development of deep learning, re-identification (ReID) has become the most direct method to solve the cross-camera matching task and has achieved remarkable results.

[0005] In the vehicle-road collaborative scenario, the vehicle-mounted camera is a mobile camera, and the roadside camera is a fixed camera. At the same time, the relative installation height and inclination angle of the vehicle-mounted camera and the roadside camera are also different, resulting in problems such as different perspectives, pose changes, and inconsistent scale sizes of the vehicle in the vehicle-mounted camera and the roadside camera. Different perspectives cause the vehicle-mounted camera to mainly obtain the front and rear image information of surrounding vehicles during driving, while the roadside camera obtains more side image information of vehicles due to perspective limitations. This will increase the difficulty of extracting common features of the same target under the mobile platform and the fixed platform. The pose changes and scale size differences of the vehicle under multi-perspective cameras make it difficult for visual sensors to capture the detailed features of the vehicle. In recent years, vehicle re-identification schemes mainly focus on the global features of the vehicle's appearance, ignoring the attention to the detailed features of the vehicle's structure, resulting in more misidentifications of vehicles with similar models and low reliability of the matching effect.

[0006] At the same time, the existing publicly available datasets based on vehicle re-identification are few in number and low in quality, and most of the existing publicly available datasets are captured under one or more fixed roadside cameras, which do not meet the conditions of the vehicle-road collaborative perception scenario. Summary of the Invention

[0007] In order to overcome the shortcomings of the prior art, the present invention proposes a multi-perspective vehicle re-identification method in the vehicle-road collaborative scenario. A simulation dataset V2I-CARLA based on the vehicle-road collaborative scenario is established, an algorithm framework for cross-camera matching in the vehicle-road collaborative scenario is proposed, a Feature Pyramid Attention module (FPA) is introduced into the feature extraction module, and it is proposed to apply Circle loss to the vehicle re-identification task to improve the multi-perspective vehicle re-identification accuracy in the vehicle-road collaborative scenario. The specific technical solutions are as follows:

[0008] A multi-perspective vehicle re-identification model in the vehicle-road collaborative scenario, including a feature extraction module, a Feature Pyramid Attention module, and a loss function module; among them, a Feature Pyramid Attention module is introduced into the feature extraction module, and Circle loss is used in the loss function module to implement the vehicle re-identification task.

[0009] Furthermore, the feature extraction module: Add an IBN block to the original resnet50 backbone. The IBN block includes Instance Normalization (IN) and Batch Normalization (BN). BN is used to improve the sensitivity of features to image content and reduce the robustness to changes in image style, while IN is used to improve the robustness to style changes. Resnet consists of 4 groups of residual blocks. IN and BN are used simultaneously in the first three residual blocks. IN and BN are added after the first convolution in the residual blocks of resnet50 and used simultaneously. For the output channels (64, 128, 256) of the first convolution, half of the channels (32, 64, 128) pass through BN, and the other half (32, 64, 128) pass through IN. At the same time, a channel attention and spatial attention module CBAM is introduced into the backbone network. CBAM is added after the first convolution and before the first residual block in resnet50. The channel attention and spatial attention in CBAM are arranged in sequence. The channel attention uses average pooling (Avg Pooling) and max pooling (MaxPooling) on the features after the first convolution to obtain the average pooling feature and the max pooling feature. Then these two features are passed to a shared network containing a multi-layer perceptron (shared MLP) and a hidden layer to obtain the channel attention map. After that, the channel attention map is input into the spatial attention part. First, a channel-based global max pooling and global average pooling are performed, and then the two obtained features are added together. After a convolution operation and a sigmoid activation function, the spatial attention feature is generated.

[0010] Furthermore, the feature pyramid attention module: Uses three convolutional kernels of different sizes, 3×3, 5×5, and 7×7, to extract pyramid features of different scales. This module contains four branches and a main branch. After the convolution operations of the three branches using different convolutional kernel sizes, deconvolution is used for upsampling respectively. The feature maps obtained from each branch are added together in sequence from top to bottom. The main branch introduces a 1×1 convolution, and the original features extracted by the feature extraction module are pixel-multiplied with the pyramid attention features generated by the three branches with different convolutional kernel sizes to obtain multi-scale features. Another branch introduces global pooling and a 1×1 convolution, and then performs upsampling to output global features. Finally, the global features are added to the previously obtained multi-scale features to obtain the final output features.

[0011] Furthermore, the loss function module: Introduce Circle Loss to replace Triplet Loss, and the Circle Loss function is designed as:

[0012]

[0013] Where: O p = 1 + m, O n = -m, Δ p = 1 - m, Δ n = m. [·] + Represents the non - negative operation, α n Is the non - negative integer weight factor for negative samples, α p Is the non - negative integer weight factor for positive samples, O n And O p Are And The optimized ideal values of, that is, O n = 0 and O p = 1, Δ n And Δ p Are the threshold parameters between classes and within classes respectively. In Circle loss, s n And s p Are asymmetric, so s n And s p Each requires a threshold respectively. The value of the boundary coefficient m is 0.25, allowing to maximize the intra - class similarity s n And minimize the inter - class similarity s p To learn at different speeds. γ is the adjustment coefficient, and its value is 256;

[0014] Use the classification loss (ID Loss) and Circle Loss together to optimize the same feature. The classification loss is calculated using cross - entropy loss:

[0015]

[0016] p i Is the predicted ID probability, q i Is the positive sample probability, and N is the number of labels;

[0017] The label form is in One - hot encoding. q i Has only two values, 0 and 1. When y ≠ i, q i = 0; when y = i, q i = 1; where y is the true ID label.

[0018] When calculating the classification loss, the method of Label Smoothing is used, that is, when i = y, When ε is a parameter with a value of 0.1;

[0019] At the same time, circle loss and classification loss are used to optimize the entire network structure, and the overall loss function is:

[0020] L sum = L circle + L cross .

[0021] Furthermore, for the training of this model: the size of the input pictures is unified to 256×128, the batchsize is set to 32, the Adam optimizer is selected during training, and the initial learning rate is set to 0.05.

[0022] Furthermore, for the testing of this model: the network model extracts the features of the input query pictures and the vehicle pictures in the gallery, and calculates the Euclidean distance between the features to compare the similarity between the query pictures and the gallery pictures. The calculation formula of the Euclidean distance is d = ||f q - f g ||2, where ||·|| is the 2-norm of the corresponding feature vector.

[0023] A dataset for the multi-view vehicle re-identification model in the vehicle-road cooperation scenario described above is made as follows:

[0024] Use the open-source simulator CARLA to build a vehicle-road cooperation dual-camera scenario to collect vehicle picture data. Specifically: Select Town05 as the experimental map, which includes a section of highway, a large number of streets and intersections. Select two scenarios, the highway and the street, to build the experimental scenario. On the highway and street sections, set a fixed camera with a pitch angle of -45° and a rotation angle of 30° at a position 3 meters above the ground at intervals of about 10m along the roadside. Set a car at the end of the road, and set a fixed camera with a pitch angle of 0° and a rotation angle of 90° on the car roof;

[0025] Set the vehicle to autonomous driving mode, and the vehicle-mounted camera and the roadside camera collect vehicle pictures at the same time. The vehicle-mounted camera mainly collects pictures of the front and back views of the vehicle, and the roadside view mainly obtains pictures of the side of the vehicle. With this experimental scenario, a total of 153 vehicles with different colors and models are collected. Cut out the vehicles in the collected pictures, and label the cut-out vehicle pictures according to the vehicle id and camera id information;

[0026] The dataset includes 1500 dual-view images with different IDs, among which 1045 are training images and 455 are test images. The images are randomly flipped, randomly erased, and cropped to enrich the picture scenarios, and the images are uniformly adjusted to 256×128.

[0027] The present invention also proposes a multi-view vehicle re-identification method in a vehicle-road collaborative scenario, designs a multi-view vehicle re-identification model in the vehicle-road collaborative scenario, and uses this model to realize multi-view vehicle re-identification in the vehicle-road collaborative scenario.

[0028] The model design stage includes:

[0029] S2. Design a feature extraction module:

[0030] Add an IBN block to the original resnet50 backbone. The IBN block includes Instance Normalization (IN) and Batch Normalization (BN). BN is used to improve the sensitivity of features to image content and reduce the robustness to changes in picture style. IN is used to improve the robustness to style changes. Resnet includes 4 groups of residual blocks. IN and BN are used simultaneously in the first three residual blocks. IN and BN are added after the first convolution of the residual blocks in renet50. IN and BN are used simultaneously, and for the output channels (64, 128, 256) of the first convolution, half of the channels (32, 64, 128) pass through BN, and the other half of the channels (32, 64, 128) pass through IN. At the same time, a channel attention and spatial attention module CBAM is introduced into the backbone network. CBAM is added after the first convolution and before the first residual block in resnet50. The channel attention and spatial attention in CBAM are arranged in sequence. The channel attention uses average pooling (Avg Pooling) and max pooling (Max Pooling) simultaneously for the features after the first convolution to obtain the average pooling feature and the max pooling feature. Then these two features are passed to a shared network containing a multi-layer perceptron (shared MLP) and a hidden layer to obtain the channel attention map. Then the channel attention map is input into the spatial attention part. First, a channel-based global max pooling (Global Max Pooling) and global average pooling (Global Average Pooling) are performed, and then the two obtained features are added. After a convolution operation and a sigmoid activation function, the spatial attention feature is generated.

[0031] S3. Design a feature pyramid attention module

[0032] Use three convolution kernels of different sizes, namely 3×3, 5×5, and 7×7, to extract pyramid features of different scales. This module contains four branches and a main branch. After convolution operations are performed on the three branches using different convolution kernel sizes, deconvolution is used for upsampling respectively. The feature maps obtained from each branch are added together in sequence from top to bottom. The main branch introduces a 1×1 convolution, and the original features extracted by the feature extraction module are pixel-multiplied with the pyramid attention features generated by the three branches with different convolution kernel sizes to obtain multi-scale features. Another branch introduces global pooling and 1×1 convolution, and then performs upsampling to output global features. Finally, the global features are added to the previously obtained multi-scale features to obtain the final output features;

[0033] S4. Design a loss function for the re-identification task

[0034] Introduce Circle Loss to replace Triplet Loss. The Circle Loss function is designed as follows:

[0035]

[0036] Where: O p = 1 + m, O n = -m, Δ p = 1 - m, Δ n = m. [·] + Represents a non-negative operation, α n Is a non-negative integer weight factor for negative samples, α p Is a non-negative integer weight factor for positive samples, O n And O p Are And The optimized ideal values of, that is, O n = 0 and O p = 1, Δ n And Δ p Are the threshold parameters between classes and within classes respectively. In Circle loss, s n And s p Are asymmetric, so s n And s p Each requires a threshold. The value of the boundary coefficient m is 0.25, allowing the maximization of the within-class similarity s n And the minimization of the between-class similarity s p To learn at different speeds. γ is a regulation coefficient, and its numerical value is 256;

[0037] Use the classification loss (ID Loss) and Circle Loss together to optimize the same feature. The classification loss is calculated using cross-entropy loss:

[0038]

[0039] p i To predict the ID probability, q i is the positive sample probability, and N is the number of labels;

[0040] The label form is in One-hot encoding form, q i has only two values, 0 and 1. When y ≠ i, q i = 0; when y = i, q i = 1; where y is the true ID label.

[0041] When calculating the classification loss, the method of Label Smoothing is used, that is, when i = y, when, ε is a parameter with a value of 0.1;

[0042] At the same time, circular loss and classification loss are used to optimize the entire network structure, and the overall loss function is:

[0043] L sum = L circle + L cross ;

[0044] Model training and testing stages:

[0045] The size of the input images is uniformly set to 256×128, the batchsize is set to 32, and the Adam optimizer is selected during training, with the initial learning rate set to 0.05;

[0046] Extract the features of the input query images and the vehicle images in the gallery, and calculate the Euclidean distance between the features to compare the similarity between the query images and the gallery images. The calculation formula for the Euclidean distance is d = ||f q - f g ||2, and ||·|| is the 2-norm of the corresponding feature vector.

[0047] Furthermore, it also includes: S1. Produce a dataset for model training and testing

[0048] Use the open-source simulator CARLA to build a vehicle-road collaborative dual-camera scenario to collect vehicle image data. Specifically: Select Town05 as the experimental map, which includes a section of highway, a large number of streets, and intersections. Select two scenarios, the highway and the street, to build the experimental scenario. On the highway and street sections, set a fixed camera with a pitch angle of -45° and a rotation angle of 30° at a position 3 meters above the ground at intervals of about 10 meters along the roadside. Set a car at the end of the road, and set a fixed camera on the car roof with a pitch angle of 0° and a rotation angle of 90°.

[0049] Set the vehicle to the autonomous driving mode, and the vehicle-mounted camera and the roadside camera collect vehicle images simultaneously. The vehicle-mounted camera mainly collects images from the front and back views of the vehicle, and the roadside view mainly obtains the side images of the vehicle. In this experimental scenario, a total of 153 vehicles with different colors and models are collected. Cut out the vehicles in the collected images, and label the cut-out vehicle images according to the vehicle ID and camera ID information.

[0050] The dataset includes 1500 dual-view images with different IDs, of which 1045 are training images and 455 are test images. Randomly flip, randomly erase, and crop the images to enrich the richness of the image scenarios, and uniformly adjust the images to 256×128.

[0051] The beneficial effects of the present invention are as follows:

[0052] (1) Aiming at the problem of insufficient dataset, use the Carla autonomous driving simulator to build a vehicle-road collaborative multi-camera scenario, collect multi-view vehicle visual information, and establish a simulated dataset V2I-CARLA based on the vehicle-road collaborative scenario;

[0053] (2) Aiming at the problem of fewer common features of vehicles in the vehicle-road collaborative scenario, based on the strongReID baseline, an algorithm framework for cross-camera matching in the vehicle-road collaborative scenario is proposed;

[0054] (3) Aiming at the problems of large differences in vehicle target scales and inability to obtain detailed features in the visual information collected by fixed platforms and mobile platforms, introduce the Feature Pyramid Attention module (FPA) into the feature extraction module to obtain multi-scale features of vehicles and improve the integrity of capturing vehicle features.

[0055] (4) Aiming at the problem of insufficient discrimination of similar detailed features, propose to apply Circle loss to the vehicle re-identification task to strengthen the discrimination of the algorithm and improve the vehicle matching effect.

[0056] (5) This algorithm achieves a high matching accuracy in the vehicle-road collaborative simulation scenario and also has a good matching effect on the publicly available vehicle re-identification dataset. Among them, the matching accuracy on the V2I-CARLA dataset reaches 82.93%, and the rank@1 matching accuracy reaches 81.7%. Brief Description of the Drawings

[0057] Figure 1 . Overall Structure Diagram of the Algorithm of the Present Invention

[0058] Figure 2 . Structure Diagram of ResNet50-IBN of the Present Invention

[0059] Figure 3 . Structure Diagram of the Attention Mechanism (CBAM) of the Present Invention

[0060] Figure 4 . Structure Diagram of the Feature Pyramid Attention Module of the Present Invention

[0061] Figure 5 . Visualization Results of the Present Invention on the Dataset Detailed Implementation Manner

[0062] The present invention will be further described below with reference to the accompanying drawings.

[0063] As Figure 1 shown is the overall structure diagram of the algorithm of the present invention. After the image data acquired by the camera is used to extract feature information through the backbone network, the feature map is input into the feature pyramid attention module. Calculate a binary classification loss and a cyclic loss for the features output by the feature pyramid attention module, reduce the feature size through a 512-dimensional fully connected layer and a batch normalization layer, and finally output the final prediction result through a fully connected layer.

[0064] The present invention first makes an explanation of the names:

[0065] ResNet50 is a neural network structure based on the residual structure. This network connects the input information and the output information through skip connections, effectively protecting the integrity of the information, solving the problem of gradient explosion / vanishing caused by common linear stacking, and improving the network's ability to extract features.

[0066] Carla is an open-source autonomous driving simulator. It simulates various driving conditions, provides open digital assets (city layout, buildings, vehicles), and supports flexible specification of sensor suites and environmental conditions.

[0067] The specific implementation process of the present invention includes the following:

[0068] Step 1: Dataset Creation

[0069] The open-source simulator CARLA for autonomous driving research is used to build a vehicle-road collaborative dual-camera scenario to collect vehicle image data. CARLA mainly targets 3D urban scenarios and has maps of 5 different urban scenarios. Among them, we choose Town05 as our experimental map. The map Town05 contains a section of highway, a large number of streets and intersections. We select two scenarios, the highway and the street, to build our experimental scenario. On the highway and street sections, we set up fixed cameras with a pitch angle of -45° and a rotation angle of 30° at a position 3 meters above the ground at intervals of about 10 meters along the roadside. At the end of the road, a car is set up, and a fixed camera with a pitch angle of 0° and a rotation angle of 90° is set on the car roof. The vehicle is set to autonomous driving mode. The vehicle-mounted camera and the roadside camera collect vehicle images simultaneously. The vehicle-mounted camera mainly collects images of the front and back views of the vehicle, and the roadside view mainly obtains the side images of the vehicle. With this experimental scenario, a total of 153 vehicles with different colors and models are collected. The vehicles in the collected images are cropped out of the images, and the cropped vehicle images are labeled according to the vehicle ID and camera ID information. The dataset includes 1500 dual-view images with different IDs, of which 1045 are training images and 455 are test images. Before feeding into the network structure, data augmentation methods such as random flipping, random erasing, and cropping are performed on the images to expand the richness of the image scenarios, and the images are uniformly adjusted to 256×128. Step 2: Generate the feature extraction module (ResNet-50+)

[0070] The feature extraction module of the present invention is based on the general ResNet-50 structure. An IBN block is added to the original resnet50 backbone. The IBN block consists of Instance Normalization (IN) and Batch Normalization (BN). BN improves the sensitivity of features to image content and reduces the robustness to changes in image style. IN improves the robustness to style changes. Resnet consists of 4 groups of residual blocks. IN and BN are used simultaneously in the first three residual blocks. Combining IN and BN and applying them to resnet50 effectively improves the classification accuracy and the robustness of the model. In the present invention, IN and BN are added after the first convolution of the residual block in renet50. IN and BN are used simultaneously. For the output channels (64, 128, 256) of the first convolution, half of the channels (32, 64, 128) are calculated through BN, and half of the channels (32, 64, 128) are calculated through IN. The specific structural schematic diagram is shown in the appendix Figure 3As shown. At the same time, channel attention and spatial attention modules CBAM are introduced into the backbone network. We add CBAM after the first convolution and before the first residual block in resnet50. The channel attention and spatial attention in CBAM are arranged in sequence. The channel attention uses both average pooling (Avg Pooling) and max pooling (Max Pooling) for the features after the first convolution to obtain the average pooling feature and the max pooling feature. Then these two features are passed into a shared network composed of a multi-layer perceptron (shared MLP) and a hidden layer to obtain the channel attention map. Then the channel attention map is input into the spatial attention module, where a channel-based global max pooling (Global Max Pooling) and global average pooling (Global Average Pooling) are first performed. Then the two obtained features are added together, and after a convolution operation and a sigmoid activation function, the spatial attention feature is generated. The specific structure is shown in the appendix Figure 2 As shown.

[0071] Step 3: Design the feature pyramid attention module

[0072] Adopt the feature pyramid attention module as shown in the appendix Figure 4 As shown. This module uses three convolutional kernels of different sizes, 3×3, 5×5, and 7×7, to extract pyramid features of different scales. This module mainly consists of four branches and one main branch. After the three branches using different convolutional kernel sizes go through the convolution operation, they are respectively upsampled using deconvolution. The feature maps obtained by each branch are added together in sequence from top to bottom. The main branch introduces a 1×1 convolution, and the original features extracted by the feature extraction module are pixel-multiplied with the pyramid attention features generated by the three branches with different convolutional kernel sizes to obtain multi-scale features. Another branch introduces global pooling, a 1×1 convolution, and then is upsampled to output global features. Finally, the global features are added to the previously obtained multi-scale features to obtain the final output features. This module can obtain feature information of different scales, can expand the range of the receptive field, and the multi-scale features have beneficial effects on the recognition of small targets.

[0073] Step 4: Design the loss function for the re-identification task

[0074] In the re-identification task, the combination of triplet loss and classification loss is often used as the loss function of the network. The present invention innovatively introduces the Circle Loss for the face recognition task to replace the Triplet Loss. The Circle Loss is a loss function improved based on the Triplet Loss. Different from the characteristic of the Triplet Loss function that equally optimizes positive and negative samples, the Circle Loss adds a weight to the positive and negative sample pairs to control the gradient contributions of the positive and negative sample pairs respectively, and dynamically adjusts the weight of backpropagation according to the classification probability. The designed Circle Loss function that enables the model to have a stronger discrimination ability when identifying different types of vehicles is as follows:

[0075]

[0076] Where: O p = 1 + m, O n = -m, Δ p = 1 - m, Δ n = m. [·] + represents the non-negative operation. Among them, α n is the non-negative integer weight factor for negative samples, α p is the non-negative integer weight factor for positive samples, L is the number of negative samples, and K is the number of positive samples. is the inter-class similarity of negative samples, is the intra-class similarity of positive samples. O n and O p are and optimized ideal values, that is, O n = 0 and O p = 1. Δ n and Δ p are the threshold parameters for inter-class and intra-class respectively. Because s n and s p are asymmetric in Circle loss, so s n and s p each requires a threshold. In the present invention, the value of the boundary coefficient m is 0.25. This allows maximizing the intra-class similarity s n and minimizing the inter-class similarity s p to learn at different speeds. γ is the adjustment coefficient, and the numerical value in the present invention is 256.

[0077] The present invention uses the classification loss (ID Loss) and the Circle Loss in combination to optimize the same feature. The classification loss is calculated using the cross-entropy loss. The calculation formula of the cross-entropy loss is:

[0078]

[0079] Since the classification category is determined by the vehicle ID, this loss is called the ID Loss. p i For predicting the ID probability, q i is the positive sample probability, and N is the number of labels. Since the label is in the form of one-hot encoding, that is, q i has only two values, 0 and 1. When y≠i, q i =0; when y = i, q i =1, where y is the true ID label. Therefore, using the cross-entropy loss function for optimization is prone to overfitting. To alleviate the occurrence of overfitting, the method of label smoothing is used in calculating the classification loss. That is, when i = y, when ε is a parameter with a value of 0.1.

[0080] The present invention simultaneously uses circle loss and classification loss to optimize the entire network structure, and the overall loss function is L sum =L circle +L cross .

[0081] Step 5: Feed the V2I-CARLA dataset into the network model of the present invention for training and test verification.

[0082] The training and test platform is a workstation with a GTX1080Ti GPU. During the experiment, the size of the input image is uniformly set to 256×128. The batch size is set to 32. When training, the Adam optimizer is selected, and the initial learning rate is set to 0.05. The present invention uses the Resnet50 network pre-trained on Imagenet as the basic feature extraction network.

[0083] Step 6: Test phase

[0084] The network extracts the features of the input query image and the vehicle images in the gallery, and compares the similarity between the query image and the gallery images by calculating the Euclidean distance between the features. The calculation formula of the Euclidean distance is d = ||f q -f g ||2, where ||·|| is the 2-norm of the corresponding feature vector. The visualization results are as Figure 5As shown. Among them, the first and third lines are the query results of the query vehicle on the Baseline network, and the second and fourth lines are the query results of the present invention. In the V2I-CARLA dataset, the front view or rear view image of the vehicle captured by the in-vehicle camera is used as the query image, and the vehicle images collected by the roadside camera are used as the query image library. The results show that the Baseline network always has misdetection when retrieving the correct vehicle. The present invention can successfully retrieve the top ten vehicles in the ranking list. The matching accuracy of the Baseline method on the V2I-CARLA dataset reaches 70.4%, and the rank@1 matching accuracy reaches 70.59%. While the matching accuracy of the present invention on the V2I-CARLA dataset reaches 82.93%, and the rank@1 matching accuracy reaches 81.7%, which shows the effectiveness of the present invention.

[0085] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation manners of the present invention, and they are not used to limit the protection scope of the present invention. Any equivalent manners or changes that do not depart from the technology created by the present invention should be included in the protection scope of the present invention.

Claims

1. A multi-view vehicle re-identification model in a vehicle-road collaborative scenario, characterized in that, It includes a feature extraction module, a feature pyramid attention module, and a loss function module. Among them, the feature pyramid attention module is introduced into the feature extraction module, and the Circle loss is used in the loss function module to implement the vehicle re-identification task. The feature extraction module: An IBN block is added to the original resnet50 backbone. The IBN block includes Instance Normalization (IN) and Batch Normalization (BN). BN is used to improve the sensitivity of features to image content and reduce the robustness to changes in picture style, while IN is used to improve the robustness to style changes. Resnet consists of 4 groups of residual blocks. IN and BN are both used in the first three residual blocks. IN and BN are added after the first convolution of the residual blocks in resnet50 and used simultaneously. For the output channels (64, 128, 256) of the first convolution, half of the channels (32, 64, 128) pass through BN, and the other half (32, 64, 128) pass through IN. At the same time, the Channel Attention and Spatial Attention Module (CBAM) is introduced into the backbone network. CBAM is added after the first convolution and before the first residual block in resnet50. The channel attention and spatial attention in CBAM are arranged in sequence. The channel attention uses both average pooling (Avg Pooling) and max pooling (Max Pooling) for the features after the first convolution to obtain the average pooling feature and the max pooling feature. Then these two features are passed to a shared network containing a multi-layer perceptron (shared MLP) and a hidden layer to obtain the channel attention map. After that, the channel attention map is input into the spatial attention part. First, a channel-based global max pooling (Global Max Pooling) and global average pooling (Global Average Pooling) are performed, and then the two obtained features are added together. After a convolution operation and a sigmoid activation function, the spatial attention feature is generated. The feature pyramid attention module: Convolution kernels of three different sizes, 3×3, 5×5, and 7×7, are used to extract pyramid features of different scales. This module contains four branches and a main branch. After the convolution operation, the three branches using different convolution kernel sizes are upsampled using transposed convolution respectively. The feature maps obtained by each branch are added together in sequence from top to bottom. The main branch introduces a 1×1 convolution, and the original features extracted by the feature extraction module are pixel-multiplied with the pyramid attention features generated by the three branches with different convolution kernel sizes to obtain multi-scale features. Another branch introduces global pooling and a 1×1 convolution, and then is upsampled to output global features. Finally, the global features are added to the previously obtained multi-scale features to obtain the final output features. Loss function module: Introduce Circle Loss to replace Triplet Loss. The Circle Loss function is designed as follows: Wherein: O p = 1 + m, O n = -m, △ p = 1 - m, △ n = m; L is the number of negative samples, and K is the number of positive samples. is the inter-class similarity of negative samples, is the intra-class similarity of positive samples, [·] + represents a non-negative operation, α n is the non-negative integer weight factor of negative samples, α p is the non-negative integer weight factor of positive samples, O n and O p are and the optimized ideal values, i.e., O n = 0 and O p = 1, △ n and △ p are the threshold parameters for inter-class and intra-class respectively. In Circle loss, s n and s p are asymmetric, so s n and s p each requires a threshold. The value of the boundary coefficient m is 0.25, allowing the maximization of the intra-class similarity s n and the minimization of the inter-class similarity s p to learn at different speeds. γ is a regulation coefficient, and its numerical value is 256; Use the classification loss (ID Loss) and Circle Loss together to optimize the same feature. The classification loss is calculated using cross-entropy loss: p i For predicting the ID probability, q i is the positive sample probability, and N is the number of labels; The label form is in the form of one-hot encoding, q i There are only two values, 0 and 1. When y ≠ i, q i = 0; when y = i, q i = 1; where y is the true ID label; When calculating the classification loss, the method of label smoothing is used, that is, when i = y, when i ≠ y, ε is a parameter with a value of 0.1; At the same time, use Circle Loss and classification loss to optimize the entire network structure. The overall loss function is: L sum = L circle + L cross 。 2. The multi-view vehicle re-identification model in a vehicle-road collaborative scenario according to claim 1, wherein Training of this model: Unify the size of the input images to 256×128, set the batch size to 32, select the Adam optimizer during training, and set the initial learning rate to 0.

05.

3. The multi-view vehicle re-identification model in a vehicle-road collaborative scenario according to claim 1, wherein, Testing of this model: The network model extracts the features of the input query image and the vehicle images in the image library, and calculates the Euclidean distance between the features to compare the similarity between the query image and the images in the image library. The calculation formula of the Euclidean distance is d = ||f q - f g ||2, where ||·|| is the 2-norm of the corresponding feature vector.

4. A dataset for a multi-view vehicle re-identification model in a vehicle-road collaborative scenario as described in claim 1, characterized in that, Use the open-source simulator CARLA to build a vehicle-road collaborative dual-camera scenario to collect vehicle image data. Specifically: Select Town05 as the experimental map, which includes a section of highway, streets, and intersections. Select two scenarios, the highway and the street, to build the experimental scenario. On the highway and street sections, set a fixed camera with a pitch angle of -45° and a rotation angle of 30° at a position 3 meters above the ground every 10 meters at an interval of 10 meters along the roadside. Set a car at the end of the road, and set a fixed camera with a pitch angle of 0° and a rotation angle of 90° on the car roof; Set the vehicle to autonomous driving mode. The vehicle-mounted camera and the roadside camera collect vehicle images at the same time. The vehicle-mounted camera mainly collects front and rear view images of the vehicle, and the roadside view mainly obtains side images of the vehicle. In this experimental scenario, a total of 153 vehicles with different colors and models are collected. Cut out the vehicles in the collected images and label the cut-out vehicle images according to the vehicle ID and camera ID information; The dataset includes 1500 dual-view images with different IDs, including 1045 training images and 455 test images. Randomly flip, randomly erase, and crop the images to enrich the image scenarios, and uniformly adjust the image size to 256×128.

5. A multi-view vehicle re-identification method in a vehicle-road collaborative scenario, characterized in that, Design a multi-view vehicle re-identification model in the vehicle-road collaborative scenario, and use this model to realize multi-view vehicle re-identification in the vehicle-road collaborative scenario. The model design stage includes: S2. Design the feature extraction module: Add the IBN block to the original ResNet50 backbone. The IBN block includes Instance Normalization (IN) and Batch Normalization (BN). BN is used to improve the sensitivity of features to image content and reduce the robustness to changes in image style, while IN is used to improve the robustness to style changes. ResNet consists of 4 groups of residual blocks. IN and BN are both used in the first three residual blocks. IN and BN are added after the first convolution in the residual blocks of ResNet50 and used simultaneously. For the output channels (64, 128, 256) of the first convolution, half of the channels (32, 64, 128) pass through BN and the other half (32, 64, 128) pass through IN. At the same time, the Channel Attention and Spatial Attention Module (CBAM) is introduced into the backbone network. CBAM is added after the first convolution and before the first residual block in ResNet50. The channel attention and spatial attention in CBAM are arranged in sequence. The channel attention uses both average pooling (Avg Pooling) and max pooling (Max Pooling) for the features after the first convolution to obtain the average pooling feature and the max pooling feature. Then these two features are passed to a shared network containing a multi-layer perceptron (shared MLP) and a hidden layer to obtain the channel attention map. After that, the channel attention map is input into the spatial attention part. First, a channel-based global max pooling (Global Max Pooling) and global average pooling (Global Average Pooling) are performed, and then the two obtained features are added together. After a convolution operation and a sigmoid activation function, the spatial attention feature is generated. S3. Design the Feature Pyramid Attention Module Use three convolutional kernels with different sizes of 3×3, 5×5, and 7×7 to extract pyramid features at different scales. This module contains four branches and a main branch. After the convolutional operation, the three branches using different convolutional kernel sizes are upsampled using transposed convolution respectively. The feature maps obtained from each branch are added together in sequence from top to bottom. The main branch introduces a 1×1 convolution, and the original features extracted by the feature extraction module are pixel-multiplied with the pyramid attention features generated by the three branches with different convolutional kernel sizes to obtain multi-scale features. Another branch introduces global pooling and a 1×1 convolution, and then is upsampled to output global features. Finally, the global features are added to the previously obtained multi-scale features to obtain the final output features. S4. Design the loss function for the re-identification task Introduce Circle Loss to replace Triplet Loss. The Circle Loss function is designed as: Wherein: O p = 1 + m, O n = -m, △ p = 1 - m, △ n = m; [·] + represents a non - negative operation, α n is a non - negative integer weight factor for negative samples, α p is a non - negative integer weight factor for positive samples, O n and O p are and 's optimized ideal values, that is, O n = 0 and O p = 1, △ n and △ p are the threshold parameters between classes and within classes respectively. In Circle loss, s n and s p are asymmetric, so s n and s p each require a threshold. The value of the boundary coefficient m is 0.25, allowing to maximize the within - class similarity s n and minimize the between - class similarity s p to learn at different speeds. γ is a regulation coefficient, and its numerical value is 256; The classification loss (ID Loss) and the circle loss are used together to optimize the same feature. The classification loss is calculated using the cross-entropy loss: where y is the true ID label, p i is the predicted ID probability, q i is the positive sample probability, and N is the number of labels; The label form is in the form of one-hot encoding, q i There are only two values, 0 and 1. When y ≠ i, q i = 0; when y = i, q i = 1; When calculating the classification loss, the method of Label Smoothing is used, that is, when i = y, when i ≠ y, ε is a parameter with a value of 0.1; Both the circle loss and the classification loss are used to optimize the entire network structure. The overall loss function is: L sum = L circle + L cross ; Model training and testing phases: The size of the input images is uniformly set to 256×128, the batch size is set to 32, and the Adam optimizer is selected during training with an initial learning rate of 0.

05. Extract the features of the input query image and the vehicle images in the image library, calculate the Euclidean distance between the features to compare the similarity between the query image and the images in the image library, where the calculation formula of the Euclidean distance is d = ||f q - f g ||2, ||·|| is the 2-norm of the corresponding feature vector.

6. A multi-view vehicle re-identification method in a vehicle-road collaborative scenario according to claim 5, characterized in that It also includes: S1. Prepare the dataset for model training and testing Use the open-source simulator CARLA to build a vehicle-road collaborative dual-camera scenario to collect vehicle image data. Specifically: Select Town05 as the experimental map, which includes a section of highway, streets, and intersections. Select two scenarios, the highway and the street, to build the experimental scenario. On the highway and street sections, set a fixed camera with a pitch angle of -45° and a rotation angle of 30° at a position 3 meters above the ground every 10 meters at a 10-meter interval along the roadside. Set a car at the end of the road, and set a fixed camera with a pitch angle of 0° and a rotation angle of 90° on the car roof. Set the vehicle to the autonomous driving mode. The vehicle-mounted camera and the roadside camera collect vehicle images simultaneously. The vehicle-mounted camera mainly captures the front and rear views of the vehicle, and the roadside view mainly obtains the side images of the vehicle. With this experimental scenario, a total of 153 vehicles with different colors and models are collected. Cut out the vehicles in the collected images and label the cut-out vehicle images according to the vehicle ID and camera ID information. The dataset includes 1500 dual-view images with different IDs, of which 1045 are training images and 455 are test images. Randomly flip, randomly erase, and crop the images to enrich the image scenarios, and uniformly adjust the image size to 256×128.

7. A multi-view vehicle re-identification method in a vehicle-road collaborative scenario according to claim 5, characterized in that, The platform for training and testing the described model is a workstation with a GTX1080Ti GPU.

Citation Information

Patent Citations

  • Vehicle re-identification method based on double-branch network feature fusion

    CN112560932A