Ground-air video vehicle re-identification method and device, electronic equipment and storage medium
Through the dual-branch nearest neighbor feature, the network decoupling and fusion of identity and perspective features are enhanced, and the vehicle re-identification problem in the ground-space span perspective is solved, and the recognition accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510237945.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The prior art is difficult to effectively deal with vehicle re-identification at ground-space-span perspectives, especially due to the in-class feature confusion caused by significant visual feature distribution inconsistencies between ground and aerial perspectives.
The dual-branch nearest neighbor feature enhancement network is adopted to decouple and fusion of identity features and viewing features, and the nearest neighbor feature enhancement module PFEM and gradient inversion layer GRL are used to extract and enhance the identity and viewing features of the vehicle respectively to improve the feature representation ability.
Effectively dealing with the problem of inconsistent distribution of visual features under the ground-space-span perspective improves the accuracy and robustness of vehicle re-identification, especially in complex scenarios.
Smart Images

Figure CN120510573A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, electronic device, and storage medium for ground-to-air video vehicle re-identification. Background Art
[0002] Vehicle Re-Identification (VeRi) involves finding the same vehicle in a collection of images captured by other cameras, given a given vehicle image. Therefore, VeRi can be considered a subtask of image retrieval. VeRi technology can accurately track the identity of vehicles using a multi-camera network, enabling automated traffic monitoring and other tasks. Therefore, VeRi holds broad application prospects in the field of intelligent transportation.
[0003] Most vehicle re-identification methods and datasets in related technologies only focus on ground perspective or aerial perspective, and it is difficult to deal with the problem of inconsistent visual feature distribution caused by ground-to-air cross-perspective.
[0004] The above background information does not necessarily constitute prior art. Summary of the Invention
[0005] The purpose of this application is to provide a method, device, electronic device and storage medium for ground-to-air video vehicle re-identification.
[0006] The technical solution of this application is achieved as follows:
[0007] In a first aspect, an embodiment of the present application provides a method for ground-to-air video vehicle re-identification, comprising:
[0008] A dual-branch neighbor feature enhancement network is used to extract vehicle identity and viewpoint features from input images. The network consists of two parallel branches adjacent to the backbone network, one for extracting vehicle identity attributes and the other for capturing differences in viewpoint changes.
[0009] By performing a weighted sum operation on the identity features and the enhanced features obtained through the neighbor feature enhancement module (PFEM), the enhanced features are connected to the gradient reversal layer and the perspective loss to decouple the vehicle's identity information and perspective information.
[0010] In some embodiments, the dual-branch neighbor feature enhancement network is used to extract the identity features and perspective features of the vehicle from the input image, including:
[0011] The identity feature extraction and view feature extraction branches share the same convolutional neural network backbone layer to generate basic features;
[0012] Through the backbone network Backbone, input image X iFirst, it is converted into basic features F;
[0013] The identity feature branch maps the basic features to the identity feature space through convolution or fully connected layers to generate features F that are strongly correlated with the vehicle identity. id ,This feature is used to distinguish different vehicle identities;
[0014] The view feature branch extracts the feature F related to the vehicle view view .
[0015] In some embodiments, the identity feature and the enhanced feature after PFEM are weighted summed, and the enhanced feature is connected to the gradient reversal layer and the perspective loss to decouple the vehicle's identity information and perspective information, including:
[0016] The neighbor feature enhancement module PFEM utilizes similar feature information within the same batch of processed data to fuse identity and view features. Based on a K-NN search process, this module retrieves the K neighbors closest to the current sample from the feature library and uses these neighbor features for feature enhancement, where K is a positive integer.
[0017] In some embodiments, for each input ID feature, Retrieve K nearest neighbor features from The process is expressed as:
[0018]
[0019] K-NN search uses metrics such as Euclidean distance or cosine similarity to find the K neighbor features closest to the current sample feature from the identity feature database.
[0020] In some embodiments, for the viewing angle feature Retrieve K nearest neighbor features from the database
[0021] After retrieving the nearest neighbor features of identity and viewpoint, PFEM first performs an intersection operation Extract the common parts between identity feature neighbors and view feature neighbors:
[0022]
[0023] The intersection operation extracts the common features between identity and view features, which represent the feature parts that have both identity discrimination and view information.
[0024] In some embodiments, feature aggregation is performed Compare these common characteristics to the original identity characteristics Fusion is performed to generate enhanced feature representations:
[0025]
[0026] Among them, the aggregation operation A weighted average is used.
[0027] In some embodiments, the identity recognition task is optimized jointly by the cross-entropy loss and the triplet loss, and the view prediction task is optimized by the cross-entropy loss. The cross-entropy loss, triplet loss, and cross-entropy loss for the identity recognition task and the view prediction task are expressed as follows:
[0028]
[0029] in, and represent the identity probability distribution and view probability distribution predicted by the model, respectively. and represents the real identity label and viewpoint label, is the positive sample feature, is the negative sample feature, d(·, ·) is the distance function between features, and m is a hyperparameter in the triple loss, which is used to control the minimum distance difference between positive and negative samples. The overall optimization goal of the model is:
[0030]
[0031] Where λ is a hyperparameter that balances the weights of the gradient reversal layer and is used to control the strength of the gradient reversal.
[0032] In a second aspect, the present application provides a ground-to-air video vehicle re-identification device, comprising:
[0033] A module that uses a dual-branch neighbor feature enhancement network to extract vehicle identity and viewpoint features from input images. The network consists of two parallel branches adjacent to the backbone network: one branch extracts vehicle identity attributes, and the other captures differences in viewpoint changes.
[0034] By performing a weighted sum operation on the identity features and the enhanced features after the neighbor feature enhancement module PFEM, the enhanced features are connected to the gradient reversal layer and perspective loss to decouple the vehicle's identity information and perspective information.
[0035] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory; when the processor executes the running program stored in the memory, the method described in any embodiment of the present application is implemented.
[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of the present application.
[0037] The vehicle re-identification method for ground-to-air video provided in this application addresses the intra-class feature confusion problem caused by perspective changes through feature enhancement and decoupling of identity and perspective features. A dual-branch neighbor feature enhancement network is adopted to extract identity and perspective features respectively, improving the feature representation capability. A neighbor feature enhancement module PFEM is introduced to fuse similar neighbor features, capture common features under different perspectives to enhance the original features, and achieve feature decoupling that is invariant to perspective. This method can better handle the problem of inconsistent visual feature distribution caused by ground-to-air cross-perspective. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A view of a related art ground-to-air camera is shown.
[0039] Figure 2 A flowchart of a method for ground-to-air video vehicle re-identification according to an embodiment of the present application is shown.
[0040] Figure 3 The figure shows the overall model diagram of the dual-branch neighbor feature enhancement network provided by one embodiment of the present application.
[0041] Figure 4 Visualization of the output feature maps of the three methods, Baseline, AdaSp, and DPFE, is shown.
[0042] Figure 5 A structural block diagram of an electronic device provided by an embodiment of the present application is shown.
[0043] Figure 6 A schematic diagram of a computer-readable storage medium provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0044] In order to further illustrate the features and technical content of the embodiments of this application, the following will further illustrate the technical solutions of this application in conjunction with the accompanying drawings and specific embodiments. The drawings are for reference only and are not intended to limit the scope of the embodiments of this application.
[0045] Unless otherwise defined, all technical and scientific terms used in the examples of this application shall have the same meaning as commonly understood by those skilled in the art. The use of terms is for the purpose of describing the examples of this application only and is not intended to limit this application in any way.
[0046] In the following description, the “some embodiments” mentioned represent a subset of all possible embodiments, but it should be understood that “some embodiments” can be the same or different subsets of all possible embodiments, and can be combined with each other without conflict. It should also be pointed out that if terms such as “first, second, third” appear in the embodiments of the present application, then these terms are only used to distinguish the differences between similar objects and do not represent a specific sort or order. Therefore, terms such as “first, second, third” can be used interchangeably where permitted, and the embodiments of the present application do not have to be executed strictly in the order described.
[0047] In recent years, with the rapid development of deep learning, traditional vehicle re-identification technology has achieved remarkable results. However, with the increasing intelligence of cities, real-world surveillance systems can appear in every corner, which is far from the traditional vehicle re-identification. There are two reasons for this: (1) Almost all vehicle re-identification datasets in related technologies come from ground surveillance cameras or aerial surveillance cameras alone, such as VeRi-776, VehicleID, VERI-Wild, CompCars, etc., which are only from the ground perspective, or VRAI, UAV-ReID, etc., which are only from the aerial perspective. This camera system emphasizes that the camera is kept at the same height and captures the same scene or object from different angles. In the real world, due to the vigorous development of drones in recent years, aerial surveillance systems have begun to be active in the public eye and form a linkage with traditional surveillance systems. This ground-air collaborative monitoring system combines the wide-area dynamic monitoring capability of aerial drones and the fixed-point high-precision monitoring capability of traditional ground cameras. It can cover a wider monitoring range and effectively improve the overall monitoring accuracy and real-time performance. Therefore, facing multi-view and multi-height monitoring scenes, ground-air cross-view vehicle re-identification poses some new challenges in matching. (2) Vehicle re-identification methods in related technologies are almost all trained using images from the same type of camera. When querying and matching, the images of vehicles taken from different angles by the same type of camera are matched, such as ground-ground perspective matching and air-air perspective matching. However, there is little research on ground-air bidirectional perspective vehicle matching. Figure 1 Vehicle images from different ground cameras and aerial cameras are displayed. Aerial cameras are aerial cameras, and ground cameras are ground cameras. Due to the significant view difference between the two perspectives, the discriminative identity features are destroyed. Coupled with factors such as different lighting conditions and environmental changes at different time periods, ground-to-air vehicle re-identification technology has become a more challenging task.
[0048] Vehicle re-identification from ground surveillance perspectives is a key area of current vehicle re-identification research. This task typically involves capturing vehicle images from various positions and angles using fixed ground surveillance cameras. Related methods are often based on ground surveillance perspective datasets, focusing primarily on the cross-angle variations of vehicles captured by ground surveillance cameras. This makes them difficult to apply to matching scenarios across ground and air perspectives.
[0049] The vehicle re-identification task under ground-air collaborative monitoring perspectives has significant differences in the appearance of the vehicle under different perspectives due to the large range of perspective changes, making traditional vehicle re-identification methods difficult to apply in this scenario. The ground perspective usually provides details from a horizontal angle, while the aerial perspective is mainly a bird's-eye view, resulting in significant differences in the visual features of the same vehicle under ground and aerial perspectives. Most vehicle re-identification methods and datasets in related technologies focus only on ground or aerial perspectives, and have difficulty dealing with the problem of inconsistent visual feature distribution caused by ground-air cross-perspective re-identification. In terms of vehicle re-identification, research on ground-air collaborative cross-perspective re-identification tasks is almost zero.
[0050] To fill the technological gap in ground-air collaborative cross-view re-ID research, this application example constructs the first ground-air cross-view vehicle re-ID dataset, GAVRI. This dataset supports the development of novel recognition technologies for vehicles under different heights of surveillance. This provides a challenging testbed for ground-air cross-view re-ID and a reliable experimental foundation for future vehicle re-ID research, thereby promoting further development in this field.
[0051] Feature decoupling is widely used in deep learning and computer vision. Its core purpose is to separate different features in the data, thereby improving model performance in multi-task learning. In domain adaptation tasks, feature decoupling is used to address the problem of inconsistent data distribution between the source and target domains. By separating domain features from task features, it enhances the model's generalization ability. This method has shown significant performance in multi-task learning, effectively preventing information confusion between tasks and improving feature learning performance for each task.
[0052] Vehicle re-identification (Re-ID) is crucial for urban surveillance and smart transportation. However, most research focuses on matching from ground viewpoints, while matching from ground to aerial viewpoints is relatively limited. This is primarily due to two reasons: first, a lack of available ground-to-air viewpoint datasets; second, the significant differences between ground and aerial viewpoints make traditional matching methods ineffective. To address this challenge, a dual-branch proximal feature enhancement network (DPFE) is proposed. DPFE addresses the intra-class feature confusion caused by viewpoint variations by enhancing features and decoupling identity and viewpoint features. DPFE employs a dual-branch architecture to extract identity and viewpoint features separately, improving feature representation. Furthermore, a proximal feature enhancement module (PFEM) is introduced, which utilizes a k-nearest neighbor algorithm to fuse similar neighbor features, capturing common features across different viewpoints and enhancing the original features. With the assistance of a gradient reversal layer (GRL), view-invariant feature decoupling is achieved. To validate the effectiveness of DPFE, a new dataset (GAVRI) containing 2,159 vehicle IDs and 17,437 images from various aerial and ground cameras is presented. Experimental results show that DPFE achieves significant performance improvement on GAVRI, verifying its robustness and superiority in complex scenarios.
[0053] In order to be more in line with the real world, especially vehicle re-identification under the ground-air camera network, a network DPFE (Dual-branch Proximal Feature Enhancement) based on dual-branch proximal feature enhancement is proposed in the embodiment of the present application to solve the view differences caused by extreme perspective changes. DPFE aims to separate perspective-related features from perspective-independent features, so as to better learn discriminative identity features from those perspective-independent features. Its workflow is roughly as follows: the input image passes through the backbone network to extract the identity features and perspective features of the vehicle. The identity features are used for vehicle identity recognition, and the perspective features capture the perspective information of the vehicle and the changes in appearance at different angles. In the proximal feature enhancement module PFEM (Proximal Feature Enhancement Module), the model uses two features to perform K-nearest neighbor searches in the feature database respectively. The nearest neighbor features of the identity feature query are strongly correlated with the current vehicle identity, while the perspective features can also retrieve vehicle images of the same identity due to their potential identity discrimination information. The intersection features of the two query results represent the features that play a key role in both identity and perspective prediction. To further decouple identity features and viewpoint features, a gradient reversal layer (GRL) is introduced into the model to eliminate the interference of viewpoint information in the intersection features on identity classification, enabling the model to learn more discriminative identity features.
[0054] In order to verify the effectiveness of the proposed dual-branch neighbor feature enhancement network, a new ground-air vehicle re-identification dataset GAVRI (Ground Areal Vehicle Re-identification) was constructed. The dataset covers vehicle images from multiple perspectives, involving a total of 2,159 vehicle identities and 17,437 images, including vehicle images taken by different ground and aerial surveillance cameras under different lighting and weather conditions. This diversity helps to improve the robustness of the model in different real-world scenarios. In addition, in the evaluation setting, the GAVRI dataset does not adopt a single matching mode (such as ground-air or air-ground), but adopts a two-way mixed matching, that is, the query set and the gallery set both contain vehicle images from ground and aerial perspectives. This design provides a richer and more challenging test environment, which is more in line with the problem of multi-perspective and cross-environment vehicle re-identification in the real world. The method of the embodiment of the present application proposes a dual-branch feature extraction and fusion network, which decouples identity and perspective features through the PFEM module and GRL, and then learns more discriminative identity features, thereby improving the recognition ability of the model.
[0055] In the embodiments of this application, a vehicle re-identification dataset is used. The dataset contains 2,159 vehicle identities and 17,437 cross-view vehicle images, covering different environmental conditions and shooting angles, to help solve the challenges in the vehicle re-identification task.
[0056] A systematic experimental evaluation was conducted on two ground-air datasets, GAVRI and AG-ReID. Compared with other state-of-the-art methods, the methods proposed in the embodiments of this application have significantly improved performance.
[0057] In the method of the present embodiment, the DFPE employed utilizes a gradient reversal layer for feature decoupling to further separate identity and viewpoint features. By introducing an adversarial learning mechanism, identity feature learning and viewpoint feature classification are pitted against each other, ultimately suppressing the interference of viewpoint information and promoting the model to learn more discriminative identity features. This strategy is particularly important in vehicle re-identification tasks across ground and air viewpoints, helping the model extract more robust identity features under extreme viewpoint variations.
[0058] In this application, a model architecture for re-identifying vehicles from both ground and aerial perspectives is proposed. This architecture aims to improve re-identification performance by decoupling identity features from viewpoint features and then enhancing and fusing neighbor features. The following describes the proposed model, focusing on its overall architecture, dual-branch feature extraction and decoupling, the neighbor feature enhancement module (PFEM), and the loss function.
[0059] refer to Figure 2 As shown, an embodiment of the present application provides a ground-to-air video vehicle re-identification method, which may include:
[0060] S10. Using a dual-branch neighbor feature enhancement network to extract vehicle identity features and viewpoint features from the input image; the network includes two parallel branches adjacent to the backbone network, one branch for extracting vehicle identity attributes and the other branch for capturing differences in viewpoint changes;
[0061] S20. Perform a weighted sum operation on the identity feature and the enhanced feature after the neighbor feature enhancement module PFEM, and connect the enhanced feature to the gradient reversal layer GRL and the perspective loss to decouple the vehicle's identity information and perspective information.
[0062] Exemplarily, the method of extracting the identity features and the perspective features of the vehicle from the input image using a dual-branch neighbor feature enhancement network may include:
[0063] The identity feature extraction and view feature extraction branches share the same convolutional neural network backbone layer to generate basic features;
[0064] Through the backbone network Backbone, input image X i First, it is converted into basic features F;
[0065] The identity feature branch maps the basic features to the identity feature space through convolution or fully connected layers to generate features F that are strongly correlated with the vehicle identity. id ,This feature is used to distinguish different vehicle identities;
[0066] The view feature branch extracts the feature F related to the vehicle view view .
[0067] Exemplarily, the weighted sum operation is performed on the identity feature and the enhanced feature after PFEM, the enhanced feature is integrated with GRL and perspective loss, and the identity information and perspective information of the vehicle are decoupled, including:
[0068] The neighbor feature enhancement module PFEM utilizes similar feature information within the same batch of processed data to fuse identity and view features. Based on a K-NN search process, this module retrieves the K neighbors closest to the current sample from the feature library and uses these neighbor features for feature enhancement, where K is a positive integer.
[0069] For example, for each input ID feature, in the feature database Retrieve K nearest neighbor features from The process is expressed as:
[0070]
[0071] K-NN search uses metrics such as Euclidean distance or cosine similarity to find the K neighbor features closest to the current sample feature from the identity feature database.
[0072] For example, for the viewing angle feature Retrieve K nearest neighbor features from the database
[0073] After retrieving the nearest neighbor features of identity and viewpoint, PFEM first performs an intersection operation Extract the common parts between identity feature neighbors and view feature neighbors:
[0074]
[0075] The intersection operation extracts the common features between identity and view features, which represent the feature parts that have both identity discrimination and view information.
[0076] For example, through feature aggregation operation Compare these common features to the original identity features Fusion is performed to generate enhanced feature representations:
[0077]
[0078] Among them, the aggregation operation A weighted average is used.
[0079] For example, the identity recognition task is optimized by the cross entropy loss and the triplet loss, and the view prediction task is optimized by the cross entropy loss. The cross entropy loss, triplet loss, and cross entropy loss for the identity recognition task and the view prediction task are expressed as follows:
[0080]
[0081] in, and represent the identity probability distribution and view probability distribution predicted by the model, respectively. and represents the real identity label and viewpoint label, is the positive sample feature, is the negative sample feature, d(·, ·) is the distance function between features, and m is a hyperparameter in the triple loss, which is used to control the minimum distance difference between positive and negative samples. The overall optimization goal of the model is:
[0082]
[0083] Where λ is a hyperparameter that balances the weights of the gradient reversal layer and is used to control the strength of the gradient reversal.
[0084] The ground-to-air video vehicle re-identification method proposed in the embodiment of the present application is a ground-to-air video vehicle re-identification method based on dual-branch nearest neighbor feature enhancement.
[0085] This embodiment proposes an image dual-branch neighbor feature enhancement network based on the K-NN search algorithm. The overall network architecture is as follows: Figure 3 As shown. Given a set of input images X, after preprocessing, a standard convolutional neural network backbone is used to extract the basic features F. The architecture of the model mainly contains two branches: one for extracting identity features F id , and the other one is used to extract the view feature F view. These two branches each extract features through a specific convolutional layer or other deep learning module, and are sent to subsequent modules respectively. In the identity feature branch, the extracted features are used to predict the identity of the vehicle and are optimized through identity loss and triplet loss. In the view feature branch, the extracted features are used to predict the view of the vehicle. In the early stage of model training, in order to train the view classifier and the mapping layer that maps the backbone features to the view feature space, the view features are directly used for supervised learning of the view loss. After training, the parameters of the view classifier and the mapping layer are fixed, and the view features are input into the neighbor feature enhancement module. PFEM inputs the identity and view features into the feature database, and obtains the top-k similar feature list through the K-NN search algorithm. These similar features contain the same identity under different view angles. The features selected from the top-k list are further fused through intersection operations and mean fusion to generate more representative enhanced features.
[0086] The extracted identity features are further integrated with the enhanced features to enhance the feature representation capability. At the same time, the enhanced features decouple identity features from viewpoint features through GRL, eliminating the influence of viewpoint on the identity prediction task and learning more discriminative identity features to improve the overall recognition accuracy of the model.
[0087] Dual-branch feature extraction and fusion:
[0088] The model incorporates two parallel branches adjacent to the backbone network, designed to deeply mine and separate vehicle identity and viewpoint features from the input image. These two branches each perform distinct functions: one focuses on extracting unique attributes related to vehicle identity, ensuring accurate recognition of different vehicles; the other is dedicated to capturing differences in viewpoint variations to account for the diverse poses of vehicles in the image. Furthermore, a weighted summation operation is performed on the identity features and the enhanced features obtained through PFEM. The enhanced features are then coupled with GRL and viewpoint loss, effectively decoupling identity and viewpoint information. Gradient reversal renders the model unable to distinguish viewpoint information, forcing the model to learn more robust identity features during backpropagation.
[0089] Dual-branch feature extraction and decoupling:
[0090] The identity feature extraction and view feature extraction branches share the same convolutional neural network backbone layer to generate basic features. i First, it is converted into basic features F.
[0091] F=Backbone(X i )#(1)
[0092] On this basis, the identity feature branch maps the basic feature F to the identity feature space through the convolution or full connection layer to generate the feature F that is strongly related to the vehicle identity. id This feature is used to distinguish different vehicle identities and is optimized through cross entropy loss and triplet loss. At the same time, the view feature branch extracts the feature F related to the vehicle view view . Due to the significant difference between the ground perspective and the aerial perspective, the basic model cannot effectively prevent the impact of these view differences on the identity representation when dealing with them, resulting in excessive intra-class feature differences that cannot be narrowed. Specifically, when learning identity features, the model may mistakenly mix view features into identity features, resulting in a coupling relationship between view features and identity features, thereby hindering the effective extraction of identity features. To solve this problem, a gradient reversal layer is introduced in the view feature extraction process. By reversing the gradient, the identity feature branch can "ignore" view changes during learning, thereby avoiding the interference of view information on identity feature learning. Through this mechanism, the identity feature branch can focus on learning more discriminative identity features without being interfered with by view features. Specifically, the introduction of GRL makes the optimization objectives of identity features and view features independent of each other during training, which helps the model to more accurately separate and extract identity features and view features.
[0093] Neighbor feature enhancement module
[0094] In order to further enhance the feature representation, a Proximal Feature Enhancement Module (PFEM) is proposed. This module uses similar feature information within the same batch of processed data to fuse identity and view features. Based on a K-NN search process, this module retrieves the K neighbors closest to the current sample from the feature library and uses these neighbor features for feature enhancement. PFEM is as follows: Figure 3 As shown in .
[0095] Specifically, for each input ID feature, in the feature database Retrieve K nearest neighbor features from The process is expressed as:
[0096]
[0097] K-NN search uses metrics such as Euclidean distance or cosine similarity to find the K neighbor features closest to the current sample feature from the identity feature database. Similarly, for the view feature Retrieve K nearest neighbor features from the database
[0098] After retrieving the nearest neighbor features of identity and viewpoint, PFEM first performs an intersection operation Extract the common parts between identity feature neighbors and view feature neighbors:
[0099]
[0100] The intersection operation extracts the common features between identity and view features, which represent the feature parts that have both identity discrimination and view information. Next, through the feature aggregation operation Compare these common features to the original identity features Fusion is performed to generate enhanced feature representations:
[0101]
[0102] Among them, the aggregation operation Weighted averaging is used, and other aggregation methods are described in Section 5.4 Ablation Experiment.
[0103] Loss Function
[0104] The identity recognition task is optimized by the cross entropy loss and the triplet loss, and the view prediction task is optimized by the cross entropy loss. The above three losses can be expressed as follows:
[0105]
[0106] in, and represent the identity probability distribution and view probability distribution predicted by the model, respectively. and represents the real identity label and viewpoint label, is the positive sample feature, is the negative sample feature, d(·,·) is the distance function between features, and m is a hyperparameter in the triplet loss, which is used to control the minimum distance difference between positive and negative samples. In summary, the overall optimization goal of the model is as follows:
[0107]
[0108] Where λ is a hyperparameter that balances the weights of the gradient reversal layer and is used to control the strength of the gradient reversal.
[0109] Dataset
[0110] As shown in Table 1, related ground-based vehicle re-identification datasets (e.g., VeRi-776 and VehicleID) provide comprehensive images from ground-based camera perspectives, but do not include images captured from drones. Related aerial drone-based vehicle re-identification datasets only include images captured from drone perspectives, lacking detailed information about vehicles in planar environments. To provide a publicly available benchmark, we combined drones and ground-based cameras to construct the GAVRI (Ground Areal Vehicle Re-identification) dataset, which better reflects real-world scenarios.
[0111] Table 1 Comparison of public vehicle re-ID datasets in related technologies
[0112]
[0113] The constructed dataset is derived from six drone and ground-based camera videos. The ground-based videos are framed every 32 frames, while the aerial videos are framed every 16 frames. After frame extraction, five lanes were selected for annotation using the labelme annotation software. The cropped vehicle images from the drone camera range in size from a maximum of 668x647 pixels to a minimum of 56x70 pixels. The cropped vehicle images from the ground-based cameras range in size from a maximum of 719x821 pixels to a minimum of 58x47 pixels. GAVRI was collected under various weather conditions and contains significant illumination variations, motion blur, and minor occlusions. The final dataset includes 17,437 images and 2,159 unique vehicles. Due to the differences between drone and ground-based views, the constructed dataset is challenging, reflecting the limitations of real-world re-identification tasks and increasing the difficulty compared to traditional multi-camera aerial or ground-based perspectives.
[0114] During the evaluation phase, a total of 9,610 images containing 1,056 vehicle identities were collected, using a ground-air dual-perspective hybrid query. This means that both the query set and the gallery set include images from both the ground and aerial perspectives, rather than a single ground-ground or air-air match. Unlike individual pedestrian matches, many different vehicles may appear very similar from certain angles, appearing to be the same vehicle. Due to the dual-perspective hybrid query, some vehicles in the gallery set have different IDs and perspectives, yet they share a high degree of visual similarity with the query image. This further increases the difficulty of distinguishing similar vehicles in the gallery. This requires the model to fully learn both the overall characteristics of the vehicle and its local fine-grained features.
[0115] experiment
[0116] Datasets and evaluation metrics
[0117] The proposed ground-to-air cross-view re-identification method is verified on the ground-to-air pedestrian re-identification dataset, and a standard data partitioning strategy is adopted for the AG-ReID dataset.
[0118] The AG-ReID dataset is derived from real-world outdoor environments and contains 21,983 images from two different cameras, including 388 pedestrian identities. The GAVRI dataset contains 17,437 images from four different cameras, including 2,159 vehicles. The cumulative match characteristic (CMC) and mean average precision (mAP) are used as evaluation metrics.
[0119] In all experiments, a ResNet-50 pre-trained on ImageNet was used as the backbone. To facilitate the integration of custom modules and enrich multi-scale information, the final downsampling of the ResNet-50 was removed. The output feature maps of the backbone were combined using global average pooling and global max pooling, and the two were concatenated and normalized using BatchNorm. For the GAVRI vehicle dataset, the input image size was set to 256×256, and the batch size was set to 128; for the AG-ReID pedestrian dataset, the input image size was set to 256×128, and the batch size was set to 64. A random region center cropping method was designed to randomly extract important local features from the center of the image, combined with random horizontal flipping and color jittering for data augmentation. The model was trained for 150 epochs with an initial learning rate of 3.5e-4, which was divided by 10 every 50 epochs.
[0120] All experiments were written in Pytorch and conducted on a server equipped with an Intel(R) Xeon(R) Gold 5118 CPU @ 2.30 GHz and an NVIDIA Tesla A40.
[0121] The proposed image dual-branch neighbor feature enhancement network DPFE is compared with several advanced ReID methods in Table 2 on GAVRI. As shown in Table 2, DPFE is compared with the current state-of-the-art methods.
[0122] Table 2 Performance comparison of DPFE and other advanced algorithms on the GAVRI dataset
[0123]
[0124] The method proposed in this paper achieves a mAP score of 65.0% and a Rank-1 accuracy of 65.6%, which outperforms other traditional methods, exceeds the baseline method by 6.6% and 7.3%, and exceeds CLIP-ReID by 4.8% and 6.5%. In addition, compared with the method that proposes a novel loss function, it exceeds AdaSP by 2.7% and 1.4%, achieving significant performance improvement. Table 3 shows the performance comparison of the AG-ReID dataset with mainstream methods under two settings. "Air→Ground" and "Ground→Air" represent two specific evaluation modes. The performance in Table 3 is derived from the results in related technologies. DPFE outperforms other methods on AG-ReID, including the Explain interpretable re-identification method proposed in related technologies. This strongly demonstrates the importance of DPFE in reducing the damage of viewpoint differences to identity discrimination features. The results show that the method of the embodiment of the present application achieved excellent performance on both GAVRI and AG-ReID datasets. Compared with traditional vehicle re-identification methods, DPFE can capture more discriminative fine-grained local information features, and by decoupling identity information from view information, it can obtain more features related to identity prediction, significantly improving the accuracy of identity prediction.
[0125] Table 3 Comparison of DPFE and existing algorithms on the AG-ReID dataset
[0126]
[0127] To validate the effectiveness of the proposed DPFE method and its various components, ablation experiments were conducted on the GAVRI dataset. Table 4 shows the ablation results of DPFE on the GAVRI dataset, comparing the Rank-1, 5, 10, and mAP values of the DPFE method with dual-branch feature extraction, a neighbor feature enhancement module, and a gradient reversal layer. The results show that the dual-branch approach to extracting identity and viewpoint features allows independent learning of the different features, avoiding confusion between the two. Without the PFEM module, the model relies primarily on global features for recognition, which can be difficult to distinguish in complex scenes (such as partial occlusion or vehicles with similar details but different identities), resulting in poor performance. When the model lacks GRL, identity and viewpoint features cannot be further decoupled, limiting generalization between ground and aerial views. The model is unable to learn more features related to vehicle identity, resulting in poor performance in cross-view scenarios. Furthermore, two different feature fusion methods are compared, combining features enhanced with PFEM with identity features: DPFE with a weighted summation of learnable parameters and a simple feature average summation (DPFE_(avg)).
[0128] Table 4 Ablation experiments of DPFE on the GAVRI dataset
[0129]
[0130] From the experimental results, it can be seen that the DPFE model using weighted summation of learnable parameters outperforms the average summation DPFE_(avg) in all evaluation indicators, which shows that the learnable weights play an important role in improving the effectiveness of feature fusion.
[0131] Parameter Analysis: To further validate the effectiveness of the PFEM module in this paper, we conducted experiments on the number of neighbor feature hyperparameters K and λ during the gradient reversal process in PFEM to analyze how they affect the model performance:
[0132] (1) Five values of K were tested, namely 3, 5, 10, 15, and 20. The experimental results are shown in Table 5. When the K value is 5, the overall performance of the model is most significantly improved. As the K value increases, the overall performance decreases. This phenomenon shows that the quality of the neighbor features is positively correlated with the quality of the enhanced features. When the K value is too small, the further screening of the intersection operation may result in the failure to query valid neighbor features and the inability to enhance the original features. When the K value is larger, the probability that the neighbor features queried contain negative samples will increase, and the influence of noise generated by other negative samples will also increase, resulting in performance degradation.
[0133] Table 5. Effect of hyperparameter K on model performance
[0134]
[0135] (2) During the gradient reversal layer, the hyperparameter λ controls the intensity of the gradient reversal. Specifically, λ determines the scaling factor when reversing the gradient, which affects the decoupling effect of the view feature and the identity feature. The hyperparameter K in PFEM is set to 5, and several values of λ are experimentally evaluated, namely 0.05, 0.1, 0.3, and 0.5. The experimental results are shown in Table 6. When λ = 0.1, the model performance reaches the best. The reason is that when the λ value is very small, the gradient reversal effect is weak, the decoupling degree between the identity feature and the view feature is insufficient, and the identity feature is still mixed with too much information unrelated to the identity, which affects the overall discrimination ability of the model. When λ is too large, the view information will be excessively weakened, which may cause the model to lose some fine-grained information related to identity distinction, resulting in performance degradation.
[0136] Visualization: Visual analysis of feature maps of different methods on the data set, Figure 4The Grad-CAM tool was used to visualize the second convolutional layer of the third bottleneck in the fourth stage of the model, demonstrating the performance of the baseline method, AdaSp, and DPFE. Heatmaps provide a visual indicator of the distribution of attention across different vehicle components. A comparison reveals that the baseline method's feature maps are primarily focused on localized areas of the vehicle, such as the headlights or license plate, demonstrating its limited effectiveness in handling complex backgrounds or occlusions, resulting in a relatively restricted area of attention. Furthermore, the AdaSp method significantly improves upon the baseline method, with its feature maps capturing a wider range of important vehicle regions, such as the roof and rear. However, there is still some localized feature concentration, failing to fully cover the entire vehicle. Compared to the previous two methods, the DPFE feature maps exhibit a more uniform feature distribution. With the help of PFEM, the representation of local features is further enhanced. Combined with feature decoupling, the feature maps highlight more information useful for vehicle identity prediction. The model provides more comprehensive coverage of key vehicle components, especially against complex backgrounds. It effectively captures global information, avoiding over-concentration of features in a small area, demonstrating a stronger ability to represent vehicle appearance. Consequently, DPFE achieves higher accuracy.
[0137] Table 6 Effect of hyperparameter λ on model performance
[0138]
[0139] Eliminating the discrepancy between ground and drone perspectives is a key factor in improving the performance of ground-to-air vehicle re-identification. DPFE reduces the model's sensitivity to viewpoint, decoupling identity features from viewpoint features, forcing the model to learn more robust identity features. By learning the correlations and differences between different viewpoints, it provides more accurate query results under complex and extreme viewpoint variations. Experiments on the newly constructed GAVRI dataset and other related datasets demonstrate that DPFE achieves significant performance improvements over baseline methods and outperforms state-of-the-art vehicle re-identification methods on the GAVRI dataset.
[0140] Another embodiment of the present application provides a ground-to-air video vehicle re-identification device, which may include:
[0141] A module that uses a dual-branch neighbor feature enhancement network to extract vehicle identity and viewpoint features from input images. The network consists of two parallel branches adjacent to the backbone network: one branch extracts vehicle identity attributes, and the other captures differences in viewpoint changes.
[0142] By performing a weighted sum operation on the identity features and the enhanced features after the neighbor feature enhancement module PFEM, the enhanced features are connected to the GRL and perspective loss to decouple the vehicle's identity information and perspective information.
[0143] Exemplarily, the dual-branch neighbor feature enhancement network is used to extract the identity features and perspective features of the vehicle from the input image, including:
[0144] The identity feature extraction and view feature extraction branches share the same convolutional neural network backbone layer to generate basic features;
[0145] Through the backbone network Backbone, input image X i First, it is converted into basic features F;
[0146] The identity feature branch maps the basic features to the identity feature space through convolution or fully connected layers to generate features F that are strongly correlated with the vehicle identity. id ,This feature is used to distinguish different vehicle identities;
[0147] The view feature branch extracts the feature F related to the vehicle view view .
[0148] Exemplarily, the weighted sum operation is performed on the identity feature and the enhanced feature after PFEM, the enhanced feature is integrated with GRL and perspective loss, and the identity information and perspective information of the vehicle are decoupled, including:
[0149] The neighbor feature enhancement module PFEM utilizes similar feature information within the same batch of processed data to fuse identity and view features. Based on a K-NN search process, this module retrieves the K neighbors closest to the current sample from the feature library and uses these neighbor features for feature enhancement, where K is a positive integer.
[0150] For example, for each input ID feature, in the feature database Retrieve K nearest neighbor features from The process is expressed as:
[0151]
[0152] K-NN search uses metrics such as Euclidean distance or cosine similarity to find the K neighbor features closest to the current sample feature from the identity feature database.
[0153] For example, for the viewing angle feature Retrieve K nearest neighbor features from the database
[0154] After retrieving the nearest neighbor features of identity and viewpoint, PFEM first performs an intersection operation Extract the common parts between identity feature neighbors and view feature neighbors:
[0155]
[0156] The intersection operation extracts the common features between identity and view features, which represent the feature parts that have both identity discrimination and view information.
[0157] For example, through feature aggregation operation Compare these common features to the original identity features Fusion is performed to generate enhanced feature representations:
[0158]
[0159] Among them, the aggregation operation A weighted average is used.
[0160] For example, the identity recognition task is optimized by the cross entropy loss and the triplet loss, and the view prediction task is optimized by the cross entropy loss. The cross entropy loss, triplet loss, and cross entropy loss for the identity recognition task and the view prediction task are expressed as follows:
[0161]
[0162] in, and represent the identity probability distribution and view probability distribution predicted by the model, respectively. and represents the real identity label and viewpoint label, is the positive sample feature, is the negative sample feature, d(·, ·) is the distance function between features, and m is a hyperparameter in the triple loss, which is used to control the minimum distance difference between positive and negative samples. The overall optimization goal of the model is:
[0163]
[0164] Where λ is a hyperparameter that balances the weights of the gradient reversal layer and is used to control the strength of the gradient reversal.
[0165] Another embodiment of the present application provides an electronic device, including a processor and a memory; when the processor executes the running program stored in the memory, the method described in any embodiment of the present application is implemented.
[0166] For example, Figure 5 As shown, Figure 5The following is a schematic diagram of the structure of an electronic device provided as a specific example, including a processor 601, a communication interface 602, a memory 603, and a communication bus 604. The processor 601, the communication interface 602, and the memory 603 communicate with each other via the communication bus 604. The memory 603 is used to store computer programs; the processor 601 is used to execute the programs stored in the memory 603, thereby implementing the methods provided in any embodiment of the present application.
[0167] In the electronic device described above, the communication bus may be a peripheral component interconnect standard bus or an extended industry standard architecture bus. The communication bus may include an address bus, a data bus, and a control bus. The communication interface is used to enable data exchange between the electronic device and other devices.
[0168] Memory 603 may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, memory 603 may include other storage devices located remotely from processor 601. Processor 601 may be a general-purpose processor, such as a central processing unit (CPU) or a network processor (NP); it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other types of programmable logic devices, discrete gate circuits or transistor logic devices, or discrete hardware components.
[0169] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in any embodiment of the present application when the computer program is executed by a processor. Figure 6 As shown, Figure 6 The computer-readable storage medium shown is an optical disc 20 , on which a computer program (ie, a program product) is stored. When the computer program is executed by a processor, the method described in any embodiment of the present application can be implemented.
[0170] The computer-readable storage medium of this embodiment can be any type of medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates multiple media. Common storage media types include magnetic media (such as floppy disks, hard disks, and magnetic tapes), optical media (such as DVDs), or semiconductor media (such as solid-state drives (SSDs)).
[0171] In the above embodiments, all or part of the functions can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the functions can be implemented by one or more computer instructions contained in a computer-readable storage medium. When a computer loads and executes these computer program instructions, it can partially or completely perform the processes or functions described in the embodiments of the present invention.
[0172] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions may be stored in a computer-readable storage medium and may be transmitted from one computer-readable storage medium to another via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. For example, computer instructions may be transmitted from a website, computer, server, or data center to another computer, website, server, or data center.
[0173] It should be noted that:
[0174] In the embodiments of the present application, the terms "include", "comprising" and any other variations thereof are intended to indicate a non-exclusive inclusion relationship, meaning that when referring to including certain elements, the process, method, article or apparatus includes not only these elements, but also other elements not explicitly listed, or inherent elements related to the process, method, article or apparatus. Without further qualification, the use of "comprising a..." to express an element does not exclude the possibility that other identical elements may exist in the process, method, article or apparatus that includes the element.
[0175] By the description of the above-mentioned embodiment method, it will be clearly understood by those skilled in the art that these methods can be implemented by software plus necessary general hardware platforms. Of course, it can also be implemented by pure hardware, but in many cases, the former is generally a more optimal implementation method. Based on this understanding, the technical solutions of the embodiments of the present application, or the technical contributions made by the present application, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk, etc.), including several instructions, so that an image display device (such as a mobile phone, a computer, a server, an air-conditioning device or a network device, etc.) executes the various methods described in the embodiments of the present application.
[0176] The above description is a specific implementation of the embodiments of the present application, but does not limit the scope of protection of the present application. Any changes or alternatives that can be conceived by any person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be determined based on the content of the claims.
Claims
1. A method for vehicle re-identification from ground-to-air video, characterized in that: include: A dual-branch neighbor feature enhancement network is used to extract vehicle identity and viewpoint features from input images. The network consists of two parallel branches adjacent to the backbone network, one for extracting vehicle identity attributes and the other for capturing differences in viewpoint changes. By performing a weighted sum operation on the identity features and the enhanced features obtained through the neighbor feature enhancement module (PFEM), the enhanced features are connected to the gradient reversal layer and the perspective loss to decouple the vehicle's identity information and perspective information.
2. The method according to claim 1, characterized in that The dual-branch neighbor feature enhancement network is used to extract the identity features and perspective features of the vehicle from the input image, including: The identity feature extraction and view feature extraction branches share the same convolutional neural network backbone layer to generate basic features; Through the backbone network Backbone, input image X i First, it is converted into basic features F; The identity feature branch maps the basic features to the identity feature space through convolution or fully connected layers to generate features F that are strongly correlated with the vehicle identity. id ,This feature is used to distinguish different vehicle identities; The view feature branch extracts the feature F related to the vehicle view view .
3. The method according to claim 2, characterized in that The method performs a weighted summation operation on the identity feature and the enhanced feature after PFEM, connects the enhanced feature to the gradient reversal layer and the perspective loss, and decouples the vehicle's identity information and perspective information, including: The neighbor feature enhancement module PFEM utilizes similar feature information within the same batch of processed data to fuse identity and view features. Based on a K-NN search process, this module retrieves the K neighbors closest to the current sample from the feature library and uses these neighbor features for feature enhancement, where K is a positive integer.
4. The method according to claim 3, characterized in that For each input ID feature, in the feature database Retrieve K nearest neighbor features from The process is expressed as: K-NN search uses metrics such as Euclidean distance or cosine similarity to find the K neighbor features closest to the current sample feature from the identity feature database.
5. The method according to claim 4, characterized in that For viewing angle features Retrieve K nearest neighbor features from the database After retrieving the nearest neighbor features of identity and viewpoint, PFEM first performs an intersection operation Extract the common parts between identity feature neighbors and view feature neighbors: The intersection operation extracts the common features between identity and view features, which represent the feature parts that have both identity discrimination and view information.
6. The method according to claim 5, characterized in that Through feature aggregation operation Compare these common features to the original identity features Fusion is performed to generate enhanced feature representations: Among them, the aggregation operation A weighted average is used.
7. The method according to claim 1, characterized in that The identity recognition task is optimized by the cross entropy loss and the triplet loss, and the view prediction task is optimized by the cross entropy loss. The cross entropy loss, triplet loss and cross entropy loss for the identity recognition task and the view prediction task are expressed as follows: in, and represent the identity probability distribution and view probability distribution predicted by the model, respectively. and represents the real identity label and viewpoint label, is the positive sample feature, is the negative sample feature, d(·,·) is the distance function between features, m is the hyperparameter in the triple loss, which is used to control the minimum distance difference between positive and negative samples. The overall optimization goal of the model is; Where λ is a hyperparameter that balances the weights of the gradient reversal layer and is used to control the strength of the gradient reversal.
8. A ground-to-air video vehicle re-identification device, characterized in that: include: A module that uses a dual-branch neighbor feature enhancement network to extract vehicle identity and view features from input images; The network includes two parallel branches adjacent to the backbone network, one branch is used to extract vehicle identity attributes, and the other branch is used to capture the difference of viewpoint changes; By performing a weighted sum operation on the identity features and the enhanced features after the neighbor feature enhancement module PFEM, the enhanced features are connected to the gradient reversal layer and perspective loss to decouple the vehicle's identity information and perspective information.
9. An electronic device, characterized in that: The method comprises a processor and a memory; when the processor executes the operating program stored in the memory, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Vehicle re-identification method in multi-view environment based on multi-center measurement loss
CN111814584A
Vehicle re-identification method based on dimension decoupling and non-local relation
CN114663861A
Target re-identification method, target re-identification model acquisition method and device, and terminal equipment
CN115424090A
Vehicle re-identification method and system based on aerial view perception of high-position camera
CN116630887A
Multi-image query vehicle re-identification method and device and storage medium
CN117708355A
Cited By
A multi-modal target detection and recognition method based on image processing
CN122391955A