A method for air-ground collaborative target recognition based on self-evolving visual cue learning
By adopting the self-evolution visual prompt learning method in the recognition of air-ground collaborative targets, dynamically adjusting the viewing feature prompts and decoupling the viewing feature features, the problem that traditional methods are difficult to adapt to the recognition of air-ground collaborative targets is solved, and accurate individual identity recognition and model generalization capabilities are improved.
Patent Information
- Application Number
- CN202510268028.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Traditional object recognition methods are difficult to adapt to the problem of extreme visual transformation and complementary information integration in air-ground collaborative target recognition, and deep learning-based methods rely on attribute annotations in the dataset, limiting the scalability of the air-ground collaborative dataset and the generalization ability of the model.
Using a method based on self-evolution visual cues learning, the feature extraction network and prompt recalibration module are used to decouple the feature extraction network through viewing angle decoupling, and the self-evolution viewing angle feature prompts are dynamically adjusted to achieve complete decoupling of viewing angle-independent features and viewing angle features, and the local feature refinement module is used to enhance the feature representation of viewing angle information.
It effectively solves the problem of visual transformation in air-ground collaborative target recognition, realizes accurate individual identity recognition across different viewing angle types, and improves the generalization ability of the model and the scalability of the data set.
Smart Images

Figure CN119785387B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to an air-to-ground collaborative target recognition method based on self-evolving visual cue learning. Background Art
[0002] Cross-platform collaborative target recognition is the core of intelligent monitoring systems, focusing on accurately identifying individuals across different camera perspectives and complex environmental changes. However, traditional target recognition methods are usually only applicable to the same perspective (such as ground perspective or aerial perspective), and fail to fully cope with the extreme visual transformation and complementary information integration problems brought about by air-ground collaborative target recognition (such as combining ground and drone cameras), which is very common in the real world.
[0003] Currently, deep learning-based methods use identity attributes to extract cross-view information. Generally speaking, individuals with similar attributes are more likely to be the same person. However, this method relies on attribute annotations in the dataset, which limits the scalability of the air-ground collaboration dataset. In addition, too much predefined external information input limits the generalization ability of the model. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides an air-ground collaborative target recognition method based on self-evolving visual cue learning to solve the problem of existing air-ground collaborative target recognition.
[0005] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: an air-ground collaborative target recognition method based on self-evolving visual cue learning, comprising the following steps:
[0006] S1: Use the view-decoupled feature extraction network to decouple the view-independent features of the input empty space image through layered decoupling and orthogonal loss constraint methods;
[0007] S2: Based on the decoupling results, the cue recalibration module is used to dynamically adjust the self-evolving view feature cues based on the cross-attention layer and the self-attention layer to obtain the calibrated view feature cues;
[0008] S3: The local feature refinement module is used to enhance the feature representation of the view information based on the calibrated view feature cues, and the air-ground collaborative target recognition based on self-evolving visual cue learning is completed.
[0009] Furthermore, the S1 includes the following sub-steps:
[0010] S11: Based on the input open space image, concatenate meta tags and view tags;
[0011] S12: Input the concatenated meta-tags and view tags into the visual transformer network, and perform subtractive separation and decoupling through each feature extraction layer of the visual transformer network;
[0012] S13: After the output of the last layer of the visual transformer network, the view tag is processed by the view classifier, and the view classification loss is used to guide the view tag to learn the view information. At the same time, the view-independent features decoupled from the view tag and the meta-tag are orthogonally separated based on the orthogonal loss to ensure that the view-independent features are completely decoupled from the view features.
[0013] Furthermore, the splicing meta tag and the perspective tag in S11 are:
[0014]
[0015] in, is a meta tag, is the perspective mark, is the local image feature, A visual transformer network with integrated view decoupling function. Indicates the tokenization of input image features. is the input image feature.
[0016] Furthermore, in S12, each feature extraction layer of the visual transformer network is subjected to subtraction separation and decoupling, and the formula is:
[0017]
[0018] in, It is the view-independent feature for subtractive separation operation.
[0019] Furthermore, the view classification loss in S13 is:
[0020]
[0021] in, is the view classification loss, is the total number of samples in the dataset, For the The true category label of the perspective of samples, For the The predicted probability that a sample belongs to the correct view class;
[0022] The orthogonal loss is:
[0023]
[0024] in, is the orthogonal loss, is the dimension of the feature space, represents the absolute value of the dot product between two eigenvectors, The perspective-invariant Dimensional features, is the view-dependent feature Dimensional features.
[0025] Furthermore, the calibrated viewing angle feature prompt in S2 is:
[0026]
[0027] in, is the calibrated viewing angle feature hint, is the linear transformation layer, is the self-attention layer, is the cross attention layer, is a view feature hint with a learnable vector, It is the view-independent feature for subtractive separation operation.
[0028] Furthermore, the local feature refinement module in S3 includes two stacked bidirectional attention modules and a feature fusion module, and specifically performs the following operations:
[0029] S31: Using two stacked bidirectional attention modules, self-attention layers and cross-attention layers are used in parallel in both the cue-to-image and image-to-cue directions to dynamically update and enhance all feature representations to obtain enhanced cue features and image features;
[0030] S32: Use the feature fusion module to deeply integrate the enhanced prompt features and image features to obtain the final recognition result of the open space image.
[0031] Furthermore, the enhanced prompt features and image features in S31 are:
[0032]
[0033]
[0034] in, To indicate the characteristics, is the image feature, is the linear transformation layer, is the self-attention layer, is the cross attention layer, is the calibrated viewing angle feature hint, is the local image feature.
[0035] Furthermore, the final recognition result of the empty ground image in S32 is:
[0036]
[0037] in, is the output token.
[0038] The beneficial effect of the present invention is that the present invention can dynamically adjust the self-evolving view feature prompts according to the input, dynamically generate and calibrate the view feature prompts that highly match the current view through real-time view information, thereby solving the problem of air-ground collaborative target recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 This is a flow chart of the air-to-ground collaborative target recognition method based on self-evolving visual cue learning of the present invention.
[0040] Figure 2 This is a schematic diagram of the main structure of the air-to-ground collaborative target recognition method based on self-evolving visual cue learning of the present invention. DETAILED DESCRIPTION
[0041] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0042] like Figure 1 As shown, a method for air-ground collaborative target recognition based on self-evolving visual cue learning includes the following steps:
[0043] S1: Use the view-decoupled feature extraction network to decouple the view-independent features of the input empty space image through layered decoupling and orthogonal loss constraint methods;
[0044] S2: Based on the decoupling results, the cue recalibration module is used to dynamically adjust the self-evolving view feature cues based on the cross-attention layer and the self-attention layer to obtain the calibrated view feature cues;
[0045] S3: The local feature refinement module is used to enhance the feature representation of the view information based on the calibrated view feature cues, and the air-ground collaborative target recognition based on self-evolving visual cue learning is completed.
[0046] The main structure of the technical solution of the present invention is as follows Figure 2 As shown in (a), it consists of a view-decoupled feature extraction network and a feature decoding network, where the feature decoding network is composed of a prompt recalibration module and a local feature refinement module.
[0047] The S1 includes the following sub-steps:
[0048] S11: Based on the input open space image, concatenate meta tags and view tags;
[0049] S12: Input the concatenated meta-tags and view tags into the visual transformer network, and perform subtractive separation and decoupling through each feature extraction layer of the visual transformer network;
[0050] S13: After the output of the last layer of the visual transformer network, the view tag is processed by the view classifier, and the view classification loss is used to guide the view tag to learn the view information. At the same time, the view-independent features decoupled from the view tag and the meta-tag are orthogonally separated based on the orthogonal loss to ensure that the view-independent features are completely decoupled from the view features.
[0051] The splicing meta tag and the perspective tag in S11 are:
[0052]
[0053] in, is a meta tag, is the perspective mark, is the local image feature, A visual transformer network with integrated view decoupling function. Indicates the tokenization of input image features. is the input image feature.
[0054] In S12, each feature extraction layer of the visual transformer network is subjected to subtraction separation and decoupling, and the formula is:
[0055]
[0056] in, It is the view-independent feature for subtractive separation operation.
[0057] The perspective classification loss in S13 is:
[0058]
[0059] in, is the view classification loss, is the total number of samples in the dataset, For the The true category label of the perspective of samples, For the The predicted probability that a sample belongs to the correct view class;
[0060] The orthogonal loss is:
[0061]
[0062] in, is the orthogonal loss, is the dimension of the feature space, represents the absolute value of the dot product between two eigenvectors, The perspective-invariant Dimensional features, is the view-dependent feature Dimensional features.
[0063] The calibrated viewing angle feature prompt in S2 is:
[0064]
[0065] in, is the calibrated viewing angle feature hint, is the linear transformation layer, is the self-attention layer, is the cross attention layer, is a view feature hint with a learnable vector, It is the view-independent feature for subtractive separation operation.
[0066] like Figure 2 As shown in (b), the prompt recalibration module maintains a set of self-evolving view feature prompts and recalibrates the view feature prompts according to the input view information through the multi-head attention mechanism. Specifically, the module initializes and maintains a set of view feature prompts with learnable vectors. During the prompt recalibration process, the module receives the view decoupled features and local features output by the view decoupled feature extraction network, effectively learns the decoupled information through the cross-attention mechanism and self-attention layer, and recalibrates the prompts to adaptively generate prompts that can better capture the cross-view information in the local features.
[0067] The local feature refinement module in S3 includes two stacked bidirectional attention modules and a feature fusion module, which specifically performs the following operations:
[0068] S31: Using two stacked bidirectional attention modules, self-attention layers and cross-attention layers are used in parallel in both the cue-to-image and image-to-cue directions to dynamically update and enhance all feature representations to obtain enhanced cue features and image features;
[0069] S32: Use the feature fusion module to deeply integrate the enhanced prompt features and image features to obtain the final recognition result of the open space image.
[0070] The enhanced prompt features and image features in S31 are:
[0071]
[0072]
[0073] in, To indicate the characteristics, is the image feature, is the linear transformation layer, is the self-attention layer, is the cross attention layer, is the calibrated viewing angle feature hint, is the local image feature.
[0074] The final recognition result of the empty ground image in S32 is:
[0075]
[0076] in, is the output token.
[0077] like Figure 2 As shown in (c), the local feature refinement module is based on two stacked bidirectional attention modules combined with a feature fusion module. Each bidirectional attention module requires four input variables: cue features and their position encodings, image features and their position encodings. Through the bidirectional attention mechanism, the module uses self-attention layers and cross-attention layers in parallel in both the cue-to-image and image-to-cue directions to dynamically update and enhance all feature representations to achieve accurate modeling of target features at different scales. After the last bidirectional attention module is processed, the cue features are pre-concatenated with a learnable output vector to introduce additional contextual information. Then, the concatenated cue features are used together with the image features as inputs to the feature fusion module. The feature fusion module adopts a transformer-decoder-like architecture, and deeply integrates the image features output by the bidirectional attention module with the query features through a carefully designed fusion mechanism.
[0078] In one embodiment of the present invention, the following specific examples are given to verify the effectiveness of the method proposed by the present invention. The specific process is as follows:
[0079] (1) Dataset selection
[0080] The present invention uses the air-to-ground system target recognition dataset to evaluate the performance of the proposed method in the air-to-ground system target recognition task. Specifically, two mainstream air-to-ground system target recognition datasets, AG-ReIDv1 and CARGO, are selected for training and evaluation. Cumulative matching characteristic curve (CMC) and mean average precision (mAP) are used as evaluation indicators to quantitatively analyze the model performance. For CMC, this embodiment calculates the percentage of correctly retrieved images in the first hit (Rank-1 accuracy).
[0081] (2) Implementation details setting
[0082] In terms of training details, before formal training, a domain generalization object re-identification model based on dynamic normalization technology is pre-trained on the ImageNet dataset. This step has two purposes: first, ImageNet is a large-scale annotated dataset suitable for image classification tasks, which can provide rich and effective parameters for the model; second, ImageNet contains thousands of categories and is highly versatile, suitable for various tasks, not limited to object re-identification.
[0083] In terms of data augmentation, each input image is pre-adjusted to 256 × 128 in both the training and testing phases. During the training phase, each image is horizontally flipped with a probability of 0.5 to enhance the diversity of the training data. In addition, image transformation strategies such as random erasing and padding are applied to achieve data augmentation.
[0084] In terms of hyperparameter setting, after the model is pre-trained on ImageNet, only fine-tuning is required in the formal training phase. The initial learning rate is set to 1.0×10 -2 , and gradually decreases according to the cosine decay schedule, with a minimum learning rate of 1.0×10 -5 , the momentum parameter is 0.9, and the decay rate is 0.1 to prevent the model from failing to converge due to excessive learning rate. The total number of training epochs is set to 120 to ensure that the model fully converges. In addition, in order to balance server performance and model convergence, the number of batch samples (batch size) is set to 8 (people) × (4 (ground view) + 4 (aerial view)) = 64.
[0085] (3) Implementation environment
[0086] The method of the present invention was implemented on a Sugon-W580-G20 server and a Linux Ubuntu 16.04.4 LTS operating system, using two NVIDIA GeForce GTX 1080 Ti graphics processors for training, and using the NVIDIA CUDA10.2 platform to accelerate training. In terms of software configuration, Python 3.8.13 (GCC 7.5.0) and PyTorch 1.6.0 were used, combined with Numpy 1.19.2, Fast-Reid 1.3, Pillow 9.0.1, Torchvision 0.7.0, cv2 4.5.5, CuDNN7.6.5 and other dependent libraries.
[0087] (4) Model application
[0088] In this stage, the input data is not augmented, and the data is only sampled to 256×128 image size. The model parameters remain fixed and no longer updated by the stochastic gradient descent algorithm. It is only used as an image feature extractor. In the actual reasoning process, the output features of the air-ground collaborative target recognition method based on self-evolving visual cue learning are used. For the query example, the features extracted after model feature inference Features of all images in the image library Perform Euclidean distance calculation:
[0089]
[0090] After obtaining the distance sequence through distance calculation, reorder it and take the L images closest to the query sample. If there is an image with the same ID as the query sample among these images, the query is considered successful. The relative baseline method of the present invention can effectively improve the extreme visual transformation and complementary information integration problems caused by air-ground collaborative target recognition (such as the combination of ground and drone cameras), and accurately identify individual identities across cameras of different viewing angles.
[0091] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the invention.
Claims
1. A method for air-ground collaborative target recognition based on self-evolving visual cue learning, characterized in that: The following steps are involved: S1: The view-decoupled feature extraction network is used to decouple the view-independent features of the input empty space image through layered decoupling and orthogonal loss constraint methods; S2: Based on the decoupling results, the cue recalibration module is used to dynamically adjust the self-evolving view feature cues based on the cross-attention layer and the self-attention layer to obtain the calibrated view feature cues; S3: Use the local feature refinement module to enhance the feature representation of view information based on the calibrated view feature cues, and complete the air-ground collaborative target recognition based on self-evolving visual cue learning; The local feature refinement module in S3 includes two stacked bidirectional attention modules and a feature fusion module, which specifically performs the following operations: S31: Using two stacked bidirectional attention modules, self-attention layers and cross-attention layers are used in parallel in both the cue-to-image and image-to-cue directions to dynamically update and enhance all feature representations to obtain enhanced cue features and image features; S32: using a feature fusion module to deeply integrate the enhanced prompt features and image features to obtain the final recognition result of the open space image; The enhanced prompt features and image features in S31 are: in, To indicate the characteristics, is the image feature, is the linear transformation layer, is the self-attention layer, is the cross attention layer, is the calibrated viewing angle feature hint, is the local image feature; The final recognition result of the empty ground image in S32 is: in, is the output token.
2. According to claim 1, a method for air-ground collaborative target recognition based on self-evolving visual cue learning is characterized in that: The S1 includes the following sub-steps: S11: Based on the input open space image, concatenate meta tags and view tags; S12: Input the concatenated meta-tags and view tags into the visual transformer network, and perform subtractive separation and decoupling through each feature extraction layer of the visual transformer network; S13: After the output of the last layer of the visual transformer network, the view tag is processed by the view classifier, and the view classification loss is used to guide the view tag to learn the view information. At the same time, the view-independent features decoupled from the view tag and the meta-tag are orthogonally separated based on the orthogonal loss to ensure that the view-independent features are completely decoupled from the view features.
3. The method for air-ground collaborative target recognition based on self-evolving visual cue learning according to claim 2 is characterized in that: The splicing meta tag and the perspective tag in S11 are: in, is a meta tag, is the perspective mark, is the local image feature, A visual transformer network with integrated view decoupling function. Indicates the tokenization of input image features. is the input image feature.
4. The method for air-ground collaborative target recognition based on self-evolving visual cue learning according to claim 3 is characterized in that: In S12, each feature extraction layer of the visual transformer network is subjected to subtraction separation and decoupling, and the formula is: in, It is the view-independent feature for subtractive separation operation.
5. The method for air-ground collaborative target recognition based on self-evolving visual cue learning according to claim 2 is characterized in that: The perspective classification loss in S13 is: in, is the view classification loss, is the total number of samples in the dataset, For the The true category label of the perspective of samples, For the The predicted probability that a sample belongs to the correct view class; The orthogonal loss is: in, is the orthogonal loss, is the dimension of the feature space, represents the absolute value of the dot product between two eigenvectors, The perspective-invariant Dimensional features, is the view-dependent feature Dimensional features.
6. The method for air-ground collaborative target recognition based on self-evolving visual cue learning according to claim 1 is characterized in that: The calibrated viewing angle feature prompt in S2 is: in, is the calibrated viewing angle feature hint, is the linear transformation layer, is the self-attention layer, is the cross attention layer, is a view feature hint with a learnable vector, It is the view-independent feature for subtractive separation operation.
Citation Information
Patent Citations
Large and small model collaborative tracking method based on time sequence-vision fusion
CN119027459A
Techniques for sharing mapping data between an unmanned aerial vehicle and a ground vehicle
US20210132612A1
Cited By
Rail transit non-inductive ticket checking method and system based on spatial-temporal trajectory and biological feature fusion
CN121725527A
A rail transit non-inductive ticket checking method and system based on spatiotemporal trajectory and biological feature fusion
CN121725527B