Air-ground pedestrian re-identification method combining multi-frame information and prompt learning

By combining multi-frame information with cue learning, and utilizing rotation-invariant attention and inter-frame information attention mechanisms, the unstable feature extraction problem in cross-view scenarios is solved, achieving high-accuracy cross-view pedestrian re-identification.

CN121884437APending Publication Date: 2026-04-17SOUTH CHINA UNIV OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively utilize inter-frame information in cross-view scenarios between air and ground, lack effective prior guidance mechanisms, resulting in low cross-view matching accuracy and unstable feature extraction when the model differs between air and ground perspectives.

Method used

By combining multi-frame information with cue learning, and by introducing rotation-invariant attention and cue learning mechanisms, we simulate aerial viewpoint rotation distortion to enhance the model's robustness to viewpoint changes. Furthermore, through inter-frame information attention and cue-guided visual attention mechanisms, we achieve the fusion and alignment of cross-viewpoint features.

Benefits of technology

It significantly improves the accuracy and generalization ability of pedestrian re-identification in aerial-ground cross-view videos, and can achieve accurate identity matching under conditions of large viewpoint differences and occlusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884437A_ABST
    Figure CN121884437A_ABST
Patent Text Reader

Abstract

The invention discloses an air-ground pedestrian re-identification method combining multi-frame information and prompt learning, and the method comprises the following steps: inputting a video sequence into a trained visual encoder model, mapping the same pedestrian at different visual angles into a consistent feature space, and achieving the cross-visual-angle pedestrian re-identification and tracking; according to the visual encoder model, a random rotation transformation strategy of structure perception is introduced in the input embedding stage of a visual encoder backbone network, pedestrian vector features of each frame are rotated, visual angle rotation distortion generated by aerial shooting is simulated, and an enhanced sequence is generated. Extracting global features of the enhanced sequence and the unenhanced sequence through a backbone network; inputting the global feature into an inter-frame information attention module for time dimension attention calculation to obtain an average feature of multi-frame fusion; and then inputting the multi-frame fusion average features into a prompting and guiding visual attention module to generate a text prompt so as to guide model discriminative character features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and intelligent video analysis, and in particular to a method for re-identifying pedestrians in open and ground environments by combining multi-frame information and cue learning. Background Technology

[0002] Person re-identification (ReID) is a key foundational technology in intelligent video surveillance and multi-camera collaborative sensing, playing a crucial role in various applications such as public safety, smart city management, traffic management, and autonomous sensing for unmanned systems. Its core task is to accurately identify the same pedestrian from different angles, lighting conditions, and background interference in a multi-camera environment. High-quality person re-identification results not only improve the continuity and reliability of cross-camera tracking systems but also provide important support for scenarios such as target retrieval and UAV patrol-assisted decision-making. With the rapid construction of city-level multi-source video systems and the development of air-ground collaborative sensing systems, the demand for video person re-identification methods with high robustness and strong generalization capabilities is becoming increasingly urgent.

[0003] Traditional pedestrian re-identification technologies primarily rely on image or video data captured by fixed ground cameras. Related research has evolved from early handcrafted features to representation learning frameworks dominated by deep convolutional networks and Transformers, achieving significant performance breakthroughs on public datasets. However, these techniques largely assume relatively stable camera views, limited scale variations, and controllable target pose changes, making them ill-suited for handling high-angle videos from aerial platforms (such as drones). Aerial views exhibit significant rotational distortion, small target presentation areas, irregular occlusion, and drastic viewpoint changes, causing traditional ground-based ReID models to suffer from feature degradation and unstable identity matching in cross-view scenarios between air and ground.

[0004] In recent years, with the rapid rise of urban air-ground integrated monitoring systems, drones have been widely used in traffic inspections, security patrols, and public area coverage. Aerial video has gradually become an important data source for pedestrian recognition and tracking. However, there are significant differences between aerial and ground perspectives in terms of imaging height, pitch angle, and target resolution, making cross-perspective pedestrian re-identification a highly challenging task. On the one hand, the differences between aerial and ground perspectives result in completely different pedestrian appearances, making it difficult for models to directly learn consistent representations. On the other hand, the scale changes, motion blur, and occlusion of targets across frames are complex and varied, making it difficult to fully utilize temporal information. Although some studies have attempted to mitigate the differences between different perspectives using temporal feature modeling or attention mechanisms, existing methods still have two main shortcomings: first, insufficient utilization of inter-frame information leads to video-level features failing to effectively improve matching stability in cross-perspective scenarios; second, models often lack sufficient domain priors when facing significant perspective differences, making it difficult to align learned pedestrian features across modalities, thus limiting recognition accuracy.

[0005] Against the backdrop of the rapid development of the cue-based deep learning paradigm, utilizing cues as additional prior knowledge for representation alignment offers a new approach to improving the performance of cross-view video ReID. However, existing research on cross-view pedestrian re-identification in both air and ground still explores cue-based learning to a very limited extent, lacking a robust method that can simultaneously integrate inter-frame information and cue guidance, and effectively adapt to differences between air and ground. Therefore, there is an urgent need for an air-to-ground cross-view video pedestrian re-identification technology that integrates temporal features and cue-based learning mechanisms to address the limitations of current models in terms of viewpoint differences, scene complexity, and temporal consistency, and to achieve more stable and accurate cross-view identity matching.

[0006] In a video-based person re-identification method (CN114973323B), the inventors detect key points of the human body in each frame and assign weights to each frame based on occlusion and recognition quality. Then, a convolutional neural network is used to extract static appearance features, while a recurrent neural network is combined to extract motion features such as gait, and these are weighted according to frame weights to form video-level features. Subsequently, the Euclidean and cosine distances between the static and motion features are calculated, and after interval transformation, they are adaptively fused according to the shooting time to obtain a comprehensive similarity, thus achieving the function of person re-identification.

[0007] Current pedestrian re-identification technologies still have significant shortcomings in cross-view scenarios involving both aerial and ground perspectives. Models trained on ground cameras struggle to adapt to the large overhead angles, scale compression, and rapid background changes brought about by drone perspectives, leading to easily degraded appearance features and a significant decrease in cross-view matching accuracy. While existing video ReID methods utilize inter-frame information, most only perform shallow temporal modeling, easily resulting in insufficient temporal consistency when there are significant differences between aerial and ground perspectives. Furthermore, existing cross-view methods lack effective prior guidance mechanisms; models can only rely on data to automatically learn alignment strategies, limiting their adaptability to complex viewpoint biases. Overall, current technologies generally lack effective means to simultaneously consider inter-frame dynamic features and viewpoint difference compensation, making it difficult to meet the demand for highly robust video pedestrian re-identification in aerial-ground fusion monitoring scenarios. Summary of the Invention

[0008] To overcome the aforementioned technical problems, this invention provides an aerial-to-ground pedestrian re-identification method that combines multi-frame information with cue learning. Specifically, this invention integrates inter-frame temporal information from aerial and ground-view video sequences, incorporates a rotation-invariant attention mechanism to improve robustness to drastic rotational changes in drone-captured images, and introduces a cue learning mechanism to enhance the generalization ability of cross-view features, thereby achieving accurate identity matching of the same pedestrian under conditions of large viewpoint differences, scale variations, and occlusion. This method integrates multi-frame information to suppress drastic rotational disturbances; it introduces cue learning to bridge the aerial-to-ground viewpoint differences, providing a practical solution for the implementation of cross-view video pedestrian re-identification. It can be widely applied in fields such as intelligent security monitoring, urban aerial-to-ground collaborative inspection, drone-assisted monitoring, and multi-source video fusion analysis.

[0009] The present invention is achieved by at least one of the following technical solutions.

[0010] A method for re-identifying pedestrians in open spaces by combining multi-frame information and cue learning includes the following steps: inputting a video sequence into a trained visual encoder model, extracting pedestrian features, and mapping the same pedestrian from different viewpoints to a consistent feature space to achieve cross-viewpoint pedestrian re-identification and tracking.

[0011] Furthermore, the structure of the visual encoder model is as follows: In the input embedding stage of the Vision Transformer backbone network, a structure-aware random rotation transformation strategy is introduced to rotate the pedestrian vector features of each frame to simulate the perspective rotation distortion caused by aerial shooting and generate an enhanced Patch Embedding sequence. Global features are extracted from both the enhanced and unenhanced Patch Embedding sequences using the Vision Transformer backbone network. These global features are then input into the inter-frame information attention module for temporal attention calculation to obtain the average features fused from multiple frames. Subsequently, the average features fused from multiple frames are input into the prompting and guidance visual attention module to generate text prompts to guide the model in discriminating character features.

[0012] Furthermore, the Vision Transformer backbone network adopts the ViT-Base structure based on CLIP pre-training. The ViT-Base architecture contains multi-layer Transformer encoders, each layer containing a multi-head self-attention mechanism and a feedforward neural network. Through the self-attention mechanism, it effectively captures the long-distance dependencies between regions in the image, achieving accurate modeling of the global context of the image.

[0013] Furthermore, the inter-frame information attention module calculates the interdependencies between visual features across multiple frames in the time dimension. By calculating the similarity between queries, keys, and values, it generates an attention weight matrix, which then performs weighted fusion of features from different frames and takes the average value to obtain the average feature of the multi-frame fusion.

[0014] Furthermore, the prompting and guiding visual attention module to generate text prompts includes the following steps: First, a three-layer fully connected network is used to convert the category vector of the average feature into a pseudo-semantic description. Then, it is embedded into a fixed-format text prompt through an image-to-text inversion network and encoded into a text feature vector by a Transformer text encoder. Subsequently, the text feature vector is fused with visual features to guide the model to extract discriminative character features.

[0015] Furthermore, the structure-aware stochastic rotation transformation strategy includes: The input image is divided into several non-overlapping image patches. Each patch is mapped to a high-dimensional vector through linear projection, forming a one-dimensional sequence of Patch Embedding serialization matrix. The serialization matrix is ​​reconstructed into a three-dimensional feature grid based on the spatial relative position relationship in the original image, restoring the spatial topological layout of the embedded features. Subsequently, a random rotation operation is performed on the 3D feature mesh representation. A rotation angle is randomly selected from the set {90°, 180°, 270°} with uniform probability, and a rigid rotation transformation is performed on the 3D feature mesh along the image plane. The rotated and updated 3D feature mesh is flattened back into a 1D sequence to obtain the enhanced PatchEmbedding sequence.

[0016] Furthermore, the loss function of the visual encoder model includes identity classification loss, triplet loss, rotation invariant constraint loss, and image-text contrast loss. By assigning corresponding weights, multiple supervision signals are integrated to guide the model to learn feature representations with strong discriminative power.

[0017] The system for implementing the above-mentioned method of re-identifying pedestrians in the air by combining multi-frame information and cue learning includes: The data processing module is used to collect and process data from the training and test sets. The visual encoder module is used to identify cross-view identities; The training module is used to train the visual encoder model.

[0018] A computer device according to the present invention includes a memory and a processor, the memory being electrically connected to the processor, the memory storing a computer program, which, when executed by the processor, causes the processor to implement the method described herein.

[0019] The present invention provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor implements the method described herein.

[0020] Compared with existing technologies, the beneficial effects of the present invention are as follows: This invention effectively simulates rotational distortion from the perspective of an aerial drone by introducing a random rotation transformation enhancement strategy, thereby improving the model's robustness to viewpoint changes. Combined with an inter-frame information attention mechanism, it fully leverages temporal correlations and multi-frame complementary features in the video sequence, enhancing the stability of the representation. Furthermore, through a cue-guided visual attention mechanism, semantic text cues are generated using image inversion and injected into the visual encoding process, providing the model with cross-viewpoint invariant prior knowledge. This guides the model to focus on key features for identity consistency, enabling it to learn highly discriminative pedestrian representations even under significant viewpoint differences. End-to-end optimization of the multi-task joint loss function synergistically improves the model's classification accuracy, inter-sample distance measurement, and modality consistency. The overall method forms a complete closed loop in data processing, feature extraction, and training strategies, significantly improving the accuracy and generalization ability of aerial-ground cross-view video pedestrian re-identification. Attached Figure Description

[0021] Figure 1 The flowchart illustrates a method for re-identifying pedestrians in open areas that combines multi-frame information with cue learning, as provided by this invention.

[0022] Figure 2 The network structure diagram of the VisionTransformer visual encoder, which integrates inter-frame information attention and cue-guided visual attention, is provided by this invention.

[0023] Figure 3 This is a structural diagram of the inter-frame information attention module provided by the present invention.

[0024] Figure 4 This is a structural diagram of the visual attention prompting module provided by the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and technical advantages of the present invention more apparent, the present invention will now be further described in conjunction with the accompanying drawings.

[0026] like Figure 1 As shown in the figure, an airborne pedestrian re-identification method combining multi-frame information and cue learning in this embodiment includes the following steps: Step S1: Collect and organize video sequences from existing public datasets that simultaneously contain aerial drone perspectives and ground-based fixed camera perspectives, ensuring coverage of typical cross-view challenge scenarios such as large pitch angles, rotational distortion, and significant perspective differences. Employ systematic preprocessing to clean redundant segments, verify the consistency of identity labeling, and divide the dataset into training and testing sets according to a predetermined ratio, providing a standardized data foundation for subsequent model training and testing.

[0027] To construct a high-quality, well-structured video dataset suitable for model training and evaluation, a systematic preprocessing workflow was adopted, including video acquisition, format standardization, content cleaning, frame sampling, naming conventions, and data deduplication. The data originated from two types of monitoring equipment: aerial videos captured by drones and horizontal or near-horizontal videos captured by fixed ground cameras. All raw videos were saved in standard MP4 format to ensure compatibility and stability.

[0028] The preprocessing in this embodiment includes the following steps: First, all videos are uniformly re-encoded to a resolution of 256×128 pixels. This size reduces computational overhead while preserving basic semantic information and ensuring spatial consistency across multiple video sources, avoiding model bias caused by resolution differences.

[0029] Next, quality screening and content cleaning were carried out to remove the following three types of video clips: (1) overall blurry, unable to identify the main outline or key features; (2) the target person is obscured by more than 50%, resulting in serious loss of posture or identity information; (3) there is a conflict of identity labels, that is, the same target is assigned multiple different IDs at the same time.

[0030] From the retained valid videos, image frames are extracted using a temporal uniform sampling method: each video segment is divided into 8 equal time periods, and one frame is taken from the center of each time period to ensure a balanced temporal distribution between frames. All retained image frames are named according to the rule of "video number-frame number-identity". This naming method enhances data traceability and supports subsequent automated partitioning of training and test sets.

[0031] Finally, to prevent data leakage, the principle of cross-set isolation is strictly enforced: all samples of the same person (identity) in the same scenario belong to only one of the training or test sets, avoiding their simultaneous appearance in both sets. This strategy is achieved through identity-based grouping, ensuring that the model evaluation results effectively test generalization ability.

[0032] Step S2: Establish the visual encoder model, as follows: Step S21: In the input embedding stage of the Vision Transformer (VIT) backbone network, this invention introduces a structure-aware random rotation transformation strategy to enhance the rotation of pedestrian vector features in each frame, simulating the viewpoint rotation distortion caused by aerial shooting. This strengthens the model's robustness to rotational deviations and enhances its robustness to distortions caused by camera pose rotation changes from an aerial perspective. This strategy does not directly apply to the original pixel space but performs feature-level spatial geometric enhancement in the patch embedding layer. The random rotation transformation strategy is implemented in the patch embedding layer, specifically as follows: First, in the standard ViT preprocessing flow, the input image is uniformly divided into several non-overlapping image patches. Each patch is mapped to a high-dimensional vector through a linear projection layer, forming a one-dimensional sequence-formed PatchEmbedding serialization matrix E∈. N×D ,in N This represents the total number of patches. D This is the embedding dimension. Based on this, the serialized matrix is ​​further reconstructed into a three-dimensional tensor structure T∈ according to its spatial relative position in the original image. H×W×D ,in H and W The height and width of the patch grid are respectively used to restore the spatial topological layout of the embedded features.

[0033] Subsequently, a random rotation operation is performed on this 3D feature mesh representation. A rotation angle is randomly selected with uniform probability from the set {90°, 180°, 270°}, and a rigid rotation transformation is performed on the 3D feature mesh along the image plane (i.e., the height-width plane) to obtain the rotated 3D feature mesh. This operation simulates the camera yaw phenomenon caused by heading adjustments or wind disturbances during the flight of an unmanned aerial vehicle platform, allowing the model to access diverse directional configuration samples. The rotated feature map retains the same dimensional structure, but its spatial arrangement has undergone geometric transformation.

[0034] The rotated and updated 3D feature mesh is then flattened back into a 1D sequence to obtain the enhanced PatchEmbedding sequence E. rot Meanwhile, the original, unrotated Embedding sequence E is preserved. orig Both inputs are fed in parallel into the subsequent visual encoder based on the Vision Transformer backbone network. This dual-path input design allows the model to simultaneously model the relationship between the original structure and multi-directional variants in the attention mechanism, improving its ability to decouple orientation-sensitive patterns.

[0035] It is worth noting that this random rotation transformation strategy is strictly limited to use during the training phase and is randomly activated with a 50% probability (i.e., a trigger probability p=0.5). During the inference phase, this mechanism is completely turned off, and all inputs are processed in their original attitudes to ensure the determinism and consistency of the prediction process, forcing the model to learn a rotation-invariant mapping for the random yaw of the aerial platform at the parameter level.

[0036] The patch embedding layer adjusts the spatial and channel dimensions of the input image by adjusting the size of the convolution kernel, stride, and output channels, making it suitable for subsequent Vision Transformer backbone operations.

[0037] Step S22: A visual encoder based on a Vision Transformer backbone network is used to achieve deep fusion and cross-view consistency modeling of multi-frame visual features. This visual encoder integrates an inter-frame information attention (IFIA) mechanism and a prompt-guided visual attention (PGVA) mechanism. The visual encoder uses the inter-frame information attention mechanism to mine the complementary correlations of multi-frame visual features in the spatiotemporal sequence, and injects cross-view priors through the prompt-guided visual attention mechanism to drive the model to learn rotation-invariant discriminative features in the spatial-ground feature space with significant viewpoint differences.

[0038] Specifically, the visual encoder first processes the enhanced Embedding sequence E from step S21 through the Vision Transformer backbone network. rot Compared with the original unrotated Embedding sequence E orig Global features are extracted to represent the visual features of each frame. The global features are input into the inter-frame information attention module (IFIA) for temporal attention calculation to obtain the average features of multi-frame fusion. Then, the average features of multi-frame fusion are input into the prompt guidance visual attention module. Based on the category vector (cls_token) of the average features, text prompts are generated through the image-to-text reverse net (I2TNet) and the Transformer text encoder to guide the model to learn visual features. This enables the model to learn discriminative pedestrian features under the significant difference between aerial and ground perspectives.

[0039] As one embodiment, the Vision Transformer backbone network adopts a CLIP-based pre-trained ViT-Base structure, in which the image is divided into fixed-size 16×16 pixel patches. Each patch is mapped to a 768-dimensional embedding space through a linear projection layer, outputting a 768-dimensional token. Simultaneously, learnable positional encodings are added to preserve the spatial structure information of the image. The ViT-Base architecture contains a 12-layer Transformer encoder, each layer including a multi-head self-attention mechanism and a feedforward neural network. The self-attention mechanism effectively captures long-distance dependencies between regions in the image, thereby achieving accurate modeling of the global context of the image.

[0040] Specifically, the Embedding sequence E generated in step S21 rot With E orig First, the features are fed into the VisionTransformer backbone network for initial feature extraction, obtaining visual feature representations for each frame. Then, these features are fed into the Inter-Frame Information Attention (IFIA) module. This module uses a Multi-Head Self-Attention (MHSA) mechanism to calculate the interdependencies between visual features across multiple frames in the temporal dimension. By calculating the similarity between queries, keys, and values, an attention weight matrix is ​​generated. The features from different frames are then weighted and fused, and the average value is taken to obtain the multi-frame fused average feature. The specific formula is as follows:

[0041]

[0042]

[0043] in, , , This represents the query, key, and value generated in the i-th frame. Representing feature dimension, Indicates attention weights. Indicates the number of frames in a video sequence. This represents the character features extracted by the inter-frame information attention module. This process effectively mines the complementary correlations of visual features across multiple frames in a spatiotemporal sequence, overcoming the limitations of single-frame feature extraction and enhancing the model's ability to understand dynamic scenes.

[0044] The multi-frame fused average features are input into the Guided Visual Attention (PGVA) module, which guides the visual encoder to learn discriminative person features through viewpoint-independent textual cues. The PGVA module first uses a three-layer fully connected network to convert the class vector cls_token of the average features into a pseudo-semantic description. This description is then embedded into a fixed-format text cue via an image-to-text inversion network and encoded into a text feature vector by a CLIP-based pre-trained Transformer text encoder. Subsequently, the text feature vector is fused with the visual features, and a cross-attention mechanism guides the model to focus on key visual feature regions related to the text cue, thereby injecting cross-viewpoint prior knowledge. The specific formula is as follows:

[0045]

[0046]

[0047] in The cls_token vector represents the average features fused from multiple frames, and MLP represents an image-to-text inversion network with a three-layer fully connected structure. This represents a pseudo-semantic description. ~ Text hints indicating a fixed template. Indicates a text encoder. This represents the text features extracted by the text encoder. Representing feature dimension, This indicates that the visual attention module extracts the features of the person indicated by the prompt. This mechanism design enables the model to effectively learn rotation-invariant discriminative features in an air-to-ground feature space with significant differences in viewpoints. Even when there are significant differences between the air viewpoint and the ground viewpoint, it can maintain the consistency of identity features, significantly improving the model's accuracy in cross-viewpoint identity recognition tasks.

[0048] Step S3: Construct and optimize the joint loss function. This function effectively integrates multiple supervision signals through a weight allocation strategy to guide the model to learn feature representations with strong discriminative power.

[0049] Specifically, the joint loss function consists of four parts, each with a predetermined weight, which are linearly superimposed: identity classification loss L id (Weight 1.0) Enhances the model's ability to distinguish between different identities; triple loss L tri (Weight 0.8) By constructing a triplet structure of anchor points, positive samples, and negative samples, the distance between samples of the same class is further reduced and the distance between samples of different classes is increased, thereby improving the discriminative power of the features; rotation-invariant constraint loss L rot (Weight 0.5) is specifically used to enhance the model's robustness to rotational changes, constraining the consistency between the original features generated in step S21 and the rotational features under different rotation angles; Image-text contrast loss L i2t (Weight 0.4) then utilizes a cross-modal semantic alignment mechanism to guide the model to capture richer semantic information, thereby maintaining the consistency of identity features even when there are significant differences between air and ground perspectives.

[0050] Regarding the training strategy, the training set constructed in step S1 includes video samples and their corresponding annotation files, and the visual encoder model is trained and optimized using a loss function as the joint loss. The trained model is then used to perform performance testing on video samples in the test machine.

[0051] As one example, the AdamW optimizer is used as the core optimization algorithm, and the initial learning rate is set to 3×10. -4 The training process consists of 120 epochs with a batch size of 64. During the model evaluation phase, this approach uses mAP (Mean Average Precision) and Rank-1 as the core evaluation metrics, and uses the trained model to perform performance tests on video samples in the test set.

[0052] The trained visual encoder model is used to extract pedestrian identity features with strong discriminative power and robustness to changes in viewpoint from videos captured by aerial drones and ground cameras. It can map the same pedestrian from different viewpoints to a consistent feature space, and can accurately match their identity even in the presence of large pitch angles, rotational distortion and significant appearance differences, thereby achieving pedestrian re-identification and tracking across viewpoints.

[0053] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, enabling those skilled in the art to better understand and utilize the invention.

Claims

1. A method for re-identifying pedestrians in open terrain by combining multi-frame information and cue learning, characterized in that, The process includes the following steps: inputting the video sequence into the trained visual encoder model, extracting pedestrian features, and mapping the same pedestrian from different viewpoints to a consistent feature space to achieve cross-viewpoint pedestrian re-identification and tracking.

2. The method for re-identifying pedestrians in open terrain by combining multi-frame information and cue learning according to claim 1, characterized in that, The structure of the visual encoder model is as follows: In the input embedding stage of the Vision Transformer backbone network, a structure-aware random rotation transformation strategy is introduced to rotate the pedestrian vector features of each frame to simulate the perspective rotation distortion caused by aerial shooting and generate an enhanced Patch Embedding sequence. Global features are extracted from both the enhanced and unenhanced Patch Embedding sequences using the Vision Transformer backbone network. The global features are input into the inter-frame information attention module to perform temporal dimension attention calculation, and the average features of multi-frame fusion are obtained. The multi-frame fused average features are then input into the prompt-guided visual attention module to generate text prompts to guide the model in discriminating human features.

3. The method for re-identifying pedestrians in open terrain by combining multi-frame information and cue learning according to claim 2, characterized in that, The Vision Transformer backbone network adopts the ViT-Base structure based on CLIP pre-training. The ViT-Base architecture contains multi-layer Transformer encoders, each layer containing a multi-head self-attention mechanism and a feedforward neural network. Through the self-attention mechanism, it effectively captures the long-distance dependencies between regions in the image, achieving accurate modeling of the global context of the image.

4. The method for re-identifying pedestrians in open terrain by combining multi-frame information and cue learning according to claim 2, characterized in that, The inter-frame information attention module calculates the interdependencies between visual features across multiple frames in the time dimension. It generates an attention weight matrix by calculating the similarity between queries, keys, and values, and then performs weighted fusion of features from different frames, taking the average value to obtain the average feature of the multi-frame fusion.

5. The method for re-identifying pedestrians in open terrain by combining multi-frame information and cue learning according to claim 2, characterized in that, The prompt-guided visual attention module generates text prompts by including the following steps: First, a three-layer fully connected network is used to convert the category vector of the average feature into a pseudo-semantic description. Then, it is embedded into a fixed-format text prompt through an image-to-text inversion network and encoded into a text feature vector by a Transformer text encoder. Subsequently, the text feature vector is fused with visual features to guide the model to extract discriminative character features.

6. The method for re-identifying pedestrians in open terrain by combining multi-frame information and cue learning according to claim 2, characterized in that, Structure-aware stochastic rotation transformation strategies include: The input image is divided into several non-overlapping image patches. Each patch is mapped to a high-dimensional vector through linear projection, forming a one-dimensional sequence of Patch Embedding serialization matrix. The serialization matrix is ​​reconstructed into a three-dimensional feature grid based on the spatial relative position relationship in the original image, restoring the spatial topological layout of the embedded features. Subsequently, a random rotation operation is performed on the 3D feature mesh representation. A rotation angle is randomly selected from the set {90°, 180°, 270°} with uniform probability, and a rigid rotation transformation is performed on the 3D feature mesh along the image plane. The rotated and updated 3D feature mesh is flattened back into a 1D sequence to obtain the enhanced PatchEmbedding sequence.

7. The method for re-identifying pedestrians in open terrain by combining multi-frame information and cue learning according to claim 2, characterized in that, The loss function of the visual encoder model includes identity classification loss, triplet loss, rotation invariant constraint loss, and image-text contrast loss. By assigning corresponding weights, multiple supervision signals are integrated to guide the model to learn feature representations with strong discriminative power.

8. A system for implementing the aerial pedestrian re-identification method combining multi-frame information and cue learning as described in claim 1, characterized in that, include: The data processing module is used to collect and process data from the training and test sets. The visual encoder module is used to identify cross-view identities; The training module is used to train the visual encoder model.

9. A computer device comprising a memory and a processor, the memory being electrically connected to the processor, the memory storing a computer program, characterized in that: When the computer program is executed by the processor, it causes the processor to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • A video-based person re-identification method

    CN114973323B

Cited By

  • Air-ground pedestrian re-identification method, system, device and storage medium

    CN122290180A