Unmanned aerial vehicle multi-modal feature fusion target tracking method and system based on natural language description
By adopting natural language description and multimodal feature fusion technology on drones, image quality and target feature pronunciation are improved, and insufficient tracking is solved due to poor image quality of drones, and more efficient and accurate target tracking is achieved.
Patent Information
- Application Number
- CN202510082300.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-20
AI Technical Summary
When the image quality collected by the drone is poor or the image characteristics are not obvious, the target tracking ability and long-term tracking ability are poor.
The multimodal feature fusion target tracking method of UAV based on natural language description is adopted to enhance image features through the scene-context feature pyramid network, and combine local alignment of visual-language dual-modal features to improve image quality and obviousness of target features.
The performance and long-term tracking capabilities of drone target tracking are improved, the adaptability to complex traffic accident scenarios and the utilization of multimodal information are enhanced, and the target discrimination ability and tracking accuracy are improved.
Smart Images

Figure QLYQS_1 
Figure BDA0005249035840000061 
Figure BDA0005249035840000141
Abstract
Description
Technical Field
[0001] A method and system for unmanned aerial vehicle multimodal feature fusion target tracking based on natural language description are used for unmanned aerial vehicle multimodal feature fusion target tracking, belonging to the technical field of computer vision and image processing. Background Art
[0002] With the reduction of the production and application costs of drones, and the rapid development of technologies such as sensors, automation, and artificial intelligence, drones can be used in the fields of coverage, tracking, confrontation, and communication by carrying cameras, GPS positioning systems, a variety of hardware sensors, mission payload systems, and other devices, and can efficiently complete various tasks that are difficult to operate by manpower. Video single target tracking is one of the basic tasks in the field of computer vision, and is widely used in intelligent industries such as robot vision, video security, and sports video analysis. The target tracking algorithm carried by drones can not only speed up the scene search time, but also save a lot of manpower and material resources. At present, the existing target tracking technology only relies on the initial visual information of the target and uses the initial appearance features of the target to locate the moving target. However, this form of target tracking brings about the performance bottleneck of the tracking algorithm and limits the application scope of single target tracking. On the one hand, the appearance features of the moving target contain rich texture information, but only have sparse target semantics, which makes the single target tracker have limited ability to distinguish the moving target; on the other hand, the appearance of the target changes continuously during the long-term movement, which is different from the initial appearance, making it impossible for the single target tracker to stably track the moving target for a long time.
[0003] CN 202410369741.0-A multimodal tracking method for unmanned aerial vehicles uses the combined features of template images, search area images and texts, performs feature extraction and modal interaction based on the Transformer coding layer, intercepts the features of the search area part and inputs them into the feedforward neural network for classification and regression, and calculates the final bounding box of the tracked target based on the obtained classification response map, offset and scale size; it solves the problem that the current target tracking technology cannot adapt to the fast change of perspective and high field of view of unmanned aerial vehicles. This application can improve the performance of target tracking and optimize the tracking effect. However, there are the following technical problems:
[0004] When the image quality collected by UAV is poor or the image features are not obvious, it is easy to cause problems of poor target ability and long-term tracking ability. This method proposes a contextual image feature enhancement module for UAV images to improve the quality of UAV images used in deep learning algorithms.
[0005] In the task of traffic accident detection and recognition, it is difficult to capture diverse visual features by relying solely on visual target trackers such as DeepSORT. In addition, traffic accident data often show a long-tail distribution, and the visual gap within the same category is still large. It is difficult to summarize the different characteristics between accidents by relying solely on a single accident category. After introducing natural language description, this paper can locate the referenced target based on a given visual language reference and use multimodal enhancement of target features.
[0006] In existing tasks of using drones to search for traffic accidents, usually only trained image features can be used, and the search target cannot be changed, or the search target can be changed instantly according to different scenarios. After the introduction of natural language descriptions, different search targets can be changed accordingly, and the search and identification actions can be made more efficient and accurate by providing different accident descriptions.
[0007] Existing natural language tracking algorithms usually solve the two tasks of visual grounding and tracking in two steps, and deploy independent grounding models and tracking models to implement these two steps respectively. There are the following technical problems:
[0008] 1. The separation framework ignores the connection between visual positioning and tracking. Natural language description provides a global basis for positioning in these two steps. That is, the separation framework regards visual positioning and tracking as two independent processes, which leads to an information gap between the two. This gap hinders the system from continuously and accurately understanding the target object, resulting in insufficient long-term target tracking capability, poor adaptability to dynamic environments, and insufficient use of multimodal information.
[0009] 2. The separated framework is difficult to train end-to-end, resulting in a large number of parameters. Summary of the invention
[0010] In response to the above-mentioned research problems, the purpose of the present invention is to provide a method and system for unmanned aerial vehicle multimodal feature fusion target tracking based on natural language description, so as to solve the problem that the prior art easily causes poor target tracking ability and long-term tracking ability when the image quality collected by the unmanned aerial vehicle is poor or the image features are not obvious.
[0011] In order to achieve the above object, the present invention adopts the following technical solution:
[0012] A multi-modal feature fusion target tracking method for unmanned aerial vehicles based on natural language description includes the following steps:
[0013] Step 1: Obtain a video of a traffic accident scene from the perspective of a drone, convert the video into an image from the perspective of a drone, annotate the image, and describe the traffic accident scene in the image in natural language to obtain language prompts;
[0014] Step 2: Construct a scene-context feature pyramid network to perform context information enhancement on the image from the drone's perspective in step 1 to obtain a feature-enhanced image;
[0015] Step 3, performing visual encoding and language encoding on the image enhanced in step 2 and the language prompt obtained in step 1, respectively, to obtain visual features and language feature vectors;
[0016] Step 4: Process the visual features into visual feature vectors and perform local alignment of visual-linguistic bimodal features with the corresponding language feature vectors;
[0017] Step 5: Fully integrate the aligned new language features obtained in step 4 with the visual features obtained in step 3 to obtain multimodal features;
[0018] Step 6: If the image from the current drone perspective is the first frame, the pre-trained historical features of traffic accidents are input into the decoder for decoding, and the decoded results are passed through the positioning head to obtain the final tracking result, and then go to step 7; otherwise, the tracking result obtained in the previous frame is used as the historical feature and is input into the decoder for decoding with the multimodal features of the previous frame, and then the decoded results are passed through the positioning head to obtain the final tracking result, and then go to step 7, where the final tracking result includes the target frame;
[0019] Step 7, when the image from the current drone's perspective is not the last frame, after tracking each frame based on the tracking results, the target area of the multimodal features is pooled according to the target frame to obtain the target area features, and then the target area features are planarized to obtain the historical features, and then go to step 1 to track the next frame, otherwise, the tracking is terminated.
[0020] Furthermore, the scene-context feature pyramid network constructed in step 2 is a multi-scale feature extraction network based on the BiFPN framework, which extracts multi-scale features through the backbone network ResNet50, and then uses the bidirectional feature fusion path of the BiFPN structure to obtain a feature-enhanced image;
[0021] The pre-trained ResNet50 is used as a feature extractor to extract multi-scale features from the input image. Four feature maps of different scales are obtained through the four layers of ResNet50, which are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 respectively. Then, 1x1 convolution is used to adjust the number of channels of the four feature maps of different scales extracted by ResNet50 to a unified number of channels of 512.
[0022] BiFPN includes top-down path feature fusion and bottom-up path feature fusion, and adds the feature maps obtained by the top-down path feature fusion and the bottom-up path feature fusion to obtain the final feature map, and converts the final feature map into the required output channel number 3 through a 3x3 convolution layer, and outputs the feature-enhanced image;
[0023] BiFPN receives feature maps of four different scales from ResNet50 and arranges them into four levels: E1, E2, E3, and E4 from large to small scales;
[0024] The top-down path feature fusion is to upsample the feature map of E4 to the same size as the feature map of E3, and then weightedly fuse the upsampled feature map of E4 with the feature map of E3 using the learnable weights self.weights1 through the ReLU activation function. The fused feature map is then passed down and performs the same weighted fusion operation with the feature maps of E2 and E1 in turn.
[0025] The bottom-up path feature fusion downsamples the feature map of E1 to the same size as the feature map of E2, and performs weighted fusion of the downsampled feature map of E1 and the feature map of E2 using the learnable weights self.weights2 through the ReLU activation function. The fused feature map continues to be passed downward and performs the same weighted fusion operation with the feature maps of E3 and E4 in turn.
[0026] Furthermore, in step 3, the encoder Swin Transformer is used as a visual encoder to encode the enhanced image, and the specific steps are as follows:
[0027] The visual encoder processes the enhanced image and retains the output feature maps of the first three layers of Swin Transformer. Each layer of the output feature map is flattened into a sequence of shape (B, H*W, C). The output features are flattened using nn.Flatten and the feature dimensions are adjusted using a linear layer. The three layers of feature maps are added together to generate the final feature map, which is the visual feature. B is the batch size, H and W are the height and width of the feature map output by the first three layers of Swin Transformer, and C is the number of channels.
[0028] The language conversion model BERT is used as the language encoder to encode the natural language description. The specific steps are as follows:
[0029] Preprocessing:
[0030] For the natural language description, BERT's WordPiece tokenizer is used to segment the words and obtain a token sequence. A special token [CLS] is added at the beginning of the token sequence and a special token [SEP] is added at the end of the token sequence. Then, the token sequence is converted to the corresponding token ID and mapped according to the vocabulary of the BERT model. The token sequence includes multiple tokens, each of which includes a query vector, a key vector, and a value vector.
[0031] Embedding layer processing:
[0032] In the embedding layer, the token ID is converted into a word embedding vector. Using the pre-trained word embedding parameters of BERT, each token ID is mapped to a word embedding vector of fixed dimension. Then, a position embedding vector corresponding to its position is added to each token in the token sequence. The final input representation is the sum of the word embedding vector and the position embedding vector, which constitutes the initial feature vector sent to the encoder layer.
[0033] Encoding layer:
[0034] Each encoder layer includes a self-attention mechanism and a feedforward neural network. In the self-attention mechanism, after the embedding layer is processed, the attention score of each token and the remaining tokens in the token sequence is calculated through the query vector and the key vector and normalized into weights. The value vector of the token is weighted and summed according to the weight, and the representation of each token is updated to capture the long-distance dependency in the token sequence. The feedforward neural network performs nonlinear transformation on the output of the self-attention mechanism.
[0035] The final output shows:
[0036] Through the above steps, BERT outputs a language feature vector for each token in the token sequence processed by the encoding layer.
[0037] Further, the specific steps of step 4 are:
[0038] First, the visual features obtained in step 3 retain the batch and channel information, and standardize the data through the StandardScaler() function, divide the feature maps of different batches and channels into different samples, perform GMM clustering, define four cluster centers for each sample, obtain the four cluster centers and convert them into feature vectors, and generate several visual feature vectors on different batches and channels through this step;
[0039] Then, the visual feature vector and the language feature vector output from step 3 are mapped to a predefined common dimension through two independent fully connected layers respectively;
[0040] After dimension matching, for each pair of visual feature vectors and language feature vectors, calculate their dot product to get a similarity score, and combine the similarity scores into an attention matrix S of shape (num_visual_features, num_text_embeddings), where num_visual_features refers to the number of visual feature vectors, num_text_embeddings refers to the number of language feature vectors, and the attention matrix S represents the pairwise similarity scores between all visual feature vectors and all language feature vectors. Each value in the attention matrix reflects the strength of the association between the visual feature vector and the language feature vector; each row of the attention matrix represents a visual feature vector, and each column represents a language feature vector;
[0041] The softmax function in the soft attention method is applied to normalize each column in the attention matrix S to generate a weight distribution. Then, the language feature vectors are weighted summed according to the weights to obtain a new language feature corresponding to the visual feature, that is, the local alignment of the visual-language bimodal features is obtained.
[0042] Further, the specific steps of step 5 are:
[0043] Based on the local alignment of the visual-language bimodal features, the aligned new language features are added to the original visual features generated in step 3 through residual connections, and then layer normalization is performed. The features after layer normalization are nonlinearly transformed through a position feedforward neural network, where the position feedforward neural network includes two linear layers and an activation function ReLU, and the specific structure is:
[0044] FFN(x)=Linear(ReLU(Linear(x))), where FFN represents a position feedforward neural network, Linear represents a linear layer, ReLU represents an activation function, and x represents input data;
[0045] Based on the output of the nonlinear transformation, the output of the position nonlinear transformation is added to the features after the first residual processing through the residual connection again, and then the layer normalization is performed, and finally the multimodal features are output, that is, the fused features are obtained.
[0046] Furthermore, the framework of the decoder in step 6 is an improved transformer decoder, and the basic architecture includes a self-attention module, a residual connection layer and a feedforward neural network layer, and a channel attention module is added between the residual connection layer and the feedforward neural network layer;
[0047] The channel attention module includes an input layer that receives the feature map transmitted from the previous layer, a global average pooling layer that compresses the spatial information of each channel of the feature map input to the input layer to generate a channel description vector, a first fully connected layer that reduces the number of output channels of the global average pooling layer to 1 / 16 of the original ratio, an activation function layer that performs linear processing on the output of the first fully connected layer, a second fully connected layer that restores the channels of the result output by the activation function layer to the original number of channels, a sigmoid function layer that generates a channel weight vector from the result output by the second fully connected layer, and an output layer that multiplies the channel weight vector with the feature map input by the input layer to obtain an output feature map; the positioning head is used to convert the decoder obtained The feature vector is mapped to the bounding box parameters of the target, that is, the positioning head uses a 1x1 convolution kernel to reduce the number of channels in the target area to 1, generates a score map, and obtains a probability score for whether each position is a target corner point. The bounding box of the target is predicted based on the score map, and four values are output, corresponding to the center coordinates (x, v) and width and height (w, h) of the bounding box. The positioning head parameters are updated by back propagation by minimizing the difference between the bounding box and the true box using the IoU loss function to obtain the final tracking result. The similarity metrics of the score map include dot product and cosine similarity. The final tracking result includes the target box, where the target represents the traffic accident location obtained in the entire traffic scene. Further, in step 7, the target area of the multimodal feature of the frame image is pooled with the target frame according to the target frame, which is to expand the target frame in the target area of the multimodal feature by 1.5, that is, the width and height of the target frame are respectively expanded by 1.5 times, and the adjusted target frame is converted from the center coordinate and width and height format to the bounding box format, and then the ROI extraction method is used to extract the target area feature from the feature map in the adjusted bounding box, wherein the formula of the ROI extraction method is:
[0048]
[0049] Among them, RoIAlign is the layer that performs the RoI Align operation, which is used to calculate the pixel value of the floating point coordinates through bilinear interpolation, "(6, 6)" represents the output size, that is, each RoI is sampled into a 6x6 feature map, spatial_scale represents the scale factor used to adjust the RoI coordinates, that is, the ratio of the original size of the feature map in the bounding box to the size of the feature map after the RoI Align operation, self.divisor represents the downsampling multiple, sampling_ratio represents the number of sampling points used in the RoI, and setting it to 2 means sampling 2 points in the height and width directions respectively. A multimodal feature fusion target tracking system for unmanned aerial vehicles based on natural language description, including:
[0050] Information acquisition module: obtain the traffic accident scene video from the drone's perspective, convert the video into an image from the drone's perspective, annotate the image, and use natural language to describe the traffic accident scene in the image to obtain language prompts;
[0051] BiFPN feature pyramid module: builds a scene-context feature pyramid network to enhance the context information of the image from the drone's perspective to obtain the feature-enhanced image;
[0052] Encoding module: Perform visual encoding and language encoding on the enhanced image and language prompts respectively to obtain visual features and language feature vectors;
[0053] Multimodal feature alignment module: processes visual features into visual feature vectors and performs local alignment of visual-linguistic bimodal features with the corresponding language feature vectors;
[0054] Multimodal feature fusion module: fully integrates the aligned new language features with the visual features to obtain multimodal features;
[0055] Decoder module: If the image from the current drone perspective is the first frame, the pre-trained historical features of traffic accidents are input into the decoder for decoding; otherwise, the tracking result obtained in the previous frame is used as the historical feature and is input into the decoder for decoding together with the multimodal features of the previous frame;
[0056] Positioning and tracking module: used to obtain the final tracking result through the positioning head after decoding, wherein the final tracking result includes the target frame;
[0057] Historical feature embedding generation module: When the image from the current drone perspective is not the last frame, after tracking each frame based on the tracking results, the target area of the multimodal features is pooled according to the target frame to obtain the target area features, and then the target area features are planarized to obtain historical features.
[0058] Compared with the prior art, the present invention has the following beneficial effects:
[0059] First, when the present invention incorporates natural language description into target tracking decoding, the description provides high-level semantic information about the target object, including the state of the object and its relationship with other objects. This global semantic information is very useful for locating targets in a specific environment, especially in complex or ambiguous traffic accident scenes, which improves the long-term target tracking capability, dynamic environment adaptability, and makes full use of multimodal information and target discrimination capabilities;
[0060] Second, the present invention combines visual language grounding (in visual language object tracking, it refers to the process of matching the text description with the actual object position in the image or video frame) with the tracking method to establish an end-to-end training framework. The same network is used to implement the two tasks of visual language grounding and tracking, making the network lightweight and reducing the number of parameters.
[0061] Third, the present invention aims at the existing task of visual target tracking from the perspective of drones, and proposes a scene-context feature pyramid network for images from the perspective of drones, thereby improving the quality of drone images for deep learning algorithms;
[0062] Fourth, the present invention introduces rich semantic information described in natural language based on visual features, which can be used to improve the target discrimination ability and long-term tracking ability of the target tracker, thereby improving the vision-language dual-modality single target tracking task;
[0063] 5. In order to fully integrate the two modal information of visual features and language features, the present invention first aligns the visual and language features, so that the generated new language features have enhanced information of image features, thereby improving the ability of language features to understand the current visual scene;
[0064] 6. The present invention improves the decoder structure and allocates attention weights between different channels to emphasize the channel features that are more important to the task, so that it can perform better on the traffic accident dataset;
[0065] 7. The drone of the present invention is equipped with a target tracking algorithm based on natural language description, which can change the drone search task while switching the natural language instruction description, thereby improving the intelligence of drone search and tracking, and can cover more tasks that are difficult to complete by manpower;
[0066] 8. In the task of traffic accident identification, the present invention finds that traffic accident data often show a long-tail distribution, and the visual gap within the same category is still large. It is difficult to summarize the different characteristics between accidents by relying solely on a single accident category. After the introduction of natural language description, the referenced target can be located based on a given visual language reference. In the task of searching for traffic accidents using drones, the present invention can provide relevant accident descriptions set in advance, making traffic accident identification actions more efficient and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is a schematic diagram of the overall framework of the present invention;
[0068] Figure 2Schematic diagram of the BiFPN structure in the present invention, where P3, P4, P5, and P6 are feature maps of different scales extracted from ResNet50, F4 and F5 are saved intermediate results, D3, D4, D5, and D6 are the outputs of the four feature maps of different scales after a bidirectional fusion process, and the arrows indicate the direction of data fusion;
[0069] Figure 3 Schematic diagram of the decoder architecture in the present invention, where Add&Norm represents the residual connection layer, Feedforward represents the feedforward neural network layer, SEAttention represents the channel attention module, MHAttention represents the multi-head attention module, and Q represents the historical result feature embedding;
[0070] Figure 4 A schematic diagram of the structure of the Swin Transformer visual encoder in the present invention that stores the first three layers of feature maps;
[0071] Figure 5 It is a structural diagram of the Bert encoder module in the present invention. DETAILED DESCRIPTION
[0072] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.
[0073] In the existing technology, people usually rely on the naked eye observation of the images taken by drones to find the target, which is time-consuming and laborious. By installing a target tracking algorithm on the drone, real-time tracking of all categories in the current drone shooting scene can be achieved, but it is difficult to convey the target that people really want to track.
[0074] The method for unmanned aerial vehicle multimodal feature fusion target tracking based on natural language description described in the present invention can help the unmanned aerial vehicle complete highly targeted target tracking under the guidance of natural language.
[0075] A multi-modal feature fusion target tracking method for unmanned aerial vehicles based on natural language description includes the following steps:
[0076] Step 1: Obtain a video of a traffic accident scene from the perspective of a drone, convert the video into an image from the perspective of a drone, annotate the image, and describe the traffic accident scene in the image in natural language to obtain language prompts;
[0077] Step 2: Construct a scene-context feature pyramid network to perform context information enhancement on the image from the drone's perspective to obtain a feature-enhanced image;
[0078] In images taken from drone perspectives, complex backgrounds, vertical perspectives, and changes in target type and size make target tracking a challenging task. In this work, considering that the type of an object is usually closely related to the scene it is in, a scene-context feature pyramid network is proposed for object detection. This network aims to strengthen the relationship between the target and the scene and solve the problem caused by the change in target size.
[0079] The constructed scene-context feature pyramid network is a multi-scale feature extraction network based on the BiFPN framework. The network extracts multi-scale features through the backbone network ResNet50, and then uses the bidirectional feature fusion path of the BiFPN structure to obtain the feature-enhanced image.
[0080] The pre-trained ResNet50 is used as a feature extractor to extract multi-scale features from the input image. Four feature maps of different scales are obtained through the four layers of ResNet50, which are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 respectively. Then, 1x1 convolution is used to adjust the number of channels of the four feature maps of different scales extracted by ResNet50 to a unified number of channels of 512.
[0081] BiFPN includes top-down path feature fusion and bottom-up path feature fusion, and adds the feature maps obtained by the top-down path feature fusion and the bottom-up path feature fusion to obtain the final feature map, and converts the final feature map into the required output channel number 3 through a 3x3 convolution layer, and outputs the feature-enhanced image;
[0082] BiFPN receives feature maps of four different scales from ResNet50 and arranges them into four levels: E1, E2, E3, and E4 from large to small scales;
[0083] The top-down path feature fusion is to upsample the feature map of E4 to the same size as the feature map of E3, and then weightedly fuse the upsampled feature map of E4 with the feature map of E3 using the learnable weights self.weights1 through the ReLU activation function. The fused feature map is then passed down and performs the same weighted fusion operation with the feature maps of E2 and E1 in turn.
[0084] The bottom-up path feature fusion downsamples the feature map of E1 to the same size as the feature map of E2, and performs weighted fusion of the downsampled feature map of E1 and the feature map of E2 using the learnable weights self.weights2 through the ReLU activation function. The fused feature map continues to be passed downward and performs the same weighted fusion operation with the feature maps of E3 and E4 in turn.
[0085] Step 3, performing visual encoding and language encoding on the image enhanced in step 2 and the language prompt obtained in step 1, respectively, to obtain visual features and language feature vectors;
[0086] The encoder Swin Transformer is used as the visual encoder to encode the enhanced image. The specific steps are as follows:
[0087] The visual encoder processes the enhanced image and retains the output feature maps of the first three layers of Swin Transformer. Each layer of the output feature map is flattened into a sequence of shape (B, H*W, C). The output features are flattened using nn.Flatten and the feature dimensions are adjusted using a linear layer. The three layers of feature maps are added together to generate the final feature map, which is the visual feature. B is the batch size, H and W are the height and width of the feature map output by the first three layers of Swin Transformer, and C is the number of channels.
[0088] The language conversion model BERT is used as the language encoder to encode the natural language description. The specific steps are as follows:
[0089] Preprocessing:
[0090] For the natural language description, BERT's WordPiece tokenizer is used to segment the words to obtain a token sequence. A special token [CLS] is added at the beginning of the token sequence, and a special token [SEP] is added at the end of the token sequence. Then, the token sequence is converted to the corresponding token ID and mapped according to the vocabulary of the BERT model. The token sequence includes multiple tokens, each of which includes a query vector, a key vector, and a value vector. For example, for the natural language description "A white carcollided with a blue truck.", BERT's WordPiece tokenizer is used to segment the words to obtain a token sequence ["A", "white", "car", "collided", "with", "a", "blue", "truck", "."]. A special token [CLS] is added at the beginning of the sequence, and a special token [SEP] is added at the end of the sequence. Then, the token sequence is converted to the corresponding token ID and mapped according to the vocabulary of the BERT model.
[0091] Embedding layer processing:
[0092] In the embedding layer, the token ID is converted into a word embedding vector. Using the word embedding parameters pre-trained by BERT, each token ID is mapped to a word embedding vector of fixed dimension. Then, a position embedding vector corresponding to its position is added to each token in the token sequence. The final input representation is the sum of the word embedding vector and the position embedding vector, which constitutes the initial feature vector sent to the encoder layer. For example, in the embedding layer, the token ID is converted into a word embedding vector. Using the word embedding parameters pre-trained by BERT, each token ID is mapped to a word embedding vector of fixed dimension. For example, the token ID corresponding to "white" in the example is mapped to a 768-dimensional word embedding vector. Then, a vector related to its position is added to each token, so that the model can capture the order information in the sequence.
[0093] Encoding layer:
[0094] Each encoder layer includes a self-attention mechanism and a feedforward neural network. In the self-attention mechanism, after the embedding layer is processed, the attention scores of each token and the remaining tokens in the token sequence are calculated through the query vector and the key vector and normalized into weights. The value vectors of the tokens are weighted and summed according to the weights, and the representation of each token is updated to capture the long-distance dependencies in the token sequence. The feedforward neural network performs a nonlinear transformation on the output of the self-attention mechanism; for example, the "collided" token may have a higher attention weight for "white car" and "blue truck" because they are closely related to the collision event; the feedforward neural network performs a nonlinear transformation on the output of the self-attention mechanism to further enhance the expressive power of the model.
[0095] In the sentence “A car collided with a truck.”, each token (such as “car”) is first converted into three vectors: query vector (Query), key vector (Key), and value vector (Value) according to the self-attention mechanism in the BERT encoder. Next, the attention score between the token and all other tokens in the sequence (including itself) is calculated, and the Softmax function is used to convert these scores into weights to ensure that they add up to 1 to form a probability distribution. Subsequently, these weights are used to perform a weighted summation of the value vectors of the token to generate a new representation of the token. This new representation incorporates the information of the entire sentence and captures richer semantic relationships and long-distance dependencies. The new representation of “car” integrates information from “collided”, “with”, “a”, “truck”, and itself through the above process, with special emphasis on the association with “truck” to better reflect its meaning in the context.
[0096] The final output shows:
[0097] Through the above steps, BERT outputs a language feature vector for each token in the token sequence processed by the encoding layer.
[0098] Step 4: Process the visual features into visual feature vectors and perform local alignment of visual-linguistic bimodal features with the corresponding language feature vectors; the specific steps are:
[0099] First, the visual features obtained in step 3 retain the batch and channel information, and standardize the data through the StandardScaler() function, divide the feature maps of different batches and channels into different samples, perform GMM clustering, define four cluster centers for each sample, obtain the four cluster centers and convert them into feature vectors, and generate several visual feature vectors on different batches and channels through this step;
[0100] Then, the visual feature vector and the language feature vector output from step 3 are mapped to a predefined common dimension through two independent fully connected layers respectively;
[0101] After dimension matching, for each pair of visual feature vectors and language feature vectors, calculate their dot product to get a similarity score, and combine the similarity scores into an attention matrix S of shape (num_visual_features, num_text_embeddings), where num_visual_features refers to the number of visual feature vectors, num_text_embeddings refers to the number of language feature vectors, and the attention matrix S represents the pairwise similarity scores between all visual feature vectors and all language feature vectors. Each value in the attention matrix reflects the strength of the association between the visual feature vector and the language feature vector; each row of the attention matrix represents a visual feature vector, and each column represents a language feature vector;
[0102] The softmax function in the soft attention method is applied to normalize each column in the attention matrix S to generate a weight distribution. Then, the language feature vectors are weighted summed according to the weights to obtain a new language feature corresponding to the visual feature, that is, the local alignment of the visual-language bimodal features is obtained.
[0103] Step 5: Fully integrate the aligned new language features obtained in step 4 with the visual features obtained in step 3 to obtain multimodal features. The specific steps are:
[0104] Based on the local alignment of the visual-language bimodal features, the aligned new language features are added to the original visual features generated in step 3 through residual connections, and then layer normalization is performed. The features after layer normalization are nonlinearly transformed through a position feedforward neural network, where the position feedforward neural network includes two linear layers and an activation function ReLU, and the specific structure is:
[0105] FFN(x)=Linear(ReLU(Linear(x))), where FFN represents a position feedforward neural network, Linear represents a linear layer, ReLU represents an activation function, and x represents input data;
[0106] Based on the output of the nonlinear transformation, the output of the nonlinear transformation is added to the features after the first residual processing through the residual connection again, and then the layer normalization is performed, and finally the multimodal features are output, that is, the fused features are obtained.
[0107] Step 6: If the image from the current drone perspective is the first frame, the pre-trained historical features of traffic accidents are input into the decoder for decoding, and the decoded results are passed through the positioning head to obtain the final tracking result, and then go to step 7; otherwise, the tracking result obtained in the previous frame is used as the historical feature and is input into the decoder for decoding with the multimodal features of the previous frame, and then the decoded results are passed through the positioning head to obtain the final tracking result, and then go to step 7, where the final tracking result includes the target frame;
[0108] The decoder framework is an improved transformer decoder. The basic architecture includes a self-attention module, a residual connection layer, and a feedforward neural network layer. A channel attention module is added between the residual connection layer and the feedforward neural network layer.
[0109] The channel attention module includes an input layer that receives the feature map transmitted from the previous layer, a global average pooling layer that compresses the spatial information of each channel of the feature map input to the input layer to generate a channel description vector, a first fully connected layer that reduces the number of output channels of the global average pooling layer to 1 / 16 of the original ratio, an activation function layer that linearly processes the output of the first fully connected layer, a second fully connected layer that restores the channels of the result output by the activation function layer to the original number of channels, a sigmoid function layer that generates a channel weight vector from the result output by the second fully connected layer, and an output layer that multiplies the channel weight vector with the feature map input by the input layer to obtain an output feature map;
[0110] Through the multi-layer attention mechanism, the decoder can interact the query vector, which refers to the historical features in this paper, with the multimodal features output by the encoder to generate a target-specific feature vector for the query vector, corresponding to a possible target. This feature vector is then fed into the localization head for specific bounding box prediction.
[0111] The localization head is used to map the feature vector obtained by the decoder into the bounding box parameters of the target, that is, the localization head uses a 1x1 convolution kernel to reduce the number of channels in the target area to 1, generates a score map, and obtains a probability score for whether each position is a target corner point. The bounding box of the target is predicted based on the score map, and four values are output, corresponding to the center coordinates (x, y) and width and height (w, h) of the bounding box. The localization head parameters are updated by back propagation by minimizing the difference between the bounding box and the true box using the IoU loss function to obtain the final tracking result. The similarity metrics of the score map include dot product and cosine similarity. The final tracking result includes the target box, where the target represents the location of the traffic accident obtained in the entire traffic scene.
[0112] Step 7, when the image from the current drone's perspective is not the last frame, after tracking each frame based on the tracking results, the target area of the multimodal features is pooled according to the target frame to obtain the target area features, and then the target area features are planarized to obtain the historical features, and then go to step 1 to track the next frame, otherwise, the tracking is terminated.
[0113] The target area of the multimodal feature of the frame image is pooled with the target frame according to the target frame, which is to expand the target frame in the target area of the multimodal feature by 1.5, that is, the width and height of the target frame are expanded by 1.5 times respectively, and the adjusted target frame is converted from the center coordinate and width and height format to the bounding box format, and then the R0I extraction method is used to extract the target area features from the feature map in the adjusted bounding box. The formula of the R0I extraction method is:
[0114]
[0115] Among them, RoIAlign is the layer that performs the RoI Align operation, which is used to calculate the pixel values of floating-point coordinates through bilinear interpolation. "(6, 6)" represents the output size, that is, each RoI is sampled into a 6x6 feature map, spatial_scale represents the scale factor used to adjust the RoI coordinates, that is, the ratio of the original size of the feature map in the bounding box to the size of the feature map after the RoI Align operation, self.divisor represents the downsampling multiple, sampling_ratio represents the number of sampling points used in the RoI, and setting it to 2 means sampling 2 points each in the height and width directions.
[0116] The above are only representative embodiments of the present invention in many specific application scopes, and do not constitute any limitation on the protection scope of the present invention. Any technical solutions formed by transformation or equivalent replacement fall within the protection scope of the present invention.
Claims
1. A multi-modal feature fusion target tracking method for unmanned aerial vehicles based on natural language description, characterized in that: The steps include: Step 1: Obtain a video of a traffic accident scene from the perspective of a drone, convert the video into an image from the perspective of a drone, annotate the image, and describe the traffic accident scene in the image in natural language to obtain language prompts; Step 2: Construct a scene-context feature pyramid network to perform context information enhancement on the image from the drone's perspective in step 1 to obtain a feature-enhanced image; Step 3, performing visual encoding and language encoding on the image enhanced in step 2 and the language prompt obtained in step 1, respectively, to obtain visual features and language feature vectors; Step 4: Process the visual features into visual feature vectors and perform local alignment of visual-linguistic bimodal features with the corresponding language feature vectors; Step 5: Fully integrate the aligned new language features obtained in step 4 with the visual features obtained in step 3 to obtain multimodal features; Step 6: If the image from the current drone perspective is the first frame, the pre-trained historical features of traffic accidents are input into the decoder for decoding, and the decoded results are passed through the positioning head to obtain the final tracking result, and then go to step 7; otherwise, the tracking result obtained in the previous frame is used as the historical feature and is input into the decoder for decoding with the multimodal features of the previous frame, and then the decoded results are passed through the positioning head to obtain the final tracking result, and then go to step 7, where the final tracking result includes the target frame; Step 7, when the image from the current drone's perspective is not the last frame, after tracking each frame based on the tracking results, the target area of the multimodal features is pooled according to the target frame to obtain the target area features, and then the target area features are planarized to obtain the historical features, and then go to step 1 to track the next frame, otherwise, the tracking is terminated.
2. The method for tracking a target using multi-modal features of a drone based on natural language description according to claim 1 is characterized in that: The scene-context feature pyramid network constructed in step 2 is a multi-scale feature extraction network based on the BiFPN framework, which extracts multi-scale features through the backbone network ResNet50, and then uses the bidirectional feature fusion path of the BiFPN structure to obtain a feature-enhanced image; The pre-trained ResNet50 is used as a feature extractor to extract multi-scale features from the input image. Four feature maps of different scales are obtained through the four layers of ResNet50, which are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 respectively. Then, 1x1 convolution is used to adjust the number of channels of the four feature maps of different scales extracted by ResNet50 to a unified number of channels of 512. BiFPN includes top-down path feature fusion and bottom-up path feature fusion, and adds the feature maps obtained by the top-down path feature fusion and the bottom-up path feature fusion to obtain the final feature map, and converts the final feature map into the required output channel number 3 through a 3x3 convolution layer, and outputs the feature-enhanced image; BiFPN receives feature maps of four different scales from ResNet50 and arranges them into four levels: E1, E2, E3, and E4 from large to small scales; The top-down path feature fusion is to upsample the feature map of E4 to the same size as the feature map of E3, and then weightedly fuse the upsampled feature map of E4 with the feature map of E3 using the learnable weights self.weights1 through the ReLU activation function. The fused feature map is then passed down and performs the same weighted fusion operation with the feature maps of E2 and E1 in turn. The bottom-up path feature fusion downsamples the feature map of E1 to the same size as the feature map of E2, and performs weighted fusion of the downsampled feature map of E1 and the feature map of E2 using the learnable weights self.weights2 through the ReLU activation function. The fused feature map continues to be passed downward and performs the same weighted fusion operation with the feature maps of E3 and E4 in turn.
3. The method for tracking a target using multi-modal features of a drone based on natural language description according to claim 2 is characterized in that: In step 3, the encoder Swin Transformer is used as a visual encoder to encode the enhanced image. The specific steps are: The visual encoder processes the enhanced image and retains the output feature maps of the first three layers of Swin Transformer. Each layer of the output feature map is flattened into a sequence of shape (B, H*W, C). The output features are flattened using nn.Flatten and the feature dimensions are adjusted using a linear layer. The three layers of feature maps are added together to generate the final feature map, which is the visual feature. B is the batch size, H and W are the height and width of the feature map output by the first three layers of Swin Transformer, and C is the number of channels. The language conversion model BERT is used as the language encoder to encode the natural language description. The specific steps are as follows: Preprocessing: For the natural language description, BERT's WordPiece tokenizer is used to segment the words and obtain a token sequence. A special token [CLS] is added at the beginning of the token sequence and a special token [SEP] is added at the end of the token sequence. Then, the token sequence is converted to the corresponding token ID and mapped according to the vocabulary of the BERT model. The token sequence includes multiple tokens, each of which includes a query vector, a key vector, and a value vector. Embedding layer processing: In the embedding layer, the token ID is converted into a word embedding vector. Using the pre-trained word embedding parameters of BERT, each token ID is mapped to a word embedding vector of fixed dimension. Then, a position embedding vector corresponding to its position is added to each token in the token sequence. The final input representation is the sum of the word embedding vector and the position embedding vector, which constitutes the initial feature vector sent to the encoder layer. Encoding layer: Each encoder layer includes a self-attention mechanism and a feedforward neural network. In the self-attention mechanism, after the embedding layer is processed, the attention score of each token and the remaining tokens in the token sequence is calculated through the query vector and the key vector and normalized into weights. The value vector of the token is weighted and summed according to the weight, and the representation of each token is updated to capture the long-distance dependency in the token sequence. The feedforward neural network performs nonlinear transformation on the output of the self-attention mechanism. The final output shows: Through the above steps, BERT outputs a language feature vector for each token in the token sequence processed by the encoding layer.
4. The method for tracking a target using multi-modal features of a drone based on natural language description according to claim 3 is characterized in that: The specific steps of step 4 are: First, the visual features obtained in step 3 retain the batch and channel information, and standardize the data through the StandardScaler() function, divide the feature maps of different batches and channels into different samples, perform GMM clustering, define four cluster centers for each sample, obtain the four cluster centers and convert them into feature vectors, and generate several visual feature vectors on different batches and channels through this step; Then, the visual feature vector and the language feature vector output from step 3 are mapped to a predefined common dimension through two independent fully connected layers respectively; After dimension matching, for each pair of visual feature vectors and language feature vectors, calculate their dot product to get a similarity score, and combine the similarity scores into an attention matrix S of shape (num_visual_features, num_text_embeddings), where num_visual_features refers to the number of visual feature vectors, num_text_embeddings refers to the number of language feature vectors, and the attention matrix S represents the pairwise similarity scores between all visual feature vectors and all language feature vectors. Each value in the attention matrix reflects the strength of the association between the visual feature vector and the language feature vector; each row of the attention matrix represents a visual feature vector, and each column represents a language feature vector; The softmax function in the soft attention method is applied to normalize each column in the attention matrix S to generate a weight distribution. Then, the language feature vectors are weighted summed according to the weights to obtain a new language feature corresponding to the visual feature, that is, the local alignment of the visual-language bimodal features is obtained.
5. The method for tracking a target using multi-modal features of a drone based on natural language description according to claim 3 is characterized in that: The specific steps of step 5 are: Based on the local alignment of the visual-language bimodal features, the aligned new language features are added to the original visual features generated in step 3 through residual connections, and then layer normalization is performed. The features after layer normalization are nonlinearly transformed through a position feedforward neural network, where the position feedforward neural network includes two linear layers and an activation function ReLU, and the specific structure is: FFN(x)=Linear(ReLU(Linear(x))), where FFN represents a position feedforward neural network, Linear represents a linear layer, ReLU represents an activation function, and x represents input data; Based on the output of the nonlinear transformation, the output of the position nonlinear transformation is added to the features after the first residual processing through the residual connection again, and then the layer normalization is performed, and finally the multimodal features are output, that is, the fused features are obtained.
6. The method for tracking a target using multi-modal features of a drone based on natural language description according to claim 3 is characterized in that: The framework of the decoder in step 6 is an improved transformer decoder, and the basic architecture includes a self-attention module, a residual connection layer and a feedforward neural network layer, and a channel attention module is added between the residual connection layer and the feedforward neural network layer; The channel attention module includes an input layer that receives the feature map transmitted from the previous layer, a global average pooling layer that compresses the spatial information of each channel of the feature map input to the input layer to generate a channel description vector, a first fully connected layer that reduces the number of output channels of the global average pooling layer to 1 / 16 of the original ratio, an activation function layer that performs linear processing on the output of the first fully connected layer, a second fully connected layer that restores the channels of the result output by the activation function layer to the original number of channels, a sigmoid function layer that generates a channel weight vector from the result output by the second fully connected layer, and an output layer that multiplies the channel weight vector with the feature map input by the input layer to obtain an output feature map; the positioning head is used to convert the decoder obtained The feature vector is mapped to the bounding box parameters of the target, that is, the positioning head uses a 1x1 convolution kernel to reduce the number of channels in the target area to 1, generates a score map, and obtains a probability score for whether each position is a target corner point. The bounding box of the target is predicted based on the score map, and four values are output, corresponding to the center coordinates (x, v) and width and height (w, h) of the bounding box. The positioning head parameters are updated by back propagation by minimizing the difference between the bounding box and the true box using the IoU loss function to obtain the final tracking result. The similarity metrics of the score map include dot product and cosine similarity. The final tracking result includes the target box, where the target represents the traffic accident location obtained in the entire traffic scene.
7. The method for tracking a target using multi-modal features of a drone based on natural language description according to claim 6 is characterized in that: In step 7, the target area of the multimodal feature of the frame image is pooled with the target frame by 1.5 bits in the target area of the multimodal feature, that is, the width and height of the target frame are respectively expanded by 1.5 times, and the adjusted target frame is converted from the center coordinate and width and height format to the bounding box format, and then the ROI extraction method is used to extract the target area feature from the feature map in the adjusted bounding box, wherein the formula of the ROI extraction method is: Among them, RoIAlign is the layer that performs the RoI Align operation, which is used to calculate the pixel values of floating-point coordinates through bilinear interpolation. "(6, 6)" represents the output size, that is, each RoI is sampled into a 6x6 feature map. Spatial_scale represents the scale factor used to adjust the RoI coordinates, that is, the ratio of the original size of the feature map in the bounding box to the size of the feature map after the RoI Align operation. Self.divisor represents the downsampling multiple. Sampling_ratio represents the number of sampling points used in the RoI. Setting it to 2 means sampling 2 points in the height and width directions respectively.
8. A multi-modal feature fusion target tracking system for unmanned aerial vehicles based on natural language description, characterized in that: include: Information acquisition module: obtain the traffic accident scene video from the drone's perspective, convert the video into an image from the drone's perspective, annotate the image, and use natural language to describe the traffic accident scene in the image to obtain language prompts; BiFPN feature pyramid module: builds a scene-context feature pyramid network to enhance the context information of the image from the drone's perspective to obtain the feature-enhanced image; Encoding module: Perform visual encoding and language encoding on the enhanced image and language prompts respectively to obtain visual features and language feature vectors; Multimodal feature alignment module: processes visual features into visual feature vectors and performs local alignment of visual-linguistic bimodal features with the corresponding language feature vectors; Multimodal feature fusion module: fully integrates the aligned new language features with the visual features to obtain multimodal features; Decoder module: If the image from the current drone perspective is the first frame, the pre-trained historical features of traffic accidents are input into the decoder for decoding; otherwise, the tracking result obtained in the previous frame is used as the historical feature and is input into the decoder for decoding together with the multimodal features of the previous frame; Positioning and tracking module: used to obtain the final tracking result through the positioning head after decoding, wherein the final tracking result includes the target frame; Historical feature embedding generation module: When the image from the current drone perspective is not the last frame, after tracking each frame based on the tracking results, the target area of the multimodal features is pooled according to the target frame to obtain the target area features, and then the target area features are planarized to obtain historical features.
Citation Information
Patent Citations
A multi-modal tracking method for unmanned aerial vehicles
CN117975314B
Natural language target tracking method based on Transform architecture
CN114372173A
Short-time natural language target tracking method based on visual language large model
CN117746024A
Video target tracking method based on language description
CN118115526A
Visual language target tracking method based on prompt learning
CN118279807A
Cited By
Visual language tracking method and system based on multi-scale adaptive cascade fusion
CN120563560A
Visual language tracking method and system based on multi-scale adaptive cascading fusion
CN120563560B
Image-fused scene adaptive video tracking system and method
CN120655675A
Water conservancy unmanned aerial vehicle inspection method based on visual language action multi-modal model
CN120704350A
Multi-view target tracking detection method and device, terminal and medium
CN120707596A