Visual language tracking method and system based on multi-scale adaptive cascading fusion

By using a multi-scale adaptive cascaded fusion tracking network model, the robustness and accuracy of visual-language tracking methods in complex scenarios are addressed, achieving high-precision tracking of targets.

CN120563560BActive Publication Date: 2025-11-28GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY +1

Patent Information

Application Number
CN202510698564.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-11-28
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing visual-language tracking methods lack robustness and accuracy in complex scenarios, and struggle to adapt to real-time changes in target state and multi-scale feature variations.

Method used

A multi-scale adaptive cascaded fusion method is adopted, which constructs a tracking network model based on multi-scale adaptive cascaded fusion through a multi-scale feature alignment module, a visual-language unified encoder, an adaptive cascaded fusion module, and a scale recovery module, so as to realize deep association and dynamic adaptation of visual and linguistic features.

Benefits of technology

It significantly improves tracking robustness and accuracy in complex scenarios, is suitable for visual-language collaborative tracking tasks in dynamic environments, and improves target tracking accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563560B_ABST
    Figure CN120563560B_ABST
Patent Text Reader

Abstract

The application discloses a visual language tracking method and system based on multi-scale adaptive cascade fusion, and the method comprises the following steps: pre-processing and text extraction processing are performed on a target video sequence to obtain pre-processed images and image text data; a multi-scale feature alignment module and a scale recovery module are introduced to construct a tracking network model based on multi-scale adaptive cascade fusion; and the tracking network model based on multi-scale adaptive cascade fusion is used for visual language tracking on the pre-processed images and image text data to obtain a target tracking result. Through multi-scale adaptive fusion and cross-modal alignment, the application can improve the tracking robustness and accuracy in a complex scene. The application can be widely applied to the technical field of computer vision tracking.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision tracking, and particularly relates to a visual language tracking method and system based on multi-scale adaptive cascade fusion. BACKGROUND

[0002] The core goal of target tracking is to continuously and accurately locate a specific target object in a video sequence. Early tracking methods are mainly based on manually designed features such as direction gradient histogram and color histogram, combined with traditional algorithms such as correlation filtering or particle filtering to realize target tracking. However, these methods have poor robustness in complex scenes and are difficult to cope with challenging problems such as target deformation and occlusion.

[0003] As a new research direction in the field of target tracking, visual-language tracking provides new solutions to the challenges faced by traditional visual tracking by introducing rich semantic information provided by natural language description. However, most existing visual-language tracking methods use simple feature splicing or shallow attention mechanisms for feature fusion, which fails to fully exploit the deep associations between visual and language modalities and is difficult to adapt to real-time state changes of the target in the video sequence. SUMMARY

[0004] To solve the above technical problems, the purpose of the present application is to provide a visual language tracking method and system based on multi-scale adaptive cascade fusion, which can improve the tracking robustness and accuracy in complex scenes through multi-scale adaptive fusion and cross-modal alignment.

[0005] The first technical solution adopted by the present application is: a visual language tracking method based on multi-scale adaptive cascade fusion, comprising the following steps:

[0006] Preprocessing and text extraction processing are performed on the target video sequence to obtain preprocessed images and image text data;

[0007] A multi-scale feature alignment module and a scale recovery module are introduced to construct a tracking network model based on multi-scale adaptive cascade fusion;

[0008] The tracking network model based on multi-scale adaptive cascade fusion performs visual language tracking on the preprocessed images and image text data to obtain the target tracking result.

[0009] Further, the step of preprocessing and text extraction processing the target video sequence to obtain preprocessed images and image text data specifically comprises:

[0010] The target video sequence is obtained and a template frame and a search frame are randomly selected according to the time sequence;

[0011] The template frame and the search frame are cropped to obtain a template image and a search region image;

[0012] The template image and the search region image are normalized and enhanced to obtain a preprocessed image;

[0013] Based on the target video sequence, its corresponding natural language description is extracted, and the text is standardized and marked to obtain image text data.

[0014] Further, the tracking network model based on multi-scale adaptive cascade fusion specifically includes a language encoder, a visual encoder, a multi-scale feature alignment module, a visual language unified encoder, an adaptive cascade fusion module, a scale recovery module, and a positioning head detection module, wherein the output ends of the language encoder and the visual encoder are connected with the input end of the multi-scale feature alignment module respectively, the multi-scale feature alignment module, the visual language unified encoder, the adaptive cascade fusion module, the scale recovery module, and the positioning head detection module are connected in sequence, and the loss function of the tracking network model based on multi-scale adaptive cascade fusion is specifically as follows:

[0015] L=L cls +λ iou L iou +λ L1 L1

[0016] L cls =-α t (1-p t ) γ log(p t )

[0017]

[0018] In the above formula, L represents the total loss function of the tracking network model based on multi-scale adaptive cascade fusion, L cls represents the focal loss function, L1 represents the L1 loss function, L iou represents the generalized intersection over union loss function, λ iou , represents an adjustable weight coefficient, α t represents a weight for balancing class imbalance, p t represents the probability of the predicted class, γ represents the focal parameter for adjusting the loss of easy-to-classify samples, y i represents the true value, represents the predicted value, n represents the number of samples, i represents the i-th sample, A represents the predicted frame, B represents the true frame, |A∩B| represents the intersection area of the predicted frame and the true frame, and |A∪B| represents the union area of the predicted frame and the true frame.

[0019] Further, the step of the tracking network model based on multi-scale adaptive cascaded fusion performing visual language tracking on the preprocessed image and image text data to obtain a target tracking result specifically comprises:

[0020] inputting the preprocessed image and image text data into the tracking network model based on multi-scale adaptive cascaded fusion;

[0021] performing language encoding processing on the image text data by the language encoder of the tracking network model based on multi-scale adaptive cascaded fusion to obtain image language feature information;

[0022] performing visual encoding processing on the preprocessed image by the visual encoder of the tracking network model based on multi-scale adaptive cascaded fusion to obtain image visual feature information;

[0023] aligning the image language feature information and the image visual feature information in multiple scales by the multi-scale feature alignment module of the tracking network model based on multi-scale adaptive cascaded fusion to obtain aligned image language feature information and aligned image visual feature information;

[0024] performing cross-modal feature interaction fusion on the aligned image language feature information and the aligned image visual feature information by the visual language unified encoder of the tracking network model based on multi-scale adaptive cascaded fusion to obtain unified cross-modal feature representation;

[0025] performing adaptive fusion on the unified cross-modal feature representation by the adaptive cascaded fusion module of the tracking network model based on multi-scale adaptive cascaded fusion to obtain multi-scale search region feature;

[0026] performing scale restoration processing on the multi-scale search region feature by the scale restoration module of the tracking network model based on multi-scale adaptive cascaded fusion to obtain search region feature;

[0027] performing tracking positioning on the search region feature by the positioning head detection module of the tracking network model based on multi-scale adaptive cascaded fusion to obtain a target tracking result.

[0028] Further, the step of the multi-scale feature alignment module of the tracking network model based on multi-scale adaptive cascaded fusion aligning the image language feature information and the image visual feature information in multiple scales to obtain aligned image language feature information and aligned image visual feature information specifically comprises:

[0029] inputting the image language feature information and the image visual feature information into the multi-scale feature alignment module of the tracking network model based on multi-scale adaptive cascaded fusion;

[0030] The image visual feature information is sequentially subjected to convolution and down-sampling operation processing through a convolution layer and a max-pooling layer to obtain multi-scale visual feature representations containing multiple spatial resolutions.

[0031] The multi-scale visual feature representations containing multiple spatial resolutions are subjected to dimension reduction processing to obtain dimension-reduced multi-scale visual feature representations.

[0032] The dimension-reduced multi-scale visual feature representations and the image language feature information are interactively fused through a cross-modal attention mechanism to obtain visual features and language features of each scale.

[0033] The visual features and language features of each scale are sequentially subjected to normalization processing, splicing processing and alignment operation to obtain aligned image language feature information and aligned image visual feature information.

[0034] Further, the visual language unified encoder of the tracking network model based on multi-scale adaptive cascaded fusion performs cross-modal feature interaction fusion on the aligned image language feature information and the aligned image visual feature information to obtain a unified cross-modal feature representation, which specifically includes:

[0035] The aligned image language feature information and the aligned image visual feature information are input into the visual language unified encoder of the multi-scale adaptive cascaded fusion tracking network model.

[0036] The aligned image language feature information and the aligned image visual feature information are respectively subjected to serialization processing to obtain serialized image language feature information and serialized image visual feature information.

[0037] The serialized image language feature information and the serialized image visual feature information are subjected to feature association based on a self-attention mechanism through a plurality of stacked Transformer modules to obtain a unified cross-modal feature representation.

[0038] Further, the adaptive cascaded fusion module of the tracking network model based on multi-scale adaptive cascaded fusion performs adaptive fusion on the unified cross-modal feature representation to obtain multi-scale search region features, which specifically includes:

[0039] The unified cross-modal feature representation is input into the adaptive cascaded fusion module of the multi-scale adaptive cascaded fusion tracking network model.

[0040] Based on a cascaded feature fusion strategy, the unified cross-modal feature representation is subjected to layer-by-layer calculation through an adaptive gating mechanism to obtain feature contribution weights.

[0041] According to the search area, the contribution weight of the feature is weighted and fused with the unified cross-modal feature representation to obtain multi-scale search area features.

[0042] Further, the scale recovery module of the tracking network model based on multi-scale adaptive cascaded fusion performs scale recovery processing on the multi-scale search area features to obtain search area features, which specifically includes:

[0043] The multi-scale search area features are input into the scale recovery module of the tracking network model based on multi-scale adaptive cascaded fusion.

[0044] Based on the cross-scale multi-head attention mechanism, the high-level abstract features and low-level detail features in the multi-scale search area features are processed for information interaction to obtain multi-scale search area features with an initial scale.

[0045] The multi-scale search area features with the initial scale are subjected to layer normalization operation, and a feedforward neural network is introduced for enhanced fusion features to obtain search area features.

[0046] Further, the positioning head detection module of the tracking network model based on multi-scale adaptive cascaded fusion performs tracking positioning on the search area features to obtain target tracking results, which specifically includes:

[0047] The search area features are input into the positioning head detection module of the tracking network model based on multi-scale adaptive cascaded fusion.

[0048] The search area features are converted to obtain two-dimensional search area features.

[0049] The two-dimensional search area features are stacked through multi-layer convolution, batch normalization and activation function, and classification score maps, bounding box size maps and offset maps are output.

[0050] The position with the highest confidence in the classification score map is selected as the target center, and the offset map and the bounding box size map are combined for calculation to obtain the target tracking result.

[0051] The second technical solution adopted by the present application is: a visual language tracking system based on multi-scale adaptive cascaded fusion, comprising:

[0052] The first module is used for pre-processing and text extraction processing of the target video sequence to obtain pre-processed images and image text data.

[0053] The second module is used for introducing a multi-scale feature alignment module and a scale recovery module to construct a tracking network model based on multi-scale adaptive cascaded fusion.

[0054] The third module is configured to perform visual language tracking on the preprocessed image and image text data based on the tracking network model based on multi-scale adaptive cascaded fusion to obtain a target tracking result.

[0055] The method and system have the following advantages: the target video sequence is preprocessed and text extraction is performed to obtain preprocessed image and image text data, a multi-scale feature alignment module and a scale recovery module are further introduced, a tracking network model based on multi-scale adaptive cascaded fusion is constructed, multi-scale adaptive fusion and cross-modal alignment are performed, the tracking robustness and accuracy in a complex scene are significantly improved, the method is suitable for visual language collaborative tracking tasks in a dynamic environment, the preprocessed image and image text data are tracked based on the tracking network model based on multi-scale adaptive cascaded fusion to obtain a target tracking result, and the tracking accuracy of the target is improved. BRIEF DESCRIPTION OF DRAWINGS

[0056] Figure 1 is a step flowchart of the visual language tracking method based on multi-scale adaptive cascaded fusion of the present application;

[0057] Figure 2 is a structural block diagram of the visual language tracking system based on multi-scale adaptive cascaded fusion of the present application;

[0058] Figure 3 is a structural diagram of the tracking network model based on multi-scale adaptive cascaded fusion provided by the embodiment of the present application. DETAILED DESCRIPTION

[0059] The present application will be further described in detail below in combination with the drawings and specific embodiments. For the step numbers in the following embodiments, only the setting is for the convenience of description, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0060] First of all, it needs to be pointed out that visual-linguistic tracking as a new research direction in the field of target tracking provides a new solution to the challenges faced by traditional visual tracking by introducing the rich semantic information provided by natural language description. The existing VLT methods can be mainly divided into three categories: early methods based on LSTM such as TNLS, which realizes the preliminary fusion of visual and language features through a recurrent neural network; improved methods based on attention mechanism such as JointNLT, which realizes more fine-grained feature interaction through cross-modal attention; and methods based on prompt tuning such as QueryNLT, which dynamically adjusts the language guidance through a learnable prompt vector. These methods have improved the semantic understanding ability of the tracking system to varying degrees, especially when the appearance of the target changes significantly, the prior knowledge provided by the language description can effectively assist the tracking decision.

[0061] However, there are still several key problems to be solved in the existing visual-linguistic tracking methods. First of all, in terms of feature fusion, most methods use simple feature concatenation or shallow attention mechanism, which fails to fully exploit the deep association between visual and language modalities. This coarse-grained fusion method is prone to lose semantic information, especially when dealing with complex and variable real-world scenarios. Secondly, in terms of dynamic adaptability, existing methods usually use fixed language descriptions as guidance, which is difficult to adapt to the real-time state changes of the target in the video sequence. When the target position or appearance changes significantly, the initial language description may no longer be accurate and may mislead the tracking process. Thirdly, in terms of multi-scale feature utilization, current methods often lack a systematic cross-scale feature alignment mechanism, making it difficult to effectively handle common challenges such as target scale changes.

[0062] Based on this, the embodiment of the present application proposes an innovative multi-scale adaptive cascaded fusion method. This method effectively solves the shortcomings of existing methods in modal interaction, dynamic adaptation and multi-scale fusion by designing a multi-level feature alignment mechanism, adaptive cascaded fusion and scale restoration architecture. Compared with the prior art, the embodiment of the present application has stronger scene adaptability and higher tracking accuracy, especially in complex and variable real-world application scenarios.

[0063] Reference Figure 1 The present application provides a visual-linguistic tracking method based on multi-scale adaptive cascaded fusion, which comprises the following steps:

[0064] S100, pre-processing and text extraction processing of the target video sequence to obtain pre-processed images and image text data;

[0065] Specifically, the target video sequence is acquired and template frames and search frames are randomly selected according to time order; the template frames and search frames are cropped to obtain template images and search region images; the template images and search region images are standardized and enhanced to obtain preprocessed images; the corresponding natural language descriptions are extracted based on the target video sequence, and the text is standardized and special tags are added to obtain image text data.

[0066] In this embodiment, during training, template frames and search frames are randomly selected from the video sequence in chronological order. Template images and search region images are then cropped based on their bounding boxes. During inference, the template image is cropped using the bounding box of the initial video frame, and the search region image is cropped in each subsequent frame based on the prediction box of the previous frame. All images undergo normalization, while the training images undergo data augmentation processing such as random flipping and scaling.

[0067] S200: Introducing a multi-scale feature alignment module and a scale recovery module to construct a tracking network model based on multi-scale adaptive cascade fusion;

[0068] In this embodiment, as Figure 3 As shown in (a), the tracking network model based on multi-scale adaptive cascaded fusion specifically includes a language encoder, a visual encoder, a multi-scale feature alignment module, a unified visual-language encoder, an adaptive cascaded fusion module, a scale recovery module, and a localization head detection module. The outputs of the language encoder and the visual encoder are respectively connected to the input of the multi-scale feature alignment module. The multi-scale feature alignment module, the unified visual-language encoder, the adaptive cascaded fusion module, the scale recovery module, and the localization head detection module are connected sequentially. The loss function of the tracking network model based on multi-scale adaptive cascaded fusion is shown below:

[0069]

[0070] L cls =-α t (1-p t ) γ log(p t )

[0071]

[0072] In the above formula, L represents the total loss function of the multi-scale adaptive cascaded fusion tracking network model. cls L1 represents the focus loss function, and L1 represents the L1 loss function. iou Let λ represent the generalized intersection-union loss function. iou , denotes the adjustable weight coefficient, a t denotes the weight for balancing the class imbalance, p t denotes the probability of predicting the class, denotes the focus parameter for adjusting the loss of easy-to-classify samples, y i denotes the true value, denotes the predicted value, n denotes the number of samples, i denotes the i-th sample, A denotes the predicted bounding box, B denotes the true bounding box, |A∩B| denotes the intersection area of the predicted bounding box and the true bounding box, and |A∪B| denotes the union area of the predicted bounding box and the true bounding box.

[0073] Further, as shown in (b) of Figure 3 the multi-scale feature alignment module includes a first branch structure, a second branch structure, and a third branch structure, wherein the first branch structure includes a linear layer and a multi-head attention mechanism, and the second branch structure and the third branch structure each include three convolutional layers and two down-sampling layers.

[0074] As shown in (c) of Figure 3 the scale recovery module includes two multi-head attention mechanisms, three residual connections, and a normalization layer, and a feedforward network.

[0075] S300, the multi-scale adaptive cascaded fusion tracking network model is used to perform visual and language tracking on the preprocessed image and image text data to obtain a target tracking result.

[0076] S310, the preprocessed image and image text data are input into the multi-scale adaptive cascaded fusion tracking network model.

[0077] S320, the language encoder of the multi-scale adaptive cascaded fusion tracking network model is used to perform language encoding processing on the image text data to obtain image language feature information.

[0078] In this embodiment, a pre-trained BERT-base model is loaded, and all parameters thereof are frozen during training; the input text is standardized and special marks are added to extract the features of the [CLS] mark; the input text mark is converted into a continuous vector representation through a word embedding layer, and position encoding is added to generate an initial word vector sequence; the multi-head attention mechanism is used in the Transformer layer to model the context dependency between words, and a feedforward neural network is used to extract deep semantic features, and a language representation containing rich context information is output.

[0079] S330, the visual encoder of the multi-scale adaptive cascaded fusion tracking network model is used to perform visual encoding processing on the preprocessed image to obtain image visual feature information.

[0080] In this embodiment, the pre-trained HiViT model parameters are loaded, and all trainable parameters thereof are frozen during the training process to maintain the general perception ability of the visual encoder; the input image is processed in blocks, and is mapped to visual tokens through linear embedding; the image blocks are modeled by using a multi-layer Transformer encoder, multi-scale visual feature representations are extracted layer by layer, the dependence between features is enhanced by means of multi-head attention mechanism and layer normalization operation, a feature map containing rich spatial details and global semantic information is generated, and is used for subsequent modal alignment and target positioning tasks.

[0081] S340, a multi-scale feature alignment module of the multi-scale adaptive cascaded fusion tracking network model is used to perform multi-scale alignment on the image language feature information and the image visual feature information, to obtain aligned image language feature information and aligned image visual feature information.

[0082] Specifically, the image language feature information and the image visual feature information are input into the multi-scale feature alignment module of the multi-scale adaptive cascaded fusion tracking network model; the image visual feature information is sequentially subjected to convolution and down-sampling operation processing by means of a convolution layer and a max-pooling layer, to obtain multi-scale visual feature representations containing multiple spatial resolutions; the multi-scale visual feature representations containing multiple spatial resolutions are subjected to dimension reduction processing, to obtain reduced multi-scale visual feature representations; the reduced multi-scale visual feature representations and the image language feature information are interactively fused by means of a cross-modal attention mechanism, to obtain visual features and language features at each scale; the visual features and language features at each scale are sequentially subjected to normalization processing, splicing processing and alignment operation, to obtain aligned image language feature information and aligned image visual feature information.

[0083] In this embodiment, the outputs of the language encoder and the visual encoder are processed by using a multi-scale cross-modal feature alignment module, the search region features are aligned with the language features at multiple scales, a 1x1 convolution layer is used in combination with a max-pooling down-sampling operation to construct multi-scale visual feature representations containing multiple spatial resolutions for the feature map output by the visual encoder; a cross-modal attention mechanism is used to realize interactive fusion of visual and language features; the visual features and language features at each scale are spliced after being subjected to normalization processing, and unified spatial dimensions are realized through feature alignment operation.

[0084] S350, a visual language unified encoder of the multi-scale adaptive cascaded fusion tracking network model is used to perform cross-modal feature interactive fusion on the aligned image language feature information and the aligned image visual feature information, to obtain unified cross-modal feature representations.

[0085] Specifically, the aligned image language feature information and the aligned image visual feature information are input into a visual language unified encoder of a multi-scale adaptive cascaded fusion tracking network model; the aligned image language feature information and the aligned image visual feature information are respectively subjected to serialization processing to obtain serialized image language feature information and serialized image visual feature information; and the serialized image language feature information and the serialized image visual feature information are subjected to feature association based on a self-attention mechanism through a plurality of stacked Transformer modules to obtain unified cross-modal feature representation.

[0086] In the embodiment, pre-trained ViT encoder weights are loaded, and all parameters thereof are fine-tuned in a training stage; the multi-scale aligned visual feature and the language feature are jointly input into the encoder and subjected to unified serialization processing; the multi-layer stacked Transformer modules are used to mine deep associations between the visual feature and the language feature by using the self-attention mechanism to realize efficient information interaction between the modalities; and the fused unified cross-modal feature representation is output as an input basis for subsequent tracking prediction.

[0087] S360, an adaptive cascaded fusion module of the multi-scale adaptive cascaded fusion tracking network model performs adaptive fusion on the unified cross-modal feature representation to obtain multi-scale search region features;

[0088] Specifically, the unified cross-modal feature representation is input into the adaptive cascaded fusion module of the multi-scale adaptive cascaded fusion tracking network model; the unified cross-modal feature representation is calculated layer by layer based on a cascaded feature fusion strategy by using an adaptive gating mechanism to obtain contribution weights of the features; and the contribution weights of the features and the unified cross-modal feature representation are weighted and fused according to the search region to obtain the multi-scale search region features.

[0089] In the embodiment, the cascaded feature fusion strategy is used to calculate the contribution weights of the features output by the cross-modal unified encoder layer by layer by using the adaptive gating mechanism; and then the search region features output by different encoding layers are weighted and fused according to different weights, and the calculation formula is as follows:

[0090]

[0091] The multi-layer visual features are weighted and fused by using the adaptive gating mechanism. Each layer of features is processed by an MLP, multiplied by a gating factor, then the weight is calculated by Softmax normalization, and finally the features of each layer are weighted and fused according to the weight to generate the final visual feature representation.

[0092] S370, a scale recovery module of the multi-scale adaptive cascaded fusion tracking network model performs scale recovery processing on the multi-scale search region features to obtain search region features.

[0093] Specifically, the multi-scale search region feature is input to a scale recovery module of the multi-scale adaptive cascaded fusion tracking network model; based on a cross-scale multi-head attention mechanism, high-level abstract features and low-level detail features in the multi-scale search region feature are processed for information interaction to obtain a multi-scale search region feature with an initial scale; a layer normalization operation is applied to the multi-scale search region feature with the initial scale, and a feedforward neural network is introduced to enhance the fused feature to obtain a search region feature.

[0094] In this embodiment, the feature information of the search region at different scales is fused to the original scale. First, the cross-scale multi-head attention mechanism is used to layer by layer interact the high-level abstract features and the low-level detail features for information, and gradually guide the features to restore to the initial scale; a layer normalization operation is applied, and a feedforward neural network is introduced to further enhance the discriminability of the fused feature, and output the search region feature at a uniform scale.

[0095] S380, a positioning head detection module of the multi-scale adaptive cascaded fusion tracking network model is used to track and locate the search region feature to obtain a target tracking result.

[0096] Specifically, the search region feature is input to a positioning head detection module of the multi-scale adaptive cascaded fusion tracking network model; the search region feature is converted to obtain a two-dimensional search region feature; the two-dimensional search region feature is stacked by a plurality of convolution, batch normalization and activation functions to output a classification score map, a bounding box size map and an offset map; the position with the highest confidence in the classification score map is selected as the target center, and the offset map and the bounding box size map are combined for calculation to obtain the target tracking result.

[0097] In this embodiment, a tracking prediction head based on a convolutional neural network is used to convert the input feature into a two-dimensional feature map; the two-dimensional feature map is stacked by a plurality of convolution, batch normalization and activation functions to output a classification score map, a bounding box size map and an offset map; the position with the highest confidence in the classification score map is selected as the target center, and the offset and the bounding box size are combined to calculate the final tracking result.

[0098] To sum up, the embodiment of the present application first pre-processes the input template image and the search region image; extracts language features by using a pre-trained language model, and extracts visual features by using a visual encoder; aligns the search region features and the language features at multiple scales by using a multi-scale cross-modal feature alignment module; performs cross-modal interaction fusion by using a visual language unified encoder; dynamically fuses the features of different layers in the unified encoder by using an adaptive cascaded fusion module; restores the multi-scale features to the original scale and fuses them; and finally outputs the target tracking result by using a tracking prediction head. By using multi-scale adaptive fusion and cross-modal alignment, the tracking robustness and accuracy in a complex scene can be significantly improved, and the visual language collaborative tracking task in a dynamic environment can be applied.

[0099] With reference to Figure 2 The visual language tracking system based on multi-scale adaptive cascaded fusion comprises:

[0100] A first module is configured to pre-process and text extraction process a target video sequence to obtain pre-processed images and image text data.

[0101] A second module is configured to introduce a multi-scale feature alignment module and a scale restoration module to construct a tracking network model based on multi-scale adaptive cascaded fusion.

[0102] A third module is configured to perform visual language tracking on the pre-processed images and image text data by using the tracking network model based on multi-scale adaptive cascaded fusion to obtain a target tracking result.

[0103] The content in the method embodiments is applicable to the system embodiments, the system embodiments achieve the same functions as the method embodiments, and achieve the same beneficial effects as the method embodiments.

[0104] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the above-mentioned embodiments, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the present application.

Claims

1. A visual language tracking method based on multi-scale adaptive cascaded fusion, characterized in that, The method comprises the following steps: preprocessing and text extraction processing are performed on the target video sequence to obtain preprocessed images and image text data; a multi-scale feature alignment module and a scale recovery module are introduced to construct a tracking network model based on multi-scale adaptive cascaded fusion; The tracking network model based on multi-scale adaptive cascaded fusion comprises a language encoder, a visual encoder, a multi-scale feature alignment module, a visual language unified encoder, an adaptive cascaded fusion module, a scale recovery module and a positioning head detection module, wherein the output ends of the language encoder and the visual encoder are connected with the input end of the multi-scale feature alignment module, and the multi-scale feature alignment module, the visual language unified encoder, the adaptive cascaded fusion module, the scale recovery module and the positioning head detection module are connected in sequence, and the loss function of the tracking network model based on multi-scale adaptive cascaded fusion is as follows: ; In the above formula, This represents the total loss function of the multi-scale adaptive cascaded fusion tracking network model. Represents the focus loss function. Describes the L1 loss function. This represents the generalized intersection-union loss function. , This represents the adjustable weighting coefficient. This indicates the weights that balance the class imbalance. This represents the probability of predicting the category. This indicates the focus parameter for adjusting the loss of easily classified samples. Represents the actual value. Indicates the predicted value. Indicates the number of samples. Indicates the first One sample, Indicates the prediction box. Represents the true bounding box. This represents the area of ​​the intersection between the predicted bounding box and the ground truth bounding box. This represents the area of ​​the union of the predicted bounding box and the ground truth bounding box; The tracking network model based on multi-scale adaptive cascaded fusion performs visual language tracking on the preprocessed images and image text data to obtain target tracking results. The multi-scale feature alignment module of the tracking network model based on multi-scale adaptive cascaded fusion performs multi-scale alignment on image language feature information and image visual feature information to obtain aligned image language feature information and aligned image visual feature information, comprising: inputting the image language feature information and the image visual feature information into the multi-scale feature alignment module of the tracking network model based on multi-scale adaptive cascaded fusion; performing convolution and downsampling operation processing on the image visual feature information in sequence through the convolution layer and the max-pooling layer to obtain multi-scale visual feature representations containing multiple spatial resolutions; dimensionality reduction processing is performed on the multi-scale visual feature representations containing multiple spatial resolutions to obtain dimensionality-reduced multi-scale visual feature representations; interacting and fusing the dimensionality-reduced multi-scale visual feature representations and the image language feature information through the cross-modal attention mechanism to obtain visual features and language features of each scale; performing normalization processing, splicing processing and alignment operation on the visual features and language features of each scale in sequence to obtain aligned image language feature information and aligned image visual feature information.

2. The visual language tracking method based on multi-scale adaptive cascade fusion according to claim 1, characterized in that, The step of preprocessing and text extraction processing on the target video sequence to obtain preprocessed images and image text data comprises: obtaining a target video sequence and randomly selecting template frames and search frames according to time sequence; performing cropping processing on the template frames and the search frames to obtain template images and search region images; performing standardization processing and image enhancement processing on the template images and the search region images to obtain preprocessed images; extracting the natural language description corresponding to the target video sequence based on the target video sequence, and performing standardization processing and adding special marks on the text to obtain image text data. 3.The visual language tracking method based on multi-scale adaptive cascade fusion according to claim 2, characterized in that, The step of performing visual language tracking on the preprocessed images and image text data by the tracking network model based on multi-scale adaptive cascaded fusion to obtain target tracking results comprises: The preprocessed image and the image text data are input into the multi-scale adaptive cascaded fusion tracking network model; The language encoder of the multi-scale adaptive cascaded fusion tracking network model performs language coding processing on the image text data to obtain image language feature information; The visual encoder of the multi-scale adaptive cascaded fusion tracking network model performs visual coding processing on the preprocessed image to obtain image visual feature information; The multi-scale feature alignment module of the multi-scale adaptive cascaded fusion tracking network model performs multi-scale alignment on the image language feature information and the image visual feature information to obtain aligned image language feature information and aligned image visual feature information; The visual language unified encoder of the multi-scale adaptive cascaded fusion tracking network model performs cross-modal feature interaction fusion on the aligned image language feature information and the aligned image visual feature information to obtain unified cross-modal feature representation; The adaptive cascaded fusion module of the multi-scale adaptive cascaded fusion tracking network model performs adaptive fusion on the unified cross-modal feature representation to obtain multi-scale search region features; The scale recovery module of the multi-scale adaptive cascaded fusion tracking network model performs scale recovery processing on the multi-scale search region features to obtain search region features; The positioning head detection module of the multi-scale adaptive cascaded fusion tracking network model performs tracking positioning on the search region features to obtain target tracking results.

4. The visual language tracking method based on multi-scale adaptive cascade fusion according to claim 3, characterized in that, The visual language unified encoder of the multi-scale adaptive cascaded fusion tracking network model performs cross-modal feature interaction fusion on the aligned image language feature information and the aligned image visual feature information to obtain unified cross-modal feature representation, which specifically includes: The aligned image language feature information and the aligned image visual feature information are input into the visual language unified encoder of the multi-scale adaptive cascaded fusion tracking network model; The aligned image language feature information and the aligned image visual feature information are respectively serialized to obtain serialized image language feature information and serialized image visual feature information; Through a plurality of stacked Transformer modules, the serialized image language feature information and the serialized image visual feature information are associated based on a self-attention mechanism to obtain unified cross-modal feature representation.

5. The visual language tracking method based on multi-scale adaptive cascade fusion according to claim 4, characterized in that, The adaptive cascaded fusion module of the multi-scale adaptive cascaded fusion tracking network model performs adaptive fusion on the unified cross-modal feature representation to obtain multi-scale search region features, which specifically includes: The unified cross-modal feature representation is input into the adaptive cascaded fusion module of the multi-scale adaptive cascaded fusion tracking network model; Based on a cascaded feature fusion strategy, the unified cross-modal feature representation is calculated layer by layer through an adaptive gating mechanism to obtain feature contribution weights; According to the search region, the feature contribution weights and the unified cross-modal feature representation are weighted and fused to obtain multi-scale search region features.

6. The visual language tracking method based on multi-scale adaptive cascade fusion according to claim 5, characterized in that, The scale recovery module of the tracking network model based on multi-scale adaptive cascaded fusion performs scale recovery processing on the multi-scale search region features to obtain search region features, which specifically includes: The scale recovery module of the tracking network model based on multi-scale adaptive cascaded fusion is inputted with the multi-scale search region features; Based on the cross-scale multi-head attention mechanism, the high-level abstract features and low-level detail features in the multi-scale search region features are processed for information interaction to obtain multi-scale search region features with an initial scale. The multi-scale search region features with an initial scale are subjected to layer normalization operation, and a feedforward neural network is introduced to enhance the fused features to obtain search region features.

7. The visual language tracking method based on multi-scale adaptive cascade fusion according to claim 6, characterized in that, The positioning head detection module of the tracking network model based on multi-scale adaptive cascaded fusion performs tracking positioning on the search region features to obtain target tracking results, which specifically includes: The positioning head detection module of the tracking network model based on multi-scale adaptive cascaded fusion is inputted with the search region features; The search region features are converted to obtain two-dimensional search region features; The two-dimensional search region features are stacked through multi-layer convolution, batch normalization and activation function to output classification score map, bounding box size map and offset map; The position with the highest confidence in the classification score map is selected as the target center, and the offset map and the bounding box size map are combined for calculation to obtain the target tracking result.

8. A visual language tracking system based on multi-scale adaptive cascaded fusion, characterized in that, It includes the following modules: The first module is used for pre-processing and text extraction processing of the target video sequence to obtain pre-processed images and image text data; The second module is used for introducing a multi-scale feature alignment module and a scale recovery module to construct a tracking network model based on multi-scale adaptive cascaded fusion; The tracking network model based on multi-scale adaptive cascaded fusion specifically includes a language encoder, a visual encoder, a multi-scale feature alignment module, a visual language unified encoder, an adaptive cascaded fusion module, a scale recovery module and a positioning head detection module, wherein the output ends of the language encoder and the visual encoder are respectively connected with the input end of the multi-scale feature alignment module, the multi-scale feature alignment module, the visual language unified encoder, the adaptive cascaded fusion module, the scale recovery module and the positioning head detection module are connected in sequence, and the loss function of the tracking network model based on multi-scale adaptive cascaded fusion is specifically as follows: ; In the above formula, This represents the total loss function of the multi-scale adaptive cascaded fusion tracking network model. Represents the focus loss function. Describes the L1 loss function. This represents the generalized intersection-union loss function. , This represents the adjustable weighting coefficient. This indicates the weights that balance the class imbalance. This represents the probability of predicting the category. This indicates the focus parameter for adjusting the loss of easily classified samples. Represents the actual value. Indicates the predicted value. Indicates the number of samples. Indicates the first One sample, Indicates the prediction box. Represents the true bounding box. This represents the area of ​​the intersection between the predicted bounding box and the ground truth bounding box. This represents the area of ​​the union of the predicted bounding box and the ground truth bounding box; The third module is used for visual language tracking of the pre-processed images and image text data based on the tracking network model based on multi-scale adaptive cascaded fusion to obtain target tracking results. The multi-scale feature alignment module of the tracking network model based on multi-scale adaptive cascaded fusion performs multi-scale alignment on image language feature information and image visual feature information to obtain aligned image language feature information and aligned image visual feature information, which includes: The multi-scale feature alignment module of the tracking network model based on multi-scale adaptive cascaded fusion is inputted with the image language feature information and the image visual feature information; The image visual feature information is sequentially subjected to convolution and down-sampling operation processing through a convolution layer and a maximum pooling layer to obtain multi-scale visual feature representations containing multiple spatial resolutions; The multi-scale visual feature representations containing multiple spatial resolutions are subjected to dimension reduction processing to obtain dimension-reduced multi-scale visual feature representations; The dimension-reduced multi-scale visual feature representations and the image language feature information are interactively fused through a cross-modal attention mechanism to obtain visual features and language features of each scale; The visual features and language features of each scale are sequentially subjected to normalization processing, splicing processing and alignment operation to obtain aligned image language feature information and aligned image visual feature information.

Citation Information

Patent Citations

  • Short-time natural language target tracking method based on visual language large model

    CN117746024A

  • Unmanned aerial vehicle multi-modal feature fusion target tracking method and system based on natural language description

    CN120013992A

Cited By

  • Text guidance and occlusion perception-based complex scene infrared ship multi-target tracking method and system

    CN122067205A