Target tracking method and system based on Mama visual hybrid module

By combining the Mamba visual hybrid module of Transformer and Mamba state space model, the global context and local detail information are extracted and superimposed, the problem of insufficient image information capture in the prior art is solved, and faster and more accurate target tracking is achieved.

CN120298458AActive Publication Date: 2025-07-11NANCHANG INST OF TECH

Patent Information

Application Number
CN202510777185.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing Mamba-based target tracking technology is difficult to capture the global information of the image when it handles long-term tracking and facing occlusion, similar objects interference, sudden disappearance, etc., resulting in tracking failure.

Method used

The Mamba visual hybrid module based on the Transformer architecture and the Mamba state space model is adopted, combining the multi-head self-attention and image state space modules to extract and superimpose global context information and local detail information, and feature enhancement is performed through feature fusion and feature improvement modules to optimize the tracking model.

Benefits of technology

While maintaining tracking performance, it improves tracking speed and robustness, can better handle occlusion and similar interference in complex scenarios, and improves the generalization ability of the tracker.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298458A_ABST
    Figure CN120298458A_ABST
Patent Text Reader

Abstract

The invention provides a target tracking method and system based on a Mama visual hybrid module, and the method comprises the steps: inputting a search image and a template image into the Mama visual hybrid module, and extracting and superposing global context information and local detail information; inputting the output of the Mama visual mixing module into a feature fusion module for fusion; performing feature enhancement on the output of the feature fusion module by using a feature enhancement module to obtain enhanced features; pre-training the optimized tracking model by using a large-scale data set, and updating parameters of the tracking model; and inputting the search image and the template image into the parameter-updated tracking model to obtain a tracking result. According to the method, a feature aggregation network is designed, historical state information of a search area is integrated by using an image state space module, the historical state information and interacted features are spliced along a channel direction, and effective aggregation of the features is realized through convolution, activation, normalization and jump connection to filter irrelevant information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and particularly to an object tracking method and system based on a Mamba vision hybrid module. Background Art

[0002] Object tracking is a basic task in computer vision applications, aiming to locate and label an object of interest in consecutive video frames given a tracking target. It is currently widely used in fields such as autonomous driving, biomedicine, indoor and outdoor surveillance, highway intersection cameras, drones and robots, and human-computer interaction. Although object tracking technology has developed rapidly in recent years, it still faces challenges such as occlusion, deformation, background blur, and interference from similar objects in practical applications. Therefore, there is still great room for improvement in object tracking technology.

[0003] The emergence and application of Transformer have further promoted the development of object tracking technology. It can obtain the context information of video images through the attention mechanism, and automatically focus on the areas of interest of the model by reducing the loss through a multi-layer network architecture and backpropagation algorithm, thereby improving the tracking performance. Currently, most tracking algorithms are based on Transformer and have achieved very good tracking results. However, the attention mechanism focuses on capturing global information and loses a lot of key local information, which has a relatively large impact on the performance of the tracker. Combining local information and global information is an important direction for improving the performance of the tracker.

[0004] The Mamba structure based on the state space model is widely used in various fields. The state space model efficiently processes data through selective scanning, and the Mamba structure based on this reduces the computational complexity and realizes low-latency inference. In some application fields, the same effect can be achieved with far fewer model parameters than Transformer. In the latest Mamba-based tracker, the tracking speed has been greatly improved while ensuring the tracking performance. However, the way it processes data in a sequential manner has limitations in spatial dependence and is difficult to capture the global information of the image. This often leads to tracking failure when dealing with long-term tracking and situations such as occlusion, interference from similar objects, and sudden disappearance. Summary of the Invention

[0005] In view of the above situation, the main purpose of the present invention is to propose an object tracking method and system based on a Mamba vision hybrid module to solve the above technical problems.

[0006] The present invention proposes an object tracking method based on a Mamba vision hybrid module, and the method includes the following steps: Step 1: Construct a Mamba visual hybrid module based on the Transformer architecture and the Mamba state space model, construct a feature fusion module based on the feature aggregation mechanism, and construct a feature enhancement module based on the feature enhancement mechanism. The Mamba visual hybrid module, the feature fusion module, and the feature enhancement module constitute the tracking model. Among them, the Mamba visual hybrid module includes multi-head self-attention and an image state space module; Step 2: Input the search image and the template image into the Mamba visual hybrid module, extract and superimpose the global context information and the local detail information to obtain the output of the Mamba visual hybrid module; Step 3: Input the output of the Mamba visual hybrid module into the feature fusion module for fusion to obtain the output of the feature fusion module; Step 4: Use the feature enhancement module to enhance the features of the output of the feature fusion module to obtain the enhanced features. Calculate the classification loss and the regression loss using the enhanced features, and optimize the tracking model using the classification loss and the regression loss to obtain the optimized tracking model; Step 5: Pre-train the optimized tracking model using a large-scale dataset and update the parameters of the tracking model to obtain the tracking model with updated parameters; Step 6: Input the search image and the template image into the tracking model with updated parameters, and repeat Steps 2 to 4 iteratively to obtain the finally enhanced features. Input the finally enhanced features into the prediction head to obtain the tracking result.

[0007] The present invention also proposes an object tracking system based on the Mamba visual hybrid module. The system includes: A construction module for: Constructing a Mamba visual hybrid module based on the Transformer architecture and the Mamba state space model, constructing a feature fusion module based on the feature aggregation mechanism, and constructing a feature enhancement module based on the feature enhancement mechanism. The Mamba visual hybrid module, the feature fusion module, and the feature enhancement module constitute the tracking model. Among them, the Mamba visual hybrid module includes multi-head self-attention and an image state space module; An extraction module for: Inputting the search image and the template image into the Mamba visual hybrid module, extracting and superimposing the global context information and the local detail information to obtain the output of the Mamba visual hybrid module; A calculation module for: Inputting the output of the Mamba visual hybrid module into the feature fusion module for fusion to obtain the output of the feature fusion module; A learning module for: The output of the feature fusion module is enhanced using the feature enhancement module to obtain enhanced features. The classification loss and regression loss are calculated using the enhanced features, and the tracking model is optimized using the classification loss and regression loss to obtain an optimized tracking model; A pre-training module, which is used for: Pre-training the optimized tracking model using a large-scale dataset and updating the parameters of the tracking model to obtain a tracking model with updated parameters; A tracking module, which is used for: Inputting the search image and the template image into the tracking model with updated parameters, iteratively using the Mamba vision mixing module to extract and superimpose global context information and local detail information, fusing them through the feature fusion module, and finally enhancing the features using the feature enhancement module to obtain finally enhanced features, and inputting the finally enhanced features into the prediction head to obtain the tracking result.

[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The Mamba vision mixing module designed in the present invention adopts an attention and image state dual-branch structure. The attention branch captures the long-range dependence relationship of the image by calculating multi-head attention to obtain the global context information of the image; the image state branch can not only dynamically select features but also retain the spatial details of the image through a symmetric convolution structure, with high parameter efficiency and capable of obtaining the local semantic information of the image. The two are superimposed to enhance the local semantic information while avoiding the loss of global information.

[0009] 2. The present invention designs a feature aggregation network, which uses an image state space module to integrate the historical state information of the search area, splices the historical state information and the interacted features along the channel direction, and realizes the effective aggregation of features through convolution, activation, normalization, and skip connection, filtering out irrelevant information.

[0010] 3. The present invention designs a feature enhancement module, which is divided into two parts and is connected in series. The first part performs multi-layer channel attention calculation on the input and then performs multi-feature aggregation; the second part performs spatial attention calculation according to the same structure. This module models the channel and spatial dependence relationships, captures both global features and local features, and enhances the corresponding region features of the target object.

[0011] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the embodiments of the present invention. Description of the Drawings

[0012] Figure 1Flowchart of the steps of the object tracking method based on the Mamba vision hybrid module proposed by the present invention.

[0013] Figure 2 Schematic diagram of the Mamba vision hybrid module of the object tracking method based on the Mamba vision hybrid module proposed by the present invention.

[0014] Figure 3 Schematic diagram of the feature enhancement module PFM of the object tracking method based on the Mamba vision hybrid module proposed by the present invention.

[0015] Figure 4 Tracking architecture diagram of the object tracking method based on the Mamba vision hybrid module proposed by the present invention.

[0016] Figure 5 Structural diagram of the object tracking system based on the Mamba vision hybrid module proposed by the present invention. Detailed implementation manners

[0017] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals are the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0018] These and other aspects of the embodiments of the present invention will be clear with reference to the following description and drawings. In these descriptions and drawings, some specific implementation manners in the embodiments of the present invention are specifically disclosed as some ways to implement the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0019] Please refer to Figure 1 , the embodiments of the present invention propose an object tracking method based on the Mamba vision hybrid module. The method includes the following steps: Step 1: Construct a Mamba vision hybrid module based on the Transformer architecture and the Mamba state space model, construct a feature fusion module based on the feature aggregation mechanism, and construct a feature enhancement module based on the feature enhancement mechanism. The Mamba vision hybrid module, the feature fusion module, and the feature enhancement module constitute a tracking model; wherein the Mamba vision hybrid module includes a multi-head self-attention and an image state space module.

[0020] Step 2: Input the search image and the template image into the Mamba vision hybrid module, extract and superimpose the global context information and the local detail information to obtain the output of the Mamba vision hybrid module.

[0021] Please refer to Figure 2and Figure 4 In step 2, the search image and the template image are input into the Mamba visual mixing module, and the global context information and local detail information are extracted and superimposed to obtain the output of the Mamba visual mixing module. The specific steps are as follows: The search image and the template image are input into the Mamba visual mixing module for initialization processing respectively to obtain the image sequence of the template area and the image sequence of the search area; The image sequences of the template area and the search area are respectively subjected to feature extraction using a convolutional network to obtain the original features of the image sequence of the template area and the original features of the image sequence of the search area; The original features of the image sequence of the template area are successively flattened and processed with position encoding added to obtain the original feature token sequence of the template image; The original features of the image sequence of the search area are successively flattened and processed with position encoding added to obtain the original feature token sequence of the search area; The original feature token sequence of the template image is input into the multi-head self-attention for information extraction to obtain the global context features of the image; The original feature token sequence of the search area image is input into the image state space module for local information extraction to obtain the local features of the image; The global context features of the image and the local features of the image are superimposed to obtain the output of the Mamba visual mixing module.

[0022] When the original feature token sequence of the template image is input into the multi-head self-attention for information extraction, the following relational expressions exist in the corresponding process: ; Among them, represents the intermediate value, represents the attention calculation, represents the input of the Mamba visual mixing module, represents the normalization function, represents the query vector, represents the transpose of the key vector, represents the square root of the key vector dimension, represents the value vector; In the step of inputting the original feature token sequence of the search area image into the image state space module for local information extraction to obtain the local features of the image, the following relational expressions exist in the corresponding process: ; Among them, represents the state space model the current moment The input, represents the first learnable matrix, represents the state space model Past moment The information set of all input images, represents the second learnable matrix, represents the third learnable matrix, represents the fourth learnable matrix, represents the matrix multiplication operation, represents the state branch of the image state space module, represents the symmetric convolutional branch of the image state space module, represents the first linear mapping process, represents the output of the state branch of the image state space module, represents the activation operation, represents the one-dimensional ordinary convolution operation, represents the second linear mapping process, represents the concatenation operation; In the step of superimposing the global context features of the image and the local features of the image to obtain the output of the Mamba vision mixing module, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the output of the Mamba vision mixing module.

[0023] Furthermore, initialize the first-frame template image and the search area images of subsequent frames to obtain an image sequence of the template and the search area; respectively extract the original features of the template and the search area image sequences through a convolutional network, and flatten the original features; then, by adding position encoding, obtain the original feature token sequence of the template image and the original feature token sequence of the search area image.

[0024] It should be noted that based on the network architecture of Transformer, the Mamba vision mixing module is used to extract features from the template area and the search area at the low layer. As Figure 2 shown, the Mamba vision mixing module consists of an image state space module and multi-head self-attention, and extracts the global context information and local detail information of the image through a dual-branch structure; the first branch of the image state space module fuses the image historical information to avoid the loss and incompleteness of the target object's information, provides global information in the second branch, enhances the modeling ability, and the two branches are combined to dynamically select key features to obtain the local information of the image.

[0025] Step 3: Input the output of the Mamba visual mixing module into the feature fusion module for fusion to obtain the output of the feature fusion module.

[0026] In Step 3, input the output of the Mamba visual mixing module into the feature fusion module for fusion to obtain the output of the feature fusion module, which specifically includes the following steps: Use the image state space module to integrate the historical state information of the search area to obtain the integrated historical state information of the search area; Input the output of the Mamba visual mixing module into the feature fusion module, and successively perform a concatenation operation, a convolutional activation operation, and a residual connection operation with the integrated historical state information of the search area to obtain the output of the feature fusion module.

[0027] Input the output of the Mamba visual mixing module into the feature fusion module, and successively perform a concatenation operation, a convolutional activation operation, and a residual connection operation with the integrated historical state information of the search area to obtain the output of the feature fusion module. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the feature output after the concatenation operation and linear mapping, represents the fully connected layer, represents the feature after the interaction between the template area and the search area, represents the feature output by the image state space module, represents the fusion process, represents the batch normalization operation, represents the convolutional operation with a convolutional kernel size of 1×1, represents the activation function, represents the output of the feature fusion module.

[0028] Step 4: Use the feature enhancement module to enhance the features of the output of the feature fusion module to obtain enhanced features, calculate the classification loss and regression loss respectively using the enhanced features, and optimize the tracking model using the classification loss and regression loss to obtain the optimized tracking model.

[0029] Please refer to Figure 3 and Figure 4 . In Step 4, use the feature enhancement module to enhance the features of the output of the feature fusion module to obtain enhanced features, which are specifically as follows: S101. Perform global average pooling on the spatial dimension of the output of the feature fusion module using the feature enhancement module to obtain the result after global average pooling. Perform channel weight assignment processing on the result after global average pooling through a convolutional layer to obtain the channel weight matrix of the input features. Multiply the channel weight matrix of the input features by the original matrix of the output of the feature fusion module to obtain the feature matrix output by the first layer of the channel branch and the matrix output by the first layer of the skip connection; S102. Use the matrix output by the first layer of the skip connection as the input and repeat step S101 iteratively to obtain the feature matrix output by the second layer of the channel branch, the feature matrix output by the third layer of the channel branch, and the feature matrix output by the fourth layer of the channel branch respectively; S103. Concatenate the feature matrix output by the first layer of the channel branch, the feature matrix output by the second layer of the channel branch, the feature matrix output by the third layer of the channel branch, and the feature matrix output by the fourth layer of the channel branch to obtain the first concatenated result. Use a convolutional layer to perform channel dimension weight assignment processing on the first concatenated result to obtain the output feature matrix of the first part of the feature enhancement module; S104. Perform global average pooling and max pooling on the channel dimension of the output feature matrix of the first part of the feature enhancement module in sequence to obtain the pooled features. Perform concatenation processing and convolutional activation processing on the pooled features along the channel dimension in sequence to obtain the weight matrix of each spatial position of the input feature matrix. Multiply the weight matrix of each spatial position of the input feature matrix by the output feature matrix of the first part of the feature enhancement module to obtain the feature matrix output by the first layer of the spatial branch and the matrix output by the first layer of the skip connection; S105. Use the matrix output by the first layer of the skip connection as the input and repeat step S104 iteratively to obtain the feature matrix output by the second layer of the spatial branch, the feature matrix output by the third layer of the spatial branch, and the feature matrix output by the fourth layer of the spatial branch respectively; S106. Concatenate the feature matrix output by the first layer of the spatial branch, the feature matrix output by the second layer of the spatial branch, the feature matrix output by the third layer of the spatial branch, and the feature matrix output by the fourth layer of the spatial branch along the channel dimension to obtain the second concatenated result. Perform channel dimension weight assignment processing on the second concatenated result through a convolutional layer to obtain the output of the second part for modeling spatial position dependencies; S107. Add the output of the second part for modeling spatial position dependencies to the output of the feature fusion module to obtain the output of the feature enhancement module.

[0030] The spatial dimension of the output of the feature fusion module is processed by the feature enhancement module through global average pooling to obtain the result after global average pooling. The result after global average pooling is processed by a convolutional layer to assign weights to the channels, obtaining the channel weight matrix of the input features. The channel weight matrix of the input features is multiplied by the original matrix of the output of the feature fusion module to obtain the feature matrix of the output of the channel branch of the first layer and the matrix of the output of the first layer skip connection. The relational expressions for the corresponding process are as follows: ; Among them, represents the feature matrix of the output of the channel branch of the first layer, represents the channel attention calculation, represents the matrix of the output of the first layer skip connection of the first part of the structure; In the step of taking the matrix of the output of the first layer skip connection as the input and repeating step S101 iteratively to obtain the feature matrix of the output of the channel branch of the second layer, the feature matrix of the output of the channel branch of the third layer, and the feature matrix of the output of the channel branch of the fourth layer respectively, the relational expressions for the corresponding process are as follows: ; Among them, represents the feature matrix of the output of the channel branch of the second layer, represents the matrix of the output of the second layer skip connection of the first part of the structure, represents the feature matrix of the output of the channel branch of the third layer, represents the matrix of the output of the third layer skip connection of the first part of the structure, represents the feature matrix of the output of the channel branch of the fourth layer; In the step of concatenating the feature matrix of the output of the channel branch of the first layer, the feature matrix of the output of the channel branch of the second layer, the feature matrix of the output of the channel branch of the third layer, and the feature matrix of the output of the channel branch of the fourth layer to obtain the first concatenated result, and using a convolutional layer to assign weights to the channel dimension of the first concatenated result to obtain the output feature matrix of the first part of the feature enhancement module, the relational expressions for the corresponding process are as follows: ; Among them, represents the output feature matrix of the first part of the feature enhancement module.

[0031] The output feature matrix of the first part of the feature enhancement module is successively subjected to global average pooling and max pooling along the channel dimension to obtain the pooled features. The pooled features are successively subjected to concatenation and convolutional activation along the channel dimension to obtain the weight matrix for each spatial position of the input feature matrix. The weight matrix for each spatial position of the input feature matrix is multiplied by the output feature matrix of the first part of the feature enhancement module to obtain the feature matrix output by the spatial branch of the first layer and the matrix output by the skip connection of the first layer. The relational expressions for the corresponding processes are as follows: ; Among them, represents the feature matrix output by the spatial branch of the first layer, represents the channel attention calculation, represents the matrix output by the skip connection of the first layer of the second part of the structure; In the step of taking the matrix output by the skip connection of the first layer as the input and repeating step S104 iteratively to obtain the feature matrix output by the spatial branch of the second layer, the feature matrix output by the spatial branch of the third layer, and the feature matrix output by the spatial branch of the fourth layer respectively, the relational expressions for the corresponding processes are as follows: ; Among them, represents the feature matrix output by the spatial branch of the second layer, represents the matrix output by the skip connection of the second layer of the second part of the structure, represents the feature matrix output by the spatial branch of the third layer, represents the matrix output by the skip connection of the third layer of the second part of the structure, represents the feature matrix output by the spatial branch of the fourth layer; In the step of concatenating the feature matrix output by the spatial branch of the first layer, the feature matrix output by the spatial branch of the second layer, the feature matrix output by the spatial branch of the third layer, and the feature matrix output by the spatial branch of the fourth layer along the channel dimension to obtain the second concatenated result, and then performing weight assignment for the channel dimension on the second concatenated result through a convolutional layer to obtain the output of the second part for modeling spatial position dependencies, the relational expressions for the corresponding processes are as follows: ; Among them, represents the output of the second part for modeling spatial position dependencies; In the step of adding the output of the second part for modeling spatial position dependencies to the output of the feature fusion module to obtain the output of the feature enhancement module, the relational expressions for the corresponding processes are as follows: ; Among them, represents the output of the feature enhancement module.

[0032] Step 5: Use a large-scale dataset to pre-train the optimized tracking model and update the parameters of the tracking model to obtain a tracking model with updated parameters.

[0033] Furthermore, use the training set parts of the GOT-10K, LaSOT, COCO2017, and TrackingNet datasets to train the network. Randomly sample 60,000 image pairs from the video frame images in the training set for 300 rounds of training. Set the batch size to 32 and the learning rate to e-4. Reduce the learning rate at the 240th round of training to prevent overfitting.

[0034] Step 6: Input the search image and the template image into the tracking model with updated parameters, and repeat Steps 2 to 4 iteratively to obtain the finally enhanced feature. Input the finally enhanced feature into the prediction head to obtain the tracking result.

[0035] In Step 6, input the search image and the template image into the tracking model with updated parameters, and repeat Steps 1 to 4 iteratively to obtain the finally enhanced feature. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the template feature after feature extraction, represents the search region feature after feature extraction, represents the feature after concatenating the template feature and the search region feature along the channel dimension, represents the template feature length of, represents the image state space module, represents the feature fusion module, represents the feature enhancement module.

[0036] Furthermore, the present invention extracts features by selecting the Mamba vision hybrid module; uses the feature aggregation module for feature fusion to reduce the impact of irrelevant features on the tracking result; designs the network as a two-stream structure to extract features from the template and the search region respectively; uses the feature enhancement module to further process the fused features, and enhances the relevant feature regions of the target object by modeling the channel and spatial dependencies in multiple layers; pre-trains the network model based on the Mamba vision hybrid module, uses the pre-trained network model to first extract features from the linearly embedded template image, and then extract features from the search region image; connects the template features and the features of the last search region layer for multi-layer feature interaction, uses the image state space module to process the features of the search region, integrates the historical state information of the search region image, and fuses this part of the features and the interacted features through the feature aggregation module; selectively enhances the relevant features of the target object by the fused features through the feature enhancement module, and then sends them to the prediction head to obtain and mark the maximum response region of the tracking target in the current frame. The present invention fully combines the advantages of Mamba and Transformer to construct a network model based on the Mamba vision hybrid module, and uses the feature aggregation module and the feature enhancement module to improve the performance of the model, so that the tracker has a faster tracking speed and stronger generalization ability, can meet the requirements of different tracking scenarios, and can achieve more accurate and robust tracking.

[0037] Please refer to Figure 5 , the present invention also proposes an object tracking system based on the Mamba vision hybrid module, and the system includes: A construction module, used for: Construct a Mamba vision hybrid module based on the Transformer architecture and the Mamba state space model, construct a feature fusion module based on the feature aggregation mechanism, and construct a feature enhancement module based on the feature enhancement mechanism. The Mamba vision hybrid module, the feature fusion module and the feature enhancement module constitute a tracking model; wherein the Mamba vision hybrid module includes multi-head self-attention and an image state space module; An extraction module, used for: Input the search image and the template image into the Mamba vision hybrid module, extract and superimpose the global context information and the local detail information to obtain the output of the Mamba vision hybrid module; A calculation module, used for: Input the output of the Mamba vision hybrid module into the feature fusion module for fusion to obtain the output of the feature fusion module; A learning module, used for: Enhance the output of the feature fusion module using the feature enhancement module to obtain enhanced features. Calculate the classification loss and regression loss using the enhanced features, and optimize the tracking model using the classification loss and regression loss to obtain an optimized tracking model; A pre-training module, for: Pre-train the optimized tracking model using a large-scale dataset and update the parameters of the tracking model to obtain a tracking model with updated parameters; A tracking module, for: Input the search image and the template image into the tracking model with updated parameters, iteratively extract and superimpose the global context information and local detail information using the Mamba visual mixing module, fuse them through the feature fusion module, and finally enhance the features using the feature enhancement module to obtain the finally enhanced features. Input the finally enhanced features into the prediction head to obtain the tracking result.

[0038] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0039] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0040] The above-described embodiments only represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the present invention's patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention's patent should be subject to the appended claims.

Claims

1. A target tracking method based on the Mamba visual mixing module, characterized in that, The method includes the following steps: Step 1: Construct a Mamba vision hybrid module based on the Transformer architecture and the Mamba state space model, construct a feature fusion module based on a feature aggregation mechanism, and construct a feature enhancement module based on a feature enhancement mechanism. The Mamba vision hybrid module, the feature fusion module, and the feature enhancement module constitute a tracking model; wherein the Mamba vision hybrid module includes a multi-head self-attention and an image state space module; Step 2: Input the search image and the template image into the Mamba vision hybrid module, extract and superimpose the global context information and the local detail information to obtain the output of the Mamba vision hybrid module; Step 3: Input the output of the Mamba vision hybrid module into the feature fusion module for fusion to obtain the output of the feature fusion module; Step 4: Use the feature enhancement module to enhance the features of the output of the feature fusion module to obtain enhanced features. Calculate the classification loss and the regression loss using the enhanced features, and optimize the tracking model using the classification loss and the regression loss to obtain an optimized tracking model; Step 5: Use a large-scale dataset to pre-train the optimized tracking model and update the parameters of the tracking model to obtain a tracking model with updated parameters; Step 6: Input the search image and the template image into the tracking model with updated parameters, and repeat Steps 2 to 4 in an iterative manner to obtain the finally enhanced features. Input the finally enhanced features into the prediction head to obtain the tracking result.

2. The target tracking method based on the Mamba vision hybrid module according to claim 1, characterized in that, In Step 2, when inputting the search image and the template image into the Mamba vision hybrid module, extracting and superimposing the global context information and the local detail information to obtain the output of the Mamba vision hybrid module, it specifically includes the following steps: Input the search image and the template image into the Mamba vision hybrid module for initialization processing respectively to obtain an image sequence of the template region and an image sequence of the search region; Use a convolutional network to extract features from the image sequence of the template region and the image sequence of the search region respectively to obtain the original features of the image sequence of the template region and the original features of the image sequence of the search region; Flatten the original features of the image sequence of the template region and then add position encoding to obtain the original feature token sequence of the template image; Flatten the original features of the image sequence of the search region and then add position encoding to obtain the original feature token sequence of the search region image; Input the original feature token sequence of the template image into the multi-head self-attention for information extraction to obtain the global context features of the image; Input the original feature token sequence of the search region image into the image state space module for local information extraction to obtain the local features of the image; Superimpose the global context features of the image and the local features of the image to obtain the output of the Mamba vision hybrid module.

3. The target tracking method based on the Mamba vision mixing module according to claim 2, characterized in that, When inputting the original feature token sequence of the template image into the multi-head self-attention for information extraction, the relational formula existing in the corresponding process is as follows: ; Among them, represents the intermediate value, represents the attention calculation, represents the input of the Mamba visual mixing module, represents the normalization function, represents the query vector, represents the transpose of the key vector, represents the square root of the key vector dimension, represents the value vector; In the step of inputting the original feature marking sequence of the search area image into the image state space module for local information extraction to obtain the local features of the image, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the state space model the current moment of the input, represents the first learnable matrix, represents the state space model the past moment the information set of all input images, represents the second learnable matrix, represents the third learnable matrix, represents the fourth learnable matrix, represents the matrix multiplication operation, represents the state branch of the image state space module, represents the symmetric convolution branch of the image state space module, represents the first linear mapping process, represents the output of the state branch of the image state space module, represents the activation operation, represents the one-dimensional ordinary convolution operation, represents the second linear mapping process, represents the concatenation operation; In the step of superimposing and processing the global context features of the image and the local features of the image to obtain the output of the Mamba visual mixing module, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the output of the Mamba vision mixing module.

4. The target tracking method based on the Mamba vision hybrid module according to claim 3, wherein, In step 3, the output of the Mamba visual mixing module is input into the feature fusion module for fusion to obtain the output of the feature fusion module, which specifically includes the following steps: The image state space module is used to integrate the historical state information of the search area to obtain the integrated historical state information of the search area; The output of the Mamba visual mixing module is input into the feature fusion module, and a concatenation operation, a convolutional activation operation, and a residual connection operation are sequentially performed with the integrated historical state information of the search area to obtain the output of the feature fusion module.

5. The target tracking method based on the Mamba vision hybrid module according to claim 4, characterized in that, The output of the Mamba visual mixing module is input into the feature fusion module, and a concatenation operation, a convolutional activation operation, and a residual connection operation are sequentially performed with the integrated historical state information of the search area to obtain the output of the feature fusion module. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the features output after splicing operation and linear mapping, represents the fully connected layer, represents the features after the interaction between the template area and the search area, represents the features output by the image state space module, represents the fusion process, represents the batch normalization operation, represents the convolution operation with a convolution kernel size of 1×1, represents the activation function, represents the output of the feature fusion module.

6. The target tracking method based on the Mamba vision hybrid module according to claim 5, wherein In step 4, the output of the feature fusion module is enhanced using the feature enhancement module to obtain enhanced features, which are specifically as follows: S101. Perform global average pooling processing on the spatial dimension of the output of the feature fusion module using the feature enhancement module to obtain the result after global average pooling processing. Perform channel weight assignment processing on the result after global average pooling processing through a convolutional layer to obtain the channel weight matrix of the input features. Multiply the channel weight matrix of the input features by the original matrix of the output of the feature fusion module to obtain the feature matrix output by the first channel branch and the matrix output by the first skip connection; S102. Use the matrix output by the first skip connection as the input and repeat step S101 iteratively to obtain the feature matrix output by the second channel branch, the feature matrix output by the third channel branch, and the feature matrix output by the fourth channel branch respectively; S103. Perform concatenation processing on the feature matrix output by the first channel branch, the feature matrix output by the second channel branch, the feature matrix output by the third channel branch, and the feature matrix output by the fourth channel branch to obtain the first concatenated result. Use a convolutional layer to perform channel dimension weight assignment processing on the first concatenated result to obtain the output feature matrix of the first part structure of the feature enhancement module; S104. Perform global average pooling and max pooling operations on the channel dimension of the output feature matrix of the first part of the feature enhancement module in sequence to obtain the pooled features. Then, perform concatenation and convolutional activation operations on the pooled features along the channel dimension in sequence to obtain the weight matrix for each spatial position of the input feature matrix. Multiply the weight matrix for each spatial position of the input feature matrix by the output feature matrix of the first part of the feature enhancement module to obtain the feature matrix output by the spatial branch of the first layer and the matrix output by the skip connection of the first layer; S105. Take the matrix output by the skip connection of the first layer as the input and repeat step S104 iteratively to obtain the feature matrix output by the spatial branch of the second layer, the feature matrix output by the spatial branch of the third layer, and the feature matrix output by the spatial branch of the fourth layer respectively; S106. Concatenate the feature matrix output by the spatial branch of the first layer, the feature matrix output by the spatial branch of the second layer, the feature matrix output by the spatial branch of the third layer, and the feature matrix output by the spatial branch of the fourth layer along the channel dimension to obtain the second concatenated result. Perform channel dimension weight assignment processing on the second concatenated result through a convolutional layer to obtain the output of the second part for modeling spatial position dependencies; S107. Add the output of the second part for modeling spatial position dependencies to the output of the feature fusion module to obtain the output of the feature enhancement module.

7. The object tracking method based on the Mamba vision mixing module according to claim 6, characterized in that Perform global average pooling on the spatial dimension of the output of the feature fusion module using the feature enhancement module to obtain the result after global average pooling. Perform channel weight assignment processing on the result after global average pooling through a convolutional layer to obtain the channel weight matrix of the input features. Multiply the channel weight matrix of the input features by the original matrix of the output of the feature fusion module to obtain the feature matrix output by the channel branch of the first layer and the matrix output by the skip connection of the first layer. The corresponding relationships in the process are as follows: ; Among them, represents the feature matrix output by the channel branch of the first layer, represents the channel attention calculation, represents the matrix output by the skip connection of the first part of the structure in the first layer; In the step of taking the matrix output by the skip connection of the first layer as the input and repeating step S101 iteratively to obtain the feature matrix output by the channel branch of the second layer, the feature matrix output by the channel branch of the third layer, and the feature matrix output by the channel branch of the fourth layer respectively, the corresponding relationships in the process are as follows: ; Among them, represents the feature matrix output by the channel branch of the second layer, represents the matrix output by the skip connection of the second layer of the first part of the structure, represents the feature matrix output by the channel branch of the third layer, represents the matrix output by the skip connection of the third layer of the first part of the structure, represents the feature matrix output by the channel branch of the fourth layer; In the step of concatenating the feature matrix output by the channel branch of the first layer, the feature matrix output by the channel branch of the second layer, the feature matrix output by the channel branch of the third layer, and the feature matrix output by the channel branch of the fourth layer to obtain the first concatenated result and using a convolutional layer to perform channel dimension weight assignment processing on the first concatenated result to obtain the output feature matrix of the first part of the feature enhancement module, the corresponding relationships in the process are as follows: ; Among them, represents the output feature matrix of the first part of the feature enhancement module.

8. The object tracking method based on the Mamba vision hybrid module according to claim 7, wherein Perform global average pooling and max pooling operations on the channel dimension of the output feature matrix of the first part of the feature enhancement module in sequence to obtain the pooled features. Then, perform concatenation and convolution activation operations on the pooled features along the channel dimension in sequence to obtain the weight matrix for each spatial position of the input feature matrix. Multiply the weight matrix for each spatial position of the input feature matrix by the output feature matrix of the first part of the feature enhancement module to obtain the feature matrix output by the spatial branch of the first layer and the matrix output by the skip connection of the first layer. The relational expressions for the corresponding process are as follows: ; Among them, represents the feature matrix output by the spatial branch of the first layer, represents the channel attention calculation, represents the matrix output by the first-layer skip connection of the second-part structure; In the step of taking the matrix output by the skip connection of the first layer as the input and repeating step S104 iteratively to obtain the feature matrix output by the spatial branch of the second layer, the feature matrix output by the spatial branch of the third layer, and the feature matrix output by the spatial branch of the fourth layer respectively, the relational expressions for the corresponding process are as follows: ; Among them, represents the feature matrix output by the spatial branch of the second layer, represents the matrix output by the skip connection of the second layer of the second part of the structure, represents the feature matrix output by the spatial branch of the third layer, represents the matrix output by the skip connection of the third layer of the second part of the structure, represents the feature matrix output by the spatial branch of the fourth layer; In the step of performing concatenation processing along the channel dimension on the feature matrix output by the spatial branch of the first layer, the feature matrix output by the spatial branch of the second layer, the feature matrix output by the spatial branch of the third layer, and the feature matrix output by the spatial branch of the fourth layer to obtain the second concatenated result, and then performing weight assignment processing on the channel dimension of the second concatenated result through a convolutional layer to obtain the output of the second part for modeling spatial position dependencies, the relational expressions for the corresponding process are as follows: ; Among them, represents the output indicating the position dependence relationship of the second part of the modeling space; In the step of adding the output of the second part for modeling spatial position dependencies to the output of the feature fusion module to obtain the output of the feature enhancement module, the relational expressions for the corresponding process are as follows: ; Among them, represents the output of the feature enhancement module.

9. The target tracking method based on the Mamba vision hybrid module according to claim 7, characterized in that In step 6, input the search image and the template image into the tracking model with updated parameters, and repeat steps 1 to 4 iteratively to obtain the finally enhanced features. The relational expressions for the corresponding process are as follows: ; Among them, represents the template features after feature extraction, represents the search area features after feature extraction, represents the features after concatenating the template features and the search area features along the channel dimension, represents the template features length, represents the image state space module, represents the feature fusion module, represents the feature enhancement module.

10. A target tracking system based on a Mamba vision hybrid module, characterized in that, The system applies the object tracking method based on the Mamba visual hybrid module as described in any one of the above claims 1 to 9. The system includes: A construction module for: Construct a Mamba visual hybrid module based on the Transformer architecture and the Mamba state space model, construct a feature fusion module based on the feature aggregation mechanism, and construct a feature enhancement module based on the feature enhancement mechanism. The Mamba visual hybrid module, the feature fusion module, and the feature enhancement module constitute the tracking model; wherein the Mamba visual hybrid module includes multi-head self-attention and an image state space module; An extraction module for: Input the search image and the template image into the Mamba visual hybrid module, extract and superimpose the global context information and local detail information to obtain the output of the Mamba visual hybrid module; A calculation module for: Input the output of the Mamba visual hybrid module into the feature fusion module for fusion to obtain the output of the feature fusion module; A learning module for: Enhance the features of the output of the feature fusion module using the feature enhancement module to obtain enhanced features, calculate the classification loss and regression loss using the enhanced features respectively, and optimize the tracking model using the classification loss and regression loss to obtain the optimized tracking model; The pre-training module is used for: Pre-training the optimized tracking model using a large-scale dataset and updating the parameters of the tracking model to obtain the tracking model with updated parameters; The tracking module is used for: Inputting the search image and the template image into the tracking model with updated parameters, iteratively extracting and superimposing global context information and local detail information using the Mamba vision mixing module, fusing them through the feature fusion module, and finally enhancing the features using the feature enhancement module to obtain the finally enhanced features, and inputting the finally enhanced features into the prediction head to obtain the tracking result.

Citation Information

Patent Citations

  • Twin network target tracking method and system based on convolutional self-attention module

    CN113705588A

  • Attention-enhanced space-time Transform visual single-target tracking method

    CN117011342A

  • Target tracking method and system based on rotation equivariant network and triple attention mechanism

    CN118096836A

  • Target tracking method and system of twin network based on recursive distraction attention

    CN118781155A

  • High-resolution remote sensing image semantic segmentation method based on multi-scale depth supervision

    CN119559403A

Cited By

  • Target tracking method and system based on Mama and attention mechanism hybrid network

    CN121661100A

  • A Target Tracking Method and System Based on a Hybrid Network of Mamba and Attention Mechanism

    CN121661100B