Object Tracking Method Based on the Fusion of RGB Data and Event Data of Visual Mamba
The visual Mamba method is used to extract, interact and fusion the feature of RGB data and event data, which solves the problem of insufficient target tracking accuracy in dynamic scenarios, and achieves efficient and accurate target tracking effects.
Patent Information
- Application Number
- CN202510613928.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-13
AI Technical Summary
In the prior art, the alignment and fusion efficiency of RGB data and event data is low, the target positioning accuracy is insufficient, making it difficult to achieve efficient and accurate target tracking in dynamic scenarios.
Using a visual Mamba-based method, through frame-level synchronization, dynamic context modeling and efficient modal fusion, combined with two-dimensional selective scanning SS2D mechanism and state space modeling, feature extraction, interaction, fusion and alignment of RGB data and event data is performed to generate cross-modal fusion features for target tracking.
It significantly improves the target tracking accuracy and robustness in dynamic scenarios, and achieves efficient multimodal target tracking performance.
Smart Images

Figure CN120125617B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to an object tracking method based on the fusion of RGB data and event data using Vision Mamba. Background Art
[0002] Object tracking is an important application direction in the field of computer vision. Its task is to accurately locate and track the position of a specified object through consecutive frames of video data, and it is widely used in scenarios such as autonomous driving, intelligent monitoring, robot navigation, and human-computer interaction. It is of great significance for the real-time perception and behavior analysis of objects in a dynamic environment. In actual use, with the increase in the complexity of video data, object tracking faces challenges such as object occlusion, fast movement, lighting changes, and background complexity. Traditional tracking methods that rely on manually designed features and filters are limited in complex scenarios.
[0003] In recent years, deep learning has provided powerful feature extraction capabilities for object tracking, enabling algorithms to exhibit high robustness and accuracy in various scenarios. However, there are still bottlenecks in terms of real-time performance and computational resource requirements. Multi-modal object tracking significantly enhances the robustness of the tracking system by fusing data from different modalities, making it more adaptable to complex scenarios such as occlusion, dynamic changes, and extreme lighting. At the same time, it reduces the complexity of single-modal processing through information extraction and interaction, achieving efficient real-time tracking. This technology has wide application value in fields such as autonomous driving, intelligent transportation, drone vision, and augmented reality. It not only has innovative significance at the algorithm level but also provides a practical basis for multi-modal perception and fusion.
[0004] An event camera is a new type of visual sensor. Different from traditional frame image cameras, it does not capture scenes at a fixed frame rate but triggers event outputs based on pixel-level brightness changes. The event camera records dynamic scene information in the form of high temporal resolution and sparse data, can capture the changes of fast-moving objects in real time, and has the advantages of low latency, high dynamic range, and low power consumption. These characteristics make it perform excellently in dynamic scene perception and extreme lighting conditions. RGB cameras and event cameras are two visual perception devices with complementary characteristics. RGB cameras can capture rich texture and color information but perform poorly in scenes with high dynamic range or low light; event cameras, with their high temporal resolution and sparse data representation characteristics, can efficiently capture motion changes but cannot record the details or color information of static objects.
[0005] In recent years, data fusion of two types of cameras has emerged to achieve more robust visual perception and object tracking. Among them, event cameras supplement motion information in dynamic scenes, while RGB cameras provide rich details of static objects. The fusion at the feature layer or decision layer of the two has significantly improved the tracking performance in extreme scenarios. However, existing methods are mostly based on the Transformer architecture. Although they have high performance, they are limited in application scenarios with high real-time requirements due to their large computational complexity and memory occupancy.
[0006] The State Space Model (SSM) originates from classical dynamic system modeling methods and can effectively describe the dynamic changes of sequential data. Compared with traditional RNN or Transformer architectures, the computational complexity of SSM is linear O(N), and it has significant computational efficiency advantages when processing long-sequence data. Therefore, SSM is widely used in fields such as natural language processing, time series analysis, and computer vision. The state equation of SSM is: x ′( t )= Fx ( t )+ Du ( t )+ ω ( t ), This equation describes the evolution of the internal state x(t) of the system over time. Among them, x′(t) is the derivative with respect to time, representing the rate of change of the system state; F is the system matrix, which describes the relationship between system states and how they naturally evolve over time (when there is no control input), and is an n × n matrix; x(t) is the system state variable, which contains all the necessary variables to describe the condition of the system at any moment and is an n-dimensional vector; D is the input matrix, which describes how the control input affects the system state and is an n×m matrix; u(t) is the control input vector, representing the influence of external input or control signals and is an m-dimensional vector; ω(t) is the process noise, which represents the uncertainty within the system and is generally assumed to be Gaussian noise. The observation equation of SSM is: y ( t )= Mx ( t )+ Gu ( t )+ e ( t ), This equation describes how the output y(t) of the system depends on the system state and control input. Among them, y(t)is the output vector or observation vector, which contains all measured or observed variables and is a p-dimensional vector; M is the output matrix, which describes how the system state affects the output and is a p ≤ n matrix; G is the direct transfer matrix, which represents the direct influence of the control input on the output and is a n×m matrix; e(t) is the measurement noise, which represents the uncertainty or error in the measurement process and is generally assumed to be Gaussian noise.
[0007] As a time series modeling framework based on the state space model, Mamba adopts a selective state space mechanism, which can efficiently model the temporal dynamic characteristics in long sequence data while retaining sufficient expressive power. The application of Mamba in visual tasks further develops the capabilities of the SSM. By introducing a cross-modal learning module, it realizes interactive feature fusion between modalities and effectively improves the performance of multi-modal tracking tasks. In terms of framework design, Mamba adopts a lightweight architecture, significantly reducing memory occupancy and computational burden, making real-time processing on resource-constrained devices possible.
[0008] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0009] The alignment and fusion efficiency of RGB data and event data is low, and the target localization accuracy is insufficient, making it difficult to achieve efficient and accurate target tracking in dynamic scenes. Summary of the Invention
[0010] The purpose of the present invention is to provide a target tracking method for fusing RGB data and event data based on visual Mamba to solve the technical problems in the prior art, such as low alignment and fusion efficiency of RGB data and event data, insufficient target localization accuracy, and difficulty in achieving efficient and accurate target tracking in dynamic scenes. The many technical effects that can be produced by the preferred technical solutions provided by the present invention are described in detail below.
[0011] To achieve the above purpose, the present invention provides the following technical solutions:
[0012] A target tracking method based on the fusion of RGB data and event data of Visual Mamba provided by the present invention includes the following steps: S100: Based on the characteristics of the RGB camera and the event camera, perform frame-level synchronization and conversion on the original RGB image of the RGB camera and the original event stream of the event camera to generate RGB input data and event input data; S200: Use the feature extraction module of the Mamba network to divide the RGB input data and event input data into image blocks of a fixed size to obtain an RGB feature map and an event feature map; S300: Perform linear transformation and depth convolution operations on the RGB feature map and the event feature map through the interaction module, and perform dynamic interaction and restoration of cross-modal features through feature map segmentation, alternating fusion, and two-dimensional selective scanning SS2D to obtain RGB features and event features; S400: Perform two-dimensional selective scanning SS2D, depth convolution, layer normalization, and residual connection operations through the fusion module to dynamically fuse the RGB features and event features within and between modalities, obtain cross-modal fusion features, and output them; S500: Perform search feature segmentation, template feature downsampling, and alternating fusion of search features and template features through the alignment module to align the template features and search features, and reconstruct the cross-modal fusion features to generate an aligned feature map; S600: Based on the aligned feature map, generate a target classification score map, local offset, and normalized bounding box size through the target tracking and detection module to complete target positioning and tracking.
[0013] Preferably, step S100 specifically includes: S110: Use a synchronized RGB camera and event camera to collect target scene data with additional timestamps. The RGB camera generates a continuous frame of the original RGB image, and the event camera generates a sparse original event stream; S120: Read each frame of the original RGB image based on the additional timestamp, extract its timestamp information, and divide the frame into corresponding time windows; S130: Adjust the image size of the original RGB image and perform standardization processing on the color space; S140: Divide the original event stream according to the timestamp of each frame of the original RGB image, filter out the original event data belonging to the time window between two frames, and generate an original event stream in the form of an image; S150: Align the original event stream with the corresponding original RGB image through the timestamp to check the consistency of the original RGB image and the original event stream in spatial resolution and pixel scale, and obtain an RGB image and an event image; S160: Output the RGB image and the event image in tensor form and save them in a standardized data format to obtain RGB input data and event input data.
[0014] Preferably, the step S200 includes: S210: dividing the RGB input data and event input data into local regions (patches) of a fixed size through two convolution operations and mapping them to the target embedding dimension; S220: mapping the embedded representation of the local region (patch) into the state space as the starting point for dynamic update; S230: performing feature extraction and transformation on each local region (patch), jointly updating the current state and the input features, and performing the interaction between the RGB input data and the event input data; S240: outputting the dynamic RGB feature map and event feature map that combine the temporal and spatial dimensions through a mapping operation.
[0015] Preferably, in the two convolution operations of the step S210, the first convolution extracts the low-dimensional features of the RGB input data and event input data and performs normalization processing, then enhances the feature expression through a non-linear activation function, and then extracts the high-dimensional features of the RGB input data and event input data through the second convolution.
[0016] Preferably, the step S300 specifically includes: S310: enhancing the features of the RGB feature map and event feature map through a linear layer and a depth convolution layer; S320: dividing the enhanced features into RGB feature blocks and event feature blocks of a fixed size, where the RGB feature blocks and event feature blocks capture the key features of the local region (patch) and retain the global context information; S330: inputting the RGB feature blocks and event feature blocks into the two-dimensional selective scan (SS2D) module in an alternating order to perform dynamic interaction between modalities and obtain the fused cross-modal features; S340: restoring the cross-modal features to the single-modal features of the event modality and RGB modality, and respectively balancing the feature distribution through a normalization operation; S350: respectively mapping the features of the normalized RGB modality and event modality through a linear layer to generate RGB features and event features.
[0017] Preferably, the step S400 specifically includes: S410: transforming the RGB feature and the event feature through a linear layer to obtain the features after linear transformation; S420: processing the features after linear transformation through depthwise separable convolution to enhance the feature expression ability within the modality, and obtaining the preliminarily enhanced RGB modality feature and the preliminarily enhanced event modality feature; S430: inputting the preliminarily enhanced RGB modality feature and the preliminarily enhanced event modality feature into the two-dimensional selective scanning SS2D module to respectively integrate spatial information and temporal information; S440: through layer normalization operation, making the features of the two modalities reach consistency in numerical distribution, and multiplying the modality features of the two modalities after layer normalization with the input features of this module pixel by pixel, and screening and highlighting the main information through the attention mechanism; S450: splicing the RGB modality feature and the event modality feature in the channel dimension to generate a unified cross-modal feature; S460: inputting the spliced cross-modal feature into the visual state space VSSB module for dynamic fusion processing and outputting the cross-modal fusion feature.
[0018] Preferably, the step S500 specifically includes: S510: linearly transforming the feature dimensions of the template feature and the search feature through a linear layer to reduce redundant information and retain the main semantic features; S520: slicing the search feature into feature blocks of a fixed size to capture local features and retain part of the global context information, and reducing the size of the template feature through downsampling to unify the feature scales of the template feature and the search feature; S530: alternately inputting the segmented search feature blocks and the downsampled template feature into the alternating fusion module to integrate feature information through the dynamic interaction mechanism between modalities; S540: performing depth integration on the features after alternating fusion through two-dimensional selective scanning SS2D to obtain the fused feature representation of the template feature and the search feature; S550: performing an upsampling operation on the template feature to restore the original size of the feature map, and directly merging the search features; S560: respectively performing layer normalization processing on the restored search feature and the template feature to balance the numerical distribution of the features, and further enhancing the features through a linear layer; S570: inputting the normalized search feature and the initial cross-modal fusion feature for a residual operation to generate an aligned feature map.
[0019] Preferably, the step S600 specifically includes: S610: Taking the search feature as the feature input of the target tracking and detection module, and reinterpret the marked filling sequence of the search area as a two-dimensional spatial feature map to provide complete spatial distribution information; S620: Inputting the aligned feature map into the fully convolutional network FCN, generating the target classification score map, local offset, and predicting the bounding box size of the target through the fully convolutional network FCN, and the output is a normalized value; S630: Through the target classification score map, select the position with the highest classification score as the target center point, and at the same time correct the position by combining the local offset to determine the final coordinates of the target.
[0020] Preferably, in the step S610, three loss functions are used to jointly optimize the target tracking and detection module: the Focal Loss function balances the positive and negative samples in target classification; the L1 Loss function is used for the accurate regression of the bounding box offset to optimize the error between the predicted box and the ground truth box; the GIoU Loss function is used to measure the overlap degree between the predicted box and the ground truth box to improve the accuracy of target localization.
[0021] Preferably, in the step S620, the fully convolutional network FCN includes L layers of convolution, batch normalization, and ReLU activation. Through multi-layer convolution operations, the features are further extracted and compressed to generate multiple target-related outputs, including the target classification score map, local offset, and normalized bounding box size.
[0022] Implementing one of the above technical solutions of the present invention has the following advantages or beneficial effects:
[0023] Through frame-level synchronization, dynamic context modeling, and efficient modality fusion, and at the same time, combined with the two-dimensional selective scanning SS2D mechanism, state space modeling, and multi-modal interaction operations, the present invention realizes efficient processing in dynamic scenarios. Through efficient modality fusion and state space model, the target tracking accuracy and robustness in dynamic scenarios are significantly improved, and good multi-modal target tracking performance is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings. In the drawings:
[0025] Figure 1 is the flowchart of the target tracking method for the fusion of RGB data and event data based on Vision Mamba in the embodiment of the present invention;
[0026] Figure 2It is a flowchart of step S100 of the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention;
[0027] Figure 3 It is a flowchart of step S200 of the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention;
[0028] Figure 4 It is a flowchart of step S300 of the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention;
[0029] Figure 5 It is a flowchart of step S400 of the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention;
[0030] Figure 6 It is a flowchart of step S500 of the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention;
[0031] Figure 7 It is a flowchart of step S600 of the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention;
[0032] Figure 8 It is a schematic diagram of the two-dimensional selective scanning SS2D module in the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention;
[0033] Figure 9 It is a schematic diagram of the visual state space VSSB module in the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention;
[0034] Figure 10 It is a schematic diagram of the interaction module in the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention;
[0035] Figure 11 It is a schematic diagram of the fusion module in the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention;
[0036] Figure 12 It is a schematic diagram of the alignment module in the object tracking method based on the fusion of RGB data and event data of Visual Mamba in the embodiments of the present invention. Detailed implementation manners
[0037] To make the objectives, technical solutions and advantages of the present invention more clearly understood, various exemplary embodiments to be described hereinafter will refer to the corresponding drawings, which form a part of the exemplary embodiments and illustrate various exemplary embodiments that may be used to implement the present invention. Unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. It should be understood that they are merely examples of processes, methods, devices, etc. consistent with some aspects of the present invention disclosed in detail in the appended claims. Other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present invention.
[0038] In the description of the present invention, it should be understood that terms such as "center", "longitudinal", "lateral", etc. indicate the orientation or positional relationship based on the drawings shown, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the elements referred to must have a specific orientation, be constructed and operated in a specific orientation. Terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the technical features indicated. The meaning of the term "plurality" is two or more. The terms "connected" and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, an integral connection, a mechanical connection, an electrical connection, a communication connection, a direct connection, an indirect connection through an intermediate medium, and may be the communication inside two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention may be understood according to specific circumstances.
[0039] To illustrate the technical solution of the present invention, the following will be described by specific embodiments, and only the parts related to the embodiments of the present invention are shown.
[0040] Embodiment:
[0041] Such as Figure 1As shown, the present invention provides an object tracking method based on the fusion of RGB data and event data of Visual Mamba. In this embodiment, Visual Mamba includes a feature extraction module, an interaction module, a fusion module, an alignment module, and an object tracking detection module. Mamba is based on the selective state space model and shows higher efficiency and performance when processing long sequences. In this embodiment, the training data collected in step S100 is used for end-to-end training. After the model training is completed, the trained model weights are used to test the test data collected in step S100. After the test is completed, it is deployed in the actual application scenario. The method includes the following steps: S100: Based on the characteristics of the RGB camera and the event camera, frame-level synchronization and conversion are performed on the original RGB image of the RGB camera and the original event stream of the event camera to generate RGB input data and event input data, that is, the original RGB image and the original event stream are collected and preprocessed. This process ensures the accurate alignment of the two modalities of data in the time and space dimensions and provides a reliable data basis for subsequent processing. S200: Operate based on the feature extraction module. Specifically, the RGB input data and the event input data are sliced into image blocks of a fixed size through a chunking operation to obtain an RGB feature map and an event feature map. S300: Operate based on the interaction extraction module. Specifically, linear transformation and depth convolution operations are performed on the RGB feature map and the event feature map, and dynamic interaction and restoration of cross-modal features are performed through feature map slicing, alternating fusion, and two-dimensional selective scanning SS2D to obtain RGB features and event features. S400: Operate based on the fusion module. Specifically, through two-dimensional selective scanning SS2D, depth convolution, layer normalization, and residual connection operations, dynamic fusion of RGB features and event features within and between modalities is performed to obtain cross-modal fusion features and output them. S500: Operate based on the alignment module. Specifically, alignment of the template feature and the search feature is performed through search feature slicing, template feature downsampling, and alternating fusion of the search feature and the template feature, and the cross-modal fusion feature is reconstructed to generate an aligned feature map. S600: Operate based on the object tracking detection module. Specifically, an object classification score map, a local offset, and a normalized bounding box size are generated based on the aligned feature map to complete object positioning and tracking. Through frame-level synchronization, dynamic context modeling, and efficient modality fusion, the present invention significantly improves the robustness and accuracy of object tracking. At the same time, combined with the two-dimensional selective scanning SS2D mechanism, state space modeling, and multi-modal interaction operations, efficient processing in dynamic scenarios is achieved. Through efficient modality fusion and state space model, the object tracking accuracy and robustness in dynamic scenarios are significantly improved, and good multi-modal object tracking performance is achieved.
[0042] As an optional implementation, such as Figure 2As shown in the figure, the S100 step specifically includes: S110: Use synchronized RGB cameras and event cameras to collect target scene data with additional timestamps. The synchronized cameras facilitate keeping the time consistent. The RGB camera is used to record the color and texture information of the target and generate the original RGB images of consecutive frames. The event camera is used to capture the changes in pixel brightness and generate a sparse original event stream. S120: Based on the additional timestamps, read each frame of the original RGB image, extract its timestamp information, and divide the frames into corresponding time windows. S130: Adjust the image size of the original RGB image and perform standardization processing on the color space to provide a consistent data basis for subsequent processing. S140: Divide the original event stream according to the timestamps of each frame of the original RGB image, filter out the original event data belonging to the time window between two frames, initialize an event map matrix matching the target resolution as the starting state, and encode the pixel points according to the event polarity. For example, assign blue to the pixel points with increasing polarity and red to the pixel points with decreasing polarity, and finally generate the original event map in the form of an image. S150: Align the original event map with the corresponding original RGB image through timestamps to ensure that the original event stream and the original RGB image are consistent in the time dimension, check the consistency of the original RGB image and the original event map in spatial resolution and pixel scale, verify the correct alignment of the multimodal data, and obtain the RGB image and the event map. S160: Output the RGB image and the event map in the form of tensors and save them in a standardized data format to obtain the RGB input data and the event input data. After adopting the standardized data format, all data is stored by frame, supporting batch loading and efficient training or inference processes, ensuring that multimodal input data can be quickly and accurately obtained in subsequent algorithms. Thus, the spatio-temporal information in the dynamic scene is captured, and unified and standardized multimodal input data is generated, facilitating subsequent processing. Preferably, a network is built on the basis of data collection and preprocessing to implement the training stage and the testing stage. In the training stage, the first frame of each video collected is used as the original template image, and subsequent image frames are randomly selected as the original search images. Each frame is divided into event-modal data and RGB-modal data. In the testing stage, the first frame of each video is used as the original template image, and all subsequent frames are used as the original search images.
[0043] As an alternative implementation, the feature extraction module of the Mamba network further includes a two-dimensional selective scanning SS2D module and a visual state space VSSB module. As shown in the figure, the schematic diagram of the two-dimensional selective scanning SS2D module for context modeling of the feature map in four directions; as shown in the figure, the visual state space VSSB module is constructed using the two-dimensional selective scanning SS2D module and is used for in-depth feature extraction; as shown in the figure, in the Mamba backbone network, in-depth feature extraction is performed through the visual state space VSSB module. As Figure 3As shown in the figure, step S200 includes: S210: The RGB input data and event input data are segmented into local regions (patches) of a fixed size through two convolution operations and mapped to the target embedding dimension. S220: Initialize the state variable, map the embedded representation of the local region (patch) into the state space as the starting point for dynamic update, providing an initial representation in the time dimension for the input data; the initial state is generated through a convolution-based mapping operation and serves as the starting point for the dynamic update of the model, providing an initial representation in the time dimension for the input data. S230: Extract features and perform transformations on each local region (patch), gradually reduce the spatial resolution and increase the channel dimension through downsampling operations, jointly update the current state and the input features, integrate the modal characteristics and global context, perform the interaction between the RGB input data and event input data, and achieve the efficient fusion of dynamic information and static details, making the state estimation more accurate. S240: Output dynamic RGB feature maps and event feature maps that combine the time and space dimensions through a mapping operation, thereby realizing the dynamic information that combines the time and space dimensions, making the output features have higher expressiveness and semantic consistency, and providing high-quality inputs for subsequent fusion and classification tasks.
[0044] As an optional implementation method, in the two convolution operations of step S210, the first convolution extracts the low-dimensional features of the RGB input data and event input data and performs normalization processing, then enhances the feature expression through a non-linear activation function, and then extracts the high-dimensional features of the RGB input data and event input data through the second convolution.
[0045] As an optional implementation method, as Figure 4 , Figure 10 shown in the figure, step S300 specifically includes: S310: The RGB feature maps and event feature maps are enhanced through a linear layer and a depth convolution layer, making the features have a more compact dimensional representation and higher semantic expression ability. S320: The enhanced features are segmented into RGB feature blocks and event feature blocks of a fixed size. The RGB feature blocks and event feature blocks capture the key features of the local region (patch) and retain the global context information. S330: The RGB feature blocks and event feature blocks are input into the two-dimensional selective scan (SS2D) module in an alternating order for dynamic interaction between modalities, obtaining the fused cross-modal features, thereby realizing the dynamic interaction between modalities. This alternating input strategy effectively fuses the features of the two modalities and improves the semantic consistency and expression ability of the multi-modal features. S340: Restore the cross-modal features to the single-modal features of the event modality and RGB modality, and balance the feature distribution through normalization operations respectively to maintain the stability and consistency of the output features; S350: Map the features of the normalized RGB modality and event modality through a linear layer respectively, and add the input features based on the residual mechanism to generate RGB features and event features.
[0046] As an alternative implementation, such as Figure 5 , Figure 11 shown, the steps of S400 specifically include: S410: Transform the RGB features and event features through a linear layer to obtain the linearly transformed features; S420: Process the linearly transformed features through depthwise separable convolution to enhance the feature expression ability within the modality, and obtain the preliminary enhanced RGB modality features and preliminary enhanced event modality features; S430: Input the preliminary enhanced RGB modality features and preliminary enhanced event modality features into the two-dimensional selective scanning SS2D module to integrate spatial information and temporal information respectively; The two-dimensional selective scanning SS2D module deeply processes the features within the modality, generates modality features with high expression ability, and retains the key characteristics of each modality, providing high-quality feature input for subsequent modality interaction. S440: Through layer normalization operation, make the features of the two modalities reach consistency in numerical distribution, and multiply the modality features of the two modalities after layer normalization with the input features of this module pixel by pixel, and screen and highlight the main information through the attention mechanism; S450: Concatenate the RGB modality and event modality features in the channel dimension to generate unified cross-modal features; S460: Input the concatenated cross-modal features into the Visual State Space Block (VSSB module). The Visual State Space Block is as Figure 9 shown, perform dynamic fusion processing, output cross-modal fusion features, retain the uniqueness of the modality features, and achieve the efficient integration of RGB features and event features.
[0047] As an alternative implementation, such as Figure 6 , Figure 12 shown, the steps of S500 specifically include: S510: Linearly transform the feature dimensions of the template features (features obtained after processing the original template image) and search features (features obtained after processing the original search image) through a linear layer to reduce redundant information and retain the main semantic features; The depthwise separable convolution can also be used to extract the local features of the image to enhance the spatial resolution ability. S520: Split the search features into feature blocks of a fixed size, capture local features and retain part of the global context information, and reduce the size of the template features through downsampling to unify the feature scales of the template features and search features. S530: Alternately input the segmented search feature blocks and the downsampled template features into the alternating fusion module to integrate the feature information through the dynamic interaction mechanism between modalities; S540: As Figure 8As shown, the depth integration of the features after alternating fusion is performed through two-dimensional selective scanning SS2D to obtain the fused feature representation of the template feature and the search feature; S550: Upsample the template feature to restore the original size of the feature map, and directly perform feature merging on the search feature; S560: Normalize the restored search feature and the template feature layer by layer to balance the numerical distribution of the features, and further enhance the features through a linear layer to highlight the key feature information. S570: Input the normalized search feature and the initial cross-modal fusion feature into a residual operation to generate an aligned feature map.
[0048] As an optional implementation, as Figure 7 shown, step S600 specifically includes: S610: Use the search feature as the feature input of the target tracking and detection module, and reinterpret the marked filling sequence of the search area as a two-dimensional spatial feature map to provide complete spatial distribution information; thus, the feature is restored to the spatial domain to provide a basis for subsequent convolution operations. S620: Input the aligned feature map into the fully convolutional network FCN, and generate the target classification score map, local offset, and predict the boundary box size of the target through the fully convolutional network FCN, and the output is a normalized value; S630: Through the target classification score map, select the position with the highest classification score as the target center point, and at the same time correct the position by combining the local offset, and use the Kalman filter to associate the detected target to determine the final coordinates of the target, thereby completing the target positioning and tracking.
[0049] As an optional implementation, in step S610, three loss functions are used to jointly optimize the target tracking and detection module: The Focal Loss function balances the positive and negative samples in target classification and improves the performance of small target detection; The L1Loss function is used for the precise regression of the boundary box offset to optimize the error between the predicted box and the ground truth box; The GIoU Loss function is used to measure the overlap degree between the predicted box and the ground truth box to improve the accuracy of target positioning.
[0050] As an alternative implementation, in step S620, the fully convolutional network FCN includes L layers of convolution, batch normalization, and ReLU activation. Through multi-layer convolution operations, features are further extracted and compressed to generate multiple outputs related to the target, including the target classification score map, local offset, and normalized bounding box size. The target classification score map is used to evaluate the probability of the target's presence at each spatial position. The score value ranges from [0, 1]. The position with the highest score is considered the center point of the target, and its maximum value corresponds to the probability of the target's presence. The local offset is used to compensate for the discretization error caused by the reduced resolution. The offset value ranges from [0, 1). By adjusting the subtle deviation of the target position, more accurate positioning is ensured. The normalized bounding box size ranges from [0, 1] and is used to provide the scale information of the target to generate the final detection box.
[0051] The embodiment is only a special case and does not indicate that the present invention has only such an implementation.
[0052] The above are only the preferred embodiments of the present invention. Those skilled in the art will know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the protection scope of the present invention.
Claims
1. A target tracking method for fusing RGB data and event data based on Vision Mamba, characterized in that It includes the following steps: S100: Based on the characteristics of the RGB camera and the event camera, perform frame-level synchronization and conversion on the original RGB image of the RGB camera and the original event stream of the event camera to generate RGB input data and event input data; S200: Use the feature extraction module of the Mamba network to slice the RGB input data and event input data into image blocks of a fixed size to obtain an RGB feature map and an event feature map; S300: Perform linear transformation and depth convolution operations on the RGB feature map and the event feature map through the interaction module, and perform dynamic interaction and restoration of cross-modal features through feature map slicing, alternating fusion, and two-dimensional selective scanning SS2D to obtain RGB features and event features; S400: Through the fusion module, perform two-dimensional selective scanning SS2D, depth convolution, layer normalization, and residual connection operations to dynamically fuse the RGB features and event features within and between modalities, obtain cross-modal fusion features and output them; S500: Through the alignment module, perform search feature slicing, template feature downsampling, and alternating fusion of the search feature and the template feature to align the template feature and the search feature, and reconstruct the cross-modal fusion feature to generate an aligned feature map; S600: Based on the aligned feature map, generate a target classification score map, local offset, and normalized bounding box size through the target tracking and detection module to complete target positioning and tracking; The specific steps of S400 include: S410: Transform the RGB features and event features through a linear layer to obtain the features after linear transformation; S420: Process the features after linear transformation through depthwise separable convolution to enhance the feature expression ability within the modality, and obtain the preliminary enhanced features of the RGB modality and the preliminary enhanced features of the event modality; S430: Input the preliminary enhanced features of the RGB modality and the preliminary enhanced features of the event modality into the two-dimensional selective scanning SS2D module to integrate spatial information and temporal information respectively; S440: Through layer normalization operation, make the features of the two modalities reach consistency in numerical distribution, and multiply the layer-normalized modality features of the two modalities with the input features of this module pixel by pixel, and screen and highlight the main information through the attention mechanism; S450: Concatenate the RGB modality and event modality features in the channel dimension to generate a unified cross-modal feature; S460: Input the concatenated cross-modal feature into the visual state space VSSB module for dynamic fusion processing and output the cross-modal fusion feature.
2. The object tracking method for fusing RGB data and event data based on Visual Mamba according to claim 1, wherein, The specific steps of S100 include: S110: Use a synchronized RGB camera and event camera to collect target scene data with additional timestamps. The RGB camera generates the original RGB images of consecutive frames, and the event camera generates a sparse original event stream; S120: Read each frame of the original RGB image based on the additional timestamp, extract its timestamp information, and divide the frame into corresponding time windows; S130: Adjust the image size of the original RGB image and perform standardization processing on the color space; S140: Divide the original event stream according to the timestamps of each frame of the original RGB image, filter out the original event data belonging to the time window between two frames, and generate an original event map in the form of an image; S150: Align the original event map with the corresponding original RGB image through timestamps to check the consistency of the original RGB image and the original event map in spatial resolution and pixel scale, and obtain the RGB map and the event map; S160: Output the RGB map and the event map in tensor form and save them in a standardized data format to obtain the RGB input data and the event input data.
3. A target tracking method for fusing RGB data and event data based on Visual Mamba according to claim 1, characterized in that, Step S200 includes: S210: Cut the RGB input data and the event input data into local regions (patches) of a fixed size through two convolutional operations and map them to the target embedding dimension; S220: Map the embedded representation of the local region (patch) into the state space as the starting point for dynamic update; S230: Extract and transform features for each local region (patch), jointly update the current state and the input features, and perform the interaction between the RGB input data and the event input data; S240: Output the dynamic RGB feature map and the event feature map that combine time and space dimensions through a mapping operation.
4. The object tracking method for fusing RGB data and event data based on Visual Mamba according to claim 3, characterized in that, In the two convolutional operations of step S210, the first convolution extracts the low-dimensional features of the RGB input data and the event input data and performs normalization processing, then enhances the feature expression through a non-linear activation function, and then extracts the high-dimensional features of the RGB input data and the event input data through the second convolution.
5. A target tracking method for fusing RGB data and event data based on Visual Mamba according to claim 1, characterized in that, Step S300 specifically includes: S310: Enhance the features of the RGB feature map and the event feature map through a linear layer and a depth convolutional layer; S320: Cut the enhanced features into RGB feature blocks and event feature blocks of a fixed size. The RGB feature blocks and the event feature blocks capture the key features of the local region (patch) and retain the global context information; S330: Input the RGB feature blocks and the event feature blocks into the two-dimensional selective scanning (SS2D) module in an alternating order for dynamic interaction between modalities to obtain the fused cross-modal features; S340: Restore the cross-modal features to single-modal features of the event modality and the RGB modality, and balance the feature distribution through normalization operations respectively; S350: Map the features of the normalized RGB modality and event modality through a linear layer respectively to generate RGB features and event features.
6. The object tracking method based on the fusion of RGB data and event data of Vision Mamba according to claim 1, characterized in that, Step S500 specifically includes: S510: Perform a linear transformation on the feature dimensions of the template feature and the search feature through a linear layer to reduce redundant information and retain the main semantic features; S520: Cut the search feature into feature blocks of a fixed size, capture the local features and retain part of the global context information, and reduce the size of the template feature through downsampling to unify the feature scales of the template feature and the search feature; S530: Alternately input the segmented search feature blocks and the downsampled template feature into the alternating fusion module to integrate the feature information through the dynamic interaction mechanism between modalities; S540: The features after alternating fusion are deeply integrated through two-dimensional selective scanning SS2D to obtain the fused feature representation of the template feature and the search feature; S550: The template feature is upsampled to restore the original size of the feature map, and the search feature is directly subjected to feature merging; S560: The restored search feature and the template feature are respectively subjected to layer normalization processing to balance the numerical distribution of the features, and the features are further enhanced through a linear layer; S570: The normalized search feature and the template feature and the initial cross-modal fusion feature input are subjected to a residual operation to generate an aligned feature map.
7. A target tracking method based on the fusion of RGB data and event data of Visual Mamba according to claim 1, characterized in that, The steps of S600 specifically include: S610: The search feature is used as the feature input of the target tracking and detection module, and the marked padding sequence in the search area is reinterpreted as a two-dimensional spatial feature map to provide complete spatial distribution information; S620: The aligned feature map is input into the fully convolutional network FCN. Through the fully convolutional network FCN, a target classification score map, a local offset, and the size of the target bounding box are generated, and the output is a normalized value; S630: Through the target classification score map, the position with the highest classification score is selected as the target center point, and the position is corrected by combining the local offset to determine the final coordinates of the target.
8. A target tracking method for fusing RGB data and event data based on Visual Mamba according to claim 7, characterized in that, In step S610, three loss functions are used to jointly optimize the target tracking and detection module: The Focal Loss function balances the positive and negative samples in target classification; The L1 Loss function is used for the precise regression of the bounding box offset to optimize the error between the predicted box and the ground truth box; The GIoU Loss function is used to measure the overlap degree between the predicted box and the ground truth box to improve the accuracy of target localization.
9. A target tracking method for fusing RGB data and event data based on Visual Mamba according to claim 7, characterized in that, In step S620, the fully convolutional network FCN includes L layers of convolution, batch normalization, and ReLU activation. Through multi-layer convolution operations, the features are further extracted and compressed to generate multiple target-related outputs, including the target classification score map, the local offset, and the normalized bounding box size.
Citation Information
Patent Citations
Multi-modal visual target tracking method based on self-distillation symmetric adapter
CN117710414A
Target tracking method fusing RGB / EVENT bimodal samples
CN118334078A