Multi-modal target tracking method based on state space model

By adopting a state-space model-based method in RGB-event multimodal visual tracking, combining RGB and event data, extracting and fusing multimodal features, and using historical information decoder, the tracking problem under harsh motion and harsh lighting conditions is solved, and efficient target tracking and robustness are achieved.

CN120070504AActive Publication Date: 2025-05-30DALIAN UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510220834.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing RGB-event multimodal visual tracking methods perform poorly under fast motion and harsh lighting conditions, and ignore the mining and application of time information, resulting in poor performance in scenarios such as severe change in target appearance or loss of field of view.

Method used

Using a multimodal visual tracking method based on the state space model, by combining RGB images and event data, using a vertical architecture attention network HiViT and a modal fusion module FM based on the state space model, multimodal features are extracted and fused, and the movement trend and appearance changes of the target are perceived through a historical information decoder.

Benefits of technology

It significantly improves the target tracking performance under fast motion and harsh lighting conditions, effectively captures the changes in target state and movement trends, and enhances the robustness of the model and its performance in long-term sequence modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070504A_ABST
    Figure CN120070504A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and relates to a multi-modal target tracking method based on a state space model. The invention provides a multi-modal visual tracking method based on a state space model for a moving target visual tracking task under the conditions of rapid movement and severe illumination, so as to realize accurate tracking of a visual target. The method fully combines the advantages of the RGB image and the event data: the RGB image provides abundant texture information, and the event camera can still capture the edge and motion information of an object in a complex scene, so that the tracking performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and relates to a method for motion target tracking based on deep learning, combining event camera and RGB camera images. Background Art

[0002] Visual target tracking is a fundamental task in computer vision. The main goal is to estimate the position and shape of a target based on its initial state in a video sequence. Due to the advantages of high temporal resolution and high dynamic range of event cameras, in recent years, multi-modal methods combining event cameras with traditional RGB cameras have received increasing attention in visual tracking tasks. These RGB-event multi-modal methods have significantly improved the model tracking performance under challenging conditions such as fast motion and low light.

[0003] Although the above methods perform well in improving the tracking effect, they often require a lot of effort to develop overly complex fusion algorithms. In addition, existing RGB image-event image tracking algorithms usually neglect the mining and application of temporal information, which is crucial for target tracking because it can capture the appearance changes of the target and its motion trends. Existing RGB visual trackers update appearance information by replacing or fusing templates, but these simple methods sometimes have difficulty effectively filtering out low-quality images, thus affecting the performance. To solve this problem, some attention mechanism-based methods introduce an autoregressive mechanism for transmitting temporal information between multiple search frames. However, the attention mechanism itself lacks a design specifically for time series modeling. Especially when dealing with long sequences, the above problems become the bottleneck of the model performance:

[0004] (1) State Space Model (SSM)

[0005] The general form of the state space model can be expressed as: mapping a one-dimensional input sequence to a one-dimensional output sequence through an implicit high-dimensional state sequence. Compared with traditional convolutional neural networks and attention mechanisms, the state space model shows unique advantages in modeling global information. Recent research such as the Mamba model innovatively introduces data-dependent state space parameters and a parallel scan selection mechanism on this basis, significantly improving the inference speed and overall performance, exceeding attention mechanism models of the same scale.

[0006] Recently, the state space model has been widely used in the field of computer vision and has attracted extensive attention. Its characteristics are naturally compatible with the transmission of temporal information in target tracking. However, there is currently no effective research applying the state space model to the field of visual target tracking.

[0007] (2) Tracking Based on Historical Temporal Information

[0008] Historical information can provide additional prior information about the movement, appearance, and its changing trends of the target, thereby enhancing the robustness of the tracker. Tracking methods that utilize historical temporal information can be roughly divided into two categories. The first category of methods processes single-frame search images, relying on dynamic filtering or additional update branches to capture the changes in the target appearance over time, and dynamically updates the template image during the tracking process. However, such methods usually require carefully designed strategies and manual adjustment according to different application scenarios. The second category of methods inputs multiple search sequences and gradually decodes the entire trajectory sequence to perform temporal modeling on the inter-frame target movement trajectory. However, such methods are limited by the limitations of the attention mechanism architecture in long-term sequence modeling.

[0009] (3) RGB-Event Information Multimodal Tracking

[0010] In recent years, due to the unreliability of RGB images in extreme scenarios, more and more research has introduced event cameras to improve the performance of object tracking. These multimodal trackers are mainly divided into two categories according to the visual backbone networks adopted. The first category of methods relies on traditional convolutional neural networks (CNNs) to extract features from RGB and event modalities and fuse them through a carefully designed cross-modal attention mechanism. The second category of methods uses the attention mechanism as the backbone network and models the relationship between RGB image information and event information using the multi-head attention mechanism. However, these methods often overemphasize the architecture design of multimodal fusion and neglect the mining of historical temporal information, resulting in poor performance in scenarios where the target appearance changes drastically or the field of view is lost. Summary of the Invention

[0011] The present invention aims at the visual tracking task of moving targets under fast motion and poor lighting conditions, and proposes a multimodal visual tracking method based on a state space model to achieve precise tracking of visual targets. This method fully combines the advantages of RGB images and event data: RGB images provide rich texture information, while event cameras can still capture the edges and motion information of objects in complex scenarios, thereby improving the tracking performance.

[0012] The technical solution of the present invention:

[0013] A multimodal object tracking method based on a state space model, the steps are as follows:

[0014] Step 1, RGB-Event Image Data Processing

[0015] First, collect image data in various challenging scenarios through a calibrated event camera and an RGB camera. Stack the event streams generated by the event camera into event image frames according to a certain time step and perform image alignment and screening according to the RGB camera frequency to obtain event images. Then, send the preprocessed RGB images and event images into the model.

[0016] Step 2: Construct a multimodal encoder

[0017] The multimodal encoder is used to extract information from RGB images and event images. The attention network HiViT with a vertical architecture is adopted as the multimodal encoder, and features are extracted layer by layer through token encoding, token merging, and self-attention mechanisms. Between the intermediate attention layers, a modality fusion module FM based on the state space model is inserted for cross-modal fusion. Finally, the features of the final layer are fed into the historical information decoder and the prediction head for the next prediction.

[0018] Step 2.1: Construction of the modality fusion module based on the state space model

[0019] Construct a modality fusion module FM based on the state space model to fuse the semantic information of multiple modalities. The target-aware scanning module and the cross-modal scanning module are respectively introduced into the FM module.

[0020] Furthermore:

[0021] (1) Target-aware scanning module

[0022] The target-aware scanning module utilizes the global modeling ability of the state space model to enhance the interaction between the template region and the search region within the modality. The target-aware scanning module adopts a bidirectional state space model scanning mechanism to calculate the correlation between the template feature and the search feature;

[0023] Definition of the target-aware scanning module:

[0024]

[0025] Wherein, represents a linear projection layer; represents a layer normalization operation; σ is the SiLU activation function; ψ is a 1D convolution operation; ⊙ represents element-wise multiplication; represents the feature of a certain layer in the multimodal encoder, respectively represent the RGB image feature and the event image feature, is 's projection; is the gating signal; represents the forward scanning state space model, represents the backward scanning state space model, Y f and Y b are the outputs through the forward and backward directions of the SSM (i.e., and ). is the output after the target perception scanning module jointly models the template region features and the search region features in the n-th layer encoder, represents the template region features in the n-th layer encoder, represents the search region features in the n-th layer encoder.

[0026] (2) Cross-modal scanning module

[0027] A cross-modal scanning module is established based on the gating mechanism of the state space model to perform cross-fusion on the two modalities. The definition of the cross-modal scanning module is as follows:

[0028]

[0029] Taking the RGB modality image as an example, where, represents the features of the RGB image after being encoded by the target perception scanning module in the n-th layer, represents the RGB image search region features in the n-th layer, Y ′ represents the intermediate features after being scanned by the state space model, represents the fused features extracted from the RGB image in the n-th layer after passing through the cross-modal scanning module. The modality-specific gating signals and are used to guide the interaction between modalities, thereby enhancing the main modality while suppressing the degraded modality.

[0030] Through the target perception scanning module and the cross-modal scanning module, the multi-modal encoder can promote intra-modal and cross-modal interactions. The RGB and event features of the last layer encoder are added together to obtain the final fused feature Final fused feature will be used to predict the target position through the historical information decoder.

[0031] Step 3: Construction of the historical information decoder

[0032] The historical information decoder perceives the motion trend and appearance changes of the target by decoding the historical information, thereby assisting the model in making accurate predictions. Each layer of the historical information decoder includes a historical state perception module and a feature sequence attention module:

[0033] Furthermore, (1) Historical state perception module

[0034] The historical state perception module is used to model the historical information in the video sequence. By using a sequence modeling module based on the state space model, the target states of N frames in the sequence are modeled as N one-dimensional learnable vectors Q(q 1 ,…,q N )∈R N×D, where D is the encoding dimension. The sequence is modeled through a state space model and Q is output ′ =(q ′ ,…,q ′ N )∈R N×D , and the formula is as follows:

[0035]

[0036] Among them, X Q represents the input feature after the projection of the query Q, and Z Q represents the gating feature after the projection of the query Q. represents the state space model. is a linear projection layer, is a normalization layer,

[0037] Output Q ′ will then be used as the input of the feature sequence attention module. The target query vector Q ′ will then be fused with the multi-modal features in step 2 through the feature sequence attention module to accurately locate the target in the current search area.

[0038] (2) Definition of the feature sequence attention module:

[0039]

[0040] Among them, Q ″ represents the historical information feature after being scanned by the historical state perception module, represents the vector concatenation operation, and H i represents the intermediate feature after self-attention. represents the attention mechanism, q represents the query of the attention, k represents the key of the attention, and v represents the value of the attention. And are all learnable parameters, d m represents the input dimension d q , d k ,, and d v represent the vector dimensions of q, k, and v respectively, n h represents the token length. is the Softmax function. After Q ″ is processed by the feed-forward network (FFN), it is input into the subsequent decoding layer to generate the query at the next moment Finally, the spatial and temporal information is combined through the following matrix multiplication method:

[0041]

[0042] Among them, F out represents the multi-modal features that incorporate historical temporal information and highlights the possible locations of the targets.

[0043] Step 4: Input the multi-modal features the output F of the historical information decoder out These features will be input into the tracking head for predicting the target bounding box.

[0044] Advantages of the present invention:

[0045] (1) Historical temporal information perception

[0046] Existing RGB-event multi-modal tracking methods usually neglect the mining of historical temporal information, which is crucial for multi-modal tracking tasks. This information can capture changes in the target state and related motion trends. Different from several past methods that design template branches or correlation filters to update template image information, this patent first uses a state space model with linear complexity to capture the temporal information between multiple search frames. By regressing historical information and associating it with the current frame information, the position of the next frame of the prediction model is assisted in an autoregressive manner, solving the problems of target loss and complex background in long-distance tracking.

[0047] (2) Efficient fusion of RGB modality and event modality

[0048] Due to the asynchrony of event data, different from past RGB-event information fusion methods based on convolutional networks or attention mechanisms, this patent first explores a method for fusing RGB data and event data based on the state space model scanning mechanism. The FM module proposed in this patent effectively extracts features from the RGB modality and event modality through a gating mechanism and fuses them according to the significance of the gating signal of the state space model. By dynamically fusing the information of the two modalities in different scenarios, this patent can effectively complement the advantages of the two sensors and thus solve the problem of target tracking under the condition of single-modal degradation. Description of the drawings

[0049] Figure 1 is the overall target tracking flowchart of the present invention.

[0050] Figure 2 is the structural diagram of the modality fusion module (FM) based on the state space model of the present invention, which includes (a) the overall model architecture, (b) the target perception scanning module (TAS), and (c) the cross-modal scanning module (CMS).

[0051] Figure 3 is the structural diagram of the historical information decoder of the present invention.

[0052] Figure 4 These are the comparative experiment results of the present invention on the VisEvent dataset.

[0053] Figure 5 These are the comparative experiment results of the present invention on the FELT dataset.

[0054] Figure 6 These are the comparative experiment results of the present invention on the FE108 dataset. Detailed Embodiments

[0055] The present invention will be further described in detail below in conjunction with the detailed embodiments, but the present invention is not limited to the detailed embodiments.

[0056] A multi-modal single-object visual tracking method based on a state space model, including the training and testing of the network model.

[0057] (1) Multi-modal Encoder

[0058] The multi-modal encoder is used to extract the information of RGB and event images. In this method, a vertical architecture self-attention network HiViT is adopted as the multi-modal encoder, and features are extracted layer by layer through token encoding, token merging, and self-attention mechanism. Between the intermediate attention layers, a modal fusion module based on the state space model is inserted for cross-modal fusion. Finally, the extracted final layer features are sent to the decoder and prediction head for the next prediction.

[0059] (2) Modal Fusion Module (FM) Based on State Space Model

[0060] This patent proposes a modal fusion module (FM) based on the state space model for fusing multi-modal semantic information. The design of the FM module is based on the following two principles: (i) The interaction between the template image features and the search image features is crucial, which can help the tracker better focus on the search area related to the target; (ii) Directly fusing the RGB modality and the event modality may ignore the problem that in some cases, a certain modality cannot provide valuable semantic information. In the FM module, a target-aware scanning module (TAS) and a cross-modal scanning module (CMS) are respectively introduced to solve the above problems.

[0061] Target-aware Scanning Module (TAS)

[0062] The target-aware scanning module enhances the interaction between the template region and the search region within the modality through the global modeling ability of the state space model. This module adopts a bidirectional scanning mechanism to calculate the correlation between the template features and the search features. After obtaining the features of a certain layer in the multi-modal encoder after which Representing RGB image features and event image features respectively, the formula of the TAS module is as follows:

[0063]

[0064] Among them, represents the linear projection layer; represents the layer normalization operation; σ is the SiLU activation function; ψ is the 1D convolution operation; ⊙ represents the element-wise multiplication; is 's projection; is the gating signal; represents the forward scanning state space model, represents the backward scanning state space model; Y f and Y b are the outputs through bidirectional scanning (i.e., and ). is the output after the target-aware scanning module jointly models the template region features and search region features in the n-th layer encoder, represents the RGB image features in the n-th layer encoder, represents the event image features in the n-th layer encoder, indicating forward.

[0065] Cross-modal Scanning Module (CMS)

[0066] In the CMS module of this model, cross-modal interaction is carried out, and a simple and effective cross-fusion scheme based on the gating mechanism is proposed to fuse semantic information and cross-modal complementarity. Different from the target-aware scanning module, the cross-modal scanning module only takes the search features as input to prevent interference from template feature information. Taking the RGB image modality as an example, given the RGB image search region features in the n-th layer, the cross-modal scanning module is defined as follows:

[0067]

[0068] Among them, the modality-specific gating signals and are used to guide the interaction between modalities, thereby enhancing the main modality and suppressing the degraded modality. Through the target-aware scanning and cross-modal scanning modules proposed in this patent, the multi-modal encoder can effectively promote intra-modal and cross-modal interactions. By adding the RGB and event features of the last layer encoder, the final fused feature is obtained. This feature will be used to predict the target position after historical decoder modeling.

[0069] (3) Construction of the historical information decoder

[0070] The historical information decoder perceives the motion trend and appearance changes of the target by decoding historical information, thereby assisting the model in making accurate predictions. Each layer of the historical information decoder proposed by this model consists of two key components: the historical state perception module and the feature sequence attention module.

[0071] Historical State Perception Module (HSA)

[0072] The historical state perception module is used to model the historical information in the video sequence. A sequence modeling module based on the state space model architecture is proposed, and the target states of N frames in the sequence are modeled as N one-dimensional learnable vectors Q(q 1 ,…,q N )∈R N×D , where D is the encoding dimension. The sequence is modeled through the state space model and Q ′ =(q ′ ,…,q ′ N )∈R N×D is output, and the formula is as follows:

[0073]

[0074] is the linear projection layer, is the normalization layer, and Q ′ is output and will subsequently be used as the input to the feature sequence attention module. Thanks to the efficient temporal modeling mechanism of the state space model, the hidden state generated by the query vector Q at each moment will be passed to the next moment and participate in the prediction of subsequent moments. In this way, the long-term effectiveness of historical clue modeling is ensured. Compared with the autoregressive model based on the attention mechanism, the inference and training speeds are significantly improved.

[0075] Feature Sequence Attention Module (FSA)

[0076] The feature sequence attention module aims to combine multi-modal spatial features and the historical temporal feature Q ′ to improve the confidence of the prediction head. The FSA module in this project adopts the multi-head attention mechanism, where the historical information sequence serves as the query, and the multi-modal features serve as the key and value. The formula is as follows:

[0077]

[0078] is Q ″ after being processed by the feed-forward network (FFN) and input into the subsequent decoding layer to generate the final temporal query Finally, we combine the spatial and temporal information through the following matrix multiplication method to avoid introducing additional parameters:

[0079]

[0080] Among them, F out represents the multi-modal features that incorporate historical temporal information and highlights the possible locations of the targets. Subsequently, F out these features are input into the tracking head for predicting the target bounding box.

[0081] Network training

[0082] For the MamTrack model proposed in the present invention, the parameters of the HiViT model pre-trained on the ImageNet dataset are used to initialize it. The batch size of the model is set to 32. During the training process, Adam is used as the optimizer to update the model parameters. The number of iterations is set to 50 rounds, the decay factor of the learning rate is set to 0.2 and it decays every 15 iterations. The learning rates of the classification head, the regression head, and the backbone network are set to 0.001, 0.001, and 0.0001 respectively. The verification metrics of this model adopt the success rate (SR) and precision rate (PR) commonly used in object tracking. Among them, the success rate focuses on the proportion of the regression box predicted by the model and the true regression box of the target that is greater than a given threshold, and the PR focuses on the proportion of the predicted target center position of the model and the true center position of the target that is less than a given threshold. Comparative experiments are carried out on the VisEvent dataset, the FE108 dataset, and the FELT dataset respectively. These three datasets are all multi-modal single-object tracking datasets based on RGB-event images, taken using the DAVIS346 temporal camera images, with an image resolution of 346×260, a dynamic range of 120dB, and a camera delay of 20 μs. The VisEvent dataset is the largest-scale RGB-event multi-modal object tracking dataset at present, containing 820 video sequences, of which 500 sequences are used for training and 320 sequences are used for testing. The VisEvent dataset is used to verify the performance of this model in outdoor scenarios; the FE108 dataset contains 108 video sequences, of which 76 sequences are used for training and 32 sequences are used for testing. The FE108 dataset is used to verify the tracking performance of this model in indoor scenarios. FELT is an RGB-event object tracking dataset for long-term tracking, containing 742 video sequences. Among them, 1,594,474 pairs of RGB-event images contain 45 categories of indoor and outdoor targets. 520 sequences are used for training and 222 video sequences are used for testing. The FELT dataset is used to verify the performance of this model in long-term tracking. The experimental results of this model are as follows:

[0083] By the appendix Figure 4As shown, the model achieved a success rate of 61.6% and an accuracy rate of 79.2% on the VisEvent dataset, exceeding the current SOTA models AQATrack and SDSTrack, proving the robustness of the model in the present invention; by the attached Figure 5 As shown, the model achieved a success rate of 48.9% and an accuracy rate of 60.8% on the FELT dataset, exceeding the current SOTA models AiATrack and BAT, verifying the effectiveness of the present invention in long-term tracking; by the attached Figure 6 As shown, the model achieved a success rate of 66.4% and an accuracy rate of 94.2% on the FE108 dataset, reaching a new SOTA effect, further verifying the generalization ability of the model in different scenarios.

Claims

1. A multimodal target tracking method based on a state space model, characterized in that: Here are the steps: Step 1: RGB-event image data processing Firstly, the image data of various challenging scenes are collected by calibrating the event camera and RGB camera. The event stream generated by the event camera is superimposed into event image frames according to a certain time step, and the images are aligned and filtered according to the RGB camera frequency to obtain the event image. Then, the RGB image and event image are pre-processed and sent to the model. Step 2: Build a multimodal encoder The multimodal encoder is used to extract information from RGB images and event images. The vertically structured attention network HiViT is used as the multimodal encoder to extract features layer by layer through token encoding, token merging, and self-attention mechanisms. The modal fusion module FM based on the state space model is inserted between the middle attention layers for cross-modal fusion. Finally, the extracted final layer features are sent to the history information decoder and prediction head for the next prediction. Step 2.1: Construction of modal fusion module based on state-space model A modal fusion module FM based on the state space model is constructed to fuse multi-modal semantic information. The target perception scanning module and the cross-modal scanning module are introduced into the FM module respectively. (1) Target perception scanning module The target perception scanning module uses the global modeling capability of the state space model to enhance the interaction between the template area and the search area within the modality. The target perception scanning module adopts a bidirectional state space model scanning mechanism to calculate the correlation between the template features and the search features. (2) Cross-modal scanning module A cross-modal scanning module is established based on the gating mechanism of the state-space model to cross-fuse the two modalities; Through the target perception scanning module and the cross-modal scanning module, the multimodal encoder can promote intra-modal and cross-modal interactions; the RGB and event features of the last layer of encoder are added to obtain the final fusion feature Final fusion features It will be used to predict the target position after passing through the historical information decoder; Step 3: Historical information decoder construction The historical information decoder perceives the target’s motion trend and appearance changes by decoding historical information, thereby assisting the model in making accurate predictions. Each layer of the historical information decoder includes a historical state perception module and a feature sequence attention module. The historical state perception module is used to model the historical information in the video sequence, using a sequence modeling module based on a state space model; the output of the historical state perception module is then used as the input of the feature sequence attention module and the multimodal feature in step 2 Fusion is performed to accurately locate the target in the current search area; Step 4: Multimodal features The output of the history information decoder F out Input to the tracking head for object bounding box prediction.

2. The multimodal target tracking method based on the state space model according to claim 1, characterized in that: Definition of target perception scanning module: in, Represents a linear projection layer; represents the layer normalization operation; σ is the SiLU activation function; ψ is the 1D convolution operation; ⊙ represents element-wise multiplication; represents the features of a layer in the multimodal encoder, Represent RGB image features and event image features respectively, yes projection; is the gating signal; represents the forward scanning state space model, represents the backward scanning state space model, Y f and Y b is through the SSM forward and backward directions (i.e. and ) output; It is the output of the target perception scanning module after jointly modeling the template area features and the search area features in the n-th layer encoder. represents the template region feature in the n-th layer encoder, Represents the search region features in the n-th layer encoder.

3. The multimodal target tracking method based on the state space model according to claim 1, characterized in that: Definition of the cross-modal scanning module: Take the RGB modality image as an example, where It represents the features of the RGB image after being encoded by the target perception scanning module at the nth layer. Represents the RGB image search area feature in the nth layer, Y ′ express The intermediate features after scanning by the state space model, Represents the fusion features extracted from the RGB image in the nth layer after passing through the cross-modal scanning module, and the modality-specific gating signal and It is used to guide the interaction between modes, thereby enhancing the main mode while suppressing the degenerate mode.

4. The multimodal target tracking method based on a state space model according to claim 1, characterized in that: The historical information decoder is specifically: (1) Historical state perception module The target state of N frames in the sequence is modeled as N one-dimensional learnable vectors Q(q1,…,q N )∈R N×D , where D is the encoding dimension; the sequence is modeled through the state space model and outputs Q ′ =(q ′ ,…,q ′ N )∈R N×D , the formula is as follows: Among them, X Q represents the input features after query Q projection, Z Q represents the gated features after query Q projection, represents the state space model; is the linear projection layer, is the normalization layer, (2) Definition of feature sequence attention module: Among them, Q ″ Represents the historical information features scanned by the historical state perception module. represents the vector concatenation operation, H i represents the intermediate features after self-attention, represents the attention mechanism, q represents the attention query, k represents the attention key, and v represents the attention value as well as are all learnable parameters, d m Represents the input dimension d q , d k and d v Represents the vector dimensions of q, k and v respectively, n h is the token length, is the Softmax function, Q ″ After being processed by the feed-forward network (FFN), it is input into the subsequent decoding layer to generate the query at the next moment Finally, the spatial and temporal information are combined via the following matrix multiplication: Among them, F out It represents a multimodal feature that integrates historical temporal information and highlights the possible location of the target.

Citation Information

Patent Citations

  • Visible light-thermal infrared target tracking method based on robust spatio-temporal context modeling of state space model

    CN119006529A

  • Lightweight multi-modal target tracking method

    CN119048874A

  • Sequence multi-modal scene recognition method based on state space model

    CN119445343A

  • Multi-scale deep reinforcement machine learning for N-dimensional segmentation in medical imaging

    US10032281B1