Single target tracking method and system based on selective state sensing space model

The time modeling process is simplified through selective state-aware spatial model (SASM), and the existing algorithms have high computational complexity and high computational overhead in long-term dynamic scenarios, achieving efficient single-objective tracking and dynamic scenario adaptability.

CN120339328APending Publication Date: 2025-07-18DEEP SPACE EXPLORATION LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510407292.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing single-objective tracking algorithm based on convolutional neural networks and attention networks has high computational complexity when processing long-term dynamic scene information, requiring additional customized modules or high computing overhead, which is difficult to meet real-time requirements.

Method used

The selective state-aware spatial model (SASM) is adopted to simplify the time modeling process through the interaction between the time scale parameters of the shared state and the hidden state, introducing time information for long time series, avoiding additional module design and high computing overhead.

Benefits of technology

It realizes efficient single-objective tracking, improves adaptability to dynamic scenes and global consistency between multiple frames, and reduces computing complexity and computing overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339328A_ABST
    Figure CN120339328A_ABST
Patent Text Reader

Abstract

The invention discloses a single target tracking method and system based on a selective state sensing spatial model, and relates to the technical field of computer vision, and the method comprises the following steps: obtaining image related parameters, extracting the input features of the image related parameters, and inputting the input features into a pre-established selective state sensing spatial model SASM; and based on a preset time sequence clue algorithm, performing training tracking on the selective state sensing space model SASM after the features are input, and outputting to obtain a single-target tracking result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, specifically a single-object tracking method and system based on a selective state perception spatial model. Background Art

[0002] Visual single-object tracking is a basic research topic in the field of computer vision. Its goal is to automatically locate the target in subsequent frames by given the target box of the first frame. Single-object tracking has wide applications in autonomous driving, intelligent monitoring, and human-computer interaction.

[0003] Tracking algorithms based on temporal scene modeling have been widely studied, especially on the basis of convolutional neural networks and attention network architectures. However, these methods have certain limitations in dealing with long-term dynamic scene information, especially in terms of computational complexity and model structure. Tracking algorithms based on convolutional neural networks usually introduce dynamic temporal scene information through target filters (such as correlation filters or convolutional filters), but this requires designing efficient filter optimizers to accelerate the online training process. In addition, some methods construct temporal interactions through manually designed cross-frame association scores or association networks, which makes the network structure complex and the computational cost high. Tracking algorithms based on attention networks introduce temporal cues by splicing the target template with the tracked frames, but due to the quadratic complexity of their attention mechanism, it is difficult to process long-term scene cues. Although some methods try to introduce customized modules to propagate temporal cues (such as temporal tokens, visual cues, etc.), these customized modules often increase the computational overhead and make the training process more complex. It can be seen that convolutional neural networks and attention networks are not naturally suitable for modeling video-level scene changes when dealing with long temporal dependencies, usually requiring additional customized modules or high computational overhead, so the temporal scene modeling ability is limited under real-time requirements. Summary of the Invention

[0004] To solve the deficiencies mentioned in the above background art, the purpose of the present invention is to provide a single-object tracking method and system based on a selective state perception spatial model.

[0005] In a first aspect, the purpose of the present invention can be achieved by the following technical solutions: A single-object tracking method based on a selective state perception spatial model, the method comprising the following steps:

[0006] Obtain image-related parameters, extract input features of the image-related parameters, and input the input features into a pre-established selective state perception spatial model SASM;

[0007] Based on a pre-set temporal cue algorithm, train and track the selective state perception spatial model SASM after the input features, and output a single-object tracking result.

[0008] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the pre-established selective state-aware space model SASM uses the time-scale parameter Δ of the shared state, and respectively extracts the time-scale parameters Δ of the state and the channel through linear projection state and Δ channel , and then adds them through broadcasting, and the formula is expressed as:

[0009]

[0010] In the formula, σ(·) represents the softplus operation, and are both linear mapping layers to generate the time-scale parameter of the state and the time-scale parameter of the channel.

[0011] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the pre-established selective state-aware space model SASM performs information exchange through a tracking algorithm, so as to be optimized from a global perspective, and the formula is expressed as:

[0012] h = Linear up (Linear down (h′))

[0013] where h′ ∈ R B×D×N represents the last hidden state of each image, B, D, N are the dimensions of batch, channel, and state respectively, Linear down (·) reduces the dimension of the state to N / 4, and Linear up (·) increases the dimension of the state to N;

[0014] makes SASM be expressed as: (F out , h last ) = SASM(F, h init )

[0015] where, F, F out ∈ R B×L×D are the input and output features, and h init , h last ∈ R B×D×N are the initial and final hidden states.

[0016] In combination with the first aspect, in some implementations of the first aspect, the method further includes: given the input feature F, the SASM block is expressed as:

[0017] F x = Linear(Norm(F)), Fz = Linear(Norm(F)),

[0018] F' f = σ(Conv1D f (F x ), F'), b = σ(Conv1D b (TCFlip(F x )),

[0019]

[0020] F out = Linear(F f ⊙ σ(F z ) + F b ⊙ σ(F z ))),

[0021] where Norm(·) is RMS normalization, σ(·) is the SiLU function, is the learnable initial hidden state for forward and backward scans, is the last hidden state for forward and backward scans, TCFlip(·) is time-causal flipping, and by simplification, the SASM block is represented as:

[0022]

[0023] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the training process of training and tracking the selective state-aware spatial model SASM based on the preset timing clue algorithm is performed by given a target template and a search area which are split and mapped to patch embeddings through linear projection and where N t , N s are the number of patches of the template and the search area, L t = H t W t / p 2 , L s = H s W s / p 2 , p is the resolution of each patch. After adding the positional embedding for spatial-aware feature learning, and then adding the temporal embedding for temporal-aware feature learning, and adding the target embedding for target-aware feature learning,

[0024]

[0025] Among them, E t , E a , E b are learnable parameters.

[0026] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the target embedding E b is generated by assigning the positions of the target region to one parameter vector and the positions outside the target region to another parameter vector. After the features of the target template and the search region are concatenated, they are input into N SASM SASM blocks to establish global interaction between the template and the search region,

[0027]

[0028] The enhanced features of the search region are used as the input of the bounding box head, and the bounding box head and the training objective are the same as those of OSTrack.

[0029] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the temporal tracking process based on a preset temporal clue algorithm:

[0030] The tracking algorithm maintains the temporal causal relationship between frames in the SASM block. The search region freely establishes interaction with the scanned template through the hidden state. Given the patch embedding of the initial target template Obtain the last hidden state of the template in each SASM block Expressed by the formula:

[0031]

[0032] The search region embedding Establishes interaction with the template through the scanned hidden state Expressed by the formula:

[0033]

[0034] The features of the search region are input into the bounding box head to locate the target, and the target template of the new tracking frame is updated to introduce temporal clues:

[0035]

[0036] In a second aspect, to achieve the above object, the present invention discloses a single-object tracking system based on a selective state-aware spatial model, including:

[0037] A feature input module, configured to obtain image-related parameters, extract input features of the image-related parameters, and input the input features into a pre-established Selective State Awareness Space Model (SASM).

[0038] A training and tracking module, configured to perform training and tracking on the Selective State Awareness Space Model (SASM) after the input features based on a pre-set time-series clue algorithm, and output a single-object tracking result.

[0039] In another aspect of the present invention, to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, the single-object tracking method based on the Selective State Awareness Space Model as described above is adopted.

[0040] In yet another aspect of the present invention, to achieve the above object, a computer-readable storage medium is disclosed. The computer-readable storage medium stores a computer program. When the computer program is loaded and executed by a processor, the single-object tracking method based on the Selective State Awareness Space Model as described above is adopted.

[0041] Advantages of the present invention:

[0042] The tracking algorithm of the present invention establishes an interaction between the search area and the time clue through the propagation and update of hidden states, avoiding the need for additional module design or high computational overhead, thereby realizing a more efficient tracking process. This design enables the algorithm to effectively introduce time information of a long time series without relying on complex calculations and additional modules. One of the core innovations of the tracking algorithm is the proposed Selective State Awareness Space Model (SASM), which captures more diverse time clues through the time-scale parameter of state awareness. This design enables the model to flexibly adapt to and extract time information in different states, further enhancing the adaptability to dynamic scenes. In addition, the interaction between hidden states is introduced in the Selective State Awareness Space Model, allowing for global state optimization, which effectively improves the global consistency and time perception ability of the model among multiple frames. Description of the Drawings

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts;

[0044] Figure 1It is a schematic diagram of the method flow of the present invention;

[0045] Figure 2 It is a schematic diagram of the design details of the selective state perception space model of the present invention;

[0046] Figure 3 It is the overall tracking framework diagram of the present invention;

[0047] Figure 4 It is a schematic diagram of the system structure of the present invention;

[0048] Figure 5 It is an example diagram for verifying the effect of the present invention. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] Embodiment 1:

[0051] As Figure 1 shown, a single-object tracking method based on a selective state perception space model, the method includes the following steps:

[0052] S101: Obtain image-related parameters, extract the input features of the image-related parameters, and input the input features into a pre-established selective state perception space model SASM;

[0053] Specifically, as Figure 2 shown,

[0054] The selective state space model uses a shared-state time-scale parameter Δ as a gating mechanism to control which features participate in state update. This shared-state design makes the features in all states either uniformly ignored or considered, thus restricting the ability of the hidden state to capture diverse cues. To overcome the above limitations, the present invention designs a selective state perception space model (SASM), as Figure 2 shown. The time-scale parameters Δ state and Δ channel of the state and the channel are respectively extracted through linear projection, and then they are added together through broadcasting. The formula is expressed as:

[0055] Δ state = Linear s (x), Δ channel = Linear c (x), Δ = σ(Δstate +Δ channel ),

[0056] where σ(·) represents the softplus operation, and Δ channel controls whether the features of each channel are ignored or considered as in the previous selective state space model. Δ state Controls whether the features in each state are ignored or considered. With this design, different hidden states can freely capture diverse scene cues, such as targets, backgrounds, or distractors. Linear s (·) and Linear c (·) are both linear mapping layers to generate the time scale parameters of the states and the time scale parameters of the channels.

[0057] Therefore, the scanned hidden states can retain more diverse temporal cues. In addition, the selective state space model learns the hidden states independently, making it impossible to optimize them from a global perspective to capture more comprehensive cues, as shown by the black line in Figure 2 . To overcome this limitation, the tracking algorithm introduces interactions between hidden states after scanning each image, as shown by the red line in Figure 2 , so as to achieve information exchange and enable them to be optimized from a global perspective. The formula is expressed as: h = Linear up (Linear down (h′)), where h′ ∈ R B×D×N represents the last hidden state of each image. B, D, and N are the dimensions of batch, channel, and state respectively. Linear down (·) reduces the dimension of the state to N / 4. Linear up (·) increases the dimension of the state to N. For simplicity, SASM can be generally expressed as: (F out , h last ) = SASM(F, h init ). Here, F, F out ∈ R B×L×D are the input and output features, and h init , h last ∈ R B×D×N are the initial and final hidden states.

[0058] S102: Train and track the selective state-aware space model SASM after the input features based on a preset temporal cue algorithm, and output a single-object tracking result.

[0059] Designed the selective state-aware space model layer and proposed a concise time modeling paradigm for visual tracking, as shown in Figure 3 .

[0060] a) Selective State Awareness Space Model Layer: In the field of natural language processing, the original selective state space model block uses a one-way scan to process causal 1D sequences, enabling each element to interact with any previously scanned sample through a compressed hidden state. In computer vision, many methods introduce multi-directional scans through flipping or transposing to process 2D images for spatial awareness feature learning. Different from previous methods, the tracking algorithm not only seeks to maintain the causal relationship between frames for efficient temporal modeling but also hopes to learn spatial awareness features in the SASM block for image understanding.

[0061] The present invention designs a temporal causal flipping method for temporal causal scanning, as Figure 3 shown (different colors represent different frames, and different shades represent different image regions). Only perform in-frame flipping on the sequence, which can maintain the temporal causality between frames. As Figure 3 shown, given the input feature F, the SASM block can be expressed as:

[0062] F x = Linear(Norm(F)), F z = Linear(Norm(F)),

[0063] F′ f = σ(Conv1D f (F x ))), F′ b = σ(Conv1D b (TCFlip(F x ))),

[0064]

[0065] F out = Linear(F f ⊙ σ(F z ) + F b ⊙ σ(F z )),

[0066] where Norm(·) is RMS normalization, σ(·) is the SiLU function, is the learnable initial hidden state for forward and backward scans, is the final hidden state for forward and backward scans, and TCFlip(·) is temporal causal flipping. For simplicity, the SASM block can be expressed as:

[0067]

[0068] Training with Temporal Clues: Conventional tracking algorithms based on convolutional neural networks or attention networks usually require a large amount of computational overhead or complex modules and processes to introduce temporal clues in a long time range for training. The present invention hopes to introduce more target templates with linear complexity to integrate temporal clues in a long time range for training without additional design. As Figure 3 shown, the target template and the search region are multi-fold cropped to the target size. Given the target template and the search region they are split and mapped to patch embeddings through linear projection and Here, N t , N s is the number of patches of the template and the search region, L t = H t W t / p 2 , L s = H s W s / p 2 , where p is the resolution of each patch. Then, position embeddings are added for spatial-aware feature learning, temporal embeddings are added for temporal-aware feature learning, and target embeddings are added for target-aware feature learning,

[0069]

[0070] where E t , E a , E b are learnable parameters. The target embedding E b is generated by assigning the positions of the target region to one parameter vector and the positions outside the target region to another parameter vector. The features of the target template and the search region are concatenated and input into N SASM SASM blocks to establish global interaction between the template and the search region,

[0071]

[0072] The enhanced features of the search region are used as the input of the bounding box head for regressing the bounding box. The bounding box head and the training objective are the same as OSTrack.

[0073] Tracking with Temporal Clues: The tracking algorithm maintains the temporal causal relationship between frames in the SASM block. Therefore, the search region can freely establish interaction with the scanned template through the hidden state without repeatedly scanning the template during the tracking process, asFigure 3 As shown, the patch embedding of the initial target template can obtain the last hidden state of the template in each SASM block The formula is expressed as:

[0074]

[0075] Then, the search area embedding can interact with the template through the scanned hidden state The formula is expressed as:

[0076]

[0077] Next, the features of the search area are input into the bounding box head to locate the target. In addition, the target template of the new tracking frame can be used to update to introduce temporal cues:

[0078]

[0079] Thanks to the effective design, the tracking algorithm can establish global interaction between the target template and the search area through the compressed hidden state. At the same time, the tracking algorithm can easily introduce temporal cues through hidden state updates without additional module design or large computational overhead, which provides a concise network structure and tracking process for temporal modeling.

[0080] The present invention can be applied to systems such as autonomous driving and intelligent monitoring to continuously and accurately locate targets in visual signals. It can be applied to embedded devices and mobile devices to provide real-time tracking results, or deployed on large computing servers to provide target location and tracking services for a large number of users.

[0081] Embodiment 2: As Figure 4 shown, to achieve the above object, the present invention discloses a single target tracking system based on a selective state-aware spatial model, including:

[0082] A feature input module, configured to obtain image-related parameters, extract input features of the image-related parameters, and input the input features into a pre-established selective state-aware spatial model SASM;

[0083] A training and tracking module, configured to perform training and tracking on the selective state-aware spatial model SASM after the input features based on a pre-set temporal cue algorithm, and output a single target tracking result;

[0084] As Figure 5 shown, the algorithm proposed by the present invention outperforms existing tracking algorithms in a large number of complex scenarios, demonstrating the effectiveness and superiority of the algorithm of the present invention.

[0085] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.

[0086] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program executes the above method when run by a processor. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.

[0087] In the description of this specification, the descriptions referring to the terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0088] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art of this industry should understand that the present disclosure is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.

Claims

1. A single-object tracking method based on a selective state perception space model, characterized in that The method includes the following steps: Obtain image-related parameters, extract the input features of the image-related parameters, and input the input features into the pre-established Selective State Awareness Space Model (SASM); Based on the pre-set temporal cue algorithm, train and track the Selective State Awareness Space Model (SASM) after the input features, and output a single-object tracking result.

2. The single-object tracking method based on a selective state perception space model according to claim 1, wherein The pre-established selective state awareness space model (SASM) uses a time scale parameter Δ for state sharing, and separately extracts the time scale parameters Δ of the state and the channel through linear projection state and Δ channel , and then adds them through broadcasting. The formula is expressed as: where, σ(·) represents the softplus operation, and are both linear mapping layers to generate the time scale parameter of the state and the time scale parameter of the channel.

3. The single-object tracking method based on the selective state perception space model according to claim 2, wherein The pre-established Selective State Awareness Space Model (SASM) exchanges information through a tracking algorithm, so as to be optimized from a global perspective. The formula is expressed as: h = Linear up (Linear down (h')) where h′ ∈ R B×D×N represents the last hidden state of each image, and B, D, N are the dimensions of batch, channel, and state respectively, reduce the dimension of the state to N / 4, increase the dimension of the state to N; such that SASM is represented as: (F out ,h last ) = SASM(F,h init ) where F,F out ∈R B×L×D are the input and output features, h init ,h last ∈R B×D×N are the initial and final hidden states.

4. The single-object tracking method based on the selective state perception space model according to claim 1, characterized in that When the pre-established Selective State Awareness Space Model (SASM) is given the input feature F, the SASM block is expressed as: F x = Linear(Norm(F)), F z = Linear(Norm(F)), F′ f = σ(Conv1D f (F x )), F′ b = σ(Conv1D b (TCFlip(F x ))), F out = Linear(F f ⊙σ(F z ) + F b ⊙σ(F z )) where Norm(·) is RMS normalization, σ(·) is the SiLU function, is the learnable initial hidden state for forward and backward scans, is the final hidden state for forward and backward scans, TCFlip(·) is time-causal flipping, and by simplification, the SASM block is represented as:

5. The single-object tracking method based on the selective state perception space model according to claim 1, characterized in that The training process of training and tracking the selective state perception space model SASM after input features based on a preset timing clue algorithm is through a given target template and a search area are split by linear projection and mapped into patch embeddings and where, L t , L s are the number of patches of the template and the search area, L t = H t W t / p 2 , L s = H s W s / p 2 , p is the resolution of each patch, N t is the number of templates, and then position embeddings are added for spatial perception feature learning, and then temporal embeddings are added for temporal perception feature learning, and target embeddings are added for target perception feature learning Among them, E t , E a , E b are learnable parameters.

6. The single-object tracking method based on the selective state perception space model according to claim 5, characterized in that, The target embedding E b is generated by assigning the positions in the target region to one parameter vector and the positions outside the target region to another parameter vector. After the features of the target template and the search region are concatenated, they are input into SASM N SASM blocks to establish the global interaction between the template and the search region. Enhanced features of the search region As the input of the bounding box head, the bounding box head and the training objective are the same as those of OSTrack.

7. The single-object tracking method based on the selective state perception space model according to claim 6, wherein The temporal tracking process based on the pre-set temporal cue algorithm: The tracking algorithm maintains the temporal causal relationship between frames in the SASM block. The search region freely establishes interactions with the scanned templates through the hidden state, given the patch embedding of the initial target template. Obtain the last hidden state of the template in each SASM block. The formula is expressed as: Search area embedding With the hidden state obtained by scanning Establish an interaction with the template, which can be expressed by the formula: The features of the search area are input into the border header to locate the target, and the target template of the new tracking frame is updated to introduce temporal cues:

8. A single-object tracking system based on a selective state-aware spatial model, characterized in that, Includes: A feature input module, which is used to obtain image-related parameters, extract the input features of the image-related parameters, and input the input features into the pre-established Selective State Awareness Space Model (SASM); A training and tracking module, which is used to train and track the Selective State Awareness Space Model (SASM) after the input features based on the pre-set temporal cue algorithm, and output a single-object tracking result.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, the single-object tracking method based on the Selective State Awareness Space Model according to any one of claims 1 to 7 is adopted.

10. A computer-readable storage medium storing a computer program therein, characterized in that, When the computer program is loaded and executed by the processor, the single-object tracking method based on the Selective State Awareness Space Model according to any one of claims 1 to 7 is adopted.