Efficient visual tracking method based on rgb-event adaptive frame deletion and insertion
Patent Information
- Application Number
- CN202410704578.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-03
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-06-03
AI Technical Summary
[0005](1)现有的跟踪算法依赖于帧级的训练,导致训练和测试之间在数据分布和任务目标方面存在不一致性
[0052]本发明构建的网络模型是序列级的训练和测试,模型考虑到当前帧中目标状态的估计受到历史记录的影响,并且会影响后一帧中的跟踪结果。保证了训练和测试之间在数据分布和任务目标方面的一致性。
Smart Images

Figure CN118691643B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of event camera technology, and more specifically to an efficient visual tracking method based on adaptive frame insertion and deletion of RGB-events. Background Technology
[0002] Visual object tracking is a crucial task in computer vision. Traditional visual object tracking tasks are based on RGB data, but RGB data is susceptible to challenging factors such as low illumination, fast motion, background interference, and severe occlusion. Compared to traditional RGB cameras, event cameras offer advantages such as low power consumption, low latency, and high dynamic range. Considering the respective strengths of RGB and event cameras, using both simultaneously and leveraging the complementary information from these two modalities for tracking has significant practical value and potential application prospects.
[0003] To improve the accuracy and robustness of visual target tracking in various scenarios, a variety of visual tracking algorithms have emerged, including those based on classification, correlation filtering, Siamese networks, and Transformers. Although machine learning is widely used in visual target tracking tasks, recent learning-based algorithms have largely ignored the fact that visual tracking is essentially a sequence-level task. The estimated target state in the current frame is influenced by the historical target state of the previous frame, which also affects the tracking results in the next frame. Existing tracking algorithms heavily rely on frame-level training, which inevitably leads to inconsistencies between training and testing in terms of data distribution and task objectives. Furthermore, most existing multimodal tracking algorithms employ various fusion methods (early fusion, mid-term fusion, late-term fusion) to model the relationship between the template and the search region. Fusion methods inevitably increase model complexity and reduce computational speed.
[0004] In summary, existing bimodal visual tracking methods have the following problems:
[0005] (1) Existing tracking algorithms rely on frame-level training, which leads to inconsistencies between training and testing in terms of data distribution and task objectives.
[0006] (2) The fusion method increases the complexity of the model and reduces the computation speed.
[0007] Based on this, the present invention designs an efficient visual tracking method based on RGB-event adaptive frame insertion and deletion to solve the above problems. Summary of the Invention
[0008] To address the aforementioned shortcomings of existing technologies, this invention provides an efficient visual tracking method based on adaptive frame insertion and deletion using RGB-events.
[0009] To achieve the above objectives, the present invention provides the following technical solution:
[0010] An efficient visual tracking method based on adaptive frame insertion and deletion using RGB-events includes the following steps:
[0011] Step (1): Input data;
[0012] The input template image and the search region image are input. The mapping layer segments the input image into a patch sequence and uses the projection layer to convert it into patch embeddings. Then, positional encoding is added to obtain token embeddings.
[0013] Step (2), feature extraction and relationship modeling;
[0014] The obtained token embeddings are input into the Transformer encoding layer. The Transformer encoding layer extracts features and models relationships between the template and the search region. Through the information interaction between the template and the search region, the features of the target object in the search region are extracted.
[0015] Step (3): The tracking head locates the target bounding box;
[0016] The features of the search area are input into the tracking head to obtain the bounding box of the target object, and the tracking result of the current frame is output.
[0017] Step (4), Adaptive Decision Module;
[0018] The token embeddings from the mapping layer, the features from the Transformer encoding layer, and the target bounding box from the tracking head are concatenated and input into the decision module to obtain the corresponding decision. The decisions are divided into three categories:
[0019] No action is taken; continue using RGB data as input for tracking.
[0020] Delete a frame and directly use the tracking results of the previous frame;
[0021] Frame interpolation is used, selecting event stream data as model input for tracking.
[0022] When the decision network outputs the no-operation option, input RGB data into the model and then execute steps (1)-(3);
[0023] When the decision network outputs the frame deletion option, proceed to step (3);
[0024] When the decision network outputs the frame interpolation option, input event stream data into the model and then execute steps (1)-(3).
[0025] Furthermore, the data input in step (1) is RGB data or event stream data.
[0026] Furthermore, the event stream size is T×X×Y×P, where T represents time, X and Y represent spatial coordinates, and P represents the polarity of the event. Spatial locations with events are represented as 1, and locations without events are marked as 0. The event stream data is divided into segments of size t×x×y according to the exposure interval t of the RGB data, with a total of T / t segments. The event streams of each segment are stacked to form an event image aligned with the RGB image.
[0027] Furthermore, the specific steps of step (1) include: inputting the template image Z. rgb / Z event and search area image X rgb / X event Then the template image Z rgb / Z event Search area image X rgb / X event Divide into patch sequences and Next, the patch sequence is projected into patch embeddings through a linear projection layer; the learnable positional encoding is added to the patch embeddings of the template and the search region, respectively, to finally obtain the corresponding token embeddings. and
[0028] Furthermore, the specific steps of step (2) include: concatenating the token embeddings of the template and the search area along the dimension of the number of tokens. Then the resulting vector The input is fed into the Transformer encoding layer, where the multi-head self-attention network in the Transformer block is used to extract features and model the relationship between the template and the search region, as shown in Equations (1), (2), (3), and (4). Through feature extraction and feature fusion of multiple Transformers, the output features of the last Transformer encoding layer are used as the final feature output F. rgb / F event ;
[0029]
[0030]
[0031]
[0032]
[0033] in, This represents the feature output by the l-th Transformer coding layer, where L represents the total number of Transformer coding layers. This represents the i-th head in the multi-head self-attention mechanism of the Transformer backbone network; h represents the total number of heads; W represents the weight parameters; Q, K, and V represent the query, key, and value in the self-attention mechanism, respectively; d k The feature dimension is represented by m ∈ {rgb, event}, which represents the RGB modality or the event modality. The softmax function is represented by softmax.
[0034] Furthermore, the specific steps of step (3) include: transferring feature F rgb / F event Features corresponding to the search area Inputting data into a specific tracking head yields a bounding box representation of the target location:
[0035] Box = {x, y, w, h}
[0036] Where x and y represent the coordinates of the top-left corner of the bounding box, and w and h represent the width and height of the bounding box.
[0037] Furthermore, the specific steps of step (4) include:
[0038] The adaptive decision-making module consists of two lightweight decision networks Φ insert and decision network Φ skip The decision network is composed of token embeddings from the mapping layer, features from the Transformer encoding layer, and the target bounding box from the tracking head as inputs to obtain Bernoulli distribution probabilities, as shown in equations (5) and (6):
[0039]
[0040]
[0041] Where p insert and p skip This represents the probability of a Bernoulli distribution;
[0042] The binary decision d is sampled using the Gumbel-Softmax operation. insert ∈{[0,1],[1,0]} and d skip∈{[0,1],[1,0]}, as shown in formulas (7) and (8):
[0043] d insert =Gumbel-Softmax(p insert (7)
[0044] d skip =Gumbel-Softmax(p skip (8)
[0045] Suppose there exists a classification distribution where the probability of the k-th class is p. k Where k = 1, ..., K; the discrete samples d that conform to the classification distribution are obtained using the Gumbel-Softmax method. The detailed calculation process is shown in formula (9):
[0046]
[0047] Where g k =-log(-logU k ) is a standard Gumbel distribution, U k is a random variable sampled from a uniform distribution U(0,1); where It is the temperature coefficient that controls the degree of dispersion of d.
[0048] Furthermore, when d insert =[1,0],d skip When = [0, 1], the decision module makes the option of interpolation, inputs event stream data into the model, and then executes steps (1)-(3);
[0049] When d insert =[0,1],d skip When = [1, 0], the decision module makes the option to delete the frame and jumps to step (3);
[0050] When d insert =[0,1],d skip =[0,1] or d insert =[1,0],d skip When the value is [1, 0], the decision module makes a no-operation option, inputs RGB data into the model, and then executes steps (1) to (3).
[0051] Beneficial effects
[0052] The network model constructed in this invention employs sequence-level training and testing. The model considers that the estimation of the target state in the current frame is influenced by historical records and will affect the tracking results in subsequent frames. This ensures consistency between training and testing in terms of data distribution and task objectives.
[0053] This invention introduces an adaptive frame insertion / deletion strategy. The decision module determines whether frame insertion or deletion is necessary based on the tracking results of the current RGB frame. This strategy not only improves the accuracy of the tracking algorithm but also increases the model's running speed. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0055] Figure 1 This is a flowchart of the efficient visual tracking method based on RGB-event adaptive frame insertion and deletion of the present invention;
[0056] Figure 2 This is a block diagram illustrating the principle of the efficient visual tracking method based on RGB-event adaptive frame insertion and deletion of the present invention. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0058] The present invention will be further described below with reference to embodiments.
[0059] Example 1
[0060] Please refer to the instruction manual appendix. Figure 1-2 An efficient visual tracking method based on adaptive frame insertion and deletion using RGB-events includes the following steps:
[0061] Step (1): Input data;
[0062] The input template image and search region image are input. The mapping layer segments the input image into a patch sequence and uses a projection layer to convert it into patch embeddings. Then, positional encoding is added to obtain token embeddings.
[0063] The data input in step (1) is either RGB data or event stream data.
[0064] The event stream size is T×X×Y×P, where T represents time, X and Y represent spatial coordinates, and P represents the polarity of the event. Spatial locations with events are represented as 1, and locations without events are marked as 0. The event stream data is divided into segments of size t×x×y according to the exposure interval t of the RGB data, with a total of T / t segments. The event streams of each segment are stacked to form an event image aligned with the RGB image.
[0065] The specific steps of step (1) include: inputting the template image Z rgb / Z event and search area image X rgb / X event Then the template image Z rgb / Z event Search area image X rgb / X event Divide into patch sequences Z r p gb / Z e p vent and X r p gb / X e p vent Next, the patch sequence is projected into patch embeddings using a linear projection layer. Learnable positional codes are added to the patch embeddings of the template and search region, respectively, ultimately yielding the corresponding token embeddings: H r z gb / H e z vent and H r x gb / H e x vent .
[0066] Step (2), feature extraction and relationship modeling;
[0067] The obtained token embeddings are input into several Transformer encoding layers (the number of layers can be set, for example, 12 layers, each layer is the same). The Transformer encoding layers act as the backbone network to extract features and model relationships between the template and the search region. Through the information interaction between the template and the search region, the features of the target object in the search region are extracted.
[0068] The specific steps of step (2) include: concatenating the token embeddings of the template and the search area along the dimension of the number of tokens to form H. r z g x b =[H r z gb H r x gb ] / H e z v x ent =[H e z vent H e x vent ], and then the resulting vector The input is fed into the Transformer encoding layer, where a multi-head self-attention network within the Transformer block is used to extract features and model the relationship between the template and the search region, as shown in equations (1), (2), (3), and (4). Through feature extraction and fusion across multiple Transformer layers, the output features of the last Transformer encoding layer are used as the final feature output F. rgb / F event .
[0069]
[0070]
[0071]
[0072]
[0073] in, This represents the feature output by the l-th Transformer coding layer, where L represents the total number of Transformer coding layers. Let represent the i-th head in the multi-head self-attention mechanism of the Transformer backbone network, h represent the total number of heads; W represent the weight parameters; Q, K, and V represent the query, key, and value in the self-attention mechanism, respectively; dk represent the feature dimension; m∈{rgb, event} represent the RGB mode or event mode; softmax represents the activation function.
[0074] Step (3): The tracking head locates the target bounding box;
[0075] The features of the search area are input into the tracking head to obtain the bounding box of the target object, and the tracking result of the current frame is output.
[0076] The specific steps of step (3) include: transferring feature F rgb / F event Features corresponding to the search area (through feature F) rgb / F event After performing a slicing operation and discarding the template features, the remaining search region corresponds to the feature. The input is given to a specific tracking head, which generates a bounding box representation of the target location: Box = {x, y, w, h}, resulting in the final tracking result. Here, x and y represent the coordinates of the top-left corner of the bounding box, and w and h represent the width and height of the bounding box.
[0077] Step (4), Adaptive Decision Module;
[0078] The token embeddings from the mapping layer, the features from the Transformer encoding layer, and the target bounding box from the tracking head are concatenated and input into the decision module to obtain the corresponding decision. The decisions are divided into three categories:
[0079] No action is taken; continue using RGB data as input for tracking.
[0080] Delete a frame and directly use the tracking results of the previous frame;
[0081] Frame interpolation is used, selecting event stream data as model input for tracking.
[0082] When the decision network outputs the no-operation option, input RGB data into the model and then execute steps (1) to (3).
[0083] When the decision network outputs the frame deletion option, proceed to step (3).
[0084] When the decision network outputs the frame interpolation option, input event stream data into the model and then execute steps (1)-(3).
[0085] The specific steps of step (4) include:
[0086] The adaptive decision-making module consists of two lightweight decision networks Φ insert and decision network Φ skipThe decision network is composed of token embeddings from the mapping layer, features from the Transformer encoding layer, and the target bounding box from the tracking head as inputs to obtain Bernoulli distribution probabilities, as shown in equations (5) and (6):
[0087]
[0088]
[0089] Where p insert and p skip This represents the probability of a Bernoulli distribution;
[0090] To ensure that the decision-making module can make appropriate decisions while simultaneously achieving end-to-end model training, the binary decision d is sampled using the Gumbel-Softmax operation. insert ∈{[0,1],[1,0]} and d skip ∈{[0,1],[1,0]}, as shown in formulas (7) and (8):
[0091] d insert =Gumbel-Softmax(p insert (7)
[0092] d skip =Gumbel-Softmax(p skip (8)
[0093] The Gumbel-Softmax operation is essentially a reparameterization technique for categorical distributions. Assume there exists a categorical distribution with the probability p of the k-th class. k Where k = 1, ..., K. The discrete samples d conforming to the classification distribution are obtained using the Gumbel-Softmax method. The detailed calculation process is shown in formula (9):
[0094]
[0095] Where g k =-log(-logU k ) is a standard Gumbel distribution, U k Let be a random variable sampled from a uniform distribution U(0,1). It is a temperature coefficient that controls the degree of dispersion of d. When As T approaches infinity, d approaches a uniformly distributed vector; when T approaches 0, d approaches a one-hot vector (a one-hot vector is an encoding method used to represent categorical variables, where only one position in the vector is 1 and the rest are 0).
[0096] When d insert =[1,0],d skip When = [O,1], the decision module makes the option of interpolation, inputs event stream data into the model, and then executes steps (1)-(3).
[0097] When d insert =[0,1],d skip When the value is [1, 0], the decision module makes the option to delete the frame and jumps to step (3).
[0098] When d insert =[0,1],d skip =[0,1] or d insert =[1,0],d skip When the value is [1, 0], the decision module makes a no-operation option, inputs RGB data into the model, and then executes steps (1) to (3).
[0099] This invention designs an adaptive decision module, whose output can adaptively select either RGB data or event stream data as model input. Specifically, when RGB data presents challenges, the model adaptively chooses to use event data for tracking. When the target object is stationary, the bounding box is directly acquired, and the model adaptively uses the tracking result from the previous frame as the current tracking result. This invention enables the model to fully utilize the advantages of different modalities of data, making the model more flexible in responding to different situations, thereby improving the accuracy, stability, and efficiency of tracking.
[0100] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An efficient visual tracking method based on adaptive frame insertion and deletion using RGB-events, characterized in that, Includes the following steps: Step (1): Input data; The input template image and the search region image are input. The mapping layer segments the input image into a patch sequence and uses the projection layer to convert it into patch embeddings. Then, positional encoding is added to obtain token embeddings. Step (2), feature extraction and relationship modeling; The obtained token embeddings are input into the Transformer encoding layer. The Transformer encoding layer extracts features and models relationships between the template and the search region. Through the information interaction between the template and the search region, the features of the target object in the search region are extracted. Step (3): The tracking head locates the target bounding box; The features of the search area are input into the tracking head to obtain the bounding box of the target object, and the tracking result of the current frame is output. Step (4), Adaptive Decision Module; The token embeddings from the mapping layer, the features from the Transformer encoding layer, and the target bounding box from the tracking head are concatenated and input into the decision module to obtain the corresponding decision. The decisions are divided into three categories: No action is taken; continue using RGB data as input for tracking. Delete a frame and directly use the tracking results of the previous frame; Frame interpolation, selecting to use event stream data as input for tracking; When the decision network outputs the no-operation option, input RGB data and then execute steps (1)-(3); When the decision network outputs the frame deletion option, proceed to step (3); When the decision network outputs the frame interpolation option, input event stream data and then execute steps (1)-(3).
2. The efficient visual tracking method based on adaptive frame insertion and deletion according to claim 1, characterized in that, in, The data input in step (1) is RGB data or event stream data.
3. The efficient visual tracking method based on adaptive frame insertion and deletion according to claim 2, characterized in that, in, The event stream size is T×X×Y×P, where T represents time, X and Y represent spatial coordinates, and P represents the polarity of the event. Spatial locations with events are represented as 1, and locations without events are marked as 0. The event stream data is divided into segments of size t×x×y according to the exposure interval t of the RGB data, with a total of T / t segments. The event stream of each segment is stacked to form an event image aligned with the RGB image.
4. The efficient visual tracking method based on adaptive frame insertion and deletion according to claim 3, characterized in that, The specific steps of step (1) include: inputting the template image Z rgb / Z event and search area image X rgb / X event Then the template image Z rgb / Z event Search area image X rgb / X event Divide into patch sequences and Next, the patch sequence is projected into patch embeddings through a linear projection layer; the learnable positional encoding is added to the patch embeddings of the template and the search region, respectively, to finally obtain the corresponding token embeddings: and 5. The efficient visual tracking method based on RGB-event adaptive frame insertion and deletion as described in claim 4, characterized in that, The specific steps of step (2) include: concatenating the token embeddings of the template and the search area along the dimension of the number of tokens. Then the resulting vector The input is fed into the Transformer encoding layer, where the multi-head self-attention network in the Transformer block is used to extract features and model the relationship between the template and the search region, as shown in Equations (1), (2), (3), and (4). Through feature extraction and feature fusion of multiple Transformers, the output features of the last Transformer encoding layer are used as the final feature output F. rgb / F event ; in, This represents the feature output by the l-th Transformer coding layer, where L represents the total number of Transformer coding layers. This represents the i-th head in the multi-head self-attention mechanism of the Transformer backbone network; h represents the total number of heads; W represents the weight parameters; Q, K, and V represent the query, key, and value in the self-attention mechanism, respectively; d k The feature dimension is represented by m ∈ {rgb, event}, which represents the RGB modality or the event modality. The softmax function is represented by softmax.
6. The efficient visual tracking method based on adaptive frame insertion and deletion according to claim 5, characterized in that, The specific steps of step (3) include: transferring feature F rgb / F event Features corresponding to the search area Inputting data into a specific tracking head yields a bounding box representation of the target location: Box = {x, y, w, h} Where x and y represent the coordinates of the top-left corner of the bounding box, and w and h represent the width and height of the bounding box.
7. The efficient visual tracking method based on adaptive frame insertion and deletion according to claim 6, characterized in that, The specific steps of step (4) include: The adaptive decision-making module consists of two lightweight decision networks Φ insert and decision network Φ skip The decision network is composed of token embeddings from the mapping layer, features from the Transformer encoding layer, and the target bounding box from the tracking head as inputs to obtain Bernoulli distribution probabilities, as shown in equations (5) and (6): Where p insert and p skip This represents the probability of a Bernoulli distribution; The binary decision d is sampled using the Gumbel-Softmax operation. insert ∈{[0,1],[1,0]} and d skip ∈{[0,1],[1,0]}, as shown in formulas (7) and (8): d insert Gumbel-Softmax(p insert (7) d skip Gumbel-Softmax(p skip (8) Suppose there exists a classification distribution where the probability of the k-th class is p. k Where k = 1, ..., K; the discrete samples d that conform to the classification distribution are obtained using the Gumbel-Softmax method. The detailed calculation process is shown in formula (9): Where g k =-log(-log U) k ) is a standard Gumbel distribution, U k is a random variable sampled from a uniform distribution U(0,1); where It is the temperature coefficient that controls the degree of dispersion of d.
8. The efficient visual tracking method based on adaptive frame insertion and deletion according to claim 7, characterized in that, When d insert =[1,0],d skip When = [0,1], the decision module makes the option to insert the frame, inputs the event stream data, and then executes steps (1)-(3); When d insert =[0,1],d skip When = [1,0], the decision module makes the option to delete the frame and jumps to step (3); When d insert =[0,1],d skip =[0,1] or d insert =[1,0],d skip When the value is [1,0], the decision module makes a no-operation option, inputs RGB data, and then executes steps (1) to (3).
Citation Information
Patent Citations
Image processing method and device
CN116134484A
Target tracking method and device
CN117152200A