Target tracking method and device based on heterogeneous graph and adaptive modal weighting
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本公开的实施例提供了一种基于异构图和自适应模态加权的目标跟踪方法、装置及存储介质,以至少解决现有技术中存在的进行目标跟踪的准确性较低的技术问题
[0010]In this embodiment, when performing joint target tracking using RGB images and event tensors, the method determines the RGB image features of the RGB image, the event tensor features of the event tensor, and the trajectory features at each time step. Furthermore, it determines the features of each RGB patch and each event patch at each time step, thereby constructing a spatiotemporal-modal heterogeneous graph. This graph represents the correlation between two image modalities, their temporal order, and the relationships between patches within the same frame. Based on this spatiotemporal-modal heterogeneous graph, a self-attention mechanism for message passing is implemented, and target tracking is performed using the results of this message passing. This improves the accuracy of target tracking using joint RGB images and event tensors.
Smart Images

Figure CN121582290B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of target tracking technology, and in particular to a target tracking method, apparatus and storage medium based on heterogeneous graphs and adaptive modal weighting. Background Technology
[0002] Currently, target tracking can be achieved by combining event cameras and RGB cameras. Event cameras are a novel type of sensor that asynchronously outputs event streams based on changes in light intensity, recording scene brightness changes with microsecond-level temporal resolution and extremely high dynamic range. Traditional RGB cameras, on the other hand, output entire frames of images at a fixed frequency. Applications such as robot obstacle avoidance, drone following, industrial inspection, and augmented reality all require stable target tracking in various environments. When tracking targets using both cameras, two modalities of data are acquired simultaneously in the same scene, and the target position is continuously estimated. A single modality is prone to failure under extreme lighting, rapid movement, occlusion, and low-texture conditions; using two complementary modalities improves system robustness.
[0003] In existing technologies, there are two main approaches to joint target tracking. The first is a decision-level fusion method, which performs target tracking independently for both RGB images and event data, and then uses Kalman filtering, Hungarian matching, or weighted voting for decision fusion. This type of method is easy to implement and interpretable, but because it ignores the intermediate feature correlation between the two modalities, errors may occur if one side is distorted, thus reducing the accuracy of target tracking. The second approach is a simple feature fusion method, which extracts features from both RGB images and event data separately, and then fuses the features of the two modalities (e.g., by concatenation, weighted summation, etc.) to obtain fused features, which are then used for target tracking. However, this method simply fuses the features of the two modalities, making it difficult to adapt to changes in lighting conditions and resulting in lower target tracking accuracy.
[0004] There is currently no effective solution to the technical problem of low accuracy in target tracking in the existing technology mentioned above. Summary of the Invention
[0005] The embodiments of this disclosure provide a target tracking method, apparatus, and storage medium based on heterogeneous graphs and adaptive modal weighting, to at least solve the technical problem of low accuracy in target tracking in the prior art.
[0006] According to one aspect of the present disclosure, a target tracking method is provided, comprising: acquiring an RGB image and an event tensor, wherein the event tensor is obtained by preprocessing event data acquired by an event camera; determining RGB image features corresponding to the RGB image, event tensor features corresponding to the event tensor, RGB patch features, and event patch features, wherein a patch includes an RGB patch and an event patch, and a patch corresponds to an image block in the RGB image or the event tensor; determining a fusion feature between the RGB image and the event tensor based on the RGB image features and the event tensor features, and determining a trajectory feature based on the fusion feature, wherein the trajectory feature is used to characterize the motion change of a target object from the past to the present moment; constructing a spatiotemporal-modal heterogeneous graph to represent the correlation and temporal order between different image modalities based on the trajectory feature, the RGB patch features, and the event patch features, wherein the spatiotemporal-modal heterogeneous graph includes a trajectory node, an RGB patch node, and an event patch node corresponding to each moment; and performing message passing based on a self-attention mechanism of the spatiotemporal-modal heterogeneous graph, determining RGB patch node features, event patch node features, and trajectory node features, and performing target tracking based on the RGB patch node features, event patch node features, and trajectory node features.
[0007] According to another aspect of the present disclosure, a storage medium is also provided, the storage medium including a stored program, wherein, when the program is executed, a processor performs any of the methods described above.
[0008] According to another aspect of the embodiments of this disclosure, a target tracking apparatus is also provided, comprising: an acquisition module, configured to acquire an RGB image and an event tensor, wherein the event tensor is obtained by preprocessing event data acquired by an event camera; a feature determination module, configured to determine RGB image features corresponding to the RGB image, event tensor features corresponding to the event tensor, RGB patch features, and event patch features, wherein a patch includes an RGB patch and an event patch, and a patch corresponds to an image block in the RGB image or the event tensor; and a fusion module, configured to determine fusion features between the RGB image and the event tensor based on the RGB image features and the event tensor features, and to determine the trajectory based on the fusion features. The system comprises several modules: a trajectory feature module, a heterogeneous graph construction module, and a target tracking module. The former is used to characterize the motion changes of a target object from its historical state to the present moment. The latter is used to construct a spatiotemporal-modal heterogeneous graph that represents the relationships and temporal order between different image modalities based on the trajectory feature, RGB patch features, and event patch features. The former contains trajectory nodes, RGB patch nodes, and event patch nodes corresponding to each moment. The latter is used to perform message passing based on the spatiotemporal-modal heterogeneous graph using a self-attention mechanism, determine the RGB patch node features, event patch node features, and trajectory node features, and perform target tracking based on these features.
[0009] According to another aspect of the present disclosure, a target tracking apparatus is also provided, comprising: a processor; and a memory connected to the processor, configured to provide the processor with instructions for processing the following steps: acquiring an RGB image and an event tensor, wherein the event tensor is obtained by preprocessing event data acquired by an event camera; determining RGB image features corresponding to the RGB image, event tensor features corresponding to the event tensor, RGB patch features, and event patch features, wherein a patch includes an RGB patch and an event patch, and a patch corresponds to an image block in the RGB image or the event tensor; and determining the relationship between the RGB image and the event tensor based on the RGB image features and the event tensor features. Features are fused, and trajectory features are determined based on these features. Trajectory features are used to characterize the motion changes of the target object from the past to the present moment. Based on the trajectory features, RGB patch features, and event patch features, a spatiotemporal-modal heterogeneous graph is constructed to represent the correlation and temporal order between different image modalities. The spatiotemporal-modal heterogeneous graph contains trajectory nodes, RGB patch nodes, and event patch nodes corresponding to each moment. Message passing based on the spatiotemporal-modal heterogeneous graph and a self-attention mechanism is performed to determine the RGB patch node features, event patch node features, and trajectory node features. Target tracking is then performed based on these features.
[0010] In this embodiment, when performing joint target tracking using RGB images and event tensors, the method determines the RGB image features of the RGB image, the event tensor features of the event tensor, and the trajectory features at each time step. Furthermore, it determines the features of each RGB patch and each event patch at each time step, thereby constructing a spatiotemporal-modal heterogeneous graph. This graph represents the correlation between two image modalities, their temporal order, and the relationships between patches within the same frame. Based on this spatiotemporal-modal heterogeneous graph, a self-attention mechanism for message passing is implemented, and target tracking is performed using the results of this message passing. This improves the accuracy of target tracking using joint RGB images and event tensors. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this application, illustrate exemplary embodiments of this disclosure and are used to explain this disclosure, but do not constitute an undue limitation of this disclosure. In the drawings: Figure 1 This is a hardware structure block diagram of a computing device for implementing the method described in Embodiment 1 of this disclosure; Figure 2 This is a flowchart illustrating the target tracking method according to the first aspect of Embodiment 1 of this disclosure; Figure 3 This is a schematic diagram illustrating the mapping relationship between patch features and images provided in Embodiment 1 of this disclosure; Figure 4 A schematic diagram of a spatiotemporal modal heterogeneity graph provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of a target tracking device according to the first aspect of Embodiment 2 of this disclosure; and Figure 6 This is a schematic diagram of a target tracking device according to the first aspect of Embodiment 3 of this disclosure. Detailed Implementation
[0012] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this disclosure.
[0013] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0014] Example 1 According to this embodiment, a target tracking method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0015] The method embodiments provided in this example can be executed on mobile terminals, computer terminals, servers, or similar computing devices. Figure 1 A hardware block diagram of a computing device for implementing a target tracking method is shown. Figure 1 As shown, a computing device may include one or more processors (processors may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include a display, keyboard, and cursor control device connected to the input / output interface. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, a computing device may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0016] Under the aforementioned operating environment, according to the first aspect of this embodiment, a target tracking method is provided, which can be... Figure 1 The computing device implementation is shown. Figure 2 A flowchart illustrating the method is shown below. (Refer to...) Figure 2 As shown, the method includes: S202: Acquire an RGB image and an event tensor, wherein the event tensor is obtained by preprocessing event data acquired by the event camera; S204: Determine the RGB image features corresponding to the RGB image, the event tensor features corresponding to the event tensor, the RGB patch features, and the event patch features; S206: Based on the RGB image features and the event tensor features, determine the fusion features between the RGB image and the event tensor, and determine the trajectory features based on the fusion features; S208: Based on the trajectory features, the RGB patch features, and the event patch features, construct a spatiotemporal-modal heterogeneous graph to represent the correlation and temporal order between different image modalities. The spatiotemporal-modal heterogeneous graph includes a trajectory node, an RGB patch node, and an event patch node corresponding to each time step; and S210: Message passing based on a self-attention mechanism using a spatiotemporal-modal heterogeneous graph, determining RGB patch node features, event patch node features, and trajectory node features, and performing target tracking based on the RGB patch node features, event patch node features, and trajectory node features.
[0017] Specifically, the computing device can acquire RGB images and event tensors (S202). The RGB images are acquired by an RGB camera, and the event tensors are obtained by preprocessing the event data acquired by the event camera. In this embodiment, RGB images and event tensors are combined for target tracking. To achieve target tracking, the RGB images and event data are acquired continuously. The computing device can align the RGB images and event data according to time, so that each moment corresponds to one RGB image and one event tensor. That is, the preprocessing of the event data aligns the event data with each frame of the RGB image, thereby obtaining the event tensor. For example, for the RGB images and event tensors from time 1 to time T ([1, ..., t, ..., T]), at each moment, such as time t, there is a corresponding three-channel RGB image I. t With a three-channel event tensor E t Both have the same dimensions; record the height (H) and width (W).
[0018] Then, the computing device can determine the RGB image features corresponding to the RGB image, the event tensor features corresponding to the event tensor, the features of each RGB patch, and the features of each event patch (S204). A patch corresponds to an image block in the RGB image or the event tensor. Patches are divided into RGB patches and event patches; an RGB patch corresponds to an image block in the RGB image, and an event patch corresponds to an image block in the event tensor. In other words, the RGB image / event tensor is effectively divided into multiple non-overlapping image blocks, and each image block can correspond to one patch (RGB patch / event patch). Figure 3 The diagram illustrates the mapping between images and patches.
[0019] Figure 3 This is a schematic diagram illustrating the mapping relationship between patch features and images provided in Embodiment 1 of this disclosure.
[0020] exist Figure 3 The original image can be understood as an RGB image or event tensor at a given moment. When the original image is an RGB image, the determined patch features are RGB patch features; when the original image is an event tensor, the determined patch features are event patch features. Figure 3 It can be seen that each patch feature corresponds one-to-one with each image block after the original image has been segmented. Figure 3 The diagram also shows the mapping relationship between the original image, a feature map of the original image, and the features of each patch. This is because the patch features can be determined from the feature map, which will be explained in detail below.
[0021] Specifically, how to determine the RGB image features f? t rgb Event tensor features f t ev Features of each RGB patch and the characteristics of each event patch This will be explained in detail later.
[0022] Then, the computing device can determine the fusion features between the RGB image and the event tensor based on the RGB image features and the event tensor features, and determine the trajectory features based on the fusion features (S206). The trajectory features are used to characterize the motion changes of the target object over a period of time. Each moment can correspond to a trajectory feature at that moment, which can be used to represent the motion changes of the target object from the past to that moment.
[0023] It should be noted that the computing device can determine the weights corresponding to the RGB image and the event tensor respectively, and then determine the fusion features between the RGB image and the event tensor based on the weights corresponding to the RGB image and the event tensor respectively. The specific method for determining the weights corresponding to the RGB image and the event tensor will be explained later.
[0024] Then, the computing device can construct a spatiotemporal-modal heterogeneous graph to represent the relationships and temporal order between different image modalities based on trajectory features, RGB patch features, and event patch features. The spatiotemporal-modal heterogeneous graph contains trajectory nodes, RGB patch nodes, and event patch nodes corresponding to each time step (S208). Specifically, each RGB patch node at a given time step corresponds one-to-one with each RGB patch at that time step, and each event patch node at a given time step corresponds one-to-one with each event patch at that time step.
[0025] Edges in a spatiotemporal-modal heterogeneous graph may include edges connecting RGB patch nodes corresponding to adjacent RGB patches in the same RGB image, edges connecting event patch nodes corresponding to adjacent event patches in the same event tensor, edges representing temporal order (edges between patch nodes of the same modality at the same position in adjacent frames), edges between trajectory nodes and each RGB patch node at the same time, and edges between trajectory nodes and each event patch node at the same time.
[0026] Then, the computing device can perform message passing based on the spatiotemporal-modal heterogeneous graph using a self-attention mechanism, determine the RGB patch node features, event patch node features, and trajectory node features, and perform target tracking based on the RGB patch node features, event patch node features, and trajectory node features (S210).
[0027] It should be noted that the construction of this spatiotemporal-modal heterogeneous graph can be a continuous process. That is, after obtaining the RGB image and event tensor at time 1, various nodes in the spatiotemporal-modal heterogeneous graph corresponding to the RGB image and event tensor at time 1 can be constructed, and the edges between the corresponding nodes can be constructed. Then, after obtaining the RGB image and event tensor at each time step, the corresponding new nodes in the spatiotemporal-modal heterogeneous graph can be constructed, and the edges between the corresponding nodes can be constructed.
[0028] As described in the background section, in existing technologies, decision-level fusion methods perform target tracking independently for both RGB images and event tensors, and then use Kalman filtering, Hungarian matching, or weighted voting for decision fusion. These methods are convenient to implement and interpretable, but because they ignore the intermediate feature correlations between the two modalities, errors may occur if one side is distorted, thus reducing the accuracy of target tracking.
[0029] In view of this, when performing joint target tracking using RGB images and event tensors in this application, the RGB image features of the RGB image and the event tensor features of the event tensor are determined at each time step, and then fused using determined weights to obtain the fused features of the RGB image and the event tensor. Based on the fused features at multiple time steps, trajectory features can be determined. Furthermore, the features of each RGB patch and each event patch at each time step are also determined, thereby constructing a spatiotemporal-modal heterogeneous graph. The spatiotemporal-modal heterogeneous graph can represent the correlation between two image modalities (indirectly represented by the connection of trajectory nodes to RGB patch nodes and event patch nodes respectively), temporal order, and image positional relationships between patches of the same modality. Then, message passing of a self-attention mechanism is performed based on this spatiotemporal-modal heterogeneous graph, and target tracking is performed based on the results of message passing. Therefore, compared with the existing technology that obtains the target tracking results from two modalities separately and then fuses the two results, this application uses a heterogeneous graph to link the image information of the two modalities and the trajectory information (trajectory features) extracted by the fusion features, and then performs target tracking. In this way, the two modalities are indirectly linked together by the heterogeneous graph for target tracking, thereby improving the accuracy of target tracking by combining RGB images and event tensors.
[0030] The details involved in the above steps are explained in detail below. It should be noted that the following explanation will take time t as an example to illustrate the methods for determining RGB image features, event tensor features, features of each RGB patch, features of each event patch, and the fusion features.
[0031] Optionally, before determining the RGB image features corresponding to the RGB image, the event tensor features corresponding to the event tensor, each RGB patch feature, and each event patch feature, the method further includes: determining a first multi-scale feature map corresponding to the RGB image and a second multi-scale feature map corresponding to the event tensor; and wherein the operation of determining the RGB image features corresponding to the RGB image and the event tensor features corresponding to the event tensor includes: determining the RGB image features corresponding to the RGB image based on the first multi-scale feature map corresponding to the RGB image, and determining the event tensor features corresponding to the event tensor based on the second multi-scale feature map corresponding to the event tensor.
[0032] Specifically, a multi-scale encoder can be deployed in the computing device to determine RGB image features, event tensor features, RGB patch features, and event patch features. It should be noted that the term "multi-scale" in this specification refers to RGB image features and event tensor features in a form similar to that of the "FPN pyramid" feature structure.
[0033] The multi-scale encoder may include a multi-scale feature map extraction network, a feature map fusion network, and a patch feature extraction module, corresponding to the RGB image and the event tensor, respectively. The computing device can input the RGB image into the multi-scale feature map extraction network to determine the first multi-scale feature map, and input the event tensor into the multi-scale feature map extraction network to determine the second multi-scale feature map. For the RGB image (or event tensor) at time t, the first multi-scale feature map (or the second multi-scale feature map) may contain four feature maps P. 2,t ~P 5,t (Hereinafter referred to as the first feature maps). The multi-scale feature map extraction network can include a 5-layer ResNet-50 network and FPN. The structure of the 5-layer ResNet-50 network is shown in Table 1.
[0034] Table 1 <![CDATA[Stage 0(C1)]]> The conv1 convolution is followed by max-pooling, with a total stride of 1 / 4. <![CDATA[Stage 1(C2)]]> conv2_x has 3 residual blocks and outputs a stride of 1 / 4. <![CDATA[Stage 2(C3)]]> The conv3_x function has 4 residual blocks and outputs a stride of 1 / 8. <![CDATA[Stage 3(C4)]]> The conv4_x function has 6 residual blocks and outputs a stride of 1 / 16. <![CDATA[Stage 4(C5)]]> conv5_x has 3 residual blocks and outputs a stride of 1 / 32. To conserve computational resources and maintain semantic depth, the multi-scale encoder can retain only four feature maps: C2, C3, C4, and C5, corresponding to the outputs of Stages 1–4 in the table above. Therefore, for time t, four stage feature maps C can be obtained for an image (RGB image or event tensor) of one modality. 2,t ~C 5,t (Hereinafter referred to as the second feature maps). Since the event tensor can be preprocessed into 3-channel data, the event tensor branch can use the same network as the RGB image branch and share parameters.
[0035] Then, the multi-scale encoder can determine the first multi-scale feature map based on the four stage feature maps corresponding to the RGB image through the FPN, and determine the second multi-scale feature map based on the four stage feature maps corresponding to the event tensor.
[0036] Since the method for determining multi-scale feature maps is the same for both RGB images and event tensors, it will be explained uniformly here. First, we can consider the second feature map C of the highest layer among the four staged feature maps. 5,t Perform a 1×1 convolution to obtain the first feature map P of the highest layer in the multi-scale feature map. 5,t Then, each layer of the feature map in the multi-scale feature map is determined sequentially using the following formula (1): (1) Where Up represents upsampling, Conv 3×3 Represents a 3×3 convolution, Conv 1×1 This represents a 1×1 convolution.
[0037] In other words, in the multi-scale feature map, C 2,t ~C4,t The corresponding first feature maps (P) 2,t ~P 4,t (), is obtained by fusing the first feature map of the higher layer and the second feature map of the current layer, for example, P 5,t With C 4,t Fusion yields P 4,t Computing devices can... 5,t After upsampling by two times, and C 4,t The 1×1 convolution results are summed element-wise, then smoothed by a 3×3 convolution to obtain P. 4,t This allows us to obtain the first multi-scale feature map corresponding to the RGB image. And the second multi-scale feature map corresponding to the event tensor. .
[0038] Then, the computing device can use a feature map fusion network to fuse the first multi-scale feature map (or the second multi-scale feature map) of the RGB image (or event tensor) into feature maps of multiple sizes, thereby obtaining RGB image features (or event tensor features) and thus solving the problem of cross-scale semantic misalignment.
[0039] To aggregate information from different resolutions into a description vector of uniform length, computing devices can process each first feature map P in the multi-scale feature maps. l,t Perform adaptive average pooling to obtain the pooled feature map. (l=2~5), where D=256. Then, the pooled features of each size are weighted and summed to obtain the RGB image features (or event tensor features), where the weights are proportional to the spatial size, ensuring that high-resolution layers contribute relatively more. The RGB image features are... The event tensor features are , .in , For the first The height and width of the high-resolution feature map. Because high-resolution layers have a large area and a large relative coefficient, detailed information accounts for a higher proportion in the result.
[0040] Optionally, the operation of determining the features of each RGB patch and each event patch includes: selecting a feature map of a predetermined scale level from the first multi-scale feature map as the first target feature map, and selecting a feature map of a predetermined scale level from the second multi-scale feature map as the second target feature map; segmenting the first target feature map and the second target feature map according to the preset size of the patch, respectively, to obtain each first segmented feature map after segmenting the first target feature map, and each second segmented feature map after segmenting the second target feature map; determining the features of each RGB patch according to the first segmented feature map, and determining the features of each event patch according to the second segmented feature map.
[0041] It can be seen that the process of determining RGB patch features and event patch features is consistent. The number of patches to be divided into for a feature map can be set, and the number and size of patches in a feature map can be fixed. The size of the event patch and the RGB patch are the same. Multiple patches can be obtained by selecting a feature map of a predetermined scale level from a multi-scale feature map and dividing it. In this application, the third-level feature map P is selected. 3,t The computing device can use the patch feature extraction module to extract P... 3,t Divided into multiple patches ( It was split into multiple RGB patches , It was split into multiple event patches ).
[0042] For details, please refer to Figure 3 ,exist Figure 3 The feature map shown can refer to a feature map selected from multi-scale feature maps at a predetermined scale level. This feature map is then subjected to, for example... Figure 3 The segmentation yields several patches, each used to determine corresponding patch features. If the feature map is an RGB image feature map, the segmented patches are RGB patches, and the corresponding RGB patch features can be further determined. If the feature map is an event tensor feature map, the segmented patches are event patches, and the corresponding event patch features can be further determined.
[0043] For example, the length and width of each patch can be set to 's', dividing a feature map into R×C patches, where R is the number of patches in a column of the feature map and C is the number of patches in a row of the feature map. The length and width of each patch are... ,thereby Where H3 is P 3,t The height of W3 is P 3,t The width. Therefore, according to the aforementioned preset patch size and number of patches, it can be... and Perform the same segmentation process to obtain each RGB patch and each event patch.
[0044] Define a patch that has been split as , 0≤r <R,0≤c<C。
[0045] Where r and c represent the row and column indices of the patch, respectively, r represents the r-th row, and c represents the c-th column. t represents time t. This indicates the starting pixel position of the current patch in the height direction. This indicates the end pixel position of the current patch in the height direction. This indicates the starting pixel position of the current patch in the width direction; This indicates the end pixel position of the current patch in the width direction. The above formula can be understood as... 3,t The feature map is sliced to extract patches at specific locations. The first dimension (height) is taken from... arrive All pixels between; the second dimension (width) is taken from... arrive The first dimension contains all pixels between the first and second dimensions; the third dimension (channels) uses a colon (:) to indicate that all channels are retrieved. The colon (:) at the end of the formula means "retrieve all elements in the current dimension". Each patch maintains 256 channels, is non-overlapping, and covers the entire image. To facilitate the subsequent construction of edges in the heterogeneous graph, the computing device can save the 3D coordinates corresponding to each patch. This allows patch nodes with the same 3D coordinates under the same modality to be directly connected when constructing spatial edges in the future.
[0046] Using the methods described above, each RGB patch (each first segmentation feature map) can be obtained. And each event patch (each second segmentation feature map) .
[0047] Then, the RGB patch characteristics corresponding to each RGB patch and the event patch characteristics corresponding to each event patch can be determined.
[0048] Specifically, for a patch Specifically, the patch characteristics corresponding to the patch can be determined using the following formula. : (2) Depthwise represents a 3×3 depthwise separable convolution (convolution only within a channel), while Pointwise represents a 1×1 pointwise convolution fused across channels. Parameters are shared between the RGB image and the event tensor. This operation has a small number of parameters, is computationally fast, and does not require storing separate sets of parameters for the RGB image and the event tensor. This is achieved through RGB patching. By performing the operation corresponding to formula (2) above, the RGB patch features can be obtained. For event patches By performing the operation corresponding to formula (2) above, the event patch characteristics can be obtained. .
[0049] Optionally, the operation of determining the fusion features between the RGB image and the event tensor based on the RGB image features and the event tensor features includes: determining the weights corresponding to the RGB image and the event tensor respectively based on the sharpness score and the motion score of the RGB image; determining the fusion features between the RGB image and the event tensor based on the weights corresponding to the RGB image and the event tensor features, wherein the operation of determining the sharpness score of the RGB image includes: determining the grayscale image of the RGB image; performing a discrete Laplacian convolution on the grayscale image to obtain a second-order gradient image; determining the sharpness score of the RGB image based on the second-order gradient image and the grayscale image, wherein the operation of determining the motion score of the RGB image includes: determining the gradient magnitude image of the RGB image based on the grayscale image of the RGB image; determining the motion score of the RGB image based on the difference between the gradient magnitude image of the RGB image and the gradient magnitude image of the previous frame of the RGB image.
[0050] Specifically, after determining the RGB image features f at time t using the aforementioned multi-scale encoder... t rgb and the event tensor features f at time t t ev Next, the features of these two modalities need to be fused according to the determined weights to obtain the fused feature f between the RGB image and the event tensor at time t. t .
[0051] The computing device can adaptively adjust the contribution ratio of two modalities—RGB image and event tensor—in the fused features based on image sharpness and motion intensity using an adaptive modal weighting fusion module. Image sharpness is represented by the sharpness score of the RGB image, and motion intensity by the motion score. The motion score indicates the degree of displacement or deformation of the scene in the RGB image compared to the previous frame; a higher motion score indicates greater displacement or deformation of the scene relative to the previous frame. The purpose of the adaptive modal weighting fusion module is to assign greater weight to the RGB image when the image is sharp and the motion is weak, and to assign greater weight to the event tensor when the image is blurry and the motion is strong. Thus, the weights corresponding to the RGB image and event tensor can be determined solely using the sharpness score and motion score of the RGB image.
[0052] Specifically, when determining the sharpness score of the RGB image of frame t, the RGB image of frame t can be... First convert it to a grayscale image (Gt) (a weighted method can be used). Then, a discrete Laplacian convolution is performed on the grayscale image. Convolution kernel take The second-order gradient map L is obtained. t = Then, the relationship with the second-order gradient map L can be determined. t The corresponding variance v t And based on the variance v t Determine the sharpness score .
[0053] Specifically, the sharpness score can be determined based on the maximum and minimum variances of the variances corresponding to the second-order gradient map and the variances corresponding to the second-order gradient maps of the previous K RGB images. , where | L t | indicates that the second-order gradient map L is... t The elements in the array are positive. Divide into S×T sub-images B j For each B j Calculate the statistical variance of pixel intensity ,in, , |B j | Represents subimage B j The number of pixels contained. The statistical variance v corresponding to each sub-image. t,j Take the mean value to obtain the variance v corresponding to the second-order gradient map of the RGB image in frame t. t .
[0054] To measure different scenes on the same scale, a sliding window buffer of length K can be maintained. This buffer stores the variance of the second-order gradient maps of the most recent K RGB images. The sharpness score is then determined using the following formula: (3) in, It is compressed to 0~1. These represent the minimum and maximum values within the sliding window buffer, i.e., the maximum and minimum variances of the variances corresponding to the second-order gradient maps of the first K RGB images. ε is a set value, which can be... To prevent the denominator from being zero. The larger the variance, the more high-frequency information there is, and the clearer the image.
[0055] Determine the motion score of the RGB image in frame t. At that time, for grayscale images Based on the Sobel convolution kernel and its transpose Determine the horizontal gradient with vertical gradient Then, the gradient magnitude image of the RGB image in frame t is determined. Next, the gradient magnitude map of the RGB image in frame t is compared with the gradient magnitude map of the RGB image in the previous frame (frame t-1) pixel by pixel. The difference is then absoluteified and averaged across the entire image to obtain the motion score m. t .
[0056] (4) in The width and height of the RGB image. The larger the value, the more displacement or deformation the scene has undergone relative to the previous frame. If... Without the previous frame, Set it to 0.
[0057] Finally, the computing device can determine the sharpness score c. t and exercise score m t Determine the weights corresponding to the RGB image of frame t and the event tensor of frame t. , .
[0058] Specifically, computing devices can convert c t and m t Concatenate into a two-dimensional vector The input is then a three-layer perceptron MLP (the first layer of the perceptron has a width of 32, the second layer has a width of 16, and the third layer has a width of 8; all hidden layers use ReLU). The output layer has two neurons, corresponding to the two modality scores. Then, the instantaneous weights of the two modes are obtained through Softmax. (5) in The two weighting meanings are: A larger value indicates a higher signal-to-noise ratio for the RGB image in this frame. The larger the value, the more reliable the event tensor information.
[0059] Since instantaneous weights fluctuate with each frame, to reduce jitter, a first-order exponent can be used to smooth historical weights, obtaining weights corresponding to the RGB image and event tensor of frame t, respectively. , A smoothing coefficient can be set. The formula is: , (6) in Initialize to 0.5. The closer the smoothing coefficient is to 1, the greater the influence of historical weights; the closer it is to 0, the more it depends on real-time scoring.
[0060] Finally, the fusion feature between the RGB image of frame t and the event tensor of frame t is determined according to the following formula. : (7) Among them, those determined by the above method It preserves the details of the high-quality modal image without excessively neglecting the compensation effect of the other modal image. (Letter) The horizontal line indicates the smoothed weights, and the multiplication sign indicates element-wise weighting.
[0061] The background section also mentions another existing technique: a simple feature fusion method. This method extracts features from the RGB image and the event tensor separately, then fuses the features of the two modalities (e.g., by stitching, weighted summation, etc.) to obtain fused features, which are then used for target tracking. However, this method typically uses fixed-weighted summation or direct stitching, making it difficult to adapt to changes in lighting conditions. In our proposed method, by determining the sharpness and motion scores of the RGB image, the weights between the RGB image and the event tensor can be adaptively adjusted. Therefore, when sudden changes in lighting conditions affect the quality of the RGB image, this method can maintain the accuracy of target tracking by adjusting the weights between the RGB image and the event tensor (e.g., when a sudden change in lighting degrades the quality of the RGB image, this method can adaptively reduce the weights of the RGB image).
[0062] Then, the computing device can determine the trajectory characteristics through the trajectory memory generation module. For each moment, it can generate the corresponding trajectory characteristics h. t Specifically, this can be achieved through the trajectory features h from the previous time step. t-1 The fusion feature f of the current moment t Based on the gated loop unit, the trajectory feature h at the current moment is determined. t .
[0063] In the trajectory memory generation module, the trajectory features are determined by a gated recurrent unit (GRU). When determining the trajectory features at time t, the GRU first calculates the update gate. Then calculate the reset door Then obtain the candidate state. .symbol This represents the Sigmoid activation function. It is a hyperbolic tangent. This represents element-wise multiplication. The trajectory characteristics at time t are: The gate vector dimension is 1. When the target object moves quickly, As the value increases, the newly observed information has a greater weight; when the target is stationary, The value is reduced, but historical information is preserved. Weight matrix With bias All updates are performed during training.
[0064] It should be noted that, in order to facilitate subsequent message passing in heterogeneous graphs, trajectory features can be... Through linear mapping Convert to dimension vector o t The vector o t This can be used as the initial feature of the trajectory node at time t, where , Then, the obtained o t With patch index Alignment yields the feature matrix of the trajectory nodes at time t. ,in This represents a vector consisting entirely of 1s, i.e., all 1s are represented by the same 'o'. t copy Next, make the i-th row correspond to the i-th patch. Expand and stack all frames in the time dimension to obtain the overall matrix: .
[0065] The rows are ordered first by time, then by patch position. This makes it easier to directly pair the trajectory node of frame t with each patch node of that frame in the heterogeneous graph, and to perform attention calculations.
[0066] Optionally, the method further includes: if no second candidate box is matched in the historical trajectory, the historical trajectory is taken as the pending trajectory: based on the trajectory features of the previous moment and the preset attenuation coefficient, pseudo features are generated, and the pseudo features are used to replace the fusion features corresponding to the pending trajectory at the current moment; based on the pseudo features, the trajectory features corresponding to the pending trajectory at the current moment are determined.
[0067] The specific methods of target tracking will be discussed later, in which a "trajectory" is generated. For each time step, the historical trajectory needs to be matched with each second candidate bounding box, thus continuing the historical trajectory up to that time step to obtain the "trajectory" for that moment. However, if, starting from time t, the historical trajectory fails to match any second candidate bounding box (this could be due to occlusion of the target object, etc.), then that historical trajectory can be considered an undetermined trajectory. Furthermore, the corresponding trajectory features are no longer generated in real-time through the fusion of features between the RGB image and the event tensor, but rather through the trajectory features h from the previous time step. t-1 And a preset attenuation coefficient γ is used to generate pseudo-features. ,in, The trajectory feature h at time t is generated using this pseudo-feature. t This avoids issues such as occlusion starting from time t affecting target tracking performance. Specifically, false features can be directly... The trajectory feature h at time t t For time t+1, if the historical trajectory still does not match any second candidate box, then a false feature will still be generated. =γ This pseudo-feature The trajectory feature h at time t+1 t+1 This process continues for subsequent time steps. If a second candidate box is matched again at time t+n, the process continues based on the gated recurrent unit, using the trajectory features h from time t-1. t-1 (Hereinafter referred to as effective trajectory features) and the fusion feature f at time t+n t+n Determine the trajectory characteristics h at time t+n. t+n .
[0068] It should also be noted that the target tracking method provided in this application can be mainly used for single target tracking. Therefore, in single target tracking, the trajectory feature h at time t... t It can be regarded as the trajectory characteristics of a single target object at time t.
[0069] The trajectory memory generation module can create a first-in-first-out (FIFO) buffer queue for active trajectories (trajectories still tracking the target at the current moment, i.e., trajectories that have not yet been deleted at the current moment), with a fixed length. This buffer stores the fused features, or pseudo-features, for each moment (each frame). When the detection head... If the frame still matches the trajectory (i.e., the trajectory matches the second candidate box), then the new frame is... Add the feature to the tail of the queue and pop the oldest feature from the head of the queue, keeping the buffer length unchanged. If a trajectory fails to match any second candidate box from a certain point in time but has not yet been deleted (this type of trajectory is called an undetermined trajectory), generate a pseudo-feature. Fill the tail of the queue, γ∈(0,1) is the attenuation coefficient (usually taken as 0.9), representing the attenuation of the trajectory feature h from the previous time step. t-1 Attenuation is applied. This ensures that even if short-term occlusion leads to missing true observations, the model still receives a smooth input; and since pseudo-features originate from the previous hidden state, they do not introduce random noise. (Symbol) This represents the pseudo-input. The queue can be in the form of {f1, f2, ..., f...}. T}, through this queue, the corresponding {h1, h2, ..., h} can be determined. T}
[0070] Finally, the trajectory memory generation module maintains the memory unit hidden_bank, with the trajectory ID as the key and the most recent valid trajectory feature as the value (i.e., if the trajectory becomes an undetermined trajectory starting from time t, then the valid trajectory feature is h). t-1 If a trajectory is lost in a frame (i.e., fails to match any second candidate box from a certain moment) and is subsequently detected again, the module will read the corresponding data from the hidden_bank. As the initial hidden vector for the GRU, it eliminates the need for re-warm-up time. That is, if a trajectory disappears in a frame (e.g., due to occlusion) but is re-detected in a subsequent frame, the previously stored valid trajectory features are read from the hidden_bank and used as input to the GRU. Without memory, the GRU would have to be initialized from a zero vector every time it reappears, resulting in information loss; with the hidden_bank, trajectory recovery is more consistent. If a trajectory is determined to have completely disappeared and is deleted, the corresponding entry is immediately removed from the hidden_bank, freeing up GPU memory / RAM.
[0071] Optionally, based on trajectory features, RGB patch features, and event patch features, an operation is performed to construct a spatiotemporal-modal heterogeneous graph representing the relationships and temporal order between different image modalities. This includes: constructing trajectory nodes corresponding to trajectory features, RGB patch nodes corresponding to each RGB patch feature, and event patch nodes corresponding to each event patch feature; constructing bidirectional edges between trajectory nodes and RGB patch nodes at the same time, and bidirectional edges between trajectory nodes and event patch nodes at the same time; constructing bidirectional edges between RGB patch nodes at the same position in RGB images at adjacent time points, and bidirectional edges between event patch nodes at the same position in event tensors at adjacent time points; constructing bidirectional edges between horizontally or vertically adjacent RGB patch nodes in RGB images at the same time point, and bidirectional edges between horizontally or vertically adjacent event patch nodes in event tensors at the same time point.
[0072] For details on the form of the spatiotemporal-modal heterogeneous graph, please refer to [reference needed]. Figure 4 .
[0073] Figure 4 This is a schematic diagram of a spatiotemporal modal heterogeneous graph provided in an embodiment of this disclosure.
[0074] It should be noted that, Figure 4 The examples of spatiotemporal-modal heterogeneous graphs are only intended to illustrate the three types of edges in a spatiotemporal-modal heterogeneous graph: spatial edges, temporal edges, and trajectory edges. Therefore, only the RGB patch nodes associated with the RGB image at time t (frame t), the event patch nodes associated with the event tensor at time t, the trajectory nodes corresponding to time t, and the RGB patch nodes associated with the RGB image at time t+1 are shown. Figure 4 As can be seen from this, at time t, there are bidirectional edges between the RGB patch nodes of the vertically or horizontally adjacent RGB patches in the RGB image (it should be noted that...). Figure 4 The arrangement of patch nodes corresponds to the arrangement of the corresponding patches in the feature map. At time t+1, there are bidirectional edges between the RGB patch nodes of vertically or horizontally adjacent RGB patches in the RGB image. At time t, there are bidirectional edges between the event patch nodes of vertically or horizontally adjacent event patches in the event tensor (all of the above are spatial edges). There are bidirectional edges between the trajectory node at time t and each event patch node of the event tensor at time t, and between the trajectory node at time t and each RGB patch node of the RGB image at time t (trajectory edges). There are bidirectional edges between the RGB patch nodes at the same position in the images at time t and time t+1 (temporal edges).
[0075] Specifically, the computing device can generate three types of edges in the spatiotemporal-modal heterogeneous graph according to three relationships through the heterogeneous graph construction module: spatial edges, temporal edges, and trajectory edges. Spatial edges only appear in the same frame. Spatial edges enable the network that subsequently extracts features from the heterogeneous graph to capture local geometry at the same time. Texture, color, and fine-grained edges are transmitted to each other along these edges, allowing the network to notice the continuity of the object's outline on the mesh, and also allowing for smooth compensation of occlusion and lighting differences within the same frame. Spatial edges are stored bidirectionally so that information can flow in both directions. Temporal edges connect nodes of the same modality and the same patch coordinates in two consecutive frames. Since these edges are adjacent in time, the network can directly compare the appearance differences at the same location to learn velocity, deformation, and occlusion order. If the target's displacement across frames is greater than one patch, information can also be indirectly propagated through spatial edges and trajectory edges, thereby maintaining a coherent motion representation. The temporal edge also stores data bidirectionally, facilitating historical feature backtracking and mitigating instantaneous errors caused by camera shake. The trajectory edge connects all patch nodes in each frame to the trajectory node for that frame. The trajectory node carries accumulated appearance and motion memories across frames, while the patch node only contains local observations. Thus, subsequent network iterations can use the trajectory edge to allow patches to inherit global identity cues (i.e., which target the patch belongs to), and the trajectory node, receiving information from the patch node, can better identify newly observed targets, thereby updating the trajectory-related feature representation. The trajectory edge provides an explicit path for the attention mechanism to adjust weights based on historical similarity (the attention mechanism compares the historical features of the trajectory node with the current features of the patch node, assigning greater weight to high similarity and less weight to low similarity), thereby reducing target tracking errors (reducing errors in the association between trajectory and target). The trajectory edge is bidirectional, allowing trajectory nodes to send information back to their corresponding patch nodes.
[0076] Therefore, this method uses three types of edges to represent the spatial relationship of patches, the relationship between different modalities indirectly connected by trajectory nodes, and the temporal order in heterogeneous graphs. In the subsequent process, a multi-head self-attention mechanism is used to learn the features of different types of nodes, which can simultaneously learn cross-modal relationships and spatiotemporal relationships, thereby improving the accuracy of fine-grained detection and association.
[0077] Then, the computing device can perform message passing with a self-attention mechanism based on the spatiotemporal-modal heterogeneous graph constructed above to determine the characteristics of RGB patch nodes, event patch nodes, and trajectory nodes.
[0078] First, the initial features of each node can be initialized. This involves using the RGB patch features determined above. As the initial feature vector for the corresponding RGB patch node, the event patch features determined above are used. As the initial feature vector of the corresponding event patch node, the vector o after transforming the determined trajectory features is...t The initial feature vectors for the corresponding trajectory nodes all have a dimension of 256.
[0079] Then, a Transformer-based graph feature extraction network can be used to perform message passing on the spatiotemporal-modal heterogeneous graph to determine the node features of each node, namely, RGB patch node features, event patch node features, and trajectory node features.
[0080] The graph feature extraction network takes a heterogeneous graph (spatiotemporal-modal heterogeneous graph) G = (V, E) as input. Here, V represents the set of nodes, and E represents the set of edges. Each node v in the node set carries an initial feature vector. Its dimension is denoted by d. Node types are represented by Greek letters. Record the writing of edge types (which can be divided into 6 types: two edges in different directions in the trajectory edge, two edges in different directions in the spatial edge, and two edges in different directions in the temporal edge). This means starting from the node type. Pointer to the end node type The network consists of L stacked layers, each layer comprising an "attention message passing block" and a "feedforward block," with residuals and layer normalization applied between them. A graph feature extraction network can be composed of L Transformer layers. The computational process within a single layer of the graph feature extraction network is explained below.
[0081] First, the mapping between the query vector, key vector, and value vector is calculated. In the... In each Transformer layer, each node generates three vectors: query, key, and value. Writing a query weight matrix for each attention head ,in , This refers to the number of attention heads. Key and value weights are related to the edge type and are written as follows: , .node The query vector is If node If there is an edge between node v and node v, then node v The key vector is The value vector is This approach encodes different types of nodes and edges using different matrices, allowing for shared dimensions while maintaining semantic distinction. (Letters) Representing neighboring nodes, Represents the header number. Representative layer number.
[0082] Next, the attention weights are calculated. After obtaining the query vector and key vector, the graph feature extraction network calculates a scaled dot product attention score for each directed edge. The formula is: The numerator is the vector inner product, and the denominator is... Used to control the value. Then, the node... Perform Softmax on the incoming edges to obtain the normalized weights of each incoming edge. .in Represents a polygon point to The set of neighbors. Letters To express "attention", so That is, the edge (u, v, r) is on the th... The importance of each neighbor in the hierarchy. In this way, the network can learn different focusing patterns within different relationships, highlighting key neighbors while suppressing noisy neighbors.
[0083] Then, messages are aggregated by edge type. The aggregation phase first performs a weighted summation of the value vectors within each head: symbol Represents "message", subscript Indicates the endpoint node, superscript This indicates the layer number, edge type, and header number. Then, group the elements of the same edge type... The individual results are concatenated column by column and then multiplied by the output matrix. get Here It's a concatenation symbol. Matrix Binding with polygons is used to integrate multi-head information and return to the original dimension. This not only utilizes the directional diversity of multiple heads, but also preserves the independence of different polygons, avoiding information confusion.
[0084] The node characteristics are then updated. Since different edge types contribute differently to different endpoint nodes, each pair... There is a corresponding adaptation matrix T r→τ (That is, the fitting matrix T for different edge types of the same endpoint node) r→τ (Different). The updated value is calculated as follows: ,in This is the intermediate vector obtained by fusing all edge information of node v. Residual connections and layer normalization are then applied. , here Layer normalization, or layer normalization, maintains a mean of 0 and a variance of 1, making deep networks easier to train. This indicates the output that has passed through the attention block but has not yet entered the feedforward block. The feedforward block then uses two fully connected layers with GELU activation. The weights of the first layer are... bias The weights of the second layer are bias .in Can be set to The calculation process is as follows: GELU is an element-wise nonlinear function that produces smooth gradients. Then, residuals and layer normalization are used again: This gives us the input for the next layer (layer l+1). .letter The superscript within the square brackets indicates the output of the feedforward block. This indicates that it belongs to the first Layers. This design allows the network to learn local feature transformations after global relation aggregation, thus improving its expressive power.
[0085] The above steps illustrate a complete Transformer layer. Stacking layers with the same structure L times yields a complete graph feature extraction network, with layer numbers ranging from 0 to L-1. Each layer independently has its own weight matrix, such as... , etc., to enhance model capacity. The final node features are denoted as... Because the entire computation repeatedly utilizes three types of edges—space, time, and trajectory—and simultaneously encodes patch proximity appearance consistency, cross-frame motion continuity, and target identity history, it is suitable for subsequent tasks such as target detection, associating trajectories with detected bounding boxes (target tracking), and re-detection.
[0086] Optionally, the target tracking operation based on RGB patch node features, event patch node features, and trajectory node features includes: fusing the RGB patch node features and event patch node features corresponding to the same image position at the same time to obtain a first fused node feature; determining a second fused node feature based on the first fused node feature using a multi-head self-attention mechanism; determining a first candidate box and the probability of target existence based on the second fused node feature; determining a second candidate box based on the first candidate box and a first multi-scale feature map; performing correlation detection between the second candidate box and historical trajectories to determine a second candidate box that matches the historical trajectory as the target box, wherein the second candidate box is correlated with the historical trajectory... The operation of performing correlation detection between historical trajectories and determining the second candidate box that matches the historical trajectory includes: determining the similarity matrix between the historical trajectory and the second candidate box based on the trajectory node features and the second fusion node features corresponding to the current time. The similarity matrix is used to represent the similarity between the second candidate box and the historical trajectory, where the historical trajectory is obtained through target tracking in the past; determining the cross-temporal spatial intersection-union ratio, which represents the intersection-union ratio between the second candidate box at the current time and the second candidate box at the previous time in the historical trajectory; determining the comprehensive cost based on the cross-temporal spatial intersection-union ratio and the similarity matrix, and determining the second candidate box that matches the historical trajectory with the goal of minimizing the comprehensive cost.
[0087] It should be noted that target tracking is a continuous process. In this embodiment, the computing device continuously acquires RGB images and event tensors to continuously track the target. Each moment corresponds to one RGB image and one event tensor. Taking the period from moment 1 to moment 3 as an example, at moment 1, the second candidate box at moment 1 can be determined by the detection head using the above method. At this time, it can be understood that a continuous trajectory has not yet been formed. At moment 2, the detection head continues to determine the second candidate box at moment 2. The association head performs association detection, matching the second candidate box at moment 2 with the second candidate box at moment 1. The successfully matched second candidate boxes form a trajectory. At moment 3, the detection head continues to determine the second candidate box at moment 3. The second candidate box at moment 3 is matched with the trajectory determined at moment 2 (i.e., the historical trajectory). If the match is successful, the trajectory is extended. That is to say, a trajectory (historical trajectory) can be composed of second candidate boxes from multiple moments.
[0088] Specifically, when determining the second candidate box using the detection head, the computing device first fuses the RGB patch node features and event patch node features corresponding to the same image position at the same time to obtain the first fused node feature. .vector After passing through two fully connected layers, each layer is followed by a ReLU activation function and layer normalization, outputting intermediate features. All locations The patch coordinates are then rearranged into a matrix, and the second fusion node features are generated using the multi-head self-attention module (MHA). The number of self-attention heads is The width of a single head is Then, Input to linear layer (weights) With bias Obtain target information Target information This indicates whether a target object exists in the patch corresponding to the feature of the second fusion node. Then, the probability P of the target existence corresponding to each patch can be output using the Softmax operation. tar,r,c,t =Softmax(s r,c,t ).
[0089] Then, the detection head can be used to detect features based on the second fusion node. This determines the corresponding second candidate box. Wherein, the probability P of the target existing is... tar,r,c,t When the value is higher than a preset value (e.g., 0.7), a corresponding second candidate box can be determined. The detection head uses a linear layer (weighted) Bias ), and the features of the second fusion node Mapped to quadruples (First candidate box), where The coordinates are relative to the center of the frame. The relative values of the frame width and height are located at... Then inspect the headband. After restoring to the original image size, the first candidate box is used in the first multi-scale feature... Extract the region of interest (perform the ROIAlign operation).
[0090] For each first candidate bounding box, the image features corresponding to the position of the first candidate bounding box are determined based on each feature map in the first multi-scale feature map. Then, the second candidate bounding box is determined based on the image features corresponding to the position of the first candidate bounding box and the features of the second fusion node. Specifically, for example, if there are 7×7 patches, each feature map in the first multi-scale feature map can be divided into 7×7 grids, and the pooling tensor can be obtained through bilinear interpolation. (Image features corresponding to the location of the first candidate box). All First, flatten each component, then concatenate them into a vector according to hierarchical order. .vector Features of the second fusion node The data is concatenated and input into a two-layer perceptron with weights W.ref Bias is b ref The second candidate box is obtained. (The refined box is relative to the first candidate box). Symbol Represents multi-scale pooling information. This indicates vector concatenation.
[0091] Then, the association header is used to detect the association between the historical trajectory and the second candidate box at the current time (e.g., the second candidate box at time t is associated with the historical trajectory at time t-1). If the historical trajectory matches the second candidate box, then the second candidate box can be used as the target box B at the current time. t .
[0092] Next, the association header will include the trajectory node features at time t. Through fully connected layers (weight matrix) Bias The vector obtained after projection The vector is converted into a trajectory embedding by L2 regularization. Then, the association head can use another fully connected layer (weight matrix W). det Bias b det ) Features of the second fusion node Projection yields vector v r,c,t Then, similarly for vector v r,c,t Perform L2 regularization to obtain the detection embedding. ,in, .
[0093] Then, the computing device can construct a similarity matrix at time t to represent the similarity between the historical trajectory and the second candidate box using the association header. Where sim(a, b) represents calculating the cosine similarity between elements a and b. This represents the index of each second candidate box during traversal. At time t, j represents the j-th second candidate box at time t.
[0094] Furthermore, the correlation header can calculate the spatial intersection-union ratio across time periods. Among them, the cross-time spatial intersection-union ratio This can refer to the intersection-union ratio (IoU) between the j-th second candidate box at the current time and the second candidate box in the historical trajectory at the previous time (t-1 time). Then, based on the similarity matrix and the cross-time spatial IoU, the comprehensive cost representing the matching degree between the historical trajectory and the second candidate boxes is determined. The formula is: (8) Where the coefficient , This represents the balance factor between appearance and geometric weights.
[0095] Specifically, the computing device can minimize this comprehensive cost using the Hungarian algorithm, thereby determining a second candidate box that matches the historical trajectory, and using the second candidate box that matches the historical trajectory as the target box B. t The comprehensive cost corresponding to each second candidate box can be determined. The minimum value can reduce the overall cost. The second candidate box with the smallest minimum value is selected as the target box B. t , the target box B t As a second candidate box that matches the historical trajectory.
[0096] If the second candidate box does not match a historical trajectory, it can be used as the starting point of a new trajectory. If a historical trajectory does not match any second candidate box at the current moment, the loss count of the historical trajectory will be recorded (or the loss count of the historical trajectory will be incremented by 1). When the loss count exceeds the third preset threshold, the historical trajectory can enter the re-detection (the re-detection module will re-match the historical trajectory with the target object).
[0097] By combining the association head and the detection head, the detection head can determine the probability of the target's presence in the image at each time step, as well as the second candidate box representing the target's location. The association head can then associate these second candidate boxes to determine the trajectory of the same target at multiple consecutive time steps, thereby achieving target tracking.
[0098] Optionally, the method further includes: if no second candidate box is matched in the historical trajectory, the historical trajectory is taken as a pending trajectory; pseudo features are generated based on the trajectory features of the previous moment and a preset attenuation coefficient, and the pseudo features are used as the trajectory features corresponding to the pending trajectory at the current moment; the method further includes: if the pending trajectory meets a first preset condition and the RGB image at the current moment meets a second preset condition, the pending trajectory is taken as a target historical trajectory, and re-detection is performed based on the target historical trajectory to determine a third candidate box, wherein the first preset condition includes that the trajectory does not match a second candidate box for multiple consecutive moments (multiple consecutive frames), and the second preset condition includes that the occlusion probability corresponding to the RGB image at the current moment is higher than a first preset threshold; and correlation detection is performed between the third candidate box and the target historical trajectory.
[0099] If the target's historical trajectory successfully matches the third candidate bounding box, then regular association detection continues using the association head and detection head. However, if no third candidate bounding box is found for the target's historical trajectory in multiple consecutive frames (the number of consecutive frames can be preset), then the target's historical trajectory is deleted. The aforementioned target existence probability P... tar,r,c,t It can be used to determine whether the second preset condition is met.
[0100] Specifically, the computing device determines that the probability O of the target corresponding to each patch not existing is... r,c,t =1-P tar,r,c,t Then determine whether the target box B exists at the previous time t-1. t-1 If B exists t-1 Then the target box B is determined. t-1 The corresponding area of the expanded frame: R t =expand(B t-1 , ρ), where ρ is the expansion coefficient, which can be set to 1.5 (i.e., according to the target box B). t-1 Enlarge the width and height by 1.5 times each, keeping the center unchanged, to obtain R. t If the target box B did not exist at the previous time t-1. t-1 Then continue using the target bounding box B from the previous time step. t−2 If no second candidate box is found up to the first frame, the occlusion probability is directly determined to be O. t =0.
[0101] Then, determine the area weight corresponding to each patch: patch r,c This refers to the area corresponding to each patch. Then, based on the area weights, the first weighted average result is determined. Second weighted average result S t S represents the set of patches. top For S t Press O r,c,t Sort the subsets from largest to smallest by the top q (q can be set to 30%). The first weighted result is obtained by directly weighting according to the area weight. The second weighted result only takes 0. r.c,t Higher-weighted patches are obtained by weighting them according to their area. Then, the first weighted average and the second weighted average are combined to obtain the base occlusion probability. The parameter η can be set to 0.6. Then, the base occlusion probability can be smoothed over time to obtain the final occlusion probability. The parameter λ can be set to 0.5, O t-1 This is the historical occlusion probability. When t is the initial time step or there is no historical occlusion probability, the base occlusion probability is directly used as the final occlusion probability: .
[0102] The second preset condition mentioned above can be an occlusion probability of 0. t Not less than the preset probability θ on Where the preset probability θ on It can be set to 0.6. When the condition is met... The target object is determined to be "occluded". A re-detection process is initiated when both the second and first preset conditions are met. Furthermore, the second preset condition can also be that the object is satisfied for M consecutive frames. (For example, the default setting M=1 is sufficient; M=2 can be set when there is high data noise). In other words, in practical applications, certain situations may lead to target tracking errors (such as offsets in the second candidate bounding box provided by the detection head, or large areas of foreground occlusion in the image causing trajectory tracking interruptions). In such cases, the re-detection module can be activated. The re-detection module uses trajectory embedding... This guides the patch features to restore the correct position of the detection box as much as possible, resulting in a third candidate box. Then, the association header is used to re-detect the association between the third candidate box and the historical trajectory through the third candidate box. This allows trajectories that failed to match the second candidate box (trajectories that failed to achieve target tracking) to continue matching the third candidate box from the interrupted part, thereby improving the accuracy of target tracking.
[0103] Specifically, when the computing device determines the third candidate box through the re-detection module, it can determine the corresponding trajectory embedding. ,in The number of feature channels is given. All patch node features at the corresponding time point (including RGB patch node features and event node features) are determined, and all patch node features (including RGB patch node features corresponding to RGB patches and event patch nodes corresponding to event patches) are stacked into a matrix. .in, , It is a matrix composed of RGB patch node features. It is a matrix composed of the features of event patch nodes. Equal to the total number of patches, and the number of columns is also equal. .
[0104] First, the re-detection module determines the similarity between the trajectory embedding and the features of each patch node by scaling dot product attention. The calculation formula is as follows: To obtain the weight vector , No. Each component The larger the value, the more likely it is a patch. The closer it is to the appearance of the trajectory, the better. Then, through the weight vector... For matrix Perform weighted summation , to obtain the context vector This context vector incorporates information from the most relevant patches of the trajectory and preserves the channel dimension. .
[0105] Secondly, to maintain a stable appearance description during long-term occlusion, the re-detection module saves the corresponding appearance template for the trajectory (historical trajectory), denoted as... The appearance template can refer to a trajectory-stable identity feature vector. If the trajectory enters the re-detection process for the first time, the appearance template is initialized to the trajectory embedding at the current moment. If the appearance template already exists, update the appearance template for that trajectory using an exponential averaging method. ,in It is a retention factor, by default. .high Preserve historical appearance, low This allows the trajectory to quickly follow the new appearance. The formula performs only one vector weighting, introduces no extra layers, and does not affect the gradient stability of the backpropagation path.
[0106] After updating the appearance template, the attention score between the appearance template and the patch node features is calculated again. and retrieve the template context. ,symbol and The meaning is the same, only the query vector is replaced with an appearance template. Next, the three features are embedded: trajectory The current context vector Template context — The vector is obtained by concatenating the data along the channel dimension. The splicing operation does not change the element values, only the shape of the tensor, making it easier for the next linear layer to process all the information at once.
[0107] Final vector Feed it into a three-layer fully connected network. First layer weights. Keeping the dimensions constant, the second layer of weights Compress the dimensions to Third layer weights Press down again Each layer is followed by ReLU and layer normalization to improve nonlinear representation while ensuring distribution stability. The last layer uses linear regression weights. Bias Output five-dimensional vector (Third candidate box), where The coordinates of the candidate box center are: Width and height are normalized relative values. This represents the confidence score.
[0108] like (Fourth preset threshold) (Set on the validation set), the third candidate box is sent to the association head. The association head re-detects the association with the target's historical trajectory based on this third candidate box; otherwise, it waits for the next frame. The module maintains a corresponding cooling counter for the trajectory. When the trajectory is continuous The re-detection trigger condition was met in all frames, but the confidence score was low. Failed to meet the fourth preset threshold This module determines that the trajectory has been lost and removes the corresponding appearance template of the trajectory from the cache (and deletes the trajectory itself), freeing up video memory. (Default) It can be adjusted according to frame rate and scene density.
[0109] Finally, the overall network structure adopted in this method may include the aforementioned multi-scale encoder (including a multi-scale feature map extraction network, a feature map fusion network, and a patch feature extraction module), the network in the adaptive modality weight fusion module, the GRU for trajectory feature extraction, the graph feature extraction network, the association head, the detection head, and the re-detection module. During the training phase, the overall network structure can be trained using pre-annotated RGB images and event tensor image pairs. The annotation of the image pairs includes whether a target object exists in the image, the bounding box representing the actual location of the target object, and the actual trajectory of the target object. The loss function used by the overall network structure during training can be expressed by the following formula: (9) Wherein, the weight λ in the formula det , λ rank and λ tri This can be set manually. For example, set it to λ. det =1、λ rank =0.5 and λ tri =0.3. Represents the loss of the detection branch. Represents the loss of the sorting branch. This represents the branch loss of the triplet.
[0110] The detection branch can be broken down into three items: L det =L focal +L L1 +L DIoU The first item L focal Using Focal Loss to suppress inter-class imbalance: (10) The modulation factor γ is used to reduce the weight of easily classified samples and increase the weight of difficult-to-classify samples; the balance factor α is used to alleviate the imbalance between positive and negative samples. (Symbol p) t∈[0,1] represents the predicted foreground probability (corresponding to the target existence probability P). tar,r,c,t For example, we can set the modulation factor. Balance factor When the samples are easy to separate or Making the gradient smaller, while keeping the gradient large for difficult samples, helps with learning.
[0111] Second item This loss term represents the difference between the bounding boxes detected by the model (such as the first, second, and third candidate boxes) and the actual bounding boxes. It is applied to the box center. With size Apply absolute error penalty: (11) Among them, in the above formula For predicted values, This represents the number of positive sample frames.
[0112] Third item : The difference between predicted and actual values is measured by the intersection-over-union ratio (IoU) and center distance. The IoU measures overlap. The distance between the centers of the predicted bounding box and the ground truth bounding box is the Euclidean distance. This is the diagonal length of the minimum bounding box. Because Taking both area and location into account, the boundary is more stable after training. Three parameters are fed back in parallel, sharing the same front-end network parameters.
[0113] The ranking branch loss ensures that the score of the second candidate box that matches the trajectory is higher than the score of the second candidate box that does not match the trajectory. Given the scores of the same trajectory, negative pair fractions of different trajectories The loss formula is: (12) letter For fixed intervals. If and Separated for more than If the loss is zero, then the gradient pushes the network to increase positive scores or decrease negative scores. Among these, scores corresponding to the same trajectory are... Negative Pair Score of Different Trajectories All sources: The same trajectory corresponds to the score. This represents the matching score between the historical trajectory and the second candidate box belonging to that historical trajectory. Negative pairing score for dissimilar trajectories. This represents the matching score between the historical trajectory and the second candidate box that does not belong to that historical trajectory. The ranking loss is only calculated between successfully matched box pairs, so Hungarian matching needs to be performed first, and then positive pairs of the same trajectory and negative pairs of different trajectories are taken from the resulting pairs.
[0114] Finally, the triplet loss is learned in the metric space.
[0115] Trajectory embedding notation (Right now, ), same trajectory patch embedding is recorded as Different trajectory patch embedding is recorded as The form of loss is: (13) Among them, symbols Represents the Euclidean norm. If the positive pair distance is already less than the negative pair distance minus... This triple does not produce a gradient; otherwise, the network would... Zoom in Simultaneously push away The active trajectory selects the patch closest to it in the current frame as the positive sample, and randomly selects three patches with different IDs in the same frame as negative samples. This allows for the construction of rich triples in one forward pass while keeping the computational cost under control. Both same-trajectory patch embedding and different-trajectory patch embedding are derived from the features of the second fusion node. Corresponding detection embedding The same trajectory patch embedding represents the detection embedding corresponding to the second candidate box that matches the historical trajectory. Different trajectory patch embeddings represent the detection embeddings corresponding to second candidate boxes that do not match historical trajectories. .
[0116] In addition, refer to Figure 1 As shown, according to a second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program is executed, a processor performs any of the methods described above.
[0117] Therefore, according to this embodiment, the method can improve the accuracy of target tracking.
[0118] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0119] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0120] Example 2 Figure 5 A target tracking device according to a first aspect of this embodiment is shown, which corresponds to the method described according to a first aspect of Embodiment 1. Reference Figure 5 As shown, the target tracking device includes: an acquisition module 510 for acquiring RGB images and event tensors, wherein the event tensor is obtained by preprocessing event data collected by an event camera; a feature determination module 520 for determining RGB image features corresponding to the RGB image, event tensor features corresponding to the event tensor, features of each RGB patch, and features of each event patch; a fusion module 530 for determining the fusion features between the RGB image and the event tensor based on the RGB image features and the event tensor features, and determining trajectory features based on the fusion features, wherein the trajectory features are used to characterize the motion changes of the target object from the past to the present moment; a heterogeneous graph construction module 540 for constructing a spatiotemporal-modal heterogeneous graph representing the correlation and temporal order between different image modalities based on the trajectory features, features of each RGB patch, and features of each event patch; and a target tracking module 550 for performing message passing with a self-attention mechanism based on the spatiotemporal-modal heterogeneous graph, determining RGB patch node features, event patch node features, and trajectory node features, and performing target tracking based on the RGB patch node features, event patch node features, and trajectory node features.
[0121] Therefore, according to this embodiment, the accuracy of target tracking can be improved.
[0122] Example 3 Figure 6 A target tracking device according to a first aspect of this embodiment is shown, which corresponds to the method described according to a first aspect of Embodiment 1. Reference Figure 6As shown, the target tracking device includes: a processor 610; and a memory 620 connected to the processor 610, used to provide the processor 610 with instructions to perform the following processing steps: acquiring an RGB image and an event tensor, wherein the event tensor is obtained by preprocessing event data acquired by an event camera; determining RGB image features corresponding to the RGB image, event tensor features corresponding to the event tensor, features of each RGB patch, and features of each event patch; determining the fusion features between the RGB image and the event tensor based on the RGB image features and the event tensor features; and determining the trajectory features based on the fusion features. This is used to characterize the motion changes of a target object from the past to the present moment; based on trajectory features, RGB patch features, and event patch features, a spatiotemporal-modal heterogeneous graph is constructed to represent the correlation and temporal order between different image modalities. The spatiotemporal-modal heterogeneous graph contains trajectory nodes, RGB patch nodes, and event patch nodes corresponding to each moment; and message passing based on the spatiotemporal-modal heterogeneous graph is performed using a self-attention mechanism to determine the features of RGB patch nodes, event patch nodes, and trajectory nodes, and target tracking is performed based on the features of RGB patch nodes, event patch nodes, and trajectory nodes.
[0123] Therefore, according to this embodiment, the accuracy of target tracking is improved.
[0124] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0125] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0126] It should be noted that the devices in Embodiments 2 and 3 belong to the same inventive concept as the method in Embodiment 1, solve the same technical problems, and achieve the same technical effects. The device in Embodiment 2 and the system in Embodiment 1 can implement all the methods in Embodiment 1. The similarities will not be repeated here.
[0127] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0128] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0129] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0130] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0131] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A target tracking method, characterized in that, include: RGB images and event tensors are acquired, wherein the event tensor is obtained by preprocessing event data acquired by the event camera; Determine the RGB image features corresponding to the RGB image, the event tensor features corresponding to the event tensor, the RGB patch features, and the event patch features, wherein the patch includes RGB patches and event patches, and one patch corresponds to an image block in the RGB image or event tensor; Based on the RGB image features and the event tensor features, a fusion feature between the RGB image and the event tensor is determined, and a trajectory feature is determined based on the fusion feature. The trajectory feature is used to characterize the motion changes of the target object from the past to the present moment. Based on the trajectory features, the RGB patch features, and the event patch features, a spatiotemporal-modal heterogeneous graph is constructed to represent the correlation and temporal order between different image modalities. The spatiotemporal-modal heterogeneous graph contains trajectory nodes, RGB patch nodes, and event patch nodes corresponding to each moment. The spatiotemporal-modal heterogeneous graph indirectly represents the correlation between two image modalities by connecting the trajectory nodes with the RGB patch nodes and the event patch nodes, respectively. The spatiotemporal-modal heterogeneous graph is also used to represent the image positional relationship between patches of the same modality. as well as The message passing mechanism based on the spatiotemporal-modal heterogeneous graph determines the RGB patch node features, event patch node features, and trajectory node features, and performs target tracking based on the RGB patch node features, event patch node features, and trajectory node features.
2. The method according to claim 1, characterized in that, Before determining the RGB image features corresponding to the RGB image, the event tensor features corresponding to the event tensor, each RGB patch feature, and each event patch feature, the method further includes: A first multi-scale feature map corresponding to the RGB image and a second multi-scale feature map corresponding to the event tensor are determined respectively; and wherein, The operation of determining the RGB image features corresponding to the RGB image and the event tensor features corresponding to the event tensor includes: Based on the first multi-scale feature map corresponding to the RGB image, determine the RGB image features corresponding to the RGB image; and based on the second multi-scale feature map corresponding to the event tensor, determine the event tensor features corresponding to the event tensor.
3. The method according to claim 2, characterized in that, The operations for determining the characteristics of each RGB patch and each event patch include: A feature map of a predetermined scale level is selected from the first multi-scale feature map as the first target feature map, and a feature map of a predetermined scale level is selected from the second multi-scale feature map as the second target feature map. According to the preset size of the patch, the first target feature map and the second target feature map are segmented respectively to obtain each first segmented feature map after segmenting the first target feature map and each second segmented feature map after segmenting the second target feature map; Based on the first segmentation feature maps, the RGB patch features are determined, and based on the second segmentation feature maps, the event patch features are determined.
4. The method according to claim 1, characterized in that, The operation of determining the fusion features between the RGB image and the event tensor based on the RGB image features and the event tensor features includes: Based on the sharpness score and motion score of the RGB image, determine the weights corresponding to the RGB image and the event tensor, respectively. Based on the weights corresponding to the RGB image and the event tensor, respectively, and based on the features of the RGB image and the event tensor, a fusion feature between the RGB image and the event tensor is determined, wherein... The operation of determining the sharpness score of the RGB image includes: Determine the grayscale image of the RGB image; The grayscale image is subjected to discrete Laplacian convolution to obtain a second-order gradient image; Based on the second-order gradient map, the sharpness score of the RGB image is determined, and wherein, The operation of determining the motion score of the RGB image includes: Based on the grayscale image of the RGB image, determine the gradient magnitude map of the RGB image; The motion score of the RGB image is determined based on the difference between the gradient magnitude map of the RGB image and the gradient magnitude map of the previous frame of the RGB image.
5. The method according to claim 1, characterized in that, The operation of constructing a spatiotemporal-modal heterogeneous graph representing the correlation and temporal order between different image modalities based on the trajectory features, the RGB patch features, and the event patch features includes: Construct trajectory nodes corresponding to the trajectory features, RGB patch nodes corresponding to each RGB patch feature, and event patch nodes corresponding to each event patch feature; Construct bidirectional edges between trajectory nodes and RGB patch nodes at the same time, and between trajectory nodes and event patch nodes at the same time; Construct bidirectional edges between RGB patch nodes at the same position in RGB images at adjacent time points, and construct bidirectional edges between event patch nodes at the same position in event tensors at adjacent time points; Construct bidirectional edges between horizontally or vertically adjacent RGB patch nodes of the RGB image at the same time, and construct bidirectional edges between horizontally or vertically adjacent event patch nodes of the event tensor at the same time.
6. The method according to claim 2, characterized in that, The target tracking operation based on the RGB patch node features, the event patch node features, and the trajectory node features includes: The RGB patch node features and event patch node features corresponding to the same image position at the same time are fused to obtain the first fused node feature; Based on the multi-head self-attention mechanism, the features of the second fusion node are determined according to the features of the first fusion node; Based on the features of the second fusion node, determine the probability of the existence of the first candidate box and the target; Based on the first candidate box and the first multi-scale feature map, a second candidate box is determined; and A correlation detection is performed between the second candidate bounding box and the historical trajectory to determine the second candidate bounding box that matches the historical trajectory, which is then used as the target bounding box. The historical trajectory is obtained through target tracking in the past. The operation of performing correlation detection between the second candidate box and the historical trajectory to determine the second candidate box that matches the historical trajectory includes: Based on the trajectory node features and the second fusion node features corresponding to the current moment, a similarity matrix between the historical trajectory and the second candidate box is determined. The similarity matrix is used to represent the similarity between the second candidate box and the historical trajectory. Determine the cross-time spatial intersection-union ratio, which represents the intersection-union ratio between the second candidate box at the current time and the second candidate box at the previous time in the historical trajectory; Based on the cross-temporal spatial intersection-union ratio and the similarity matrix, a comprehensive cost is determined, and a second candidate box matching the historical trajectory is determined with the goal of minimizing the comprehensive cost.
7. The method according to claim 6, characterized in that, The method further includes: If no second candidate box is matched in the historical trajectory at the current moment, the historical trajectory will be regarded as a pending trajectory: Based on the trajectory characteristics of the previous moment and the preset attenuation coefficient, generate pseudo features; The pseudo-features are used as the trajectory features corresponding to the undetermined trajectory at the current time. Furthermore, the method further includes: If the undetermined trajectory meets a first preset condition and the RGB image at the current moment meets a second preset condition, the undetermined trajectory is taken as the target historical trajectory, and re-detection is performed based on the target historical trajectory to determine a third candidate box. The first preset condition includes that the historical trajectory has not matched a second candidate box within multiple consecutive time periods, and the second preset condition includes that the occlusion probability corresponding to the RGB image at the current moment is higher than a first preset threshold. The third candidate box is re-correlated with the target historical trajectory.
8. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, the method described in any one of claims 1 to 7 is performed by a processor.
9. A target tracking device, characterized in that, include: The acquisition module is used to acquire RGB images and event tensors, wherein the event tensors are obtained by preprocessing event data collected by the event camera; The feature determination module is used to determine the RGB image features corresponding to the RGB image, the event tensor features corresponding to the event tensor, the RGB patch features, and the event patch features, wherein the patch includes RGB patches and event patches, and one patch corresponds to an image block in the RGB image or event tensor; The fusion module is used to determine the fusion features between the RGB image and the event tensor based on the RGB image features and the event tensor features, and to determine the trajectory features based on the fusion features, wherein the trajectory features are used to characterize the motion changes of the target object from the past to the present moment; A heterogeneous graph construction module is used to construct a spatiotemporal-modal heterogeneous graph representing the correlation and temporal order between different image modalities based on the trajectory features, the RGB patch features, and the event patch features. The spatiotemporal-modal heterogeneous graph includes trajectory nodes, RGB patch nodes, and event patch nodes corresponding to each time step. The spatiotemporal-modal heterogeneous graph indirectly represents the correlation between two image modalities by connecting trajectory nodes to RGB patch nodes and event patch nodes, respectively. The spatiotemporal-modal heterogeneous graph is also used to represent the image positional relationship between patches of the same modality. The target tracking module is used for message passing based on the spatiotemporal-modal heterogeneous graph using a self-attention mechanism, determining RGB patch node features, event patch node features, and trajectory node features, and performing target tracking based on the RGB patch node features, the event patch node features, and the trajectory node features.
10. A target tracking device, characterized in that, include: processor; as well as A memory, connected to the processor, for providing the processor with instructions to perform the following processing steps: RGB images and event tensors are acquired, wherein the event tensor is obtained by preprocessing event data acquired by the event camera; Determine the RGB image features corresponding to the RGB image, the event tensor features corresponding to the event tensor, the RGB patch features, and the event patch features, wherein the patch includes RGB patches and event patches, and one patch corresponds to an image block in the RGB image or event tensor; Based on the RGB image features and the event tensor features, a fusion feature between the RGB image and the event tensor is determined, and a trajectory feature is determined based on the fusion feature. The trajectory feature is used to characterize the motion changes of the target object from the past to the present moment. Based on the trajectory features, the RGB patch features, and the event patch features, a spatiotemporal-modal heterogeneous graph is constructed to represent the correlation and temporal order between different image modalities. The spatiotemporal-modal heterogeneous graph contains trajectory nodes, RGB patch nodes, and event patch nodes corresponding to each moment. The spatiotemporal-modal heterogeneous graph indirectly represents the correlation between two image modalities by connecting the trajectory nodes with the RGB patch nodes and the event patch nodes, respectively. The spatiotemporal-modal heterogeneous graph is also used to represent the image positional relationship between patches of the same modality. as well as The message passing mechanism based on the spatiotemporal-modal heterogeneous graph determines the RGB patch node features, event patch node features, and trajectory node features, and performs target tracking based on the RGB patch node features, event patch node features, and trajectory node features.
Citation Information
Patent Citations
Target identification tracking method and system based on multi-source fusion imaging
CN120182323A
Pedestrian trajectory tracking method and system, and related apparatus
WO2023206904A1