Object Tracking Method Based on Causality and Attention Mechanism
By introducing causal and or logical graph and graph convolution networks into the target tracking method, the causal logical relationship and spatiotemporal information are integrated, and the tracking error problem of vehicle targets in the occlusion situation is solved, achieving higher tracking accuracy and robustness.
Patent Information
- Application Number
- CN202411385757.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-09-30
AI Technical Summary
The existing target tracking method is prone to tracking errors when vehicle targets are blocked in road environments, and cannot effectively overcome the impact of occlusion, resulting in missed inspection and missed detection.
The goal tracking method based on causal relationship and attention mechanism is adopted to fuse the causal logical relationship with space-time information through the convolutional network of OR logic graphs and graphs to realize high-level semantic guidance status prediction.
It effectively reduces the dependence on independent and homogeneous distribution assumptions, improves the tracking accuracy of vehicle targets in the case of occlusion, and reduces the number of tracking error matches.
Smart Images

Figure CN119359762B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target tracking, and particularly relates to the tracking technology of road vehicles. Specifically, a target tracking method based on causal relationship and attention mechanism is disclosed. Background Art
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art or existing technology.
[0003] Vehicle target tracking is a key step in realizing the autonomous driving function. In recent years, with the development of artificial intelligence technology, significant progress has been made in vehicle target tracking methods. However, in the vehicle target tracking scenario on the road, during the target driving process, it interacts with other vehicles, pedestrians, and static buildings, resulting in occlusion, which makes the target visibility change frequently, leading to target missed detection and false detection. Therefore, there is an urgent need to design a high-precision real-time tracking method that can overcome the influence of occlusion.
[0004] Early target tracking methods include methods based on feature matching (such as template matching method), methods based on statistical laws (such as Kalman filter, particle filter), etc. With the rapid development of deep neural networks, due to their more expressive features, the tracking accuracy and robustness have been improved. In particular, two-stage target tracking methods based on Convolutional Neural Network (CNN) have been continuously proposed. The two-stage target tracking method first detects the position of the target box, and then associates the target box with the existing tracking trajectory. The two-stage method can select the most advanced models for each stage to be combined, and usually can improve the overall tracking performance. The representative algorithm is the DeepSort model that combines the YOLO detection model, Kalman filter, and Hungarian associator. However, in the case of frequent / long-term occlusion of the target, the CNN-based detection model cannot provide effective target boxes, resulting in a decrease in the IoU matching accuracy of the association model, thus leading to tracking failure. The above-mentioned methods cannot solve the missed detection and false detection caused by occlusion. The fundamental reasons are as follows: First, they are all based on the assumption that the prior data and the actual data are independently and identically distributed. When the actual data does not conform to the prior data distribution, the parameters obtained based on the original distribution fail, and thus the method fails. For example, when the target is occluded by an object with a completely different appearance, the probability distribution of the appearance features of the target position in the image changes, or when the target is occluded by an object with a similar appearance but opposite motion direction, the probability distribution of the motion model changes. Second, separating detection and tracking will make the detector unable to perceive the association between frames, only considering the spatial / appearance features of the target image and ignoring the temporal causal relationship of the moving target. Therefore, a vehicle target tracking method that can relax the constraint on the independent and identically distributed assumption and at the same time consider the spatio-temporal features of the target is needed to overcome the tracking failure caused by occlusion in target tracking.
[0005] The AOG (And-Or Graph), first proposed by Pearl in 1984, is a hierarchical symbolic graph structure representation, consisting of And-nodes, Or-nodes, Terminal-nodes, and the edges connecting these nodes. If all child-node events are completed before the parent node can be completed, the parent node is an And-node; if any child-node event completion allows the parent node to be completed, the parent node is an Or-node; multi-level relationships are represented by vertical arrows from top to bottom; different logical subgraphs can be obtained by selecting different child nodes at Or-nodes. The steps for prediction using AOG are to first generate multiple logical subgraphs based on the jump probabilities between nodes within the range of interest, then select the optimal subgraph using an optimization method or evaluation criterion, and finally use this subgraph for inference and prediction. In 2018, Xu et al. proposed the causal and-or graph (C-AOG) object tracking method, where And-nodes represent different perspectives of a vehicle and Or-nodes represent different components of the vehicle. The transition probabilities on the Or-nodes in the graph are known a priori, and the subgraph with the maximum posterior probability is taken as the optimal subgraph. This method can learn the causal relationship between nodes from high-level semantics to low-level features, effectively improving the tracking accuracy. However, due to the lack of expression and calculation of deep appearance features, when the similarity between targets is high, the posterior probability value of non-targets may be greater than that of the true target, resulting in judgment failure.
[0006] To overcome the defect that the target temporal relationship is ignored, the multi-object tracker TransMOT proposed by Chu in 2023 combines the CNN detection model with the attention mechanism calculation structure of Transformer. By transposing the time dimension of the feature matrix of the spatial graph encoding layer to the first dimension and then inputting it into the standard encoder layer, the calculation of the attention scoring function for temporal information is achieved; a sparse weighted graph is constructed using the spatial relationship between targets, and then a graph convolutional model is used for information aggregation. This method can effectively handle multi-object tracking problems, but has the following defects: (1) Only by changing the matrix arrangement to solve the temporal causal problem cannot fully express the causal relationship of mutual constraints between targets in the time series and cannot learn long-term time dependencies; (2) In the spatial graph encoding layer, the target positions are input into the graph convolutional network (GCN, Graph Convolutional Network, abbreviated as GCN) and two other linear layers simultaneously, and handling the relationship of target positions in the same frame makes the ability of GCN to extract features from the data structure not fully utilized.
[0007] In summary, in existing object tracking, there is a problem that tracking errors are likely to occur when the tracked target is occluded. Summary of the Invention
[0008] The object of the present invention is to propose an object tracking method based on causal relationship and attention mechanism to overcome the tracking failure when the vehicle object is occluded, aiming at the deficiencies of the above-mentioned existing AOG-based object tracking method and TransMOT method.
[0009] Design principle: The inherent causal logical relationship between things will not change due to the change of the input of the external environment and does not depend on the independent and identically distributed hypothesis. Therefore, the present invention introduces causal logical relationship to reduce the dependence on hypothesis conditions. The causal relationship can be represented by a directed acyclic graph. Considering the application scenario of road vehicle object tracking, there is a natural hierarchical relationship between the driving task and the target state. The And-Or Graph (AOG) is selected to represent the hierarchical logical relationship from high-level semantics to low-level features. The overall design scheme of this application is as follows.
[0010] An object tracking method, the object tracking method is based on a causal and-or logic graph and a spatio-temporal attention mechanism, and the object tracking method includes:
[0011] S1. Use an object feature extraction model to perform object detection and extract its feature vector.
[0012] S2. Calculate the fitting curve parameters of each trajectory.
[0013] S3. Calculate the relative motion features of each trajectory including the state of T frames.
[0014] S4. Construct an and-or logic graph according to the trajectory curve parameters and relative motion features.
[0015] S5. Construct an object feature vector and a graph adjacency matrix, and input them into the graph convolutional layer of the spatial graph encoding layer.
[0016] S6. Input the obtained trajectory features into the linear layer of the spatial graph encoding layer after position encoding.
[0017] S7. The output of the spatial graph encoding layer is used as the input of the temporal encoding layer, and the output of the temporal encoding layer is the output of the encoder module.
[0018] S8. Input the features of the detected object in the current frame and the corresponding adjacency matrix into the decoder.
[0019] S9. Use the output of the decoder as the input of the fully connected prediction layer. The fully connected prediction layer outputs the object tracking results: the trajectory attributes of the object, the object position box, and feeds back the object tracking results to the inputs of S2 and S3 to update the trajectory information at the input end.
[0020] Further, the object feature extraction model in step S1 is constructed based on the YOLOv10-x network, and includes a trajectory feature extraction sub-network and a color feature extraction sub-network, which are used for object detection and feature vector extraction of images.
[0021] Further, step S1 includes: S11. Obtain the target detection result: Input the image into the trained YOLOv10-x network, and the network outputs the coordinate values of the upper left and lower right corners of the bounding box of each detected target, the attribute confidence score, and the attribute classification result; S12. Calculate the color features in the HSV space: According to the target bounding boxes obtained in S11, convert all the pixels within each target bounding box from the RGB color space to the HSV space, and calculate the differences Δh and Δs between the mean and median of the H and S channels within the target box respectively.
[0022] Among them, HSV (Hue, Saturation, Value) is a color space created based on the intuitive characteristics of colors, also known as the hexagonal pyramid model.
[0023] Further, in step S2, calculate the fitting curve parameters of each trajectory according to the tracking result of the previous moment, including: S21. Input the position matrix of the trajectory in the T-frame time period, and perform cubic polynomial fitting using the least squares method; S22. Output the three-dimensional parameter vector after removing the constant term from the parameters calculated in S21.
[0024] Further, in step S3, calculate the relative motion characteristics of each trajectory including the T-frame state according to the tracking result of the previous moment, including: S31. Calculate the relative position including the center position and image coordinates of the subsequent frame of the trajectory with respect to the previous frame; S32. Output the relative position vector.
[0025] Further, step S4 includes: S41. According to the curve parameters obtained in S2, when the first two dimensions of the vector are not 0, it is judged as turning, when the first two dimensions are both 0, it is judged as straight driving, and when all three dimensions are 0, it is judged as stopping; S42. After the judgment result in S41, judge the vehicle driving scenario according to the relative motion characteristics output in S3: Calculate the relative position between the last two frames and the relative angle between the previous two frames. According to the angle in different quadrants and it being a turning type, judge whether it is a left turn, a right turn, a right turn and reverse, or a left turn and reverse; for the straight driving type, calculate the relative displacement. If the displacement is positive, it is straight forward, otherwise it is straight backward; S43. Construct the AND-OR logic graph AOG.
[0026] Further, step S5 includes: S51. Concatenate the target features, relative motion features, curve parameters, and the numerical values of the AND-OR logic graph adjacency matrix in sequence to generate the state vector of each target, forming a state matrix containing all targets; S52. Use the trajectory as a node, the sum of the AND-OR logic graphs AOG of the trajectory in the previous few moments as the time causal weight, the intersection over union between different targets at the same moment as the spatial weight, and the element value of the total target adjacency matrix is the sum of the time causal weight and the spatial weight.
[0027] S53. Input the state matrix of the target and the total adjacency matrix into the graph convolutional neural network GCN of the spatial graph encoding layer. The output of the graph convolutional neural network GCN serves as the Key vector of the standard Transformer model.
[0028] Further, step S6 includes: S61. Perform standard position encoding on the position vector, and then add it to the original position vector to generate an encoded vector; S62. Input the encoded vector into the linear layers of the spatial graph encoding layer as the Query vector and the Value vector in the standard Transformer model respectively. The Query vector is the query vector Q in the attention mechanism, and the Value vector is the value vector V in the self-attention mechanism. Additionally, in the Transformer model, there is also a Key vector, which is the key vector K in the self-attention mechanism, corresponding to three matrices: the query matrix W Q , the value matrix W V and the key matrix W K .
[0029] Further, step S7 includes: S71. Transpose the first two dimensions of the output of the spatial graph encoding layer and input it into the time encoding layer; S72. The structure of the time encoding layer is the same as that of the standard Transformer encoder structure.
[0030] Further, the structure of the spatial graph decoding layer module is the same as that of the spatial graph encoding layer module. In step S8, input the features of the detected target in the current frame and the corresponding adjacency matrix into the spatial decoder, output the tracking result of the target in the current frame, and feedback the tracking result back to the input end to update the trajectory information at the input end. Specifically, it includes: 81. The input of the graph convolutional layer is the graph structure of the target detection result in the current frame: using the target as a node, the position coordinates of the target box as the feature vector, and the IoU (intersection over union) between target boxes as the corresponding element values of the adjacency matrix; S82. The output of the spatial decoder serves as the Query vector, and calculates cross-attention with the Key vector and the Value vector output by the time encoding layer; S83. Output after passing through the softmax layer. The softmax function is used to normalize the attention weights, that is, it is a normalized exponential function.
[0031] Further, step S9 includes: Input the output of step S83 into the prediction layer; the prediction layer contains two two-layer fully connected layers, which respectively output the trajectory attributes of the target and the coordinates of the target position box; feedback the trajectory attributes and the coordinates of the target position box back to the input of the spatial graph encoding layer to calculate the trajectory-related features.
[0032] Compared with the prior art, the beneficial effects of the present invention are as follows: In view of the problem that tracking errors are likely to occur when the target is occluded in vehicle target tracking in a road environment, a target tracking method based on a causal AND-OR logic graph and a spatio-temporal attention mechanism is designed. The causal logical relationship between the vehicle driving task and the target state is aggregated into the spatial information of the trajectory through a graph convolutional network, and the attention mechanism of the Transformer architecture is used to fuse the causal logical relationship, spatial information, and temporal information, realizing the function of guiding state prediction with high-level semantics. Compared with the traditional method of determining whether the target is occluded by the number of unupdated frames, the present invention more flexibly processes the tracking prediction problem of occluded targets, reduces the number of incorrect target tracking matches, and improves the tracking accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 FIG. is a schematic diagram of the overall process of the target tracking method based on causal relationship and attention mechanism of the present invention;
[0034] Figure 2 FIG. is a causal AND-OR logic graph of vehicle driving;
[0035] Figure 3 FIG. is a structural diagram of a spatial encoder. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0037] Target tracking method
[0038] A target tracking method, see Figures 1 - 3 , the method is based on a causal relationship (AND-OR logic graph) and a spatio-temporal attention mechanism, and is suitable for tracking road rigid targets, such as formal vehicles. The technical idea of the present invention is to calculate the trajectory fitting parameters and motion features according to the target features in the video frames extracted from the network by feature extraction, and construct an AND-OR logic graph of the driving task. The trajectory, target features, and the adjacency matrix of the AND-OR graph are input into the graph convolutional network of the spatial graph encoding layer, and the target features are input into the linear layer of the spatial graph encoding layer. The output is used as the input of the time encoding layer. Finally, the decoder generates the target tracking result of the current frame and updates the trajectory information. The target tracking method includes the following steps, and the following steps do not necessarily constitute a sequential order.
[0039] Step S1: Use the target feature extraction model to perform target detection and extract its feature vectors.
[0040] Among them, the target feature extraction model is constructed based on the YOLOv10-x network, including a trajectory feature extraction sub-network and a color feature extraction sub-network, which are used to perform target detection and extract feature vectors from images.
[0041] S11: Obtain the target detection result: Input the image into the trained YOLOv10-x network, and the network output is the upper left and lower right coordinate values of the bounding box of each detected target, the attribute confidence score, and the attribute classification result; specific examples are as follows.
[0042] Input a frame of image I in the video t into the trained YOLOv10-x network, and output the detection result matrix of N0 detected targets. The matrix dimension is 300×6. 300 refers to the number of predicted bounding boxes output, and 6 is the dimension. The 6 dimensions are respectively: for the i-th target the attribute confidence score and the attribute classification result are the upper left and lower right coordinates of the bounding box Output the detection of N0 targets. Among them, score represents the score value, cls represents the classification value, t represents the t-th moment, and i represents the i-th target (i is a positive integer).
[0043] S12: Calculate the color features in the HSV space: According to the target bounding boxes obtained in S11, convert all the pixels within each target bounding box from the RGB color space to the HSV space, and calculate the differences Δh and Δs between the mean and median of the H and S channels within the target box respectively.
[0044] HSV (Hue, Saturation, Value) is a color space created based on the intuitive characteristics of colors, also known as the hexagonal pyramid model. In the HSV space, H is the least sensitive to light, and V is very sensitive. Therefore, the V channel is discarded, and only the feature values of the H and S channels are calculated. Convert all the pixels within each bounding box from the RGB color space to the HSV space according to the following formula:
[0045] V = max(R, G, B);
[0046] Δ = V - min(R, G, B);
[0047]
[0048] Calculate the differences Δh and Δs between the mean and median of the H and S channels within the target box respectively. Among them, V, H, and S respectively represent the feature values of the three channels in the HSV space, and R, G, and B respectively represent the feature values of the three channels in the RGB space.
[0049] Step S2: Calculate the fitting curve parameters of each trajectory. Specifically, calculate the fitting curve parameters of each trajectory according to the tracking results of the previous moment, including: S21: Input the position matrix of the trajectory in the T-frame time period, and perform cubic polynomial fitting using the least squares method; S22: Output the three-dimensional parameter vector after removing the constant term from the parameters calculated in S21.
[0050] In a specific example, the calculation process of the fitting curve parameters is as follows.
[0051] (2.1) If the target in the t-th frame belongs to the k-th trajectory then its position coordinates at the t-th frame are the upper left and lower right coordinates of the bounding box output in step S11 Calculate the geometric center coordinates of the target as:
[0052] (2.2) The cubic matrix of the X-axis coordinate of the center position of the trajectory Tra k in the T-frame time period is The Y-axis coordinate matrix is where k represents the k-th trajectory and is a positive integer.
[0053] (2.3) Fit the trajectory curve with a cubic polynomial and calculate the coefficients of each term using the least squares method. Solve the parameter vector X in the equation AX = T according to the formula X = (A T A) -1 A T T. Among them, are the coefficients of the cubic term, quadratic form, and linear term respectively.
[0054] (2.4) Finally, output the first three dimensions of the parameter vector X
[0055] Step S3: Calculate the relative motion characteristics of each trajectory including the T-frame state. Specifically, calculate the relative motion characteristics of each trajectory including the T-frame state according to the tracking results of the previous moment, including the following steps.
[0056] S31: Calculate the relative position of the subsequent frame of the trajectory to the previous frame, including the center position and image coordinates; a specific example is: calculate the relative position of the k-th trajectory at the t-th frame to the previous frame:
[0057] S32: Output the relative position vector:
[0058] Step S4: Construct an AND / OR logic diagram based on the trajectory curve parameters and relative motion characteristics;
[0059] S41. Based on the curve parameters obtained in S2, when the first two dimensions of the vector are not zero, it is determined as a turn; when both of the first two dimensions are zero, it is determined as straight driving; when all three dimensions are zero, it is a stop. In a specific example, the curve parameter vector is Determine the target driving category according to the element values in this vector:
[0060] (a) When the driving category is a turn;
[0061] (b) When the driving category is straight driving;
[0062] (c) When the driving category is a stop.
[0063] S42. After the judgment result of S41, judge the vehicle driving scenario according to the relative motion characteristics output by S3: Calculate the relative position between the last two frames and the relative angle between the first two frames. When the angle is in different quadrants and it is a turning category, judge whether it is a left turn, a right turn, a right turn and reverse, or a left turn and reverse; for the straight driving category, calculate the relative displacement. When the displacement is positive, it is straight forward, otherwise it is straight backward. In a specific example: Based on the driving category, judge the vehicle driving scenario according to the relative motion characteristics:
[0064] (a) Calculate the relative displacement, cosine of the angle, and relative motion respectively:
[0065] The relative displacement is:
[0066] The cosine of the angle is:
[0067] The relative motion is:
[0068] (b) Solve the change angle θ of the motion direction between the end point of the trajectory segment and the starting point:
[0069]
[0070] (c) Judge the driving scenario according to the driving category, relative displacement, and change angle of the motion direction according to Table 1:
[0071] Table 1 Driving scenario judgment
[0072]
[0073]
[0074] S43. Construct an And-Or Graph (AOG). The logical relationships in the driving task are divided into three layers from top to bottom: driving scenarios, trajectory behaviors, and target states. The nodes in these logical relationships are represented as AOG nodes. The AOG structure is as follows Figure 2 shown. The top-level root node is the driving task of the vehicle itself. Each driving task contains multiple different driving scenarios. The driving scenarios have an "or" relationship with the driving task node. Therefore, the driving task is an or-node. Each driving scenario is also an or-node that contains different trajectory behaviors. Trajectory behaviors refer to the trajectory curve parameters and the relative motion characteristics of the trajectory. The trajectory behavior node is composed of the bottom-level target states. The chronological order of the target states indicates an "and" relationship between the states. Therefore, the trajectory behavior node is an and-node.
[0075] Step S5. Construct the target feature vector and the graph adjacency matrix, and input them into the graph convolutional layer of the input space graph encoding layer.
[0076] S51. Concatenate the target features, relative motion features, curve parameters, and the numerical values of the AOG adjacency matrix of the And-Or Logic Graph in sequence to generate the state vector of each target, forming a state matrix that includes all targets. In a specific example, the or-nodes in the And-Or Graph represent the classification labels of sub-layers. The attribute label value I k is the classification label of the driving scenario corresponding to the trajectory; concatenate the target features, relative motion features, curve parameters, and the attribute label to obtain the feature vector of the target trajectory Tra k at the t-th frame:
[0077] S52. Use the trajectory as a node, the sum of the AOGs of the trajectory in the previous few moments as the time causal weight, and the intersection-over-union ratio between different targets at the same moment as the spatial weight. The element value of the total target adjacency matrix is the sum of the time causal weight and the spatial weight. Specifically, establish a graph structure G(δ, E) with the target trajectory points at the t-th frame as nodes. The node set The boundary matrix is E, which represents the relationship between the trajectory points, and establish the adjacency matrix A t of the graph:
[0078] (a) The sum of the AOGs of the trajectory in the previous few moments is used as the time causal weight matrix A Tra ;
[0079] (b) The intersection-over-union ratio between different targets in the same frame is used as the position weight matrix A o of the adjacency matrix in the same frame.
[0080] The graph adjacency matrix A is the sum of the time causal matrix and the position weight matrix: A t = A Tra + A o .
[0081] S53. Input the state matrix of the target and the total adjacency matrix into the graph convolutional neural network GCN (Graph Convolutional Network) of the spatial graph encoding layer. The graph convolutional neural network GCN outputs the Key vector for the standard Transformer model. The movement of the target is affected not only by its own state at the previous moment but also by the states of other targets. Each node in the two-layer GCN can collect information from its neighbor nodes and non-neighbor nodes, so a two-layer GCN network is selected accordingly. The feature matrix containing N g trajectory Input the target feature matrix Δ t and the adjacency matrix A t into the two-layer GCN network. Among them, σ represents the feature vector of the trajectory at time t.
[0082] Step S6. Input the obtained trajectory features into the linear projection layer of the spatial graph encoding layer after position encoding.
[0083] S61. Perform standard position encoding on the position vector:
[0084] where is the coordinate position, k is the trajectory label, and d model is the total number of trajectories.
[0085] Then add it to the original position vector to generate the encoded vector for each target After arrangement, an N g ×4-dimensional encoded matrix
[0086] S62. Input the encoded vector into the linear layers of the spatial graph encoding layer as the Query vector and the Value vector in the standard Transformer model respectively. The Query vector is the query vector Q in the self-attention mechanism, and the Value vector is the value vector V in the self-attention mechanism.
[0087] First are input into two linear projection layers respectively, and Q t and V t are output respectively; then H t 、Q t 、V t are used as the Key vector, the Query vector, and the Value vector respectively to calculate the dot product attention value: Then, input the dot product attention value and the trajectory feature vector into the residual connection and normalization layer, and its output passes through the feed-forward layer and another residual connection and normalization to obtain the output of the spatial graph encoding layer
[0088] Step S7: The output of the spatial graph encoding layer is used as the input of the temporal encoding layer, and the output of the temporal encoding layer is the output of the encoder module. It includes the following steps.
[0089] S71: The first two dimensions of the output of the spatial graph encoding layer are transposed and then input into the temporal encoding layer; S72: The structure of the temporal encoding layer is the same as that of the standard Transformer encoder structure. A specific example is as follows.
[0090] Output of the spatial graph encoding layer Indicates that the T frames representing the time series are in the second dimension. After transposing its first two dimensions, we get Placing the T frames representing the time series in the first dimension.
[0091] Then Input into the standard Transformer encoder layer, whose structure is as Figure 3 shown.
[0092] Finally, the output of the temporal encoding layer
[0093] Step S8: The features of the detection target of the current frame and the corresponding adjacency matrix are input into the decoder. The spatial decoder module has the same structure as the spatial graph encoding layer module. In step S8, the features of the detection target of the current frame and the corresponding adjacency matrix are input into the spatial decoder, and the tracking result of the target of the current frame is output, and the tracking result is fed back to the input end to update the trajectory information at the input end. It specifically includes the following steps.
[0094] S81: The input of the graph convolutional layer is the graph structure G(O, E) of the detection result of the target of the current frame: the targets are used as nodes, the position coordinates of the target boxes are used as feature vectors, and the IoU between the target boxes is used as the corresponding element value of the adjacency matrix; Specific example: Using the target as a node, the feature matrix composed of the position coordinates of the target box as the feature vector is Adjacency matrix The corresponding element value is the IoU between the target boxes; The input of the linear layer is m t . Among them, the uppercase bold O represents the node set, and the lowercase bold o in the set is all the target feature vectors at time t.
[0095] S82: The output of the spatial decoder is used as the Query vector, and the output by the temporal encoding layer is used as the Key vector and the Value vector to calculate the cross attention:
[0096] Where dtemp represents the temporary value of the decoder at the current moment, out represents the output; in ten, t represents time, and en represents the encoder. yes The dimension of .
[0097] S83, output through softmax layer
[0098] Step S9: The decoder output is used as the input of the fully connected prediction layer, and the fully connected prediction layer outputs the target tracking result: the trajectory attribute of the target, the target position box, and feeds the target tracking result back to the input terminals S2 and S3 to update the trajectory information of the input terminals. The specific example includes the following steps.
[0099] S91, the output of step S83 Input prediction layer;
[0100] S92, the prediction layer contains two two-layer fully connected layers, which output the trajectory attributes of the target respectively and the target location frame coordinates
[0101] S93, feeding back the trajectory attributes and the target position frame coordinates to the spatial graph encoding layer input to calculate trajectory related features.
[0102] In summary, the present invention is a target tracking method based on causal AND / OR logic graph and spatiotemporal attention mechanism, and the implementation scheme is: first, the target features are obtained by using the feature extraction subnetwork, the trajectory features are calculated according to the target tracking results of the previous moment, and the AND / OR logic graph of the driving task is constructed; the adjacency matrix is initialized according to the distance between the target detected in the current frame and the trajectory, and then the target feature vector and the adjacency matrix are input into the graph convolution layer, and the target feature vector is input into the linear layer after position encoding; the feature vector of the target detected in the current frame and the corresponding adjacency matrix are input into the decoder, and the decoder output is output through the prediction layer to output the tracking result of the target in the current frame: the trajectory attribute of the target and the target position box; finally, the trajectory information is updated and fed back to the encoder input for prediction at the next moment. The present invention introduces the causal logic relationship from the driving task to the target state as a high-level semantics, aggregates the causal logic relationship with the target trajectory information through the graph convolution network, enhances the network's ability to extract the spatiotemporal and causal dependencies of the target movement at different times, and effectively avoids missed detection and false detection when the target is blocked for a long time.
[0103] Target tracking system
[0104] A target tracking system for executing the aforementioned target tracking method, the target tracking system comprising:
[0105] The position encoding adds the obtained trajectory features after position encoding to the original position vector to generate the encoding vector of each target, and inputs it into the linear projection layer of the spatial graph encoding layer, which is used to represent the position relationship between this trajectory and other trajectories in the current frame.
[0106] The spatial decoder has the same structure as the spatial graph encoding layer. Its main components include a graph convolutional network and two linear projection layers, which are used to map the trajectory and target features calculated by the spatio-temporal encoder into the association matrix between the target and the trajectory.
[0107] The prediction layer adopts a double-layer fully connected prediction layer, including a target trajectory attribute prediction layer and a target position box prediction layer, which are used to predict the target position and trajectory attributes.
[0108] The spatio-temporal encoder includes a spatial graph encoding layer and a temporal encoding layer, which are used to extract the features of the trajectory and the target and aggregate them with their corresponding causal logical relationships.
[0109] Through the above methods and systems, the problem that the prior art is prone to tracking errors when the tracked target is occluded can be effectively solved.
[0110] Computer-readable storage medium
[0111] The present invention also provides a computer-readable storage medium, on which computer instructions are stored. When the computer instructions run, they execute the steps of the foregoing method. Among them, for the method, please refer to the detailed introduction in the foregoing part, and details will not be repeated here.
[0112] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium. Computer-readable media include permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0113] The computer program codes required for the operations of various parts of this application can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C, VisualBasic, Fortran2003, Perl, COBOL2002, PHP, ABAP, dynamic programming languages such as Python, Ruby, and Groovy, or other programming languages. The program codes can run entirely on the user's computer, or run on the user's computer as an independent software package, or run partially on the user's computer and partially on a remote computer, or run entirely on a remote computer or processing device. In the latter case, the remote computer can be connected to the user's computer through any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (for example, through the Internet), or in a cloud computing environment, or used as a service such as software as a service (SaaS).
[0114] Terminal
[0115] The present invention also provides a terminal, including a memory and a processor. The memory stores data provider information and computer instructions that can run on the processor. When the processor runs the computer instructions, it executes the steps of the foregoing method. Among them, for the description of the method, please refer to the detailed introduction in the foregoing part, and it will not be elaborated here.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A target tracking method based on causality and attention mechanism, characterized in that: The target tracking method comprises: S1. Target detection and feature vector extraction using target feature extraction model; S2, calculating the fitting curve parameters of each trajectory according to the tracking result at the previous moment; S3, calculating the relative motion characteristics of each track including the T frame state according to the tracking result at the previous moment; S4, constructing an AND / OR logic diagram according to trajectory curve parameters and relative motion characteristics; S5, construct the state matrix and total adjacency matrix of the target, and input them into the graph convolution layer of the spatial graph encoding layer; S6, inputting the acquired trajectory features into the linear layer of the spatial graph encoding layer after position encoding; S7, the output of the spatial image coding layer is used as the input of the temporal coding layer, and the output of the temporal coding layer is the output of the encoder module; S8, input the features of the current frame detection target and the corresponding adjacency matrix into the decoder; S9: The decoder output is used as the input of the fully connected prediction layer. The fully connected prediction layer outputs the target tracking result: the trajectory attributes of the target, the target position box, and feeds the target tracking result back to S2 and S3 as input to update the input trajectory information.
2. The target tracking method according to claim 1, characterized in that: The target feature extraction model in step S1 is built based on the YOLOv10-x network, including a trajectory feature extraction subnetwork and a color feature extraction subnetwork, which are used to perform target detection and feature vector extraction on the image.
3. The target tracking method according to claim 2, characterized in that: Step S1 includes: S11. Obtain target detection results: Input the image into the trained YOLOv10-x network. The network outputs the coordinate values of the upper left corner and lower right corner of the bounding box of each detected target, the attribute confidence score, and the attribute classification result. S12. Calculate the color features of the HSV space: According to the target bounding box obtained in S11, convert all pixels in each target bounding box from the RGB color space to the HSV space, and calculate the difference between the mean and median of the H and S channels in the target box respectively. and .
4. The target tracking method according to claim 1, characterized in that: Step S2 includes: S21, input the position matrix of the trajectory in the T frame time period, and perform cubic polynomial fitting using the least squares method; S22, output a three-dimensional parameter vector after removing the constant term from the parameters calculated in S21.
5. The target tracking method according to claim 1, characterized in that: Step S3 includes: S31, calculating the relative position of the next frame of the trajectory to the previous frame, including the center position and image coordinates; S32. Output relative position vector.
6. The target tracking method according to claim 1, characterized in that: Step S4 includes: S41, according to the curve parameters obtained in S2, when the first two dimensions of the vector are not 0, it is judged as turning, when the first two dimensions are 0, it is judged as straight driving, and when the three dimensions are all 0, it is stopped; S42, after the result of S41 is determined, the vehicle driving scene is determined according to the relative motion characteristics output by S3: the relative position between the last two frames and the relative angle between the first two frames are calculated, and if the angles are in different quadrants and it is a turning type, it is determined to be a left turn, a right turn, a right turn and backward, or a left turn and backward; if it is a straight-line driving type, the relative displacement is calculated, and the displacement is a straight-line forward, otherwise it is a straight-line backward; S43. Construct an AND / OR logic graph AOG.
7. The target tracking method according to claim 1, characterized in that: Step S5 includes: S51, sequentially concatenating the target features, relative motion features, curve parameters, and the values of the AND or logic graph adjacency matrix to generate a state vector for each target, thereby forming a state matrix containing all targets; S52, taking the trajectory as a node, the sum of the AND / OR logic graphs AOG of the trajectory at the previous moments is taken as the temporal causal weight, the intersection and union ratio between different targets at the same moment is taken as the spatial weight, and the element value of the total adjacency matrix of the target is the sum of the temporal causal weight and the spatial weight; S53. Input the state matrix and total adjacency matrix of the target into the graph convolutional neural network GCN of the spatial graph encoding layer, and the output of the graph convolutional neural network GCN is used as the Key vector of the standard Transformer model.
8. The target tracking method according to claim 1, characterized in that: Step S6 includes: S61, performing standard position encoding on the position vector, and then adding the encoded vector to the original position vector to generate an encoded vector; S62. Input the encoded vector into the linear layer of the spatial graph encoding layer as the Query vector and Value vector in the standard Transformer model respectively.
9. The target tracking method according to claim 1, characterized in that: Step S7 includes: S71, the first two dimensions of the output of the spatial image coding layer are transposed and input into the temporal coding layer; S72, the structure of the temporal coding layer is the same as the standard Transformer encoder structure.
10. The target tracking method according to claim 1, characterized in that: The spatial decoder module has the same structure as the spatial graph coding layer module. In step S8, the features of the detected target in the current frame and the corresponding adjacency matrix are input into the spatial decoder, the tracking result of the target in the current frame is output, and the tracking result is fed back to the input end to update the trajectory information of the input end, which specifically includes: S81, the input of the graph convolution layer is the graph structure of the target detection result of the current frame: the target is the node, the position coordinates of the target frame are the feature vector, and the IoU between the target frames is the corresponding element value of the adjacency matrix; S82, the output of the spatial decoder is used as the Query vector, and the cross attention is calculated with the Key vector and Value vector output by the temporal encoding layer; S83, output through the softmax layer.
11. The target tracking method according to claim 10, characterized in that: Step S9 includes: The output of step S83 is input into the prediction layer; The prediction layer consists of two two-layer fully connected layers, which output the target trajectory attributes and the target position box coordinates respectively; The trajectory attributes and target location box coordinates are fed back to the spatial graph encoding layer input to calculate trajectory related features.
Citation Information
Patent Citations
User intention prediction method and device using man-machine object space-time interaction relationship
CN112580550A
Track prediction method and device based on time attention convolutional network
CN114116944A