Dynamic multi-target tracking and trajectory prediction method
Through dynamic threshold dual-stream neural network and self-motion decoupling mechanism, the accuracy of multi-object tracking and trajectory prediction in dynamic environments is solved, efficient and stable object detection and prediction of robots in complex environments is achieved, and obstacle avoidance capabilities of sweeping robots are improved.
Patent Information
- Application Number
- CN202510431623.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The prior art lacks the ability to track and track prediction of multiple targets in dynamic environments, especially in the case of robots' own movements, it is difficult to achieve high-precision target detection and stable tracking.
A dynamic threshold dual-stream neural network is used to combine self-motion decoupling mechanism and a two-layer decision-making mechanism. By processing color images and depth images in parallel, complementary features are extracted, fusion weights are adaptively adjusted, trajectory prediction is performed by combining environmental constraints and social behaviors, and time-varying impact maps are generated for planning and decision-making.
Improve the accuracy and completeness of target detection, ensure robustness and high-precision target tracking in complex environments, accurately predict target behavior and perceive potential collision risks in advance, and achieve harmonious interaction.
Smart Images

Figure CN120355754A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-target tracking, and specifically provides a dynamic multi-target tracking and trajectory prediction method. Background Art
[0002] In the field of image analysis, multi-target tracking technology represents and identifies targets by extracting visual information such as texture features, edge features, and local features in images. Current mainstream technologies use feature extraction methods such as histogram of oriented gradients, scale-invariant feature transform, and bag of visual words to obtain the appearance information of targets from image sequences, and achieve inter-frame association of targets through feature matching. During target recognition, comparative analysis based on local feature descriptors can effectively distinguish different targets and achieve accurate target classification. In dynamic scene analysis, spatio-temporal features of targets are extracted through image processing technology, and a motion model of targets is constructed using visual feature matching algorithms to simultaneously track multiple targets. Target trajectory prediction is based on the visual motion features extracted from the image sequence, analyzes the historical motion patterns of targets, and predicts their future positions. In addition, visual elements such as texture information and edge structures in the scene are also used as context clues to assist in improving target recognition and tracking performance. With the development of image analysis technology, target representation methods based on deep feature learning can automatically extract hierarchical visual features from image data, further improving the accuracy of target recognition and tracking.
[0003] The integrated application of image analysis and processing technologies provides strong support for realizing dynamic multi-target tracking and trajectory prediction in complex scenarios. However, existing technologies still lack the ability to actively track and predict dynamic objects in their own motion state. Therefore, a dynamic multi-target tracking and trajectory prediction method is proposed. Summary of the Invention
[0004] The object of the present invention is to provide a dynamic multi-target tracking and trajectory prediction method, including: collecting color images and depth images and performing preprocessing; constructing a dynamic threshold two-stream neural network to extract features from the preprocessed color images and depth images to obtain a visual feature set and a depth feature set, and fusing them through an adaptive attention mechanism to obtain a fused feature set, and combining with a dynamic threshold to obtain a list of observation information of the targets; updating the list of observation information through a self-motion decoupling mechanism to obtain a list of real information; establishing spatial and temporal associations between targets and trajectories based on the list of real information to obtain the latest trajectory set; analyzing target characteristics, environmental constraints, social behaviors, and historical trajectories, generating prediction trajectories for each target and corresponding multi-factor scores, generating a time-varying influence map, and adopting a two-layer decision-making mechanism for planning and decision-making.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A dynamic multi - target tracking and trajectory prediction method, comprising:
[0007] Collect color images and depth images and perform pre - processing;
[0008] Construct a dynamic - threshold two - stream neural network to extract features from the pre - processed color images and depth images, obtain a visual feature set and a depth feature set, and fuse the visual feature set and the depth feature set through an adaptive attention mechanism to obtain a fused feature set, and combine with a dynamic threshold to obtain a list of observation information of the target;
[0009] Obtain the self - motion state vector, update the list of observation information through a self - motion decoupling mechanism to obtain a list of real information;
[0010] Establish spatial and temporal associations between the target and the trajectory based on the list of real information to obtain the latest trajectory set;
[0011] Analyze target characteristics, environmental constraints, social behaviors, and historical trajectories to generate a predicted trajectory for each target and the corresponding multi - factor scores;
[0012] Generate a time - varying influence diagram according to the predicted trajectory and the multi - factor scores, and adopt a two - layer decision - making mechanism for planning and decision - making.
[0013] Furthermore, the dynamic - threshold two - stream neural network includes:
[0014] A feature extraction layer, including: the visual feature extraction branch adopts a feature pyramid network structure to process the pre - processed color image to obtain the visual feature set; the depth feature extraction branch adopts a three - dimensional convolutional neural network to process the pre - processed depth image to obtain the depth feature set;
[0015] An adaptive attention layer, including: a lightweight evaluation module analyzes the quality scores of visual features and depth features; an attention fusion module adjusts the fusion weights according to the quality scores and fuses the visual feature set and the depth feature set to obtain the fused feature set;
[0016] A detection and segmentation layer, including: the detection head branch outputs initial detection information based on the fused feature set; the segmentation head branch outputs initial segmentation information based on the fused feature set;
[0017] A post - processing output layer, including: a dynamic - threshold adjustment module calculates a dynamic detection threshold and a dynamic segmentation threshold according to the environmental illumination conditions and the target motion speed; an output module screens the initial detection information and the initial segmentation information in combination with the dynamic detection threshold and the dynamic segmentation threshold to obtain the list of observation information.
[0018] Further, the self-motion decoupling mechanism specifically includes: fusing the data of the inertial measurement unit by using an extended Kalman filter to estimate the self-motion state vector of the robot; calculating a coordinate system transformation matrix according to the self-motion state vector; extracting the target position based on the observation information of each target and transforming it to the world coordinate system through the coordinate system transformation matrix; obtaining the true motion vector of the target by combining the world coordinate position in the previous frame; and updating the observation information list based on the true motion vector to obtain the true information list.
[0019] Further, establishing the spatial and temporal associations between the target and the trajectory based on the true information list, and obtaining the latest trajectory set includes:
[0020] Extracting the depth appearance features for each target detected in the current frame and each existing trajectory;
[0021] Calculating an appearance similarity matrix based on the depth appearance features, calculating a spatial position similarity matrix based on the world coordinate positions, and calculating a velocity consistency similarity matrix based on the estimated velocity of the target and the predicted velocity of the trajectory;
[0022] Weightedly fusing the appearance similarity matrix, the spatial position similarity matrix, and the velocity consistency similarity matrix to obtain a comprehensive similarity matrix;
[0023] Using the Hungarian algorithm to solve the optimal matching relationship for the comprehensive similarity matrix, and according to the optimal matching relationship, creating new trajectories, updating trajectories, and terminating to obtain the latest trajectory set, specifically including: for the targets that are not matched to trajectories, if the confidence of the new trajectory is greater than the threshold, create new trajectories; for the trajectories that are matched to targets, update the trajectory state through a Kalman filter; for the trajectories that are not matched to targets, if the number of consecutive unmatched frames is less than the unmatched frame number threshold, use a long short-term memory network to predict the trajectory state, otherwise terminate the tracking of this trajectory.
[0024] Furthermore, analyze the target characteristics, environmental constraints, social behaviors, and historical trajectories to generate the predicted trajectories of each target and the corresponding multi-factor scores, including: select the corresponding motion model according to the target category to calculate the estimated motion parameters; use the depth information to construct an indoor environment grid map, calculate the static obstacle distance field, and define the environmental cost function, and calculate the environmental feasibility score according to the environmental cost function; calculate the total social influence of each target by all other targets; use the gated recurrent unit network to obtain the historical trajectory encoding of the target; based on the estimated motion parameters of the target, the indoor environment grid map, the total social influence, and the historical trajectory encoding, generate multiple initial predicted trajectories through a conditional variational autoencoder, and combine the environmental feasibility score, the total social influence, and the acceleration of the predicted trajectory to obtain the multi-factor score of each initial predicted trajectory; sort the initial predicted trajectories from largest to smallest according to the multi-factor scores, and select the top m trajectories as the predicted trajectories.
[0025] Furthermore, generate a time-varying influence map according to the predicted trajectory and the multi-factor score, including: initialize the static influence map according to the indoor environment grid map; calculate the influence contribution of each predicted trajectory of the target to each grid in the static influence map; summarize the influence contributions of all targets and trajectories to obtain the dynamic influence map; combine the static influence map and the dynamic influence map to obtain the time-varying influence map.
[0026] Furthermore, adopt a two-layer decision-making mechanism for planning and decision-making, including: establish a three-dimensional spatio-temporal search space based on the time-varying influence map, and design an evaluation function by comprehensively considering the time cost, space cost, and influence cost; set the global planning period, and perform heuristic search according to the evaluation function to obtain the global planning path; set the local adjustment period, sample the feasible speed space to obtain speed samples; calculate the comprehensive score of each speed sample, and select the speed sample with the highest comprehensive score to generate the local execution trajectory.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] 1. The dynamic threshold dual-stream neural network structure achieves more accurate object detection and segmentation by processing color images and depth images in parallel. The feature extraction layer contains a visual feature extraction branch and a depth feature extraction branch, which extract complementary features from different sensor data respectively; the adaptive attention layer dynamically analyzes the quality of different features through a lightweight evaluation module, and adaptively adjusts the fusion weights according to environmental conditions, effectively solving the problem of insufficient reliability of single-modal data under different lighting and scene conditions; the dynamic threshold adjustment module calculates the dynamic detection threshold and the dynamic segmentation threshold according to the environmental lighting conditions and the target movement speed, improving the robustness of the system in complex environments. This structure significantly improves the accuracy and integrity of object detection, can maintain high-quality object perception ability in various lighting conditions and complex movement environments, provides a reliable observation basis for subsequent dynamic object tracking and prediction, and realizes high-precision object detection in dynamic environments.
[0029] 2. The self-motion decoupling mechanism and the multi-object spatio-temporal association method solve the problem of accurately estimating the true motion of objects under the robot's own motion state. The self-motion decoupling mechanism estimates the robot's motion state by fusing IMU data through an extended Kalman filter, calculates the coordinate system transformation matrix, converts the target observation position to the world coordinate system, and separates the true motion vector of the target by combining historical information, which can effectively eliminate the interference of the robot's motion on the target position observation and ensure accurate estimation of the target motion when the robot is moving. The multi-object spatio-temporal association mechanism constructs a comprehensive similarity matrix based on depth appearance features, spatial position similarity, and speed consistency, and solves the optimal matching relationship through the Hungarian algorithm, which can accurately process and manage objects and trajectories, ensuring that the robot can stably track multiple dynamic objects while moving itself, providing continuous and consistent object trajectory information, and laying a solid foundation for accurately predicting the future behavior of objects.
[0030] 3. In terms of trajectory prediction, considering environmental constraints, social behaviors, and historical trajectories comprehensively, multiple possible future trajectories are generated through a conditional variational autoencoder, and the multi-factor scores of each trajectory are calculated. This prediction method not only considers the physical motion laws but also incorporates social behavior patterns and environmental interaction factors, and can more accurately predict the complex behavior patterns of different objects such as humans and pets. In terms of decision-making, a time-varying influence graph is constructed based on the predicted trajectories and multi-factor scores, the environmental space is represented as an influence distribution that changes over time, and a two-layer decision-making mechanism is adopted for path planning, enabling the sweeping robot to perceive potential collision risks in advance, plan the optimal obstacle avoidance path, and dynamically adjust the cleaning strategy according to the actual behavior of the target, being able to intelligently anticipate the behavior of dynamic objects in the environment and react in advance, fundamentally solving the lag problem of traditional passive obstacle avoidance strategies and realizing harmonious interaction with the dynamic environment. Description of the Drawings
[0031] Figure 1 Schematic diagram of a dynamic multi - target tracking and trajectory prediction method of the present invention;
[0032] Figure 2 Schematic diagram of the hierarchical structure of the dynamic - threshold two - stream neural network of the present invention;
[0033] Figure 3 Schematic diagram of the detailed structure of the dynamic - threshold two - stream neural network of the present invention. Specific embodiments
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.
[0035] Please refer to Figures 1 to 3 , the present invention provides a dynamic multi - target tracking and trajectory prediction method, and the technical solution is as follows:
[0036] Embodiment 1:
[0037] A dynamic multi - target tracking and trajectory prediction method, as Figure 1 shown, includes:
[0038] Collect color images and depth images and perform pre - processing;
[0039] Construct a dynamic - threshold two - stream neural network, extract features from the pre - processed color images and depth images to obtain a visual feature set and a depth feature set, and fuse the visual feature set and the depth feature set through an adaptive attention mechanism to obtain a fused feature set, and combine with a dynamic threshold to obtain a list of observation information of the target;
[0040] Obtain the self - motion state vector, and update the list of observation information through a self - motion decoupling mechanism to obtain a list of real information;
[0041] Based on the list of real information, establish spatial and temporal associations between the target and the trajectory to obtain the latest trajectory set;
[0042] Analyze target characteristics, environmental constraints, social behaviors, and historical trajectories to generate prediction trajectories and corresponding multi - factor scores for each target;
[0043] Generate a time - varying influence diagram according to the prediction trajectories and the multi - factor scores, and adopt a two - layer decision - making mechanism for planning and decision - making.
[0044] Furthermore, as Figure 2 andFigure 3 As shown in Figure 3 , the dynamic threshold two-stream neural network includes:
[0045] A feature extraction layer, including: The visual feature extraction branch adopts a feature pyramid network structure to process the preprocessed color image to obtain the visual feature set; The depth feature extraction branch adopts a three-dimensional convolutional neural network to process the preprocessed depth image to obtain the depth feature set;
[0046] An adaptive attention layer, including: A lightweight evaluation module analyzes the quality scores of visual features and depth features; An attention fusion module adjusts the fusion weights according to the quality scores and fuses the visual feature set and the depth feature set to obtain the fused feature set;
[0047] A detection and segmentation layer, including: The detection head branch outputs initial detection information based on the fused feature set; The segmentation head branch outputs initial segmentation information based on the fused feature set;
[0048] A post-processing output layer, including: A dynamic threshold adjustment module calculates the dynamic detection threshold and the dynamic segmentation threshold according to the environmental light condition and the target movement speed; An output module combines the dynamic detection threshold and the dynamic segmentation threshold to screen the initial detection information and the initial segmentation information to obtain the observation information list.
[0049] The dynamic threshold two-stream neural network adopts a parallel structure of visual and depth feature extraction branches, which can simultaneously extract the appearance information of RGB images and the three-dimensional structure information of depth images, avoiding the limitations of single features. The adaptive attention mechanism can dynamically adjust the weights of visual and depth features to ensure stable detection performance in complex environments. The dynamic threshold adjustment can adaptively adjust the detection and segmentation thresholds according to the environmental light and the target movement speed, effectively solving problems such as missed detection under low light conditions and blurred fast-moving targets in traditional fixed thresholds, and improving the robustness and reliability in complex movement environments.
[0050] In this embodiment, the visual feature extraction branch first uses ResNet-50 to extract multi-scale feature maps, and then extracts visual features of different scales through a feature pyramid network to obtain the visual feature set; The depth feature extraction branch first converts the depth image into a voxel representation, and then applies a three-dimensional convolutional neural network to extract features to obtain the depth feature set, which includes depth features of multiple scales and maintains the same spatial size as the visual features of multiple scales; The lightweight evaluation module consists of two layers of convolution and global average pooling, and outputs the quality scores of the visual features and the depth features at the s-th scale and the quality scores of the depth features Calculate the fusion weights and perform feature fusion:
[0051]
[0052] Among them, represents the fusion weight of the visual features at the s-th scale, and exp() represents the natural exponential function. and respectively represent the fused features, visual features, and depth features at the s-th scale. represents the result of 1×1 convolution after concatenating the visual features and depth features at the s-th scale along the channel dimension;
[0053] The calculation methods of the dynamic detection threshold and the dynamic segmentation threshold are to multiply the detection base threshold and the segmentation base threshold by the light adjustment function and the motion adjustment function respectively. The detection base threshold and the segmentation base threshold are selected by balancing the precision and recall under standard conditions. These base thresholds serve as the starting points for dynamic threshold adjustment and will be fine-tuned according to the test results in the actual deployment environment to ensure the stability of the system performance under various lighting and motion conditions; the light adjustment function f light (L) and the motion adjustment function f motion (V) are as follows:
[0054]
[0055] f motion (V)=1 + α m ·min(1, V / V th );
[0056] Among them, α l and α m respectively represent the adjustment coefficients of the light adjustment function and the motion adjustment function, and can be determined by fitting experimental data to find the relationship curve between the optimal threshold and the light to determine α l , and can be determined by analyzing the relationship between the motion blur degree at different speeds and the optimal detection threshold to determine α m ; L represents the environmental illuminance, calculated based on the color image, and L opt represents the optimal illuminance, which can be determined by testing the detection accuracy under different lighting conditions in the controlled environment to determine L opt ; V represents the target motion speed, which is obtained by associating the radar points with the target bounding box, finding the set of radar points belonging to the same target, and calculating the average motion speed of these points to obtain the target motion speed estimate. V th represents the speed threshold, which can be determined by analyzing the detection performance at different speeds and finding the speed point where the performance starts to significantly decline to determine V th ; min() represents the minimum value function;
[0057] The initial detection information includes the category, bounding box, and confidence of the target; the initial segmentation information includes the category probability of each pixel.
[0058] The post - processing output layer first applies the non - maximum suppression algorithm to remove duplicate detection boxes, then only retains the detection results with confidence higher than the dynamic detection threshold, and selects the category with the highest probability as the target category. For each pixel point in the image, if its highest probability value belonging to any category exceeds the dynamic segmentation threshold, the pixel is marked as the category with the highest probability; otherwise, the pixel is marked as the background.
[0059] According to the filtered initial detection information and initial segmentation information, an observation information list O = {o1, o1,..., o n} is obtained, where the observation information o i of each target includes: category, bounding box, confidence, segmentation mask, pixel - level depth information, speed information; n represents the number of filtered targets.
[0060] Furthermore, the self - motion decoupling mechanism specifically includes: using an extended Kalman filter to fuse inertial measurement unit data to estimate the self - motion state vector of the robot; calculating a coordinate system transformation matrix according to the self - motion state vector; extracting the target position based on the observation information of each target and transforming it to the world coordinate system through the coordinate system transformation matrix; obtaining the true motion vector of the target by combining the world coordinate position in the previous frame; and updating the observation information list based on the true motion vector to obtain the true information list.
[0061] The self - motion decoupling mechanism effectively avoids the target motion perception error caused by the movement of the sweeping robot by separating the robot's self - motion and the target's true motion, eliminates the interference brought by self - motion, can accurately distinguish the apparent motion and true motion of the target, thereby providing reliable motion information for subsequent trajectory association and prediction, and improving the tracking accuracy in a dynamic environment.
[0062] In this embodiment, the self - motion state vector of the robot at time t is where the dimension of each component is 3, is the position of the robot at time t, is the speed of the robot at time t, is the quaternion representing rotation at time t, is the angular velocity of the robot at time t. The coordinate system transformation matrix is expressed as:
[0063]
[0064] where, represents the coordinate system transformation matrix from the robot coordinate system to the world coordinate system at time t, represents the rotation matrix corresponding to the quaternion at time t; T t-1→tRepresents the relative transformation between two moments of the robot; Represents the coordinate system transformation matrix at time t-1;
[0065] The target position of the i-th target at time t Is transformed into the world coordinate system to obtain the world coordinate position of the target
[0066] The true motion vector of the i-th target Where, Represents the world coordinate position of the i-th target at time t-1, M ego,i Represents the displacement caused by the self-motion of the robot, and the calculation method is: the world coordinate position of the target at time t-1 is inversely transformed to the robot coordinate system through the coordinate system transformation matrix to obtain Then calculate the position of the i-th target in the robot coordinate system at time t Then Is transformed back to the world coordinate system and compared with Subtracted to obtain M ego,i ;
[0067] The true information list includes the observation information list, the world coordinate position of the target, and the true speed, and the true speed is the ratio of the true motion vector to the time interval.
[0068] Further, based on the true information list, establish the spatial and temporal associations between the target and the trajectory, and obtain the latest trajectory set including:
[0069] Extract the depth appearance features for each target detected in the current frame and each existing trajectory;
[0070] Calculate the appearance similarity matrix based on the depth appearance features, calculate the spatial position similarity matrix based on the world coordinate position, and calculate the speed consistency similarity matrix based on the estimated speed of the target and the predicted speed of the trajectory;
[0071] Weightedly fuse the appearance similarity matrix, the spatial position similarity matrix, and the speed consistency similarity matrix to obtain the comprehensive similarity matrix;
[0072] The Hungarian algorithm is used to solve the optimal matching relationship for the comprehensive similarity matrix. According to the optimal matching relationship, new trajectory creation, trajectory update, and termination are performed to obtain the latest trajectory set, which specifically includes: for the targets that are not matched to a trajectory, if the confidence of the new trajectory is greater than the threshold, a new trajectory is created; for the trajectories that are matched to a target, the trajectory state is updated through a Kalman filter; for the trajectories that are not matched to a target, if the number of consecutive unmatched frames is less than the unmatched frame threshold, a long short-term memory network is used to predict the trajectory state, otherwise the tracking of this trajectory is terminated. Each trajectory includes the current state of the trajectory and historical information. The current state of the trajectory includes the trajectory identification ID, world coordinate position, true speed, scale, category, appearance feature, and survival frame number; the historical information includes the position history sequence, speed history sequence, and appearance feature history sequence.
[0073] By comprehensively considering the similarities in the three dimensions of appearance, position, and speed, the problem of association errors easily caused by a single feature is effectively solved. By using the Hungarian algorithm to solve the optimal matching relationship, efficient target-trajectory association is achieved. The complete trajectory management strategy can effectively handle complex situations such as the appearance and disappearance of targets. This spatio-temporal association mechanism significantly improves the ability of continuous target tracking and provides stable and continuous target motion history information for subsequent trajectory prediction.
[0074] In this embodiment, deep appearance feature extraction is implemented through a pre-trained convolutional neural network; the deep appearance feature of the target is extracted from the color image of the corresponding target area; the appearance feature of the trajectory is extracted when the target is first detected and a new trajectory is created;
[0075] The calculation methods of the elements in the three similarity matrices are as follows:
[0076] Calculate the dot product of the deep appearance features of each detected target and the trajectory, and then divide it by the product of the norms of the deep appearance features of the target and the trajectory to obtain an appearance similarity score ranging from [-1, 1]. The larger the value, the more similar the appearance.
[0077] Square the Euclidean distance between the world coordinate position of the detected target and the current predicted position of the trajectory, divide it by twice the position variance, take the negative value, and then calculate through the natural exponential function to obtain a spatial position similarity score ranging from (0, 1]. The closer the distance, the closer the value is to 1.
[0078] Square the Euclidean distance between the motion speed of the detected target and the predicted speed of the trajectory, divide it by twice the speed variance, take the negative value, and then calculate through the natural exponential function to obtain a speed consistency similarity score ranging from (0, 1]. The more consistent the speed, the closer the value is to 1.
[0079] Furthermore, by analyzing the target characteristics, environmental constraints, social behaviors, and historical trajectories, the predicted trajectories of each target and the corresponding multi-factor scores are generated, including: selecting the corresponding motion model according to the target category to calculate the estimated motion parameters; constructing an indoor environmental grid map using depth information, calculating the static obstacle distance field, and defining an environmental cost function, and calculating the environmental feasibility score according to the environmental cost function; calculating the total social influence of each target by all other targets; using a gated recurrent unit network to obtain the historical trajectory encoding of the target; based on the estimated motion parameters, the indoor environmental grid map, the total social influence, and the historical trajectory encoding of the target, generating multiple initial predicted trajectories through a conditional variational autoencoder, and combining the environmental feasibility score, the total social influence, and the acceleration of the predicted trajectory to obtain the multi-factor score of each initial predicted trajectory; sorting the initial predicted trajectories from largest to smallest according to the multi-factor scores, and selecting the top m trajectories as the predicted trajectories.
[0080] The multi-factor trajectory prediction method realizes accurate trajectory prediction for different types of targets by comprehensively analyzing target characteristics, environmental constraints, social behaviors, and historical trajectories. It uses a conditional variational autoencoder to generate diverse predicted trajectories and screens the most likely trajectories through multi-factor scoring, effectively handling the uncertainty problem of target motion. It not only improves the prediction accuracy but also provides more comprehensive information for subsequent decision-making by generating multiple possible trajectories.
[0081] The indoor environmental grid map \(G\in\{0,1\}\), where \(G = 1\) indicates that there is an obstacle at the grid \((u, v)\), and \(G = 0\) indicates that the grid \((u, v)\) is a passable area, and \(M\times N\) represents the number of grids; the static obstacle distance field \(D\) represents the distance from each grid to the nearest obstacle, and the environmental cost function \(C(P)\) is expressed as: M×N , \(G\) u,v where \(G = 1\) indicates that there is an obstacle at the grid \((u, v)\), and \(G = 0\) indicates that the grid \((u, v)\) is a passable area, and \(M\times N\) represents the number of grids; the static obstacle distance field \(D\) represents the distance from each grid to the nearest obstacle, and the environmental cost function \(C(P)\) is expressed as: u,v where \(G = 1\) indicates that there is an obstacle at the grid \((u, v)\), and \(G = 0\) indicates that the grid \((u, v)\) is a passable area, \(M\times N\) represents the number of grids; the static obstacle distance field \(D\) represents the distance from each grid to the nearest obstacle, and the environmental cost function \(C(P)\) is expressed as: env is expressed as:
[0082]
[0083] where \(P\) represents the position in the world coordinate system, \(D(P)\) represents the value of the static obstacle distance field of the grid corresponding to the position \(P\), \(\infty\) represents infinity, \(r\) min represents the safety distance threshold, \(\sigma\) env represents the control coefficient used to control the influence range of the obstacle, and \(\exp()\) represents the natural exponential function; the calculation method of the environmental feasibility score of the trajectory is: taking the negative value of the environmental cost function value and then calculating through the natural exponential function;
[0084] Calculate the attention weight \(a\) of target \(i\) to each other target \(j\), which is weighted and summed by the distance and speed correlation between targets, and then normalized by the softmax function to obtain \(a\) i,j , which is weighted and summed by the distance and speed correlation between targets, and then normalized by the softmax function to obtain \(a\) i,j;Input the position, velocity, and category information of two targets into a multi-layer perceptron, multiply the output by the corresponding attention weights to obtain the individual influence of target j on target i; sum up the individual influences of target i from all other targets to obtain the total social influence;
[0085] Input the environmental feasibility score, total social influence, and acceleration of the predicted trajectory into a multi-layer perceptron to output a multi-factor score.
[0086] Further, generating a time-varying influence map based on the predicted trajectory and the multi-factor score includes: initializing a static influence map according to the indoor environment grid map; calculating the influence contribution of each predicted trajectory of the target to each grid in the static influence map; summarizing the influence contributions of all targets and trajectories to obtain a dynamic influence map; combining the static influence map and the dynamic influence map to obtain a time-varying influence map.
[0087] The influence contribution is obtained by multiplying the normalized multi-factor score of the predicted trajectory of the target by the Gaussian function of the distance from the target to the grid. Summarizing the influence contributions of all targets and trajectories to obtain a dynamic influence map:
[0088]
[0089] Among them, represents the dynamic influence map, represents the influence contribution of each predicted trajectory k of target i to the grid (u, v) at time t, n represents the number of targets, m represents the number of predicted trajectories of each target, and the time-varying influence map is obtained by selecting the value in the static influence map and the dynamic influence map through the maximum function.
[0090] The construction of the time-varying influence map integrates the influences of static obstacles and dynamic targets into a unified representation. By considering the uncertainty and diversity of the predicted trajectory, reasonable influence contributions are assigned to each spatio-temporal grid, avoiding the limitations of the traditional binary occupancy grid map, being able to intuitively reflect the change of influence in the environment over time, being suitable for the obstacle avoidance task in a dynamic environment, enabling the sweeping robot to foresee potential collision risks, adjust the path in advance, and achieve a safer and more efficient obstacle avoidance effect.
[0091] Further, adopting a two-layer decision-making mechanism for planning and decision-making includes: establishing a three-dimensional spatio-temporal search space based on the time-varying influence map, designing an evaluation function by comprehensively considering the time cost, space cost, and influence cost; setting a global planning period, and performing heuristic search according to the evaluation function to obtain a global planning path; setting a local adjustment period, sampling the feasible speed space to obtain speed samples; calculating the comprehensive score of each speed sample, and selecting the speed sample with the highest comprehensive score to generate a local execution trajectory.
[0092] The evaluation function is obtained by weighted summation of the actual cost, estimated cost, and impact cost. The actual cost combines the time cost (number of elapsed time steps) and the space cost (travel distance and turning cost); usually, the Euclidean distance or Manhattan distance is used to calculate the estimated cost from a node to the target point; the impact cost is calculated by weighted summation of the impact contribution values at all time points on the path where the node is located with time discounting.
[0093] The comprehensive score of each speed sample considers four factors: the heading factor uses the cosine value of the angle between the heading angle and the target point, the distance factor is based on the minimum distance to the obstacle, the speed factor is the ratio of the linear speed to the maximum speed, and the impact factor is the maximum impact contribution value at different times on the trajectory. The comprehensive score is obtained by weighted summation of the four factors, and the weight adjustment mechanism changes dynamically according to the environmental complexity and target density. For example, in a dense target area, the weight of the impact factor is increased to prioritize safety, and in an open area, the weight of the speed factor is increased to improve efficiency.
[0094] The two-layer decision-making mechanism divides path planning into two levels: global planning and local adjustment, realizing the organic combination of long-term goal orientation and real-time response, forming a complementary structure of low-frequency global decision-making and high-frequency local response, significantly improving the obstacle avoidance ability and path planning efficiency of the sweeping robot in a dynamic environment, and enabling a smoother and more efficient cleaning operation while ensuring safety.
[0095] The present invention establishes a complete technical link from multi-modal perception, target detection, self-motion compensation, multi-target tracking to trajectory prediction and obstacle avoidance decision-making, solving the core problems of the sweeping robot in a dynamic environment. The dynamic threshold dual-stream neural network can fuse the complementary information of RGB images and depth images and dynamically adjust the detection threshold according to environmental conditions, effectively overcoming the limitations of a single sensor in a complex environment; secondly, the self-motion decoupling mechanism and spatio-temporal correlation mechanism achieve stable tracking of targets in the robot's motion scenario; thirdly, the trajectory prediction method based on multi-factor analysis can generate diverse possible trajectories and give reasonable scores, improving the prediction accuracy while retaining the uncertainty information of motion; finally, the time-varying influence map and the two-layer decision-making mechanism achieve early perception of future collision risks and multi-scale responses, enabling the sweeping robot to actively and smoothly avoid moving objects while ensuring cleaning efficiency. The present invention significantly improves the running safety and intelligent level of the sweeping robot in a complex dynamic environment.
[0096] Embodiment 2:
[0097] This embodiment is applied to the scenario of an intelligent floor-sweeping robot in a home environment, which contains various dynamic objects, such as moving people, pets, and children's toys. During the operation of traditional floor-sweeping robots, these dynamic objects often lead to collision events or a decrease in cleaning efficiency. This embodiment applies a dynamic multi-object tracking and trajectory prediction method, including:
[0098] Collect color images and depth images and perform preprocessing;
[0099] Construct a dynamic threshold two-stream neural network to extract features from the preprocessed color images and depth images, obtain a visual feature set and a depth feature set, and fuse the visual feature set and the depth feature set through an adaptive attention mechanism to obtain a fused feature set, and combine with a dynamic threshold to obtain a list of observation information of the target;
[0100] Obtain its own motion state vector, and update the list of observation information through a self-motion decoupling mechanism to obtain a list of real information;
[0101] Establish spatial and temporal associations between the target and the trajectory based on the list of real information to obtain the latest trajectory set;
[0102] Analyze target characteristics, environmental constraints, social behaviors, and historical trajectories to generate a predicted trajectory for each target and the corresponding multi-factor scores;
[0103] Generate a time-varying influence diagram based on the predicted trajectory and the multi-factor scores, and adopt a two-layer decision-making mechanism for planning and decision-making.
[0104] Furthermore, the dynamic threshold two-stream neural network includes:
[0105] A feature extraction layer, including: the visual feature extraction branch adopts a feature pyramid network structure to process the preprocessed color image to obtain the visual feature set; the depth feature extraction branch adopts a three-dimensional convolutional neural network to process the preprocessed depth image to obtain the depth feature set;
[0106] An adaptive attention layer, including: a lightweight evaluation module analyzes the quality scores of visual features and depth features; an attention fusion module adjusts the fusion weights according to the quality scores and fuses the visual feature set and the depth feature set to obtain the fused feature set;
[0107] A detection and segmentation layer, including: a detection head branch outputs initial detection information based on the fused feature set; a segmentation head branch outputs initial segmentation information based on the fused feature set;
[0108] The post-processing output layer includes: a dynamic threshold adjustment module that calculates a dynamic detection threshold and a dynamic segmentation threshold based on the ambient light condition and the target motion speed; an output module that combines the dynamic detection threshold and the dynamic segmentation threshold to screen the initial detection information and the initial segmentation information to obtain the list of observation information.
[0109] Further, the self-motion decoupling mechanism specifically includes: using an extended Kalman filter to fuse the inertial measurement unit data to estimate the self-motion state vector of the robot; calculating a coordinate system transformation matrix according to the self-motion state vector; extracting the target position according to the observation information of each target and transforming it to the world coordinate system through the coordinate system transformation matrix; obtaining the true motion vector of the target by combining the world coordinate position in the previous frame; and updating the list of observation information based on the true motion vector to obtain the list of true information.
[0110] Further, establishing the spatial association and temporal association between the target and the trajectory based on the list of true information to obtain the latest trajectory set includes:
[0111] Extracting the depth appearance features for each target detected in the current frame and each existing trajectory;
[0112] Calculating an appearance similarity matrix based on the depth appearance features, calculating a spatial position similarity matrix based on the world coordinate positions, and calculating a speed consistency similarity matrix based on the estimated speed of the target and the predicted speed of the trajectory;
[0113] Weightedly fusing the appearance similarity matrix, the spatial position similarity matrix, and the speed consistency similarity matrix to obtain a comprehensive similarity matrix;
[0114] Using the Hungarian algorithm to solve the optimal matching relationship for the comprehensive similarity matrix, and according to the optimal matching relationship, creating new trajectories, updating trajectories, and terminating to obtain the latest trajectory set, specifically including: for the targets that are not matched to a trajectory, if the confidence of the new trajectory is greater than the threshold, create a new trajectory; for the trajectories that are matched to a target, update the trajectory state through a Kalman filter; for the trajectories that are not matched to a target, if the number of consecutive unmatched frames is less than the unmatched frame number threshold, use a long short-term memory network to predict the trajectory state, otherwise terminate the tracking of this trajectory.
[0115] Further, analyze the target characteristics, environmental constraints, social behaviors, and historical trajectories to generate the predicted trajectories for each target and the corresponding multi-factor scores, including: selecting the corresponding motion model according to the target category to calculate the estimated motion parameters; using depth information to construct an indoor environmental grid map, calculating the static obstacle distance field, and defining an environmental cost function, and calculating the environmental feasibility score according to the environmental cost function; calculating the total social influence of each target by all other targets; using a gated recurrent unit network to obtain the historical trajectory encoding of the target; based on the estimated motion parameters of the target, the indoor environmental grid map, the total social influence, and the historical trajectory encoding, generating multiple initial predicted trajectories through a conditional variational autoencoder, and combining the environmental feasibility score, the total social influence, and the acceleration of the predicted trajectory to obtain the multi-factor scores of each initial predicted trajectory; sorting the initial predicted trajectories from large to small according to the multi-factor scores, and selecting the top m trajectories as the predicted trajectories.
[0116] Table 1 shows the trajectory prediction performance for different types of targets. By selecting appropriate motion models for each target and considering environmental constraints, the system achieves high prediction accuracy. Targets with strong motion regularity, such as adults and toy cars, have higher prediction accuracy, while targets with irregular motion, such as children and cats, have relatively larger prediction errors, but still remain within an acceptable range. Overall, the average displacement error (ADE) for long-term prediction (3 seconds) is controlled within 0.6 meters, and the trajectory coverage rate (the probability that the true trajectory falls within the range of the predicted trajectory set within 3 seconds) reaches 83.12% on average.
[0117] Table 1 Comparison table of trajectory prediction performance for different target types
[0118]
[0119] Further, generate a time-varying influence map according to the predicted trajectory and the multi-factor scores, including: initializing a static influence map according to the indoor environmental grid map; calculating the influence contribution of each predicted trajectory of the target to each grid in the static influence map; summarizing the influence contributions of all targets and trajectories to obtain a dynamic influence map; combining the static influence map and the dynamic influence map to obtain a time-varying influence map.
[0120] Further, adopt a two-layer decision-making mechanism for planning and decision-making, including: establishing a three-dimensional spatio-temporal search space based on the time-varying influence map, and designing an evaluation function by comprehensively considering the time cost, space cost, and influence cost; setting a global planning period, and performing heuristic search according to the evaluation function to obtain a global planning path; setting a local adjustment period, sampling the feasible speed space to obtain speed samples; calculating the comprehensive score of each speed sample, and selecting the speed sample with the highest comprehensive score to generate a local execution trajectory.
[0121] Table 2 evaluates the obstacle avoidance performance under different home scenarios. The data shows that the system maintains a high obstacle avoidance success rate in various environments while ensuring reasonable cleaning efficiency. The obstacle avoidance time reflects predictability, meaning the ability to identify potential collision risks in advance and make adjustments.
[0122] Table 2 Obstacle Avoidance Path Planning Performance Evaluation Table
[0123]
[0124] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A dynamic multi-object tracking and trajectory prediction method, characterized in that, Including: Collecting color images and depth images and performing preprocessing; Constructing a dynamic threshold two-stream neural network to extract features from the preprocessed color images and depth images, obtaining a visual feature set and a depth feature set, and fusing the visual feature set and the depth feature set through an adaptive attention mechanism to obtain a fused feature set, and combining the dynamic threshold to obtain a list of observation information of the target; Obtaining the self-motion state vector, and updating the list of observation information through a self-motion decoupling mechanism to obtain a list of real information; Based on the list of real information, establishing spatial and temporal associations between the target and the trajectory to obtain the latest trajectory set; Analyzing target characteristics, environmental constraints, social behaviors, and historical trajectories to generate a predicted trajectory for each target and the corresponding multi-factor score; Generating a time-varying influence diagram according to the predicted trajectory and the multi-factor score, and adopting a two-layer decision-making mechanism for planning and decision-making.
2. The dynamic multi-object tracking and trajectory prediction method according to claim 1, characterized in that, The dynamic threshold two-stream neural network includes: A feature extraction layer, including: the visual feature extraction branch adopts a feature pyramid network structure to process the preprocessed color image to obtain the visual feature set; the depth feature extraction branch adopts a three-dimensional convolutional neural network to process the preprocessed depth image to obtain the depth feature set; An adaptive attention layer, including: a lightweight evaluation module analyzes the quality scores of visual features and depth features; an attention fusion module adjusts the fusion weights according to the quality scores and fuses the visual feature set and the depth feature set to obtain the fused feature set; A detection and segmentation layer, including: the detection head branch outputs initial detection information based on the fused feature set; the segmentation head branch outputs initial segmentation information based on the fused feature set; A post-processing output layer, including: a dynamic threshold adjustment module calculates a dynamic detection threshold and a dynamic segmentation threshold according to the environmental illumination condition and the target motion speed; an output module combines the dynamic detection threshold and the dynamic segmentation threshold to screen the initial detection information and the initial segmentation information to obtain the list of observation information.
3. The dynamic multi-object tracking and trajectory prediction method according to claim 1, characterized in that, The self-motion decoupling mechanism specifically includes: using an extended Kalman filter to fuse inertial measurement unit data to estimate the self-motion state vector of the robot; calculating a coordinate system transformation matrix according to the self-motion state vector; extracting the target position from the observation information of each target and converting it to the world coordinate system through the coordinate system transformation matrix; combining the world coordinate position in the previous frame to obtain the real motion vector of the target; updating the list of observation information based on the real motion vector to obtain the list of real information.
4. A dynamic multi-target tracking and trajectory prediction method according to claim 1, characterized in that, Establishing spatial and temporal associations between the target and the trajectory based on the list of real information to obtain the latest trajectory set includes: Extracting depth appearance features for each target detected in the current frame and each existing trajectory; Calculating an appearance similarity matrix based on the depth appearance features, calculating a spatial position similarity matrix based on the world coordinate positions, and calculating a speed consistency similarity matrix based on the estimated speed of the target and the predicted speed of the trajectory; Weightedly fusing the appearance similarity matrix, the spatial position similarity matrix, and the speed consistency similarity matrix to obtain a comprehensive similarity matrix; The Hungarian algorithm is used to solve the optimal matching relationship for the comprehensive similarity matrix. According to the optimal matching relationship, new trajectory creation, trajectory update, and termination are performed to obtain the latest trajectory set, which specifically includes: for a target that has not been matched to a trajectory, if the confidence of the new trajectory is greater than the threshold, a new trajectory is created; for a trajectory that has been matched to a target, the trajectory state is updated through a Kalman filter; for a trajectory that has not been matched to a target, if the number of consecutive unmatched frames is less than the unmatched frame number threshold, the long short-term memory network is used to predict the trajectory state, otherwise, the tracking of this trajectory is terminated.
5. A dynamic multi-object tracking and trajectory prediction method according to claim 1, characterized in that Analyze the target characteristics, environmental constraints, social behaviors, and historical trajectories to generate the predicted trajectories of each target and the corresponding multi-factor scores, including: selecting the corresponding motion model according to the target category to calculate the estimated motion parameters; using depth information to construct an indoor environment grid map, calculating the static obstacle distance field, and defining the environmental cost function, and calculating the environmental feasibility score according to the environmental cost function; calculating the total social influence of each target by all other targets; using a gated recurrent unit network to obtain the historical trajectory encoding of the target; based on the estimated motion parameters, the indoor environment grid map, the total social influence, and the historical trajectory encoding of the target, generating multiple initial predicted trajectories through a conditional variational autoencoder, and combining the environmental feasibility score, the total social influence, and the acceleration of the predicted trajectory to obtain the multi-factor scores of each initial predicted trajectory; sorting the initial predicted trajectories from largest to smallest according to the multi-factor scores, and selecting the top m trajectories as the predicted trajectories.
6. A dynamic multi-target tracking and trajectory prediction method according to claim 1, characterized in that, Generate a time-varying influence map according to the predicted trajectories and the multi-factor scores, including: initializing the static influence map according to the indoor environment grid map; calculating the influence contribution of each predicted trajectory of the target to each grid in the static influence map; summarizing the influence contributions of all targets and trajectories to obtain the dynamic influence map; combining the static influence map and the dynamic influence map to obtain the time-varying influence map.
7. A dynamic multi-object tracking and trajectory prediction method according to claim 1, characterized in that Adopt a two-layer decision-making mechanism for planning and decision-making, including: establishing a three-dimensional spatio-temporal search space based on the time-varying influence map, and designing an evaluation function by comprehensively considering the time cost, space cost, and influence cost; setting a global planning period, and performing heuristic search according to the evaluation function to obtain the global planning path; setting a local adjustment period, sampling the feasible speed space to obtain speed samples; calculating the comprehensive score of each speed sample, and selecting the speed sample with the highest comprehensive score to generate the local execution trajectory.
Citation Information
Patent Citations
Scene flow digital twinning method and system based on dynamic trajectory flow
CN114970321A
Real-time accurate trajectory prediction method and system in unstructured human-computer interaction environment
CN115659275A
Pedestrian trajectory prediction method combining spatio-temporal information and social interaction features
CN115829171A
Multi-moving target detection and trajectory prediction method in field environment
CN116385493A
Target tracking and trajectory prediction method and device, equipment and storage medium
CN117890922A
Cited By
Target tracking method and device
CN120526406A
Moving target trajectory data generation method
CN120765694A
A mobile target trajectory data generation method
CN120765694B
Fish migration trajectory tracking method
CN120782823A
A fish migration trajectory tracking method
CN120782823B