Tracking method, device and equipment for self-balancing robot and medium
By generating a continuous and smooth target trajectory through high-frequency extrapolation and combining it with physical constraints to generate control torque, the jitter and instability problems caused by frequency mismatch in the self-balancing robot are solved, achieving a more stable tracking task.
Patent Information
- Application Number
- CN202511263303.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-05
AI Technical Summary
In the tracking task, the existing self-balancing robot has problems such as discontinuous control instructions due to the mismatch between the sensor frequency and the control frequency, which causes posture jitter and instability.
By acquiring target pose data at a frequency higher than the sensor frequency, a continuous and smooth target trajectory is generated by extrapolation, and the control torque is generated in combination with the robot's own physical constraints to ensure that the robot maintains balance at high frequencies.
The smoothness and stability of the self-balancing robot in tracking tasks are improved, control command delays and jumps caused by frequency mismatch are avoided, and the reliability and safety of the robot in complex environments are ensured.
Smart Images

Figure CN120742909A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robot control technology, and in particular to a tracking method, device, equipment and medium for a self-balancing robot. Background Art
[0002] When a self-balancing robot performs tracking tasks, its control system requires high-frequency, continuous commands to maintain its dynamic balance in real time. However, in existing tracking solutions, the sensors used to perceive the target's position and posture typically provide discrete data periodically at a relatively low frequency. When this low-frequency, discontinuous position data is directly used to generate control commands, it can lead to sudden and inconsistent commands, which can easily cause the self-balancing robot's posture to oscillate, or even cause instability during dynamic tracking due to uneven control.
[0003] Therefore, there is an urgent need for a tracking method for self-balancing robots that can effectively convert low-frequency, discrete target pose measurements into continuous and smooth target trajectories that can be used by high-frequency balancing controllers. Summary of the Invention
[0004] In view of this, the present application provides a tracking method, device, equipment and medium for a self-balancing robot to solve the problems of discontinuous control instructions and dynamic instability caused by mismatch between sensing and control frequencies in the prior art.
[0005] In a first aspect, the present application provides a tracking method for a self-balancing robot, the method being performed by a self-balancing robot serving as a slave vehicle, the method being used to track a moving target serving as a master vehicle, the method comprising: Periodically acquiring position and posture data of a mobile target at a first frequency to obtain time-series position and posture data of the mobile target; Determine the motion state of the moving target based on the time series posture data; the motion state includes position, velocity and acceleration; Extrapolating the motion state to generate a target trajectory of the self-balancing robot at a second frequency; the second frequency is higher than the first frequency; Based on the target trajectory, the control torque of the self-balancing robot is determined so that the self-balancing robot maintains balance while tracking along the target trajectory.
[0006] The tracking method for a self-balancing robot provided in this application determines the complete motion state of a moving target, including position, velocity, and acceleration, based on time-series pose data. This provides richer and more accurate physical model input for subsequent trajectory prediction and extrapolation, laying the foundation for high-quality trajectory generation. The determined motion state is extrapolated to generate the target trajectory at a second frequency, much higher than the data acquisition frequency (the first frequency). Low-frequency, discrete state input points are converted into a high-frequency, continuous trajectory command stream. Therefore, even in the interval between two acquisitions of raw pose data, the control system of the self-balancing robot can continuously obtain smooth and coherent target guidance, fundamentally solving the problem of control command delays and jumps caused by data frequency mismatch. Ultimately, the self-balancing robot determines its control torque based on this high-frequency, smooth target trajectory. Because the reference trajectory used as the control basis is continuous and changes smoothly, the calculated control torque is also continuous and stable, avoiding the severe impact on the robot's posture caused by sudden command changes. This enables the robot to accurately track moving targets while effectively maintaining its own dynamic balance, thereby significantly improving the smoothness and stability of the entire tracking process, and solving the problem in existing technologies where self-balancing robots are prone to shaking or even becoming unstable and falling due to discontinuous control.
[0007] In an optional embodiment, obtaining the position data of the moving target includes: The camera mounted on the self-balancing robot detects the visual marker installed on the moving target, calculates the position of the visual marker relative to the camera, and obtains the position data of the moving target.
[0008] The tracking method for a self-balancing robot provided in this application utilizes visual markers as passive beacons, making the position acquisition process completely independent of external wireless communications (such as Wi-Fi and GPS). This method is therefore capable of stable operation in communication-restricted or signal-free environments, such as indoors and in tunnels. It also reduces the system's reliance on communication hardware, thereby lowering costs and power consumption, and fundamentally avoiding the risk of tracking failures caused by communication delays or interruptions.
[0009] In an optional implementation, determining the motion state of the moving target includes: The time series pose data is used to determine the relative pose of the moving target relative to the self-balancing robot; Based on the relative posture and the self-balancing robot's own posture in the world coordinate system, the time series position data of the moving target in the world coordinate system is obtained through coordinate system conversion; Based on the time series position data, the velocity and acceleration of the moving target are calculated by backward difference to determine the motion state of the moving target.
[0010] The tracking method for a self-balancing robot provided in this application enables the robot's motion planning and control to be performed in a unified global coordinate system through coordinate transformation; and through differential calculation, the system can grasp the dynamic trend (speed and acceleration) of the target, not just its instantaneous position, providing an indispensable data foundation for subsequent precise prediction and control.
[0011] In an optional embodiment, before extrapolating the motion state, the method further includes: The motion state is smoothed by exponential weighted filtering.
[0012] The tracking method for a self-balancing robot provided in this application effectively suppresses data jitter and glitches caused by sensor measurement noise and differential calculations by introducing exponentially weighted filtering. This makes the motion state data input into the subsequent extrapolation model smoother and more stable, closer to the actual physical motion of the target. As a result, the quality of the generated final target trajectory is significantly improved, avoiding sudden changes in the trajectory due to noise interference, thus laying a solid foundation for achieving smoother robot control.
[0013] In an optional embodiment, the motion state is extrapolated, including: The motion state is input into a preset constant acceleration prediction model for extrapolation to generate a target trajectory; wherein the target trajectory is composed of a series of target points, and the time sampling frequency of the target points is the second frequency.
[0014] In an optional embodiment, the constant acceleration prediction model is extrapolated according to a second-order polynomial, and the expression is: ; ; Where, and They represent the target position and speed of the following vehicle predicted at any control time t; 、 、 Represent the latest position, velocity, and acceleration at the time of the last visual detection update; Represents the current control time t and the last visual update time The time difference between .
[0015] The tracking method for a self-balancing robot provided in this application not only considers the target's position and velocity but also incorporates its acceleration into the prediction process. Compared to zero-order (maintaining position) or first-order (constant velocity) predictions, this method can more accurately predict the target's motion trend over short periods of time, especially when the target is accelerating or decelerating. Using this model to generate high-frequency (second-frequency) trajectory points ensures that the generated trajectory is not only high-frequency and dense, but also physically more reasonable and accurate.
[0016] In an optional embodiment, periodically acquiring the position and posture data of the mobile target at a first frequency further includes: When the moving target is successfully identified continuously and the linear velocity of the moving target reaches a preset non-zero threshold, the position and posture data of the moving target is periodically acquired at a first frequency.
[0017] The tracking method for a self-balancing robot provided in this application effectively prevents two typical failure scenarios: first, initiating tracking of a non-existent or incorrect target due to transient sensor misidentification or environmental interference; and second, attempting to track a completely stationary target, which can cause the robot to experience unnecessary jitter or oscillation during in-situ fine-tuning. Therefore, this feature significantly improves the robustness and safety of the system during startup, ensuring that tracking tasks only begin after the target is confirmed to be a stable and valid dynamic object.
[0018] In an optional embodiment, the method further includes: When the pose data of the moving target is lost continuously for a preset number of frames, or the error between the predicted target trajectory and the actual control state exceeds a preset threshold, the self-balancing robot switches to the preset backup trajectory control mode.
[0019] The tracking method for a self-balancing robot provided in this application prevents the robot from blindly executing a predicted trajectory that has become invalid or has accumulated excessive errors after losing its target. This blind execution is highly likely to cause the robot to lose control or crash. By switching to a backup mode (such as maintaining balance in place or executing a preset safe path), the present invention ensures that the robot can maintain its safety and stability even when visual perception is interrupted, greatly enhancing the reliability and practicality of the method in complex environments.
[0020] In an optional embodiment, determining the control torque of the self-balancing robot includes: The control torque is generated based on the underactuated dynamic constraints and ground reaction force constraints of the self-balancing robot.
[0021] The tracking method for a self-balancing robot provided in this application directly incorporates the robot's inherent physical constraints (underactuated characteristics, ground friction, etc.) into the calculation of the control torque, ensuring that every control instruction generated is physically feasible. This avoids invalid instructions that may be issued by traditional simple controllers that exceed the robot's physical limits (for example, requiring the robot to instantly translate sideways). Therefore, the effectiveness and safety of control instructions can be fundamentally guaranteed, allowing the robot to remain stable within its physical limits when performing high-speed, large-maneuver tracking tasks, thereby maximizing the robot's performance while ensuring its balance and stability.
[0022] In summary, the tracking method for a self-balancing robot provided in this application first uses a camera to detect visual markers to acquire the relative pose of the target. Combined with the robot's own pose, coordinate transformation and differential calculations are used to determine the target's complete motion state in a global coordinate system, including position, velocity, and acceleration. This method establishes a data processing chain from raw visual information to the complete physical state that is independent of external communication, ensuring that the system can obtain the basic data required for subsequent predictions in any environment. Furthermore, to improve data quality, an exponentially weighted filter is introduced to smooth the motion state, effectively eliminating sensor noise and computational errors, providing a more reliable and smooth data source for subsequent predictions. This is a prerequisite for achieving high-quality tracking. Next, using the smoothed motion state, a constant acceleration prediction model is applied for extrapolation, generating a high-density, continuous, and smooth target trajectory at a second frequency much higher than the data acquisition frequency. This method bridges the gap between low-frequency sensing and high-frequency control, fundamentally addressing the problem of control command jumps caused by data frequency mismatch. Furthermore, to further enhance robustness in practical applications, safety policies are added at both ends of the process. On the front end, a triggering judgment mechanism ensures that tracking is initiated only after the target is confirmed to be a valid dynamic object, avoiding misjudgments and instability in the initial phase. On the back end, a fault-tolerant switching mechanism enables a safe switch to a backup mode in the event of an emergency, such as target loss. This prevents loss of control due to perception interruptions and significantly improves reliability. Finally, this target trajectory is input into a controller that fully accounts for the robot's physical constraints. This ensures that each resulting motor control torque not only accurately drives the robot along the trajectory but also remains well within the robot's physical capabilities.
[0023] In a second aspect, the present application provides a tracking device for a self-balancing robot, the device being executed by a self-balancing robot as a slave vehicle, the device being used to track a moving target as a master vehicle, the device comprising: An acquisition module, configured to periodically acquire the position and posture data of the mobile target at a first frequency to obtain time-series position and posture data of the mobile target; The state confirmation module is used to determine the motion state of the mobile target based on the time series posture data; the motion state includes position, velocity and acceleration; a trajectory generation module, configured to extrapolate the motion state and generate a target trajectory of the self-balancing robot at a second frequency; the second frequency being higher than the first frequency; The control module is used to determine the control torque of the self-balancing robot based on the target trajectory so that the self-balancing robot maintains balance while tracking along the target trajectory.
[0024] In a third aspect, the present application provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to execute the tracking method for a self-balancing robot according to the first aspect or any corresponding embodiment thereof.
[0025] In a fourth aspect, the present application provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the tracking method for a self-balancing robot according to the first aspect or any corresponding embodiment thereof.
[0026] In a fifth aspect, the present application provides a computer program product comprising computer instructions, wherein the computer instructions are used to enable a computer to execute the tracking method for a self-balancing robot according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 is a flow chart of a tracking method for a self-balancing robot according to an embodiment of the present application; Figure 2 is a schematic diagram of the workflow of a tracking system for a two-wheeled vehicle according to an embodiment of the present application; Figure 3 is a structural block diagram of a tracking device for a self-balancing robot according to an embodiment of the present application; Figure 4 It is a schematic diagram of the hardware structure of the computer device of an embodiment of the present application. DETAILED DESCRIPTION
[0029] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0030] In autonomous tracking applications for self-balancing robots, existing technologies primarily rely on two approaches: communication-based formation control and deep learning-based visual detection. Communication-based solutions exchange state information, such as pose and velocity, between vehicles via wireless networks (such as Wi-Fi and Bluetooth). While enabling precise formation control, they rely heavily on communication quality and fail completely in signal-limited environments (such as indoors and in tunnels). They also suffer from inherent drawbacks such as communication latency, high costs, and safety risks. Deep learning-based visual detection solutions (such as YOLO and SSD) don't require communication, but their computational complexity is extremely high, making them difficult to implement in real-time at high frame rates on resource-constrained self-balancing robot embedded systems. Furthermore, these algorithms are sensitive to lighting variations, object occlusion, and appearance changes, resulting in insufficient detection stability. Their output is typically a bounding box, requiring complex post-processing to estimate the precise six-degree-of-freedom pose. These factors can easily lead to instability and falls of self-balancing robots due to object loss or pose errors.
[0031] To overcome these shortcomings, a more stable and reliable solution is to use a camera to detect preset visual markers (such as AprilTags) to determine the target pose. However, even with this seemingly superior solution, a series of profound technical challenges remain and need to be addressed. First, the camera's sampling frequency (e.g., 30 Hz) is far lower than the frequency required for self-balancing robot balance control (e.g., above 500 Hz). This significant frequency mismatch is the root cause of discontinuous control commands and jittery robot operation. Second, when attempting to estimate velocity and acceleration by differencing low-frequency, discrete position data, the sensor's inherent measurement noise is dramatically amplified, resulting in an estimated motion state full of glitches and interference, making it unsuitable for direct use in precise control. Furthermore, as a dynamically unstable system, a self-balancing robot is extremely sensitive to anomalies in its initial state and during operation. Existing technologies generally lack targeted startup stabilization strategies and fault-tolerance mechanisms for target loss, making it highly susceptible to loss of control at the start of a mission or when perception is interrupted. Finally, general controllers usually fail to fully consider the unique physical dynamic constraints of self-balancing robots, such as under-actuation and ground friction, and are prone to output control instructions that are physically impossible to achieve, thereby limiting the performance of the robot and possibly causing instability.
[0032] Therefore, the existing technology has not yet provided a complete and highly robust autonomous tracking solution that can systematically solve all of the above problems.
[0033] In this embodiment, a tracking method for a self-balancing robot is provided. The method is performed by a self-balancing robot as a slave vehicle. The method is used to track a moving target as a master vehicle. Figure 1 is a flow chart of a tracking method for a self-balancing robot according to an embodiment of the present application, such as Figure 1 As shown, the process includes the following steps: S101. Periodically acquire position and posture data of a moving target at a first frequency to obtain time-series position and posture data of the moving target.
[0034] Specifically, a self-balancing robot is a special robot whose physical structure makes it unable to maintain stability naturally when it is stationary or moving, and must maintain balance through active and continuous control.
[0035] The master vehicle refers to the moving target being tracked, and the slave vehicle refers to the self-balancing robot that executes this method.
[0036] Pose data is a combination of the position and attitude of an object. Position refers to the specific coordinates of the object in three-dimensional space (e.g., x, y, z), while attitude describes the orientation of the object (e.g., roll, pitch, and yaw, which are the rotation angles around the x, y, and z axes).
[0037] The first frequency refers to the rate at which target pose data is collected from the vehicle's perception system (such as a camera). It is a fixed, low frequency, for example, 30 times per second (30Hz).
[0038] Since time series pose data is collected periodically, a series of pose data points with timestamps are obtained. These data points are arranged in chronological order, forming a historical trajectory that records the pose of the target at different times.
[0039] S102. Determine the motion state of the moving target based on the time-series position data; the motion state includes position, velocity, and acceleration.
[0040] Specifically, the motion state is a complete description of how the target moves. It is not enough to have only the position information, but also its motion trend.
[0041] Determining the motion state is to process and calculate the original time-series pose data to derive richer motion information.
[0042] The position is obtained directly from the pose data obtained in the previous step.
[0043] By comparing the position change and time interval between two consecutive time points, the speed and direction of the target's movement can be calculated.
[0044] Acceleration, similarly, by comparing the speed change and time interval between two consecutive time points, the rate of change of the target speed can be calculated.
[0045] S103 , extrapolating the motion state to generate a target trajectory of the self-balancing robot at a second frequency; the second frequency is higher than the first frequency.
[0046] Specifically, extrapolation is one of the core steps of the method in this embodiment. Because the first frequency (perception frequency) is low, the acquired motion state is discrete and spaced. Extrapolation is a prediction technique that uses known motion states (position, velocity, acceleration) to fill in the gaps between two known data points and predict the state in the very near future.
[0047] The second frequency is a frequency much higher than the first frequency, for example 500 times per second (500 Hz), which is the command frequency required by the balance control system of the self-balancing robot.
[0048] The target trajectory is generated through extrapolation, allowing the system to generate a dense, continuous series of target points at a very high second frequency. These points together form a smooth path that the robot can follow at high frequency. This step essentially converts low-frequency, sparse perception data into high-frequency, dense control commands.
[0049] S104: Determine a control torque of the self-balancing robot based on the target trajectory, so that the self-balancing robot maintains balance while tracking along the target trajectory.
[0050] Specifically, the control torque is the physical driving force applied to the wheel motors as calculated by the robot controller. This torque directly determines the speed and direction of the wheels.
[0051] The controller continuously compares the robot's current state with the next target point on the high-frequency target trajectory and calculates how much control torque is needed to accurately move to that target point.
[0052] Maintaining balance during tracking is the ultimate goal and constraint. The calculated control torque must simultaneously satisfy two conditions: first, it drives the robot to follow the target trajectory generated in the previous step; second, it maintains the robot's dynamic balance to prevent it from tipping over.
[0053] The tracking method for a self-balancing robot provided in an embodiment of the present application determines the complete motion state of a moving target, including position, velocity, and acceleration, based on time-series pose data. This provides richer and more accurate physical model input for subsequent trajectory prediction and extrapolation, laying the foundation for high-quality trajectory generation. The determined motion state is extrapolated to generate a target trajectory at a second frequency that is significantly higher than the data acquisition frequency (the first frequency). Low-frequency, discrete state input points are converted into a high-frequency, continuous trajectory instruction stream. Therefore, even in the interval between two acquisitions of raw pose data, the control system of the self-balancing robot can continuously obtain smooth and coherent target guidance, fundamentally solving the problem of control command delays and jumps caused by data frequency mismatch. Ultimately, the self-balancing robot determines its control torque based on this high-frequency, smooth target trajectory. Because the reference trajectory used as the control basis is continuous and changes smoothly, the calculated control torque is also correspondingly continuous and stable, avoiding the severe impact on the robot's posture caused by sudden command changes. This enables the robot to accurately track moving targets while effectively maintaining its own dynamic balance, thereby significantly improving the smoothness and stability of the entire tracking process, and solving the problem in existing technologies where self-balancing robots are prone to shaking or even becoming unstable and falling due to discontinuous control.
[0054] In an optional embodiment, obtaining the position data of the moving target includes: The camera mounted on the self-balancing robot detects the visual marker installed on the moving target, calculates the position of the visual marker relative to the camera, and obtains the position data of the moving target.
[0055] By using visual markers as passive beacons, the pose acquisition process is completely independent of external wireless communications (such as Wi-Fi and GPS). This allows for stable operation in communication-restricted or signal-free environments, such as indoors and in tunnels. This reduces the system's reliance on communication hardware, lowering costs and power consumption, and fundamentally avoiding the risk of tracking failures caused by communication delays or interruptions.
[0056] In an optional implementation, determining the motion state of the moving target includes: The time series pose data is used to determine the relative pose of the moving target relative to the self-balancing robot; Based on the relative posture and the self-balancing robot's own posture in the world coordinate system, the time series position data of the moving target in the world coordinate system is obtained through coordinate system conversion; Based on the time series position data, the velocity and acceleration of the moving target are calculated by backward difference to determine the motion state of the moving target.
[0057] Through coordinate transformation, the robot's motion planning and control can be carried out in a unified global coordinate system; and through differential calculation, the system can grasp the dynamic trend (speed and acceleration) of the target, not just its instantaneous position, providing an essential data foundation for subsequent precise prediction and control.
[0058] In an optional embodiment, before extrapolating the motion state, the method further includes: The motion state is smoothed by exponential weighted filtering.
[0059] By introducing exponentially weighted filtering, we can effectively suppress data jitter and glitches caused by sensor measurement noise and differential calculations. This makes the motion state data input into the subsequent extrapolation model smoother and more stable, closer to the target's actual physical motion. As a result, the quality of the generated final target trajectory is significantly improved, avoiding sudden trajectory changes caused by noise interference, laying a solid foundation for smoother robot control.
[0060] In an optional embodiment, the motion state is extrapolated, including: The motion state is input into a preset constant acceleration prediction model for extrapolation to generate a target trajectory; wherein the target trajectory is composed of a series of target points, and the time sampling frequency of the target points is the second frequency.
[0061] In an optional embodiment, the constant acceleration prediction model is extrapolated according to a second-order polynomial, and the expression is: ; ; Where, and They represent the target position and speed of the following vehicle predicted at any control time t; 、 、 Represent the latest position, velocity, and acceleration at the time of the last visual detection update; Represents the current control time t and the last visual update time The time difference between .
[0062] This model not only considers the target's position and velocity, but also incorporates its acceleration into the prediction process. Compared to zero-order (maintaining position) or first-order (constant velocity) predictions, it can more accurately predict the target's motion over short periods of time, especially when the target is accelerating or decelerating. Using this model to generate high-frequency (second-frequency) trajectory points ensures that the generated trajectory is not only high-frequency and dense, but also physically reasonable and accurate.
[0063] In an optional embodiment, periodically acquiring the position and posture data of the mobile target at a first frequency further includes: When the moving target is successfully identified continuously and the linear velocity of the moving target reaches a preset non-zero threshold, the position and posture data of the moving target is periodically acquired at a first frequency.
[0064] This effectively prevents two typical failure scenarios: first, initiating tracking of a nonexistent or incorrect target due to transient sensor misidentification or environmental interference; and second, attempting to track a completely stationary target, which can cause unnecessary jitter or oscillation when the robot fine-tunes in place. Therefore, this feature significantly improves the robustness and safety of the system during startup, ensuring that tracking tasks only begin after the target is confirmed to be a stable and valid dynamic object.
[0065] In an optional embodiment, the method further includes: When the pose data of the moving target is lost continuously for a preset number of frames, or the error between the predicted target trajectory and the actual control state exceeds a preset threshold, the self-balancing robot switches to the preset backup trajectory control mode.
[0066] This prevents the robot from blindly executing a predicted trajectory that has become invalid or has accumulated excessive errors after losing its target, which could easily lead to loss of control or collision. By switching to a backup mode (such as maintaining balance in place or executing a preset safe path), the present invention ensures that the robot remains safe and stable even when visual perception is interrupted, greatly enhancing the reliability and practicality of the method in complex environments.
[0067] In an optional embodiment, determining the control torque of the self-balancing robot includes: The control torque is generated based on the underactuated dynamic constraints and ground reaction force constraints of the self-balancing robot.
[0068] By directly incorporating the robot's inherent physical constraints (such as underactuated characteristics and ground friction) into the calculation of control torques, the system ensures that every generated control command is physically feasible. This avoids invalid commands that exceed the robot's physical limits, which might be issued by traditional simple controllers (for example, requiring the robot to instantly translate laterally). This fundamentally guarantees the effectiveness and safety of control commands, allowing the robot to operate stably within its physical limits even during high-speed, highly maneuverable tracking tasks, maximizing its performance while ensuring its balance and stability.
[0069] In summary, the tracking method for a self-balancing robot provided in the embodiments of the present application first uses a camera to detect visual markers to acquire the relative pose of the target. Combined with the robot's own pose, coordinate transformation and differential calculations are used to obtain the target's complete motion state in a global coordinate system, including position, velocity, and acceleration. This method constructs a data processing link from raw visual information to the complete physical state that is independent of external communication, ensuring that the system can obtain the basic data required for subsequent predictions in any environment. Furthermore, to improve data quality, an exponentially weighted filter is introduced to smooth the motion state, effectively eliminating sensor noise and computational errors, providing a more reliable and smooth data source for subsequent predictions. This is a prerequisite for achieving high-quality tracking. Next, using the smoothed motion state, a constant acceleration prediction model is applied for extrapolation, generating a high-density, continuous, and smooth target trajectory at a second frequency far higher than the data acquisition frequency. This method bridges the gap between low-frequency sensing and high-frequency control, fundamentally resolving the problem of control command jumps caused by data frequency mismatch. Furthermore, to further enhance robustness in practical applications, safety policies are added at both ends of the process. On the front end, a triggering judgment mechanism ensures that tracking is initiated only after the target is confirmed to be a valid dynamic object, avoiding misjudgments and instability in the initial phase. On the back end, a fault-tolerant switching mechanism enables a safe switch to a backup mode in the event of an emergency, such as target loss. This prevents loss of control due to perception interruptions and significantly improves reliability. Finally, this target trajectory is input into a controller that fully accounts for the robot's physical constraints. This ensures that each resulting motor control torque not only accurately drives the robot along the trajectory but also remains well within the robot's physical capabilities.
[0070] For example, a specific example will be used below to specifically illustrate the method of the above embodiment.
[0071] Based on the method of the above embodiment, taking a two-wheeled vehicle as an example, a tracking system for a two-wheeled vehicle is provided.
[0072] With the rapid development of robotics and unmanned driving technologies, two-wheeled vehicles have a wide range of applications in logistics distribution, environmental inspection, and entertainment and competition due to their simple structure, flexible movement, and low energy consumption. In multi-vehicle collaborative scenarios, the autonomous tracking technology of two-wheeled vehicles can be used to achieve functions such as convoy formation and target following, and to perform tasks such as surveillance, reconnaissance, and transportation. In existing technologies, two-wheeled vehicle tracking mainly relies on two methods: communication-based convoy technology [1] and deep learning-based visual detection methods, but both have significant limitations.
[0073] Formation-based tracking technology primarily relies on wireless communication for information exchange. Through communication, the leading vehicle transmits its position, speed, and other status information to the trailing vehicle, which then performs path planning and motion control based on the received data. However, this approach carries the risk of system instability due to communication delays and signal interference, and is completely ineffective in enclosed scenarios such as tunnels. In recent years, deep learning-based target detection algorithms such as YOLO and SSD have performed well in visual tracking and can be used to identify and locate the leading vehicle. However, these methods face challenges in application, including high computational complexity and difficulty running in real time on embedded devices; sensitivity to changes in lighting and object appearance; insufficient detection stability, which can easily cause two-wheeled vehicles to fall due to target loss; and the algorithm requires additional processing to obtain accurate position and motion status, increasing the algorithmic burden.
[0074] Currently, a more stable solution uses a camera to detect the markings of the preceding vehicle, estimates the preceding vehicle's relative position through image processing algorithms, and combines this with control algorithms to achieve tracking. However, due to the structural characteristics of two-wheeled vehicles, they must maintain their own balance while tracking. Furthermore, visual detection alone cannot meet the real-time requirements of high-frequency control for two-wheeled vehicles. To address these shortcomings, this embodiment proposes an autonomous tracking algorithm for two-wheeled vehicles based on visual detection and interpolation optimization. This algorithm achieves efficient and stable estimation of the preceding vehicle's motion state in a non-communication environment, enabling smooth and stable tracking.
[0075] The system workflow of the tracking system for two-wheeled vehicles provided in this embodiment is as follows: Figure 2 As shown in the figure, the state of the preceding vehicle is estimated by combining visual detection with differential estimation. The real-time trajectory information of the preceding vehicle is used in conjunction with interpolation optimization methods to bridge the frequency gap. The distributed trajectory tracking controller is used to execute control commands to achieve multi-vehicle tracking and ensure real-time and stable tracking. The process includes the following: Each two-wheeled vehicle is equipped with an IMU (Inertial Measurement Unit) sensor and an AprilTag (such as the 36h11 series, 10 cm long) mounted on the rear of each vehicle to ensure clear identification under varying lighting conditions and viewing angles. A Realsense camera is mounted on the rear vehicle, fixed to the front of the vehicle at an aligned height with the AprilTag, with the lens facing forward.
[0076] The leading vehicle is the tracked vehicle in the tracking system. It periodically generates its own trajectory and moves continuously. Its position, velocity, and acceleration serve as the detection and reference objects for the following vehicle. To ensure the maneuverability of the two-wheeled vehicle, the drive adopts torque control. The trajectory information includes the center of mass position, velocity, and acceleration of the main vehicle: ; This formula defines the motion state of the leading vehicle (the vehicle being tracked) at any time t. This is a column vector containing the complete motion information of the leading vehicle's center of mass, which is the target of the following vehicle's tracking. r(t) represents the complete trajectory information or state vector of the leading vehicle at time t; q(t) represents the position of the leading vehicle's center of mass at time t; v(t) represents the velocity of the leading vehicle's center of mass at time t; and a(t) represents the acceleration of the leading vehicle's center of mass at time t.
[0077] The vehicle trajectory setting requires full consideration of the structural characteristics of the two-wheeled vehicle to avoid tipping over and loss of control during operation. Therefore, a segmented cubic spline interpolation method is used to generate the center of mass trajectory of the leading vehicle, ensuring that the center of mass trajectory is continuous and differentiable at the position and velocity levels, as well as continuous at the acceleration level.
[0078] The camera on the rear vehicle is pre-calibrated to obtain camera intrinsic parameters and distortion parameters, and then calibrated using the Zhang Zhengyou calibration method. The rear vehicle camera captures images of the front field at a constant frequency, then calls the optimized AprilTag detection algorithm to quickly locate the four corner points of the tag in each grayscale image frame and solve its edge ID. The pre-calibrated camera intrinsic parameters (focal length, principal point, distortion coefficient) are then used to calculate the tag coordinate system through the PnP (Perspective-n-Point) algorithm combined with RANSAC iterative optimization. Relative to the camera coordinate system The homogeneous transformation matrix of : ; This formula represents the homogeneous transformation matrix of the AprilTag coordinate system (T) relative to the following vehicle's camera coordinate system (C). This matrix is obtained using the PnP algorithm combined with RANSAC iterative optimization. It fully describes the AprilTag's 3D spatial position and posture within the camera's field of view.
[0079] Represents the 4x4 homogeneous transformation matrix from the Tag coordinate system (T) to the camera coordinate system (C).
[0080] It is a 3x3 rotation matrix that describes the rotation of the Tag coordinate system relative to the camera coordinate system.
[0081] It is a 3x1 translation vector that describes the three-dimensional position of the origin of the tag coordinate system relative to the origin of the camera coordinate system.
[0082] Add a system timestamp after each solution is completed The trajectory is stored in the cache queue. If multiple consecutive frames fail to detect, a detection failure is triggered. The subsequent logic decides whether to switch back to the backup trajectory to ensure control stability.
[0083] The real-time position of the following vehicle in the world coordinate system is obtained by fusing the IMU and odometer , obtain the static transformation of the camera in the rear vehicle coordinate system through mechanical installation measurement , and the static transformation of Apriltag relative to the preceding vehicle , combined with the AprilTag detection algorithm to obtain the dynamic transformation from camera to AprilTag , using timestamp synchronization to ensure the consistency of received messages and avoid posture conversion errors. Further, by building a coordinate conversion chain to calculate the posture of the front vehicle in the world coordinate system : ; Then, according to the preceding vehicle at time , The speed of the preceding vehicle is estimated by position difference at the continuous sampling points of ; Estimate the acceleration of the preceding vehicle by velocity difference: ; The speed and acceleration of the preceding vehicle are estimated by the backward difference method. Since visual inspection can only obtain the position information of the preceding vehicle and cannot directly obtain the speed and acceleration information, these two quantities are approximately calculated by the position and speed changes of two consecutive sampling points.
[0084] is The estimated speed of the vehicle ahead at all times. is The estimated acceleration of the vehicle ahead at any moment. and The front car is at the current moment and the previous sampling moment The three-dimensional position of . This position information is calculated from the above Extracted from. At the last sampling moment Estimated speed of the vehicle ahead. and is the system timestamp of the current and previous sampling points.
[0085] After obtaining the leading vehicle's 3D position, velocity, and acceleration based on AprilTag visual detection and differential estimation, the following vehicle will not immediately execute the following control mode like other intelligent agents. Instead, it will first perform tracking according to the preset and enter the visual observation waiting state. Tracking will only start when the following conditions are met: AprilTag continuous recognition is successful and the target posture solution is stable; the leading vehicle's linear velocity reaches a non-zero threshold to avoid false locking or misidentification of stationary objects. When the above conditions are met and the following control mode is entered, the target offset of the following vehicle in the leading vehicle's coordinate system is set: ; This formula defines the target expected offset of the following vehicle relative to the leading vehicle and describes the ideal following position of the following vehicle. is the desired relative position offset, is the desired longitudinal (forward) offset, and the next two values represent the target lateral and vertical offsets of the following vehicle in the coordinate system of the leading vehicle, respectively. For example, [-2, 0, 0] means the following vehicle wants to remain 2 meters behind the leading vehicle. A desired lateral and vertical offset of 0 means the following vehicle aims to maintain the same trajectory as the leading vehicle, tracking only longitudinally.
[0086] Combined with the target offset of the following vehicle in the coordinate system of the leading vehicle, the expected tracking trajectory of the following vehicle can be calculated: ; This formula calculates the final desired tracking trajectory of the following vehicle. Combining the real-time state of the leading vehicle and the desired relative offset, it provides the following vehicle controller with a clear motion target to be executed.
[0087] in, represents the expected tracking trajectory of the following vehicle at time t. 、 、 are the position, velocity, and acceleration that the following vehicle expects to achieve. 、v 、a are the real-time position, velocity and acceleration of the preceding vehicle, which are obtained by the previous visual detection and differential estimation. R represents the attitude rotation matrix of the preceding vehicle, which is calculated from the above Extracted from the Convert from the front vehicle coordinate system to the world coordinate system.
[0088] During vehicle motion, it is expected that the motion is smooth and the trajectory transitions smoothly. Considering that the visual detection outputs discrete state samples and there are frequency differences in the actual detection and tracking control systems, directly using the desired trajectory generated above may make real-time tracking difficult. In order to support high-frequency control requirements, compensation for detection frequency and tracking trajectory is required. First, the downsampling rate parameters and the number of detection threads of the Apriltag detection algorithm are modified so that the algorithm detection frequency reaches the maximum detection frequency of the camera. Interpolation scheduling is further used to achieve smoothness in trajectory tracking. After the velocity and acceleration are estimated by backward difference based on the detected data, an exponentially weighted low-pass filter is used at any time to smooth out the fluctuations: ; ; These two formulas are used to smooth the differentially estimated velocity and acceleration to eliminate fluctuations caused by measurement noise and discrete sampling.
[0089] in, and Represent the smoothed velocity and acceleration values after filtering, respectively. and They represent the original velocity and acceleration values estimated by the k-difference at the current moment respectively. and They represent the original velocity and acceleration values estimated by the k-1 difference at the previous moment.
[0090] α represents the smoothing factor, which ranges from 0 to 1. The larger α is, the greater the weight of the historical value at the previous moment, and the smoother the filtering result, but the response will be slower. Conversely, the smaller α is, the greater the weight of the current measurement value, the faster the response, but the poorer the smoothing effect.
[0091] For any control moment, if the visual measurement is not updated or continuous frame loss occurs, the constant acceleration model is extrapolated according to the second-order polynomial: ; ; These two formulas are the key to bridging the gap between low-frequency detection and high-frequency control in the present embodiment. When the high-frequency controller needs instructions but the new visual detection results have not yet arrived, the second-order polynomial based on the constant acceleration kinematic model is used for extrapolation prediction. and They represent the target position and speed of the following vehicle predicted at any control time t. 、 、 Represents the latest position, velocity, and acceleration at the time of the last visual detection update. Represents the current control time t and the last visual update time The time difference between The target position, velocity, and acceleration of the following vehicle are instantly generated for the controller to track smoothly and efficiently, ensuring that the system trajectory input is continuous and directional, and the control response is smooth.
[0092] A multi-task control framework for two-wheeled vehicles is established based on QP. The objective function is set to include spatial trajectory tracking, complete constraint tracking, and non-complete constraint tracking in the task space. The control torque is solved by considering constraints including the underactuated dynamics equation (dynamic constraint, obtained according to the first-kind Lagrange equation), ground reaction force constraints (normal unilateral constraint, tangential friction cone constraint), and upper and lower limit constraints of joint torque / acceleration. The control algorithm is highly versatile and ensures the balance of the vehicle body under various desired trajectories.
[0093] In summary, the tracking system for two-wheeled vehicles provided in this embodiment realizes real-time relative positioning without communication conditions by deploying Apriltag tags at the rear of the two-wheeled vehicle and combining real-time detection of the rear vehicle camera and coordinate transformation chain solution, so that the slave vehicle can accurately obtain the position and posture of the leading vehicle without relying on a wireless link; based on the differential estimation and polynomial interpolation prediction of discrete detection data, it bridges the time delay gap between low-frequency detection and high-frequency control, thereby avoiding tracking instability caused by low-frequency detection; in the system startup phase and tracking process, preset path driving and out-of-detection fault-tolerant switching strategies are introduced respectively to ensure the balance and stability of the two-wheeled vehicle in the initial stage and abnormal state; finally, combined with high-speed WBC closed-loop control, the output joint torque is calculated to implement driving, so that the slave vehicle can accurately follow the position, speed and acceleration of the leading vehicle, which not only reduces the calculation burden of the master vehicle, but also has high robustness, high real-time performance and low communication overhead.
[0094] This embodiment also provides a tracking method and apparatus for a self-balancing robot. This apparatus is used to implement the above-mentioned embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0095] This embodiment provides a tracking method and device for a self-balancing robot, such as Figure 3 As shown, the device is executed by a self-balancing robot as a slave vehicle, and the device is used to track a moving target as a master vehicle. The device includes: An acquisition module 301 is configured to periodically acquire position and posture data of a mobile target at a first frequency to obtain time-series position and posture data of the mobile target; The state confirmation module 302 is used to determine the motion state of the mobile target based on the time series posture data; the motion state includes position, velocity and acceleration; The trajectory generation module 303 is used to extrapolate the motion state and generate a target trajectory of the self-balancing robot at a second frequency; the second frequency is higher than the first frequency; The control module 304 is configured to determine a control torque of the self-balancing robot based on the target trajectory, so that the self-balancing robot maintains balance while tracking the target trajectory.
[0096] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0097] The tracking method device for the self-balancing robot in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0098] The present application also provides a computer device having the above Figure 3 The tracking method and apparatus for a self-balancing robot is shown.
[0099] See also Figure 4 , Figure 4 This is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present application. Figure 4 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 A processor 10 is taken as an example.
[0100] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0101] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0102] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0103] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0104] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0105] The embodiments of the present application also provide a computer-readable storage medium. The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0106] Part of the present application may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0107] Although the embodiments of the present application have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the appended claims.
Claims
1. A tracking method for a self-balancing robot, characterized in that: The method is performed by a self-balancing robot as a slave vehicle, and is used to track a moving target as a master vehicle. The method includes: Periodically acquiring position and posture data of a mobile target at a first frequency to obtain time-series position and posture data of the mobile target; Determining the motion state of the mobile target based on the time-series posture data; the motion state includes position, velocity and acceleration; Extrapolating the motion state to generate a target trajectory of the self-balancing robot at a second frequency; the second frequency is higher than the first frequency; Based on the target trajectory, a control torque of the self-balancing robot is determined so that the self-balancing robot maintains balance while tracking along the target trajectory.
2. The method according to claim 1, characterized in that Get the pose data of the moving target, including: The camera carried by the self-balancing robot detects the visual marker installed on the moving target, calculates the position and posture of the visual marker relative to the camera, and obtains the position and posture data of the moving target.
3. The method according to claim 2, characterized in that Determining the motion state of the moving target includes: Determine the relative position of the moving target relative to the self-balancing robot using the time series position data; Based on the relative posture and the self-balancing robot's own posture in the world coordinate system, the time series position data of the moving target in the world coordinate system is obtained through coordinate system conversion; Based on the time series position data, the velocity and acceleration of the moving target are obtained by backward difference calculation to determine the motion state of the moving target.
4. The method according to claim 3, characterized in that Before extrapolating the motion state, the method further includes: The motion state is smoothed by exponential weighted filtering.
5. The method according to claim 4, characterized in that Extrapolating the motion state includes: The motion state is input into a preset constant acceleration prediction model for extrapolation to generate a target trajectory; wherein the target trajectory is composed of a series of target points, and the time sampling frequency of the target points is the second frequency.
6. The method according to claim 5, characterized in that The constant acceleration prediction model is extrapolated according to a second-order polynomial, and the expression is: ; ; Where, and They represent the target position and speed of the following vehicle predicted at any control time t; 、 、 Represent the latest position, velocity, and acceleration at the time of the last visual detection update; Represents the current control time t and the last visual update time The time difference between .
7. The method according to any one of claims 1 to 6, characterized in that The method of periodically acquiring the position and posture data of the mobile target at the first frequency further includes: When the moving target is successfully identified continuously and the linear velocity of the moving target reaches a preset non-zero threshold, the position and posture data of the moving target is periodically acquired at a first frequency.
8. The method according to claim 7, characterized in that The method further comprises: When the position data of the mobile target are lost continuously for a preset number of frames, or the error between the predicted target trajectory and the actual control state exceeds a preset threshold, the self-balancing robot switches to a preset backup trajectory control mode.
9. The method according to claim 8, characterized in that Determining the control torque of the self-balancing robot includes: A control torque is generated based on the underactuated dynamics constraints and the ground reaction force constraints of the self-balancing robot.
10. A tracking device for a self-balancing robot, characterized in that: The device is executed by a self-balancing robot as a slave vehicle, and is used to track a moving target as a master vehicle. The device includes: An acquisition module, configured to periodically acquire the position and posture data of the mobile target at a first frequency to obtain time-series position and posture data of the mobile target; A state confirmation module is used to determine the motion state of the mobile target based on the time-series posture data; the motion state includes position, velocity and acceleration; a trajectory generation module, configured to extrapolate the motion state to generate a target trajectory of the self-balancing robot at a second frequency; the second frequency being higher than the first frequency; The control module is used to determine the control torque of the self-balancing robot based on the target trajectory, so that the self-balancing robot maintains balance when tracking along the target trajectory.
11. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the tracking method for a self-balancing robot according to any one of claims 1 to 9 when executing the computer program.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the tracking method for a self-balancing robot according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Multi-target tracking method based on SRCK-GMCPHD filtering
CN106372646A
Robot following method and device based on pedestrian re-identification and mobile robot
CN112989983A
Robot, robot following method and device and storage medium
CN116661505A
Aircraft trajectory prediction method and device, electronic equipment and storage medium
CN118410655A
Unmanned aerial vehicle visual tracking method and system based on neural network
CN120219997A