Foot type robot control method and device, robot and medium

By integrating a movable UWB beacon and multimodal sensors onto a legged robot, a dual-mode switching between automatic following and autonomous recharging is achieved. This solves the problems of separation of recharging and accompanying functions and susceptibility of a single sensor to environmental interference in existing technologies, thereby improving positioning accuracy and system reliability.

CN121411264APending Publication Date: 2026-01-27INTELLIGENT BODY TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511574088.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing legged robots often have separate recharging and accompanying functions, and rely on a single sensor, making them susceptible to environmental interference and resulting in unstable positioning performance. In particular, they are prone to losing their target in complex environments, affecting the reliability and safety of the system.

Method used

By employing a mobile UWB beacon and combining multimodal data fusion from UWB, vision, and LiDAR, the system enables switching between automatic following and autonomous recharging modes. It acquires positioning data in multiple modes through UWB positioning, visual positioning, and radar ranging, and performs fusion processing to generate a following control strategy.

Benefits of technology

It improves positioning accuracy and following performance in complex environments, enhances the reliability and safety of robot control, reduces hardware costs, and improves scene adaptability and human-computer interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121411264A_ABST
    Figure CN121411264A_ABST
Patent Text Reader

Abstract

The invention provides a foot type robot control method and device, a robot and a medium. The foot type robot is provided with an ultra wide band (UWB) receiving part, a laser radar and a camera. The UWB receiving component can receive UWB signals sent by a UWB beacon, the UWB beacon is a movable UWB beacon and has two working modes including an automatic following mode and an autonomous recharging mode, and the method comprises the steps that when the UWB beacon is in the automatic following mode, the UWB beacon is in the autonomous recharging mode, and the UWB beacon is in the autonomous recharging mode; a UWB signal is received through a UWB receiving part to carry out UWB positioning, visual positioning is carried out through a camera, radar ranging is carried out through a laser radar, and positioning data in multiple modes are determined; performing fusion processing on the positioning data in the multiple modes to obtain fused position data, and generating a following control strategy based on the fused position data; according to the following control strategy, motion control and path planning are carried out on the foot-type robot, and therefore the movable UWB beacon capable of being switched between the two modes is utilized, and autonomous following is achieved through multi-mode data fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of robotics, and more specifically, to a control method, apparatus, robot, and medium for a legged robot. Background Technology

[0002] In recent years, with the rapid development of robotics technology, legged robots have shown broad application prospects in disaster relief, logistics and transportation due to their excellent terrain adaptability and flexibility. Among them, the recharging function and the following function are important capabilities of quadruped robots. The recharging function enables the robot to automatically go to the charging station, and the following function enables the robot to follow a specific target (such as an operator or other moving object) in real time and maintain a stable relative position.

[0003] However, the above functions are often separate and their implementation mostly relies on a single sensor, making them susceptible to environmental interference, which affects the positioning effect and leads to the loss of the target being followed. Summary of the Invention

[0004] This disclosure provides at least one legged robot control method, device, robot, and medium, which achieves dual-mode switching of automatic following and autonomous recharging through a movable UWB beacon, and improves the following effect by employing multimodal fusion.

[0005] Specifically, this application is implemented through the following technical solution: According to a first aspect of this application, a control method for a legged robot is provided. The legged robot is equipped with an Ultra Wide Band (UWB) receiver, a lidar, and a camera. The UWB receiver is capable of receiving UWB signals transmitted from a UWB beacon, which is a movable UWB beacon with two operating modes: an automatic following mode and an autonomous recharging mode. The method includes: When the UWB beacon is in automatic following mode, the UWB receiver receives UWB signals for UWB positioning, the camera performs visual positioning, and the lidar performs radar ranging to determine positioning data in multiple modes; the UWB beacon is carried by the target object followed by the legged robot in automatic following mode. The positioning data under the multiple modes are fused to obtain fused position data, and a following control strategy is generated based on the fused position data; According to the following control strategy, motion control and path planning are performed on the legged robot.

[0006] The mobile UWB beacon used in this solution has two operating modes: automatic following mode and autonomous recharging mode, which can be switched between. Users do not need to carry additional traction equipment when using the legged robot, reducing hardware costs and facilitating flexible application in multiple scenarios, improving scene adaptability and human-machine interaction efficiency. In automatic following mode, the UWB beacon accurately identifies targets and achieves dynamic following through the fusion of multi-modal data from UWB, vision, and LiDAR; it can adapt to complex terrain and environmental changes, enhancing positioning accuracy, improving obstacle avoidance capabilities, improving following performance, and ensuring the reliability and safety of robot control.

[0007] In one optional implementation, the method further includes: When the UWB beacon is in autonomous recharging mode, positioning data via UWB positioning is acquired; the UWB beacon is fixed to the charging device that charges the legged robot in the autonomous recharging mode. Based on the positioning data obtained from the UWB positioning, the legged robot is controlled to move to the charging range of the charging device, and the positioning is assisted by other sensors and the QR code on the charging device to complete the docking of the charging port.

[0008] In this solution, under autonomous recharging mode, a UWB beacon is fixed to the charging device. UWB can be directly used to match the legged robot with the charging device, effectively establishing an index connection between the two. UWB positioning enables the legged robot to quickly and coarsely align with the charging area, significantly reducing the lateral error range. Combined with other sensors and the QR code on the charging device for fine alignment, accurate docking of the charging port is completed. This effectively improves the positioning efficiency and docking success rate of the recharging process, reduces the complexity of operation steps, and enhances the reliability and stability of autonomous recharging.

[0009] In one optional implementation, the legged robot is further equipped with an inertial measurement unit (IMU) and a joint encoder; the method further includes: The joint encoder measures the leg joint movement of the legged robot, and the IMU obtains the angular velocity and acceleration of the legged robot. The body pose information of the legged robot is determined by fusing the output results of the IMU and the joint encoder through Kalman filtering. The process of fusing the positioning data under the multiple modalities to obtain fused location data includes: Based on the body pose information of the legged robot, the positioning data under the multiple modalities are fused to obtain the fused position data.

[0010] In this scheme, the leg movements, body angular velocity, and acceleration of the legged robot are acquired in real time through joint encoders and an IMU, and then fused using Kalman filtering to obtain the body pose information. Combining the body pose information with the multimodal positioning data for fusion processing effectively eliminates errors caused by body jitter and motion distortion, achieving efficient coupling between external observation and internal motion state. This helps improve the stability, continuity, and dynamic accuracy of the fused position data, ensuring that the legged robot maintains accurate and smooth following performance even under high-speed walking or sudden terrain changes.

[0011] In one optional implementation, receiving UWB signals through the UWB receiving component for UWB positioning includes: Based on the time-of-flight (TOF) of the UWB signal transmitted between the UWB beacon and the UWB receiving component, as well as the air humidity compensation and temperature compensation, the distance and azimuth of the legged robot relative to the UWB beacon are determined. The air humidity compensation and temperature compensation are used to compensate for the influence of air humidity and temperature on the speed of light, and the speed of light is used to calculate the distance and azimuth angle in conjunction with the TOF.

[0012] In this scheme, the speed of light is corrected in real time by introducing air humidity compensation and temperature compensation. Combined with the flight time of the UWB signal, the relative distance and azimuth between the legged robot and the UWB beacon are determined. This effectively eliminates the impact of temperature and humidity changes on the signal propagation speed, improves the accuracy and stability of UWB positioning, and ensures that reliable UWB positioning data can still be obtained under complex environmental conditions, thereby enhancing the environmental adaptability and robustness of the overall positioning.

[0013] In one optional implementation, the step of fusing the positioning data from the multiple modalities to obtain fused location data includes: An attention module is used to preprocess UWB positioning data, visual positioning data, and radar ranging data to obtain preprocessed multimodal data. During the preprocessing, attention weights are assigned according to the data validity of each modality to prioritize the more reliable modal data. The preprocessed multimodal data are labeled and encoded, and then mapped to the same spatial coordinate system; The Transformer model is used to fuse multiple modal data in the same spatial coordinate system after label encoding to determine the fused location data.

[0014] In this scheme, the attention module dynamically allocates weights based on the validity of each modality's data, prioritizing reliable information. Then, through label encoding, the multi-source data is uniformly mapped to the same spatial coordinate system, and deep fusion is achieved with the help of the Transformer model. When the localization data of a certain modality is disturbed, other modal localization data are used to supplement it in time, forming a complementary fusion mechanism. This significantly improves the continuity, robustness, and accuracy of the localization results, ensuring that the legged robot can maintain stable and accurate following performance in complex environments.

[0015] In one optional implementation, the step of fusing multiple modal data in the same spatial coordinate system after label encoding using a Transformer model to determine the fused location data includes: The Transformer model is used to fuse multiple modal data in the same spatial coordinate system after label encoding to obtain a multimodal fusion result; the multimodal fusion result indicates the relative position information of the target object relative to the legged robot, which is output multiple times in succession; The Kalman filter algorithm is used to smooth the multimodal fusion results to obtain the final fused position data.

[0016] In this scheme, the Transformer model is used to perform deep temporal fusion on the multimodal data after unifying the coordinates to obtain the continuous relative position information of the target object relative to the legged robot. Then, Kalman filtering is used to smooth the sequence, effectively suppressing abrupt noise and instantaneous jumps, and outputting stable, coherent and highly reliable fused position data, thereby ensuring the stability and accuracy of the following control in dynamic scenarios.

[0017] In one optional implementation, the step of performing motion control and path planning on the legged robot according to the following control strategy includes: If, during the fusion processing of the positioning data under the multiple modalities, it is determined that there is missing detection data from the target sensor, the legged robot is controlled to turn so that the target object returns to the detection range of the target sensor.

[0018] In this solution, when a target sensor data is found to be missing during the fusion process, a turning command can be triggered in the following control strategy to enable the legged robot to actively adjust its own orientation and quickly return the target object to the detection range of the sensor. This maintains the integrity of multimodal information, avoids the decrease in positioning accuracy due to the failure of a single sensor, and ensures that the following task is executed continuously, smoothly, and reliably.

[0019] According to a second aspect of this application, a legged robot control device is provided, wherein the legged robot is equipped with an ultra-wideband (UWB) receiver, a lidar, and a camera; the UWB receiver is capable of receiving UWB signals transmitted by a UWB beacon, the UWB beacon being a movable UWB beacon having two operating modes: an automatic following mode and an autonomous recharging mode; the device includes: The positioning module is used to receive UWB signals through the UWB receiving component for UWB positioning, perform visual positioning through the camera, and perform radar ranging through the lidar when the UWB beacon is in automatic following mode, thereby determining positioning data in multiple modes; the UWB beacon is carried by the target object followed by the legged robot in the automatic following mode. The fusion module is used to fuse the positioning data under the multiple modes to obtain fused position data, and generate a following control strategy based on the fused position data; The control module is used to perform motion control and path planning for the legged robot according to the following control strategy.

[0020] According to a third aspect of this application, a legged robot is provided, which is equipped with an ultra-wideband (UWB) receiver, a lidar, a camera, and a controller; the UWB receiver is capable of receiving UWB signals transmitted by a UWB beacon, which is a movable UWB beacon and has two working modes: an automatic following mode and an autonomous recharging mode. The controller is used to receive UWB signals through the UWB receiving component for UWB positioning, perform visual positioning through the camera, and perform radar ranging through the LiDAR when the UWB beacon is in automatic following mode, thereby determining positioning data under multiple modalities. The UWB beacon is carried by the target object followed by the legged robot in automatic following mode. The positioning data under the multiple modalities are fused to obtain fused position data, and a following control strategy is generated based on the fused position data. According to the following control strategy, the legged robot is subjected to motion control and path planning.

[0021] According to a fourth aspect of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, performs the steps of the legged robot control method described in the first aspect.

[0022] For a description of the effects of the aforementioned legged robot control device, legged robot, and computer-readable storage medium, please refer to the description of the aforementioned legged robot control method; it will not be repeated here.

[0023] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure.

[0024] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0026] Figure 1 A flowchart of a legged robot control method provided in an embodiment of this disclosure is shown; Figure 2 A schematic diagram of a legged robot provided in an embodiment of this disclosure is shown; Figure 3 A schematic diagram of coarse alignment of a legged robot in autonomous recharging mode provided by an embodiment of this disclosure is shown; Figure 4 A schematic diagram of a legged robot in autonomous recharging mode provided by an embodiment of the present disclosure is shown. Figure 5 A schematic diagram of a legged robot charging in autonomous recharging mode, provided in an embodiment of this disclosure, is shown. Figure 6 This illustration shows a data fusion and strategy generation process provided by an embodiment of the present disclosure; Figure 7 One of the schematic diagrams of a legged robot control device provided in an embodiment of this disclosure is shown; Figure 8 A second schematic diagram of a legged robot control device provided in an embodiment of this disclosure is shown; Figure 9 An architectural diagram of a legged robot provided in an embodiment of this disclosure is shown. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0028] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0029] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0030] Furthermore, the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein.

[0031] Research has revealed that the recharging and companion functions of legged robots are often implemented separately. In the recharging function, beacons are typically fixed to base stations, failing to adapt to dynamic scenarios. In the companion function, users often need to carry additional equipment, making operation cumbersome and resulting in a fragmented user experience. Furthermore, both recharging and companion functions currently rely heavily on single sensors, making them susceptible to environmental interference and affecting positioning accuracy. In environments with strong light, low light, dust, or obstruction by grass, they are prone to losing their target, leading to positioning drift, docking failure, or collisions with obstacles. Especially in scenarios with frequent dynamic obstacles, system response lag and path planning failures occur frequently, severely restricting the reliability and safety of legged robots in real-world task environments.

[0032] Based on the above research, this disclosure provides a legged robot control method that uses a movable UWB beacon and has two working modes: automatic following mode and autonomous recharging mode. It can switch between the two working modes to improve scene adaptability and human-computer interaction efficiency. When the UWB beacon is in automatic following mode, it can accurately identify targets and improve the following effect by fusing multiple modal data from UWB, vision, and LiDAR.

[0033] To facilitate understanding of this embodiment, a detailed description of a legged robot control method disclosed in this disclosure will be provided first. The executor of the legged robot control method provided in this disclosure is generally a legged robot, such as a robot dog or robot cat, which is also referred to as a quadruped robot. It should be noted that within the scope of protection of this disclosure, the number of legs is not limited. For the purpose of use, fewer or more legs can be configured, such as two legs, six legs, eight legs, etc.

[0034] In other embodiments, the executing entity may also be a legged robot controller, which may be a device including a processor and a memory, and is not limited thereto. In some embodiments, the legged robot control method can be implemented by the processor calling computer-readable instructions stored in the memory.

[0035] Furthermore, this legged robot control method can also be applied to implementation environments consisting of a legged robot and a server, or to implementation environments consisting of a legged robot controller and a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms.

[0036] The following description, in conjunction with the accompanying drawings, illustrates a legged robot control method provided in an embodiment of this application.

[0037] See Figure 1 The diagram shown is a flowchart of a legged robot control method provided in an embodiment of this disclosure. The legged robot is equipped with an ultra-wideband (UWB) receiver, a lidar, and a camera. The UWB receiver can receive UWB signals transmitted from a UWB beacon, which is a movable UWB beacon with two operating modes: automatic following mode and autonomous recharging mode. Figure 1 As shown in the figure, the legged robot control method provided in this embodiment includes steps S101 to S103, wherein: S101: When the UWB beacon is in automatic following mode, the UWB signal is received by the UWB receiving component for UWB positioning, the camera is used for visual positioning, and the lidar is used for radar ranging to determine positioning data in multiple modes; the UWB beacon is carried by the target object followed by the legged robot in the automatic following mode.

[0038] Here, for a better understanding of the embodiments of this disclosure, exemplary references can be made to... Figure 2 This is a schematic diagram of a legged robot provided in an embodiment of this disclosure. Figure 2 As shown, when setting up the camera and lidar, they can be configured separately on the head and chest of the legged robot 200, such as... Figure 2 As shown in 201, when setting up the UWB receiver, it can be set on the back of the legged robot 200. This forms an orderly multimodal sensor layout, which allows the UWB signal reception to avoid being blocked by the robot body and maintain continuous pointing to the UWB beacon. At the same time, the sensors on the head and chest can ensure the field of vision and improve the positioning effect.

[0039] In the automatic following mode, the target object typically moves while holding the UWB beacon, and the legged robot can achieve dynamic following with the help of the UWB beacon.

[0040] When receiving UWB signals through the UWB receiving component for UWB positioning, the straight-line distance between the legged robot and the UWB beacon can be obtained by the flight time of the transmitted UWB signal. Then, the azimuth angle is calculated by combining multi-beacon cooperation or single-beacon attitude assistance, and finally the relative pose relationship between the legged robot and the UWB beacon is output.

[0041] Specifically, the distance and azimuth of the legged robot relative to the UWB beacon can be determined based on the time of flight (TOF) of the UWB signal transmitted between the UWB beacon and the UWB receiving component, as well as the air humidity compensation and temperature compensation. The air humidity compensation and temperature compensation are used to compensate for the influence of air humidity and temperature on the speed of light, and the speed of light is used to calculate the distance and azimuth in conjunction with the TOF.

[0042] Here, the UWB beacon can have a built-in UWB radio frequency module that can actively and periodically transmit UWB signals, for example, 100-500 times per second. The signal transmission frequency can be adjusted according to the positioning accuracy requirements, and no specific limitation is made here.

[0043] In practical applications, the TOF algorithm requires high time measurement accuracy, generally reaching the nanosecond level. Clock deviation will directly lead to distance calculation error. Therefore, the UWB beacon can also integrate a high-precision clock module to ensure the stability of signal transmission time.

[0044] The UWB receiving component includes at least two UWB receiving modules that are matched with the UWB radio frequency module and can accurately capture the UWB signal transmitted by the beacon. For example, it may include two to three UWB receiving modules for azimuth angle calculation.

[0045] Optionally, the legged robot can also integrate a high-precision clock module, and synchronize its time with the high-precision clock module integrated with the UWB beacon. This can be achieved by having the UWB beacon first transmit a synchronization signal, which the legged robot then receives and calibrates its local clock, thus avoiding time measurement deviations caused by clock asynchrony.

[0046] The legged robot can also be equipped with a signal processing unit, which can quickly calculate the transmission-reception time difference of UWB signals and provide raw time data for the TOF algorithm.

[0047] When determining the distance between the legged robot and the UWB beacon based on the Time-of-Flight (TOF) algorithm, specifically, this can be achieved by measuring the time difference between the UWB signal transmitted from the UWB beacon and the reception by the legged robot, combined with the speed of electromagnetic wave propagation in air (approximately 3 × 10⁻⁶). 8 The relative distance between the legged robot and the UWB beacon is determined by the distance (m / s). In specific implementations, this can be divided into one-way Time-of-Flight (TOF) and two-way Time-of-Flight (TOF). Since the UWB beacon in this embodiment is movable and requires dynamic positioning, two-way TOF can be used to assist in improving accuracy.

[0048] For example, the UWB beacon actively transmits a UWB positioning signal at time T1, and simultaneously records its own transmission time T1. The UWB positioning signal may carry information such as the UWB beacon's identity document (ID) and transmission timestamp T1. The legged robot's UWB receiving component captures this signal at time T2, records its local reception time T2, and immediately generates a UWB response signal. This UWB response signal carries the legged robot's ID and reception timestamp T2, and is transmitted back to the UWB beacon at time T3. The UWB beacon receives the UWB response signal from the legged robot at time T4 and records its local reception time T4. Therefore, the total round-trip time ΔT can be determined as ΔT = (T4 - T1) - (T3 - T2), where (T4 - T1) is the total time taken for the UWB beacon to transmit the UWB positioning signal and receive the UWB response signal, and (T3 - T2) is the time taken for the legged robot to transmit the UWB response signal from receiving the UWB positioning signal, i.e., the processing delay time. Hardware calibration and subtraction of this time help avoid the delay affecting the accuracy of the time difference. Next, the one-way flight time of the UWB signal can be determined as t = ΔT / 2. Since the round-trip distance is the same, the one-way time is half of the total round-trip time. Substituting the one-way flight time into the formula "distance d = speed of light c × one-way flight time t", the distance between the legged robot and the UWB beacon can be obtained. For example, if the measured one-way time t = 10 ns (1 ns = 10... -9 (seconds), then d = 3 × 10 8 m / s×10×10 -9 s=3m, which matches the UWB ranging range of 0.1-30m in practical applications.

[0049] Optionally, the average value can be obtained by taking multiple measurements, such as 10 consecutive measurements, removing outliers, and then calculating the average. The influence of air humidity and temperature on the speed of light can be compensated according to the air humidity compensation and temperature compensation. For example, for every 1°C change in temperature, the speed of light deviates by about 0.006 m / s. This can be corrected in real time by the environmental sensors (including temperature sensors and humidity sensors) set on the legged robot, ultimately achieving a distance measurement accuracy of less than 10 cm and improving the accuracy of UWB positioning.

[0050] When determining the azimuth angle of the legged robot relative to the UWB beacon, the azimuth angle refers to the horizontal angle of the UWB beacon relative to the legged robot in the robot's coordinate system. For example, 0° represents directly in front, and 90° represents directly to the right. This cannot be directly calculated using a single UWB receiver module; therefore, it requires the collaboration of multiple UWB receiver modules.

[0051] Here, the UWB receiving modules can be positioned on either side of the legged robot's head or torso, with two UWB receiving modules (hereinafter referred to as module A and module B for ease of understanding) deployed at a fixed distance L. This ensures that the lines connecting the two UWB receiving modules and the UWB beacon form a triangle, allowing the azimuth angle to be calculated using the Time Difference of Arrival (TDOA). The fixed distance L is, for example, 10cm, and its specific value can be adjusted according to positioning accuracy requirements; no specific limitation is made here. Optionally, the relative positions of the two UWB modules can be pre-calibrated, for example, module A at +5cm on the X-axis of the legged robot's coordinate system and module B at -5cm on the X-axis, and recorded in the legged robot's positioning algorithm as the basic parameters for azimuth angle calculation.

[0052] For the same UWB signal emitted by the UWB beacon, module A receives it at time TA, and module B receives it at time TB. The arrival time difference between the two modules can be determined as Δt = |TA - TB|. Since the distance L between the two modules is known, the distance difference between the UWB signals reaching the two modules is Δd = c × Δt, where c is the speed of light. A coordinate system is established with the center of the legged robot as the origin O, with the X-axis pointing forward and the Y-axis pointing to the right. The position of the UWB beacon is set as point P. Modules A and B can form an isosceles triangle OA = OB = L / 2 with the origin O. According to the geometric relationship cosθ = Δd / L, where θ is the horizontal deflection angle of the UWB beacon relative to the X-axis of the legged robot's coordinate system, i.e., the azimuth angle. Therefore, the azimuth angle θ = arccos(Δd / L) can be determined. For example, if L = 10cm, Δt = 100ps (1ps = 10...). -12 (seconds), then Δd = 3 × 10 8 m / s×100×10 -12 s=0.03m=3cm, cosθ=3cm / 10cm=0.3, θ≈72.5°, that is, the UWB beacon is located about 17.5° to the right and slightly forward of the legged robot.

[0053] In this way, by introducing air humidity compensation and temperature compensation to correct the speed of light in real time, and combining the flight time of the UWB signal, the relative distance and azimuth between the legged robot and the UWB beacon can be determined. This effectively eliminates the impact of temperature and humidity changes on the signal propagation speed, improves the accuracy and stability of UWB positioning, and ensures that reliable UWB positioning data can still be obtained under complex environmental conditions, thereby enhancing the environmental adaptability and robustness of the overall positioning.

[0054] When performing visual positioning using the camera, specifically, the camera on the legged robot can be used to identify the feature markers of the target object and determine the relative positional relationship between the legged robot and the target object. This mode can promptly compensate for positioning deviations caused by the complex electromagnetic environment affecting the positioning quality of UWB signals and / or radar point clouds, thereby improving the positioning effect.

[0055] Here, the UWB beacon is pre-paired with the legged robot. The camera on the legged robot can be an RGB-D camera, which can simultaneously output a color image (RGB channel) and a depth image (Depth channel). The color image is used to identify the feature markers of the target object, and the depth image is used to determine the vertical distance between the UWB marker and the camera. A pre-trained feature marker detection model, such as a lightweight object detection model based on YOLOv8, can detect the position of the feature markers in real time at a frame rate ≥15fps and output the pixel coordinate bounding box of the marker in the image.

[0056] Based on the output of the RGB-D camera, the relative positional relationship between the legged robot and the target object can be obtained through three stages: image feature determination, depth and distance measurement, and coordinate system transformation. This includes relative distance, relative angle, and relative height.

[0057] In specific implementation, in the first stage, image feature localization can be performed to determine the pixel position of the UWB marker in the camera's field of view. Marker detection and bounding box selection can be performed. The color image from the RGB-D camera is input into the legged robot's vision processing module. Through a pre-trained detection model, the feature markers on the UWB beacon are identified and bounded, and the pixel coordinate range of the UWB marker in the image is output. For example, the pixel coordinate range can be defined by the top-left pixel (x1, y1) and the bottom-right pixel (x2, y2). Then, the marker's center point can be determined by selecting the geometric center of the pixel coordinate frame (x0, y0) = [(x1+x2) / 2, (y1+y2) / 2]. This center point represents the visual center of the feature marker, and the relative angle can be calculated based on this point subsequently. If objects with similar characteristics to UWB beacons exist in the environment, such as black and white checkered floor tiles, false detections can be filtered out through scale verification of the marker (e.g., the detected bounding box size must match the preset "20cm×20cm" ratio with a deviation of ≤10%) and dynamic tracking (e.g., the displacement of the marker's center point in two consecutive frames of images is ≤5 pixels), ensuring that the located object is a real UWB beacon and reducing environmental interference.

[0058] In the second stage, depth distance measurement can be performed to obtain the straight-line distance between the UWB marker and the legged robot. Specifically, depth value extraction can be performed by mapping the depth image of the RGB-D camera to the pixels of the color image one by one. Based on the pixel coordinates (x0, y0) of the marker center point obtained in the first stage, the depth value z (in meters) of that pixel in the depth image can be directly read, thus obtaining the vertical straight-line distance between the feature marker and the camera.

[0059] In practical applications, depth images often contain noise in edge areas or low-light environments. To improve positioning accuracy, distance correction can be performed.

[0060] Optionally, distance correction can be performed by averaging the regions. This involves taking the depth values ​​of all pixels within the coordinate frame of the identified pixel, removing outliers that are outside the UWB ranging range (e.g., 0.1-30m), and then averaging the values ​​to obtain a more stable depth value z_avg.

[0061] Alternatively, distance correction can be performed through scale conversion verification. If the actual size of the feature marker is known (e.g., the side length is 20cm), the distance can be calculated by comparing the pixel size of the marker in the image with the actual size using the ratio of distance = actual size × camera focal length / image pixel size. This distance is then cross-validated with the depth value z_avg. If the deviation is less than or equal to a preset threshold (e.g., 5%), z_avg is used; otherwise, the scale conversion result is used to ensure that the distance measurement accuracy is ≤5cm and to reduce positioning errors.

[0062] In the third stage, coordinate system transformation can be performed to output relative positional relationships. Specifically, a three-dimensional coordinate system can be established with the center of the legged robot as the origin. This coordinate system can be defined as follows: X-axis forward (the direction of movement of the legged robot), Y-axis to the right (the right side of the legged robot), and Z-axis upward (vertical to the ground). Based on this, relative angles can be determined, including the horizontal angle θ and the vertical angle φ. Regarding the determination of the horizontal angle θ, in the camera's image coordinate system, the horizontal offset of the center point (x0, y0) is related to the image's horizontal resolution. The corresponding expression is θ = arctan[(x0 - W / 2) × tan(FOV_h / 2) / (W / 2)], where W represents the camera's horizontal resolution (e.g., 1920 pixels), FOV_h represents the camera's horizontal field of view (e.g., 60°), θ is positive when the UWB marker is on the right side of the legged robot, and negative when the UWB marker is on the left side of the legged robot, with an accuracy ≤ 1°. The vertical angle φ can be determined by the center point (y0-H / 2) (where H is the camera's vertical resolution, such as 1080 pixels) and the vertical field of view (FOV_v). When φ is positive, it means that the UWB beacon is above the camera, for example, when the target object holds the UWB beacon high. When φ is negative, it means that the UWB beacon is below the camera, for example, when the target object bends down to hold the UWB beacon. This can compensate for the positional deviation caused by changes in the height at which the target object holds the beacon. This allows us to obtain the relative 3D coordinate output. Combined with the corrected depth distance z_avg, the distance-angle is converted into a 3D relative position (x, y, z) in the legged robot's coordinate system. Here, x = z_avg × cosφ × cosθ (the distance along the X-axis, positive for front and negative for back), y = z_avg × cosφ × sinθ (the distance along the Y-axis, positive for right and negative for left), and z = z_avg × sinφ (the height difference along the Z-axis, positive for the marker being higher than the robot and negative for lower). The final output (x, y, z) represents the relative position between the legged robot and the target object. For example, (x = 1.5m, y = 0.3m, z = 0.2m) indicates that the target object is 1.5m in front of the legged robot and 0.3m to its right, and the UWB beacon is 0.2m higher than the legged robot.

[0063] When performing radar ranging using the aforementioned lidar, specifically, target features (such as leg features of a target object at close range) can be extracted from the lidar point cloud. When the distance is too close, causing visual perception to be unable to fully acquire the body features of the target person, or when the target leaves the effective field of view (FoV) of the camera, this modality can supplement the target distance information, avoiding positioning interruption due to the failure of a single visual modality. Thus, in scenarios where visual perception fails (e.g., incomplete features due to close distance, or the target leaving the camera's FoV field of view), target features can be extracted from the lidar point cloud to supplement the target distance information. By utilizing the lidar's characteristics of being unaffected by lighting / occlusion and being able to directly acquire 3D point cloud structures, the extraction and matching of key human features (such as legs) at close range can be achieved, enabling accurate compensation of distance information.

[0064] In practical applications, legged robots are typically positioned close to their targets, usually within 3 meters for close-range compensation. Therefore, it's crucial to first define the hardware deployment and point cloud preprocessing rules for the LiDAR to ensure data validity. For LiDAR hardware deployment, it's advisable to place the LiDAR at the head of the legged robot to ensure the scanning range covers the leg area of ​​the target object when it's standing or moving at close range. This avoids missing leg areas due to excessive height and interference from ground debris due to insufficient height. Alternatively, a short-range, high-precision LiDAR with a scanning frequency ≥10Hz, ranging accuracy ±2cm, and horizontal field of view ≥120° can be used to effectively match close-range detection scenarios of 0.1-3 meters and quickly output point cloud data to meet real-time compensation requirements.

[0065] To ensure accurate localization, point cloud preprocessing can be performed, such as denoising filtering and ground segmentation. During denoising filtering, voxel lattice downsampling can be used to reduce point cloud density and computational load. Statistical filtering can be employed to remove ground noise (such as pebbles and weeds) and outliers (such as sudden interference points), retaining only valid target point clouds. For ground segmentation, the RANSAC algorithm can be used to fit the ground plane, classifying points with a z-axis height <0.1m (i.e., within 10cm of the ground) as ground level and removing them. Only target point clouds above the ground are retained, allowing focus on non-ground targets such as humans and reducing interference for subsequent feature extraction.

[0066] In scenarios where the distance information is incomplete due to insufficient visual perception of human body features when the distance between the legged robot and the target object is too close (e.g., ≤1m), the camera may truncate the target object's body features due to the limited field of view. For example, it may only be able to capture a part of the legs / torso, and cannot recognize the complete visual identifier or human outline. In this case, the lidar can determine the distance by extracting the leg point cloud features.

[0067] When extracting close-range leg point cloud features, feature filtering can be performed first. Based on the preprocessed point cloud, point cloud clusters with a z-axis height of 0.1-1.2m (covering the human leg area) and an x-axis distance of 0.3-3m (matching close-range scenes) can be selected. These point cloud clusters can correspond to the leg area of ​​the target object. Then, clustering and segmentation are performed. Euclidean clustering algorithm is used to cluster the selected point cloud according to spatial distance (for example, adjacent points ≤5cm are considered as the same target), resulting in one or more leg point cloud clusters (if it is a single person scene, usually 2 clusters are output, corresponding to the left and right legs). Next, key feature extraction is performed. For each leg point cloud cluster, the coordinates of the cluster center point (x1, y1, z1) and the bounding box size of the cluster are calculated (e.g., width 15-25cm, height 80-100cm, matching the size of an adult's leg) to ensure that the extracted features are effective leg features, rather than objects of similar size in the environment (such as table and chair legs).

[0068] After extracting the close-range leg point cloud features, the target distance can be determined. Taking the center of the legged robot as the origin, the midpoint of the line connecting the center points of two leg point cloud clusters is taken as the core positioning point of the target object, and the distance between this positioning point and the origin of the legged robot is determined. Here, the influence of the z-axis height on the horizontal distance is negligible.

[0069] To ensure accuracy, accuracy verification can be performed. Specifically, if the bounding box size of the leg point cloud cluster deviates from the preset adult leg size by ≤10%, the distance d is deemed valid (error ≤2cm) and is directly used as the compensated target distance. If the deviation is large (e.g., children's legs are thinner), the distance calculation is corrected by the ratio of cluster height to width (e.g., the smaller the width, the closer the distance may be) to ensure that the distance accuracy meets the requirements.

[0070] When the visual output is incomplete due to the close proximity, the leg point cloud distance calculated by the LiDAR can be used to replace the visual distance data. This ensures that the legged robot maintains a stable close following distance (0.5-1m) with the target object and avoids collisions or loss of distance control due to visual failure.

[0071] For scenarios where the distance information of the target leaves the effective field of view of the camera is compensated, when the target object moves outside the effective field of view of the camera, such as when the legged robot turns around and the target object has not yet entered the camera's field of view, or when the target object moves around to the back of the legged robot, the vision cannot output distance information at all. At this time, the LiDAR determines the distance by combining historical feature matching with real-time point cloud tracking.

[0072] In order to detect the scene in real time, the out-of-field tracking trigger condition can be preset, and the output of the vision module can be monitored in real time. If no UWB beacon feature or complete human body outline is detected for 3 consecutive frames (about 0.3 seconds), the out-of-field compensation mode of the lidar is triggered.

[0073] In practical implementation, target association can be performed based on historical point cloud features. Since the LiDAR has recorded the feature template of the human leg point cloud (such as the relative position, size ratio, and historical trajectory of the center point of the two leg clusters) and stored it in the cache before the target object leaves the field of view, real-time point cloud matching can be performed. The LiDAR continuously scans possible areas outside the field of view (such as the side and rear of the legged robot, the turning direction, etc.) and performs similarity matching between the real-time acquired point cloud clusters and the historical feature templates (e.g., the relative position deviation of the clusters is ≤10cm and the size ratio deviation is ≤15% as a successful match). If the match is successful, the point cloud cluster can be determined to be the leg of the target object that has left the field of view, avoiding mistaking other environmental targets (such as passers-by or furniture) as the tracking object. If the match fails (e.g., the target object temporarily leaves the scanning range), the distance is initially estimated by combining the UWB distance (if UWB is effective), and updated after the point cloud is successfully matched again.

[0074] Using the above method, the compensation distance can be continuously output. For the successfully matched leg point cloud clusters, the real-time distance can be determined according to the method corresponding to the above scenario. When the vision detects the target object again (that is, the target object returns to the camera's FoV field of view), it can automatically switch back to the multimodal fusion of vision + LiDAR + UWB, stop the LiDAR compensation alone, and ensure the continuity and accuracy of the distance information.

[0075] In some possible implementations, the method further includes: When the UWB beacon is in autonomous recharging mode, positioning data via UWB positioning is acquired; the UWB beacon is fixed to the charging device that charges the legged robot in the autonomous recharging mode. Based on the positioning data obtained from the UWB positioning, the legged robot is controlled to move to the charging range of the charging device, and the positioning is assisted by other sensors and the QR code on the charging device to complete the docking of the charging port.

[0076] In the above steps, when the UWB beacon is in autonomous recharging mode, positioning data obtained through UWB positioning can be acquired. Coarse alignment is performed based on the UWB positioning data, and the legged robot is controlled to move to the charging range of the charging device (e.g., a charging pile), for example, within a lateral error range of <30cm. Then, fine alignment is performed using other sensors (e.g., cameras and LiDAR) and the QR code on the charging device to assist in positioning and complete the docking of the charging port.

[0077] The charging station can be equipped with a magnetic UWB beacon slot, supporting quick installation and removal. The charging contacts adopt a self-aligning spring pin design, allowing for a positional tolerance of ±5cm.

[0078] For example, please see Figure 3 This is a schematic diagram of coarse alignment of a legged robot in autonomous recharging mode, as provided in an embodiment of this disclosure. Figure 3 As shown, the legged robot 200 uses its UWB receiver 203 to position the charging device 300, reaches the charging range of the charging device 300, completes coarse alignment pre-aiming, and meets the starting accuracy required for subsequent fine alignment.

[0079] Next, please see Figure 4 This is a schematic diagram illustrating the precise alignment of a legged robot in autonomous recharging mode, as provided in an embodiment of this disclosure. Figure 4 As shown, within the charging range of the charging device 300, the charging device 300 is repositioned using other sensors (such as cameras, LiDAR, etc.) and a QR code on the charging device 300 to reach the designated charging position, achieving precise alignment and positioning. Here, cameras and LiDAR have higher measurement accuracy than UWB, which helps to achieve precise pile alignment.

[0080] See again Figure 2 The charging location is set on the abdomen of the legged robot 200, such as... Figure 2 As shown in 202, specifically, charging electrodes can be provided.

[0081] For example, please see Figure 5 This is a schematic diagram illustrating the charging of a legged robot in autonomous recharging mode, as provided in an embodiment of this disclosure. Figure 5 As shown, the legged robot 200 starts in a standing posture above the charging device 300 and moves to a crawling posture. When the legged robot 200 lies on the charging base, a short circuit is formed using the large electrode plate on its abdomen. The main controller of the charging device 300 determines whether the positive and negative electrodes are making good contact by collecting the voltage ADC values ​​at the two ports of the positive and negative electrodes. Here, the charging base can have fine-tuning capabilities to meet the required charging precision.

[0082] In practical applications, due to the lack of a corresponding index connection between legged robots and charging devices, for example, if a user sets up multiple legged robots and deploys multiple charging devices, and the user wants a specific legged robot to complete the charging task on a specific charging device, it requires operation through auxiliary software such as host computer software. However, in this embodiment, by fixing the UWB beacon to the charging device in autonomous recharging mode, UWB can be directly used to match the legged robot with the charging device. Using UWB positioning, the legged robot can quickly and coarsely align with the charging area, significantly reducing the lateral error range. Combined with other sensors and the QR code on the charging device for fine alignment, accurate docking of the charging port is completed, effectively improving the positioning efficiency and docking success rate of the recharging process, reducing the complexity of operation steps, and improving the reliability and stability of autonomous recharging.

[0083] S102: The positioning data under the multiple modes are fused to obtain fused position data, and a following control strategy is generated based on the fused position data.

[0084] In this step, multimodal positioning data from UWB, vision, and LiDAR can be fused to obtain a unified, continuous, and highly reliable fused positioning result. Based on this, a following control strategy can be generated to effectively suppress the impact of single sensor errors or transient failures, thereby improving the following smoothness, path rationality, and motion safety of legged robots in complex environments.

[0085] In some possible implementations, the fusion processing of the positioning data under the multiple modalities to obtain fused location data includes: An attention module is used to preprocess UWB positioning data, visual positioning data, and radar ranging data to obtain preprocessed multimodal data. During the preprocessing, attention weights are assigned according to the data validity of each modality to prioritize the more reliable modal data. The preprocessed multimodal data are labeled and encoded, and then mapped to the same spatial coordinate system; The Transformer model is used to fuse multiple modal data in the same spatial coordinate system after label encoding to determine the fused location data.

[0086] For example, see Figure 6 This is a schematic diagram illustrating a data fusion and strategy generation process provided in an embodiment of this disclosure, such as... Figure 6As shown, multimodal sensors such as LiDAR, cameras, and UWB collect raw data (e.g., point clouds, images, UWB signal arrival times recorded by the UWB receiver). This data is first input into the Attention Model, which assigns attention weights based on the validity of the data from each sensor, prioritizing more reliable modal data. Additionally, the Pretrained Model performs pre-training feature extraction and other preprocessing operations on the relevant data, providing a foundation for subsequent fusion. The preprocessed data then enters the Tag Encoding module for label encoding, transforming the data into an encoding format suitable for subsequent processing. Next, the encoded data enters the Common Space Projection module, mapping the labeled data to the same coordinate system, achieving spatial alignment of the multi-source data and providing more accurate and robust foundational data for subsequent weight allocation and feature fusion. Finally, a Transformer model is used to assign weights to the projected labeled data onto the same spatial coordinate system and fuse them to obtain fused positional data. The visual positioning information obtained through the camera is mainly used to acquire semantic information, while the radar is mainly used to acquire geometric information. The UWB receiving unit, due to the presence of the tag, can acquire the semantic information of the UWB beacon and simultaneously determine geometric information through ranging. During fusion, the weight of the UWB signal is higher when it is stable. When electromagnetic interference causes inaccurate measurements of the UWB signal, the visual positioning data and radar ranging data are used to compensate for the positioning error of the UWB signal. This means reducing the weight of the UWB signal and increasing the weight of the visual and lidar signals. Thus, the visual positioning information and lidar ranging can supplement the UWB positioning data, thereby improving measurement robustness when the UWB is interfered with.

[0087] In this way, the attention module dynamically allocates weights based on the validity of each modality's data, prioritizing reliable information. Then, through label encoding, the multi-source data is uniformly mapped to the same spatial coordinate system, and deep fusion is achieved with the help of the Transformer model. When the localization data of a certain modality is disturbed, other modal localization data are used to supplement it in time, forming a complementary fusion mechanism. This significantly improves the continuity, robustness, and accuracy of the localization results, ensuring that the legged robot can maintain stable and accurate following performance in complex environments.

[0088] In some possible implementations, the step of using a Transformer model to fuse multiple modal data in the same spatial coordinate system after label encoding to determine the fused location data includes: The Transformer model is used to fuse multiple modal data in the same spatial coordinate system after label encoding to obtain a multimodal fusion result; the multimodal fusion result indicates the relative position information of the target object relative to the legged robot, which is output multiple times in succession; The Kalman filter algorithm is used to smooth the multimodal fusion results to obtain the final fused position data.

[0089] See again Figure 6 As can be seen, the multimodal data in the same spatial coordinate system after label encoding enters the Transformer encoder for feature-level fusion. For example, fusion weights for each modality are assigned based on factors such as UWB signal occlusion and visual perception failure. Feature fusion is then performed to extract the features after multimodal fusion, resulting in Encoder Outputs. These Encoder Outputs are then input to the Transformer decoder for decoding, yielding Decoder Outputs, which are the multimodal fusion results. This allows determination of the relative position information of the target object relative to the legged robot in multiple consecutive outputs. A Kalman filtering algorithm is used to smooth the positional relationships of the consecutive outputs, avoiding abrupt jumps in output, thus obtaining the final fused position data.

[0090] In this way, by performing deep temporal fusion on the multimodal data after unifying coordinates using the Transformer model, the continuous relative position information of the target object relative to the legged robot is obtained. Then, the sequence is smoothed by Kalman filtering, which effectively suppresses abrupt noise and instantaneous jumps, and outputs stable, coherent and highly reliable fused position data, thereby ensuring the stability and accuracy of the following control in dynamic scenarios.

[0091] Combined again Figure 6 As can be seen, the fused position data is input into the Path Planning and Motion Control Module (PNCModule) to generate a following control strategy. This module includes a Path Planning submodule and a Motion Control submodule, thereby realizing path planning and motion control for the legged robot. Simultaneously, a yaw angle adaptation mechanism can be combined to transmit the following control strategy to Tag Encoding. This allows the previous motion state of the legged robot to be incorporated into subsequent target encoding, helping to improve the stability and consistency of the legged robot's following.

[0092] In some possible implementations, the legged robot is further equipped with an inertial measurement unit (IMU) and a joint encoder; the method also includes: The joint encoder measures the leg joint movement of the legged robot, and the IMU obtains the angular velocity and acceleration of the legged robot. The body pose information of the legged robot is determined by fusing the output results of the IMU and the joint encoder through Kalman filtering. The process of fusing the positioning data under the multiple modalities to obtain fused location data includes: Based on the body pose information of the legged robot, the positioning data under the multiple modalities are fused to obtain the fused position data.

[0093] In the above steps, each joint encoder can be set at each joint of the legged robot. The instantaneous angle of each joint encoder can be read, and the position of the four legs in the body coordinate system can be determined by forward kinematics. This position is compared with the position of the legs in the previous cycle, and the motion velocity of the legs in the body coordinate system is estimated by differential method. The IMU can provide the raw values ​​of angular velocity and acceleration in the body coordinate system. After subtracting the pre-calibrated zero bias, the initial attitude value is estimated based on the direction of gravity. The acceleration is integrated into the velocity change, and the angular velocity is integrated into the attitude change, forming a preliminary prediction of the robot's position and orientation.

[0094] To suppress integral drift, the physical constraint that "the instantaneous velocity of the foot of the supporting leg relative to the ground is zero" is used. When a leg is detected to be in the supporting phase, the velocity of the foot of that leg in the body coordinate system is regarded as the reverse observation of the motion of the body itself. As long as one of the four legs is supporting, a velocity observation value can be obtained.

[0095] The position, velocity, and orientation predicted by the IMU are used as time updates for the Kalman filter. The velocity observations provided by the supporting leg and the zero displacement constraint of the foot position at the moment of contact are used as measurement updates to correct the prediction results. The corrected state is the position, velocity, and attitude of the body in the world coordinate system at the current moment, which is the final body pose information.

[0096] The entire process is executed cyclically in every millisecond cycle. The IMU is responsible for high-frequency prediction, the joint encoder is responsible for generating zero-velocity or zero-displacement constraints, and the Kalman filter integrates the two to effectively realize the leg odometry function. It retains the wideband response of the IMU and eliminates integral drift by using leg constraints, thereby outputting a stable body pose without cumulative error, effectively achieving highly robust pose estimation of the legged robot body.

[0097] In this way, the leg movements, body angular velocity, and acceleration of the legged robot are acquired in real time through the joint encoder and IMU, and high-precision body pose information is obtained by Kalman filtering and fusion, realizing the legged odometry function. Then, the multimodal positioning data is fused in a unified manner with the body pose information, effectively eliminating the errors caused by body jitter and motion distortion, realizing efficient coupling between external observation and internal motion state, which helps to improve the stability, continuity and dynamic accuracy of fused position data, and ensures that the legged robot can maintain accurate and smooth following performance under high-speed walking or terrain change conditions.

[0098] When fusing the positioning data from the various modalities based on the body pose information of the legged robot to obtain the fused position data, for details, please refer again to... Figure 6 The locomotion module determines the body pose information of the legged robot. The body pose information enters the Common Space Projection module, where the tagged data and the body pose information are mapped to the same spatial coordinate system. Then, the Transformer model is used to fuse them to obtain Decoder Outputs. The Kalman filtering algorithm is used to smooth the body pose information. The smoothed body pose information and the fused position data are then entered into the PNC Module to generate a following control strategy.

[0099] S103: Perform motion control and path planning on the legged robot according to the following control strategy.

[0100] In this step, gait commands and trajectories that conform to the terrain and dynamic obstacles can be generated in real time according to the following control strategy. This enables the legged robot to smoothly accelerate, decelerate or turn while maintaining a safe distance, avoiding collisions and sudden stops, improving the smoothness of the following process, energy efficiency and terrain adaptability, while ensuring the stability of the body posture and the uniform distribution of joint torque, extending the hardware life and enhancing the overall operational reliability.

[0101] In some possible implementations, the step of performing motion control and path planning on the legged robot according to the following control strategy includes: If, during the fusion processing of the positioning data under the multiple modalities, it is determined that there is missing detection data from the target sensor, the legged robot is controlled to turn so that the target object returns to the detection range of the target sensor.

[0102] In the above steps, before each fusion positioning, the system first checks the validity of the data from each sensor. If no valid frames are received from a target sensor within several consecutive cycles, the sensor data is determined to be missing, and it is immediately marked as invalid. This sensor can be temporarily removed from the weight allocation, and the legged robot's steering is controlled. This pauses the normal straight-line or following trajectory and instead executes an in-situ search and steering process. The robot slowly rotates to one side at a preset small angular velocity while continuously monitoring the sensor output. If consecutive valid frames are received again during the rotation, the steering stops and normal following resumes. If the robot does not recover after rotating to a limited angle, it rotates in the opposite direction to continue searching until the target object re-enters the sensor's field of view. Then, the steering speed is reduced to zero, the sensor is reinstated into the fusion weights, and normal following continues. The entire process is implemented through a periodic state machine. The steering amplitude and speed are adaptively adjusted according to the robot's dynamic capabilities and the sensor's field of view, ensuring that the robot's balance is maintained while resuming detection, avoiding instability or collisions caused by sudden turns.

[0103] In this way, when a target sensor data is found to be missing during the fusion process, a turning command can be triggered in the follow control strategy, so that the legged robot can actively adjust its own orientation and bring the target object back to the detection range of the sensor. This maintains the integrity of multimodal information, avoids the decrease in positioning accuracy due to the failure of a single sensor, and ensures that the follow task is executed continuously, smoothly and reliably.

[0104] The legged robot control method provided in this disclosure adopts a movable UWB beacon and has two working modes: automatic following mode and autonomous recharging mode. It can switch between the two working modes, eliminating the need for users to carry additional equipment, reducing hardware costs, and facilitating flexible application in multiple scenarios. It also improves scenario adaptability and human-computer interaction efficiency. When the UWB beacon is in automatic following mode, it accurately identifies targets through the fusion of multiple modal data from UWB, vision, and LiDAR, achieving dynamic companionship and dynamically adapting to complex terrain and environmental changes. This enhances positioning accuracy, improves obstacle avoidance capabilities, improves following performance, and ensures the reliability and safety of robot control.

[0105] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0106] Based on the same inventive concept, this disclosure also provides a legged robot control device corresponding to the legged robot control method. Since the principle of the legged robot control device in this disclosure is similar to that of the above-mentioned legged robot control method in this disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0107] Please see Figure 7 and Figure 8 , Figure 7 This is one of the schematic diagrams of a legged robot control device provided in an embodiment of this disclosure. Figure 8 This is a second schematic diagram of a legged robot control device provided in an embodiment of this disclosure. Figure 7 As shown in the figure, the legged robot control device 700 provided in this embodiment includes: The positioning module 701 is used to receive UWB signals through the UWB receiving component for UWB positioning, perform visual positioning through the camera, and perform radar ranging through the lidar when the UWB beacon is in automatic following mode, thereby determining positioning data under multiple modes; the UWB beacon is carried by the target object followed by the legged robot in the automatic following mode. The fusion module 702 is used to fuse the positioning data under the multiple modes to obtain fused position data, and generate a following control strategy based on the fused position data; The control module 703 is used to perform motion control and path planning for the legged robot according to the following control strategy.

[0108] In one alternative implementation, such as Figure 8 As shown, the legged robot control device 700 further includes a recharging module 704, which is used for: When the UWB beacon is in autonomous recharging mode, positioning data via UWB positioning is acquired; the UWB beacon is fixed to the charging device that charges the legged robot in the autonomous recharging mode. Based on the positioning data obtained from the UWB positioning, the legged robot is controlled to move to the charging range of the charging device, and the positioning is assisted by other sensors and the QR code on the charging device to complete the docking of the charging port.

[0109] In one alternative implementation, such as Figure 8 As shown, the legged robot control device 700 further includes a detection module 705, which is used for: The joint encoder measures the leg joint movement of the legged robot, and the IMU obtains the angular velocity and acceleration of the legged robot. The body pose information of the legged robot is determined by fusing the output results of the IMU and the joint encoder through Kalman filtering. When the fusion module 702 performs fusion processing on the positioning data under the multiple modalities to obtain fused location data, it is specifically used for: Based on the body pose information of the legged robot, the positioning data under the multiple modalities are fused to obtain the fused position data.

[0110] In an optional implementation, when the positioning module 701 is used to receive UWB signals through the UWB receiving component for UWB positioning, it is specifically used for: Based on the time of flight (TOF) of the UWB signal transmitted between the UWB beacon and the UWB receiving component, as well as the air humidity compensation and temperature compensation, the distance and azimuth of the legged robot relative to the UWB beacon are determined. The air humidity compensation and temperature compensation are used to compensate for the influence of air humidity and temperature on the speed of light, and the speed of light is used to calculate the distance and azimuth angle in conjunction with the TOF.

[0111] In one optional implementation, when the fusion module 702 performs fusion processing on the positioning data under the multiple modalities to obtain fused location data, it is specifically used for: An attention module is used to preprocess UWB positioning data, visual positioning data, and radar ranging data to obtain preprocessed multimodal data. During the preprocessing, attention weights are assigned according to the data validity of each modality to prioritize the more reliable modal data. The preprocessed multimodal data are labeled and encoded, and then mapped to the same spatial coordinate system; The Transformer model is used to fuse multiple modal data in the same spatial coordinate system after label encoding to determine the fused location data.

[0112] In one optional implementation, when the fusion module 702 is used to fuse multiple modal data in the same spatial coordinate system after label encoding using a Transformer model to determine the fused position data, it is specifically used for: The Transformer model is used to fuse multiple modal data in the same spatial coordinate system after label encoding to obtain a multimodal fusion result; the multimodal fusion result indicates the relative position information of the target object relative to the legged robot, which is output multiple times in succession; The Kalman filter algorithm is used to smooth the multimodal fusion results to obtain the final fused position data.

[0113] In one optional implementation, when the control module 703 performs motion control and path planning for the legged robot according to the following control strategy, it is specifically used for: If, during the fusion processing of the positioning data under the multiple modalities, it is determined that there is missing detection data from the target sensor, the legged robot is controlled to turn so that the target object returns to the detection range of the target sensor.

[0114] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0115] The legged robot control device provided in this embodiment adopts a movable UWB beacon and has two working modes: automatic following mode and autonomous recharging mode. It can switch between the two working modes, eliminating the need for users to carry additional equipment, reducing hardware costs, and facilitating flexible application in multiple scenarios. It also improves scenario adaptability and human-computer interaction efficiency. When the UWB beacon is in automatic following mode, it accurately identifies targets through the fusion of multiple modal data from UWB, vision, and LiDAR, achieving dynamic companionship and dynamically adapting to complex terrain and environmental changes. This enhances positioning accuracy, improves obstacle avoidance capabilities, improves following performance, and ensures the reliability and safety of robot control.

[0116] like Figure 9 The diagram shown is an architectural diagram of a legged robot provided in an embodiment of this disclosure. The legged robot is equipped with an ultra-wideband (UWB) receiver 91, a lidar 92, a camera 93, and a controller 94. The UWB receiver 91 is capable of receiving UWB signals transmitted by a UWB beacon. The UWB beacon is a movable UWB beacon with two working modes: an automatic following mode and an autonomous recharging mode. The controller is used to receive UWB signals through the UWB receiving component for UWB positioning, perform visual positioning through the camera, and perform radar ranging through the LiDAR when the UWB beacon is in automatic following mode, thereby determining positioning data under multiple modalities. The UWB beacon is carried by the target object followed by the legged robot in automatic following mode. The positioning data under the multiple modalities are fused to obtain fused position data, and a following control strategy is generated based on the fused position data. According to the following control strategy, the legged robot is subjected to motion control and path planning.

[0117] Among them, the ultra-wideband (UWB) receiver 91, lidar 92, camera 93, and controller 94 can all be mounted on the legged robot body.

[0118] The specific control process of the above-mentioned ultra-wideband (UWB) receiver 91, lidar 92, camera 93, and controller 94 can be referred to the above embodiments, and will not be repeated here.

[0119] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the legged robot. In other embodiments of this application, the legged robot may include more or fewer components than illustrated, or combine some components, or separate some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0120] Furthermore, this disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the legged robot control method described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0121] This disclosure also provides a computer program product, which stores a computer program. When the computer program is run by a processor, it executes the steps of the legged robot control method provided in any of the above embodiments of this disclosure. For details, please refer to the above method embodiments, which will not be repeated here.

[0122] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium, which can be a volatile or non-volatile computer-readable storage medium. In another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0123] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices and apparatuses described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0124] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0125] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0126] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0127] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A control method for a legged robot, characterized in that, The legged robot is equipped with an ultra-wideband (UWB) receiver, a lidar, and a camera; the UWB receiver can receive UWB signals transmitted by a UWB beacon, which is a movable UWB beacon with two operating modes: automatic following mode and autonomous recharging mode; the method includes: When the UWB beacon is in automatic following mode, the UWB receiver receives UWB signals for UWB positioning, the camera performs visual positioning, and the lidar performs radar ranging to determine positioning data in multiple modes; the UWB beacon is carried by the target object followed by the legged robot in automatic following mode. The positioning data under the multiple modes are fused to obtain fused position data, and a following control strategy is generated based on the fused position data; According to the following control strategy, motion control and path planning are performed on the legged robot.

2. The method according to claim 1, characterized in that, The method further includes: When the UWB beacon is in autonomous recharging mode, positioning data via UWB positioning is acquired; the UWB beacon is fixed to the charging device that charges the legged robot in the autonomous recharging mode. Based on the positioning data obtained from the UWB positioning, the legged robot is controlled to move to the charging range of the charging device, and the positioning is assisted by other sensors and the QR code on the charging device to complete the docking of the charging port.

3. The method according to claim 1, characterized in that, The legged robot is also equipped with an inertial measurement unit (IMU) and a joint encoder; the method further includes: The joint encoder measures the leg joint movement of the legged robot, and the IMU obtains the angular velocity and acceleration of the legged robot. The body pose information of the legged robot is determined by fusing the output results of the IMU and the joint encoder through Kalman filtering. The process of fusing the positioning data under the multiple modalities to obtain fused location data includes: Based on the body pose information of the legged robot, the positioning data under the multiple modalities are fused to obtain the fused position data.

4. The method according to claim 1, characterized in that, The step of receiving UWB signals through the UWB receiving component for UWB positioning includes: Based on the time of flight (TOF) of the UWB signal transmitted between the UWB beacon and the UWB receiving component, as well as the air humidity compensation and temperature compensation, the distance and azimuth of the legged robot relative to the UWB beacon are determined. The air humidity compensation and temperature compensation are used to compensate for the influence of air humidity and temperature on the speed of light, and the speed of light is used to calculate the distance and azimuth angle in conjunction with the TOF.

5. The method according to claim 1, characterized in that, The process of fusing the positioning data under the multiple modalities to obtain fused location data includes: An attention module is used to preprocess UWB positioning data, visual positioning data, and radar ranging data to obtain preprocessed multimodal data. During the preprocessing, attention weights are assigned according to the data validity of each modality to prioritize the more reliable modal data. The preprocessed multimodal data are labeled and encoded, and then mapped to the same spatial coordinate system; The Transformer model is used to fuse multiple modal data in the same spatial coordinate system after label encoding to determine the fused location data.

6. The method according to claim 5, characterized in that, The process of fusing multiple modal data in the same spatial coordinate system after label encoding using the Transformer model to determine the fused location data includes: The Transformer model is used to fuse multiple modal data in the same spatial coordinate system after label encoding to obtain a multimodal fusion result; the multimodal fusion result indicates the relative position information of the target object relative to the legged robot, which is output multiple times in succession; The Kalman filter algorithm is used to smooth the multimodal fusion results to obtain the final fused position data.

7. The method according to claim 1, characterized in that, The step of performing motion control and path planning on the legged robot according to the following control strategy includes: If, during the fusion processing of the positioning data under the multiple modalities, it is determined that there is missing detection data from the target sensor, the legged robot is controlled to turn so that the target object returns to the detection range of the target sensor.

8. A control device for a legged robot, characterized in that, The legged robot is equipped with an ultra-wideband (UWB) receiver, a lidar, and a camera; the UWB receiver can receive UWB signals transmitted by a UWB beacon, which is a movable UWB beacon with two operating modes: automatic following mode and autonomous recharging mode; the device includes: The positioning module is used to receive UWB signals through the UWB receiving component for UWB positioning, perform visual positioning through the camera, and perform radar ranging through the lidar when the UWB beacon is in automatic following mode, thereby determining positioning data in multiple modes; the UWB beacon is carried by the target object followed by the legged robot in the automatic following mode. The fusion module is used to fuse the positioning data under the multiple modes to obtain fused position data, and generate a following control strategy based on the fused position data; The control module is used to perform motion control and path planning for the legged robot according to the following control strategy.

9. A legged robot, characterized in that, The legged robot is equipped with an ultra-wideband (UWB) receiver, a lidar, a camera, and a controller. The UWB receiver can receive UWB signals transmitted by a UWB beacon. The UWB beacon is a movable UWB beacon with two working modes: automatic following mode and autonomous recharging mode. The controller is used to receive UWB signals through the UWB receiving component for UWB positioning, perform visual positioning through the camera, and perform radar ranging through the LiDAR when the UWB beacon is in automatic following mode, thereby determining positioning data under multiple modalities. The UWB beacon is carried by the target object followed by the legged robot in automatic following mode. The positioning data under the multiple modalities are fused to obtain fused position data, and a following control strategy is generated based on the fused position data. According to the following control strategy, the legged robot is subjected to motion control and path planning.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the legged robot control method as described in any one of claims 1 to 7.