Quadruped robot motion information determination method and device and quadruped robot
By using binocular cameras and depth cameras on a quadruped robot for multimodal data fusion, the high cost and low efficiency problems of existing technologies are solved, and accurate and stable motion information determination is achieved, which is suitable for navigation tasks on urban sidewalks and in complex environments.
Patent Information
- Application Number
- CN202510673138.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Existing quadruped robot mapping and positioning relies on high-precision lidar pre-mapping and scanning, which has high hardware costs and makes it difficult to balance real-time performance and multimodal data fusion efficiency.
Using binocular cameras and depth cameras as multimodal sensors, combined with MLP and Transformer encoders, the obstacle avoidance path and gait parameters of the quadruped robot are determined through multimodal data fusion, reducing hardware costs and improving data fusion efficiency.
This method reduces hardware costs while improving the accuracy and stability of the quadruped robot's motion information, making it suitable for urban sidewalk navigation and navigation tasks in complex environments.
Smart Images

Figure CN120705537A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to robot control technology, and more specifically, to a method and device for determining motion information of a quadruped robot, and a quadruped robot. Background Art
[0002] Existing quadruped robot mapping and localization rely on pre-built, high-precision lidar scanning and deployment, resulting in high hardware costs and difficulty generalizing to different scenarios. Existing algorithms for determining quadruped robot motion information struggle to balance real-time performance with efficient multimodal data fusion. Summary of the Invention
[0003] An object of the present invention is to provide a new technical solution for a method of determining motion information of a quadruped robot.
[0004] According to a first aspect of the present invention, a method for determining motion information of a quadruped robot is provided, comprising:
[0005] Acquire data collected by a multimodal sensor, wherein the multimodal sensor includes a binocular camera and a depth camera, the binocular camera is arranged on the front side of the quadruped robot, and the depth camera is symmetrically arranged on the front and rear sides of the quadruped robot;
[0006] Determining the BEV characteristics corresponding to the current moment based on the data collected by the multimodal sensor;
[0007] Obtain BEV characteristics corresponding to historical moments;
[0008] Determining a time series BEV feature according to the BEV feature corresponding to the historical moment and the BEV feature corresponding to the current moment;
[0009] Determining, based on the data collected by the multimodal sensor, a multimodal fusion feature after spatial alignment at a current moment;
[0010] The obstacle avoidance path and gait parameters of the quadruped robot are determined according to the time-series BEV features and the multimodal fusion features after spatial alignment at the current moment.
[0011] Optionally, determining the obstacle avoidance path and gait parameters of the quadruped robot according to the time-series BEV features and the multimodal fusion features after spatial alignment at the current moment includes:
[0012] Extracting motion information of dynamic obstacles from the time-series BEV features, and extracting spatial position information of dynamic obstacles, spatial position information of static obstacles, and environmental information from the multimodal fusion features after spatial alignment at the current moment;
[0013] Constructing a cost map based on the motion information of the dynamic obstacle, the spatial position information of the dynamic obstacle, and the spatial position information of the static obstacle;
[0014] Obtain the current position information and target point position information of the quadruped robot;
[0015] According to the current position information of the quadruped robot and the target point position information, the obstacle avoidance path of the quadruped robot is determined from the cost map, and according to the environmental information, the gait parameters of the quadruped robot at each landing point on the obstacle avoidance path are determined.
[0016] Optionally, determining, based on the environmental information, gait parameters of the quadruped robot at each landing point on the obstacle avoidance path includes:
[0017] Get the speed information of the quadruped robot;
[0018] determining a landing point of the quadruped robot on the obstacle avoidance path according to the speed information of the quadruped robot;
[0019] Determining environmental information corresponding to a landing point of the quadruped robot on the obstacle avoidance path;
[0020] According to the environmental information corresponding to the landing points of the quadruped robot on the obstacle avoidance path, the gait parameters of the quadruped robot at each landing point on the obstacle avoidance path are determined.
[0021] Optionally, determining the BEV characteristics corresponding to the current moment based on the data collected by the multimodal sensor includes:
[0022] The data collected by the multimodal sensor is input into the MLP to perform nonlinear fusion on the data collected by the multimodal sensor and map it to the BEV space to obtain the BEV features corresponding to the current moment.
[0023] Optionally, determining the time series BEV feature according to the BEV feature corresponding to the historical moment and the BEV feature corresponding to the current moment includes:
[0024] The BEV features corresponding to the historical moments and the BEV features corresponding to the current moment are input into a Transformer encoder to obtain a temporal BEV feature determined by temporal splicing, wherein the temporal BEV feature records the motion information of the dynamic obstacle.
[0025] Optionally, determining, based on the data collected by the multimodal sensor, the multimodal fusion features after spatial alignment at the current moment, includes:
[0026] The data collected by the multimodal sensor is input into the MLP to fuse the data collected by the multimodal sensor to obtain the multimodal fusion features after spatial alignment at the current moment, wherein the multimodal fusion features after spatial alignment at the current moment record the spatial position information of dynamic obstacles, the spatial position information of static obstacles, and environmental information.
[0027] Optionally, the multimodal sensor also includes a fisheye camera, a lidar and a telephoto camera, the fisheye camera is arranged on the left side, right side and rear side of the quadruped robot, the lidar is arranged on the top of the quadruped robot, and the telephoto camera is arranged on the front side of the quadruped robot.
[0028] According to a second aspect of the present invention, there is provided a device for determining motion information of a quadruped robot, comprising:
[0029] A first acquisition module is configured to acquire data collected by a multimodal sensor, wherein the multimodal sensor includes a binocular camera and a depth camera, the binocular camera is disposed on the front side of the quadruped robot, and the depth camera is symmetrically disposed on the front and rear sides of the quadruped robot;
[0030] A BEV feature determination module, configured to determine the BEV feature corresponding to the current moment based on the data collected by the multimodal sensor;
[0031] The second acquisition module is used to obtain the BEV characteristics corresponding to the historical moment;
[0032] a time-series BEV feature determination module, configured to determine a time-series BEV feature based on the BEV feature corresponding to the historical moment and the BEV feature corresponding to the current moment;
[0033] A multimodal fusion feature determination module, configured to determine the multimodal fusion features after spatial alignment at the current moment based on the data collected by the multimodal sensor;
[0034] The motion information determination module is used to determine the obstacle avoidance path and gait parameters of the quadruped robot based on the time-series BEV features and the multimodal fusion features after spatial alignment at the current moment.
[0035] According to a third aspect of the present invention, a device for determining motion information of a quadruped robot is provided, comprising a memory and a processor, wherein the memory stores a computer program for controlling the processor to operate to execute the method according to any one of the first aspects.
[0036] According to a fourth aspect of the present invention, there is provided a quadruped robot comprising the quadruped robot motion information determination device as described in the second aspect or the third aspect.
[0037] The method for determining the motion information of a quadruped robot provided by the present invention is that the multimodal sensors installed on the quadruped robot are binocular cameras and depth cameras, which are relatively low in cost. In addition, based on the data collected by the multimodal sensors, the parameters required for the motion information of the quadruped robot are determined from two branches respectively, one branch is to determine the time series BEV features, and the other branch is to determine the multimodal fusion features after spatial alignment at the current moment, thereby improving the accuracy of the operation information of the quadruped robot and ensuring the stability of the quadruped robot's motion.
[0038] Features and advantages of the embodiments of the present specification will become apparent from the following detailed description of exemplary embodiments of the present specification with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the specification and, together with the description, serve to explain the principles of the embodiments of the specification.
[0040] Figure 1 4 is a flow chart of a method for determining motion information of a quadruped robot according to an embodiment of the present invention.
[0041] Figure 2 Schematic diagram of a multimodal sensor provided on a quadruped robot according to one embodiment of the present invention.
[0042] Figure 3 Schematic diagram of a multimodal sensor provided on a quadruped robot according to one embodiment of the present invention.
[0043] Figure 4 It is a schematic block diagram of a principle of a device for determining motion information of a quadruped robot according to an embodiment of the present invention.
[0044] Figure 5 4 is a structural block diagram of a device for determining motion information of a quadruped robot according to an embodiment of the present invention.
[0045] Figure 6 1 is a schematic structural diagram of a quadruped robot according to an embodiment of the present invention. DETAILED DESCRIPTION
[0046] Various exemplary embodiments of the present specification will now be described in detail with reference to the accompanying drawings.
[0047] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the embodiments of this specification, its application, or uses.
[0048] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0049] In order to solve the above technical problems, the present invention provides a method for determining the motion information of a quadruped robot. The multimodal sensors installed on the quadruped robot are binocular cameras and depth cameras, which are relatively low in cost. In addition, based on the data collected by the multimodal sensors, the parameters required for the motion information of the quadruped robot are determined from two branches respectively. One branch is to determine the time series BEV features, and the other branch is to determine the multimodal fusion features after spatial alignment at the current moment, which improves the accuracy of the operation information of the quadruped robot and ensures the stability of the quadruped robot's motion.
[0050] In one embodiment of the present invention, a method for determining motion information of a quadruped robot is provided. Figure 1 As shown, the method for determining motion information of a quadruped robot in this embodiment includes the following steps S110 to S140.
[0051] Step S110, acquiring data collected by a multimodal sensor, wherein the multimodal sensor includes a binocular camera and a depth camera, the binocular camera is set on the front side of the quadruped robot, and the depth camera is symmetrically set on the front and back sides of the quadruped robot.
[0052] The resolution of the binocular camera is 1920×1080. Figure 2 , the field of view is 120°×90°, the baseline distance is 15cm, and the detection distance is 30m. Figure 2 There are two depth cameras, which are symmetrically set on the front and back sides of the quadruped robot. The field of view is 120°×90° and the detection distance is 0.1-5m.
[0053] In some embodiments, the multimodal sensor also includes a fisheye camera, a laser radar, and a telephoto camera. The fisheye cameras are located on the left, right, and rear sides of the quadruped robot, the laser radar is located on the top of the quadruped robot, and the telephoto camera is located on the front side of the quadruped robot. This allows the multimodal sensor to collect richer data, providing more comprehensive and accurate data for determining the quadruped robot's motion information.
[0054] See Figure 3 The three fisheye cameras have a horizontal field of view of ≥180° and a detection range of 0-10m. The lidar is an 8-line lidar with a vertical angular resolution of 2° and a horizontal angular resolution of 0.1°, and a detection range of 0.2-20m. The telephoto camera has a resolution of 1920×1080, a field of view of 120°×30°, and a detection range of 100m.
[0055] When the multimodal sensor includes a binocular camera and a depth camera, the data collected by the multimodal sensor includes RGB image data and depth image data. The RGB image data is collected by the binocular camera, and the depth image data is collected by the depth camera.
[0056] When the multimodal sensor includes a binocular camera, a depth camera, a fisheye camera, a lidar, and a telephoto camera, the data collected by the multimodal sensor includes RGB image data, depth image data, and point cloud data. The RGB image data is collected by the binocular camera, the fisheye camera, and the telephoto camera, the depth image data is collected by the depth camera, and the point cloud data is collected by the lidar.
[0057] RGB image data I t ∈R H×W×3 In [1], H is the image height (Height), the unit is pixel (Pixel), W is the image width (Width), the unit is pixel (Pixel), 3 is the color channel, corresponding to the RGB red, green and blue channels respectively, t is the timestamp, corresponding to the current frame is the t-th frame in the time series data, R is the pixel value, usually normalized to [0, 1] or [0, 255]. For example, the tensor shape of an RGB image with a resolution of 1080×1920 is R 1080×1920×3 .
[0058] Depth image data D t ∈R H×W In the example, H and W represent the height and width consistent with the RGB image, and the purpose is to align with the RGB image. t Each pixel value represents the depth value at the corresponding position, in meters or millimeters. The depth map is single-channel data and has no color information. To align with the RGB image, the pixel coordinates can be mapped to 3D space using the camera intrinsic parameters.
[0059] Point cloud data t ∈R N×3 Where N is the number of points in the point cloud, 3 is the three-dimensional coordinates (X, Y, Z) of each point in meters, and R is the coordinate value, which is a real number (including positive and negative, depending on the specific position in the coordinate system).
[0060] Step S120 : determining the BEV characteristics corresponding to the current moment based on the data collected by the multimodal sensor.
[0061] In some embodiments, before executing the above step S120, the method also includes: normalizing and histogram equalizing the RGB image data to obtain preprocessed RGB image data; filling the depth image data with invalid values, mapping the single-channel depth data to three-channel color data, and adjusting the resolution corresponding to the depth image data to be consistent with the resolution corresponding to the RGB image data to obtain preprocessed depth image data; performing denoising filtering on the point cloud data, voxelizing the point cloud data, and retaining the point cloud data corresponding to non-empty voxels to obtain preprocessed point cloud data; and performing timestamp alignment on the preprocessed RGB image data, the preprocessed depth image data, and the preprocessed point cloud data to obtain aligned RGB image data, depth image data, and point cloud data.
[0062] Normalize the RGB image data, specifically by converting the original pixel value I raw (x, y, c) are mapped from integer range (0-255) to floating point range (0-1),
[0063]
[0064] Among them, (x, y) is the pixel coordinate, c∈{R,G,B}, specifically a color channel.
[0065] Perform histogram equalization on the RGB image data, specifically: split the RGB image into 8×8 grids, calculate the histogram for each grid independently, apply histogram equalization to each grid, limit the contrast increase (Clip Limit = 2.0),
[0066] I eq (x,y,c)=T(I norm (x,y,c))
[0067] Where T(·) is the mapping function of the cumulative distribution function (CDF).
[0068] Normalization and histogram equalization are performed on the RGB image data to improve the accuracy of the RGB image data and enhance the features of low-light areas, such as obstacles in shadows.
[0069] Bilateral filtering interpolation is used to fill invalid values in the depth image data. Specifically,
[0070]
[0071] Among them, (x, y) is the coordinate of the pixel to be filled, (i, j) is the coordinate of all adjacent pixels centered at (x, y), and D filled (x, y) is the depth value corresponding to the pixel to be filled, Draw (i, j) is the depth value corresponding to a certain adjacent pixel centered at (x, y), w spatial is the spatial Gaussian kernel (σ s =5 pixels), w range is the depth difference Gaussian kernel (σ r =0.1 m), W is the normalization coefficient.
[0072] Fill the depth image data with invalid values to reduce interference in invalid areas.
[0073] Using Jet color mapping, the single-channel depth data is mapped to three-channel color data, specifically,
[0074]
[0075] Where (x, y) is the coordinate of the pixel to be mapped in the depth image, D filled (x, y) is the depth value corresponding to the pixel to be mapped, c is the color channel, representing R, B, G3 channels, d min =0.2m,d max =10m is the effective detection range of the depth camera.
[0076] If the depth map resolution is lower than the RGB image resolution, bilinear interpolation upsampling is used to adjust the resolution corresponding to the depth image data to the same resolution as the RGB image data.
[0077] The point cloud data is processed by denoising and filtering. Specifically, the noise points are removed by statistical outlier method, and each point p i , calculate the mean μ and standard deviation σ of the points in its neighborhood (for example, radius r = 0.1m), and remove the points that satisfy the following formula:
[0078] |p i -μ|>3σ.
[0079] Perform voxelization on the point cloud data and retain the point cloud data corresponding to the non-empty voxels. Specifically, define the voxel grid size (for example, 0.05m×0.05m×0.05m) and divide the point cloud data into the X×Y×Z voxel space. Use sparse tensor storage to only retain the point cloud data corresponding to the non-empty voxels, that is,
[0080]
[0081] Among them, (x i ,y i , z i , f i )(x i ,y i , z i) is the position information corresponding to the voxel, fi is the feature corresponding to the voxel, N v is the number of non-empty voxels, and 10 is the feature dimension. i The feature corresponding to the voxel is the density feature, specifically
[0082] f density =log(1+N points )
[0083] Among them, N points is the number of points inside the voxel.
[0084] The point cloud data is subjected to denoising and filtering processing, and the point cloud data corresponding to non-empty voxels is retained to reduce the computational complexity of the point cloud data.
[0085] The timestamps of the preprocessed RGB image data, preprocessed depth image data, and preprocessed point cloud data are aligned to ensure the temporal consistency of the multimodal data and avoid errors caused by temporal misalignment during multimodal data fusion.
[0086] In some embodiments, step S120 specifically includes: inputting the data collected by the multimodal sensor into the MLP to perform nonlinear fusion on the data collected by the multimodal sensor and map it to the BEV space to obtain the BEV features corresponding to the current moment.
[0087] Specifically, the data collected by the multimodal sensors are projected into the BEV (Bird's-Eye View) coordinate system respectively, and then the MLP is used to perform nonlinear fusion on the data collected by the multimodal sensors projected into the BEV coordinate system and map them into the BEV space to obtain the BEV features corresponding to the current moment.
[0088] Before projecting the data collected by the multimodal sensor into the BEV (Bird's-Eye View) coordinate system, feature extraction is performed on the data collected by the multimodal sensor.
[0089] For RGB image data, the original input dimension is H×W×3. To preserve semantics and compress computational complexity, the processed dimension is H f ×W f ×256. Since ResNet-50 is selected as the backbone network and feature extractor, the number of channels of the feature map output by the last residual block (layer4) of ResNet-50 is 2048. Then, through 1×1 convolution dimensionality reduction, the number of channels can be compressed to 256. This maximizes the retention of features while reducing the amount of subsequent calculations. The RGB image data features extracted by ResNet-50 are:
[0090]
[0091] Among them, H f =H / 16,W f =W / 16.
[0092] For depth image data, the input original dimension is H×W×1, and for aligned RGB image data, the processed dimension is H f ×W f ×256. The deep image data features extracted based on ResNet-50 are:
[0093]
[0094] For point cloud data, the original input dimension is N×3, and the point cloud data features extracted by MLP and pooling operations are H B ×W B ×64, where H B , W B is the spatial resolution of the BEV grid, corresponding to the division of the bird's-eye view. Specifically, the point cloud data is divided into H B ×W B For the point cloud in each voxel, local features are extracted through MLP, that is,
[0095] f local =MLP([x,y,z,r])∈R 64
[0096] Where (x, y, z) are the coordinates of the point cloud data, and r is the reflection intensity. Global pooling is then performed to max-pool all the point cloud features within each voxel, yielding a single 64-dimensional feature. This 64-dimensional feature is sufficient to encode geometric and density information and can be aligned with the 256-channel features of the RGB and depth image data through MLP projection.
[0097] The features of point cloud data can be specifically expressed as:
[0098]
[0099] Among them, H B , W B is the spatial resolution of the BEV grid, and 64 is the number of channels.
[0100] Step S130: Obtain BEV features corresponding to historical moments.
[0101] The determination of the BEV features corresponding to the historical moments can refer to the BEV features corresponding to the current moment. The BEV features corresponding to the historical moments can be the BEV features corresponding to one historical moment or the BEV features corresponding to multiple historical moments.
[0102] Step S140 , determining a time series BEV feature based on the BEV feature corresponding to the historical moment and the BEV feature corresponding to the current moment.
[0103] In some embodiments, step S140 specifically includes: inputting the BEV features corresponding to the historical moments and the BEV features corresponding to the current moment into the Transformer encoder to obtain the temporal BEV features determined by temporal splicing, wherein the temporal BEV features record the motion information of the dynamic obstacles.
[0104] Specifically, the BEV features at the current moment and the BEV features corresponding to k historical moments are input into the same multi-head self-attention module, with the aim of allowing the Transformer encoder to learn the association between the BEV features at the current moment and the BEV features corresponding to k historical moments.
[0105] The Transformer encoder architecture is based on key-value pair construction, linear projection, multi-head self-attention, residual connections, and layer normalization. Specifically, the input data is projected onto three sets of vectors: query, key, and value. Multi-head splitting is then performed, which cuts the projected large channel dimension into h segments (for example, h = 8). Attention calculations are performed on each segment, followed by concatenation, and finally linear mapping to obtain time-series BEV features.
[0106] Step S150: Determine the multimodal fusion features after spatial alignment at the current moment based on the data collected by the multimodal sensor.
[0107] In some embodiments, step S150 specifically includes: inputting the data collected by the multimodal sensor into the MLP to fuse the data collected by the multimodal sensor to obtain the multimodal fusion features after spatial alignment at the current moment, wherein the multimodal fusion features after spatial alignment at the current moment record the spatial position information of dynamic obstacles, the spatial position information of static obstacles and environmental information.
[0108] Specifically, the input data is RGB image data features, depth image data features, and point cloud data features, and the output data is the multimodal fusion features after spatial alignment at the current moment.
[0109] The RGB image data feature is represented as: The depth image data feature is represented as: The feature representation of point cloud data is: The multimodal fusion feature after spatial alignment at the current moment is expressed as: Among them, H f , W fis the spatial resolution of the multimodal fusion feature after spatial alignment (usually 1 / 16 of the input image), and C is the number of fusion feature channels, usually set to 256. The extraction methods of RGB image data features, depth image data features, and point cloud data features can refer to the above embodiments and will not be elaborated on here.
[0110] Step S160 , determining the obstacle avoidance path and gait parameters of the quadruped robot based on the time-series BEV features and the multimodal fusion features after spatial alignment at the current moment.
[0111] In some embodiments, step S160 includes steps S161 to S164.
[0112] Step S161 : extracting the motion information of the dynamic obstacle from the time-series BEV features, and extracting the spatial position information of the dynamic obstacle, the spatial position information of the static obstacle, and the environmental information from the multimodal fusion features after spatial alignment at the current moment.
[0113] The motion information of a dynamic obstacle includes the speed and acceleration of the dynamic obstacle.
[0114] The environmental information includes road surface type information, which is one of a flat road surface, a road surface with a slope, and a road surface with steps.
[0115] Step S162: construct a cost map based on the motion information of the dynamic obstacles, the spatial position information of the dynamic obstacles, and the spatial position information of the static obstacles.
[0116] The cost map provided in this embodiment is divided into two-dimensional cost grids. Each two-dimensional cost grid stores a cost value, which represents the difficulty or risk of the quadruped robot passing through the area. The higher the cost value, the greater the difficulty or risk of the quadruped robot passing through the area.
[0117] Step S163: Obtain the current position information of the quadruped robot and the target point position information.
[0118] In step S164, the obstacle avoidance path of the quadruped robot is determined from the cost map according to the current position information of the quadruped robot and the target point position information, and the gait parameters of the quadruped robot at each landing point on the obstacle avoidance path are determined according to the environmental information.
[0119] The static environment optimal algorithm (for example, the A* algorithm) and the dynamic environment optimal algorithm (for example, the D*Lite algorithm) are used to determine the obstacle avoidance path of the quadruped robot from the current position to the target point position.
[0120] In this embodiment, a high-precision obstacle avoidance path is provided for the quadruped robot through the motion information of dynamic obstacles, the spatial position information of dynamic obstacles, and the spatial position information of static obstacles. At the same time, the gait parameters of the quadruped robot are adjusted in real time based on environmental information to ensure the stability of the quadruped robot's movement.
[0121] In this embodiment, determining the gait parameters of the quadruped robot at each landing point on the obstacle avoidance path based on the environmental information in step S164 specifically includes: obtaining the speed information of the quadruped robot; determining the landing point of the quadruped robot on the obstacle avoidance path based on the speed information of the quadruped robot; determining the environmental information corresponding to the landing point of the quadruped robot on the obstacle avoidance path; and determining the gait parameters of the quadruped robot at each landing point on the obstacle avoidance path based on the environmental information corresponding to the landing point of the quadruped robot on the obstacle avoidance path.
[0122] The speed information of the quadruped robot is pre-stored information and can be directly obtained.
[0123] The landing points of the quadruped robot on the obstacle avoidance path form a landing point sequence.
[0124] Gait parameters include at least one of the following: gait cycle, stride length, and step height. The gait cycle is the time it takes to complete a full gait cycle. Stride length is the distance a single leg travels in one cycle. Step height is the maximum height of the swinging leg from the ground.
[0125] After the obstacle avoidance path and gait parameters of the quadruped robot are determined, the quadruped robot is controlled to move according to the obstacle avoidance path and gait parameters of the quadruped robot.
[0126] When the multimodal sensors installed on the quadruped robot include binocular cameras and depth cameras, the quadruped robot is suitable for urban sidewalk navigation. When the multimodal sensors installed on the group robot include binocular cameras, depth cameras, fisheye cameras, lidar and telephoto cameras, the quadruped robot is not only suitable for urban sidewalk navigation, but also for indoor cross-story scenarios, such as disaster rescue in ruined terrain, logistics and distribution scenarios for indoor and outdoor navigation in multi-story buildings, and security inspections in complex lighting environments.
[0127] An embodiment of the present invention further provides a device for determining motion information of a quadruped robot. Figure 4 As shown, the quadruped robot motion information determination device 400 includes a first acquisition module 410, a BEV feature determination module 420, a second acquisition module 430, a time series BEV feature determination module 440, a multimodal fusion feature determination module 450 and a motion information determination module 460.
[0128] The first acquisition module 410 is used to acquire data collected by a multimodal sensor, wherein the multimodal sensor includes a binocular camera and a depth camera. The binocular camera is set on the front side of the quadruped robot, and the depth camera is symmetrically set on the front and back sides of the quadruped robot.
[0129] The BEV feature determination module 420 is used to determine the BEV feature corresponding to the current moment based on the data collected by the multimodal sensor.
[0130] The second acquisition module 430 is used to acquire BEV features corresponding to historical moments.
[0131] The time-series BEV feature determination module 440 is configured to determine the time-series BEV features according to the BEV features corresponding to the historical moments and the BEV features corresponding to the current moment.
[0132] The multimodal fusion feature determination module 450 is used to determine the multimodal fusion features after spatial alignment at the current moment based on the data collected by the multimodal sensor.
[0133] The motion information determination module 460 is used to determine the obstacle avoidance path and gait parameters of the quadruped robot based on the time-series BEV features and the multimodal fusion features after spatial alignment at the current moment.
[0134] In some embodiments, the motion information determination module 460 is also used to extract the motion information of dynamic obstacles from the time-series BEV features, and to extract the spatial position information of dynamic obstacles, the spatial position information of static obstacles and environmental information from the multimodal fusion features after spatial alignment at the current moment; construct a cost map based on the motion information of dynamic obstacles, the spatial position information of dynamic obstacles and the spatial position information of static obstacles; obtain the current position information and target point position information of the quadruped robot; determine the obstacle avoidance path of the quadruped robot from the cost map based on the current position information and target point position information of the quadruped robot, and determine the gait parameters of the quadruped robot at each landing point on the obstacle avoidance path based on the environmental information.
[0135] In some embodiments, the motion information determination module 460 is also used to obtain the speed information of the quadruped robot; determine the landing point of the quadruped robot on the obstacle avoidance path based on the speed information of the quadruped robot; determine the environmental information corresponding to the landing point of the quadruped robot on the obstacle avoidance path; and determine the gait parameters of the quadruped robot at each landing point on the obstacle avoidance path based on the environmental information corresponding to the landing point of the quadruped robot on the obstacle avoidance path.
[0136] In some embodiments, the BEV feature determination module 420 is used to input the data collected by the multimodal sensor into the MLP to perform nonlinear fusion on the data collected by the multimodal sensor and map it to the BEV space to obtain the BEV feature corresponding to the current moment.
[0137] In some embodiments, the time series BEV feature determination module 440 is used to input the BEV features corresponding to the historical moments and the BEV features corresponding to the current moments into the Transformer encoder to obtain the time series BEV features determined by time series splicing, wherein the time series BEV features record the motion information of dynamic obstacles.
[0138] In some embodiments, the multimodal fusion feature determination module 450 is used to input the data collected by the multimodal sensor into the MLP to fuse the data collected by the multimodal sensor to obtain the multimodal fusion features after spatial alignment at the current moment, wherein the multimodal fusion features after spatial alignment at the current moment record the spatial position information of dynamic obstacles, the spatial position information of static obstacles and environmental information.
[0139] In some embodiments, the multimodal sensor also includes a fisheye camera, a lidar, and a telephoto camera. The fisheye camera is set on the left, right, and rear sides of the quadruped robot, the lidar is set on the top of the quadruped robot, and the telephoto camera is set on the front side of the quadruped robot.
[0140] An embodiment of the present invention further provides a device for determining motion information of a quadruped robot, such as Figure 5 The quadruped robot motion information determination device 500 includes a memory 520 and a processor 510. The memory 520 stores a computer program, which is used to control the processor 510 to operate so as to execute the quadruped robot motion information determination method provided by any of the above embodiments.
[0141] One embodiment of the present invention further provides a quadruped robot, such as Figure 6 As shown, it includes a quadruped robot motion information determination device as provided in any of the above embodiments.
[0142] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0143] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0144] The embodiments of this specification may be systems, methods, and / or computer program products. The computer program product may include a computer-readable storage medium carrying computer instructions for causing a processor to implement various aspects of the embodiments of this specification.
[0145] A computer-readable storage medium can be a tangible device that can hold and store computer instructions for use by a computer instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which computer instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0146] The computer instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer instructions from the network and forwards the computer instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0147] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architectures, functions and operations of the systems, methods and computer program products according to multiple embodiments of this specification. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of a computer instruction, and the module, program segment or part of a computer instruction contains one or more executable computer instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are all equivalent.
[0148] The embodiments of the present specification have been described above. The above description is illustrative and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for determining motion information of a quadruped robot, characterized in that: include: Acquire data collected by a multimodal sensor, wherein the multimodal sensor includes a binocular camera and a depth camera, the binocular camera is arranged on the front side of the quadruped robot, and the depth camera is symmetrically arranged on the front and rear sides of the quadruped robot; Determining the BEV characteristics corresponding to the current moment based on the data collected by the multimodal sensor; Obtain BEV characteristics corresponding to historical moments; Determining a time series BEV feature according to the BEV feature corresponding to the historical moment and the BEV feature corresponding to the current moment; Determining, based on the data collected by the multimodal sensor, a multimodal fusion feature after spatial alignment at a current moment; The obstacle avoidance path and gait parameters of the quadruped robot are determined according to the time-series BEV features and the multimodal fusion features after spatial alignment at the current moment.
2. The method according to claim 1, characterized in that The determining of the obstacle avoidance path and gait parameters of the quadruped robot according to the time-series BEV features and the multimodal fusion features after spatial alignment at the current moment includes: Extracting motion information of dynamic obstacles from the time-series BEV features, and extracting spatial position information of dynamic obstacles, spatial position information of static obstacles, and environmental information from the multimodal fusion features after spatial alignment at the current moment; Constructing a cost map based on the motion information of the dynamic obstacle, the spatial position information of the dynamic obstacle, and the spatial position information of the static obstacle; Obtain the current position information and target point position information of the quadruped robot; According to the current position information of the quadruped robot and the target point position information, the obstacle avoidance path of the quadruped robot is determined from the cost map, and according to the environmental information, the gait parameters of the quadruped robot at each landing point on the obstacle avoidance path are determined.
3. The method according to claim 2, characterized in that Determining the gait parameters of the quadruped robot at each landing point on the obstacle avoidance path based on the environmental information includes: Get the speed information of the quadruped robot; determining a landing point of the quadruped robot on the obstacle avoidance path according to the speed information of the quadruped robot; Determining environmental information corresponding to a landing point of the quadruped robot on the obstacle avoidance path; According to the environmental information corresponding to the landing points of the quadruped robot on the obstacle avoidance path, the gait parameters of the quadruped robot at each landing point on the obstacle avoidance path are determined.
4. The method according to claim 1, wherein Determining the BEV characteristics corresponding to the current moment based on the data collected by the multimodal sensor includes: The data collected by the multimodal sensor is input into the MLP to perform nonlinear fusion on the data collected by the multimodal sensor and map it to the BEV space to obtain the BEV features corresponding to the current moment.
5. The method according to claim 1, wherein The determining of the time series BEV features according to the BEV features corresponding to the historical moments and the BEV features corresponding to the current moment includes: The BEV features corresponding to the historical moments and the BEV features corresponding to the current moment are input into a Transformer encoder to obtain a temporal BEV feature determined by temporal splicing, wherein the temporal BEV feature records the motion information of the dynamic obstacle.
6. The method according to claim 1, characterized in that Determining the multimodal fusion features after spatial alignment at the current moment based on the data collected by the multimodal sensor includes: The data collected by the multimodal sensor is input into the MLP to fuse the data collected by the multimodal sensor to obtain the multimodal fusion features after spatial alignment at the current moment, wherein the multimodal fusion features after spatial alignment at the current moment record the spatial position information of dynamic obstacles, the spatial position information of static obstacles, and environmental information.
7. The method according to any one of claims 1 to 6, characterized in that: The multimodal sensor also includes a fisheye camera, a laser radar and a telephoto camera. The fisheye camera is arranged on the left, right and rear sides of the quadruped robot, the laser radar is arranged on the top of the quadruped robot, and the telephoto camera is arranged on the front side of the quadruped robot.
8. A device for determining motion information of a quadruped robot, characterized in that: include: A first acquisition module is configured to acquire data collected by a multimodal sensor, wherein the multimodal sensor includes a binocular camera and a depth camera, the binocular camera is disposed on the front side of the quadruped robot, and the depth camera is symmetrically disposed on the front and rear sides of the quadruped robot; A BEV feature determination module, configured to determine the BEV feature corresponding to the current moment based on the data collected by the multimodal sensor; The second acquisition module is used to obtain the BEV characteristics corresponding to the historical moment; a time-series BEV feature determination module, configured to determine a time-series BEV feature based on the BEV feature corresponding to the historical moment and the BEV feature corresponding to the current moment; A multimodal fusion feature determination module, configured to determine the multimodal fusion features after spatial alignment at the current moment based on the data collected by the multimodal sensor; The motion information determination module is used to determine the obstacle avoidance path and gait parameters of the quadruped robot based on the time-series BEV features and the multimodal fusion features after spatial alignment at the current moment.
9. A device for determining motion information of a quadruped robot, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the computer program is used to control the processor to operate so as to perform the method according to any one of claims 1 to 7.
10. A quadruped robot, characterized in that: It includes the quadruped robot motion information determination device as described in claim 8 or 9.
Citation Information
Patent Citations
Autonomous exploration and track monitoring method for quadruped robot
CN115356743A
Robot motion planning method and device based on multi-modal information fusion
CN115617036A
Target detection method and device, electronic equipment and medium
CN116844133A
Quadruped robot complex scene sensing method and system based on multimode fusion
CN117830991A
Autonomous navigation method based on aerial view space-time contrast reinforcement learning
CN117968703A
Cited By
Quadruped robot parkour navigation method and system based on multi-modal feature fusion
CN121384037A
Four-legged robot parkour navigation method and system based on multi-modal feature fusion
CN121384037B