Robot Observation Data Completion Model Training Method and Device
The loss function optimization of the ladder terrain elevation diagram through a two-dimensional convolutional neural network solves the problem of poor ladder completion effect in the existing technology, and realizes efficient observation data completion and precise motion control under CPU conditions.
Patent Information
- Application Number
- CN202510527006.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing technology is in the completion of robot observation data, especially the elevation map completion and identification of ladder terrain, which leads to falls or safety accidents during robot motion control. The existing methods are expensive to calculate and rely on GPU, making it difficult to effectively deploy under CPU conditions.
A two-dimensional convolutional neural network is used to complete the step topographic elevation map containing the broken areas. By calculating the loss function of the elevation map and the step edge map, it is combined with noise data training to reduce the calculation complexity and improve the learning accuracy of the terrain boundary feature.
Under CPU conditions, efficient observation data completion and ladder recognition have been achieved, which improves the robot's perception ability and motion control accuracy of complex terrain, reduces the computational complexity and enhances the deployability of the model.
Smart Images

Figure CN120068952B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of sensors and robotics, and relates to a method and device for training a robot observation data completion model. Background Art
[0002] In robot perception and motion planning, high-precision environmental observation data is crucial for adaptability to complex terrains. The Elevation Map is an important data structure for robot environmental perception and can be used for tasks such as path planning, obstacle avoidance, and stable walking. However, due to factors such as sensor view occlusion and measurement errors, the elevation maps collected by robots usually have missing data and need to be restored by a completion algorithm to improve the integrity and accuracy of environmental modeling.
[0003] The paper "Neural Scene Representation for Locomotion on Structured Terrain" (arXiv:2206.08077v1 [cs.RO], 2022.06.16) proposed a method for calculating elevation maps based on point cloud completion. This method first fills in the surrounding point cloud data and then calculates the elevation map using the filled-in point cloud data. While maintaining the original spatial structure, this solution can complete complex three-dimensional details and improve the integrity of the elevation map. However, this method has problems such as large computational overhead and long training time, and relies on a GPU for inference. Currently, most robots still mainly use CPU computing and do not have GPU devices installed. Therefore, the applicability of this method in practical applications is limited.
[0004] In some existing technologies, such as the Chinese patent applications with publication numbers CN117788546A and CN118552806A, the point cloud data is converted into an actual depth map and directly input into a pre-trained U-Net depth estimation model to predict the estimated depth map of the current scene for depth completion. Taking CN117788546A as an example, to train the depth estimation model, first a simulation scene is constructed, and RGB images and their corresponding depth information are collected using a virtual camera to form a training set. Subsequently, pre-training is performed on the depth estimation backbone network based on the U-Net structure, and the model is optimized in combination with a dedicated loss function to make it better adapt to different object materials and lighting conditions. Among them, the calculation of the loss function is based on the matching degree between the estimated depth map and the actual depth map, and the model parameters are adjusted by minimizing the error between the two.
[0005] Compared with the point cloud level completion method in the above paper, the technical solution in the above patent directly uses a pre-trained two-dimensional convolutional neural network with a U-Net structure to complete the elevation map, which effectively reduces the amount of calculation and reduces the parameter scale of the network model, thereby being able to complete the inference calculation under the condition of only CPU, thereby improving the deployability of the solution on the robot.
[0006] Then, the inventor discovered that the technical solution in the above patent still has room for improvement in the completion effect of elevation maps containing stepped terrain and the stepped terrain recognition effect, which may lead to falls, errors and even personal safety accidents due to the inability to obtain accurate perception data during robot motion control. Summary of the invention
[0007] The present invention provides a robot observation data completion model training method and device for improving robot observation data completion and step terrain recognition effects.
[0008] Additional aspects and advantages of the disclosure will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the disclosure.
[0009] According to a first aspect of the present disclosure, a method for training a robot observation data completion model is provided, comprising:
[0010] Input the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain the completed elevation map and the predicted step edge map;
[0011] The first loss is calculated using the completed elevation map and the actual complete elevation map; the second loss is calculated using the predicted step edge map and the actual step edge map;
[0012] Back-propagation optimization of the 2D convolutional neural network is performed based on the first loss and the second loss.
[0013] In an exemplary embodiment of the present disclosure, the method further includes:
[0014] Obtaining point cloud data of stepped terrain observation;
[0015] Based on the stepped terrain observation point cloud data, a stepped terrain elevation map including the incomplete area is obtained.
[0016] In an exemplary embodiment of the present disclosure, a stepped terrain elevation map including an incomplete area is obtained according to stepped terrain observation point cloud data, including:
[0017] Converting the stepped terrain observation point cloud data into a first elevation map;
[0018] The first elevation map is downsampled to obtain a stepped terrain elevation map including the incomplete area; the resolution of the stepped terrain elevation map including the incomplete area is lower than the resolution of the first elevation map.
[0019] In an exemplary embodiment of the present disclosure, the resolution of the first elevation map is a×b, where 0.005m≤a≤0.02m, 0.005m≤b≤0.02m; the resolution of the stepped terrain elevation map containing the incomplete area is (n·a)×(n·b), where 4≤n≤10.
[0020] In an exemplary embodiment of the present disclosure, the resolution of the first elevation map is 0.01 m×0.01 m, and the resolution of the stepped terrain elevation map including the incomplete area is 0.05 m×0.05 m.
[0021] In an exemplary embodiment of the present disclosure, the method further includes:
[0022] Noise data is added to the step terrain observation point cloud data; wherein the noise data includes randomly added plane point clouds, and the area of the plane point clouds is positively correlated to the foot coverage area of the robot.
[0023] In an exemplary embodiment of the present disclosure, the noise data further includes: Gaussian noise, a plurality of randomly added external points around a certain point, and / or randomly cut data.
[0024] In an exemplary embodiment of the present disclosure, the two-dimensional convolutional neural network model includes:
[0025] A first encoder is used for performing dimensionality reduction processing on the stepped terrain elevation map containing the incomplete area to extract image features;
[0026] The first decoder is used to perform dimensionality restoration according to image features; and the decoder is connected to a depth head for outputting a completed elevation map and a segmentation head for outputting a predicted step edge map.
[0027] In an exemplary embodiment of the present disclosure, the predicted step edge map output by the segmentation head is also used to input to the depth head.
[0028] In an exemplary embodiment of the present disclosure, the two-dimensional convolutional neural network model includes:
[0029] The second encoder is used to perform dimensionality reduction processing on the stepped terrain elevation map containing the incomplete area to extract image features;
[0030] The second decoder is used to restore the dimension according to the image features and output the completed elevation map;
[0031] The third decoder is used to restore the dimension according to the image features and output a predicted step edge map.
[0032] In an exemplary embodiment of the present disclosure, calculating a first loss by using the completed elevation map and the actual complete elevation map includes:
[0033] According to:
[0034]
[0035] Calculate to obtain the first loss;
[0036] Wherein, is the first loss, is the true height value of the i-th pixel in the actual complete elevation map, is the predicted height value of the i-th pixel in the completed elevation map, and n is the number of all pixel points in the actual complete elevation map and the completed elevation map.
[0037] In an exemplary embodiment of the present disclosure, calculating a second loss by using the predicted step edge map and the actual step edge map includes:
[0038] According to:
[0039]
[0040] Calculate to obtain the second loss;
[0041] Wherein, BCE Loss is the second loss, is the true label corresponding to the i-th pixel in the actual step edge map, is the probability that the i-th pixel in the predicted step edge map belongs to the edge, and N is the number of all pixel points in the predicted step edge map and the actual step edge map.
[0042] In an exemplary embodiment of the present disclosure, backpropagation optimization is performed on the two-dimensional convolutional neural network according to the first loss and the second loss, including:
[0043] Calculate the comprehensive loss according to the first loss and the second loss;
[0044] Taking the minimum of the comprehensive loss as the goal, perform backpropagation optimization on the model parameters in the two-dimensional convolutional neural network.
[0045] In an exemplary embodiment of the present disclosure, the comprehensive loss includes a first comprehensive loss; calculating the comprehensive loss according to the first loss and the second loss includes:
[0046] According to:
[0047]
[0048] Calculate to obtain the first comprehensive loss;
[0049] Among them, is the first comprehensive loss, is the first loss, BCE Loss is the second loss, and are weights.
[0050] In an exemplary embodiment of the present disclosure, the comprehensive loss includes a second comprehensive loss; calculating the comprehensive loss according to the first loss and the second loss includes:
[0051] Generating a corresponding weight mask according to the image region whose height value exceeds a preset height threshold;
[0052] Adjusting the first loss and the second loss using the weight mask, and calculating the second comprehensive loss according to the adjusted first loss and second loss.
[0053] In an exemplary embodiment of the present disclosure, adjusting the first loss and the second loss using the weight mask, and calculating the second comprehensive loss according to the adjusted first loss and second loss includes:
[0054] According to:
[0055]
[0056] Calculating to obtain the second comprehensive loss;
[0057] Among them, is the second comprehensive loss, Mask(i) is the weight mask, Mask(i)=α(0<α<1) means assigning a loss weight α to the i-th pixel point whose height value exceeds the preset height threshold, is the first loss, is the first loss adjusted using the weight mask is the second loss, is the second loss adjusted using the weight mask and are weights.
[0058] According to the second aspect of the present disclosure, a method for completing robot observation data is provided, including:
[0059] Collecting a target stepped terrain elevation map;
[0060] Inputting the target stepped terrain elevation map into a pre-trained robot observation data completion model to obtain a completed elevation map and a predicted stepped edge map;
[0061] Among them, the pre-trained robot observation data completion model is obtained according to the robot observation data completion model training method in the first aspect of the present disclosure.
[0062] According to a third aspect of the present disclosure, there is provided a robot motion control method, comprising:
[0063] The completed elevation map and the predicted step edge map are input into the robot motion control model pre-trained by the deep reinforcement learning algorithm to obtain the motion control parameters;
[0064] Control the robot to go up or down the stairs according to the motion control parameters;
[0065] Among them, the completed elevation map and the predicted step edge map are obtained according to the robot observation data completion method in the second aspect of the present disclosure.
[0066] According to a fourth aspect of the present disclosure, a robot observation data completion model training device is provided, comprising:
[0067] An image processing module is used to input the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain a completed elevation map and a predicted stepped edge map;
[0068] A loss calculation module, used to calculate a first loss using the completed elevation map and the actual complete elevation map; and calculate a second loss using the predicted step edge map and the actual step edge map;
[0069] The network optimization module is used to perform back-propagation optimization on the two-dimensional convolutional neural network according to the first loss and the second loss.
[0070] According to a fifth aspect of the present disclosure, a robot observation data completion device is provided, comprising:
[0071] An image acquisition module is used to acquire a target stepped terrain elevation map;
[0072] A data completion module is used to input the target step terrain elevation map into the pre-trained robot observation data completion model to obtain the completed elevation map and the predicted step edge map;
[0073] Among them, the pre-trained robot observation data completion model is obtained according to the robot observation data completion model training method in the first aspect of the present disclosure.
[0074] According to a sixth aspect of the present disclosure, there is provided a robot motion control device, comprising:
[0075] A parameter determination module, used to input the completed elevation map and the predicted step edge map into a robot motion control model pre-trained by a deep reinforcement learning algorithm to obtain motion control parameters;
[0076] A motion control module for controlling a robot to climb or descend stairs according to motion control parameters;
[0077] Among them, the completed elevation map and the predicted stair edge map are obtained according to the robot observation data completion method in the second aspect of the present disclosure.
[0078] According to a seventh aspect of the present disclosure, there is provided an electronic device, including:
[0079] A processor; and
[0080] A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods of the above embodiments are implemented.
[0081] According to an eighth aspect of the present disclosure, there is provided a robot, including:
[0082] A processor; and
[0083] A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods of the above embodiments are implemented.
[0084] In an exemplary embodiment of the present disclosure, the robot includes any one of a legged robot, a quadruped robot, a biped robot, a wheeled robot, a wheel-legged robot, a four-wheeled legged robot, a humanoid robot, a cleaning robot, a transportation robot, a mobile robot, and a robotic arm.
[0085] According to a ninth aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program code instructions are stored, and when the computer program code instructions are called by a processor of a robot, the robot is caused to execute the methods of the above embodiments.
[0086] From the above technical solutions, it can be seen that the present disclosure has at least one of the following advantages and positive effects:
[0087] The robot observation data completion model training method disclosed in the present invention inputs the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained, outputs the completed elevation map and the predicted step edge map, and uses the completed elevation map and the actual complete elevation map to calculate the first loss, and uses the predicted step edge map and the actual step edge map to calculate the second loss, and then optimizes the two-dimensional convolutional neural network according to the first loss and the second loss to obtain the robot observation data completion model. Compared with the existing technology, on the one hand, it avoids the point cloud completion link, and directly completes the completion based on the elevation map, making the data preprocessing process simpler, reducing the overall calculation complexity, and reducing the parameter scale of the network model, so that efficient reasoning can be completed even with only a CPU, improving the deployability of the model on the robot. On the other hand, the present invention also introduces step edge information as an auxiliary supervision signal. During the training process, the elevation completion error and the step edge error are calculated for joint optimization, thereby guiding the model to learn more accurate terrain boundary features, ensuring that the completed data conforms to the large-scale terrain structure while maintaining the integrity of key local features. Therefore, both the observation data completion effect and the step terrain recognition effect will be better, which can provide better data support for robot motion control and achieve precise control. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0089] Figure 1 A system architecture diagram is shown to which the robot observation data completion model training method, the robot observation data completion method, and the robot motion control method in the embodiments of the present disclosure can be applied.
[0090] Figure 2 A flow chart of a robot observation data completion model training method in an embodiment of the present disclosure is shown.
[0091] Figure 3 A schematic diagram of a stepped terrain observation point cloud in an embodiment of the present disclosure is shown.
[0092] Figure 4 Another schematic diagram of a stepped terrain observation point cloud in an embodiment of the present disclosure is shown.
[0093] Figure 5 A schematic diagram of another stepped terrain observation point cloud in an embodiment of the present disclosure is shown.
[0094] Figure 6 A schematic diagram of a process for obtaining a stepped ground elevation map including an incomplete area in an embodiment of the present disclosure is shown.
[0095] Figure 7 A schematic diagram of the structure of a two-dimensional convolutional neural network in an embodiment of the present disclosure is shown.
[0096] Figure 8 A schematic diagram of the structure of another two-dimensional convolutional neural network in an embodiment of the present disclosure is shown.
[0097] Figure 9 A schematic diagram of the structure of another two-dimensional convolutional neural network in an embodiment of the present disclosure is shown.
[0098] Figure 10 A schematic diagram of another stepped terrain observation point cloud in an embodiment of the present disclosure is shown.
[0099] Figure 11 A schematic diagram of a stepped ground elevation map in an embodiment of the present disclosure is shown.
[0100] Figure 12 A schematic diagram of an actual complete elevation map in an embodiment of the present disclosure is shown.
[0101] Figure 13 A schematic diagram of a completed elevation map in an embodiment of the present disclosure is shown.
[0102] Figure 14 A schematic diagram of an actual step edge graph in an embodiment of the present disclosure is shown.
[0103] Figure 15 A schematic diagram of a predicted step edge graph in an embodiment of the present disclosure is shown.
[0104] Figure 16 A schematic diagram of a two-dimensional convolutional neural network optimization process in an embodiment of the present disclosure is shown.
[0105] Figure 17 A schematic flow chart of a robot observation data completion method in an embodiment of the present disclosure is shown.
[0106] Figure 18 A schematic flow chart of a robot motion control method in an embodiment of the present disclosure is shown.
[0107] Figure 19 A block diagram of a robot observation data completion model training device in an embodiment of the present disclosure is shown.
[0108] Figure 20 A block diagram of a robot observation data completion device in an embodiment of the present disclosure is shown.
[0109] Figure 21 The block diagram of a robot motion control device in an embodiment of the present disclosure is shown.
[0110] Figure 22 A schematic diagram of a robot in an embodiment of the present disclosure is shown.
[0111] Figure 23 Another schematic diagram of a robot in an embodiment of the present disclosure is shown.
[0112] Figure 24 Still another schematic diagram of a robot in an embodiment of the present disclosure is shown.
[0113] Figure 25 The schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure is shown.
[0114] Figure 26 The schematic diagram of a computer-readable storage medium in some embodiments of the present disclosure is shown. Detailed implementation manners
[0115] In the description of the present disclosure, the terms "first" and "second" are only used for description and do not indicate relative importance or imply the number of technical features. Therefore, the features of "first" and "second" may explicitly or implicitly include at least one of such features. The meaning of "a plurality" is at least two, unless otherwise clearly defined.
[0116] Figure 1 The system architecture diagram to which the robot observation data completion model training method, the robot observation data completion method, and the robot motion control method in the embodiments of the present disclosure can be applied is shown.
[0117] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a robot 102, a network 103, and a server 104. Among them, the terminal device 101 includes, but is not limited to, a desktop computer, a portable computer, a smart phone, a tablet computer, and the like. The terminal device 101 may serve as an interaction interface, provide a visualization function to display the running state of the robot 102, the elevation map completion result, the motion trajectory, etc., and at the same time support sending motion control instructions or adjusting motion control parameters to the robot 102.
[0118] The robot 102 has the capabilities of observation data completion, motion control, and neural network model training, and can autonomously execute the motion tasks of going up or down the stairs. For example, the robot 102 receives sensor inputs and performs elevation map completion, stair edge detection, and motion parameter inference based on the local CPU. Since the embodiments of the present disclosure avoid the point cloud completion link and reduce the computational complexity, the CPU can still complete efficient inference, so as to ensure that the robot 102 can completely rely on its own intelligent control system to adjust the motion state during the task execution process.
[0119] The server 104 can be used to train the robot observation data completion model and update the robot motion control model, and regularly send the optimized model parameters to the robot 102 to improve its motion control performance. At the same time, the server 104 supports model lightweight processing, so that the optimized model can adapt to the computing power of the robot 102, thereby ensuring the effective deployment and operation of the model on the robot 102.
[0120] The network 103 is used as a medium to provide a communication link between the terminal device 101, the robot 102 and the server 104. The network 103 may include various connection types, such as wired or wireless communication links or optical fiber cables. By connecting various devices, the network 103 ensures that the robot 102 can obtain the latest model update from the server 104 when necessary, and ensures that the robot 102 can perform independent reasoning locally, reduce dependence on the server and reduce data transmission requirements, and improve the real-time and deployability of the system. The robot 102 in the system architecture 100 can still operate efficiently in an environment with limited computing resources, while reducing the overall computing and communication costs of the system.
[0121] It should be understood that Figure 1 The number and type of terminal devices, robots, networks and servers in the embodiment are only for illustration purposes. Any number and type of terminal devices, robots, networks and servers may be provided as required.
[0122] The present disclosure provides a method for training a robot observation data completion model. Figure 2 As shown, the method may include the following steps S201 to S203:
[0123] Step S201, inputting the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain a completed elevation map and a predicted stepped edge map;
[0124] Step S202, calculating a first loss using the completed elevation map and the actual complete elevation map; calculating a second loss using the predicted step edge map and the actual step edge map;
[0125] Step S203, performing back propagation optimization on the two-dimensional convolutional neural network according to the first loss and the second loss.
[0126] Compared with the prior art, the robot observation data completion model training method provided by the example implementation of the present disclosure avoids the point cloud completion link and directly completes the completion based on the elevation map, making the data preprocessing process simpler, reducing the overall computational complexity, and reducing the parameter scale of the network model, so that efficient reasoning can be completed even with only a CPU, improving the deployability of the model on the robot. On the other hand, the present disclosure also introduces step edge information as an auxiliary supervision signal, and during the training process, the elevation completion error and the step edge error are calculated for joint optimization, thereby guiding the model to learn more accurate terrain boundary features, ensuring that the completed data not only conforms to the large-scale terrain structure, but also maintains the integrity of key local features; therefore, both the observation data completion effect and the step terrain recognition effect will be better, thereby providing better data support for robot motion control and achieving precise control.
[0127] Next, the robot observation data completion model training method in this example embodiment will be described in detail.
[0128] In step S201, the stepped terrain elevation map including the incomplete area is input into the two-dimensional convolutional neural network to be trained to obtain a completed elevation map and a predicted stepped edge map.
[0129] In the example implementation of the present disclosure, the step terrain elevation map refers to a two-dimensional array that represents the height distribution of the step terrain in a grid manner, wherein the value of each grid cell represents the ground height of the area. The step terrain is usually composed of multiple steps, each of which has a large area in the horizontal direction and a more obvious height change in the vertical direction. Therefore, the step terrain elevation map will show an obvious hierarchical structure, that is, the grid cells on the same step have similar height values, while the grid cells at the edge of the step show a sudden height difference. The step terrain elevation map can help the robot perceive the terrain undulations, enable the robot to recognize the hierarchical structure of the steps, and thus make reasonable decisions in path planning and motion control.
[0130] A stepped terrain elevation map with incomplete areas refers to an elevation map where elevation data of some areas is missing due to sensor acquisition limitations, robot view occlusion, or environmental obstacles. When actually acquiring elevation data, the height values of some grid cells cannot be measured correctly due to factors such as sensor view, low-reflectivity materials, ambient light changes, and obstacle occlusion, thus forming incomplete areas. Especially in stepped terrain, the missing areas are more obvious due to sudden height changes, sharp transitions at the edge of the steps, and perspective issues. For stepped terrain elevation maps with incomplete areas, some grid cells do not have valid height values, which can be represented by default values such as NaN or zero.
[0131] When the stepped terrain elevation map contains incomplete regions, it will affect robot navigation, path planning, and terrain analysis. Therefore, in the exemplary embodiments of the present disclosure, the stepped terrain elevation map can be completed through a neural network model, so that the incomplete regions can be reasonably inferred based on the known elevation data around, thereby generating a complete elevation map, and synchronously identifying and retaining the stepped edge information to ensure the robot's accurate perception and adaptation ability to the terrain.
[0132] Correspondingly, the completed elevation map refers to a complete elevation map generated by using a neural network model to complete the missing height data of the stepped terrain elevation map. The predicted stepped edge map refers to a binary image predicted by a neural network for identifying the position of the staircase edge, where the value of each pixel point in the stepped edge map represents the probability that the point belongs to the stepped edge. For example, a binary representation of 0 and 1 can be used, where 1 represents that the point is the edge region of the stepped terrain, and 0 represents the non-edge region.
[0133] It should be noted that the stepped edge information is crucial for robot navigation. It can help the robot identify the walkable area, detect obstacles, and adjust its gait to adapt to complex terrain changes. Moreover, the edge prediction function of the neural network helps to enhance the clarity of the step structure and improve the robot's terrain recognition accuracy.
[0134] The two-dimensional convolutional neural network to be trained is a deep learning model for processing image-like data. It can extract local features through convolutional operations and learn higher-level representations layer by layer. Exemplarily, the two-dimensional convolutional neural network can adopt the U-Net architecture, which has a small computational amount and is suitable for CPU computing. The U-Net architecture consists of an encoder and a decoder. Among them, the encoder is composed of multiple convolutional layers and pooling layers, which are responsible for feature extraction of the input stepped terrain elevation map containing incomplete regions and reducing the spatial dimension, enabling the model to learn higher-level terrain features. The decoder is composed of multiple deconvolutional layers, which are used to gradually restore the spatial dimension of the elevation map and generate a completed elevation map with the same size as the input. Importantly, the decoder in the exemplary embodiments of the present disclosure can also generate a stepped edge map.
[0135] For example, a completed elevation map is output by one decoder, and at the same time, this decoder also additionally includes an independent output layer for outputting the stepped edge map. For another example, the U-Net architecture includes one encoder and two decoders, and the completed elevation map and the predicted stepped edge map are output in parallel by the two decoders. The specific structure of the U-Net architecture in the embodiments of the present disclosure is not limited, as long as it can output a completed elevation map and a predicted stepped edge map.
[0136] Of course, according to actual needs, the attention mechanism can also be introduced on the U-Net architecture. The two-dimensional convolutional neural network can also adopt ResNet (residual network) + FPN (Feature Pyramid Network), DeepLabV3+ (image semantic segmentation) network, etc., which is not limited in this disclosure.
[0137] Before inputting the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained, it is necessary to first obtain the stepped terrain elevation map containing the incomplete area. Exemplarily, the stepped terrain elevation map containing the incomplete area can be obtained by obtaining the stepped terrain observation point cloud data and obtaining the stepped terrain elevation map containing the incomplete area according to the stepped terrain observation point cloud data.
[0138] Specifically, the robot's laser radar, depth camera or structured light sensor and other devices are used to scan the environment around the robot to obtain three-dimensional point cloud data containing stepped terrain, that is, step terrain observation point cloud data. Among them, three-dimensional point cloud data is a set of discrete coordinate points, each point contains three spatial coordinate information of X, Y, and Z, and can also contain attributes such as reflectivity, color or timestamp, which is used to describe the three-dimensional shape of the terrain.
[0139] refer to Figures 3 - 5 As shown in the figure, three different step terrain observation point cloud diagrams are given. Figure 3 The observed point cloud of the stepped terrain shown is sparse, and point data is missing in some areas; Figure 4 Compared with the observed point cloud of stepped terrain shown in Figure 3 The step terrain observation point cloud shown is denser and has color changes to distinguish the height differences between point data; Figure 3 , Figure 4 The observed point cloud of stepped terrain is shown. Figure 5 The observation point cloud of the stepped terrain shown is the densest, and the color gradient changes are more obvious, that is, the height information is clearer and more complete.
[0140] After obtaining the point cloud data of the stepped terrain observation, preprocessing operations such as noise filtering, ground segmentation and coordinate transformation can be performed. Among them, noise filtering refers to the use of statistical methods to remove abnormal points introduced by sensor errors or environmental interference to make the point cloud smoother and more reliable. Statistical methods include one or more of statistical filtering, mean filtering, Gaussian filtering, K-nearest neighbor average filtering, etc. For example, statistical filtering can be used to remove outliers first, and then combined with Gaussian filtering to smooth the data to ensure that the generated elevation map is more stable and accurate. Ground segmentation refers to removing non-topographic points (such as walls, railings, pedestrians) and retaining the real stepped terrain data. Coordinate transformation refers to converting the point cloud from the sensor coordinate system to the world coordinate system or the robot's own coordinate system so that the data can be calculated under a unified reference system.
[0141] In an exemplary embodiment, referring to Figure 6 as shown, the process of converting the preprocessed stepped terrain observation point cloud data into a stepped terrain elevation map including incomplete regions includes step S601 and step S602:
[0142] Step S601: Convert the stepped terrain observation point cloud data into a first elevation map.
[0143] Among them, the first elevation map is a regularly gridded two-dimensional array, and its accuracy is determined by the set grid size. A smaller grid size means a higher spatial resolution, making the terrain details clearer. Therefore, in step S601, a fixed grid resolution can be set first, such as 0.01m × 0.01m, and the stepped terrain observation point cloud data is projected onto the two-dimensional grid. Because the high-resolution grid is finer, even if the point cloud points are unevenly distributed, each grid will contain more accurate points, and the value of each grid cell represents the ground height at that position. Since the point cloud data contains multiple irregularly distributed points, the method of taking the highest point can be adopted during the conversion process, that is, the maximum value of the Z coordinate is selected as the elevation value of the grid within each grid range. Thus, while converting the stepped terrain observation point cloud data into the first elevation map, key terrain features such as the edges of the steps can be ensured to be retained.
[0144] Step S602: Perform downsampling on the first elevation map to obtain a stepped terrain elevation map including incomplete regions; the resolution of the stepped terrain elevation map including incomplete regions is lower than that of the first elevation map.
[0145] After obtaining the first elevation map, in order to reduce the computational complexity and simulate the resolution limitation of real sensors, it is necessary to perform downsampling on the first elevation map, that is, reduce the spatial resolution of the data to obtain a stepped terrain elevation map including incomplete regions.
[0146] In the exemplary embodiment of the present disclosure, the resolution of the first elevation map is a × b, where 0.005m ≤ a ≤ 0.02m, 0.005m ≤ b ≤ 0.02m. Such a resolution setting can fully retain the detailed information of the terrain surface. The resolution of the stepped terrain elevation map including incomplete regions is (n·a) × (n·b), where 4 ≤ n ≤ 10, ensuring that the resolution of the stepped terrain elevation map is lower than that of the first elevation map, meeting the low-resolution simulation requirements, and since the value of n ranges from 4 to 10, the accuracy difference between the incomplete map and the complete map can be flexibly controlled.
[0147] Exemplarily, the resolution of the first elevation map is 0.01m × 0.01m. The process of downsampling the first elevation map refers to re - dividing the grid according to a larger grid size, such as 0.05m × 0.05m, and calculating the corresponding elevation value by taking the highest point within each new grid cell. This can retain higher terrain change features, such as the edge information of steps, while removing some details. During the downsampling process, since a larger grid covers multiple points with different elevations, some areas are missing because there are too few data points or no points fall into them, thus forming a stepped terrain elevation map containing incomplete areas.
[0148] For example, the height distribution of the point cloud in the first elevation map is as follows: 0.12 0.15 0.18 0.20 0.11 0.14 0.17 0.19 0.10 0.13 0.16 0.18 0.10 0.12 0.15 0.17
[0153] Among them, the resolution of the first elevation map is 0.01m × 0.01m, that is, each grid cell represents an area of 0.01 meters, and the size of the first elevation map is 4×4.
[0154] By merging the 4×4 first elevation map into a 2×2 one and performing downsampling by taking the highest point, the height distribution of the point cloud in the stepped terrain elevation map containing incomplete areas is as follows: 0.15 0.20 0.13 0.18
[0157] Among them, the resolution of the stepped terrain elevation map with incomplete areas is 0.05m × 0.05m, that is, each new grid cell represents an area of 0.05 meters, and the size of this elevation map is 2×2. It can be seen that the highest 0.20 at the staircase edge is retained.
[0158] Since the resolution of the stepped terrain elevation map containing incomplete areas is lower than that of the first elevation map, compared with the stepped terrain observation point cloud data, the spatial information of the finally obtained stepped terrain elevation map containing incomplete areas is sparser, which can effectively reduce the amount of calculation and improve the efficiency of subsequent processing. At the same time, this elevation map still retains the basic contour of the stepped terrain, providing basic data for subsequent elevation completion and terrain analysis.
[0159] In this exemplary embodiment, on the one hand, by converting complex three-dimensional point clouds into regular grid-like two-dimensional data, subsequent processing becomes more efficient and computable. At the same time, converting high-dimensional point cloud data into a two-dimensional representation reduces the computational amount, enabling the robot observation data completion model to run under limited computing power conditions (such as using only a CPU). On the other hand, by constructing an elevation map by selecting the highest points, key terrain information is retained, ensuring the integrity of the step edge and terrain undulation information, thereby improving the robot's recognition ability for complex terrains. On yet another hand, by analyzing the data vacancy parts in the elevation map, it is possible to effectively determine which areas need to be completed and provide a basis for subsequent neural network processing, thus enhancing the robot's adaptability under incomplete terrain data conditions.
[0160] In addition, during the actual deployment process of the robot model, it is found that the planes output by the model often fluctuate. Among them, the plane output by the model refers to the flat area in the elevation map generated by the robot observation data completion model, and this plane corresponds to the flat area of the ground, platform, or step where the robot walks. For example, when the robot's two feet are within the camera's field of view, the point cloud of the foot surface is a plane, but due to the noise in the related technology not considering plane noise, the model identifies the position that was originally the foot surface as a step.
[0161] Therefore, in this exemplary embodiment of the present disclosure, when training the robot observation data completion model, noise data can be added to the step terrain observation point cloud data, such as randomly adding plane point cloud noise. By artificially increasing the point cloud data of a part of the plane area in the simulation environment, it helps the model better learn how to maintain a stable terrain structure and improve the model's adaptability to the real environment.
[0162] It should be noted that the method of adding plane point cloud noise is to randomly generate a certain number of plane point clouds near the robot's walking path according to the foot coverage area of the robot. Among them, the plane point cloud refers to a set of point cloud data with the same height and dense distribution. The Z coordinate values of these point clouds are the same, presenting a local flat area, so as to simulate the influence of the local flat area of the robot's sole or the ground.
[0163] Furthermore, the area size of the plane point cloud is determined by the foot coverage area of the robot. For example, the area of the plane point cloud is positively correlated with the foot coverage area of the robot. The larger the foot coverage area, the larger the area of the added plane point cloud, so as to more realistically simulate the influence of the robot's standing or moving on terrain observation. For example, if the contact area of the robot's foot on the ground is a circular or rectangular area, the distribution range of the plane point cloud can adopt the same shape, and the point cloud data is randomly distributed within this range with a certain probability to simulate the terrain observation error under different gaits or movement conditions.
[0164] It should be noted that the noise data added to the stepped terrain observation point cloud data also includes Gaussian noise, multiple outliers randomly added around a certain point, and / or randomly removed data, so as to further simulate the errors in the real robot observation environment, making the robot observation data completion model more robust to uncertain factors, improving the processing ability of uncertain factors in the real environment, and thus more accurately completing the stepped terrain data.
[0165] Among them, Gaussian noise is to add random perturbations conforming to the normal distribution to the three-dimensional coordinates of the stepped terrain observation point cloud data, so that the coordinate values of each point cloud data are no longer fixed, but fluctuate within a certain range around the true measurement value. Gaussian noise can be used to simulate sensor measurement errors. After adding Gaussian noise, the height values of the data points fluctuate slightly, and then the model can learn how to remove random errors in the observation data during training, improve the error tolerance of the model, and improve the adaptability to the real environment.
[0166] Randomly adding multiple outliers around a certain point means introducing local abnormal data points to enhance the model's ability to identify and correct outliers. Specifically, multiple offset points around a certain center point are randomly generated at certain positions in the point cloud data. The distribution of these offset points can be random or follow a certain pattern, such as uniform distribution or Gaussian distribution, making the point cloud data in some areas sparser or more concentrated. The formed local outliers can be used to simulate the measurement errors generated by the sensor on high-reflectivity or low-reflectivity surfaces. For example, key points in the point cloud data such as stepped edges or ground points are selected, and then multiple virtual points with offset amounts less than the set range are randomly generated around them to form a local abnormal area, so that the model can learn how to ignore these invalid points to improve the stability of data completion.
[0167] Randomly removing data means artificially deleting data points in some areas of the point cloud data to simulate the data loss situation caused by perspective occlusion, environmental interference, or sensor blind spots during the robot's observation process, thereby enhancing the model's recovery ability in the face of large missing areas and enabling it to learn how to infer the reasonable height of the missing area using the surrounding known data during training. There are various ways to randomly remove data. For example, it can be cropped based on regular shapes, such as randomly selecting a rectangular or circular area in the elevation map and deleting the point cloud data in it, or randomly deleting some points in the point cloud data based on probability methods, making the point cloud density uneven. Especially in the processing of stepped terrain point cloud data, randomly removing the observation points of some steps can simulate the situation where the robot cannot observe the stepped edge at certain angles, so that the model can learn to more accurately complete the missing data and improve its prediction ability for complex terrains.
[0168] It should be noted that these three methods of adding noise data are respectively for different types of observation errors and data missing situations, so different noise data can be added according to actual needs, and this disclosure does not limit this. Among them, Gaussian noise can enhance the robustness of the model to sensor measurement errors, randomly adding multiple external points around a certain point can improve the model's ability to identify and process outliers, and randomly removing data can enhance the model's ability to recover from data missing.
[0169] In the example implementation of the present disclosure, by adding noise data to the step terrain observation point cloud data, the diversity of training data can be effectively enhanced, so that the model can more stably predict and complete the elevation map when encountering terrain interference or misdetection, and improve the adaptability to complex terrain, especially in the step edge detection and completion tasks, it can reduce the misjudgment caused by the influence of the robot's feet and improve the generalization ability of the model to the actual observation data. Moreover, by introducing noise data, the model can be made close to the complexity of the real environment during the training process, so that it can show stronger adaptability in actual deployment and improve the recognition and completion accuracy of step terrain.
[0170] refer to Figure 7 As shown, a schematic diagram of the structure of a two-dimensional convolutional neural network in an embodiment of the present disclosure is shown. Figure 7 It can be seen that the two-dimensional convolutional neural network includes a first encoder 701 and a first decoder 702. Among them, the first encoder 701 is used to perform dimensionality reduction processing on the stepped terrain elevation map containing incomplete areas to extract image features; the first decoder 702 is used to restore dimensions according to image features, and the decoder is connected to a depth head 7021 for outputting a completed elevation map and a segmentation head 7022 for outputting a predicted stepped edge map. Among them, the completed elevation map is used to fill in the incomplete areas in the original input to make the generated terrain data more complete and accurate. The predicted step edge map provides additional edge information to assist in optimizing the elevation completion effect and enhance the model's recognition ability of stepped terrain. In addition, features are transferred between the first encoder 701 and the first decoder 702 through jump connections, such as Figure 7 As shown by the dotted arrow in , the feature map output by the first encoder 701 will be passed to the first decoder 702 by the jump connection module for splicing.
[0171] Specifically, the depth head 7021 is mainly responsible for the elevation completion task. By receiving the feature information extracted by the first decoder 702, it outputs a completed elevation map with the same size as the original elevation map. For example, the depth head 7021 uses convolutional layers and regression activation functions to predict the height value of each grid, ensuring that the missing areas are reasonably filled, making the completed terrain data more complete and smooth. Among them, the regression activation function can be the ReLU (Rectified Linear Unit) function, linear activation function, etc.
[0172] The segmentation head 7022 is mainly responsible for the step edge detection task. By processing the features output by the first decoder 702, it generates a binary step edge map. For example, the segmentation head 7022 uses convolutional layers and Sigmoid activation functions or Softmax activation functions to calculate the probability that each pixel belongs to the edge.
[0173] It should be noted that the depth head 7021 and the segmentation head 7022 can be used simultaneously in the training stage and the inference stage of the network. Correspondingly, the model will output the completed elevation map and the predicted step edge map during inference. Of course, according to the actual model accuracy requirements, it is also possible to choose to disable the segmentation head 7022 in the inference stage and only use it during training. At this time, the edge segmentation task is only an auxiliary task in the training stage, used to help the model learn better depth features, but is removed during inference. The model only generates the completed elevation map during inference, simplifying the inference calculation amount.
[0174] Reference Figure 8 As shown, a schematic structural diagram of another two-dimensional convolutional neural network in the embodiments of the present disclosure is shown. Different from the Figure 7 network structure shown, Figure 8 in the network structure shown, features are transmitted between the depth head 7021 and the segmentation head 7022 through skip connections. Therefore, the predicted step edge map output by the segmentation head 7022 is also used as input to the depth head 7021. By using the edge information as auxiliary information for the completed elevation map, providing constraints during the completion process, ensuring that the completion result of the elevation map not only conforms to the overall trend but also maintains clear terrain changes in key edge areas.
[0175] As Figure 7 , Figure 8 shown, the two-dimensional convolutional neural network structure can effectively enhance the accuracy of elevation map completion. Especially when dealing with step terrains, through the guidance of edge information, the model can more accurately restore the height changes of steps, prevent edge blurring or step shape distortion, and improve the navigation and perception capabilities of the robot in complex terrain environments.
[0176] Reference Figure 9As shown, a schematic structural diagram of another two-dimensional convolutional neural network in an embodiment of the present disclosure is shown. From Figure 9 it can be seen that the two-dimensional convolutional neural network includes a second encoder 901, a second decoder 902, and a third decoder 903. The second encoder 901 is used to perform dimensionality reduction on the stepped terrain elevation map containing the incomplete region to extract image features; the second decoder 902 is used to restore the dimension according to the image features and output the completed elevation map; the third decoder 903 is used to restore the dimension according to the image features and output the predicted stepped edge map.
[0177] Among them, the second encoder 901 can be composed of multiple convolutional layers and pooling layers, gradually reducing the spatial dimension and enhancing the feature expression ability, so that the model can learn the overall structure and local details of the terrain. The image features extracted by the second encoder 901 are then transmitted to the second decoder 902 and the third decoder 903 for separate decoding. The second decoder 902 restores the original spatial resolution according to the image features extracted by the second encoder 901 and generates a completed elevation map. The second decoder 902 can use a transposed convolutional layer or upsampling + convolutional layer to gradually expand the size of the feature map, so that the output image matches the size of the input elevation map. Finally, the output of the second decoder 902 is a completed elevation map with the same size as the original elevation map, filling in the height values of the incomplete regions in the input data, making the terrain data more complete and smooth. The third decoder 903 learns the shape of the step edge from the image features extracted by the second encoder 901 and generates a probability distribution map through a Sigmoid activation function or a Softmax classifier. The output edge map is a binary image. In this way, the model can accurately identify the step boundary and improve the accuracy of elevation completion.
[0178] Figure 9 The two-dimensional convolutional neural network model shown in
[0179] adopts a dual-decoder design, that is, the second decoder 902 and the third decoder 903 independently decode the elevation map and the edge map respectively, which can reduce the mutual interference between the two tasks and make the elevation completion and edge prediction more accurate. At the same time, this structure can optimize the completion and segmentation tasks respectively at different levels of feature expression, improving the adaptability of the model to complex terrains.
[0180] In the exemplary embodiments of the present disclosure, a two-dimensional convolutional neural network is optimized through a loss function, enabling the model to accurately complete the elevation map and precisely predict the step edges, thereby enhancing the robot's perception ability of complex terrains. Among them, the first loss refers to the elevation completion loss. The completed elevation map is the result generated by the network, while the actual complete elevation map is the corresponding real terrain data. The purpose of calculating the first loss is to make the completed result output by the model as close as possible to the real elevation. For example, the first loss can adopt the L1 loss (Mean Absolute Error) or the L2 loss (Mean Squared Error). The L1 loss can reduce the influence of outliers in the completion result, while the L2 loss can penalize larger errors. Of course, the L1 loss and the L2 loss can be combined for use to improve the completion accuracy, and the present disclosure does not limit this. The second loss refers to the step edge prediction loss. The predicted step edge map is a binary image output by the network, while the actual step edge map is the edge label obtained through manual annotation or calculated based on the real elevation map. The purpose of calculating the second loss is to ensure that the model can correctly identify the step edges and avoid misjudgment or edge blurring. For example, the binary cross-entropy loss function can be used to calculate the second loss, or the Dice (Dice similarity) loss function, the IoU (Intersection over Union) loss function, etc. can be used to calculate the second loss. The present disclosure does not limit the types of loss functions used for calculating the first loss and the second loss.
[0181] Reference Figure 10 As shown, another schematic diagram of the step terrain observation point cloud is shown.
[0182] Reference Figure 11 As shown, a schematic diagram of a step terrain elevation map is shown. Figure 11 The step terrain elevation map in Figure 10 is processed based on the step terrain observation point cloud shown in Figure 11 As can be seen, this step terrain elevation map contains incomplete regions.
[0183] Reference Figure 12 As shown, a schematic diagram of an actual complete elevation map is shown. Figure 12 The actual complete elevation map in Figure 11 is the complete elevation map corresponding to the step terrain elevation map shown in
[0184] Reference Figure 13 As shown, a schematic diagram of a completed elevation map is shown. Figure 13 The completed elevation map in Figure 11 takes the step terrain elevation map shown in
[0185] Reference Figure 14 As shown, a schematic diagram of an actual stepped edge map is shown.
[0186] Reference Figure 15 As shown, a schematic diagram of a predicted stepped edge map is shown. Figure 15 The predicted stepped edge map in Figure 11 is also an image output by the two-dimensional convolutional neural network in the embodiments of the present disclosure with the stepped terrain elevation map shown as the input.
[0187] Exemplarily, the first loss can be calculated by using the completed elevation map and the actual complete elevation map according to:
[0188] (1)
[0189] The first loss can be calculated; where is the first loss, is the true height value of the i-th pixel in the actual complete elevation map, is the predicted height value of the i-th pixel in the completed elevation map, and n is the number of all pixel points in the actual complete elevation map and the completed elevation map.
[0190] The second loss can be calculated by using the predicted stepped edge map and the actual stepped edge map according to:
[0191] (2)
[0192] The second loss can be calculated; where BCE Loss is the second loss, is the true label corresponding to the i-th pixel in the actual stepped edge map, is the probability that the i-th pixel in the predicted stepped edge map belongs to the edge, and N is the number of all pixel points in the predicted stepped edge map and the actual stepped edge map.
[0193] In this step, by calculating the first loss, the elevation completion ability of the model in the missing area can be optimized, so that the completed elevation map is as close as possible to the actual complete elevation map. Since there are missing areas in the original observation data, the optimization of this loss can guide the model to learn the global and local variation laws of elevation data, thereby improving the accuracy of the completion result, reducing the height mutation or unreasonable terrain change caused by data missing, improving the integrity of terrain data, and enabling the robot to perform path planning and terrain adaptation based on more accurate elevation information. By calculating the second loss, the edge detection ability of the model can be enhanced, so that the predicted step edge map can more accurately identify the true position of the steps. The step edge is an important structural feature that the robot must identify during walking and planning. Precise edge information helps to improve the resolution of the step height, avoid misjudging the starting point or boundary of the steps during the completion process by the model, and thus reduce the error of terrain perception.
[0194] In step S203, the two-dimensional convolutional neural network is optimized by backpropagation according to the first loss and the second loss.
[0195] Among them, the first loss and the second loss are combined, the gradient of the total loss with respect to the model parameters is calculated by the backpropagation algorithm, and the optimizer updates the network weights according to the gradient, so that the model simultaneously optimizes the elevation completion accuracy and the edge structure prediction ability in subsequent iterations, and finally enables the two-dimensional convolutional neural network to collaboratively repair the terrain numerical values and capture the step geometric features.
[0196] Reference Figure 16 As shown, the process of optimizing the two-dimensional convolutional neural network by backpropagation according to the first loss and the second loss may include step S1601 and step S1602:
[0197] In step S1601, the comprehensive loss is calculated according to the first loss and the second loss.
[0198] In an example implementation, the comprehensive loss includes the first comprehensive loss. The first comprehensive loss refers to the overall loss obtained by weighted summing the first loss and the second loss, which is used to optimize the overall performance of the model, enabling it to accurately complete the missing elevation data and accurately identify the step edge information.
[0199] At this time, the first comprehensive loss is calculated according to the first loss and the second loss, and there is:
[0200] (3)
[0201] Among them, is the first comprehensive loss, is the first loss, BCE Loss is the second loss, and is the weight, which is used to balance the influence of the two losses and ensure that the model does not overly favor a certain task. For example, if the model has a higher precision requirement for elevation completion, can be appropriately increased to make the elevation completion loss account for a larger proportion in the optimization process. If the accuracy requirement for edge recognition is higher, can be increased to enhance the learning of edge information.
[0202] As can be seen from formula (3), the optimization objective of the first comprehensive loss is to minimize the value, making the elevation completion more accurate, while ensuring the integrity of edge information, thereby improving the overall performance of the model on complex terrains and enhancing the adaptability of the robot to step terrains.
[0203] In another exemplary embodiment, the comprehensive loss includes a second comprehensive loss, which refers to introducing a weight mask when calculating the comprehensive loss to adjust the first loss and the second loss for specific height regions, thereby optimizing the learning effect of the model in different terrain regions.
[0204] For example, in actual training, it is found that there are deviations in the height prediction of the model at cliff or wall positions. For example, the height of these regions is estimated to be similar to the surrounding horizontal plane, resulting in the situation where the height of the stair plane is erroneously extended to the cliff region in the prediction result. Therefore, a weight mask can be added at the cliff or wall position, and a smaller (but not 0) weight can be set for the losses at these positions through the weight mask.
[0205] At this time, when calculating the comprehensive loss based on the first loss and the second loss, a corresponding weight mask can be generated according to the image region where the height value exceeds the preset height threshold. When the height of a certain region exceeds the set threshold (such as cliff, wall or high step region), the weight for calculating the loss of this region can be adjusted so that it receives stronger attention or less influence in the model optimization process. For example, in regions with higher heights, due to larger errors in point cloud data acquisition, the model is more likely to generate miscompletion phenomena. Therefore, the weight of this region can be reduced to reduce the model's dependence on unreliable regions and prevent terrain distortion caused by miscompletion. For lower steps or key passage regions, the weight of this region can be increased to allow the model to focus on optimizing the elevation completion and edge prediction accuracy of these regions to improve the adaptability of the robot in the actual terrain.
[0206] Furthermore, the first loss and the second loss are adjusted using the weight mask, and the second comprehensive loss is calculated based on the adjusted first loss and second loss. For example, the second comprehensive loss can be calculated according to formula (4):
[0207] (4)
[0208] where is the second comprehensive loss, Mask(i) is the weight mask, and Mask(i) = α (0 < α < 1) means that the loss weight α is assigned to the i-th pixel point whose height value exceeds the preset height threshold. is the first loss, is the first loss adjusted by using the weight mask is the second loss, is the second loss adjusted by using the weight mask and are the weights.
[0209] In this exemplary embodiment, after introducing the weight mask, the calculations of the first loss and the second loss will be dynamically adjusted according to the weight mask, that is, different weights are assigned to different regions during the loss calculation, making the final loss more in line with the actual terrain characteristics. The adjusted first loss and second loss are used to calculate the second comprehensive loss, making the model optimization process more adaptive. It not only focuses on the overall elevation completion and edge detection accuracy but also can perform more reasonable optimization for specific height regions to reduce the miscompletion of abnormal regions and improve the recognition accuracy of key terrains (such as low steps, ramps, and step edges), thereby enhancing the robot's perception ability and walking stability in complex environments.
[0210] In step S1602, with the goal of minimizing the comprehensive loss, the model parameters in the two-dimensional convolutional neural network are optimized by backpropagation.
[0211] In the backpropagation optimization stage, first, the gradient of the comprehensive loss function is calculated, that is, the partial derivatives of it with respect to all parameters in the network are obtained. The gradient calculation follows the chain rule and propagates backward from the loss function to the parameters of each layer, including the weights and bias terms of the convolutional layer. The calculated gradient is used to guide the parameter update. The update method usually adopts an adaptive optimization algorithm, such as Adam (Adaptive Moment Estimation) or SGD (Stochastic Gradient Descent). Among them, the Adam algorithm can automatically adjust the learning rate to improve the training stability, while the SGD algorithm helps to avoid the local optimum problem. The present disclosure does not limit this. By gradually updating the model parameters, the network weights are adjusted, enabling the model to more effectively minimize the comprehensive loss, improve the accuracy of the completed elevation map, and optimize the recognition ability of the step edge.
[0212] During the training process, the back propagation optimization process will continue for multiple iterations until the comprehensive loss converges, that is, the error between the model's prediction results and the real data tends to be stable, indicating that the network has learned the optimal elevation completion and edge detection strategy, or when the preset number of iterations is met, the training of all model parameters is completed. Ultimately, this optimization method enables the network to more accurately restore the elevation information of the missing area, while ensuring the clarity of the step edges, reducing erroneous extension or blur, and improving the robot's navigation and walking capabilities in complex terrain environments.
[0213] In this step, the joint optimization of the first loss and the second loss can achieve a complementary effect, that is, while optimizing the completed elevation, the constraint of edge information is introduced, so that the elevation completion not only conforms to the overall trend, but also can accurately retain the key step structure, making the boundary of the step clearer, and avoiding edge blur or distortion caused by elevation completion. Especially on complex stepped terrain, this joint optimization can improve the stability of elevation completion and make the model more suitable for actual robot navigation tasks. Moreover, it also improves the robustness and generalization ability of the model, so that the model can adapt to stepped terrain in different environments, reduce dependence on single features, improve the accuracy and reliability of overall terrain perception, and further enhance the robot's motion planning and navigation capabilities on complex terrain.
[0214] The exemplary embodiment of the present disclosure also provides a robot observation data completion method, referring to Figure 17 As shown, the method may include the following steps S1701 and S1702:
[0215] Step S1701, collecting a target stepped terrain elevation map;
[0216] The target stepped terrain elevation map is obtained by scanning the stepped terrain in the target environment in real time through sensing devices such as laser radar, depth camera, structured light sensor or stereo vision system carried by the robot, obtaining three-dimensional point cloud data of the area, and converting the three-dimensional point cloud data into a stepped terrain elevation map. The target stepped terrain elevation map can be a stepped terrain elevation map including an incomplete area or a complete stepped terrain elevation map, which is not limited in the present disclosure.
[0217] Step S1702, inputting the target stepped terrain elevation map into the pre-trained robot observation data completion model to obtain a completed elevation map and a predicted stepped edge map.
[0218] Among them, the pre-trained robot observation data completion model is trained according to the robot observation data completion model training method described in detail in steps S201 to S203 of another embodiment of the present disclosure, which will not be elaborated here. The pre-trained robot observation data completion model can process the input elevation map and repair the missing areas caused by sensor occlusion, data loss, or measurement errors. The model first extracts the features of the input elevation map through an encoder and gradually restores the terrain information through a decoder, and finally outputs the completed elevation map and the predicted step edge map. The inference process of the robot observation data completion model can be performed on the CPU to meet the real-time computing requirements of the robot.
[0219] Executing the robot observation data completion method provided by the exemplary embodiment of the present disclosure can directly complete the completion based on the elevation map, making the data preprocessing process simpler, reducing the overall computational complexity, and reducing the parameter scale of the network model. Thus, efficient inference can be completed even under the condition of only a CPU, improving the deployability of the model on the robot. In addition, while performing elevation completion, the model can also predict the step edge map to ensure that the robot can accurately perceive and identify the staircase structure, preventing misjudgment of the starting point of the step or blurred boundaries. The extraction of this edge information can not only optimize the elevation completion result, making the height change of the step clearer, but also assist the robot in more accurate gait control, avoiding safety problems such as unstable walking or falling due to misperceiving the position of the step.
[0220] The exemplary embodiment of the present disclosure also provides a robot motion control method, as shown in Figure 8 The method may include the following steps S1801 to S1802:
[0221] Step S1801, input the completed elevation map and the predicted step edge map into the robot motion control model pre-trained by the deep reinforcement learning algorithm to obtain motion control parameters.
[0222] Among them, the completed elevation map is used as the terrain input of the robot's traveling path, enabling the robot to accurately perceive the height, width, and slope of the steps. At the same time, combined with the predicted step edge map, it ensures that the robot accurately identifies the boundaries of the steps to avoid incorrect gait planning.
[0223] In the exemplary embodiments of the present disclosure, the robot motion control model can be pre-trained using a deep reinforcement learning algorithm. Through a large number of trainings in a simulated environment, the robot learns the optimal gait control strategy under different step heights, slopes, and surface materials. The deep reinforcement learning algorithm can be optimized based on a reward mechanism, that is, the robot tries different gait patterns during training, and rewards or punishments are given according to factors such as the stability, energy consumption, and balance of successfully ascending / descending the steps, so that it forms an optimal motion control strategy during continuous learning. The deep reinforcement learning algorithm can also adopt a teacher-student deep reinforcement learning framework, where the teacher model is a high-precision reinforcement learning model that has been fully trained in a simulated environment or a real environment and can generate optimal robot motion control strategies, mainly including key parameters such as gait adjustment, step length planning, step height control, body posture optimization, and foot landing point prediction. The role of the teacher model is to provide expert demonstrations or optimal strategy guidance for training the student model. The student model learns the strategy provided by the teacher model and performs reinforcement learning optimization during its own training process, so that it can adapt to different environmental changes, such as different step heights, different slopes, and complex terrains.
[0224] During the training process, the student model can imitate the optimal strategy of the teacher through imitation learning or generative adversarial imitation learning, so that the student model can quickly master the basic gait control ability. The student model gradually tries different motion control parameters during the training process and can also be optimized by combining the reward mechanism. For example, when the robot successfully and stably completes the ascending / descending step task, a higher reward is given; if the robot topples, has an unstable gait, or misses a step, etc., a punishment is given, thereby prompting the student model to optimize the gait control strategy. The trained student model can autonomously adjust the gait strategy under different environments and calculate the motion control parameters according to the input completed elevation map and the predicted step edge map.
[0225] Among them, the motion control parameters include key parameters such as gait patterns (such as walking, jumping, slow movement), step length, step height, joint angles, speed, and foot landing points to ensure that the robot can stably execute the step walking task.
[0226] Step S1802, control the robot to ascend or descend the steps according to the motion control parameters.
[0227] After obtaining the motion control parameters, the motion control parameters can be used to adjust the robot's movement behavior on the stairs in real time through the robot control system instructions to ensure that it can successfully complete the task of going up / down the stairs. Specifically, the robot will adjust key links such as foot movement trajectory, center of gravity balance, and posture adjustment according to the terrain data. For example, when going up the stairs, the robot needs to raise its feet to cross the height of the steps and adjust the center of gravity to prevent falling back or falling. When going down the stairs, it is necessary to appropriately lower the gait height and optimize the landing point to ensure a stable landing. In addition, the robot can also combine real-time sensor feedback such as inertial measurement units and pressure sensors to make dynamic adjustments during movement to cope with changes in the environment.
[0228] It should be noted that the completed elevation map and the predicted step edge map are obtained according to the robot observation data completion method described in detail in step S1701 and step S1702 of another embodiment of the present disclosure, and will not be repeated here.
[0229] In the robot motion control method provided in the example implementation of the present disclosure, by using the completed elevation map and the predicted step edge map as input, the robot can accurately identify the stair structure even when there is a lack of sensor data, thereby reducing motion errors caused by incomplete data. Secondly, through the motion control model trained by the deep reinforcement learning algorithm, the robot can autonomously learn the optimal gait, adapt to different stair environments, improve the success rate of passage, and reduce energy consumption. Furthermore, based on the precise control of motion control parameters, it is ensured that the robot can adjust its posture when going up / down the stairs, prevent falls, and improve motion stability. In addition, the method also has strong environmental adaptability and can cope with different types of stairs, including stairs of different heights, widths, slopes, and materials, so that the robot has stronger universality and practical value.
[0230] Furthermore, in the exemplary implementation of the present disclosure, a robot observation data completion model training device is also provided. Figure 19 As shown, the robot observation data completion model training device 1900 may include an image processing module 1901, a loss calculation module 1902 and a network optimization module 1903, wherein:
[0231] The image processing module 1901 is used to input the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain a completed elevation map and a predicted stepped edge map;
[0232] The loss calculation module 1902 is used to calculate the first loss using the completed elevation map and the actual complete elevation map; and calculate the second loss using the predicted step edge map and the actual step edge map;
[0233] The network optimization module 1903 is used to perform backpropagation optimization on the two-dimensional convolutional neural network according to the first loss and the second loss.
[0234] The specific details of each module in the above robot observation data completion model training device have been described in detail in the corresponding robot observation data completion model training method, so they will not be elaborated here.
[0235] In an exemplary embodiment of the present disclosure, a robot observation data completion device is further provided. Refer to Figure 20 As shown, the robot observation data completion device 2000 may include an image acquisition module 2001 and a data completion module 2002, where:
[0236] The image acquisition module 2001 is used to acquire the elevation map of the target stepped terrain;
[0237] The data completion module 2002 is used to input the elevation map of the target stepped terrain into the pre-trained robot observation data completion model to obtain the completed elevation map and the predicted stepped edge map;
[0238] Among them, the pre-trained robot observation data completion model is obtained according to the robot observation data completion model training method in the embodiments of the present disclosure.
[0239] The specific details of each module in the above robot observation data completion device have been described in detail in the corresponding robot observation data completion method, so they will not be elaborated here.
[0240] In an exemplary embodiment of the present disclosure, a robot motion control device is further provided. Refer to Figure 21 As shown, the robot motion control device 2100 may include a parameter determination module 2101 and a motion control module 2102, where:
[0241] The parameter determination module 2101 is used to input the completed elevation map and the predicted stepped edge map into the robot motion control model pre-trained by the deep reinforcement learning algorithm to obtain motion control parameters;
[0242] The motion control module 2102 is used to control the robot to climb or descend the steps according to the motion control parameters;
[0243] Among them, the completed elevation map and the predicted stepped edge map are obtained according to the robot observation data completion method in the embodiments of the present disclosure.
[0244] The specific details of each module in the above robot motion control device have been described in detail in the corresponding robot motion control method, so they will not be elaborated here.
[0245] In the exemplary embodiments of the present disclosure, a robot is further provided. The robot includes a processor and a memory, and computer-readable instructions are stored on the memory. When the computer-readable instructions are executed by the processor, the above-mentioned method is implemented. Among them, the robot includes any one of a legged robot, a quadruped robot, a biped robot, a wheeled robot, a wheel-legged robot, a four-wheel-legged robot, a humanoid robot, a cleaning robot, a transportation robot, a mobile robot, and a robotic arm. Refer to Figures 22 - 24 As shown, schematic diagrams of three different robots are respectively shown.
[0246] Refer to Figure 25 As shown, an electronic device capable of implementing the above-mentioned method is further provided. Among them, the electronic device 2500 includes a processor 2501 and a memory 2502, and computer-readable instructions are stored on the memory 2502. When the computer-readable instructions are executed by the processor 2501, the method in the embodiments of the present disclosure is implemented.
[0247] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is further provided, on which computer program code instructions are stored. When the computer program code instructions are called by the processor of the robot, the robot is enabled to execute the method as in the embodiment.
[0248] Refer to Figure 26 As shown, a program product 2600 for implementing the above-mentioned method according to an embodiment of the present disclosure is described. It can adopt a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0249] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described here can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0250] Finally, the above preferred embodiments are only used to illustrate the technical solutions of the present application and are not restrictive. Although the present application has been described in detail, those skilled in the art should understand that changes in form and details can be made to it without departing from the scope defined by the claims of the present application. The dimensions of the drawings have nothing to do with the specific physical objects, and the physical object dimensions can be changed arbitrarily.
Claims
1. A method for training a robot observation data completion model, characterized in that, Including: Input the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain the completed elevation map and the predicted stepped edge map; wherein, the predicted stepped edge map is a binary image predicted by the neural network for identifying the position of the stepped edge. Calculate the first loss by using the completed elevation map and the actual complete elevation map; calculate the second loss by using the predicted stepped edge map and the actual stepped edge map. Perform backpropagation optimization on the two-dimensional convolutional neural network according to the first loss and the second loss.
2. The method for training a robot observation data completion model according to claim 1, wherein The method further includes: Obtain the stepped terrain observation point cloud data. Obtain the stepped terrain elevation map containing the incomplete area according to the stepped terrain observation point cloud data.
3. The method for training a robot observation data completion model according to claim 2, wherein, The obtaining the stepped terrain elevation map containing the incomplete area according to the stepped terrain observation point cloud data includes: Convert the stepped terrain observation point cloud data into a first elevation map. Perform downsampling processing on the first elevation map to obtain the stepped terrain elevation map containing the incomplete area; the resolution of the stepped terrain elevation map containing the incomplete area is lower than the resolution of the first elevation map.
4. The method for training a robot observation data completion model according to claim 3, wherein The resolution of the first elevation map is a×b, where 0.005m≤a≤0.02m, 0.005m≤b≤0.02m; the resolution of the stepped terrain elevation map containing the incomplete area is (n·a)×(n·b), where 4≤n≤10.
5. The method for training a robot observation data completion model according to claim 4, wherein The resolution of the first elevation map is 0.01m×0.01m, and the resolution of the stepped terrain elevation map containing the incomplete area is 0.05m×0.05m.
6. The method for training a robot observation data completion model according to claim 2, wherein The method further includes: Add noise data to the stepped terrain observation point cloud data; wherein, the noise data includes randomly added planar point clouds, and the area of the planar point clouds is positively correlated with the foot coverage area of the robot.
7. The method for training a robot observation data completion model according to claim 6, wherein The noise data further includes: Gaussian noise, multiple outliers randomly added around a certain point, and / or randomly removed data.
8. The method for training a robot observation data completion model according to claim 1, wherein The two-dimensional convolutional neural network model includes: A first encoder for performing dimensionality reduction processing on the stepped terrain elevation map containing the incomplete area to extract image features. A first decoder for performing dimensionality restoration according to the image features; and the decoder is connected to a depth head for outputting the completed elevation map and a segmentation head for outputting the predicted stepped edge map.
9. The method for training a robot observation data completion model according to claim 8, wherein The predicted stepped edge map output by the segmentation head is also used for input to the depth head.
10. The method for training a robot observation data completion model according to claim 1, wherein The two-dimensional convolutional neural network model includes: A second encoder for performing dimensionality reduction processing on the stepped terrain elevation map containing the incomplete area to extract image features. A second decoder for performing dimensionality restoration according to the image features and outputting the completed elevation map. A third decoder for performing dimensionality restoration according to the image features and outputting the predicted stepped edge map.
11. The method for training a robot observation data completion model according to claim 1, wherein The calculating the first loss by using the completed elevation map and the actual complete elevation map includes: According to: Calculate to obtain the first loss. Among them, is the first loss, is the true height value of the i-th pixel in the actual complete elevation map, is the predicted height value of the i-th pixel in the complemented elevation map, and n is the number of all pixel points in the actual complete elevation map and the complemented elevation map.
12. The method for training a robot observation data completion model according to claim 1, wherein The calculating the second loss by using the predicted stepped edge map and the actual stepped edge map includes: According to: Calculate to obtain the second loss. Among them, the BCE Loss is the second loss, is the true label corresponding to the i-th pixel in the actual step edge map, is the probability that the i-th pixel in the predicted step edge map belongs to the edge, and N is the number of all pixel points in the predicted step edge map and the actual step edge map.
13. The method for training a robot observation data completion model according to any one of claims 1 to 12, characterized in that The performing back propagation optimization on the two-dimensional convolutional neural network according to the first loss and the second loss includes: Calculate the comprehensive loss based on the first loss and the second loss; With the goal of minimizing the comprehensive loss, back-propagation optimization is performed on the model parameters in the two-dimensional convolutional neural network.
14. The method for training a robot observation data completion model according to claim 13, wherein The comprehensive loss includes a first comprehensive loss; and the calculation of the comprehensive loss based on the first loss and the second loss includes: according to: The first comprehensive loss is calculated; Among them, is the first comprehensive loss, is the first loss, BCE Loss is the second loss, and are weights.
15. The method for training a robot observation data completion model according to claim 13, wherein The comprehensive loss includes a second comprehensive loss; and the calculation of the comprehensive loss based on the first loss and the second loss includes: Generate a corresponding weight mask according to the image area whose height value exceeds the preset height threshold; The first loss and the second loss are adjusted using the weight mask, and the second comprehensive loss is calculated based on the adjusted first loss and the second loss.
16. The method for training a robot observation data completion model according to claim 15, wherein The step of adjusting the first loss and the second loss by using the weight mask, and calculating the second comprehensive loss according to the adjusted first loss and the second loss, includes: according to: Calculate and obtain the second comprehensive loss; Among them, is the second comprehensive loss, Mask(i) is the weight mask, and Mask(i) = α (0 < α < 1) means that the loss weight α is assigned to the i-th pixel point whose height value exceeds the preset height threshold. is the first loss, is the first loss adjusted by using the weight mask ; is the second loss, is the second loss adjusted by using the weight mask ; and are weights.
17. A method for completing robot observation data, characterized in that, include: Collect target stepped terrain elevation map; Inputting the target stepped terrain elevation map into a pre-trained robot observation data completion model to obtain a completed elevation map and a predicted stepped edge map; Wherein, the pre-trained robot observation data completion model is obtained according to the robot observation data completion model training method according to any one of claims 1 to 16.
18. A robot motion control method, characterized in that, include: The completed elevation map and the predicted step edge map are input into the robot motion control model pre-trained by the deep reinforcement learning algorithm to obtain the motion control parameters; Control the robot to go up or down stairs according to the motion control parameters; The completed elevation map and the predicted step edge map are Obtained according to the robot observation data completion method described in claim 17.
19. A training device for a robot observation data completion model, characterized in that, include: An image processing module, used for inputting a stepped terrain elevation map containing an incomplete area into a two-dimensional convolutional neural network to be trained, to obtain a completed elevation map and a predicted step edge map; wherein the predicted step edge map is a binary image predicted by a neural network for identifying the step edge position; A loss calculation module, used to calculate a first loss using the completed elevation map and the actual complete elevation map; and calculate a second loss using the predicted step edge map and the actual step edge map; A network optimization module is used to perform back propagation optimization on the two-dimensional convolutional neural network according to the first loss and the second loss.
20. A robot observation data completion device, characterized in that, include: An image acquisition module is used to acquire a target stepped terrain elevation map; A data completion module, used for inputting the target stepped terrain elevation map into a pre-trained robot observation data completion model to obtain a completed elevation map and a predicted stepped edge map; Wherein, the pre-trained robot observation data completion model is obtained according to the robot observation data completion model training method according to any one of claims 1 to 16.
21. A robot motion control device, characterized in that, include: A parameter determination module, used to input the completed elevation map and the predicted step edge map into a robot motion control model pre-trained by a deep reinforcement learning algorithm to obtain motion control parameters; A motion control module, configured to control the robot to climb or descend stairs according to the motion control parameters; wherein, the completed elevation map and the predicted stair edge map are obtained by the method for completing robot observation data according to claim 17.
22. An electronic device, characterized in that, Comprising: a processor; and a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 16 or 17 or 18 is implemented.
23. A robot, characterized in that, Comprising: a processor; and a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 16 or 17 or 18 is implemented.
24. The robot according to claim 23, wherein, The robot includes any one of a legged robot, a quadruped robot, a biped robot, a wheeled robot, a wheel-legged robot, a four-wheel-legged robot, a humanoid robot, a cleaning robot, a transportation robot, a mobile robot, and a robotic arm.
25. A computer-readable storage medium, characterized in that, Computer program code instructions are stored on the computer-readable storage medium, and when the computer program code instructions are called by the processor of the robot, the robot is caused to execute the method according to any one of claims 1 to 16 or 17 or 18.
Citation Information
Patent Citations
Image depth completion method and device, computer equipment and storage medium
CN117788546A
Method for training elevation map completion model
CN118552806A
Depth completion method applied to sparse map densification
CN110097589A
Model training method, figure image completion method and apparatus, and electronic device
CN112488284A