Robot observation data completion model training method and device
By training a two-dimensional convolutional neural network, the missing data in the elevation map in the robot perception and predict the ladder edges are solved, and the problem of poor data completion effect in the existing technology is achieved, efficient and accurate data completion and terrain recognition are achieved, and the reliability of robot motion control is improved.
Patent Information
- Application Number
- CN202510527006.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
In the robot perception and motion planning, it is difficult to effectively fill in the data missing in the elevation map, especially when the ladder terrain is included, which may cause falls, errors or personal safety accidents during the robot's motion control.
By inputting the step topographic elevation map containing the broken areas to the two-dimensional convolutional neural network to be trained, the completed elevation map and predicted step edge map are output, and the loss is calculated using these output results and backpropagation optimization is performed to improve the effect of the completion model.
It realizes efficient completion of elevation maps and identification of ladder terrain under CPU only, improves the accuracy of robot observation data completion and motion control, reduces computational complexity and enhances the deployability of the model.
Smart Images

Figure CN120068952A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of sensors and robotics, and relates to a method and device for training a robot observation data completion model. Background Art
[0002] In robot perception and motion planning, high-precision environmental observation data is crucial for adaptability to complex terrains. The elevation map is an important data structure for robot environmental perception and can be used for tasks such as path planning, obstacle avoidance, and stable walking. However, limited by factors such as the perspective occlusion of sensors and measurement errors, the elevation maps collected by robots usually have missing data and need to be restored by a completion algorithm to improve the integrity and accuracy of environmental modeling.
[0003] The paper "Neural Scene Representation for Locomotion on Structured Terrain" (arXiv:2206.08077v1 [cs.RO], 2022.06.16) proposed a method for calculating an elevation map based on point cloud completion. This method first fills in the surrounding point cloud data and then calculates the elevation map using the filled-in point cloud data. While maintaining the original spatial structure, this solution can complete complex three-dimensional details and improve the integrity of the elevation map. However, this method has problems such as large computational overhead and long training time, and relies on a GPU for inference. Currently, most robots still mainly use CPU computing and do not have GPU devices installed. Therefore, the applicability of this method in practical applications is limited.
[0004] In some existing technologies, such as the Chinese patent applications with publication numbers CN117788546A and CN118552806A, the point cloud data is converted into an actual depth map and directly input into a pre-trained U-Net depth estimation model to predict the estimated depth map of the current scene and then perform depth completion. Taking CN117788546A as an example, to train the depth estimation model, first, a simulation scene is constructed, and RGB images and their corresponding depth information are collected using a virtual camera to form a training set. Subsequently, pre-training is performed on the depth estimation backbone network based on the U-Net structure, and the model is optimized in combination with a dedicated loss function to make it better adapt to different object materials and lighting conditions. Among them, the calculation of the loss function is based on the matching degree between the estimated depth map and the actual depth map, and the model parameters are adjusted by minimizing the error between the two.
[0005] Compared with the point cloud level completion method in the above paper, the technical solution in the above patent directly uses a pre-trained two-dimensional convolutional neural network with a U-Net structure to complete the elevation map, which effectively reduces the amount of calculation and reduces the parameter scale of the network model, thereby being able to complete the inference calculation under the condition of only CPU, thereby improving the deployability of the solution on the robot.
[0006] Then, the inventor discovered that the technical solution in the above patent still has room for improvement in the completion effect of elevation maps containing stepped terrain and the stepped terrain recognition effect, which may lead to falls, errors and even personal safety accidents due to the inability to obtain accurate perception data during robot motion control. Summary of the invention
[0007] The present invention provides a robot observation data completion model training method and device for improving robot observation data completion and step terrain recognition effects.
[0008] Additional aspects and advantages of the disclosure will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the disclosure.
[0009] According to a first aspect of the present disclosure, a method for training a robot observation data completion model is provided, comprising: Input the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain the completed elevation map and the predicted step edge map; The first loss is calculated using the completed elevation map and the actual complete elevation map; the second loss is calculated using the predicted step edge map and the actual step edge map; Back-propagation optimization of the 2D convolutional neural network is performed based on the first loss and the second loss.
[0010] In an exemplary embodiment of the present disclosure, the method further includes: Obtaining point cloud data of stepped terrain observation; Based on the stepped terrain observation point cloud data, a stepped terrain elevation map including the incomplete area is obtained.
[0011] In an exemplary embodiment of the present disclosure, a stepped terrain elevation map including an incomplete area is obtained according to stepped terrain observation point cloud data, including: Converting the stepped terrain observation point cloud data into a first elevation map; The first elevation map is downsampled to obtain a stepped terrain elevation map including the incomplete area; the resolution of the stepped terrain elevation map including the incomplete area is lower than the resolution of the first elevation map.
[0012] In an exemplary embodiment of the present disclosure, the resolution of the first elevation map is a×b, where 0.005m ≤ a ≤ 0.02m and 0.005m ≤ b ≤ 0.02m; the resolution of the stepped terrain elevation map including the incomplete area is (n·a)×(n·b), where 4 ≤ n ≤ 10.
[0013] In an exemplary embodiment of the present disclosure, the resolution of the first elevation map is 0.01m×0.01m, and the resolution of the stepped terrain elevation map including the incomplete area is 0.05m×0.05m.
[0014] In an exemplary embodiment of the present disclosure, the method further includes: Adding noise data to the stepped terrain observation point cloud data; wherein, the noise data includes randomly added planar point clouds, and the area of the planar point clouds is positively correlated with the foot coverage area of the robot.
[0015] In an exemplary embodiment of the present disclosure, the noise data further includes: Gaussian noise, multiple outliers randomly added around a certain point, and / or randomly removed data.
[0016] In an exemplary embodiment of the present disclosure, the two-dimensional convolutional neural network model includes: A first encoder for performing dimensionality reduction on the stepped terrain elevation map including the incomplete area to extract image features; A first decoder for performing dimensionality restoration according to the image features; and the decoder is connected to a depth head for outputting the completed elevation map and a segmentation head for outputting the predicted stepped edge map.
[0017] In an exemplary embodiment of the present disclosure, the predicted stepped edge map output by the segmentation head is further used as input to the depth head.
[0018] In an exemplary embodiment of the present disclosure, the two-dimensional convolutional neural network model includes: A second encoder for performing dimensionality reduction on the stepped terrain elevation map including the incomplete area to extract image features; A second decoder for performing dimensionality restoration according to the image features and outputting the completed elevation map; A third decoder for performing dimensionality restoration according to the image features and outputting the predicted stepped edge map.
[0019] In an exemplary embodiment of the present disclosure, calculating a first loss using the completed elevation map and the actual complete elevation map includes: According to: Calculating to obtain the first loss; Wherein, is the first loss, is the true height value of the i-th pixel in the actual complete elevation map, is the predicted height value of the i-th pixel in the completed elevation map, and n is the number of all pixel points in the actual complete elevation map and the completed elevation map.
[0020] In an exemplary embodiment of the present disclosure, calculating a second loss using the predicted step edge map and the actual step edge map includes: According to: calculate to obtain the second loss; wherein, BCE Loss is the second loss, is the true label corresponding to the i-th pixel in the actual step edge map, is the probability that the i-th pixel in the predicted step edge map belongs to the edge, and N is the number of all pixel points in the predicted step edge map and the actual step edge map.
[0021] In an exemplary embodiment of the present disclosure, backpropagation optimization of the two-dimensional convolutional neural network according to the first loss and the second loss includes: Calculating a comprehensive loss according to the first loss and the second loss; Taking the minimum of the comprehensive loss as the goal, performing backpropagation optimization on the model parameters in the two-dimensional convolutional neural network.
[0022] In an exemplary embodiment of the present disclosure, the comprehensive loss includes a first comprehensive loss; calculating the comprehensive loss according to the first loss and the second loss includes: According to: calculate to obtain the first comprehensive loss; wherein, is the first comprehensive loss, is the first loss, BCE Loss is the second loss, and are weights.
[0023] In an exemplary embodiment of the present disclosure, the comprehensive loss includes a second comprehensive loss; calculating the comprehensive loss according to the first loss and the second loss includes: Generating a corresponding weight mask according to the image region where the height value exceeds the preset height threshold; Adjusting the first loss and the second loss using the weight mask, and calculating the second comprehensive loss according to the adjusted first loss and second loss.
[0024] In an exemplary embodiment of the present disclosure, the first loss and the second loss are adjusted using a weight mask, and a second comprehensive loss is calculated according to the adjusted first loss and the second loss, including: according to: The second comprehensive loss is calculated; in, is the second comprehensive loss, Mask(i) is the weight mask, and Mask(i)=α(0<α<1) means that the loss weight α is assigned to the i-th pixel whose height value exceeds the preset height threshold. For the first loss, To use weight mask Adjusted first loss, For the second loss, To use weight mask Adjusted second loss, and is the weight.
[0025] According to a second aspect of the present disclosure, a robot observation data completion method is provided, comprising: Collect target stepped terrain elevation map; Input the target step terrain elevation map into the pre-trained robot observation data completion model to obtain the completed elevation map and the predicted step edge map; Among them, the pre-trained robot observation data completion model is obtained according to the robot observation data completion model training method in the first aspect of the present disclosure.
[0026] According to a third aspect of the present disclosure, there is provided a robot motion control method, comprising: The completed elevation map and the predicted step edge map are input into the robot motion control model pre-trained by the deep reinforcement learning algorithm to obtain the motion control parameters; Control the robot to go up or down the stairs according to the motion control parameters; Among them, the completed elevation map and the predicted step edge map are obtained according to the robot observation data completion method in the second aspect of the present disclosure.
[0027] According to a fourth aspect of the present disclosure, a robot observation data completion model training device is provided, comprising: An image processing module is used to input the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain a completed elevation map and a predicted stepped edge map; A loss calculation module, used to calculate a first loss using the completed elevation map and the actual complete elevation map; and calculate a second loss using the predicted step edge map and the actual step edge map; A network optimization module for performing backpropagation optimization on a two-dimensional convolutional neural network according to a first loss and a second loss.
[0028] According to a fifth aspect of the present disclosure, there is provided a robot observation data completion device, including: An image acquisition module for acquiring a target stepped terrain elevation map; A data completion module for inputting the target stepped terrain elevation map into a pre-trained robot observation data completion model to obtain a completed elevation map and a predicted stepped edge map; Wherein, the pre-trained robot observation data completion model is obtained according to the robot observation data completion model training method in the first aspect of the present disclosure.
[0029] According to a sixth aspect of the present disclosure, there is provided a robot motion control device, including: A parameter determination module for inputting the completed elevation map and the predicted stepped edge map into a robot motion control model pre-trained by a deep reinforcement learning algorithm to obtain motion control parameters; A motion control module for controlling the robot to climb or descend stairs according to the motion control parameters; Wherein, the completed elevation map and the predicted stepped edge map are obtained according to the robot observation data completion method in the second aspect of the present disclosure.
[0030] According to a seventh aspect of the present disclosure, there is provided an electronic device, including: A processor; and A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods in the above embodiments are implemented.
[0031] According to an eighth aspect of the present disclosure, there is provided a robot, including: A processor; and A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods in the above embodiments are implemented.
[0032] In an exemplary embodiment of the present disclosure, the robot includes any one of a legged robot, a quadruped robot, a biped robot, a wheeled robot, a wheel-legged robot, a four-wheel-legged robot, a humanoid robot, a cleaning robot, a transportation robot, a mobile robot, and a robotic arm.
[0033] According to a ninth aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program code instructions are stored, and when the computer program code instructions are called by a processor of a robot, the robot is caused to execute the methods in the above embodiments.
[0034] It can be seen from the above technical solution that the present disclosure has at least one of the following advantages and positive effects: The robot observation data completion model training method disclosed in the present invention inputs the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained, outputs the completed elevation map and the predicted step edge map, and uses the completed elevation map and the actual complete elevation map to calculate the first loss, and uses the predicted step edge map and the actual step edge map to calculate the second loss, and then optimizes the two-dimensional convolutional neural network according to the first loss and the second loss to obtain the robot observation data completion model. Compared with the existing technology, on the one hand, it avoids the point cloud completion link, and directly completes the completion based on the elevation map, making the data preprocessing process simpler, reducing the overall calculation complexity, and reducing the parameter scale of the network model, so that efficient reasoning can be completed even with only a CPU, improving the deployability of the model on the robot. On the other hand, the present invention also introduces step edge information as an auxiliary supervision signal. During the training process, the elevation completion error and the step edge error are calculated for joint optimization, thereby guiding the model to learn more accurate terrain boundary features, ensuring that the completed data conforms to the large-scale terrain structure while maintaining the integrity of key local features. Therefore, both the observation data completion effect and the step terrain recognition effect will be better, which can provide better data support for robot motion control and achieve precise control. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0036] Figure 1 A system architecture diagram is shown to which the robot observation data completion model training method, the robot observation data completion method, and the robot motion control method in the embodiments of the present disclosure can be applied.
[0037] Figure 2 A flow chart of a robot observation data completion model training method in an embodiment of the present disclosure is shown.
[0038] Figure 3 A schematic diagram of a stepped terrain observation point cloud in an embodiment of the present disclosure is shown.
[0039] Figure 4 Another schematic diagram of a stepped terrain observation point cloud in an embodiment of the present disclosure is shown.
[0040] Figure 5A schematic diagram of another stepped terrain observation point cloud in an embodiment of the present disclosure is shown.
[0041] Figure 6 A schematic diagram of a process for obtaining a stepped ground elevation map including an incomplete area in an embodiment of the present disclosure is shown.
[0042] Figure 7 A schematic diagram of the structure of a two-dimensional convolutional neural network in an embodiment of the present disclosure is shown.
[0043] Figure 8 A schematic diagram of the structure of another two-dimensional convolutional neural network in an embodiment of the present disclosure is shown.
[0044] Figure 9 A schematic diagram of the structure of another two-dimensional convolutional neural network in an embodiment of the present disclosure is shown.
[0045] Figure 10 A schematic diagram of another stepped terrain observation point cloud in an embodiment of the present disclosure is shown.
[0046] Figure 11 A schematic diagram of a stepped ground elevation map in an embodiment of the present disclosure is shown.
[0047] Figure 12 A schematic diagram of an actual complete elevation map in an embodiment of the present disclosure is shown.
[0048] Figure 13 A schematic diagram of a completed elevation map in an embodiment of the present disclosure is shown.
[0049] Figure 14 A schematic diagram of an actual step edge graph in an embodiment of the present disclosure is shown.
[0050] Figure 15 A schematic diagram of a predicted step edge graph in an embodiment of the present disclosure is shown.
[0051] Figure 16 A schematic diagram of a two-dimensional convolutional neural network optimization process in an embodiment of the present disclosure is shown.
[0052] Figure 17 A schematic flow chart of a robot observation data completion method in an embodiment of the present disclosure is shown.
[0053] Figure 18 A schematic flow chart of a robot motion control method in an embodiment of the present disclosure is shown.
[0054] Figure 19 A block diagram of a robot observation data completion model training device in an embodiment of the present disclosure is shown.
[0055] Figure 20Shows a block diagram of a robot observation data completion device in an embodiment of the present disclosure.
[0056] Figure 21 Shows a block diagram of a robot motion control device in an embodiment of the present disclosure.
[0057] Figure 22 Shows a schematic diagram of a robot in an embodiment of the present disclosure.
[0058] Figure 23 Shows another schematic diagram of a robot in an embodiment of the present disclosure.
[0059] Figure 24 Shows yet another schematic diagram of a robot in an embodiment of the present disclosure.
[0060] Figure 25 Shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing an embodiment of the present disclosure.
[0061] Figure 26 Shows a schematic diagram of a computer-readable storage medium in some embodiments of the present disclosure. Detailed implementation manners
[0062] In the description of the present disclosure, the terms "first" and "second" are only used for description and do not indicate relative importance or imply the number of technical features. Therefore, the features of "first" and "second" may explicitly or implicitly include at least one of such features. The meaning of "a plurality" is at least two, unless otherwise clearly defined.
[0063] Figure 1 Shows a system architecture diagram to which the robot observation data completion model training method, the robot observation data completion method, and the robot motion control method in the embodiments of the present disclosure can be applied.
[0064] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a robot 102, a network 103, and a server 104. Among them, the terminal device 101 includes, but is not limited to, a desktop computer, a portable computer, a smart phone, a tablet computer, and the like. The terminal device 101 may serve as an interaction interface, provide a visualization function to display the running state of the robot 102, the elevation map completion result, the motion trajectory, etc., and at the same time support sending motion control instructions or adjusting motion control parameters to the robot 102.
[0065] The robot 102 has the ability to complete observation data, motion control, and neural network model training, and can autonomously perform the motion task of going up or down stairs. For example, the robot 102 receives sensor input and performs elevation map completion, stair edge detection, and motion parameter reasoning based on the local CPU. The embodiment of the disclosure avoids the point cloud completion link, reduces the computational complexity, and enables the CPU to still complete efficient reasoning, thereby ensuring that the robot 102 can completely rely on its own intelligent control system to adjust the motion state during the task execution process.
[0066] The server 104 can be used to train the robot observation data completion model and update the robot motion control model, and regularly send the optimized model parameters to the robot 102 to improve its motion control performance. At the same time, the server 104 supports model lightweight processing, so that the optimized model can adapt to the computing power of the robot 102, thereby ensuring the effective deployment and operation of the model on the robot 102.
[0067] The network 103 is used as a medium to provide a communication link between the terminal device 101, the robot 102 and the server 104. The network 103 may include various connection types, such as wired or wireless communication links or optical fiber cables. By connecting various devices, the network 103 ensures that the robot 102 can obtain the latest model update from the server 104 when necessary, and ensures that the robot 102 can perform independent reasoning locally, reduce dependence on the server and reduce data transmission requirements, and improve the real-time and deployability of the system. The robot 102 in the system architecture 100 can still operate efficiently in an environment with limited computing resources, while reducing the overall computing and communication costs of the system.
[0068] It should be understood that Figure 1 The number and type of terminal devices, robots, networks and servers in the embodiment are only for illustration purposes. Any number and type of terminal devices, robots, networks and servers may be provided as required.
[0069] The present disclosure provides a method for training a robot observation data completion model. Figure 2 As shown, the method may include the following steps S201 to S203: Step S201, inputting the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain a completed elevation map and a predicted stepped edge map; Step S202, calculating a first loss using the completed elevation map and the actual complete elevation map; calculating a second loss using the predicted step edge map and the actual step edge map; Step S203, performing back propagation optimization on the two-dimensional convolutional neural network according to the first loss and the second loss.
[0070] Compared with the prior art, the robot observation data completion model training method provided by the example implementation of the present disclosure avoids the point cloud completion link and directly completes the completion based on the elevation map, making the data preprocessing process simpler, reducing the overall computational complexity, and reducing the parameter scale of the network model, so that efficient reasoning can be completed even with only a CPU, improving the deployability of the model on the robot. On the other hand, the present disclosure also introduces step edge information as an auxiliary supervision signal, and during the training process, the elevation completion error and the step edge error are calculated for joint optimization, thereby guiding the model to learn more accurate terrain boundary features, ensuring that the completed data not only conforms to the large-scale terrain structure, but also maintains the integrity of key local features; therefore, both the observation data completion effect and the step terrain recognition effect will be better, thereby providing better data support for robot motion control and achieving precise control.
[0071] Next, the robot observation data completion model training method in this example embodiment will be described in detail.
[0072] In step S201, the stepped terrain elevation map including the incomplete area is input into the two-dimensional convolutional neural network to be trained to obtain a completed elevation map and a predicted stepped edge map.
[0073] In the example implementation of the present disclosure, the step terrain elevation map refers to a two-dimensional array that represents the height distribution of the step terrain in a grid manner, wherein the value of each grid cell represents the ground height of the area. The step terrain is usually composed of multiple steps, each of which has a large area in the horizontal direction and a more obvious height change in the vertical direction. Therefore, the step terrain elevation map will show an obvious hierarchical structure, that is, the grid cells on the same step have similar height values, while the grid cells at the edge of the step show a sudden height difference. The step terrain elevation map can help the robot perceive the terrain undulations, enable the robot to recognize the hierarchical structure of the steps, and thus make reasonable decisions in path planning and motion control.
[0074] A stepped terrain elevation map with incomplete areas refers to an elevation map where elevation data of some areas is missing due to sensor acquisition limitations, robot view occlusion, or environmental obstacles. When actually acquiring elevation data, the height values of some grid cells cannot be measured correctly due to factors such as sensor view, low-reflectivity materials, ambient light changes, and obstacle occlusion, thus forming incomplete areas. Especially in stepped terrain, the missing areas are more obvious due to sudden height changes, sharp transitions at the edge of the steps, and perspective issues. For stepped terrain elevation maps with incomplete areas, some grid cells do not have valid height values, which can be represented by default values such as NaN or zero.
[0075] When the stepped terrain elevation map contains incomplete regions, it will affect robot navigation, path planning, and terrain analysis. Therefore, in the exemplary embodiments of the present disclosure, the stepped terrain elevation map can be completed through a neural network model, so that the incomplete regions can be reasonably inferred based on the known elevation data around, thereby generating a complete elevation map, and synchronously identifying and retaining the stepped edge information to ensure the robot's accurate perception and adaptation ability to the terrain.
[0076] Correspondingly, the completed elevation map refers to the complete elevation map generated by using a neural network model to complete the missing height data of the stepped terrain elevation map. The predicted stepped edge map refers to a binary image predicted by a neural network for identifying the position of the staircase edge, where the value of each pixel point in the stepped edge map represents the probability that the point belongs to the stepped edge. For example, a binary representation of 0 and 1 can be used, where 1 represents that the point is the edge region of the stepped terrain, and 0 represents the non-edge region.
[0077] It should be noted that the stepped edge information is crucial for robot navigation. It can help the robot identify the walkable area, detect obstacles, and adjust its gait to adapt to complex terrain changes. Moreover, the edge prediction function of the neural network helps to enhance the clarity of the step structure and improve the robot's terrain recognition accuracy.
[0078] The two-dimensional convolutional neural network to be trained is a deep learning model for processing image data. It can extract local features through convolution operations and learn higher-level representations layer by layer. Exemplarily, the two-dimensional convolutional neural network can adopt the U-Net architecture, which has a small computational amount and is suitable for CPU computing. The U-Net architecture consists of an encoder and a decoder. Among them, the encoder is composed of multiple convolutional layers and pooling layers, which are responsible for feature extraction of the input stepped terrain elevation map containing incomplete regions and reducing the spatial dimension, so that the model can learn more advanced terrain features. The decoder is composed of multiple transposed convolutional layers, which are used to gradually restore the spatial dimension of the elevation map and generate a completed elevation map with the same size as the input. Importantly, the decoder in the exemplary embodiments of the present disclosure can also generate a stepped edge map.
[0079] For example, a completed elevation map is output by one decoder, and at the same time, this decoder also additionally includes an independent output layer for outputting the stepped edge map. Another example is that the U-Net architecture includes an encoder and two decoders, and the completed elevation map and the predicted stepped edge map are output in parallel by the two decoders. The specific structure of the U-Net architecture in the embodiments of the present disclosure is not limited, as long as it can output the completed elevation map and the predicted stepped edge map.
[0080] Of course, according to actual needs, the attention mechanism can also be introduced on the U-Net architecture. The two-dimensional convolutional neural network can also adopt ResNet (residual network) + FPN (Feature Pyramid Network), DeepLabV3+ (image semantic segmentation) network, etc., which is not limited in this disclosure.
[0081] Before inputting the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained, it is necessary to first obtain the stepped terrain elevation map containing the incomplete area. Exemplarily, the stepped terrain elevation map containing the incomplete area can be obtained by obtaining the stepped terrain observation point cloud data and obtaining the stepped terrain elevation map containing the incomplete area according to the stepped terrain observation point cloud data.
[0082] Specifically, the robot's laser radar, depth camera or structured light sensor and other devices are used to scan the environment around the robot to obtain three-dimensional point cloud data containing stepped terrain, that is, step terrain observation point cloud data. Among them, three-dimensional point cloud data is a set of discrete coordinate points, each point contains three spatial coordinate information of X, Y, and Z, and can also contain attributes such as reflectivity, color or timestamp, which is used to describe the three-dimensional shape of the terrain.
[0083] refer to Figures 3 - 5 As shown in the figure, three different step terrain observation point cloud diagrams are given. Figure 3 The observed point cloud of the stepped terrain shown is sparse, and point data is missing in some areas; Figure 4 Compared with the observed point cloud of stepped terrain shown in Figure 3 The step terrain observation point cloud shown is denser and has color changes to distinguish the height differences between point data; Figure 3 , Figure 4 The observed point cloud of stepped terrain is shown. Figure 5 The observation point cloud of the stepped terrain shown is the densest, and the color gradient changes are more obvious, that is, the height information is clearer and more complete.
[0084] After obtaining the point cloud data of the stepped terrain observation, preprocessing operations such as noise filtering, ground segmentation and coordinate transformation can be performed. Among them, noise filtering refers to the use of statistical methods to remove abnormal points introduced by sensor errors or environmental interference to make the point cloud smoother and more reliable. Statistical methods include one or more of statistical filtering, mean filtering, Gaussian filtering, K-nearest neighbor average filtering, etc. For example, statistical filtering can be used to remove outliers first, and then combined with Gaussian filtering to smooth the data to ensure that the generated elevation map is more stable and accurate. Ground segmentation refers to removing non-topographic points (such as walls, railings, pedestrians) and retaining the real stepped terrain data. Coordinate transformation refers to converting the point cloud from the sensor coordinate system to the world coordinate system or the robot's own coordinate system so that the data can be calculated under a unified reference system.
[0085] In an exemplary embodiment, with reference to Figure 6 as shown, the process of converting the preprocessed stepped terrain observation point cloud data into a stepped terrain elevation map including incomplete regions includes step S601 and step S602: Step S601: Convert the stepped terrain observation point cloud data into a first elevation map.
[0086] Among them, the first elevation map is a regularly gridded two-dimensional array, and its accuracy is determined by the set grid size. A smaller grid size means a higher spatial resolution, making the terrain details clearer. Therefore, in step S601, a fixed grid resolution can be set first, such as 0.01m×0.01m, and the stepped terrain observation point cloud data is projected onto the two-dimensional grid. Because the high-resolution grid is finer, even if the point cloud points are unevenly distributed, each grid will contain more accurate points, and the value of each grid cell represents the ground height at that position. Since the point cloud data contains multiple irregularly distributed points, the method of taking the highest point can be used during the conversion process, that is, the maximum value of the Z coordinate is selected as the elevation value of the grid within each grid range. Thus, while converting the stepped terrain observation point cloud data into the first elevation map, key terrain features such as the edges of the steps can be ensured to be retained.
[0087] Step S602: Perform downsampling on the first elevation map to obtain a stepped terrain elevation map including incomplete regions; the resolution of the stepped terrain elevation map including incomplete regions is lower than that of the first elevation map.
[0088] After obtaining the first elevation map, in order to reduce the computational complexity and simulate the resolution limitation of real sensors, it is necessary to perform downsampling on the first elevation map, that is, reduce the spatial resolution of the data to obtain a stepped terrain elevation map including incomplete regions.
[0089] In the exemplary embodiment of the present disclosure, the resolution of the first elevation map is a×b, where 0.005m≤a≤0.02m and 0.005m≤b≤0.02m. Such a resolution setting can fully retain the detailed information on the terrain surface. The resolution of the stepped terrain elevation map including incomplete regions is (n·a)×(n·b), where 4≤n≤10, ensuring that the resolution of the stepped terrain elevation map is lower than that of the first elevation map, meeting the low-resolution simulation requirements. Moreover, since the value of n ranges from 4 to 10, the accuracy difference between the incomplete map and the complete map can be flexibly controlled.
[0090] Exemplarily, the resolution of the first elevation map is 0.01m × 0.01m. The process of downsampling the first elevation map refers to re - dividing the grid according to a larger grid size, such as 0.05m × 0.05m, and calculating the corresponding elevation value by taking the highest point within each new grid cell. This can retain higher terrain change features, such as the edge information of steps, while removing some details. During the downsampling process, since a larger grid covers multiple points with different elevations, some areas are missing due to too few data points or no points falling into them at all, thus forming a stepped terrain elevation map containing incomplete areas.
[0091] For example, the height distribution of the point cloud in the first elevation map is as follows: 0.12 0.15 0.18 0.20 0.11 0.14 0.17 0.19 0.10 0.13 0.16 0.18 0.10 0.12 0.15 0.17 Among them, the resolution of the first elevation map is 0.01m × 0.01m, that is, each grid cell represents an area of 0.01 meters, and the size of the first elevation map is 4×4.
[0092] By merging the 4×4 first elevation map into a 2×2 one and performing downsampling by taking the highest point, the height distribution of the point cloud in the stepped terrain elevation map containing incomplete areas is as follows: 0.15 0.20 0.13 0.18 Among them, the resolution of the stepped terrain elevation map with incomplete areas is 0.05m × 0.05m, that is, each new grid cell represents an area of 0.05 meters, and the size of this elevation map is 2×2. It can be seen that the highest 0.20 at the edge of the stairs is retained.
[0093] Since the resolution of the stepped terrain elevation map containing incomplete areas is lower than that of the first elevation map, compared with the stepped terrain observation point cloud data, the spatial information of the finally obtained stepped terrain elevation map containing incomplete areas is sparser, which can effectively reduce the amount of calculation and improve the efficiency of subsequent processing. At the same time, this elevation map still retains the basic contour of the stepped terrain, providing basic data for subsequent elevation completion and terrain analysis.
[0094] In this exemplary embodiment, on the one hand, by converting complex three-dimensional point clouds into regular grid-like two-dimensional data, subsequent processing becomes more efficient and computable. At the same time, converting high-dimensional point cloud data into a two-dimensional representation reduces the computational amount, enabling the robot observation data completion model to run under limited computing power conditions (such as only using a CPU). On the other hand, by constructing an elevation map by selecting the highest points, key terrain information is retained, ensuring the integrity of the step edge and terrain undulation information, thereby improving the robot's recognition ability for complex terrains. On yet another hand, by analyzing the data gaps in the elevation map, it is possible to effectively determine which areas need to be completed and provide a basis for subsequent neural network processing, thus enhancing the robot's adaptability under incomplete terrain data conditions.
[0095] In addition, during the actual deployment of the robot model, it is found that the planes output by the model often fluctuate. Among them, the plane output by the model refers to the flat area in the elevation map generated by the robot observation data completion model, and this plane corresponds to the flat area of the ground, platform, or steps where the robot walks. For example, when the robot's two feet are within the camera's field of view, the point cloud of the foot surface is a plane. However, due to the noise in the related technology not considering plane noise, the model identifies the position that was originally the foot surface as a step.
[0096] Therefore, in this exemplary embodiment of the present disclosure, when training the robot observation data completion model, noise data can be added to the point cloud data of the step terrain observation, such as randomly adding plane point cloud noise. By artificially increasing the point cloud data of a part of the plane area in the simulation environment, it helps the model better learn how to maintain a stable terrain structure and improve the model's adaptability to the real environment.
[0097] It should be noted that the way to add plane point cloud noise is to randomly generate a certain number of plane point clouds near the robot's walking path according to the foot coverage area of the robot. Among them, the plane point cloud refers to a set of point cloud data with the same height and dense distribution. The Z coordinate values of these point clouds are the same, presenting a local flat area, thereby simulating the influence of the local flat area of the robot's sole or the ground.
[0098] Furthermore, the area size of the plane point cloud is determined by the foot coverage area of the robot. For example, the area of the plane point cloud is positively correlated with the foot coverage area of the robot. The larger the foot coverage area, the larger the area of the added plane point cloud, so as to more realistically simulate the influence of the robot's standing or moving on terrain observation. For example, if the contact area of the robot's foot on the ground is a circular or rectangular area, the distribution range of the plane point cloud can adopt the same shape, and the point cloud data is randomly distributed within this range according to a certain probability to simulate the terrain observation error under different gaits or movement conditions.
[0099] It should be noted that the noise data added to the stepped terrain observation point cloud data also includes Gaussian noise, multiple outliers randomly added around a certain point, and / or randomly removed data, so as to further simulate the errors in the real robot observation environment, making the robot observation data completion model more robust to uncertain factors, improving the processing ability of uncertain factors in the real environment, and thus more accurately completing the stepped terrain data.
[0100] Among them, Gaussian noise is to add random perturbations conforming to the normal distribution to the three-dimensional coordinates of the stepped terrain observation point cloud data, so that the coordinate values of each point cloud data are no longer fixed, but fluctuate within a certain range around the true measurement value. Gaussian noise can be used to simulate sensor measurement errors. After adding Gaussian noise, the height values of the data points fluctuate slightly, and then the model can learn how to remove the random errors in the observation data during training, improve the error tolerance of the model, and improve the adaptability to the real environment.
[0101] Randomly adding multiple outliers around a certain point means introducing local abnormal data points to enhance the model's ability to identify and correct outliers. Specifically, multiple offset points around a certain center point are randomly generated at certain positions in the point cloud data. The distribution of these offset points can be random or follow a certain pattern, such as uniform distribution or Gaussian distribution, making the point cloud data in some areas sparser or more concentrated. The formed local outliers can be used to simulate the measurement errors generated by the sensor on high-reflectivity or low-reflectivity surfaces. For example, key points in the point cloud data such as the stepped edge or ground points are selected, and then multiple virtual points with offset amounts less than the set range are randomly generated around them to form a local abnormal area, so that the model can learn how to ignore these invalid points to improve the stability of data completion.
[0102] Randomly removing data means artificially deleting some area of data points in the point cloud data to simulate the data loss situation caused by perspective occlusion, environmental interference, or sensor blind spots during the robot's observation process, so as to enhance the model's recovery ability in the face of large-area missing areas and make it learn how to infer the reasonable height of the missing area using the surrounding known data during training. There are various ways of random removal. For example, it can be cropped based on regular shapes, such as randomly selecting a rectangular or circular area in the elevation map and deleting the point cloud data in it, or randomly deleting some points in the point cloud data based on probability methods, making the point cloud density uneven. Especially in the processing of stepped terrain point cloud data, randomly removing the observation points of some steps can simulate the situation where the robot cannot observe the stepped edge at certain angles, so that the model can learn to more accurately complete the missing data and improve its prediction ability for complex terrains.
[0103] It should be noted that these three methods of adding noise data are respectively for different types of observation errors and data missing situations, so different noise data can be added according to actual needs, and this disclosure does not limit this. Among them, Gaussian noise can enhance the robustness of the model to sensor measurement errors, randomly adding multiple external points around a certain point can improve the model's ability to identify and process outliers, and randomly removing data can enhance the model's ability to recover from data missing.
[0104] In the example implementation of the present disclosure, by adding noise data to the step terrain observation point cloud data, the diversity of training data can be effectively enhanced, so that the model can more stably predict and complete the elevation map when encountering terrain interference or misdetection, and improve the adaptability to complex terrain, especially in the step edge detection and completion tasks, it can reduce the misjudgment caused by the influence of the robot's feet and improve the generalization ability of the model to the actual observation data. Moreover, by introducing noise data, the model can be made close to the complexity of the real environment during the training process, so that it can show stronger adaptability in actual deployment and improve the recognition and completion accuracy of step terrain.
[0105] refer to Figure 7 As shown, a schematic diagram of the structure of a two-dimensional convolutional neural network in an embodiment of the present disclosure is shown. Figure 7 It can be seen that the two-dimensional convolutional neural network includes a first encoder 701 and a first decoder 702. Among them, the first encoder 701 is used to perform dimensionality reduction processing on the stepped terrain elevation map containing incomplete areas to extract image features; the first decoder 702 is used to restore dimensions according to image features, and the decoder is connected to a depth head 7021 for outputting a completed elevation map and a segmentation head 7022 for outputting a predicted stepped edge map. Among them, the completed elevation map is used to fill in the incomplete areas in the original input to make the generated terrain data more complete and accurate. The predicted step edge map provides additional edge information to assist in optimizing the elevation completion effect and enhance the model's recognition ability of stepped terrain. In addition, features are transferred between the first encoder 701 and the first decoder 702 through jump connections, such as Figure 7 As shown by the dotted arrow in , the feature map output by the first encoder 701 will be passed to the first decoder 702 by the jump connection module for splicing.
[0106] Specifically, the depth head 7021 is mainly responsible for the elevation completion task. By receiving the feature information extracted by the first decoder 702, it outputs a completed elevation map with the same size as the original elevation map. For example, the depth head 7021 uses convolutional layers and regression activation functions to predict the height value of each grid, ensuring that the missing areas are reasonably filled, making the completed terrain data more complete and smooth. Among them, the regression activation function can be the ReLU (Rectified Linear Unit) function, linear activation function, etc.
[0107] The segmentation head 7022 is mainly responsible for the step edge detection task. By processing the features output by the first decoder 702, it generates a binary step edge map. For example, the segmentation head 7022 uses convolutional layers and the Sigmoid activation function or Softmax activation function to calculate the probability that each pixel belongs to the edge.
[0108] It should be noted that the depth head 7021 and the segmentation head 7022 can be used simultaneously in both the training stage and the inference stage of the network. Correspondingly, the model will output the completed elevation map and the predicted step edge map during inference. Of course, according to the actual model accuracy requirements, it is also possible to choose to disable the segmentation head 7022 during the inference stage and only use it during training. At this time, the edge segmentation task is only an auxiliary task during the training stage, used to help the model learn better depth features, but is removed during inference. The model only generates the completed elevation map during inference, simplifying the inference calculation amount.
[0109] Reference Figure 8 As shown, a schematic structural diagram of another two-dimensional convolutional neural network in an embodiment of the present disclosure is shown. Different from Figure 7 the network structure shown, Figure 8 in the network structure shown, features are transmitted between the depth head 7021 and the segmentation head 7022 through skip connections. Therefore, the predicted step edge map output by the segmentation head 7022 is also used as input to the depth head 7021. By using the edge information as auxiliary information for the completed elevation map, providing constraints during the completion process, ensuring that the completion result of the elevation map not only conforms to the overall trend but also maintains clear terrain changes in key edge areas.
[0110] As Figure 7 , Figure 8 shown, the two-dimensional convolutional neural network structure can effectively enhance the accuracy of elevation map completion. Especially when dealing with stepped terrains, through the guidance of edge information, the model can more accurately restore the height changes of the steps, prevent edge blurring or distortion of the step shape, and improve the navigation and perception capabilities of the robot in complex terrain environments.
[0111] Reference Figure 9FIG. 1 is a schematic diagram showing the structure of another two-dimensional convolutional neural network in an embodiment of the present disclosure. Figure 9 It can be seen that the two-dimensional convolutional neural network includes a second encoder 901, a second decoder 902 and a third decoder 903, wherein the second encoder 901 is used to perform dimensionality reduction processing on the stepped terrain elevation map containing incomplete areas to extract image features; the second decoder 902 is used to perform dimensionality restoration according to image features and output a completed elevation map; the third decoder 903 is used to perform dimensionality restoration according to image features and output a predicted stepped edge map.
[0112] Among them, the second encoder 901 can be composed of multiple convolutional layers and pooling layers, gradually reducing the spatial dimension and enhancing the feature expression ability, so that the model can learn the overall structure and local details of the terrain. The image features extracted by the second encoder 901 are then passed to the second decoder 902 and the third decoder 903 for decoding respectively. The second decoder 902 restores the original spatial resolution according to the image features extracted by the second encoder 901, and generates a completed elevation map. The second decoder 902 can use a deconvolution layer or an upsampling + convolution layer to gradually expand the feature map size so that the output image matches the input elevation map size. Finally, the output of the second decoder 902 is a completed elevation map of the same size as the original elevation map, which fills the height value of the incomplete area in the input data, making the terrain data more complete and smooth. The third decoder 903 learns the morphology of the step edge from the image features extracted by the second encoder 901, and generates a probability distribution map through a Sigmoid activation function or a Softmax classifier, and the output edge map is a binary image. In this way, the model can accurately identify the step boundary and improve the accuracy of elevation completion.
[0113] Figure 9 The two-dimensional convolutional neural network model shown in adopts a dual decoder design, that is, the second decoder 902 and the third decoder 903 independently decode the elevation map and the edge map, which can reduce the mutual interference between the two tasks and make the elevation completion and edge prediction more accurate. At the same time, this structure can optimize the completion and segmentation tasks at different levels of feature expression, and improve the model's adaptability to complex terrain.
[0114] In step S202, a first loss is calculated using the completed elevation map and the actual complete elevation map; and a second loss is calculated using the predicted step edge map and the actual step edge map.
[0115] In the exemplary embodiments of the present disclosure, a two-dimensional convolutional neural network is optimized through a loss function, enabling the model to accurately complete the elevation map and precisely predict the step edges, thereby enhancing the robot's perception ability of complex terrains. Among them, the first loss refers to the elevation completion loss. The completed elevation map is the result generated by the network, while the actual complete elevation map is the corresponding real terrain data. The purpose of calculating the first loss is to make the completed result output by the model as close as possible to the real elevation. For example, the first loss can adopt the L1 loss (Mean Absolute Error) or the L2 loss (Mean Squared Error). The L1 loss can reduce the influence of outliers in the completed result, while the L2 loss can penalize large errors. Of course, the L1 loss and the L2 loss can be combined for use to improve the completion accuracy, and the present disclosure does not make any limitations in this regard. The second loss refers to the step edge prediction loss. The predicted step edge map is a binary image output by the network, while the actual step edge map is the edge label obtained through manual annotation or calculated based on the real elevation map. The purpose of calculating the second loss is to ensure that the model can correctly identify the step edges and avoid misjudgment or edge blurring. For example, the binary cross-entropy loss function can be used to calculate the second loss, or the Dice (Dice similarity) loss function, the IoU (Intersection over Union) loss function, etc. can be used to calculate the second loss. The present disclosure does not make any limitations on the types of loss functions used for calculating the first loss and the second loss.
[0116] Reference Figure 10 As shown, another schematic diagram of the step terrain observation point cloud is shown.
[0117] Reference Figure 11 As shown, a schematic diagram of a step terrain elevation map is shown. Figure 11 The step terrain elevation map in Figure 10 is processed based on the step terrain observation point cloud shown in Figure 11 As can be seen, this step terrain elevation map contains incomplete regions.
[0118] Reference Figure 12 As shown, a schematic diagram of an actual complete elevation map is shown. Figure 12 The actual complete elevation map in Figure 11 is the complete elevation map corresponding to the step terrain elevation map shown in
[0119] Reference Figure 13 As shown, a schematic diagram of a completed elevation map is shown. Figure 13 The completed elevation map in Figure 11 takes the step terrain elevation map shown in
[0120] Reference Figure 14 As shown, a schematic diagram of an actual stepped edge map is shown.
[0121] Reference Figure 15 As shown, a schematic diagram of a predicted stepped edge map is shown. Figure 15 The predicted stepped edge map in Figure 11 is also an image output by a two-dimensional convolutional neural network in this disclosure embodiment with the stepped terrain elevation map shown as the input.
[0122] Exemplarily, the first loss can be calculated by using the complemented elevation map and the actual complete elevation map according to: (1) The first loss can be calculated; where is the first loss, is the true height value of the i-th pixel in the actual complete elevation map, is the predicted height value of the i-th pixel in the complemented elevation map, and n is the number of all pixel points in the actual complete elevation map and the complemented elevation map.
[0123] The second loss can be calculated by using the predicted stepped edge map and the actual stepped edge map according to: (2) The second loss can be calculated; where BCE Loss is the second loss, is the true label corresponding to the i-th pixel in the actual stepped edge map, is the probability that the i-th pixel in the predicted stepped edge map belongs to the edge, and N is the number of all pixel points in the predicted stepped edge map and the actual stepped edge map.
[0124] In this step, by calculating the first loss, the elevation complementing ability of the model in the missing area can be optimized, so that the complemented elevation map is as close as possible to the actual complete elevation map. Since there are missing areas in the original observed data, the optimization of this loss can guide the model to learn the global and local change laws of elevation data, thereby improving the accuracy of the complemented result, reducing the height mutation or unreasonable terrain change caused by data missing, improving the integrity of terrain data, and enabling the robot to perform path planning and terrain adaptation based on more accurate elevation information. By calculating the second loss, the edge detection ability of the model can be enhanced, so that the predicted stepped edge map can more accurately identify the true position of the steps. The stepped edge is an important structural feature that the robot must identify during walking and planning. Precise edge information helps to improve the resolution of the step height, avoid misjudging the starting point or boundary of the steps during the complementing process of the model, and thus reduce the error of terrain perception.
[0125] In step S203, the two-dimensional convolutional neural network is optimized by backpropagation according to the first loss and the second loss.
[0126] Among them, the first loss and the second loss are combined, and the gradient of the total loss with respect to the model parameters is calculated through the backpropagation algorithm. The optimizer updates the network weights according to the gradient, so that the model can optimize the elevation completion accuracy and the edge structure prediction ability simultaneously in subsequent iterations, and finally enables the two-dimensional convolutional neural network to collaboratively repair the terrain values and capture the stepped geometric features.
[0127] Reference Figure 16 As shown, the process of optimizing the two-dimensional convolutional neural network according to the first loss and the second loss may include step S1601 and step S1602: In step S1601, the combined loss is calculated according to the first loss and the second loss.
[0128] In an exemplary embodiment, the combined loss includes a first combined loss, which refers to the overall loss obtained by weighted summation of the first loss and the second loss, and is used to optimize the overall performance of the model, enabling it to accurately complete the missing elevation data and accurately identify the stepped edge information.
[0129] At this time, the first combined loss is calculated according to the first loss and the second loss, and there is: (3) Among them, is the first combined loss, is the first loss, BCE Loss is the second loss, and are weights, which are used to balance the influence of the two losses and ensure that the model does not overly favor a certain task. For example, if the model has a higher requirement for the accuracy of elevation completion, can be appropriately increased so that the elevation completion loss accounts for a larger proportion in the optimization process. If the accuracy requirement for edge recognition is higher, can be increased to enhance the learning of edge information.
[0130] It can be seen from formula (3) that the optimization objective of the first combined loss is to minimize the value, making the elevation completion more accurate, while ensuring the integrity of the edge information, thereby improving the overall performance of the model on complex terrains and enhancing the adaptability of the robot to stepped terrains.
[0131] In another exemplary embodiment, the combined loss includes a second combined loss, which refers to introducing a weight mask when calculating the combined loss to adjust the first loss and the second loss for specific height regions, thereby optimizing the learning effect of the model in different terrain regions.
[0132] For example, in actual training, it is found that there are biases in the height prediction of the model at cliff or wall positions. For example, the height of these areas is estimated to be similar to the surrounding horizontal plane, resulting in a situation where the height of the stair plane in the prediction result is erroneously extended to the cliff area. Therefore, a weight mask can be added at the cliff or wall position, and a relatively small (but not 0) weight is set for the loss at these positions through the weight mask.
[0133] When calculating the comprehensive loss based on the first loss and the second loss at this time, a corresponding weight mask can be generated according to the image area where the height value exceeds the preset height threshold. When the height of a certain area exceeds the set threshold (such as cliff, wall or high step area), the loss calculation weight of this area can be adjusted so that it receives stronger attention or less influence in the model optimization process. For example, in areas with higher heights, due to larger point cloud data acquisition errors, the model is more likely to produce false filling phenomena. Therefore, the weight of this area can be reduced to reduce the model's dependence on unreliable areas and prevent terrain distortion caused by false filling. For lower steps or key passage areas, the weight of this area can be increased to allow the model to focus on optimizing the elevation completion and edge prediction accuracy of these areas, so as to improve the adaptability of the robot in the actual terrain.
[0134] Furthermore, the first loss and the second loss are adjusted using the weight mask, and a second comprehensive loss is calculated based on the adjusted first loss and second loss. For example, the second comprehensive loss can be calculated according to formula (4): (4) where is the second comprehensive loss, Mask(i) is the weight mask, Mask(i) = α (0 < α < 1) means that the loss weight α is assigned to the i-th pixel point whose height value exceeds the preset height threshold, is the first loss, is the first loss adjusted using the weight mask , is the second loss, is the second loss adjusted using the weight mask , and are weights.
[0135] In this exemplary embodiment, after introducing the weight mask, the calculations of the first loss and the second loss will be dynamically adjusted according to the weight mask, that is, different weights are assigned to different regions during loss calculation, making the final loss more in line with the actual terrain characteristics. The adjusted first loss and second loss are used to calculate the second comprehensive loss, making the model optimization process more adaptive. It not only focuses on the overall elevation completion and edge detection accuracy, but also can perform more reasonable optimization for specific height regions to reduce the miscompletion of abnormal regions and improve the recognition accuracy of key terrains (such as low steps, ramps, and stair edges), thereby enhancing the robot's perception ability and walking stability in complex environments.
[0136] In step S1602, with the goal of minimizing the comprehensive loss, backpropagation optimization is performed on the model parameters in the two-dimensional convolutional neural network.
[0137] In the backpropagation optimization stage, first, the gradient of the comprehensive loss function is calculated, that is, the partial derivatives of it with respect to all parameters in the network are obtained. The gradient calculation follows the chain rule and propagates backward from the loss function to the parameters of each layer, including the weights and bias terms of the convolutional layer. The calculated gradient is used to guide the parameter update. The update method usually adopts an adaptive optimization algorithm, such as Adam (Adaptive Moment Estimation) or SGD (Stochastic Gradient Descent). Among them, the Adam algorithm can automatically adjust the learning rate to improve the training stability, while the SGD algorithm helps to avoid local optimum problems. This disclosure does not make any limitations in this regard. By gradually adjusting the network weights through model parameter updates, the model can more effectively minimize the comprehensive loss, improve the accuracy of the completed elevation map, and optimize the recognition ability of stair edges.
[0138] During the training process, the backpropagation optimization process will continue for multiple iteration rounds until the comprehensive loss converges, that is, the error between the model's prediction result and the real data tends to be stable, indicating that the network has learned the optimal elevation completion and edge detection strategy, or when the preset number of iterations is reached, the training of all model parameters is completed. Finally, this optimization method enables the network to more accurately restore the elevation information of the missing region, while ensuring clear stair edges, reducing false extensions or blurring phenomena, and improving the robot's navigation and walking ability in complex terrain environments.
[0139] In this step, the joint optimization of the first loss and the second loss can achieve a complementary effect, that is, while optimizing the completed elevation, the constraint of edge information is introduced, so that the elevation completion not only conforms to the overall trend, but also can accurately retain the key step structure, making the boundary of the step clearer, and avoiding edge blur or distortion caused by elevation completion. Especially on complex stepped terrain, this joint optimization can improve the stability of elevation completion and make the model more suitable for actual robot navigation tasks. Moreover, it also improves the robustness and generalization ability of the model, so that the model can adapt to stepped terrain in different environments, reduce dependence on single features, improve the accuracy and reliability of overall terrain perception, and further enhance the robot's motion planning and navigation capabilities on complex terrain.
[0140] The exemplary embodiment of the present disclosure also provides a robot observation data completion method, referring to Figure 17 As shown, the method may include the following steps S1701 and S1702: Step S1701, collecting a target stepped terrain elevation map; The target stepped terrain elevation map is obtained by scanning the stepped terrain in the target environment in real time through sensing devices such as laser radar, depth camera, structured light sensor or stereo vision system carried by the robot, obtaining three-dimensional point cloud data of the area, and converting the three-dimensional point cloud data into a stepped terrain elevation map. The target stepped terrain elevation map can be a stepped terrain elevation map including an incomplete area or a complete stepped terrain elevation map, which is not limited in the present disclosure.
[0141] Step S1702, inputting the target stepped terrain elevation map into the pre-trained robot observation data completion model to obtain a completed elevation map and a predicted stepped edge map.
[0142] Among them, the pre-trained robot observation data completion model is trained according to the robot observation data completion model training method described in detail in steps S201 to S203 of another embodiment of the present disclosure, which will not be repeated here. The pre-trained robot observation data completion model can process the input elevation map and repair the missing areas caused by sensor occlusion, data loss or measurement errors. The model first extracts the features of the input elevation map through the encoder, and gradually restores the terrain information through the decoder, and finally outputs the completed elevation map and the predicted step edge map. The reasoning process of the robot observation data completion model can be performed on the CPU to meet the real-time computing needs of the robot.
[0143] Implementing the robot observation data completion method provided by the exemplary embodiments of the present disclosure can directly perform completion based on the elevation map, making the data preprocessing process simpler, reducing the overall computational complexity, and reducing the parameter scale of the network model. Thus, efficient inference can be completed even under the condition of only having a CPU, improving the deployability of the model on the robot. Additionally, while performing elevation completion, the model can also predict the step edge map to ensure that the robot can accurately perceive and identify the staircase structure, preventing misjudgment of the starting point of the step or blurred boundaries. The extraction of this edge information can not only optimize the elevation completion result, making the height change of the step clearer, but also assist the robot in more precise gait control, avoiding safety issues such as unstable walking or falling due to incorrect perception of the step position.
[0144] The exemplary embodiments of the present disclosure also provide a robot motion control method. Referring to Figure 8 as shown, this method may include the following steps S1801 to step S1802: Step S1801, input the completed elevation map and the predicted step edge map into a robot motion control model pre-trained by a deep reinforcement learning algorithm to obtain motion control parameters.
[0145] Among them, the completed elevation map is used as the terrain input of the robot's traveling path, enabling the robot to accurately perceive the height, width, and slope of the staircase. At the same time, combined with the predicted step edge map, it ensures that the robot accurately identifies the boundary of the step to avoid incorrect gait planning.
[0146] In the exemplary embodiments of the present disclosure, the robot motion control model can be pre-trained using a deep reinforcement learning algorithm. Through a large number of trainings in a simulated environment, the robot learns the optimal gait control strategy under different step heights, slopes, and surface materials. The deep reinforcement learning algorithm can be optimized based on a reward mechanism, that is, the robot tries different gait patterns during training and is given rewards or punishments according to factors such as the stability, energy consumption, and balance of successfully going up / down the steps, enabling it to form an optimal motion control strategy during continuous learning. The deep reinforcement learning algorithm can also adopt a teacher-student deep reinforcement learning framework. The teacher model is a high-precision reinforcement learning model that has been fully trained in a simulated environment or a real environment and can generate the optimal robot motion control strategy, mainly including key parameters such as gait adjustment, step length planning, step height control, body posture optimization, and foot landing point prediction. The role of the teacher model is to provide expert demonstrations or optimal strategy guidance for training the student model. The student model learns the strategy provided by the teacher model and performs reinforcement learning optimization during its own training process to enable it to adapt to different environmental changes, such as different step heights, different slopes, and complex terrains.
[0147] During the training process, the student model can imitate the teacher's optimal strategy through imitation learning or generative adversarial imitation learning, so that the student model can quickly master the basic gait control ability. The student model gradually tries different motion control parameters during the training process, and can also be optimized by combining a reward mechanism. For example, when the robot successfully and stably completes the task of going up / down stairs, a higher reward is given; if the robot falls, has an unstable gait, or steps on empty space, a penalty is given, thereby prompting the student model to optimize the gait control strategy. The trained student model can autonomously adjust the gait strategy in different environments, and calculate the motion control parameters based on the input completed elevation map and the predicted step edge map.
[0148] Among them, motion control parameters include gait patterns (such as walking, jumping, and moving slowly), step length, step height, joint angles, speed, foot landing, and other key parameters to ensure that the robot can stably perform stair walking tasks.
[0149] Step S1802, controlling the robot to go up or down stairs according to the motion control parameters.
[0150] After obtaining the motion control parameters, the motion control parameters can be used to adjust the robot's movement behavior on the stairs in real time through the robot control system instructions to ensure that it can successfully complete the task of going up / down the stairs. Specifically, the robot will adjust key links such as foot movement trajectory, center of gravity balance, and posture adjustment according to the terrain data. For example, when going up the stairs, the robot needs to raise its feet to cross the height of the steps and adjust the center of gravity to prevent falling back or falling. When going down the stairs, it is necessary to appropriately lower the gait height and optimize the landing point to ensure a stable landing. In addition, the robot can also combine real-time sensor feedback such as inertial measurement units and pressure sensors to make dynamic adjustments during movement to cope with changes in the environment.
[0151] It should be noted that the completed elevation map and the predicted step edge map are obtained according to the robot observation data completion method described in detail in step S1701 and step S1702 of another embodiment of the present disclosure, and will not be repeated here.
[0152] In the robot motion control method provided in the example implementation of the present disclosure, by using the completed elevation map and the predicted step edge map as input, the robot can accurately identify the stair structure even when there is a lack of sensor data, thereby reducing motion errors caused by incomplete data. Secondly, through the motion control model trained by the deep reinforcement learning algorithm, the robot can autonomously learn the optimal gait, adapt to different stair environments, improve the success rate of passage, and reduce energy consumption. Furthermore, based on the precise control of motion control parameters, it is ensured that the robot can adjust its posture when going up / down the stairs, prevent falls, and improve motion stability. In addition, the method also has strong environmental adaptability and can cope with different types of stairs, including stairs of different heights, widths, slopes, and materials, so that the robot has stronger universality and practical value.
[0153] Furthermore, in the exemplary implementation of the present disclosure, a robot observation data completion model training device is also provided. Figure 19 As shown, the robot observation data completion model training device 1900 may include an image processing module 1901, a loss calculation module 1902 and a network optimization module 1903, wherein: The image processing module 1901 is used to input the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain a completed elevation map and a predicted stepped edge map; The loss calculation module 1902 is used to calculate the first loss using the completed elevation map and the actual complete elevation map; and calculate the second loss using the predicted step edge map and the actual step edge map; The network optimization module 1903 is used to perform back propagation optimization on the two-dimensional convolutional neural network according to the first loss and the second loss.
[0154] The specific details of each module in the above-mentioned robot observation data completion model training device have been described in detail in the corresponding robot observation data completion model training method, so they will not be repeated here.
[0155] In an exemplary embodiment of the present disclosure, a robot observation data completion device is also provided. Figure 20 As shown, the robot observation data completion device 2000 may include an image acquisition module 2001 and a data completion module 2002, wherein: Image acquisition module 2001, used for acquiring target stepped terrain elevation map; The data completion module 2002 is used to input the target step terrain elevation map into the pre-trained robot observation data completion model to obtain the completed elevation map and the predicted step edge map; Among them, the pre-trained robot observation data completion model is obtained according to the robot observation data completion model training method in the embodiments of the present disclosure.
[0156] The specific details of each module in the above robot observation data completion device have been described in detail in the corresponding robot observation data completion method, so they will not be elaborated here.
[0157] In the exemplary embodiment of the present disclosure, a robot motion control device is also provided. Refer to Figure 21 As shown, the robot motion control device 2100 may include a parameter determination module 2101 and a motion control module 2102, where: The parameter determination module 2101 is configured to input the completed elevation map and the predicted step edge map into a robot motion control model pre-trained by a deep reinforcement learning algorithm to obtain motion control parameters; The motion control module 2102 is configured to control the robot to climb or descend the steps according to the motion control parameters; Among them, the completed elevation map and the predicted step edge map are obtained according to the robot observation data completion method in the embodiments of the present disclosure.
[0158] The specific details of each module in the above robot motion control device have been described in detail in the corresponding robot motion control method, so they will not be elaborated here.
[0159] In the exemplary embodiment of the present disclosure, a robot is also provided. The robot includes a processor and a memory. Computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, the above method is implemented. Among them, the robot includes any one of a legged robot, a quadruped robot, a biped robot, a wheeled robot, a wheel-legged robot, a four-wheeled legged robot, a humanoid robot, a cleaning robot, a transportation robot, a mobile robot, and a robotic arm. Refer to Figures 22 - 24 As shown, schematic diagrams of three different robots are respectively shown.
[0160] Refer to Figure 25 As shown, an electronic device capable of implementing the above method is also provided. Among them, the electronic device 2500 includes a processor 2501 and a memory 2502. Computer-readable instructions are stored on the memory 2502, and when the computer-readable instructions are executed by the processor 2501, the method in the embodiments of the present disclosure is implemented.
[0161] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which computer program code instructions are stored. When the computer program code instructions are called by the processor of the robot, the robot is caused to execute the method as in the embodiment.
[0162] Refer toFigure 26 As shown, a program product 2600 for implementing the above method according to an embodiment of the present disclosure is described. It can adopt a portable compact disc read-only memory (CD-ROM), include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0163] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by the way of software combined with necessary hardware. Therefore, the technical solution according to the embodiment of the present disclosure can be embodied in the form of a software product, and this software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiment of the present disclosure.
[0164] Finally, the above preferred embodiments are only used to illustrate the technical solution of the present application and are not restrictive. Although the present application has been described in detail, those skilled in the art should understand that changes in form and details can be made to it without departing from the scope defined by the claims of the present application. The dimensions of the drawings have nothing to do with the specific physical objects, and the physical dimensions can be arbitrarily changed.
Claims
1. A robot observation data completion model training method, characterized in that: include: Input the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain the completed elevation map and the predicted step edge map; Calculating a first loss using the completed elevation map and the actual complete elevation map; calculating a second loss using the predicted step edge map and the actual step edge map; The two-dimensional convolutional neural network is back-propagated and optimized according to the first loss and the second loss.
2. The robot observation data completion model training method according to claim 1, characterized in that: The method further comprises: Obtaining point cloud data of stepped terrain observation; The stepped terrain elevation map including the incomplete area is obtained according to the stepped terrain observation point cloud data.
3. The robot observation data completion model training method according to claim 2, characterized in that: The step of obtaining the step terrain elevation map including the incomplete area according to the step terrain observation point cloud data comprises: Converting the stepped terrain observation point cloud data into a first elevation map; The first elevation map is downsampled to obtain the stepped terrain elevation map including the incomplete area; the resolution of the stepped terrain elevation map including the incomplete area is lower than the resolution of the first elevation map.
4. The robot observation data completion model training method according to claim 3, characterized in that: The resolution of the first elevation map is a×b, where 0.005m≤a≤0.02m and 0.005m≤b≤0.02m; the resolution of the stepped terrain elevation map including the incomplete area is (n·a)×(n·b), where 4≤n≤10.
5. The robot observation data completion model training method according to claim 4, characterized in that: The resolution of the first elevation map is 0.01m×0.01m, and the resolution of the stepped terrain elevation map including the incomplete area is 0.05m×0.05m.
6. The robot observation data completion model training method according to claim 2, characterized in that: The method further comprises: Noise data is added to the step terrain observation point cloud data; wherein the noise data includes a randomly added plane point cloud, and the area of the plane point cloud is positively correlated with the foot coverage area of the robot.
7. The robot observation data completion model training method according to claim 6, characterized in that: The noise data also includes: Gaussian noise, a plurality of randomly added external points around a certain point and / or randomly cut data.
8. The robot observation data completion model training method according to claim 1, characterized in that: The two-dimensional convolutional neural network model includes: A first encoder is used to perform dimensionality reduction processing on the stepped terrain elevation map containing the incomplete area to extract image features; A first decoder is used for performing dimensionality restoration according to the image features; and the decoder is connected to a depth head for outputting the completed elevation map and a segmentation head for outputting the predicted step edge map.
9. The robot observation data completion model training method according to claim 8, characterized in that: The predicted step edge map output by the segmentation head is also used to input to the depth head.
10. The robot observation data completion model training method according to claim 1, characterized in that: The two-dimensional convolutional neural network model includes: A second encoder is used to perform dimensionality reduction processing on the stepped terrain elevation map containing the incomplete area to extract image features; A second decoder, configured to perform dimensionality restoration according to the image features and output the completed elevation map; The third decoder is used to perform dimensionality restoration according to the image features and output the predicted stair-step edge map.
11. The robot observation data completion model training method according to claim 1, characterized in that: The calculating the first loss using the completed elevation map and the actual complete elevation map comprises: according to: Calculate and obtain the first loss; in, For the first loss, is the true height value of the ith pixel in the actual complete elevation map, is the predicted height value of the i-th pixel in the completed elevation map, and n is the number of all pixels in the actual complete elevation map and the completed elevation map.
12. The robot observation data completion model training method according to claim 1, characterized in that: The calculating the second loss using the predicted step edge map and the actual step edge map comprises: according to: Calculate and obtain the second loss; Among them, BCE Loss is the second loss, is the true label corresponding to the i-th pixel in the actual step edge map, is the probability that the i-th pixel in the predicted step edge map belongs to the edge, and N is the number of all pixels in the predicted step edge map and the actual step edge map.
13. The robot observation data completion model training method according to any one of claims 1 to 12, characterized in that: The performing back propagation optimization on the two-dimensional convolutional neural network according to the first loss and the second loss includes: Calculate the comprehensive loss based on the first loss and the second loss; With the goal of minimizing the comprehensive loss, back-propagation optimization is performed on the model parameters in the two-dimensional convolutional neural network.
14. The robot observation data completion model training method according to claim 13, characterized in that: The comprehensive loss includes a first comprehensive loss; and the calculation of the comprehensive loss based on the first loss and the second loss includes: according to: The first comprehensive loss is calculated; in, is the first comprehensive loss, is the first loss, BCE Loss is the second loss, and is the weight.
15. The robot observation data completion model training method according to claim 13, characterized in that: The comprehensive loss includes a second comprehensive loss; and the calculation of the comprehensive loss based on the first loss and the second loss includes: Generate a corresponding weight mask according to the image area whose height value exceeds the preset height threshold; The first loss and the second loss are adjusted using the weight mask, and the second comprehensive loss is calculated based on the adjusted first loss and the second loss.
16. The robot observation data completion model training method according to claim 15, characterized in that: The step of adjusting the first loss and the second loss by using the weight mask, and calculating the second comprehensive loss according to the adjusted first loss and the second loss, includes: according to: Calculate and obtain the second comprehensive loss; in, is the second comprehensive loss, Mask(i) is the weight mask, and Mask(i)=α(0<α<1) means that the loss weight α is assigned to the i-th pixel whose height value exceeds the preset height threshold. For the first loss, To use weight mask Adjusted first loss, For the second loss, To use weight mask Adjusted second loss, and is the weight.
17. A robot observation data completion method, characterized in that: include: Collect target stepped terrain elevation map; Inputting the target stepped terrain elevation map into a pre-trained robot observation data completion model to obtain a completed elevation map and a predicted stepped edge map; Wherein, the pre-trained robot observation data completion model is obtained according to the robot observation data completion model training method according to any one of claims 1 to 16.
18. A robot motion control method, characterized in that: include: The completed elevation map and the predicted step edge map are input into the robot motion control model pre-trained by the deep reinforcement learning algorithm to obtain the motion control parameters; Control the robot to go up or down stairs according to the motion control parameters; The completed elevation map and the predicted step edge map are Obtained according to the robot observation data completion method described in claim 17.
19. A robot observation data completion model training device, characterized in that: include: An image processing module is used to input the stepped terrain elevation map containing the incomplete area into the two-dimensional convolutional neural network to be trained to obtain a completed elevation map and a predicted stepped edge map; A loss calculation module, used to calculate a first loss using the completed elevation map and the actual complete elevation map; and calculate a second loss using the predicted step edge map and the actual step edge map; A network optimization module is used to perform back propagation optimization on the two-dimensional convolutional neural network according to the first loss and the second loss.
20. A robot observation data completion device, characterized in that: include: An image acquisition module is used to acquire a target stepped terrain elevation map; A data completion module, used for inputting the target stepped terrain elevation map into a pre-trained robot observation data completion model to obtain a completed elevation map and a predicted stepped edge map; Wherein, the pre-trained robot observation data completion model is obtained according to the robot observation data completion model training method according to any one of claims 1 to 16.
21. A robot motion control device, characterized in that: include: A parameter determination module, used to input the completed elevation map and the predicted step edge map into a robot motion control model pre-trained by a deep reinforcement learning algorithm to obtain motion control parameters; A motion control module, used to control the robot to go up or down stairs according to the motion control parameters; The completed elevation map and the predicted step edge map are Obtained according to the robot observation data completion method described in claim 17.
22. An electronic device, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method as claimed in any one of claims 1 to 16 or 17 or 18.
23. A robot, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method as claimed in any one of claims 1 to 16 or 17 or 18.
24. The robot according to claim 23, characterized in that The robot includes any one of a legged robot, a quadruped robot, a bipedal robot, a wheeled robot, a wheel-legged robot, a quadrupedal robot, a humanoid robot, a cleaning robot, a transport robot, a mobile robot and a robotic arm.
25. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program code instructions, and when the computer program code instructions are called by a processor of the robot, the robot executes the method as described in any one of claims 1 to 16 or 17 or 18.
Citation Information
Patent Citations
Image depth completion method and device, computer equipment and storage medium
CN117788546A
Depth completion method applied to sparse map densification
CN110097589A
Model training method, figure image completion method and apparatus, and electronic device
CN112488284A
Image completion method based on content attention mechanism and mask prior
CN112686816A
Depth edge extraction method and computer readable storage medium
CN116563321A
Cited By
Radar data blind area filling method and system based on 3D-GANs
CN121010708A