Training Method and Device for Terrain Perception Model of Legged Robot
By training the terrain-aware model, obtaining sample ground images and determining the label semantic category and flatness values of pixel points, the safety problem of foot robots moving in complex terrain is solved, and safe and stable movement is achieved.
Patent Information
- Application Number
- CN202311720705.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-12-13
AI Technical Summary
The prior art is difficult to effectively provide key information for foot robots to move safely and smoothly in complex terrain environments, including terrain semantic information and flatness information.
By training the terrain-aware model, obtaining sample ground images, determining the label semantic category and flatness value of pixel points, using the predicted flatness value and label flatness value for loss calculation, and training the model to output flatness value with semantic category and edge-aware attributes.
It has achieved a more reasonable and safe feasible path in complex terrain to ensure the safe and stable movement of foot robots on the road surface.
Smart Images

Figure CN117765389B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot vision technology, and in particular, to a method and device for training a terrain perception model based on a legged robot. Background Art
[0002] With the development of technology, artificial intelligence technology has developed rapidly, and intelligent robots obtained based on artificial intelligence technology are widely used in different fields. Among them, the legged robot is one of many types of robots. Because it has wheeled feet and tracked feet, it still has high mobility even in complex terrain environments.
[0003] Generally, when a legged robot performs a mobile operation, it is necessary to collect the current terrain image through an image acquisition device, and based on the terrain image, analyze the physical geometric information, spatial shape information, texture information, etc. of the current terrain to establish a real and accurate physical environment to ensure the safety of the legged robot during movement. Among them, the flatness information of the ground and the semantic information of the ground are key information to ensure that the legged robot can move safely and smoothly. Therefore, how to provide terrain semantic information and terrain flatness information for the legged robot is an important issue.
[0004] Based on this, the specification of this application provides a method and device for training a terrain perception model based on a legged robot. Summary of the Invention
[0005] This specification provides a method, device, storage medium, and electronic device for training a terrain perception model based on a legged robot to at least partially solve the above problems existing in the prior art.
[0006] This specification adopts the following technical solutions:
[0007] This specification provides a method for training a terrain perception model based on a legged robot, and the method includes:
[0008] Obtain a sample ground image;
[0009] For each pixel point in the sample ground image, determine the label semantic category of the pixel point;
[0010] Determine a plane formed by a plurality of pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point;
[0011] According to the plane, determine the label flatness value of the pixel point;
[0012] Input the sample ground image into the terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model;
[0013] Determine a loss based on the predicted flatness value and the labeled flatness value, and train the terrain perception model according to the loss.
[0014] Optionally, determining the labeled flatness value of the pixel point specifically includes:
[0015] Use the distance from the pixel point to the plane as the labeled flatness value of the pixel point.
[0016] Optionally, inputting the sample ground image into the terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model, specifically including:
[0017] Input the sample ground image into the terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model, and obtain the predicted semantic category of each pixel point in the sample ground image output by the terrain perception model.
[0018] Optionally, training the terrain perception model specifically includes:
[0019] Determine a first loss based on the predicted flatness value and the labeled flatness value, and determine a second loss based on the predicted semantic category and the labeled semantic category;
[0020] Train the terrain perception model according to the first loss and the second loss.
[0021] Optionally, the terrain perception model includes: an encoder, a first decoder, a second decoder, a first prediction layer, and a second prediction layer;
[0022] Inputting the sample ground image into the terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model, specifically including:
[0023] Input the sample ground image into the encoder to obtain encoded image features;
[0024] Input the encoded image features into the first decoder to obtain first decoded image features; and input the encoded image features and the first decoded image features into the first prediction layer to obtain the predicted semantic category of each pixel point in the sample ground image;
[0025] Input the encoded image features into the second decoder to obtain second decoded image features; and input the encoded image features and the second decoded image features into the second prediction layer to obtain the predicted flatness value of each pixel point in the sample ground image.
[0026] Optionally, determining a plane formed by a plurality of pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point specifically includes:
[0027] Performing smoothing processing on each pixel point in the sample ground image, and for each pixel point after the smoothing processing, determining a plane formed by a plurality of pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point.
[0028] Optionally, the method further includes:
[0029] Obtaining a current ground image;
[0030] Inputting the current ground image into a trained terrain perception model to obtain the flatness value of each pixel point in the current ground image output by the terrain perception model;
[0031] According to the obtained flatness value, determining a feasible path of the legged robot in the current ground image.
[0032] This specification provides a training device for a terrain perception model based on a legged robot, including:
[0033] An acquisition module, configured to acquire a sample ground image;
[0034] A semantic determination module, configured to determine the label semantic category of each pixel point in the sample ground image;
[0035] A plane determination module, configured to determine a plane formed by a plurality of pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point;
[0036] A flatness determination module, configured to determine the label flatness value of the pixel point according to the plane;
[0037] An input module, configured to input the sample ground image into the terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model;
[0038] A training module, configured to determine a loss according to the predicted flatness value and the label flatness value, and train the terrain perception model according to the loss.
[0039] Optionally, the flatness determination module is specifically configured to use the distance from the pixel point to the plane as the label flatness value of the pixel point.
[0040] Optionally, the input module is specifically configured to input the sample ground image into the terrain perception model, obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model, and obtain the predicted semantic category of each pixel point in the sample ground image output by the terrain perception model.
[0041] Optionally, the training module is specifically configured to determine a first loss according to the predicted flatness value and the labeled flatness value, and determine a second loss according to the predicted semantic category and the labeled semantic category; train the terrain perception model according to the first loss and the second loss.
[0042] Optionally, the terrain perception model includes: an encoder, a first decoder, a second decoder, a first prediction layer, and a second prediction layer;
[0043] The input module is specifically configured to input the sample ground image into the encoder to obtain encoded image features; input the encoded image features into the first decoder to obtain first decoded image features; and input the encoded image features and the first decoded image features into the first prediction layer to obtain the predicted semantic category of each pixel point in the sample ground image; input the encoded image features into the second decoder to obtain second decoded image features; and input the encoded image features and the second decoded image features into the second prediction layer to obtain the predicted flatness value of each pixel point in the sample ground image.
[0044] Optionally, the plane determination module is specifically configured to perform smoothing processing on each pixel point in the sample ground image, and for each pixel point after smoothing processing, determine a plane composed of a plurality of pixel points whose distance from the pixel point is within a preset range and whose labeled semantic category is different from that of the pixel point.
[0045] Optionally, the device further includes an application module;
[0046] The application module is specifically configured to obtain the current ground image; input the current ground image into the trained terrain perception model to obtain the flatness value of each pixel point in the current ground image output by the terrain perception model; determine the feasible path of the legged robot in the current ground image according to the obtained flatness value.
[0047] This specification provides a computer-readable storage medium, and the storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned training method of the terrain perception model based on a legged robot.
[0048] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned training method of the terrain perception model based on the legged robot is implemented.
[0049] The above-mentioned at least one technical solution adopted in this specification can achieve the following beneficial effects:
[0050] In the training method of the terrain perception model based on the legged robot provided in this specification, a sample ground image can be obtained first. For each pixel point in the sample ground image, the label semantic category of the pixel point is determined, and a plane composed of multiple pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point is determined, so as to determine the label flatness value of the pixel point according to the plane. Then the sample ground image is input into the terrain perception model, and the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model is obtained. Finally, according to the predicted flatness value and the label flatness value, the loss is determined, and the terrain perception model is trained according to the loss.
[0051] When determining the flatness value, this method considers the semantic category information and the road surface edge information, that is, the information at the junction of different semantic categories, so that the obtained flatness value has semantic category and edge perception attributes, so as to analyze the influence of the changes of different semantic categories and the edges on the passing ability of the legged robot on the ground, and then a more reasonable and safe feasible path can be planned to ensure the safe and stable passing of the legged robot on the road surface. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The schematic embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation to this specification. In the drawings:
[0053] Figure 1 is a schematic flowchart of a training method of a terrain perception model based on a legged robot in this specification;
[0054] Figure 2 is a schematic diagram of a plane in this specification;
[0055] Figure 3 is a schematic structural diagram of a terrain perception model in this specification;
[0056] Figure 4 is a schematic diagram of a plane in this specification;
[0057] Figure 5 is a schematic diagram of a training device of a terrain perception model based on a legged robot provided in this specification;
[0058] Figure 6 The schematic diagram of the electronic device provided for this specification corresponding to Figure 1 is shown as follows. Detailed implementation manners
[0059] For a legged robot, on the one hand, its passing ability on different types of ground with the same flatness is different. For example, on asphalt roads and grasslands with the same flatness, the passing ability of a bipedal robot on the asphalt road is stronger than that on the grassland. On the other hand, its passing ability on roads with the same type but different flatness is also different. For example, on an asphalt road, the flatter the ground, the stronger the passing ability of the bipedal robot, and the more uneven the ground, the weaker the passing ability of the bipedal robot. It can be seen that the flatness of the ground and the type of the ground have a great influence on the passing ability of the legged robot. Therefore, the flatness of the ground and the type of the ground are important reference information for planning the feasible path of the legged robot.
[0060] Based on this, the specification of this application provides a training method for a terrain perception model based on a legged robot, which can train a terrain perception model that can simultaneously output the semantic category of the ground and the flatness value of the ground. When obtaining the flatness value, the semantic category information and the road edge information, that is, the information at the junction of different semantic categories, are considered, so that the obtained flatness value has semantic category and edge perception attributes, thereby analyzing the influence of the changes in different semantic categories and the edges on the passing ability of the legged robot on the ground, and then planning a more reasonable and safe feasible path to ensure the safe and stable passing of the legged robot on the road surface.
[0061] To make the purpose, technical solution and advantages of this specification clearer, the technical solution of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0062] The following will detail the technical solutions provided by each embodiment of this specification with reference to the drawings.
[0063] Figure 1 The flowchart of a training method for a terrain perception model based on a legged robot provided for this specification is shown as follows, and specifically may include the following steps:
[0064] S100: Obtain a sample ground image.
[0065] S102: For each pixel point in the sample ground image, determine the label semantic category of the pixel point.
[0066] The execution entity implementing the technical solution of this specification can be a server that controls the movement of a legged robot and has capabilities such as communication and computing.
[0067] First of all, the server can obtain the sample ground image, and for each pixel point in the sample ground image, determine the label semantic category of the pixel point. Among them, the sample ground image can include a binocular color image and a depth map. Of course, the sample ground image can also be a three-dimensional image. The semantic category can be set in advance, and the semantic category represents the category of the ground. In one or more embodiments of this specification, the semantic category at least includes: asphalt road, brick road, wooden road, grassland, gravel ground, and sandy ground. Of course, it can also include a water surface road, and can also include an icy road surface, etc. What the semantic category specifically includes is not limited in this specification and can be set according to the specific scenario and specific requirements.
[0068] S104: Determine a plane formed by multiple pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point.
[0069] S106: Determine the label flatness value of the pixel point according to the plane.
[0070] Secondly, the server can determine a plane formed by multiple pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point. In other words, the server can first determine multiple pixel points whose distance from the pixel point is within a preset range and whose semantic label category is different from that of the pixel point, and determine the plane formed by the multiple pixel points.
[0071] Furthermore, based on the obtained plane, determine the label flatness value of the pixel point. Specifically, the server can use the distance from the pixel point to the plane as the label flatness value of the pixel point. As shown in the following formula:
[0072]
[0073] Among them, R is the flatness value of the pixel point, (a, b, c) is the surface normal vector of the plane formed by multiple pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point, and d is the plane coefficient.
[0074] In one or more embodiments of the present specification, when determining a plane formed by a plurality of pixel points whose distances from the pixel point are within a preset range and whose label semantic categories are different from that of the pixel point, the least squares method can be used for plane fitting, or the RANSAC algorithm can be used for fitting. Specifically, the present specification does not make any restrictions as long as the plane is obtained by fitting based on a number of pixel points whose semantic labels are different from that of the pixel point and whose distances from the pixel point are within a preset range.
[0075] Based on the method for obtaining the plane described above, it can be known that the plane has a semantic category and edge perception. Because when fitting the plane, the pixel points whose distances from the pixel point are within a preset range, or in other words, the pixel points adjacent to the pixel point, have different semantic categories from the pixel point. That is, the plane obtained by fitting is a plane that divides the pixel point from the pixel points corresponding to other semantic categories. As Figure 2 shown, it is a schematic diagram of a plane provided in the specification of the present application.
[0076] S108: Input the sample ground image into the terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model.
[0077] S110: Determine the loss according to the predicted flatness value and the labeled flatness value, and train the terrain perception model according to the loss.
[0078] Finally, the server can input the sample ground image into the terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model. The server can determine the loss based on the predicted flatness value and the labeled flatness value, so as to train the terrain perception model based on this loss. This method can enable the trained terrain perception model to output the flatness value of each pixel point in the ground image, and the flatness value has the semantics and edge perception described in the above steps S104 - S106. As mentioned before, the semantic information and flatness information of the ground have a great impact on the movement of the legged robot. Therefore, the flatness value with semantics and edge perception output by the terrain perception model can better help plan the movement path of the legged robot and ensure the movement safety of the legged robot.
[0079] In one or more embodiments of the present specification, when determining the loss based on the predicted flatness value and the labeled flatness value, the loss can be the L1 loss.
[0080] Based on Figure 1In the above-mentioned training method of the terrain perception model based on the legged robot provided in this specification, a sample ground image can be obtained first. For each pixel point in the sample ground image, the label semantic category of the pixel point is determined, and a plane formed by multiple pixel points whose distances from the pixel point are within a preset range and whose label semantic categories are different from that of the pixel point is determined, so as to determine the label flatness value of the pixel point according to the plane. Then, the sample ground image is input into the terrain perception model, and the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model is obtained. Finally, according to the predicted flatness value and the label flatness value, the loss is determined, and the terrain perception model is trained according to the loss. When determining the flatness value, this method considers the semantic category information and the road surface edge information, that is, the information at the junction of different semantic categories, so that the obtained flatness value has the semantic category and edge perception attributes, thereby analyzing the influence of the changes of different semantic categories and the edges on the passing ability of the legged robot on the ground, and then planning a more reasonable and safe feasible path to ensure the safe and stable passing of the legged robot on the road surface.
[0081] In addition, the terrain perception model can also output the semantic category of the ground. That is, in the above step S108, when the sample ground image is input into the terrain perception model, not only the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model can be obtained, but also the predicted semantic category of each pixel point in the sample ground image output by the terrain perception model can be obtained.
[0082] Then, when training the terrain perception model, the server can not only determine the first loss according to the predicted flatness value and the label flatness value, but also determine the second loss according to the predicted semantic category and the label semantic category, so as to train the terrain perception model according to the first loss and the second loss.
[0083] In one or more embodiments of this specification, as described above, the first loss can be the L1 loss, and the second loss can be the cross-entropy loss. Of course, other loss functions can also be set, and this specification does not make specific limitations.
[0084] As Figure 3 shown, it is a schematic structural diagram of a terrain perception model provided in the specification of this application. It can be seen that in this specification, the terrain perception model can include an encoder, a first decoder, a second decoder, a first prediction layer, and a second prediction layer.
[0085] When inputting the sample ground image into the terrain perception model, the server may first input the sample ground image into the encoder to obtain encoded image features. Then, the encoded image features are input into the first decoder to obtain first decoded image features, and the encoded image features and the first decoded image features are input into the first prediction layer to obtain the predicted semantic category of each pixel point in the sample ground image. When inputting the encoded image features into the first decoder, the encoded image features can be input into the second decoder to obtain second decoded image features, and the encoded image features and the second decoded image features are input into the second prediction layer to obtain the predicted flatness value of each pixel point in the sample ground image.
[0086] Further, in this specification, the encoder may include 4 downsampling modules. Any one of the downsampling modules may include a convolutional layer, a convolutional block Conv3×3 for adjusting the number of channels. The convolutional block Conv3×3 includes a convolutional layer with a convolution kernel of 3×3, a BN layer for implementing batch normalization, and a Leaky ReLU activation function with a parameter of 0.2, and a max pooling layer.
[0087] When inputting the sample ground image into the encoder, the sample ground image is first processed by the first convolutional layer to obtain first primary features, and then enters the convolutional block Conv3×3 to obtain second primary features. Then, the first primary features and the second primary features can be concatenated to obtain fused features, and the fused features are input into a max pooling layer, so that the first-level image features processed by the first downsampling module can be obtained. Then, using the method described above, the sample ground image can be processed by the second downsampling module to obtain second-level image features, processed by the third downsampling module to obtain third-level image features, and processed by the fourth downsampling module to obtain fourth-level image features. Then, the first-level image features, second-level image features, third-level image features, and fourth-level image features are the encoded image features output by the encoder.
[0088] The encoded image features can be respectively input into the first decoder and the second decoder. Any one of the decoders may include: two convolutional blocks, and both of the two convolutional blocks may include a convolutional layer with a convolution kernel of 3×3 and a Leaky ReLU activation function. Then, the first decoder and the second decoder can respectively output first decoded image features and second decoded image features.
[0089] The server can input the first decoded image feature and the encoded image feature into a first prediction layer, which can be used to predict the semantic category of the ground, and input the second decoded image feature and the encoded image feature into a second prediction layer, which can be used to predict the flatness value of the ground. Any of the prediction layers may include 4 upsampling modules, and each upsampling module can receive the input from the upper layer of the network, and the feature map that makes a skip connection with the downsampling module of the encoder, that is, the encoded image feature.
[0090] Specifically, when any prediction layer processes the input features, the nearest neighbor interpolation with a coefficient of 2 can be first used on the first decoded image feature by an upsampling function, and the feature map after the nearest neighbor interpolation upsampling is concatenated with the encoded image feature. The concatenated features can be input into the first convolutional block, which may include a convolutional layer with a 3×3 convolutional kernel, a BN layer for batch normalization, and a Leaky ReLU activation function. Then it enters the second convolutional block, and the structure of this convolutional block is the same as that of the first convolutional block. Next, the same upsampling module structure is used to continue the feature recovery, that is, after being processed by the second downsampling module, including nearest neighbor interpolation, concatenation operation with the three-level image feature in the encoded image feature, and two convolutional blocks, after being processed by the third downsampling module, including nearest neighbor interpolation, concatenation operation with the two-level image feature in the encoded image feature, and two convolutional blocks, after being processed by the fourth downsampling module, including nearest neighbor interpolation, concatenation operation with the one-level image feature in the encoded image feature, and two convolutional blocks, so as to obtain the predicted image features. Finally, the predicted features are input into a convolutional layer with a 3×3 convolutional kernel and a ReLU activation function, so as to obtain the prediction result, and the prediction result can be a predicted feature map. In this specification, when the prediction layer is the first prediction layer, the prediction result is the predicted semantic category, and when the prediction layer is the second prediction layer, the prediction result is the predicted flatness value.
[0091] In addition, when training the terrain perception model, the backpropagation algorithm can be used to train the terrain perception model, and the model parameters are adjusted by calculating the loss. The Adam optimizer can be used during training, and the parameters of this optimizer can be β1 = 0.9, β2 = 0.999, ∈ = 1e-8, the original learning rate can be 1e-3, and the learning rate decays by 0.5 after every 50,000 iterations. The number of single-sample ground images can be 12, and the total number of iterations can be 200,000.
[0092] In this specification, in the above step S104, when determining a plane formed by a plurality of pixel points whose distances from the pixel point are within a preset range and whose label semantic categories are different from that of the pixel point, each pixel point in the sample ground image can be smoothed first, and then for each pixel point after smoothing, a plane formed by a plurality of pixel points whose distances from the pixel point are within a preset range and whose label semantic categories are different from that of the pixel point can be determined. Smoothing the pixel points can improve the accuracy of the flatness value determined in the subsequent steps, that is, make the distances from each determined pixel point to the plane more accurate.
[0093] Specifically, when smoothing the pixel points, the server can, for each pixel point in the sample ground image, determine the physical distance and pixel distance from each other pixel point in the sliding window to the pixel point according to a sliding window of a specified size centered on the pixel point. Then, using the pixel distances of the pixel points in the sliding window as the first weights, the physical distances of the pixel points in the sliding window are weighted and averaged to obtain the second weight of the pixel point. Finally, according to the second weights of each pixel point in the obtained sample ground image, each pixel point in the sample ground image is smoothed to obtain each pixel point after smoothing.
[0094] It should be noted that in this specification, the flatness value of a pixel point is non - negative, and the minimum value of the flatness value can be 0. When the flatness value is 0, it means that the pixel point is on the plane obtained by fitting. The maximum value of the flatness value cannot exceed the depth distance from the pixel point to the camera.
[0095] In addition, in this specification, the above - mentioned pixel point can be a spatial point in a three - dimensional point cloud obtained based on the sample ground image. That is, based on the sample ground image, a three - dimensional point cloud corresponding to the ground image is obtained. In one or more embodiments of this specification, the three - dimensional point cloud can also be completed to obtain a dense and accurate three - dimensional point cloud, so that the fitted plane and the determined flatness value, etc. are more accurate.
[0096] This specification also provides an application method for the trained terrain perception model. In this specification, the server can obtain the current ground image, then input the current ground image into the trained terrain perception model, obtain the flatness value of each pixel point in the current ground image output by the terrain perception model, and determine the feasible path of the legged robot in the current ground image according to the obtained flatness value.
[0097] When planning a feasible path for a legged robot, the difference in the flatness values of the pixel points within a specified range can be determined. When this difference is greater than a preset threshold, it indicates that the road surface within the specified range is uneven. When this difference is not greater than the preset threshold, it indicates that the road surface within the specified range is flat. Then, the road surface within the specified range can be determined as one of the feasible paths for the legged robot.
[0098] Of course, the semantic category of each pixel point in the current ground image output by the terrain perception model can also be obtained. Then, based on the obtained semantic category and flatness value, a feasible path for the legged robot can be planned in the current ground image. When planning a feasible path for the legged robot, flatness thresholds for different semantic categories can be preset. For example, when the semantic category of the ground is grassland, the flatness difference threshold for the pixel points in this ground image can be X. When the semantic category of the ground is an icy ground, the flatness difference threshold for the pixel points in this ground image can be Y. Then, the server can determine the semantic category of the current road surface, determine the flatness difference threshold corresponding to this semantic category, and determine the flatness difference of the road surface within a specified range in the current road surface, and compare it with the flatness difference threshold. When this difference is greater than the preset threshold, it indicates that the road surface within the specified range is uneven. When this difference is not greater than the preset threshold, it indicates that the road surface within the specified range is flat. Then, the road surface within the specified range can be determined as one of the feasible paths for the legged robot.
[0099] Furthermore, in one or more embodiments of this specification, in order to further improve the accuracy of the flatness value output by the terrain perception model, or rather, to further enhance the edge perception information possessed by the flatness value, in the above steps S104 - S106, that is, after determining the label semantic category of each pixel point in the sample ground image, the server can first use a multi-plane segmentation algorithm to segment each pixel point in the sample ground image to obtain each first plane. Then, for each pixel point in the sample ground image, the server can determine multiple pixel points whose distance from this pixel point is within a preset range and whose label semantic category is different from this pixel point, and use them as each candidate pixel point, determine the first plane where each candidate pixel point is located, and determine the candidate pixel points where the first planes are different. Finally, based on the determined candidate pixel points where the first planes are different, fit a second plane, and according to this second plane, determine the label flatness value of this pixel point, so as to train the terrain perception model in the manner described in the above steps S108 - S110 based on this label flatness value.
[0100] As Figure 4 shown, it is a schematic diagram of a plane provided in the specification of this application. As can be seen in Figure 4 it, the locally fitted plane is located at the junction of the ground corresponding to two semantic categories, taking into account semantic information and edge information.
[0101] Among them, when fitting the second plane from each candidate pixel point that is different from the determined first plane where it is located, a specified number of pixel points can be selected from each candidate pixel point that is different from the determined first plane where it is located, and plane fitting is performed to obtain the second plane. The method used for plane fitting can be the least squares method, the RANSAC algorithm, etc. as described in the above steps S104 - S106. Specifically, this specification does not make any restrictions.
[0102] Since the category of the ground, the flatness of the ground, and the changes in the ground edge have a greater impact on the legged robot, based on the above method, when determining the flatness value, the semantic category information of the ground and the height of the ground or the changes in the ground edge are considered simultaneously, so that the fitted plane has semantic category and edge perception attributes, improving the accuracy of the flatness value pre-output by the model, and further improving the safety and stability of the legged robot's movement.
[0103] Based on the training method of the terrain perception model for legged robots described above, the embodiments of this specification also correspondingly provide a schematic diagram of a training device for the terrain perception model for legged robots, as Figure 5 shown.
[0104] Figure 5 The following is a schematic diagram of a training device for the terrain perception model for legged robots provided by the embodiments of this specification. The device includes:
[0105] An acquisition module 500, configured to acquire a sample ground image;
[0106] A semantic determination module 502, configured to determine the label semantic category of each pixel point in the sample ground image;
[0107] A plane determination module 504, configured to determine a plane formed by a plurality of pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point;
[0108] A flatness determination module 506, configured to determine the label flatness value of the pixel point according to the plane;
[0109] An input module 508, configured to input the sample ground image into the terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model;
[0110] A training module 510, configured to determine a loss according to the predicted flatness value and the label flatness value, and train the terrain perception model according to the loss.
[0111] Optionally, the flatness determination module 506 is specifically configured to use the distance from the pixel to the plane as the label flatness value of the pixel.
[0112] Optionally, the input module 508 is specifically configured to input the sample ground image into the terrain perception model, obtain the predicted flatness value of each pixel in the sample ground image output by the terrain perception model, and obtain the predicted semantic category of each pixel in the sample ground image output by the terrain perception model.
[0113] Optionally, the training module 510 is specifically configured to determine a first loss according to the predicted flatness value and the label flatness value, and determine a second loss according to the predicted semantic category and the label semantic category; train the terrain perception model according to the first loss and the second loss.
[0114] Optionally, the terrain perception model includes: an encoder, a first decoder, a second decoder, a first prediction layer, and a second prediction layer;
[0115] The input module 508 is specifically configured to input the sample ground image into the encoder to obtain encoded image features; input the encoded image features into the first decoder to obtain first decoded image features; and input the encoded image features and the first decoded image features into the first prediction layer to obtain the predicted semantic category of each pixel in the sample ground image; input the encoded image features into the second decoder to obtain second decoded image features; and input the encoded image features and the second decoded image features into the second prediction layer to obtain the predicted flatness value of each pixel in the sample ground image.
[0116] Optionally, the plane determination module 504 is specifically configured to perform smoothing processing on each pixel in the sample ground image, and for each pixel after smoothing processing, determine a plane formed by a plurality of pixels whose distance from the pixel is within a preset range and whose label semantic category is different from that of the pixel.
[0117] Optionally, the device further includes an application module 512;
[0118] The application module 512 is specifically configured to obtain the current ground image; input the current ground image into the trained terrain perception model to obtain the flatness value of each pixel in the current ground image output by the terrain perception model; and determine the feasible path of the legged robot in the current ground image according to the obtained flatness value.
[0119] An embodiment of this specification also provides a computer-readable storage medium storing a computer program that can be used to execute the training method of the terrain perception model based on a legged robot described above.
[0120] Based on the training method of the terrain perception model based on a legged robot described above, an embodiment of this specification also proposes Figure 6 the schematic structural diagram of the electronic device shown. As Figure 6 , at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the training method of the terrain perception model based on a legged robot described above.
[0121] Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.
[0122] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0123] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0124] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0125] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0126] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0127] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block of the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to the processors of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart Figure 1 for one or more of the flows and / or blocks Figure 1 and / or means for implementing the functions specified in one or more of the blocks.
[0128] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flowchart Figure 1 for one or more of the flows and / or blocks Figure 1 and / or means for implementing the functions specified in one or more of the blocks.
[0129] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart Figure 1 for one or more of the flows and / or blocks Figure 1 and / or means for implementing the functions specified in one or more of the blocks.
[0130] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0131] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0132] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0133] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0134] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0135] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0136] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment.
[0137] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this application.
Claims
1. A training method for a terrain perception model based on a legged robot, characterized in that, The method includes: Obtaining a sample ground image; For each pixel point in the sample ground image, determining the label semantic category of the pixel point; Determining a plane formed by a plurality of pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point; Determining the label flatness value of the pixel point according to the plane; Inputting the sample ground image into a terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model; Determining a loss according to the predicted flatness value and the label flatness value, and training the terrain perception model according to the loss.
2. The method according to claim 1, wherein Determining the label flatness value of the pixel point specifically includes: Taking the distance from the pixel point to the plane as the label flatness value of the pixel point.
3. The method according to claim 1, wherein Inputting the sample ground image into a terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model specifically includes: Inputting the sample ground image into a terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model, and obtaining the predicted semantic category of each pixel point in the sample ground image output by the terrain perception model.
4. The method according to claim 3, wherein Training the terrain perception model specifically includes: Determining a first loss according to the predicted flatness value and the label flatness value, and determining a second loss according to the predicted semantic category and the label semantic category; Training the terrain perception model according to the first loss and the second loss.
5. The method according to claim 3, wherein The terrain perception model includes: an encoder, a first decoder, a second decoder, a first prediction layer, and a second prediction layer; Inputting the sample ground image into a terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model specifically includes: Inputting the sample ground image into the encoder to obtain encoded image features; Inputting the encoded image features into the first decoder to obtain first decoded image features; and inputting the encoded image features and the first decoded image features into the first prediction layer to obtain the predicted semantic category of each pixel point in the sample ground image; Inputting the encoded image features into the second decoder to obtain second decoded image features; and inputting the encoded image features and the second decoded image features into the second prediction layer to obtain the predicted flatness value of each pixel point in the sample ground image.
6. The method according to claim 1, characterized in that, Determining a plane formed by a plurality of pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point specifically includes: Performing smoothing processing on each pixel point in the sample ground image, and for each pixel point after the smoothing processing, determining a plane formed by a plurality of pixel points whose distance from the pixel point is within a preset range and whose label semantic category is different from that of the pixel point.
7. The method according to claim 1, wherein The method further includes: Obtaining a current ground image; Input the current ground image into the trained terrain perception model to obtain the flatness value of each pixel point in the current ground image output by the terrain perception model; According to the obtained flatness value, determine the feasible path of the legged robot in the current ground image.
8. A training device for a terrain perception model based on a legged robot, characterized in that, The device specifically includes: An acquisition module, configured to acquire a sample ground image; A semantic determination module, configured to determine the label semantic category of each pixel point in the sample ground image; A plane determination module, configured to determine a plane composed of a plurality of pixel points whose distances from the pixel point are within a preset range and whose label semantic categories are different from that of the pixel point; A flatness determination module, configured to determine the label flatness value of the pixel point according to the plane; An input module, configured to input the sample ground image into the terrain perception model to obtain the predicted flatness value of each pixel point in the sample ground image output by the terrain perception model; A training module, configured to determine a loss according to the predicted flatness value and the label flatness value, and train the terrain perception model according to the loss.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1-7 above is implemented.
10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in any one of claims 1-7 above is implemented.
Citation Information
Patent Citations
Remote sensing image segmentation method and system based on multi-scale feature fusion
CN112801109A
Image processing method and device, storage medium and electronic equipment
CN115496930A