System and method for assisting navigation of a mobile system

By combining camera and rangefinder data, using convolutional neural networks to generate a passability map of the terrain, the navigation problem of robots or autonomous vehicles moving on terrain in the prior art on non-road paths is solved, achieving more efficient and reliable navigation.

CN120167033APending Publication Date: 2025-06-17SAFRAN SA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380077199.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-07
Filing Date
2023-11-07
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of navigation of robots or autonomous vehicles moving on terrain on non-road paths, especially the challenges of identifying and evaluating terrain’s passability.

Method used

By combining the optical images acquired by the camera and the 3D point cloud data acquired by the rangefinder, a convolutional neural network is used to process the depth images and uncertainty masks to generate semantic maps, depth maps and confidence maps to finally determine the passability maps of the terrain.

Benefits of technology

Accurate assessment of the terrain passability of non-road paths is achieved, and the navigation reliability and efficiency of mobile systems on the terrain is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120167033A_ABST
    Figure CN120167033A_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for assisting navigation of a mobile system, the method comprising: obtaining (RGB) an optical image of a scene captured by a camera embedded on the mobile system; obtaining (LiDAR) 3D points of the scene acquired by a rangefinder embedded on the mobile system; projecting the 3D points and the uncertainty associated with each of the 3D points into a camera system in a 2D manner to provide a depth image and an uncertainty mask, respectively; determining (CNN1) a semantic map (MS) of the scene, a depth map (MD) of the scene and a confidence map (MT) of the depth map from the optical image, the depth image and the uncertainty mask; a trafficability map (CT) of the scene is determined (CNN2) by merging the semantic map (MS), the depth map (MD), and a confidence map (MT) of the depth map by the mobile system.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention

[0001] The field of the present invention is the field of assisting the navigation of mobile systems of the type of robots or autonomous vehicles moving on terrain, and more specifically the field of generating trajectories that can be navigated by the mobile system on the terrain. Background Art

[0002] In the field of mobile system navigation, methods are known for detecting the presence of roads in images acquired by cameras embedded in the mobile system. These methods use visual cues such as vanishing points, textures or reliefs to mark the outline of the road on the image, or directly pose the problem as a problem of segmenting the road in the image. However, these methods do not solve the path problem in the broad sense of the term "path", especially non-road paths that are not necessarily paved or properly marked, and even less solve the broader topic of navigability corresponding to identifying in the acquired images the terrain areas on which the mobile system can move.

[0003] There are also methods based on neural network models that can identify the type of ground on which a vehicle is traveling. For example, document WO 2019 / 241022 A1 describes a solution that uses a pre-trained deep neural network to detect navigable roads that are not necessarily marked by road markings. Summary of the Invention

[0004] The object of the present invention is to propose a solution that is both reliable and efficient for generating navigable trajectories for a mobile system moving on terrain.

[0005] To this end, the present invention proposes a computer-implemented method for assisting the navigation of a mobile system, the method comprising: obtaining an optical image of a scene acquired by a camera embedded on the mobile system; obtaining a 3D point cloud of the scene acquired by a rangefinder embedded on the mobile system; projecting the 3D points of the point cloud and the uncertainty associated with the measurement value of each 3D point in the 3D points of the point cloud in a 2D manner into the reference frame of the camera to respectively provide a depth image and an uncertainty mask of the depth image; determining a semantic map of the scene, a depth map of the scene and a confidence map of the depth map based on the optical image, the depth image and the uncertainty mask of the depth image; determining a navigability map of the scene by the mobile system by merging the semantic map, the depth map and the confidence map of the depth map.

[0006] Some preferred but non-limiting aspects of the method are as follows.

[0007] Determining the semantic map of the scene, the depth map of the scene, and the confidence map of the depth map includes: processing an optical image, a depth image, and an uncertainty mask of the depth image through a first convolutional neural network, the first convolutional neural network including a series of convolutional layers, each convolutional layer including a first convolutional block capable of estimating a semantic attribute map, a second convolutional block capable of estimating a depth attribute map, and a third convolutional block capable of estimating a confidence attribute map.

[0008] The second convolutional block of the (N + 1)-th convolutional layer in the series of convolutional layers is configured to: calculate the product of the confidence attribute map estimated by the third convolutional block of the N-th convolutional layer in the series of convolutional layers and the depth attribute map estimated by the second convolutional block of the N-th convolutional layer in the series of convolutional layers; calculate a first convolution result by applying a convolution kernel to the product; calculate a second convolution result by applying a convolution kernel to the confidence attribute map estimated by the third convolutional block of the N-th convolutional layer in the series of convolutional layers; and calculate the ratio of the first correlation result and the second correlation result.

[0009] The second convolutional block of the (N + 1)-th convolutional layer in the series of convolutional layers is configured to: calculate the product of the confidence attribute map estimated by the third convolutional block of the N-th convolutional layer in the series of convolutional layers and a concatenated map, the concatenated map being generated by concatenating the semantic attribute map estimated by the first convolutional block of the N-th convolutional layer in the series of convolutional layers and the depth attribute map estimated by the second convolutional block of the N-th convolutional layer in the series of convolutional layers; calculate a first convolution result by applying a convolution kernel to the product; calculate a second convolution result by applying a convolution kernel to the confidence attribute map estimated by the third convolutional block of the N-th convolutional layer in the series of convolutional layers; and calculate the ratio of the first correlation result and the second correlation result.

[0010] The second convolutional block of the (N + 1)-th convolutional layer in the series of convolutional layers is further configured to add a bias to the ratio of the first correlation result and the second correlation result.

[0011] The first convolutional block of the (N + 1)-th convolutional layer in the series of convolutional layers takes the concatenated map as input, the concatenated map being generated by concatenating the semantic attribute map estimated by the first convolutional block of the N-th convolutional layer in the series of convolutional layers and the depth attribute map estimated by the second convolutional block of the N-th convolutional layer in the series of convolutional layers.

[0012] Merging the semantic map, the depth map, and the confidence map of the depth map includes: determining a concatenated map by concatenating the semantic map, the depth map, and the confidence map of the depth map, and processing the concatenated map through a second convolutional neural network.

[0013] The present invention also relates to a computer program product comprising instructions which, when executed by a computer, cause the computer to implement the steps of the method according to the present invention. The present invention also extends to a topographic mapping device intended to be embedded in a mobile system, the topographic mapping device comprising a processor configured to implement the steps of the method according to the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Other aspects, objects, advantages and features of the present invention will become more apparent by reading the following detailed description of the preferred embodiments of the present invention given by way of non - limiting examples and with reference to the accompanying drawings, in which:

[0015] Figure 1 is a diagram showing a possible embodiment of the method according to the present invention;

[0016] Figure 2 represents operations that can be performed by the convolutional layer of a first convolutional neural network that can be used by the present invention;

[0017] Figure 3 more specifically represents operations that can be performed by the second and third convolutional blocks of the convolutional layer of a first convolutional neural network that can be used by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0018] The present invention particularly relates to a topographic mapping device intended to be embedded in a mobile system, the mobile system being, for example, an all - terrain land mobile system such as a robot, a drone or an autonomous vehicle.

[0019] The device comprises a passability estimation unit configured to generate a trajectory along which the mobile system can pass based on an image stream from a camera and depth measurements from a rangefinder.

[0020] Since generating a reliable trajectory requires perception of the geometry and semantics of the terrain, the passability estimation unit advantageously combines a geometric solution for estimating the 3D of the terrain (in this case, the depth of the terrain) with a solution for semantically segmenting the terrain. The passability estimation unit also uses a confidence map associated with the reliability of the prediction of the geometric solution, which makes it possible to greatly improve performance.

[0021] The passability estimation unit transmits a passability map (for example, a binary map) in which each point of the terrain imaged by the camera is identified as passable or non - passable for the mobile system, or the passability estimation unit transmits a map as follows, in which the probability of passability is associated with each point of the terrain.

[0022] The passability estimation unit is configured to implement the method that will be described below with reference to Figure 1 the description.

[0023] The method includes an RGB step of obtaining an optical image of a scene (in this case, the terrain on which the mobile system moves) acquired by a camera embedded in the mobile system. The camera is, for example, a monocular camera. The images continuously acquired by the camera are typically RGB images of the terrain, thus ensuring functionality under visible light. In a variant of the embodiment, night operation is ensured by exploiting another wavelength range (e.g., infrared).

[0024] The method also includes a LiDAR step of obtaining a 3D point cloud of the scene acquired by a rangefinder embedded in the mobile system. The rangefinder is, for example, a laser rangefinder, such as a Light Detection and Ranging (LiDAR). Then, the method includes the steps of projecting the 3D points of the point cloud and the uncertainties associated with the measurements of each of the 3D points of the point cloud in a 2D manner into the reference frame of the camera to respectively provide a depth image and an uncertainty mask of the depth image (i.e., a map in which the uncertainties associated with the determination of the depth are associated with each point of the terrain).

[0025] The rangefinder provides sparse depth measurements, and the sparse depth measurements are typically artificially densified by encoding the unobserved pixels. In addition, by using the power (amplitude) of the signal received by the rangefinder, the uncertainties associated with the depth measurements can be derived, where the power (amplitude) of the signal corresponds, for example, to the amount of light that returns to the sensor after transmission. In fact, the amount of light received by the sensor is directly related to the material on which the light is projected and provides information related to the reliability of the distance calculated at that point.

[0026] Then, the method includes the step of determining a semantic map MS of the scene, a depth map MD of the scene, and a confidence map MT of the depth map based on the optical image, the depth image, and the uncertainty mask of the depth image. This step is implemented, for example, by a first Convolutional Neural Network CNN1 appropriately pre-trained for this purpose.

[0027] This step simultaneously performs the derivation of the 3D (depth map) and semantics (semantic map) of the image. This enables better prediction of these two modalities with minimized computational time. In addition, this step uses the uncertainties determined a priori from the telemetry data to estimate the reliability (confidence map) of the prediction.

[0028] The method proceeds to the step of determining a traversability map CT of the scene by combining the semantic map, the depth map, and the confidence map of the depth map by the mobile system. This step is implemented, for example, by a second Convolutional Neural Network CNN2 appropriately pre-trained for this purpose. This step exploits two modalities (3D and semantic) and uses the confidence as a weight to combine these two modalities (3D and semantic).

[0029] ReferenceFigure 2 , the first convolutional neural network CNN1 includes a series of convolutional layers C N , C N+1 , and each convolutional layer may include a FMS capable of estimating the semantic attribute graph N 、FMS N+1 The first convolutional block B1 N 、B1 N+1 , can estimate the depth attribute map FMD N 、FMD N+1 The second convolution block B2 N 、B2 N+1 And can estimate the confidence attribute map FMT N 、FMT N+1 The third convolutional block B3 N 、B3 N+1 .

[0030] In one possible embodiment, the first convolution block B1 of the N+1th convolution layer in a series of convolution layers N+1 The first convolution block B1 of the Nth convolutional layer in a series of convolutional layers N Estimated Semantic Property Graph FMS N This applies to integers where N is greater than or equal to 1, and the first convolutional block of the first convolutional layer in a series of convolutional layers takes the optical image as input.

[0031] exist Figure 2 In an alternative embodiment shown, the first convolution block B1 of the N+1th convolution layer in a series of convolution layers N+1 The concatenated image is taken as input and passed through the first convolutional block B1 of the Nth convolutional layer in a series of convolutional layers. N Estimated Semantic Property Graph FMS N and the second convolutional block B2 of the Nth level convolutional layer N Estimated depth attribute map FMD N splicing (by Figure 2 This is applicable for integers N greater than or equal to 1, and the first convolutional block of the first convolutional layer in a series of convolutional layers takes the concatenation of the optical image and the depth image as input.

[0032] In this alternative embodiment, the first convolutional neural network thus comprises a first branch (a series of first convolutional blocks) that estimates the semantics of the scene by utilizing optical information from the camera and depth information from the rangefinder. This improves semantic segmentation.

[0033] In one possible implementation, the second convolution block B2 of the N+1th convolution layer in a series of convolution layers N+1The second convolutional block B2 of the N-th convolutional layer in a series of convolutional layers N estimates the depth attribute map FMD N and the third convolutional block B3 of the N-th convolutional layer in a series of convolutional layers N estimates the confidence attribute map FMS N are used as inputs. This applies to integers N greater than or equal to 1, while the second convolutional block of the first convolutional layer in a series of convolutional layers takes as inputs the depth image and the uncertainty mask of the depth map.

[0034] In Figure 2 an alternative embodiment represented by N+1 the second convolutional block B2 of the (N + 1)-th convolutional layer in a series of convolutional layers takes as input on the one hand a concatenated map produced by concatenating (identified by the label ct) the semantic attribute map FMS N estimated by the first convolutional block B1 of the N-th convolutional layer in a series of convolutional layers N and the depth attribute map FMD N estimated by the second convolutional block B2 of the N-th convolutional layer; and on the other hand takes as input the confidence attribute map FMS N estimated by the third convolutional block B3 of the N-th convolutional layer in a series of convolutional layers N This applies to integers N greater than or equal to 1, while the second convolutional block of the first convolutional layer in a series of convolutional layers takes as input on the one hand the concatenation of the optical image and the depth image, and on the other hand the uncertainty mask of the depth map. N In this alternative embodiment, the first convolutional neural network thus includes a second branch (a series of second convolutional blocks) that estimates the depth of the scene by leveraging depth information from a rangefinder and optical information from a camera. This improves semantic depth estimation.

[0035] Furthermore, in the two embodiments mentioned previously, the prior uncertainty associated with the rangefinder measurements propagates through the series of convolutional layers, which enables obtaining a confidence related to the quality and reliability of the output prediction.

[0036]

[0037] Figure 3 Figure 3

[0038] represents a possible embodiment of the operations implemented by the second and third convolutional blocks of the convolutional layers of the first convolutional neural network. In this

[0038] , · corresponds to pointwise multiplication, * corresponds to convolution, / corresponds to division, and + corresponds to addition. Γ(W) represents the kernel of the convolution.

[0038] Consider X as a tensor representing an input signal, C as a positive scalar function representing the confidence (or certainty) of each value of X, B as a tensor representing the basis of a filtering operator, B* as the conjugate of B, and A as a positive scalar function representing the applicability of each value of B. The normalized convolution can be written as follows:

[0039] Y N = N -1 D (1),

[0040] where:

[0041] D = AB*CX (2),

[0042] N = (ABB * )*C (3).

[0043] In equation (1), N is the normalization factor. For example, consider the case where the confidence C is constant and B = 1, then equation (1) becomes:

[0044]

[0045] where the convolution parameter A' is the normalized version of A.

[0046] In the framework of the present invention, the learning of the first neural network is performed as follows: Determine the parameters corresponding to the product AB to complete the task of generating a depth map from sparse input data associated with prior confidence. More specifically, set the basis B to be equal to the tensor 1 and learn the applicability function A during the network training phase.

[0047] Refer to Figure 3 , the applicability function A corresponds to the convolution parameter. Since the applicability must remain a positive function, it is necessary to ensure that the convolution weights are positive. Therefore, the softplus function Γ(.) can be applied to the weights W of the convolution. Based on equation (1), the depth propagation becomes:

[0048]

[0049] Therefore, the second convolution block B2 of the (N + 1)-th convolutional layer in a series of convolutional layers N+1 can be configured to:

[0050] compute the product of the confidence attribute map FMT estimated by the third convolution block of the N-th convolutional layer in a series of convolutional layers N with the concatenated map (through element-wise multiplication ·), the concatenated map being generated by the concatenation of the semantic attribute map FMS N estimated by the first convolution block of the N-th convolutional layer in a series of convolutional layers N and the depth attribute map FMD estimated by the second convolution block of the N-th convolutional layer in a series of convolutional layers;

[0051] The first convolution result is calculated by applying the convolution kernel Γ(W) to the product.

[0052] The second convolution result is calculated by applying the convolution kernel Γ(W) to the confidence attribute map estimated by the third convolution block of the Nth convolutional layer in a series of convolutional layers.

[0053] The ratio of the first correlation result and the second correlation result is calculated by division / .

[0054] As seen previously, in another possible implementation, the second convolution block only takes the depth attribute FMD N as input, and does not take the concatenation result of the depth attribute FMD N and the semantic attribute FMS N as input. This another possible embodiment is shown in Figure 3 , according to which the second convolution block B2 of the (N + 1)th convolutional layer in a series of convolutional layers N+1 is configured to:

[0055] Calculate the product (by element-wise multiplication ·) of the confidence attribute map FMT estimated by the third convolution block of the Nth convolutional layer in a series of convolutional layers N and the depth attribute map FMD estimated by the second convolution block of the Nth convolutional layer in a series of convolutional layers N ;

[0056] The first convolution result is calculated by applying the convolution kernel Γ(W) to the product.

[0057] The second convolution result is calculated by applying the convolution kernel Γ(W) to the confidence attribute map estimated by the third convolution block of the Nth convolutional layer in a series of convolutional layers.

[0058] The ratio of the first correlation result and the second correlation result is calculated by division / .

[0059] Furthermore, in any of the embodiments mentioned above, and as also Figure 3 represented, each second convolution block may also be configured to add the bias term BS to the ratio result of the first correlation result and the second correlation result. This bias term increases the capacity of the first neural network.

[0060] Figure 3 The third convolution block B3 N+1 is also shown. This block performs a conventional convolution for the propagation of confidence. This block may include a ReLU (Rectified Linear Unit) activation function to ensure positivity and maintain the dimension between the confidence attribute map and the depth attribute map.

[0061] Similarly, the first convolutional block for determining the semantic attribute graph can take the form of a conventional convolutional block.

[0062] A possible embodiment of the learning of the first convolutional neural network uses the following cost function to learn, so that the depth regresses and models the reciprocal of the uncertainty (i.e., confidence). Let S be a set of coordinates, where the depth value input is the ground truth, log(C (uv) ) is the predicted log-confidence, y (u,v) is the depth ground truth, is the predicted depth. The cost function can be defined as follows:

[0063]

[0064] Pen = log(C (u,v) ) (7),

[0065]

[0066] In equation (8), λ is a hyperparameter, is the regression error defined by equation (6), Pen is the penalty term defined by equation (7), which prevents the output confidence from being equal to 0. In this equation (8), the left-hand side term is the product of the regression error and the confidence. The p-norm will be replaced by the desired regression error.

[0067] Through this multiplication, the confidence serves as the weight of the regression error, thus having an overall and relative impact on the learning speed. First, overall, since the value of the average confidence decreases when λ decreases, the learning speed decreases overall. Relatively speaking, since the entropy of the confidence distribution is larger, the impact on the learning speed will vary according to the spatial position. Therefore, the choice of λ controls the entropy of the distribution and the average confidence, thus affecting the learning.

[0068] In practice, the prediction of the log-confidence can be performed to improve the learning stability. In addition, in order to keep the confidence output within the interval [0, 1] for easy interpretation of the results, activation (-1)×ReLU can be performed on the last layer to obtain the negative log-confidence, which enables the generation of the final confidence output within the interval [0, 1].

[0069] The first convolutional neural network outputs a semantic map MS, a depth map MD, and a confidence map MT of the depth map. The traversability map CT of the scene determined by the mobile system can include: determining a concatenated map by concatenating the semantic map, the depth map, and the confidence map of the depth map, and processing the concatenated map through a second convolutional neural network CNN2. This second network can be a convolutional network with a traditional architecture.

Claims

1. A computer-implemented method for assisting in the navigation of a mobile system, comprising: Obtain an optical image of a scene acquired by a camera embedded on the mobile system (RGB); Obtain a 3D point cloud of the scene acquired by a rangefinder embedded on the mobile system (LiDAR); Project the 3D points of the point cloud and the uncertainties associated with the measurements of each 3D point of the point cloud in a 2D manner into the reference frame of the camera to respectively provide a depth image and an uncertainty mask of the depth image; Determine a semantic map (MS) of the scene, a depth map (MD) of the scene, and a confidence map (MT) of the depth map based on the optical image, the depth image, and the uncertainty mask of the depth image. The determination includes processing the optical image, the depth image, and the uncertainty mask of the depth image through a first convolutional neural network (CNN1), the first convolutional neural network including a series of convolutional layers, each convolutional layer including a first convolutional block (B1 N , FMS N+1 ) capable of estimating a semantic attribute map (FMS N , B1 N+1 ), a second convolutional block (B2 N , FMD N+1 ) capable of estimating a depth attribute map (FMD N , B2 N+1 ), and a third convolutional block (B3 N , FMT N+1 ) capable of estimating a confidence attribute map (FMT N , B3 N+1 ); Determine a traversability map (CT) of the scene by merging the semantic map (MS), the depth map (MD), and the confidence map (MT) of the depth map through the mobile system.

2. The method according to claim 1, wherein, The second convolutional block (B2 N+1 ) of the (N + 1)-th convolutional layer (C N+1 ) in the series of convolutional layers is configured to: Calculate the product of the confidence attribute map (FMT N ) estimated by the third convolutional block (B3 N ) of the Nth convolutional layer (C N ) in the series of convolutional layers and the depth attribute map (FMD N ) estimated by the second convolutional block (B2 N ) of the Nth convolutional layer (C N ) in the series of convolutional layers; Calculate a first convolution result by applying a convolution kernel to the product; Calculate a second convolution result by applying the convolution kernel to the confidence attribute map estimated by the third convolution block of the N-th level convolution layer in the series of convolution layers; Calculate the ratio of the first correlation result and the second correlation result.

3. The method according to claim 1, wherein, The second convolutional block (B2 N+1 ) of the (N + 1)-th convolutional layer (C N+1 ) in the series of convolutional layers is configured to: Calculate the product of the confidence attribute map (FMT N ) estimated by the third convolution block (B3 N ) of the N-th level convolutional layer (C N ) in the series of convolutional layers and the concatenated map, where the concatenated map is produced by concatenating the semantic attribute map (FMS N ) estimated by the first convolution block (B1 N ) of the N-th level convolutional layer (C N ) in the series of convolutional layers and the depth attribute map (FMD N ) estimated by the second convolution block (B2 N ) of the N-th level convolutional layer (C N ) in the series of convolutional layers; Calculate a first convolution result by applying a convolution kernel to the product; Calculate a second convolution result by applying the convolution kernel to the confidence attribute map estimated by the third convolution block of the N-th level convolution layer in the series of convolution layers; Calculate the ratio of the first correlation result and the second correlation result.

4. The method according to claim 2 or 3, wherein, The second convolutional block (B2 N+1 ) of the (N + 1)-th convolutional layer (C N+1 ) in the series of convolutional layers is further configured to add a bias (BS) to the ratio of the first correlation result and the second correlation result.

5. The method according to claim 1, wherein, The first convolutional block (B1 N+1 ) of the (N + 1)-th convolutional layer (C N+1 ) in the series of convolutional layers takes the concatenated graph as input, and the concatenated graph is generated by concatenating the semantic attribute map (FMS N ) estimated by the first convolutional block (B1 N ) of the N-th convolutional layer (C N ) in the series of convolutional layers and the depth attribute map (FMD N ) estimated by the second convolutional block (B2 N ) of the N-th convolutional layer in the series of convolutional layers.

6. The method according to any one of claims 1 to 5, wherein, Merging the semantic map, the depth map, and the confidence map of the depth map includes: determining a stitched map by stitching the semantic map, the depth map, and the confidence map of the depth map, and processing the stitched map through a second convolutional neural network (CNN2).

7. A computer program product comprising instructions that, when executed by a computer, cause the computer to implement the steps of the method according to any one of claims 1 to 6.

8. A topographic mapping device designed to be embedded in a mobile system, comprising a processor configured to implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Path detection for autonomous machines using deep neural networks

    WO2019241022A1