Vision-based Motion Planning Method and System for Unmanned Platforms in Unstructured Environments
By using deep neural networks in the unmanned platform motion planning system to generate passability cost maps and combining MPPI algorithm for action sampling, the problem of low motion planning efficiency in unstructured environments is solved, and efficient and safe motion planning is achieved.
Patent Information
- Application Number
- CN202411744926.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-12-02
AI Technical Summary
Existing vision-based motion planning technology is difficult to effectively evaluate passability in an unstructured environment, and because path planning and motion tracking are divided into two continuous stages, the motion planning efficiency is inefficient and parallel processing cannot be achieved.
The deep neural network model is used to perform pixel-level semantic classification of RGB images in unstructured environments, generate a passability cost map, and combine the MPPI algorithm to sample the action, and determine the optimal action through cost calculation.
Fast and accurate motion planning in an unstructured environment is achieved, the motion efficiency and safety of the unmanned platform are improved, and collisions with obstacles are avoided.
Smart Images

Figure CN119229263B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and particularly to a vision-based motion planning method and system for unmanned platforms in unstructured environments. Background Art
[0002] In unstructured environments, vision-based local motion planning exhibits significant application potential. This technology can perform adaptive adjustment and dynamic path tracking by leveraging real-time vision data, thereby achieving effective navigation and motion decision-making in complex and unpredictable environments.
[0003] However, current vision-based motion planning technologies still face two major challenges. First, their generalization ability for assessing passability in unstructured scenarios is limited. Second, existing motion planning methods are typically divided into two consecutive stages: "path planning" and "motion tracking". This sequential execution mode limits the efficiency of motion planning and cannot achieve parallel processing. Summary of the Invention
[0004] The objective of the present application is to provide a vision-based motion planning method and system for unmanned platforms in unstructured environments, which can achieve more efficient motion planning and control in unstructured environments.
[0005] To achieve the above objective, the present application provides the following solutions:
[0006] In a first aspect, the present application provides a vision-based motion planning method for unmanned platforms in unstructured environments, including:
[0007] Obtain a target RGB image; the target RGB image is an RGB image in an unstructured scenario.
[0008] Input the target RGB image into a trained passability cost map model to obtain a passability cost map with the same height and width as the target RGB image; the trained passability cost map model is a deep neural network model trained with sample RGB images in unstructured environments as inputs and pixel-level semantic classification annotations corresponding to the sample RGB images as labels.
[0009] According to the passability cost map, use the MPPI algorithm to perform action sampling in a set motion space to obtain a sampling trajectory.
[0010] Calculate the cost of the sampling trajectory with the passability of the unstructured environment and the distance to the end point as indicators, and determine the optimal action for passing through the unstructured environment based on minimizing the cost.
[0011] Optionally, the semantics in the pixel-level semantic classification annotation include: concrete road, asphalt road, gravel road, grassland, dirt road, sand road, rock, bedrock, puddle, background, signboard, obstacle, tree, and building.
[0012] Optionally, the passability cost values corresponding to different semantics specifically include:
[0013] The passability cost values of concrete roads and asphalt roads are 0; the passability cost values of gravel roads, grasslands, dirt roads, and sand roads are 0.25; the passability cost values of rocks and bedrocks are 0.5; the passability cost value of a puddle is 0.75; the passability cost values of the background, signboard, obstacle, tree, and building are 1.
[0014] Optionally, the loss function used for training the passability cost map model is the L1 loss function; the formula expression of the L1 loss function is:
[0015] .
[0016] Among them, n is the number of training samples, y i is the true passability cost value, is the predicted passability cost value.
[0017] Optionally, before performing action sampling in the set motion space according to the passability cost map using the MPPI algorithm to obtain the sampling trajectory, it further includes:
[0018] Construct a vehicle dynamics model; the formula expression of the vehicle dynamics model is:
[0019] .
[0020] Among them, are respectively the t position and yaw angle of the vehicle at x time, y is the position of the vehicle at time t + 1, x is the position of the vehicle at time t + 1, y is the yaw angle of the vehicle at time t + 1 in the x direction, is the t yaw angle of the vehicle at x time in the v t , ω t are respectively the linear velocity and angular velocity, To control the time interval.
[0021] Optionally, according to the passability cost map, using the MPPI algorithm, action sampling is performed in the set motion space to obtain a sampling trajectory, specifically including:
[0022] According to the passability cost map, a strategy based on the MPPI algorithm is adopted, and the action sampling process in the motion space is optimized by combining the Numba library and Nvidia GPU acceleration technology.
[0023] Optionally, before calculating the cost of the sampling trajectory with the passability of the unstructured environment and the distance to the end point as indicators and determining the optimal action for passing through the unstructured environment based on minimizing the cost, it further includes:
[0024] Using known camera parameters to convert the point (x, y) in the real-world coordinates to the image plane, specifically including:
[0025] Determine the intrinsic parameters of the camera as K ; where f x and f y are respectively the focal lengths in the x and y directions, and ( c x , c y ) is the intersection point of the optical axis and the image plane.
[0026] According to the formula , project the point onto the image coordinates; where h is the height from the ground to the camera.
[0027] According to the formula , convert the image coordinates to pixel coordinates; where ([[]] u p , v p ) are the pixel coordinates on the image plane.
[0028] Optionally, the cost function for calculating the cost of the sampling trajectory includes a running cost function and a terminal cost function .
[0029] Optionally, the formula expression of the running cost function is: .
[0030] The formula expression of the terminal cost function is: .
[0031] Wherein, is the end target position; p t is the position of the vehicle at time t ; s default is the default speed used to estimate the time it takes for the vehicle to reach the end point; is an indicator function that returns 1 when any state x τ (0 ≤ τ ≤ T ) reaches , and returns 0 otherwise; is the traversability cost at position p t ; w dist > 0 and w tra > 0 are the weights for penalizing the distance to the end point and the traversability cost, respectively.
[0032] In a second aspect, the present application provides a vision-based unmanned platform motion planning system in an unstructured environment, including:
[0033] An image acquisition module for acquiring a target RGB image; the target RGB image is an RGB image in an unstructured scene.
[0034] An image conversion module for inputting the target RGB image into a trained traversability cost map model to obtain a traversability cost map having the same height and width as the target RGB image; the trained traversability cost map model is a deep neural network model trained with sample RGB images in an unstructured environment as input and pixel-level semantic classification annotations corresponding to the sample RGB images as labels.
[0035] A trajectory sampling module for performing action sampling in a set motion space according to the traversability cost map using the MPPI algorithm to obtain a sampled trajectory.
[0036] A calculation module for calculating the cost of the sampled trajectory with the traversability of the unstructured environment and the distance to the end point as metrics, and determining the optimal action for passing through the unstructured environment based on minimizing the cost.
[0037] According to the specific embodiments provided by the present application, the following technical effects are disclosed:
[0038] The present application provides a vision-based motion planning method and system for an unmanned platform in an unstructured environment. By inputting the target RGB image into a trained deep neural network model, a passability cost map can be quickly obtained, thereby realizing real-time motion planning. The deep neural network model trained by pixel-level semantic classification annotation can accurately identify obstacles and feasible regions in the unstructured environment. By using the MPPI algorithm for action sampling and combining cost calculation, an optimized trajectory with the highest passability and the shortest distance to the end point in the unstructured environment can be found, improving the motion efficiency of the unmanned platform. By calculating the cost of the sampled trajectories and determining the optimal action based on minimizing the cost, collisions between the unmanned platform and obstacles during motion can be effectively avoided, improving motion safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0040] Figure 1 It is a schematic flowchart of a vision-based motion planning method for an unmanned platform in an unstructured environment provided by an embodiment of the present application.
[0041] Figure 2 It is a schematic structural diagram of a vision-based motion planning system for an unmanned platform in an unstructured environment provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0043] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0044] Embodiment 1
[0045] As Figure 1 shown, this embodiment provides a vision-based motion planning method for an unmanned platform in an unstructured environment, including:
[0046] Step 101: Obtain a target RGB image; the target RGB image is an RGB image in an unstructured scenario.
[0047] Step 102: Input the target RGB image into a trained passability cost map model to obtain a passability cost map with the same height and width as the target RGB image; the trained passability cost map model is a deep neural network model trained with sample RGB images in an unstructured environment as input and pixel-level semantic classification annotations corresponding to the sample RGB images as labels.
[0048] Step 103: According to the passability cost map, use the MPPI algorithm to perform action sampling in a set motion space to obtain a sampling trajectory.
[0049] Step 104: Take the passability of the unstructured environment and the distance to the end point as metrics, calculate the cost of the sampling trajectory, and determine the optimal action to pass through the unstructured environment based on minimizing the cost.
[0050] Among them, when performing steps 101 - 102, it can be specifically as follows:
[0051] Obtain a target RGB image, which corresponds to an RGB image in an unstructured scenario; input the target RGB image into a trained passability cost map model to generate a passability cost map with the same size as the target RGB image.
[0052] Among them, the training of the trained passability cost map model is as follows:
[0053] Passability estimation takes an RGB image as the original input (assuming the shape of the image is 3×H×W, where 3 are the three channels of the R, G, and B primary colors of the image, and H and W are the height and width of the image respectively), uses a deep neural network as the image feature extractor (such as the ResUNet framework can be used), and finally the output of the neural network is a passability cost map with the same height and width as the RGB image (that is, the output shape is 1×H×W, corresponding to each pixel of the image, there is a corresponding passability cost value). To implement the training of the above network, a training set and an objective function need to be given.
[0054] The training set can be obtained by improving the existing unstructured environment semantic segmentation dataset (such as the RUGD dataset). In the RUGD dataset, it contains a series of images in an unstructured environment and the corresponding pixel-level semantic classification annotations. Define the passability cost corresponding to different semantics (the higher the cost, the more difficult it is for the ground unmanned platform to pass through this area), so as to realize the conversion of the dataset from a classification problem to a regression problem.
[0055] Specifically, the semantics in the pixel-level semantic classification annotation include: concrete road, asphalt road, gravel road, grassland, dirt road, sand road, rock, bedrock, puddle, background, signboard, obstacle, tree, and building. The passability cost values corresponding to different semantics specifically include:
[0056] The passability cost values of concrete roads and asphalt roads are 0; the passability cost values of gravel roads, grasslands, dirt roads, and sand roads are 0.25; the passability cost values of rocks and bedrocks are 0.5; the passability cost value of puddles is 0.75; the passability cost values of background, signboards, obstacles, trees, and buildings are 1. Specifically as shown in Table 1 below:
[0057] Table 1 Passability Costs Corresponding to Different Semantics
[0058]
[0059] It should be noted that setting the passability cost of the above obstacles / background to 1 is to make it close to the passability costs of other categories for convenient network training. However, in the subsequent motion control cost function, the cost of obstacles / background will be set to 1000 to prevent ground unmanned vehicles from entering these areas.
[0060] For the training function of the network, the L1 loss function is adopted here, and its definition is:
[0061] .
[0062] where n is the number of training samples, y i is the true passability cost value, is the predicted passability cost value.
[0063] In some embodiments, before performing step 103, it further includes:
[0064] Constructing a vehicle dynamics model; the formula expression of the vehicle dynamics model is:
[0065] .
[0066] where, are respectively the x , y position and yaw angle of the vehicle at time t, is the x position of the vehicle at time t + 1, is the y position of the vehicle at time t + 1, is the yaw angle of the vehicle at time t + 1 in the x direction, is the vehicle t at time xYaw angle in the direction, v t , ω t which are the linear velocity and angular velocity respectively, is the control time interval.
[0067] In some embodiments, when performing step 203, it specifically includes:
[0068] According to the passability cost map, adopt a strategy based on the MPPI algorithm, and combine the Numba library with the Nvidia GPU acceleration technology to optimize the action sampling process in the motion space.
[0069] Specifically, the basic process of the MPPI algorithm is as follows: Assume that the system dynamics is described by the following continuous-time state-space equation:
[0070] .
[0071] Where x(t) is the state vector, u(t) is the control input vector, f is the non-linear function of the system dynamics, w(t) is the process noise following a Gaussian distribution with mean 0 and variance Q.
[0072] The goal of MPPI is to minimize the cost function within the finite prediction horizon T, and the cost function is defined as:
[0073] .
[0074] Where is a trajectory containing state and control input pairs, is the running cost function used to evaluate the instantaneous cost at time t . is the terminal state cost function that evaluates the cost of the final state .
[0075] To minimize the cost function, MPPI samples control trajectories from the control sequence of the noise perturbation distribution. Specifically, N control trajectories are generated by adding Gaussian noise:
[0076] .
[0077] Where is Gaussian noise with mean 0 and variance ∑ , and is independently sampled for each trajectory and each time step.
[0078] For each sampled trajectory , simulate the system dynamics and calculate the cost function , and then, based on the cost function, use the Boltzmann distribution to calculate the weight of each trajectory:
[0079] .
[0080] where w i is the weight of the i-th trajectory. > 0 is the temperature parameter that controls the sensitivity of the weight to the cost.
[0081] The optimal control quantity is obtained by weighted averaging the sampled control trajectories:
[0082] .
[0083] Therefore, MPPI is a control method with a rolling time domain.
[0084] Among them, in the MPPI algorithm, it is necessary to sample the control sequence to obtain the optimal control with the goal of minimizing the cost function. Due to the independence of each control sequence sampling, this embodiment uses the Numba library combined with the Nvidia GPU to accelerate the sampling process. Numba is an immediate compiler for Python, which converts a subset of Python and NumPy code into efficient machine code at runtime, thus achieving a significant performance improvement. By utilizing the computing and parallel processing capabilities of the Nvidia GPU, the efficiency and speed of the MPPI algorithm are improved. The Numba and Nvidia GPU parallel acceleration technologies provide a fast solution for sampling a large number of control sequences and calculating the corresponding costs in the MPPI algorithm, enabling more efficient motion planning and control in complex, dynamic, and unstructured environments.
[0085] At each time step, the algorithm samples the control trajectory with Gaussian noise based on the current state, then simulates the system dynamics of each sampled trajectory, calculates the corresponding cost function, finally calculates the weights based on these cost functions, obtains the optimal control sequence using the weighted average, and applies the first control input to the system.
[0086] In some embodiments, before calculating the cost of the sampled trajectory with the passability of the unstructured environment and the distance to the end point as metrics and determining the optimal action for passing through the unstructured environment based on minimizing the cost, it further includes:
[0087] To derive the passability cost at the position , it is necessary to use the known camera parameters to convert the point ( x , y ) in the real-world coordinates to the image plane, specifically including:
[0088] Determine the intrinsic parameters of the camera asK ; wherein, f x and f y are respectively x and y the focal lengths in the c x , c y ) is the intersection point of the optical axis and the image plane.
[0089] According to the formula , project the point onto the image coordinates; wherein, h is the height from the ground to the camera.
[0090] According to the formula , convert the image coordinates to pixel coordinates; wherein, ([[]]END]] u p , v p ) are the pixel coordinates on the image plane, u , v , w are homogeneous coordinates.
[0091] After converting the points in the real world to the image plane, the corresponding passability cost can be obtained according to the passability cost map .
[0092] In some embodiments, when performing step 104, it can be specifically as follows:
[0093] Determine the cost function for calculating the cost of the sampling trajectory, and the cost function includes the running cost function and the terminal state cost function .
[0094] Wherein, the formula expression of the running cost function is: .
[0095] The formula expression of the terminal state cost function is: .
[0096] In the formula, is the end target position. p t is the position of the vehicle at time t . s default is the default speed used to estimate the time for the vehicle to reach the end point. is an indicator function, when any state x τ (0 ≤ τ ≤ T) reaches When it is [condition], it returns 1, otherwise it returns 0. is the position p t the traversability cost at the [position]. w dist > 0 and w tra > 0 are respectively the weights for penalizing the distance to the end point and the traversability cost.
[0097] Embodiment 2
[0098] As Figure 2 shown, this embodiment provides a vision-based unmanned platform motion planning system in an unstructured environment, including:
[0099] An image acquisition module 201, configured to acquire a target RGB image; the target RGB image is an RGB image in an unstructured scene.
[0100] An image conversion module 202, configured to input the target RGB image into a trained traversability cost map model to obtain a traversability cost map having the same height and width as the target RGB image; the trained traversability cost map model is a deep neural network model trained with a sample RGB image in an unstructured environment as the input and the pixel-level semantic classification annotation corresponding to the sample RGB image as the label.
[0101] A trajectory sampling module 203, configured to perform action sampling in a set motion space according to the traversability cost map by using the MPPI algorithm to obtain a sampling trajectory.
[0102] A calculation module 204, configured to calculate the cost of the sampling trajectory by using the traversability of the unstructured environment and the distance to the end point as indicators, and determine the optimal action for passing through the unstructured environment based on minimizing the cost.
[0103] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0104] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A vision-based unmanned platform motion planning method in an unstructured environment, characterized in that: The unmanned platform motion planning method comprises: Acquire a target RGB image; the target RGB image is an RGB image in an unstructured scene; The target RGB image is input into a trained passability cost map model to obtain a passability cost map having the same height and width as the target RGB image; the trained passability cost map model is a deep neural network model trained by taking a sample RGB image in an unstructured environment as input and using pixel-level semantic classification annotations corresponding to the sample RGB image as labels; According to the passability cost map, using the MPPI algorithm, action sampling is performed in the set motion space to obtain a sampling trajectory; Taking the passability of the unstructured environment and the distance to the end point as indicators, the cost of the sampled trajectory is calculated, and based on minimizing the cost, the optimal action to pass through the unstructured environment is determined; Before calculating the cost of the sampled trajectory based on the passability of the unstructured environment and the distance to the end point as indicators and determining the optimal action for passing through the unstructured environment based on minimizing the cost, the method further includes: Transform a point (x, y) in real-world coordinates to the image plane using known camera parameters, including: Determine the intrinsic parameters of the camera as K ;in, f x and f y They are x and y The focal length in the direction, ( c x , c y ) is the intersection of the optical axis and the image plane; According to the formula , project the point onto the image coordinates; where, h is the height from the ground to the camera; According to the formula , convert image coordinates into pixel coordinates; where ( u p , v p ) is the pixel coordinate on the image plane, u , v , w are homogeneous coordinates; The cost function for calculating the cost of the sampled trajectory includes running the cost function and the final cost function ; The running cost function The formula expression is: ; The final state cost function The formula expression is: ; in, is the final target position; p t For vehicles at time t location; s default is the default speed used to estimate the time it takes for the vehicle to reach the destination; is an indicator function. When any state x τ achieve It returns 1 when , otherwise it returns 0; For location p t The cost of accessibility; w dist >0 and w tra >0 are the weights of penalty distance to the end point and passability cost; Among them, x T is the final state, 0≤τ≤T.
2. According to the vision-based unmanned platform motion planning method in an unstructured environment according to claim 1, it is characterized in that: The semantics in the pixel-level semantic classification annotation include: concrete road, asphalt road, gravel road, grass, dirt road, sand road, rock, rock bed, puddle, background, signboard, obstacle, tree and building.
3. The method for motion planning of an unmanned platform based on vision in an unstructured environment according to claim 2, characterized in that: The passability cost values corresponding to different semantics include: The passability cost of concrete roads and asphalt roads is 0; the passability cost of gravel roads, grass, dirt roads, and sandy roads is 0.25; the passability cost of rocks and rock beds is 0.5; the passability cost of puddles is 0.75; the passability cost of backgrounds, signs, obstacles, trees, and buildings is 1.
4. The method for motion planning of an unmanned platform based on vision in an unstructured environment according to claim 1, characterized in that: The loss function used in the training of the passability cost graph model is L1 Loss function; L1 The formula of the loss function is: ; in, n is the number of training samples, y i is the true passability proxy value, is the predicted passability cost.
5. The vision-based unmanned platform motion planning method in an unstructured environment according to claim 1, characterized in that: Before sampling actions in a set motion space using the MPPI algorithm according to the passability cost map to obtain a sampling trajectory, the method further includes: Construct a vehicle dynamics model; the formula expression of the vehicle dynamics model is: ; in, For vehicles t Moment x , y Position and yaw angle, is the x position of the vehicle at time t+1, is the vehicle at time t+1 y Location, The time of vehicle t+1 x The yaw angle of the direction, For vehicles t time x The yaw angle of the direction, v t ,ω t are the linear velocity and angular velocity respectively, To control the time interval.
6. The vision-based unmanned platform motion planning method in an unstructured environment according to claim 1, characterized in that: According to the passability cost map, the MPPI algorithm is used to perform action sampling in the set motion space to obtain a sampling trajectory, which specifically includes: According to the passability cost map, a strategy based on the MPPI algorithm is adopted to optimize the action sampling process in the motion space by combining the Numba library with the Nvidia GPU acceleration technology.
7. A vision-based unmanned platform motion planning system in an unstructured environment, used to implement the unmanned platform motion planning method according to any one of claims 1 to 6, characterized in that: include: An image acquisition module is used to acquire a target RGB image; The target RGB image is an RGB image in an unstructured scene; An image conversion module is used to input the target RGB image into a trained passability cost map model to obtain a passability cost map having the same height and width as the target RGB image; the trained passability cost map model is a deep neural network model trained by taking a sample RGB image in an unstructured environment as input and using pixel-level semantic classification annotations corresponding to the sample RGB image as labels; A trajectory sampling module, used to perform action sampling in a set motion space according to the passability cost map and use the MPPI algorithm to obtain a sampling trajectory; The calculation module is used to calculate the cost of the sampled trajectory based on the passability of the unstructured environment and the distance to the end point as indicators, and determine the optimal action to pass through the unstructured environment based on minimizing the cost.
Citation Information
Patent Citations
Method and system for navigating with drivable area detection, and storage medium
CN115729228A