Unmanned vehicle obstacle sensing method and system for complex cross-country environment
By integrating vehicle-ground coupled dynamics modeling and global-local path planning strategies, and utilizing deep learning and Bayesian networks for mobility quantification and obstacle recognition, the problem of dynamic accuracy and obstacle recognition for unmanned off-road vehicles in complex environments was solved, achieving more efficient path planning and decision-making and improving task completion rate.
Patent Information
- Application Number
- CN202511152717.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies for dynamic modeling of unmanned off-road vehicles suffer from insufficient accuracy, difficulty in real-time identification of non-standard obstacles, and lack of effective risk assessment. This results in insufficient intelligence, reliability, and real-time performance of planning and decision-making mechanisms, making it impossible to effectively cope with the challenges of complex off-road environments.
The system integrates vehicle-ground coupled dynamics modeling, non-standard obstacle perception and recognition, and global-local joint path planning strategies. It quantifies mobility indicators through deep learning and Bayesian neural networks, combines a binocular camera stereo vision system for obstacle recognition, and employs maximum entropy deep reinforcement learning and sequential convex optimization for path planning.
It significantly improves the accuracy and adaptability of dynamic modeling, enables real-time identification and risk assessment of non-standard obstacles, enhances the intelligence and real-time nature of planning and decision-making, and improves the survivability and mission completion rate of unmanned off-road vehicles in complex environments.
Smart Images

Figure CN120871871A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent vehicle technology, specifically relating to an obstacle perception method and system for unmanned vehicles in complex off-road environments. Background Technology
[0002] With the rapid development of intelligent unmanned system technology, the application potential of unmanned off-road vehicles in complex field environments such as material transportation and disaster relief is becoming increasingly prominent. Compared with structured roads or urban environments, off-road vehicles face more complex and unpredictable operating and driving environments, often accompanied by challenges such as sudden terrain changes, loose soil, dense vegetation, and water disturbances. These severe factors not only significantly reduce the accuracy of environmental modeling and perception, but also significantly raise the technical threshold for autonomous operation of unmanned vehicles.
[0003] Existing research on vehicle dynamics modeling typically employs two main approaches: simplified physical mechanism models and high-fidelity finite element models or multibody dynamics simulations. While simplified physical mechanism models excel in computational efficiency and can quickly yield results, their predictive accuracy is significantly limited when dealing with complex vehicle-ground interaction problems. These models often describe vehicle motion by simplifying mechanical principles, making it difficult to accurately capture the effects of ground irregularities, soil property variations, and complex nonlinear interactions between tires and the ground. For example, patent CN118114354A, "Modeling Method, Decision Planning Method, System, and Vehicle for Single-Coupled Dynamics Models," aims to significantly reduce computational complexity, improve real-time performance, and enhance applicability under high-speed, low-speed, and large-angle conditions by constructing a "single-coupled dynamics model," while maintaining model accuracy. However, this simplification strategy may still sacrifice some model accuracy under certain extreme or complex conditions.
[0004] In contrast, high-fidelity finite element modeling (FEM) or multibody dynamics simulation (MBD) can more accurately characterize vehicle dynamic response characteristics, such as suspension deformation, vehicle attitude changes, and stress distribution on the tire-ground contact surface. However, their main drawbacks are high computational costs and long modeling cycles. Building a high-precision finite element model requires significant time for mesh generation, material parameter setting, and boundary condition definition, while multibody dynamics simulation also requires detailed component modeling and constraint definition. This makes them difficult to meet real-time requirements, especially in off-road scenarios requiring rapid decision-making and planning. Furthermore, these models are typically built for specific vehicles and terrains, resulting in poor generalization ability and difficulty in directly applying them to different vehicle models or varying terrain environments. For example, patent CN116415352A, "A Method for Multibody Dynamics Modeling and Braking Performance Analysis of Special Vehicles," employs multibody dynamics modeling for special vehicles, enabling a more detailed description of the interactions between vehicle components, thereby improving the accuracy of braking performance analysis. However, multibody dynamics models are usually time-consuming to model and computationally intensive, which may affect their real-time application efficiency; moreover, their parameters are difficult to obtain and calibrate, and require a large amount of experimental data for verification.
[0005] Furthermore, while traditional visual semantic segmentation methods can classify ground types (e.g., identifying roads, grass, water surfaces, etc.), their main limitation lies in their inability to quantitatively assess the ground's support force or traction performance. This means that even if the system identifies "mud," it cannot determine the specific load-bearing capacity of that mud, whether the vehicle can easily pass through or risks getting stuck, nor can it assess the tire's grip on different surfaces. This qualitative rather than quantitative assessment method prevents vehicles from performing refined motion planning and control based on ground characteristics. Taking patent CN114549542A, "Visual Semantic Segmentation Method, Apparatus, and Device," as an example, it performs visual semantic segmentation by fusing single-frame images and multi-frame temporal point cloud data. Utilizing complementary information from multi-source data, it aims to improve the accuracy and robustness of semantic segmentation and obtain dense environmental point cloud information containing static semantics, providing richer spatial context. Although this method is innovative in improving the accuracy of semantic segmentation, its core still lies in the classification of scenes and does not delve into the quantitative assessment of ground physical properties (such as load-bearing capacity and friction coefficient). This makes it impossible for vehicles to carry out refined and risk-avoidance motion planning based on the actual physical characteristics of the ground.
[0006] Traditional image recognition methods that rely on geometric information also face significant challenges in identifying non-geometric obstacles. These obstacles include mud, soft areas, and terrain prone to vehicle slippage, all characterized by blurred boundaries and a lack of clear structural features. For example, a muddy area may lack clear boundaries, and its depth and viscosity may vary depending on location. Traditional methods typically rely on geometric features such as edge detection and shape recognition, making it difficult to accurately detect and identify these non-geometric obstacles, thus posing potential dangers to vehicles during off-road driving.
[0007] Existing path planning and decision-making methods are mostly based on ideal scenario settings, generally ignoring the impact of terrain uncertainty on vehicle performance. They typically assume the ground is uniform and has sufficient load-bearing capacity, failing to consider the potential performance degradation or risks caused by the complex and varied terrain characteristics (such as slope, soil type, obstacle distribution, etc.) in real off-road environments. More importantly, these methods lack a joint feedback mechanism for maneuverability uncertainty assessment and obstacle perception information. This means that when planning a path, the system does not fully consider the vehicle's own maneuverability in the current terrain (e.g., whether it can climb slopes, whether it will slip), nor does it promptly feed back real-time ground perception information (such as the discovery of muddy areas) into the path planning. This information gap can easily lead to off-road vehicles losing maneuverability due to getting stuck or slipping during mission execution, thereby interrupting the mission or even causing vehicle damage.
[0008] Finally, directly solving local path planning optimization problems with nonlinear constraints is computationally complex, making them unsuitable for off-road scenarios with high real-time requirements. Local path planning needs to generate the optimal path based on the current vehicle state and environmental information within a short time. However, optimization problems involving nonlinear constraints (such as vehicle dynamics constraints, terrain adaptability constraints, and obstacle avoidance constraints) typically require iterative solutions, resulting in enormous computational demands and making it difficult to provide solutions under millisecond-level real-time requirements. This limits the ability of off-road vehicles to make rapid and robust path adjustments and decisions in complex and dynamic environments. Summary of the Invention
[0009] This invention aims to overcome the shortcomings and limitations of existing technologies, such as insufficient accuracy and adaptability in dynamic modeling, difficulty in real-time identification of non-standard obstacles, lack of effective risk assessment, and insufficient intelligence, reliability, and real-time performance of planning and decision-making mechanisms. The core solution proposed in this invention is to integrate vehicle-ground coupled dynamic modeling, non-standard obstacle perception and identification, and a global-local joint path planning strategy, forming an integrated closed-loop mechanism of "perception-prediction-planning." Through this invention, the accuracy and adaptability of dynamic modeling can be significantly improved, real-time identification and risk assessment of non-standard obstacles can be effectively achieved, and the intelligence, reliability, and real-time performance of planning and decision-making mechanisms can be enhanced, thereby systematically improving the survivability and mission completion rate of off-road vehicles in complex unstructured environments.
[0010] To achieve the above objectives, the present invention provides the following solution: an obstacle perception method for unmanned vehicles in complex off-road environments, comprising the following steps:
[0011] S1. Collect terrain parameters and vehicle control parameters to construct a deep network model; the terrain parameters include: geographic elevation and geological information; the vehicle control parameters include: speed and status;
[0012] S2. Based on the terrain parameters, the vehicle control parameters, and the deep network model, obtain the vehicle's mobility index; and quantify the uncertainty and risk of the mobility index; the mobility index includes:
[0013]
[0014] In the formula, The model predicts the mobility index, h(x,y) is the elevation variation function; s represents soil physical properties; u represents the input control variable; f DNN This is a calculation function for the mobility index fitted by a neural network.
[0015] S3. Acquire image data based on a binocular camera stereo vision system, process the image data, and assess the passability risk probability based on the processed image data to complete the non-standard obstacle identification.
[0016] S4. Perform global path planning and local dynamic planning for the off-road vehicle; the global path planning includes: constructing a path and assigning tasks on a coarse-grained topographic map based on the principle of minimizing uncertainty of the mobility index; the local dynamic planning includes: avoiding and adjusting risk points in the path based on the results of non-standard obstacle identification and dynamic environmental perception data.
[0017] More preferably, the method for quantifying the uncertainty of the mobility index includes:
[0018] A Bayesian neural network is used to model the probability distribution of mobility prediction, with the modeling objective being: in, This represents the training sample data, p() is the Monte Carlo simulation process; the output is the posterior distribution of the vehicle's mobility index under specific terrain and conditions;
[0019] Obtain the mean and variance of the maneuverability forecast, and then complete the uncertainty quantification;
[0020] The methods for quantifying the risk of the aforementioned mobility index include:
[0021] The elevation variation function and soil physical properties were input into the BNN model for several Monte Carlo simulations to obtain the prediction results:
[0022]
[0023] In the formula, f BNN This represents a network formed by a Bayesian neural network; N represents the number of Monte Carlo simulations. Indicates the expected speed; h i s represents the elevation change function for the i-th sampling; i This represents the physical properties of the soil sampled in the i-th sampling.
[0024] Calculate the probability of an off-road vehicle losing maneuverability:
[0025]
[0026] In the formula, P risk This indicates the probability that the road surface ahead will affect the vehicle's maneuverability; v th Indicates the mobility threshold; This indicates an indicator function.
[0027] More preferably, the method for non-standard obstacle identification includes:
[0028] The image data is subjected to depth map extraction and semantic segmentation to obtain depth information and semantic segmentation information of pixels; the method for performing semantic segmentation on the image data includes: extracting visual features from the image data and mapping the visual features to soil physical properties;
[0029] Spatial alignment and geological data matching are performed on the depth information to obtain the geometric embedding features of the pixels;
[0030] The terrain category corresponding to each pixel is obtained based on the geometric embedding features and the semantic segmentation information;
[0031] The visual features are modeled using a Gaussian mixture model to obtain the training distribution;
[0032] Calculate the log-likelihood of the training distribution to determine whether the terrain category belongs to the known terrain;
[0033] Based on the soil physical properties extracted after image segmentation and the uncertainty distribution, the prediction results of the mobility index are obtained.
[0034] Based on the prediction results, the passability risk probability is calculated, and non-standard obstacle identification is completed based on the passability risk probability.
[0035] More preferably, the method for obtaining the prediction results of the mobility index based on the soil physical properties s extracted after image segmentation and the uncertainty distribution δs of the soil physical properties includes:
[0036] v pred ~p(v|s,h(x,y),δs);
[0037] In the formula, v pred f BNN The predicted velocity sampled in the Monte Carlo simulation; p() represents the Monte Carlo simulation process;
[0038] The method for calculating the probability of passability risk based on the prediction results includes:
[0039] P fail =P(v pred <v th );
[0040] In the formula, P fail P(v) represents the probability of success or failure; pred <v th ) represents the probability that the predicted speed is less than the maneuverability threshold.
[0041] More preferably, the global path planning employs a maximum entropy deep learning reinforcement algorithm for path planning; the method includes:
[0042] Transform the global path planning problem into a Markov decision process:
[0043]
[0044] In the formula, For state space; Let be the action space; P be the state transition function; r(s,a) be the reward function; and γ be the discount factor.
[0045] Employing a maximum entropy deep learning reinforcement algorithm to maximize the expected return of entropy regularization:
[0046]
[0047] In the formula, π * π represents the optimal policy; π represents the policy function. R represents the expected value of the state-action trajectory distribution under policy π; t represents time; α is the entropy weighting coefficient, and r(s) represents the expected value of the state-action trajectory distribution under policy π. t ,a t ) indicates that at time step t, state s t Next, execute action a t The instant rewards received; The strategy distribution entropy;
[0048] The reward function embeds the vehicle mobility risk index R. mob :
[0049] r(s,a)=w1·r goal +w2·r smooth -w3·R mob (s);
[0050] Where: r goal To incentivize off-road vehicles to approach the target area; r smooth For path smoothness constraints; w i These are weighting coefficients;
[0051] Policy iteration using soft Q-function loss:
[0052]
[0053] in,
[0054]
[0055] In the formula, This represents the loss value used for model training; represents the expected value; Q(s,a) represents the current Q-function's estimate of the state-action pair; Indicates the target Q value; denoted by ; Q(s′,a′) represents the Q-value of Q-learning reinforcement learning; a′ and s′ represent the action value under s′ and the next state reached after executing the action, respectively; logπ(a′|s′) represents the logarithmic probability value of policy π.
[0056] More preferably, the local dynamic programming is modeled as a constrained optimization problem:
[0057]
[0058] The constraints are:
[0059] Dynamic model: x {t+1} =f(x) t ,ut );
[0060] Control input restrictions:
[0061] Mobility limit: v(x) t )≥v min ;
[0062] Obstacle avoidance restrictions:
[0063] In the formula: Denotes the control sequence; λ represents the risk penalty coefficient; x t Vehicle status; x represents the desired state of the reference trajectory at time t; {t+1} Indicates the vehicle state at time t+1; u t For control input; High-risk areas identified by non-standard barriers; This represents the vehicle's control command set; The set of obstacles perceived at the current time; v min The minimum tolerable driving speed; v(x) t ) indicates that the car is at x t The velocity under the given state; f(·) represents the vehicle dynamics model.
[0064] More preferably, the optimization method of the local dynamic programming includes:
[0065] Step 1: Using the current optimal trajectory Using the expansion point as the Taylor first-order approximation for all nonlinear constraints;
[0066] Step 2: Construct a new subproblem with linear constraints / objectives, and convert it into a standard convex quadratic programming format;
[0067] Step 3: Solve for the new trajectory Determine if convergence has occurred; if not, iterate and update to form a closed loop. Here, δ represents the direction of the variable increment, which comes from the solution of the linear subproblem.
[0068] The present invention also provides an obstacle perception system for unmanned vehicles in complex off-road environments, comprising:
[0069] The parameter acquisition module is used to collect terrain parameters and vehicle control parameters to construct a deep network model; the terrain parameters include geographic elevation and geological information; the vehicle control parameters include speed and status.
[0070] The dynamics modeling module is used to obtain the vehicle's maneuverability index based on the terrain parameters, the vehicle control parameters, and the deep network model; and to quantify the uncertainty and risk of the maneuverability index.
[0071] The non-standard obstacle recognition module is used to acquire image data based on a binocular camera stereo vision system, process the image data, and perform a passability risk probability assessment based on the processed image data, thereby completing the non-standard obstacle recognition.
[0072] The global task planning and local dynamic decision-making module is used to perform global path planning and local dynamic planning for the off-road vehicle. The global path planning includes: constructing a path and allocating tasks on a coarse-grained topographic map based on the principle of minimizing uncertainty of the mobility index. The local dynamic planning includes: avoiding and adjusting risk points in the path based on the results of non-standard obstacle identification and dynamic environmental perception data.
[0073] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0074] This invention integrates vehicle-ground coupled dynamics modeling, non-standard obstacle perception and recognition, and a global-local joint path planning strategy, compared with existing technologies:
[0075] 1. Significantly improves the accuracy and adaptability of dynamic modeling: Compared with traditional dynamic models based on physical rules, the data-driven modeling method proposed in this invention can adapt to various off-road terrain media. By introducing parameterized terrain features and control inputs, it constructs a coupled nonlinear dynamic model that is more in line with real operating conditions, significantly improving the accuracy and wide adaptability of mobility prediction.
[0076] 2. Effectively achieve real-time identification and risk assessment of non-standard obstacles: This invention combines the mapping capability from terrain features to physical attributes with Bayesian uncertainty modeling to achieve visual modeling and confidence output of non-geometric obstacles such as slipping and sinking, solving the problem that traditional vision systems have difficulty in identifying non-standard obstacles and improving the robustness of environmental perception.
[0077] 3. Enhanced intelligence, reliability, and real-time performance of planning and decision-making mechanisms: In global path planning, this invention introduces a deep reinforcement learning framework based on a maximum entropy strategy, enabling unmanned off-road vehicles to autonomously learn the optimal path to avoid high-risk areas in uncertain terrain environments, thus improving the global optimality and safety of the path. In local dynamic programming, a sequential convex optimization method is used to quickly transform complex nonlinear constraints into a convex optimization problem that can be solved in real time, ensuring high-efficiency real-time avoidance control under computationally limited conditions.
[0078] 4. Systematically improve the survivability and mission completion rate of off-road vehicles in complex unstructured environments: The integrated perception-modeling-prediction-decision system established by this invention enables unmanned off-road vehicles to cope with various complex challenges such as sudden terrain changes and non-standard obstacle interference, effectively improving the intelligence and reliability of the system, and is widely applicable to mission execution in extreme environments such as disaster relief, unmanned battlefields, and polar exploration. Attached Figure Description
[0079] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0080] Figure 1 This is a flowchart illustrating the overall framework of an embodiment of the present invention;
[0081] Figure 2 This is a flowchart of the vehicle-ground coupling dynamics modeling method according to an embodiment of the present invention;
[0082] Figure 3 This is a non-standard obstacle recognition structure diagram according to an embodiment of the present invention;
[0083] Figure 4 This is a flowchart of the global path planning based on deep reinforcement learning in an embodiment of the present invention;
[0084] Figure 5 This is a schematic diagram of the local dynamic obstacle avoidance mechanism based on sequential convex optimization in an embodiment of the present invention. Detailed Implementation
[0085] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0086] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0087] Example 1:
[0088] like Figure 1As shown, this embodiment provides an obstacle perception method for unmanned vehicles in complex off-road environments. The dataset is constructed by collecting data in a simulation environment. The two fundamental elements of the simulation environment include geographical elevation and geology, and vehicle platform dynamic parameters, thereby reproducing the vehicle's driving results in the real world. Vehicle-ground coupled dynamic modeling is mainly used to predict the vehicle's maneuverability and its uncertainties, and to assess driving risks. Non-standard obstacle recognition is mainly used to effectively identify adverse ground conditions caused by terrain, vegetation cover, and weather. The planning and decision-making process uses the previous results as input. First, reinforcement learning is used to generate policies and complete the global planning process. Then, local planning is performed based on some constraints.
[0089] Specifically, it includes the following steps:
[0090] S1. Collect terrain parameters and vehicle control parameters to construct a deep network model; in this embodiment, terrain parameters include: geographic elevation and geological information; specifically: elevation variation function h(x,y), soil physical properties (such as viscosity, moisture) s=[s1,s2,...,s n The vehicle control parameters are the input control quantities u, including: speed v, state θ, and φ.
[0091] S2. Based on terrain parameters, vehicle control parameters, and deep network models, the vehicle's mobility index is obtained; and the uncertainty and risk of the mobility index are quantified.
[0092] Among them, the vehicle's mobility indicators include:
[0093]
[0094] In the formula, This represents the maneuverability metrics predicted by the model, such as linear velocity, slip ratio, and passability score; f DNN This is a calculation function for the mobility index fitted by a neural network.
[0095] This embodiment uses a deep neural network for dynamic modeling, and learns the dynamic response mapping relationship of off-road vehicles under different terrain conditions through a data-driven approach, specifically as follows: Figure 2 As shown in the diagram. The deep neural network model was trained using training data consisting of multiple terrain-vehicle-response triples, primarily sourced from a high-fidelity vehicle-ground coupling simulation platform, containing thousands of terrain sample combinations. Vehicle parameters were selected and adjusted using the Carsim platform. The collected data was then input into the deep neural network model for transfer reinforcement learning, and test data was recorded.
[0096] Methods for quantifying the uncertainty of mobility indicators include: Since terrain elevation and soil conditions are affected by sensing accuracy, remote sensing resolution, and environmental disturbances, they possess a certain degree of uncertainty. Therefore, probabilistic modeling of the prediction results is necessary to estimate their risk range. To this end, this embodiment introduces a Bayesian neural network to model the probability distribution of mobility prediction. The modeling objective is: in, represents the training sample data; p() represents the Monte Carlo simulation process; the output is the posterior distribution of the vehicle mobility index under specific terrain and conditions.
[0097] Bayesian modeling allows us to obtain the mean and variance of vehicle mobility predictions, thus measuring uncertainty. An active learning mechanism continuously optimizes the DNN model structure and sampling strategy, guiding the model to focus more on regions with higher prediction errors, thereby improving overall prediction accuracy and reliability.
[0098] Methods for quantifying the risk of mobility indicators include: using the elevation change function obtained from the i-th sampling. Soil physical properties Input BNN model (where Multiple Monte Carlo simulations were performed on the set of elevation variation functions and the set of soil physical properties, respectively. The statistical prediction results are as follows:
[0099]
[0100] In the formula, f BNN This represents the BNN model; N represents the number of Monte Carlo simulations. Indicates the expected speed.
[0101] The calculated probability of the off-road vehicle "losing mobility" is:
[0102]
[0103] In the formula, P risk This indicates the probability that the road surface ahead will affect the vehicle's maneuverability; v th This represents the mobility threshold, which is the minimum acceptable speed threshold set. This indicates an indicator function.
[0104] This allows us to obtain the probability of mobility risk for off-road vehicles on a specific route or in a particular area, providing a quantitative basis for subsequent route selection and risk avoidance strategies.
[0105] S3. Acquire image data based on a binocular camera stereo vision system, process the image data, and assess the passability risk probability based on the processed image data to complete the non-standard obstacle recognition.
[0106] Unmanned off-road vehicles face highly diverse and uncertain unstructured natural environments. Ground types such as grassland, sand, and wetlands not only differ significantly in visual appearance, but their corresponding physical characteristics also directly affect vehicle passability. Muddy, soft areas, and sections prone to vehicle slippage are typical non-geometric obstacles. Due to their blurred boundaries and lack of clear structural features, they are difficult to accurately detect and identify using traditional image recognition methods that rely on geometric information. To address this issue, this embodiment proposes a non-standard obstacle detection method that integrates deep visual perception and vehicle-ground dynamics modeling, based on the semantic understanding capabilities of image segmentation and the physical mechanisms by which ground mechanical parameters affect vehicle maneuverability. This method effectively identifies and predicts the risks of off-road vehicle passage under complex terrain conditions. Figure 3 The paper describes the process of acquiring image data using a binocular vision system, performing terrain semantic segmentation and 3D reconstruction, and combining terrain physical characteristics prediction with Bayesian classification to identify potential non-geometric obstacle areas such as slippage and sinkholes.
[0107] To reduce system complexity and sensor deployment costs, this embodiment employs a binocular camera stereo vision system and uses a lightweight network model to extract depth maps and semantic segmentation information. Specifically, for image data, feature extraction operators are used to obtain left-eye view features and right-eye view features: I_L(x,y) and I_R(x,y). Binocular stereo matching is then performed to obtain the disparity map D(x,y), thereby obtaining pixel depth information.
[0108]
[0109] In the formula, f represents the camera focal length; B represents the baseline distance; and Z(x,y) represents the depth at the pixel.
[0110] By training a multi-scale stereo matching network, the geometric embedding features φ(x,y) of each pixel are extracted, and then combined with semantic segmentation information to obtain the terrain category corresponding to each pixel.
[0111] Traditional visual semantic segmentation only classifies ground types but cannot quantitatively assess ground support or traction performance. Therefore, this embodiment introduces a "visual-physical parameter mapping model" to map visual features φ(x,y) extracted from images, such as texture, color, and structure, to soil physical properties, such as surface cohesion c, friction coefficient μ, and density ρ.
[0112] s = f map (φ(x,y);θ); (5)
[0113] In the formula, f mapA multilayer perceptron trained from experimental calibration / simulation data is used to estimate terrain parameters closely related to mobility; θ represents the set of parameters of the multilayer perceptron.
[0114] To evaluate whether the currently detected terrain belongs to the learned distribution of the training set, a Gaussian Mixture Model (GMM) is introduced to model the feature distribution, resulting in the training distribution:
[0115]
[0116] In the formula, p(φ(x,y)) represents the probability density value of the feature under the training distribution; k represents the k-th Gaussian component; K represents the number of Gaussian components in the GMM; π k This represents the mixing weight of the k-th Gaussian component; μ represents the probability density function of the k-th Gaussian distribution; k Σ represents the mean vector of the k-th Gaussian distribution; k Let represent the covariance matrix of the k-th Gaussian distribution.
[0117] The log-likelihood of the training distribution, logp(φ(x,y)), is calculated to determine whether it belongs to a known terrain type (known terrain is used to correspond to known soil parameters and other information). The threshold is set according to the percentile method. The log-likelihood value p(φ(x,y)) of all samples on the training dataset is calculated, and then a lower percentile (such as 5%) is set as the threshold: θ=Percentile({logp(φ(x,y))},5). A low likelihood indicates that there is model uncertainty, and the risk assessment level of the subsequent decision-making system should be increased.
[0118] The soil physical properties s extracted after image segmentation, and their uncertainty distribution δs, are input into the established Bayesian dynamics model to obtain the prediction results of the mobility index:
[0119] v pred ~p(v|s,h(x,y),δs); (7)
[0120] In the formula, v pred f BNN The prediction speed sampled by the function in the Monte Carlo simulation.
[0121] And calculate the probability of success:
[0122] P fail =P(v pred <v th (8)
[0123] In the formula, P fail P(v) represents the probability of success or failure; pred <vth ) represents the probability that the predicted speed is less than the maneuverability threshold.
[0124] If a certain area has high P fail (If 50% is set as the threshold, it can be adjusted according to the urgency of the task in the actual situation) then it can be marked as a non-standard obstacle area and fed back to the local planning module in real time for path adjustment.
[0125] S4. Perform global path planning and local dynamic planning for off-road vehicles. Global path planning includes: constructing paths and assigning tasks on coarse-grained terrain maps based on the principle of minimizing uncertainty of mobility indicators. Local dynamic planning includes: avoiding and adjusting risk points in the path based on the results of non-standard obstacle identification and dynamic environmental perception data.
[0126] To enhance the operational stability and safety of unmanned off-road vehicles in complex and dynamic environments, this embodiment proposes a global-local collaborative intelligent decision-making mechanism that integrates deep reinforcement learning and sequential convex optimization. This mechanism aims to simultaneously satisfy global optimality in long-range path planning and high real-time response in near-field obstacle avoidance. It comprehensively leverages the long-term policy optimization capabilities of deep reinforcement learning for environmental states and the efficient solution advantages of sequential convex optimization under nonlinear constraints, enabling unmanned off-road vehicles to achieve both forward-looking and adaptive intelligent decision-making capabilities in varied terrains.
[0127] In this embodiment, the global planning problem is formalized as a Markov decision process (MDP) in reinforcement learning:
[0128]
[0129] In the formula: A is the state space, which includes the off-road vehicle's position, speed, attitude, and maneuverability risk indicators; A is the action space, defined as continuous variables such as throttle, braking, and steering; P is the state transition function, provided by the vehicle-ground coupled dynamics model; r(s,a) is the reward function, which combines factors such as the target task, path shortestness, and risk avoidance; γ is the discount factor.
[0130] To improve learning efficiency and policy stability in a continuous action space, this embodiment employs the maximum entropy deep reinforcement learning algorithm SoftActor-Critic (SAC), whose optimization objective is to maximize the expected reward of entropy regularization.
[0131]
[0132] In the formula, π * Let represent the optimal policy, which maximizes the weighted sum of expected reward and policy entropy; π represents the policy function, defined in a given state s. tTake action a t The probability distribution; The expected value of the state-action trajectory distribution under policy π is represented by t; t represents time t; α is the entropy weighting coefficient, and r(s) t ,a t ) indicates that at time step t, state s t Next, execute action a t The instant reward received; Let be the policy distribution entropy.
[0133] Figure 4 The paper describes the construction of the state space and action space in the SAC algorithm, introduces a mobility risk index into the reward function, and realizes the generation of the optimal safe path for the unmanned off-road vehicle in the global scope through experience pool training and reinforcement learning policy network update.
[0134] To enhance the ability to avoid low-maneuverability areas in off-road terrain, a vehicle maneuverability risk index R is embedded in the reward function. mob :
[0135] r(s,a)=w1·r goal +w2·r smooth -w3·R mob (s); (11)
[0136] Where: r goal To incentivize off-road vehicles to approach the target area; r smooth For path smoothness constraints; w i These are the weighting coefficients.
[0137] During training, the experience triplet of off-road vehicles (s t a t r t s t+1 The policy is stored in the experience pool and updated accordingly. Soft Q-function loss is used for policy iteration.
[0138]
[0139] in,
[0140]
[0141] In the formula, This represents the loss value used for model training; represents the expected value; Q(s,a) represents the current Q-function's estimate of the state-action pair; The target Q-value represents the objective of supervised training. Let represent the expectation of sampling action a′ from s′ under policy π; Q(s′,a′) represents the Q value of Q-learning reinforcement learning; a′ and s′ represent the action value under s′ and the next state reached after executing the action, respectively; logπ(a′|s′) represents the log probability value of policy π, which is used to measure the uncertainty of the policy for this action.
[0142] Through continuous updates to the policy network and value network, the global planning strategy can be gradually optimized to its optimum in simulation or real-world environments.
[0143] While global path planning considers macroscopic obstacles and maneuverability information, limitations in real-time terrain perception accuracy mean that vehicles are still prone to encountering non-standard obstacles or sudden terrain changes during travel, necessitating dynamic avoidance. Local path planning is modeled as the following constrained optimization problem:
[0144]
[0145] The constraints are:
[0146] Dynamic model: x {t+1} =f(x) t ,u t );
[0147] Control input restrictions:
[0148] Mobility limit: v(x) t )≥v min ;
[0149] Obstacle avoidance restrictions:
[0150] In the formula: λ represents the control sequence, from time 1 to T; λ represents the risk penalty coefficient, which controls the intensity of the penalty for entering the risk zone. The larger the value, the more sensitive the system is to the risk zone. t Vehicle status; This represents the desired state of the reference trajectory (target trajectory) at time t; x {t+1} Indicates the vehicle state at time t+1; u t For control input; High-risk areas identified by non-standard barriers; This represents the vehicle's control command set; The set of obstacles perceived at the current time; v min The minimum tolerable driving speed; v(x) t ) indicates that the car is at x t The velocity under the given state; f(·) represents the vehicle dynamics model; Indicates the indicator function: if x tFalling into the "risk zone" Returns 1; otherwise, returns 0.
[0151] Due to f(·), v(x t These are all nonlinear functions, and directly solving this optimization problem results in high computational complexity, making it unsuitable for real-time requirements. To improve the real-time performance and deployability of local programming, this embodiment employs a sequential convex optimization method, linearizing the original non-convex programming problem into a series of convex subproblems that are solved iteratively:
[0152] Step 1: Using the current optimal trajectory Using the expansion point as the base, apply the Taylor first-order approximation to all nonlinear constraints.
[0153] Step 2: Construct a new subproblem with linear constraints / objectives and convert it into a standard convex quadratic programming (QP) format.
[0154] Step 3: Solve for the new trajectory Determine if convergence has occurred. If not, iterate and update to form a closed loop. Here, δ represents the direction of the variable increment, which comes from the solution of the linear subproblem.
[0155] Repeat the above steps until the trajectory stabilizes or the termination condition is met, and finally obtain obstacle avoidance control commands that can be executed in real time. Figure 5 This illustrates how complex nonlinear constraints (dynamic model, obstacle location, mobility threshold) are linearized into a convex optimization problem during local path correction, constructing a fast solvable path and supporting dynamic trajectory updates at runtime to cope with sudden obstacles.
[0156] Example 2:
[0157] This embodiment provides an obstacle perception system for unmanned vehicles in complex off-road environments, including: a parameter acquisition module for acquiring terrain parameters and vehicle control parameters, and constructing a deep network model; terrain parameters include geographic elevation and geological information; vehicle control parameters include speed and state; a dynamic modeling module for obtaining vehicle mobility indicators based on terrain parameters, vehicle control parameters, and the deep network model, and quantifying the uncertainty and risk of the mobility indicators; a non-standard obstacle recognition module for acquiring image data based on a binocular camera stereo vision system, processing the image data, and assessing the probability of passability risk based on the processed image data, thereby completing the non-standard obstacle recognition; and a global task planning and local dynamic decision-making module for performing global path planning and local dynamic planning for the off-road vehicle; global path planning includes constructing a path and allocating tasks on a coarse-grained terrain map based on the principle of minimizing the uncertainty of the mobility indicators; local dynamic planning includes avoiding and adjusting risk points in the path based on the results of non-standard obstacle recognition and dynamic environmental perception data.
[0158] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for obstacle perception in unmanned vehicles operating in complex off-road environments, characterized in that: Includes the following steps: S1. Collect terrain parameters and vehicle control parameters to construct a deep network model; The terrain parameters include: geographic elevation and geological information; the vehicle control parameters include: speed and status. S2. Based on the terrain parameters, the vehicle control parameters, and the deep network model, obtain the vehicle's mobility index; and quantify the uncertainty and risk of the mobility index; the mobility index includes: In the formula, The model predicts the mobility index, h(x,y) is the elevation variation function; s represents the soil physical properties; u represents the input control variable; f DNN This is a calculation function for the mobility index fitted by a neural network. S3. Acquire image data based on a binocular camera stereo vision system, process the image data, and assess the passability risk probability based on the processed image data to complete the non-standard obstacle identification. S4. Perform global path planning and local dynamic planning for the off-road vehicle; the global path planning includes: constructing a path and assigning tasks on a coarse-grained topographic map based on the principle of minimizing uncertainty of the mobility index; the local dynamic planning includes: avoiding and adjusting risk points in the path based on the results of non-standard obstacle identification and dynamic environmental perception data.
2. The obstacle perception method for unmanned vehicles in complex off-road environments according to claim 1, characterized in that, The method for quantifying the uncertainty of the aforementioned mobility index includes: A Bayesian neural network is used to model the probability distribution of mobility prediction, with the modeling objective being: in, This represents the training sample data, p() is the Monte Carlo simulation process; the output is the posterior distribution of the vehicle's mobility index under specific terrain and conditions; Obtain the mean and variance of the maneuverability prediction, and then complete the uncertainty quantification; The methods for quantifying the risk of the aforementioned mobility index include: The elevation variation function and soil physical properties were input into the BNN model for several Monte Carlo simulations to obtain the prediction results: In the formula, f BNN This represents a network formed by a Bayesian neural network; N represents the number of Monte Carlo simulations. Indicates the expected speed; h i s represents the elevation change function for the i-th sampling; i This represents the physical properties of the soil sampled in the i-th sampling. Calculate the probability of an off-road vehicle losing maneuverability: In the formula, P risk This indicates the probability that the road surface ahead will affect the vehicle's maneuverability; v th Indicates the mobility threshold; Indicates an indicator function.
3. The obstacle perception method for unmanned vehicles in complex off-road environments according to claim 1, characterized in that, Non-standard obstacle identification methods include: The image data is subjected to depth map extraction and semantic segmentation to obtain depth information and semantic segmentation information of pixels; the method for performing semantic segmentation on the image data includes: extracting visual features from the image data and mapping the visual features to soil physical properties; Spatial alignment and geological data matching are performed on the depth information to obtain the geometric embedding features of the pixels; The terrain category corresponding to each pixel is obtained based on the geometric embedding features and the semantic segmentation information; The visual features are modeled using a Gaussian mixture model to obtain the training distribution; Calculate the log-likelihood of the training distribution to determine whether the terrain category belongs to the known terrain; Based on the soil physical properties extracted after image segmentation and the uncertainty distribution, the prediction results of the mobility index are obtained. Based on the prediction results, the passability risk probability is calculated, and non-standard obstacle identification is completed based on the passability risk probability.
4. The obstacle perception method for unmanned vehicles in complex off-road environments according to claim 3, characterized in that, Methods for obtaining prediction results of mobility indices based on soil physical properties s extracted after image segmentation and the uncertainty distribution δs of soil physical properties include: in pred ~p(v|s,h(x,y),δs); In the formula, v pred f BNN The predicted velocity sampled in the Monte Carlo simulation; p() represents the Monte Carlo simulation process; The method for calculating the probability of passability risk based on the prediction results includes: P fail =P(in pred <in th ); In the formula, P fail P(v) represents the probability of success or failure; pred <v th ) represents the probability that the predicted speed is less than the maneuverability threshold.
5. The obstacle perception method for unmanned vehicles in complex off-road environments according to claim 1, characterized in that, The global path planning employs a maximum entropy deep learning reinforcement algorithm for path planning; the method includes: Transform the global path planning problem into a Markov decision process: In the formula, For state space; Let be the action space; P be the state transition function; r(s,a) be the reward function; and γ be the discount factor. Employing a maximum entropy deep learning reinforcement algorithm to maximize the expected return of entropy regularization: In the formula, π * π represents the optimal policy; π represents the policy function. R represents the expected value of the state-action trajectory distribution under policy π; t represents time; α is the entropy weighting coefficient, and r(s) represents the expected value of the state-action trajectory distribution under policy π. t ,a t ) indicates that at time step t, state s t Next, execute action a t The instant rewards received; The strategy distribution entropy; The reward function embeds the vehicle mobility risk index R. mob : r(s,a)=w1·r goal +w2·r smooth -w3·R mob (s); Where: r goal To incentivize off-road vehicles to approach the target area; r smooth For path smoothness constraints; w i These are weighting coefficients; Policy iteration using soft Q-function loss: in, In the formula, This represents the loss value used for model training; represents the expected value; Q(s,a) represents the current Q-function's estimate of the state-action pair; Indicates the target Q value; denoted by ; Q(s′,a′) represents the Q-value of Q-learning reinforcement learning; a′ and s′ represent the action value under s′ and the next state reached after executing the action, respectively; logπ(a′|s′) represents the logarithmic probability value of policy π.
6. The obstacle perception method for unmanned vehicles in complex off-road environments according to claim 1, characterized in that, The local dynamic programming is modeled as a constrained optimization problem: The constraints are: Dynamic model: x {t+1} =f(x) t ,u t ); Control input restrictions: Mobility limit: v(x) t )≥v min ; Obstacle avoidance restrictions: In the formula: Denotes the control sequence; λ represents the risk penalty coefficient; x t Vehicle status; x represents the desired state of the reference trajectory at time t; {t+1} Indicates the vehicle state at time t+1; u t For control input; High-risk areas identified by non-standard barriers; This represents the vehicle's control command set; The set of obstacles perceived at the current time; v min The minimum tolerable driving speed; v(x) t ) indicates that the car is at x t The velocity under the given state; f(·) represents the vehicle dynamics model.
7. The obstacle perception method for unmanned vehicles in complex off-road environments according to claim 6, characterized in that, The optimization methods for local dynamic programming include: Step 1: Using the current optimal trajectory Using the expansion point as the Taylor first-order approximation for all nonlinear constraints; Step 2: Construct a new subproblem with linear constraints / objectives, and convert it into a standard convex quadratic programming format; Step 3: Solve for the new trajectory Determine if convergence has occurred; if not, iterate and update to form a closed loop. Here, δ represents the direction of the variable increment, which comes from the solution of the linear subproblem.
8. An obstacle perception system for unmanned vehicles in complex off-road environments, the system being used to implement the method described in any one of claims 1-7, characterized in that, include: The parameter acquisition module is used to collect terrain parameters and vehicle control parameters to build a deep network model; The terrain parameters include: geographic elevation and geological information; the vehicle control parameters include: speed and status. The dynamics modeling module is used to obtain the vehicle's maneuverability index based on the terrain parameters, the vehicle control parameters, and the deep network model; and to quantify the uncertainty and risk of the maneuverability index. The non-standard obstacle recognition module is used to acquire image data based on a binocular camera stereo vision system, process the image data, and perform a passability risk probability assessment based on the processed image data, thereby completing the non-standard obstacle recognition. The global task planning and local dynamic decision-making module is used to perform global path planning and local dynamic planning for the off-road vehicle. The global path planning includes: constructing a path and allocating tasks on a coarse-grained topographic map based on the principle of minimizing uncertainty of the mobility index. The local dynamic planning includes: avoiding and adjusting risk points in the path based on the results of non-standard obstacle identification and dynamic environmental perception data.
Citation Information
Patent Citations
Visual semantic segmentation method, device and equipment
CN114549542A
Ground material recognition method and device and sweeping robot
CN113362355A
Path planning method for unmanned off-road vehicle in complex terrain environment
CN116793380A
Autonomous vehicle posture transformation control system and method based on terrain perception
CN117382362A
Vehicle curve safety early warning monitoring method, device and system
CN118343153A