Off-road vehicle subsidence escape method and system

By constructing a pre-trained policy network, extracting features from vehicle state and control action data, and combining reinforcement learning to train an escape strategy, the problem of perception and control when off-road vehicles get stuck is solved, achieving a fast and safe escape effect.

CN120922133AActive Publication Date: 2025-11-11TONGJI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511461387.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-11-11
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

Existing driver assistance systems struggle to accurately perceive and model off-road environments, lacking low-speed high torque and precise slip control, resulting in off-road vehicles being unable to quickly and effectively escape when stuck. Furthermore, strategies relying on human driving data have significant limitations.

Method used

By using a pre-trained policy network to extract traction features from vehicle state and control action data, an observation space, action space, and reward function are constructed. This is combined with reinforcement learning to train the escape strategy network, coordinating the vehicle's steering and drive system to achieve intelligent escape from difficult situations.

Benefits of technology

It can quickly and safely assist vehicles in getting out of trouble in complex off-road environments, improving the safety and efficiency of off-road driving and reducing the risk of human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120922133A_ABST
    Figure CN120922133A_ABST
Patent Text Reader

Abstract

The invention provides an off-road vehicle subsidence escape method and system, and relates to the technical field of vehicle stability control, and the method comprises the following steps: obtaining state data and control action data of an off-road vehicle when the off-road vehicle subsides, extracting characteristics capable of representing traction force from the state data and the control action data; determining a corresponding observation space, an action space and a reward function; constructing an off-road vehicle escape strategy network, and learning and training the off-road vehicle escape strategy network based on the observation space, the action space and the reward function combination; acquiring state data and control action data when the actual off-road vehicle sinks, and inputting the state data and the control action data into the off-road vehicle escape strategy network for processing to obtain a corresponding escape strategy; according to the invention, a specific trapped scene does not need to be identified, and intelligent de-trapping can be realized through a pre-training strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle stability control technology, and more specifically to a method and system for off-road vehicles to get out of trouble when they get stuck. Background Technology

[0002] Currently, off-road environments are complex and varied, making it easy for vehicles to get stuck on low-traction surfaces such as mud, sand, snow, or potholes. Once a vehicle is stuck, manual control by the driver is often insufficient to extricate it using traditional methods (throttle control, pushing, towing), and may even cause further damage or personal injury. Developing intelligent traction control systems can proactively intervene when a vehicle loses traction, helping it quickly regain its driving ability, thereby significantly improving the safety and efficiency of off-road driving.

[0003] However, existing driver assistance systems (ADAS) only consider everyday passenger vehicle use when dealing with off-road vehicle entrapment, primarily focusing on structured road environments. Their perception relies on typical traffic features such as lane lines, vehicles, and pedestrians, and their control logic prioritizes comfort and economy. However, in off-road and entrapment scenarios, roads lack rule constraints, and terrain and adhesion conditions are complex and variable. Existing systems struggle to accurately perceive and model these conditions and lack features such as low-speed high torque and precise slip control. These limitations result in insufficient adaptability in off-road environments and may even exacerbate the predicament due to improper intervention. Therefore, they cannot meet the stability and entrapment capabilities required by off-road vehicles. For example, existing patent CN 116588118A proposes an automatic entrapment method for vehicles. This method requires collecting and classifying driving data from professional off-road drivers to design different entrapment schemes. It requires identifying the current entrapment type based on sensors and selecting the corresponding entrapment scheme. However, this approach cannot exhaustively cover all real-world entrapment scenarios, and human operation cannot precisely and quickly coordinate and control the vehicle's drive and steering subsystems. Therefore, entrapment strategies relying on human driving data have significant limitations.

[0004] Therefore, how to provide a method for off-road vehicles to get out of trouble and solve the above problems is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a method and system for off-road vehicles to get out of trouble, which does not rely on human driving data and does not require identification of specific trapped scenarios. Intelligent escape can be achieved through pre-trained strategies.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A method for an off-road vehicle to get out of trouble includes the following steps: S1: When the off-road vehicle gets stuck, acquire the off-road vehicle's status data and control action data, and extract features that can characterize traction from the status data and control action data. S2: Determine the corresponding observation space, action space, and reward function; S3: Construct an off-road vehicle escaping strategy network, and learn and train the off-road vehicle escaping strategy network based on a combination of observation space, action space, and reward function; S4: Obtain the status data and control action data of the actual off-road vehicle when it gets stuck, and input the status data and control action data into the off-road vehicle get-out strategy network for processing to obtain the corresponding get-out strategy.

[0007] Preferably, S1 includes: S11: Acquire status data and control action data of off-road vehicles; S12: Construct a longitudinal dynamics model of the off-road vehicle and a data-driven longitudinal dynamics model of the off-road vehicle based on the state data and control action data; S13: Based on the data, the longitudinal dynamics model of the off-road vehicle is used to obtain features that can characterize traction.

[0008] Preferably, S12 includes: S121: Construct a longitudinal dynamics model of an off-road vehicle and a data-driven longitudinal dynamics model of an off-road vehicle based on the state data and control action data, wherein the longitudinal dynamics model of the off-road vehicle includes a longitudinal dynamics model of the off-road vehicle and a wheel rolling dynamics model, and the data-driven longitudinal dynamics model of the off-road vehicle includes a data-driven wheel rolling model and a data-driven longitudinal dynamics model of the off-road vehicle. S122: Establish a data-driven wheel rolling model based on the wheel rolling dynamics model, and estimate the error of the data-driven wheel rolling model; S123: Based on the data-driven wheel rolling model and the error estimation results obtained in S122, construct the corresponding data-driven longitudinal dynamics model of the off-road vehicle.

[0009] Preferably, S2 includes: S21: Determine the observation space and action space; S22: Determine the reward function, wherein the reward function includes: reward for facilitating escaping difficulty, reward for preventing wheel slippage, and penalty for controlling the action.

[0010] Preferably, S3 includes: S31: Construct an off-road vehicle escaping strategy network, wherein the off-road vehicle escaping strategy network includes an environment model, a strategy model, and a value model; S32: The environment model, policy model, and value model are learned and trained based on the observation space, action space, and reward function, respectively. S33: Construct the loss functions corresponding to the environment model, strategy model, and value model, and minimize the loss functions during the learning and training process to obtain the final off-road vehicle escaping strategy network.

[0011] The present invention also provides an off-road vehicle getting stuck and freeing system, comprising: The feature extraction module is used to acquire the state data and control action data of the off-road vehicle when the off-road vehicle gets stuck, and to extract features that can characterize traction from the state data and the control action data. The determination module is used to determine the corresponding observation space, action space, and reward function; The model training module is used to construct the off-road vehicle traction strategy network and learn and train the off-road vehicle traction strategy network based on the combination of observation space, action space and reward function. The processing module is used to acquire the status data and control action data of the actual off-road vehicle when it gets stuck, and input the status data and control action data into the off-road vehicle get-out strategy network for processing to obtain the corresponding get-out strategy.

[0012] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method and system for off-road vehicles to get out of trouble. It uses a dynamic model learned from vehicle driving data to simulate or predict the state changes of the vehicle in a complex off-road environment, thereby assisting the reinforcement learning agent in coordinating the vehicle steering and drive subsystems, learning efficient and safe escape strategies, and further accelerating the training speed and escape efficiency of reinforcement learning. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0014] Figure 1 This is an overall flowchart of a method for an off-road vehicle to get out of trouble, provided by an embodiment of the present invention; Figure 2 This invention provides a structural principle block diagram of an off-road vehicle getting stuck and escaping a difficult situation. Figure 3 This is a schematic diagram of vehicle speed under full torque drive provided in an embodiment of the present invention; Figure 4This is a schematic diagram of vehicle acceleration under full torque drive provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the vehicle position during full torque drive provided in an embodiment of the present invention; Figure 6 A schematic diagram showing the vehicle's front wheel steering angle, four-wheel drive torque, wheel speed, longitudinal speed, slip ratio, and position for the MBRL traction control strategy provided in this embodiment of the invention. Figure 7 A schematic diagram of vehicle acceleration for the MBRL escaping strategy provided in an embodiment of the present invention; Figure 8 A schematic diagram showing the vehicle's front wheel steering angle, four-wheel drive torque, wheel speed, longitudinal speed, slip ratio, and position for the DMBRL traction control strategy provided in this embodiment of the invention. Figure 9 A schematic diagram of vehicle acceleration for the DMBRL traction control strategy provided in an embodiment of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] See Figure 1 As shown in the figure, an embodiment of the present invention discloses a method for an off-road vehicle to get out of trouble, including the following steps: S1: When an off-road vehicle gets stuck, acquire the vehicle's status data and control action data, and extract features that can characterize traction from the status data and control action data. The status data may include the vehicle's speed, wheel speed, etc., and the control action data may include the steering wheel angle, drive torque, etc. S2: Determine the corresponding observation space, action space, and reward function; S3: Construct an off-road vehicle escaping strategy network, and learn and train the off-road vehicle escaping strategy network based on a combination of observation space, action space, and reward function; S4: Acquire the status data and control action data of the actual off-road vehicle when it gets stuck, and input the status data and control action data into the off-road vehicle get-out strategy network for processing to obtain the corresponding get-out strategy.

[0017] In one specific embodiment, S1 includes: S11: Acquire status data and control action data of off-road vehicles; S12: Constructing a longitudinal dynamics model of off-road vehicles based on state data and control action data, and a data-driven longitudinal dynamics model of off-road vehicles; S13: Based on the data-driven longitudinal dynamics model of off-road vehicles, characteristics that can characterize traction are obtained.

[0018] In one specific embodiment, S12 includes: S121: Construct a longitudinal dynamics model and a data-driven longitudinal dynamics model of off-road vehicles based on state data and control action data. The longitudinal dynamics model of off-road vehicles includes a longitudinal dynamics model of off-road vehicles and a wheel rolling dynamics model. The data-driven longitudinal dynamics model of off-road vehicles includes a data-driven wheel rolling model and a data-driven longitudinal dynamics model of off-road vehicles. S122: Establish a data-driven wheel rolling model based on the wheel rolling dynamics model, and estimate the error of the data-driven wheel rolling model; S123: Based on the data-driven wheel rolling model and the error estimation results obtained in S122, construct the corresponding data-driven longitudinal dynamics model of the off-road vehicle.

[0019] Specifically, when an off-road vehicle gets stuck, generating sufficient traction is a necessary condition for escaping the predicament. Therefore, in this embodiment of the invention, features characterizing traction are first manually extracted based on the vehicle's measurable state and actuator inputs, and these features, along with the vehicle's state, are used as inputs to a neural network. Thus, the specific process of S1 includes: A. Vehicle longitudinal dynamics model: The specific expressions for the traditional vehicle longitudinal dynamics and wheel rolling dynamics models are as follows: (1) (2) In the formula, and Indicates the longitudinal and lateral speeds of the vehicle. Indicates yaw rate. Indicates vehicle mass. This indicates the front wheel steering angle; for each wheel, the control input is the motor torque. The state is the rotational speed. The rolling radius of the wheel is The moment of inertia of the wheel is , The longitudinal traction force for each wheel, This represents each wheel, in the following order: front left wheel, front right wheel, rear left wheel, and rear right wheel.

[0020] B. Data-driven model

[0021] Traditional models (1) and (2) struggle to capture key vehicle characteristics, namely longitudinal forces. Therefore, this invention provides a data-driven method for extracting features to accelerate the training process of reinforcement learning. For the rolling dynamics model of a wheel, the specific expression of the corresponding data-driven model is as follows: (3) (4) in, (5) (6) (7) (8) (9) In the formula, ,correspond The rotational speed of each wheel, It can characterize longitudinal force. , Let t be the torque of the four-wheel motor. The number of data points collected. for An identity matrix of dimension 1, elements (where i = 1, 2, ..., N2) represents the influence coefficient of motor torque on wheel speed, which can be obtained through offline step response experiments.

[0022] C. Model error estimation: In models (3)-(4), This can be obtained through torque step testing, and requires real-time estimation based on the vehicle's online input and output data. Meanwhile, the estimated It can characterize the longitudinal force of each wheel in the traditional model (2). The estimation method is as follows: (10)

[0023] (11) because Equations (10) and (11) can be simplified to: (12) (13) D. Data-driven vehicle longitudinal model: Similar to the data-driven wheel rolling model in steps A and B, a data-driven vehicle longitudinal model can be established, with the specific expression as follows: (14) (15) in, Indicates the vehicle's current speed; For the defined system state, (16) (17) (18) (19) in It can characterize the longitudinal resultant force on a vehicle. The number of data points collected. Similar to (12) and (13), we can obtain... and The estimated values ​​are as follows: (20) (twenty one) In the formula, Represents an N-dimensional identity matrix. It indicates the speed of the vehicle at the next moment.

[0024] Therefore, the embodiments of the present invention extract the following through the above steps: and As feature data, along with vehicle state and control actions, it serves as input to subsequent models, simultaneously learning environmental dynamics models and training reinforcement learning strategies.

[0025] In one specific embodiment, S2 includes: S21: Determine the observation space and action space, where the observation space is defined as: (twenty two) In the formula, v and v represent the wheel speed and vehicle speed, respectively.

[0026] To achieve efficient escape from difficult situations, it is necessary to coordinate the steering and drive subsystems; therefore, the action space is defined as: (twenty three) In the formula, Indicates the steering angle of the vehicle's front wheels. This indicates the torque for four-wheel drive.

[0027] S22: To enable off-road vehicles to extricate themselves from getting stuck, a well-defined reward function is needed. This reward function includes: rewards for facilitating extrication, rewards for preventing wheel slippage, and penalties for controlled actions. Specifically, it includes: 1) Rewards to facilitate getting out of trouble: When a vehicle gets stuck, its speed is close to zero, its position changes little, and its wheels may slip significantly. To keep the vehicle moving forward, the rewards for its speed and position are crucial. The two reward functions are defined as follows: (twenty four) (25) In the formula, Represents the speed at the current moment velocity compared to the previous moment The difference in absolute value, This represents the difference in absolute vertical position between the current time and the previous time. It represents the difference between the absolute value of the lateral position at the current time and the previous time.

[0028] 2) Prevent wheel slippage: When a vehicle sinks, its speed is close to zero, and the wheel speed tends to increase, potentially resulting in a high slip ratio. A high slip ratio can prevent the wheels from gaining sufficient traction, even causing the vehicle to sink further. Therefore, large slip should be minimized. Considering that the slip ratio is a high-frequency variable and difficult to measure accurately, a discrete reward function is used. (26) In the formula, and For the set slip ratio threshold, when the slip ratio Less than Given a smaller penalty When slip ratio exist and In between, a moderate penalty is given. When slip ratio Greater than When a larger penalty is given .

[0029] 3) Controlled action punishment: During the escape process, the control strategy coordinates the steering and drive subsystems. To prevent high-frequency oscillations in the actuators, changes in control actions are penalized. (27) In the formula, Represents the normalized action at time t and time t-1. , The change in quantity.

[0030] in, (28)

[0031] In the formula, 'a' represents the normalized control action. and These represent the maximum front wheel steering angle and the maximum wheel torque, respectively.

[0032] 4) Reward function integration: The three reward functions are weighted and combined to obtain the final reward function for off-road vehicles getting stuck and extricating themselves from difficult situations: (29) In the formula, The weighting of rewards is as follows: the reward component that promotes getting out of trouble has a positive weighting, while the penalty component that affects slip rate and control actions has a negative weighting.

[0033] In one specific embodiment, S3 includes: S31: Construct an off-road vehicle escaping strategy network, which includes an environment model, a strategy model, and a value model. S32: The environment model, policy model, and value model are learned and trained based on the observation space, action space, and reward function, respectively. S33: Construct the loss functions corresponding to the environment model, strategy model, and value model, and minimize the loss functions during the learning and training process to obtain the final off-road vehicle escaping strategy network.

[0034] Specifically, this embodiment of the invention employs a model-based reinforcement learning method to learn an efficient vehicle extrication strategy, mainly comprising the following three parts: training an environment model, a policy model, and a value model simultaneously based on measurements from the observation space, specifically including the following processes: 1) Environment model learning: This invention uses a cyclic state-space model as the environment model, specifically including: Circular model: (30) Encoder: (31) Predictive model: (32) Reward Prediction: (33) Continuous forecasting: (34) Decoder: (35) The loop model uses the variables from the previous step. Predicting the next hidden variable During the training phase, the encoder incorporates latent variables. And vehicle state sampled posterior random variables Simultaneously, prior random variables are sampled based on latent variables. To minimize the Kullback-Leibler (KL) divergence between two random variables, in addition to minimizing the error between the reward prediction, continuous prediction, and decoder and the true value, a loss function for the environment model is learned, expressed as follows: (36) in, Indicates distribution The mean is calculated through sampling. For a random distribution represented by a neural network, For the parameters of the neural network, , and The loss represents the predicted loss, the dynamic loss, and the weights representing the loss. , and This represents predicted loss, dynamic loss, and representational loss.

[0035] (37) (38) (39) In the formula, To represent logarithmic calculation, , and These represent decoding operations, reward prediction, and continuous value prediction, respectively. The vehicle status for reconstruction. To pass through random distribution Sampling random variables, Here are the hidden variables predicted by the recurrent neural network. KL[] represents the divergence between two different random distributions. The sg() operator indicates that the gradient within the operator is stopped. The environment model is learned by minimizing the loss function.

[0036] 2) Policy model training: The loss function of the policy model is defined as: (40) In the formula, , The reward at time t, The defined policy model is represented by a neural network with the following network parameters: Its input is The output is a normalized action. , The value model is represented by a neural network with the following parameters: Its input is , The entropy of the policy model, The calculation formula is as follows: (41) In the formula, For the output of the policy network, This indicates the vehicle's current speed. As a reward value, As a discount factor, Indicates whether to continue calculation. Indicates the weighting coefficient; 3) Value model training: The loss function of the value model is defined as follows: (42) In the formula, Let the random distribution be represented by a neural network, and the network parameters be... .

[0037] During the training process, minimizing the loss function (36)(40)(42) simultaneously allows the reinforcement learning strategy to be trained while learning the environment model, ultimately learning the off-road vehicle's escape strategy when it gets stuck.

[0038] See Figure 2 As shown, this embodiment of the invention also provides a system for a method of getting an off-road vehicle out of trouble using any of the above embodiments, comprising: The feature extraction module is used to acquire the state data and control action data of the off-road vehicle when it gets stuck, and to extract features that can characterize traction from the state data and control action data. The determination module is used to determine the corresponding observation space, action space, and reward function; The model training module is used to construct the off-road vehicle traction strategy network and learn and train the off-road vehicle traction strategy network based on the combination of observation space, action space and reward function. The processing module is used to acquire the status data and control action data of the actual off-road vehicle when it gets stuck, and input the status data and control action data into the off-road vehicle get-out strategy network for processing to obtain the corresponding get-out strategy.

[0039] To verify the performance of the method provided in this embodiment of the invention, the high-fidelity off-road simulation software PyChrono was used. PyChrono is a Python library that wraps the Chrono C++ simulation software. The simulation step size was 0.005s, and the control cycle was 0.01s. The test scenario was sandy soil, with the vehicle's four wheels sinking 15cm. The test vehicle was the HMMWV vehicle included with Chrono, weighing approximately 2500kg, with a four-wheel drive torque of 800Nm each and a wheel radius of 0.467m.

[0040] To illustrate the effectiveness of the proposed method in off-road obstacle avoidance, it is compared with two other methods. The first is full torque drive without an obstacle avoidance strategy, which is the method typically used by drivers when a vehicle becomes stuck. The second method does not add the manual features extracted in step 2, i.e., the observation space is... This is denoted as a model-based reinforcement learning strategy (MBRL). The method proposed in this embodiment of the invention for adding data-driven features to the observation space is denoted as a data-driven model-based reinforcement learning strategy (DMBRL).

[0041] Figure 3 , Figure 4 and Figure 5 This provides vehicle speed, acceleration, and position information under full torque drive. Since the vehicle is submerged to a depth of 15cm, it cannot be propelled forward in only one direction; therefore, the vehicle speed is 0 and its position remains unchanged.

[0042] Figure 6 and Figure 7 This represents the vehicle's state when using the MBRL (Maintaining Motor Assistance and Retrieval) strategy, including information such as front wheel steering angle, four-wheel drive torque, wheel speed, longitudinal speed, slip ratio, position, and acceleration. It can be seen that the trained MBRL strategy can assist the vehicle in escaping trouble by coordinating front wheel steering and four-wheel drive. The strategy gradually increases the vehicle's speed and helps it leave the stuck position by repeatedly applying positive and negative torque and turning the steering wheel left and right.

[0043] Figure 8 and Figure 9 This invention proposes a DMBRL strategy that incorporates manually extracted data-driven features into the observation space. Compared to MBRL, DMBRL achieves more efficient vehicle detachment. MBRL takes over 8 seconds to detach the vehicle, with a large slip ratio and higher torque oscillation. In contrast, DMBRL can detach the vehicle within 4 seconds using only two cycles of driving torque, and the slip ratio is smaller than that of MBRL both before and after detachment. This demonstrates the advantage of the method provided in this invention, which combines data-driven features with model-based reinforcement learning.

[0044] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0045] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for an off-road vehicle to get out of trouble when it gets stuck, characterized in that, Includes the following steps: S1: When the off-road vehicle gets stuck, acquire the off-road vehicle's status data and control action data, and extract features that can characterize traction from the status data and control action data. S2: Determine the corresponding observation space, action space, and reward function; S3: Construct an off-road vehicle escaping strategy network, and learn and train the off-road vehicle escaping strategy network based on the combination of observation space, action space and reward function; S4: Obtain the status data and control action data of the actual off-road vehicle when it gets stuck, and input the status data and control action data into the off-road vehicle get-out strategy network for processing to obtain the corresponding get-out strategy.

2. The method for an off-road vehicle to get out of trouble as described in claim 1, characterized in that, S1 includes: S11: Acquire status data and control action data of off-road vehicles; S12: Construct a longitudinal dynamics model of the off-road vehicle and a data-driven longitudinal dynamics model of the off-road vehicle based on the state data and control action data; S13: Based on the data, the longitudinal dynamics model of the off-road vehicle is used to obtain features that can characterize traction.

3. A method for an off-road vehicle to get out of trouble as described in claim 2, characterized in that, S12 includes: S121: Construct a longitudinal dynamics model of an off-road vehicle and a data-driven longitudinal dynamics model of an off-road vehicle based on the state data and control action data, wherein the longitudinal dynamics model of the off-road vehicle includes a longitudinal dynamics model of the off-road vehicle and a wheel rolling dynamics model, and the data-driven longitudinal dynamics model of the off-road vehicle includes a data-driven wheel rolling model and a data-driven longitudinal dynamics model of the off-road vehicle. S122: Establish a data-driven wheel rolling model based on the wheel rolling dynamics model, and estimate the error of the data-driven wheel rolling model; S123: Based on the data-driven wheel rolling model and the error estimation results obtained in S122, construct the corresponding data-driven longitudinal dynamics model of the off-road vehicle.

4. The method for an off-road vehicle to get out of trouble as described in claim 1, characterized in that, S2 includes: S21: Determine the observation space and action space; S22: Determine the reward function, wherein the reward function includes: reward for facilitating escaping difficulty, reward for preventing wheel slippage, and penalty for controlling the action.

5. A method for an off-road vehicle to get out of trouble as described in claim 1, characterized in that, S3 includes: S31: Construct an off-road vehicle escaping strategy network, wherein the off-road vehicle escaping strategy network includes an environment model, a strategy model, and a value model; S32: The environment model, policy model, and value model are learned and trained based on the observation space, action space, and reward function, respectively. S33: Construct the loss functions corresponding to the environment model, strategy model, and value model, and minimize the loss functions during the learning and training process to obtain the final off-road vehicle escaping strategy network.

6. A system utilizing the method for escaping a stuck off-road vehicle as described in any one of claims 1-5, characterized in that, include: The feature extraction module is used to acquire the state data and control action data of the off-road vehicle when the off-road vehicle gets stuck, and to extract features that can characterize traction from the state data and the control action data. The determination module is used to determine the corresponding observation space, action space, and reward function; The model training module is used to construct the off-road vehicle traction strategy network and learn and train the off-road vehicle traction strategy network based on the combination of observation space, action space and reward function. The processing module is used to acquire the status data and control action data of the actual off-road vehicle when it gets stuck, and input the status data and control action data into the off-road vehicle get-out strategy network for processing to obtain the corresponding get-out strategy.

Citation Information

Patent Citations

  • Vehicle quick escape method and system, medium and electronic equipment

    CN120270223A

  • Vehicular driving force controller

    JP2010081720A

  • Vehicle escape control method and apparatus, and storage medium

    WO2023000986A1

  • Intelligent multi-mode hybrid assembly and intelligent connected electric heavy truck

    WO2024022141A1