State-limited offline reinforcement learning control method for automatic driving scene
By employing a state-constrained offline reinforcement learning method, and utilizing multimodal sensor data and a reward model, the problems of low data efficiency and insufficient safety in autonomous driving are addressed, enabling efficient and safe driving in complex environments.
Patent Information
- Application Number
- CN202511022067.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-11
AI Technical Summary
Existing reinforcement learning methods face problems such as low data efficiency, insufficient safety, and poor generalization ability in autonomous driving, making it difficult to navigate safely and effectively in complex environments.
We employ a state-constrained offline reinforcement learning approach, which involves constructing a multimodal vehicle driving dataset, performing three-level data preprocessing, training forward and inverse dynamics models, training a reward model, and training a policy network based on an actor-critic framework. This approach is combined with sensors such as cameras, radar, lidar, and inertial measurement units to build a safe and efficient driving strategy.
It improves the generalization ability in new environments, ensuring the safety and efficiency of autonomous vehicles in complex and ever-changing traffic environments, and guides policy learning through a reward model that comprehensively considers safety factors.
Smart Images

Figure CN120930714A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a state-constrained offline reinforcement learning control method that integrates complex problems, uncertainties, and dynamic environments, and belongs to the fields of reinforcement learning and autonomous driving. Background Technology
[0002] With advancements in sensor technology, computing power, and artificial intelligence algorithms, autonomous driving technology has become a crucial development direction for the automotive industry. Autonomous vehicles need to be able to navigate safely and effectively under various traffic and weather conditions. Reinforcement learning, as a machine learning method, has been widely used in the decision-making systems of autonomous vehicles. It allows vehicles to learn optimal driving strategies through interaction with the environment. However, existing reinforcement learning methods face challenges in data efficiency, security, and generalization ability in practical applications. In the complex environment of autonomous driving, the vehicle's state (such as position, speed, and distance to surrounding objects) is strictly limited to ensure safety and compliance with traffic rules. At the same time, due to the high cost and risks of actual road testing, a method that can learn from historical data in an offline environment is needed.
[0003] Therefore, to ensure the safety of autonomous vehicles, it is necessary to develop offline reinforcement learning methods that can handle state constraints, learn optimal behavior in specific states, and comply with traffic rules and safety standards. Summary of the Invention
[0004] To address the shortcomings of the above-mentioned technologies, the present invention aims to provide a state-constrained offline reinforcement learning control method for autonomous driving scenarios. This method can effectively utilize offline data and, considering state constraints, learn a safe and efficient autonomous driving strategy. This solves the problems of low data efficiency, insufficient security, and poor generalization ability faced by existing reinforcement learning methods in autonomous driving applications. To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0005] 1. A state-constrained offline reinforcement learning control method for autonomous driving scenarios, comprising the following:
[0006] (1) Construction of multimodal vehicle driving dataset
[0007] Let the collected autonomous driving data be Where s t This represents the vehicle state at time t, with each state s... t The information composition based on multimodal sensor fusion technology, a t s represents the driver's action at time t. t+1 r represents the vehicle state at time t+1. t Indicates from state s t Take action at Transition to state s t+1 The reward received, c t Represented as state constraints, these reflect the limitations in the environment, such as security requirements.
[0008] (2) Three-level data preprocessing
[0009] The three-level data preprocessing workflow includes low-level cleaning, mid-level feature extraction, and high-level data augmentation. Low-level cleaning mainly involves denoising, aligning, normalizing, and imputing missing values in sensor-perceived data to ensure data quality and consistency. Mid-level feature extraction uses methods such as trajectory coding, road topology analysis, and object detection to construct the state representation required for reinforcement learning and establish a state constraint model to ensure safety. High-level data augmentation uses techniques such as trajectory perturbation, inverse data generation, and domain adaptation to improve the generalization ability of the data and construct a state-action-reward format that meets the requirements of offline reinforcement learning, ultimately optimizing the training effect of the reinforcement learning model.
[0010] (3) Forward dynamics model and inverse dynamics model train
[0011] Forward dynamics model Used to predict in state s t Perform action a t The next state s t+1 Inverse dynamics model Used to predict from state s t to state s t+1 The action taken a t The forward dynamics model loss is defined as The loss of the inverse dynamics model is defined as Stochastic gradient descent is used to update the model parameters ω1 and ω2 according to the loss function in order to minimize the loss;
[0012] (4) Reward Model train
[0013] Reward Model Used to estimate from state s t Transition to state s t+1 The reward, and the corresponding loss function is defined as Update the reward model parameter ω3 using an optimization algorithm;
[0014] (5) Policy network π based on actor-critic framework φ and Value Network Q θ train
[0015] Critics Network Q θ (s t ,s t+1 Estimate from state s t Transition to state s t+1 The value of actor network π φ (s t According to the current state s t Generate action a t Based on the forward dynamics model and the inverse dynamics model, state s is defined. t+1 Relative to state s t The reachability is s t+1 ∈SR M (s t ), if and only if (ε is a small positive number used to tolerate model prediction errors), the commentator network training loss function is defined as: Where, θ ' For the target commentator network parameters, θ is updated softly. ' ←τθ+(1-τ)θ ' (τ is the soft update coefficient) gradually approaches θ, and the actor network training loss function is defined as: in, λ is the regularization coefficient used to balance the exploration and exploitation of the policy. An optimization algorithm is used to update the policy network parameters φ and the value network parameters θ, respectively. For the policy network parameter update: For value network parameter updates:
[0016] (6) Online decision-making and control
[0017] During the operation of an autonomous vehicle, the current vehicle status s is acquired in real time. cur The actor network generates action a based on the current state. cur =π φ (s cur The vehicle performs action a cur The environment provides feedback on the next state s next and corresponding reward r cur New experiences (s) cur ,a cur ,s next ,r cur The data is stored in the experience replay buffer B, and a batch of data is periodically sampled from the experience replay buffer B. (m is the batch size). The above model network is updated using an online update mechanism based on temporal difference learning. The parameters of the deep neural network are iteratively optimized periodically through a priority experience replay strategy to achieve continuous evolution of the driving strategy, so as to adapt to environmental changes and optimize the driving strategy.
[0018] 2. The hardware device includes:
[0019] (1) Camera: responsible for capturing visual images around the vehicle, covering information such as road scenes, traffic signs, other vehicles and pedestrians, providing rich visual cues for environmental perception;
[0020] (2) Radar: By transmitting and receiving electromagnetic waves, it measures the distance and relative speed between the vehicle and surrounding objects, generates point cloud data, and helps to monitor the position and motion status of objects in real time.
[0021] (3) LiDAR: to acquire high-precision distance information and build a more accurate three-dimensional environment model;
[0022] (4) Inertial Measurement Unit: measures the vehicle’s acceleration and angular velocity in real time to determine the vehicle’s attitude and motion state.
[0023] 3. The reward for the aforementioned vehicle: R car =r collision +r rule +r speed +r path , where r collision Defined as a collision avoidance reward, a positive reward is given when the vehicle maintains a safe distance from other objects (including other vehicles, pedestrians, obstacles, etc.). Let the distance between the vehicle and the nearest object be d, and the safe distance threshold be d. safe Then r collision Defined as:
[0024]
[0025] Where k safe and k unsafe Positive numbers represent the reward magnitude under safe and unsafe conditions, respectively;
[0026] r rule To reward compliance with traffic rules, let the current speed of the vehicle be v, and the speed limit be v. lim If the traffic light status is a (green light = 1, red light = 0), and whether the vehicle is driving in the correct lane is b (correct = 1, incorrect = 0), then r rule It can be designed as: Where k rule A positive number indicates a reward coefficient for obeying traffic rules;
[0027] r speed To incentivize vehicles to maintain appropriate speeds while ensuring safety, thereby improving overall traffic efficiency, let the desired speed be v. desired Then r speed It can be designed as: k speed =-k speed ×|vv desired |, where k speed A positive number indicates a penalty coefficient for speeds deviating from the desired speed;
[0028] r path As a route planning reward, a bonus is given for choosing the shortest and smoothest travel route. Let the length of the vehicle's current travel route be L, and the length of the shortest path from the origin to the destination be L. min Then the path planning reward r path It can be designed as: Where k path A positive number indicates the reward coefficient for path planning.
[0029] 4. The vehicle status information s cur and action information a cur In autonomous driving scenarios, various onboard sensors can be used to acquire state information. Cameras are used to identify traffic signs, lane lines, and the position and movement of surrounding vehicles and pedestrians; radar and lidar can accurately measure the distance and relative speed to surrounding objects; and inertial measurement units provide the vehicle's own acceleration and angular velocity. These sensor data are processed and fused to form a comprehensive perception of the vehicle's state and environment, namely state information. Steering actions and acceleration together constitute the vehicle's motion information.
[0030] 5. The network structure described: Raw data such as vehicle state and surrounding environment collected by sensors are preprocessed and then flow into the network. First, the forward dynamics model network predicts the next state after the action is performed. The result is used to compare with the actual state and to transmit to other modules. The inverse dynamics model network receives the current and target state data to infer the action. The critic network combines the current state, the reachable state output by the forward dynamics model, and the reward information to evaluate the value and feeds it back to the actor network to guide it in generating the action. The actor network drives the vehicle based on the action generated by the current state, and the newly generated experience data is stored in the experience playback buffer. During training, data is sampled from the buffer and input into each network to update the parameters. This process is repeated to enable the network to continuously learn and optimize, ensuring the efficiency and safety of autonomous driving.
[0031] Compared with the prior art, the advantages of the present invention are as follows:
[0032] (1) Training with offline data enables the model to learn a variety of different driving scenarios and road conditions. These historical data include various weather conditions, different traffic flows and diverse road types. By learning from these rich data in the offline stage, the model can extract more general driving strategies and rules, thereby improving its generalization ability in new environments or unseen scenarios.
[0033] (2) The reward model takes into account safety factors to guide strategy learning. Unlike some existing technologies that may simply make decisions based on rules or a single goal, the reward model of this technology can provide positive incentives for safe driving behaviors. For example, behaviors such as maintaining a safe distance and obeying traffic rules can be rewarded, prompting the vehicle to learn safer driving strategies. This reward-based guidance method can better adapt to complex and ever-changing traffic environments and ensure the safety of autonomous driving from multiple perspectives. Attached Figure Description
[0034] Figure 1 This is a flowchart of the present invention;
[0035] Figure 2 This is a data flow diagram of the present invention. Detailed Implementation
[0036] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0037] This invention employs a state-constrained offline reinforcement learning method, focusing on autonomous vehicles. The invention aims to achieve safe and efficient driving decisions for vehicles through the collaborative work of multiple modules, such as... Figure 1 and Figure 2 As shown, the process described in this invention is as follows:
[0038] Step 1: Autonomous vehicles are equipped with various sensors, including cameras, radar, lidar, and inertial measurement units (IMUs). Cameras capture visual images of the vehicle's surroundings, covering road scenes, traffic signs, other vehicles, and pedestrians, providing rich visual cues for environmental perception. Their data format is defined as a two-dimensional matrix I(x,y) (where x and y represent the horizontal and vertical coordinates of image pixels, respectively). Radar measures the distance and relative velocity between the vehicle and surrounding objects by emitting and receiving electromagnetic waves, generating point cloud data, which can be represented as P. r ={p r1 ,p r2 ,…,p rn}, each point pri =(x ri ,y ri ,z ri ,v ri ) includes spatial coordinates (x ri ,y ri ,z ri ) and relative velocity v ri Information helps monitor the position and motion of objects in real time. LiDAR, by emitting laser beams and receiving reflected light, obtains high-precision distance information, enabling the construction of more accurate 3D environment models. The IMU measures the vehicle's acceleration a = (a... x ,a y ,a z ) and angular velocity ω=(ω x ,ω y ,ω z The vehicle's position information r(t) can be obtained by performing a double integral on the acceleration. For the initial position r0 and initial velocity v0, the velocity over the time interval [0,t] is: Location: Used to determine the vehicle's attitude and motion state. In addition, map data also serves as an important input, containing information such as road layout, traffic rules, and speed limit zones, providing a global reference for vehicle path planning and decision-making. Raw data acquired from various sensors typically contains noise, outliers, and different scales and formats, making it unsuitable for direct algorithm processing. Therefore, data preprocessing is necessary. For camera image data, convolutional neural networks (CNNs) are often used for feature extraction, converting the image into representative feature vectors through operations such as convolutional layers and pooling layers. Point cloud data from radar and lidar undergoes filtering to remove noise and identify different objects. For each point p in the point cloud data... i The filtered point p ' i can be represented as Where N is point p i The number of points in the neighborhood, w ij The weights are calculated based on the Gaussian function. IMU data is integrated to obtain the vehicle's position and attitude information. Simultaneously, all types of data undergo normalization to ensure a uniform scale and format, facilitating subsequent data fusion and model training.
[0039] Step 2: The forward dynamics network, using preprocessed and fused vehicle state information and current possible action inputs, predicts the state at the next moment through nonlinear transformation and feature extraction via a fully connected layer, aided by the vehicle dynamics model and learned parameters. The prediction result is compared with the actual state to calculate the error and update the parameters, and is also passed to the critic network to evaluate the value of the state transition. The inverse dynamics network receives current and target state information, with the target state set based on the planned path or an emergency. Based on dynamic characteristics and learned knowledge, it processes the input through a fully connected layer to infer the action to reach the target state. The inferred action is passed to the actor network to assist in decision-making, and the parameters are adjusted to improve accuracy by comparing the error with the actual action. The reward model integrates the current state, the predicted state from the forward dynamics network, and the action information from the actor network, processes it through a fully connected layer, and calculates the output reward value according to preset rules (such as obeying traffic regulations, maintaining a safe distance, and driving efficiently). This reward value reflects the "value" brought by the vehicle performing the action in the current state, and is used to guide the learning process of the critic network and the actor network, prompting the vehicle to learn driving strategies that are more in line with safety and efficiency requirements. The critic network takes the current vehicle state (from the data fusion module), the next reachable state predicted by the forward dynamics model network, and the reward value output by the reward model as input. Based on the Bellman equation, we can obtain... (Where γ is the discount factor), which represents the value of transitioning from the current state to the next state. This value assessment result represents an estimate of the long-term cumulative reward for taking a specific action to transition from the current state to the next state under the current policy. The critic network transmits this value assessment information to the actor network, providing an important basis for the actor network's policy update, while minimizing the mean squared error between the predicted value and the target value: (Where N is the number of samples), continuously updating its network parameters to optimize the accuracy of value assessment. The actor network receives the current vehicle state (from the data fusion module), the actions inferred by the inverse dynamics model network, and the value assessment information output by the critic network. Based on the policy gradient algorithm: (where π) θ (a|s) represents the policy distribution, and J(θ) represents the expected cumulative reward. These input data are processed through multiple fully connected layers to generate the driving action a in the current state. cur The specific acceleration and steering angle are determined. The generated driving actions are output to the vehicle execution module to control the actual driving of the vehicle. Simultaneously, the actor network continuously adjusts its network parameters and optimizes its strategy based on feedback from the critic network and policy gradient update rules, enabling the vehicle to obtain the maximum cumulative reward over long-term operation. Furthermore, the actor network generates experience data {(s)} during vehicle operation. t ,a t ,s t+1 ,rt )} i It will be stored in the experience replay buffer for subsequent network training.
[0040] Step 3: During actual driving, after each action is performed, the current state, the action performed, the next state obtained after the action, and the corresponding reward information are packaged into an experience data point and stored in the experience replay buffer. This buffer is similar to a memory bank, saving the vehicle's driving experience at different times and providing data support for subsequent network training.
[0041] Step 4: During training, a batch of empirical data is randomly sampled from the experience replay buffer. This sampled data is then fed into the forward dynamics model network, the inverse dynamics model network, the critic network, and the actor network, respectively. By training these networks, the forward dynamics model network can further optimize the accuracy of predicting the next state based on the sampled data; the inverse dynamics model network can improve the accuracy of action inference; the critic network can better fit the value evaluation function; and the actor network can improve the strategy and generate better driving actions. By continuously repeating the process of sampling data from the experience replay buffer and training the networks, each network gradually converges to better parameters, improving the overall performance of the algorithm in autonomous driving and enabling the vehicle to make more reasonable and safer decisions in various complex driving scenarios.
[0042] Step 5: The vehicle execution module receives driving action commands generated by the actor network, such as control signals for the accelerator pedal and steering motor. These commands are translated into actual vehicle operations, controlling the vehicle's speed, direction, and other states to enable the vehicle to move on the road. During the vehicle's execution of actions, sensors monitor the actual state changes of the vehicle in real time and feed the new state data back into the entire algorithm flow. This real-time feedback data serves as new input, participating in the next round of network computation and decision-making, enabling the algorithm to continuously adjust its driving strategy based on the vehicle's real-time state, ensuring that the vehicle maintains a safe and efficient driving state in dynamic traffic environments.
Claims
1. A state-constrained offline reinforcement learning control method for autonomous driving scenarios, comprising the following: (1) Construction of multimodal vehicle driving dataset Let the collected autonomous driving data be Where s t This represents the vehicle state at time t, with each state s... t The information composition based on multimodal sensor fusion technology, a t s represents the driver's action at time t. t+1 r represents the vehicle state at time t+1. t Indicates from state s t Take action a t Transition to state s t+1 The reward received, c t Represented as state constraints, these reflect the limitations in the environment, such as security requirements. (2) Three-level data preprocessing The three-level data preprocessing process includes low-level cleaning, intermediate-level feature extraction, and high-level data augmentation. Low-level cleaning mainly involves denoising, aligning, normalizing, and filling missing values in sensor-sensed data to ensure data quality and consistency. Intermediate feature extraction uses methods such as trajectory encoding, road topology analysis, and object detection to construct the state representation required for reinforcement learning and establish a state constraint model to ensure safety. Advanced data augmentation uses techniques such as trajectory perturbation, inverse data generation, and domain adaptation to improve the generalization ability of data and construct a state-action-reward format that meets the requirements of offline reinforcement learning, ultimately optimizing the training effect of the reinforcement learning model. (3) Forward dynamics model and inverse dynamics model train Forward dynamics model Used to predict in state s t Perform action a t The next state s t+1 Inverse dynamics model Used to predict from state s t to state s t+1 The action taken a t The forward dynamics model loss is defined as The loss of the inverse dynamics model is defined as Stochastic gradient descent is used to update the model parameters ω1 and ω2 according to the loss function in order to minimize the loss; (4) Reward Model train Reward Model Used to estimate from state s t Transition to state s t+1 The reward, and the corresponding loss function is defined as Update the reward model parameter ω3 using an optimization algorithm; (5) Policy network π based on actor-critic framework φ and Value Network Q θ train Critics Network Q θ (s t ,s t+1 Estimate from state s t Transition to state s t+1 The value of actor network π φ (s t According to the current state s t Generate action a t Based on the forward dynamics model and the inverse dynamics model, state s is defined. t+1 Relative to state s t The reachability is s t+1 ∈SR M (s t ), if and only if (ε is a small positive number used to tolerate model prediction errors), the commentator network training loss function is defined as: Where, θ ' For the target commentator network parameters, θ is updated softly. ' ←τθ+(1-τ)θ ' (τ is the soft update coefficient) gradually approaches θ, and the actor network training loss function is defined as: in, λ is the regularization coefficient used to balance the exploration and exploitation of the policy. An optimization algorithm is used to update the policy network parameters φ and the value network parameters θ, respectively. For the policy network parameter update: For value network parameter updates: (6) Online decision-making and control During the operation of an autonomous vehicle, the current vehicle status s is acquired in real time. cur The actor network generates action a based on the current state. cur =π φ (s cur The vehicle performs action a cur The environment provides feedback on the next state s next and corresponding reward r cur New experiences (s) cur ,a cur ,s next ,r cur The data is stored in the experience replay buffer B, and a batch of data is periodically sampled from the experience replay buffer B. (m is the batch size). The above model network is updated using an online update mechanism based on temporal difference learning. The parameters of the deep neural network are iteratively optimized periodically through a priority experience replay strategy to achieve continuous evolution of the driving strategy, so as to adapt to environmental changes and optimize the driving strategy.
2. The method according to claim 1, characterized in that, The hardware device includes: (1) Camera: responsible for capturing visual images around the vehicle, covering information such as road scenes, traffic signs, other vehicles and pedestrians, providing rich visual cues for environmental perception; (2) Radar: By transmitting and receiving electromagnetic waves, it measures the distance and relative speed between the vehicle and surrounding objects, generates point cloud data, and helps to monitor the position and motion status of objects in real time. (3) LiDAR: to acquire high-precision distance information and build a more accurate three-dimensional environment model; (4) Inertial Measurement Unit: measures the vehicle’s acceleration and angular velocity in real time to determine the vehicle’s attitude and motion state.
3. The method according to claim 1, characterized in that: The reward for the vehicle mentioned: R car =r collision +r rule +r speed +r path , where r collision Defined as a collision avoidance reward, a positive reward is given when the vehicle maintains a safe distance from other objects (including other vehicles, pedestrians, obstacles, etc.). Let the distance between the vehicle and the nearest object be d, and the safe distance threshold be d. safe Then r collsion Defined as: Where k safe and k unsafe Positive numbers represent the reward magnitude under safe and unsafe conditions, respectively; r rule To reward compliance with traffic rules, let the current speed of the vehicle be v, and the speed limit be v. lim If the traffic light status is a (green light = 1, red light = 0), and whether the vehicle is driving in the correct lane is b (correct = 1, incorrect = 0), then r rule It can be designed as: Where k rule A positive number indicates a reward coefficient for obeying traffic rules; r speed To incentivize vehicles to maintain appropriate speeds while ensuring safety, thereby improving overall traffic efficiency, let the desired speed be v. desired Then r speed It can be designed as: r speed =-k speed ×|vv desired |, where k speed A positive number indicates a penalty coefficient for speeds deviating from the desired speed; r path To reward route planning, a bonus is awarded for choosing the shortest and smoothest travel route. Let the length of the vehicle's current travel path be L, and the length of the shortest path from the starting point to the destination be L'. min The path planning reward r path It can be designed as: Where k path A positive number indicates the reward coefficient for path planning.
4. The method according to claim 1, wherein the vehicle status information s cur and action information a cur In autonomous driving scenarios, various onboard sensors can be used to acquire state information. Cameras are used to identify traffic signs, lane lines, and the position and movement of surrounding vehicles and pedestrians; radar and lidar can accurately measure distance and relative speed to surrounding objects; and inertial measurement units provide the vehicle's own acceleration and angular velocity. This sensor data is processed and fused to form a comprehensive perception of the vehicle's state and environment—that is, state information. Steering actions and acceleration together constitute the vehicle's motion information.
5. The method according to claim 1, wherein the network structure is as follows: the raw data such as vehicle state and surrounding environment collected by the sensor are preprocessed and then flow into the network. First, they enter the forward dynamics model network to predict the next state after the action is performed, and the result is used to compare with the actual state and transmit it to other modules. The inverse dynamics model network receives current and target state data to infer actions; the critic network combines the current state, reachable states output by the forward dynamics model, and reward information to evaluate the value and feeds it back to the actor network to guide its action generation. The actor network generates actions based on the current state to drive the vehicle, and the newly generated experience data is stored in the experience playback buffer. During training, data is sampled from the buffer and then input into each network to update the parameters. This process is repeated to enable the network to continuously learn and optimize, ensuring the efficiency and safety of autonomous driving.
Citation Information
Cited By
Deep learning-based electroencephalogram slow wave real-time feedback transcranial electrical stimulation method and system
CN121731667A