A safe envelope boundary-based intelligent driving self-supervised learning evolution method and system

CN122379575BActive Publication Date: 2026-09-29UNIV OF SCI & TECH OF CHINA +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610850859.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-12
Publication Date
2026-09-29
Estimated Expiration
2046-06-12

AI Technical Summary

Technical Problem

但该方法是强化学习通过设计奖励函数来间接地优化安全性,在训练过程中并不直接考虑安全包络边界,适合具有复杂交互行为的场景,可以有效处理多种动态和静态障碍,但对数据的依赖较大

Benefits of technology

1.本发明通过引入双重安全包络边界,能够精准控制车辆在驾驶过程中的行为,确保车辆始终在安全的行驶区域内。实时监控车辆的状态与运动轨迹,当车辆接近安全包络边界时,系统会自动调整决策和控制策略,从而避免危险发生。与传统方法相比,本发明具备更强的前瞻性和可靠性,能够有效减少碰撞、侧翻等风险,特别是在复杂环境下,如狭窄道路和复杂交叉口。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122379575B_ABST
    Figure CN122379575B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent driving and discloses an intelligent driving self-supervised learning evolution method and system based on a safety envelope boundary. The method comprises the following steps: designing a safety reward function according to the distance between the vehicle state and two safety envelope boundaries, designing a comfort reward function by using the acceleration change rate, and then obtaining a comprehensive reward; collecting driver behavior data and vehicle states, and constructing a positive and negative sample data set according to the comprehensive reward; based on a deep deterministic policy gradient algorithm, gradually optimizing a driving strategy through the interaction between an agent and an environment; and using a self-supervised learning mechanism to correct the deep deterministic policy gradient algorithm, adjusting the state input to the deep deterministic policy gradient algorithm, and adjusting the driving strategy. The application improves the intelligence, safety and comfort level of the driving system, enables the driving strategy to self-adjust in a complex environment, and improves the stability and convergence speed of training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent driving technology, specifically to an intelligent driving self-supervised learning evolution method and system based on a safety envelope boundary. Background Technology

[0002] In the field of intelligent driving technology, self-supervised learning has broad application prospects. By generating targets from unlabeled data, intelligent driving systems can not only reduce their reliance on manually labeled data but also improve their perception, decision-making, and control capabilities. Self-supervised learning helps systems adapt to complex and changing driving environments, enhancing their safety, robustness, and generalization capabilities, making it a crucial component of future intelligent driving technology. With the continuous development of algorithms and hardware, self-supervised learning will play an increasingly important role in intelligent driving, propelling autonomous driving to higher levels.

[0003] In past research, some scholars have proposed using reinforcement learning (RL) frameworks, where intelligent driving systems learn decision-making strategies through interaction with the environment. However, this method indirectly optimizes safety by designing reward functions, without directly considering the safety envelope boundary during training. It is suitable for scenarios with complex interactive behaviors and can effectively handle various dynamic and static obstacles, but it is highly dependent on data.

[0004] In contrast, the self-supervised learning evolutionary method for intelligent driving based on safety envelope boundaries aims to generate a continuous sequence of behaviors that conforms to human-like driving logic. This method emphasizes generating a series of reasonable driving behaviors in continuous time steps through autonomous learning, avoiding the problem of traditional methods making discrete decisions only at key decision points. The core of this method is to use safety envelope boundaries to constrain the learning process, ensuring that the autonomous driving system always follows safety rules in complex and ever-changing driving scenarios, and gradually optimizes its decision-making strategy based on this.

[0005] By using a safe envelope boundary, the agent can not only effectively avoid potential dangers but also continuously optimize its behavioral sequences in various driving environments, resulting in output strategies with high stability and adaptability. Ultimately, this method can generate continuous action sequences throughout the entire driving task, enhancing the autonomous driving system's decision-making capabilities in complex environments and improving the system's understanding of safety and human driving logic.

[0006] Chinese patent document CN110745136A discloses a driving adaptive control method. The method includes acquiring a historical driving dataset and dividing it into training, testing, and validation sets; constructing a network model for driving control using a deep reinforcement learning algorithm based on a deep convolutional neural network; training the network model using the training set data and iteratively training the network model using the gradient of the cost function to obtain an optimized network model; validating the performance of the optimized network model using the testing and validation sets, and using the network model that meets the performance requirements as the adaptive decision model; and processing the currently collected real-time environmental data using the adaptive decision model to make driving decisions. This application, however, focuses on a self-supervised learning evolutionary method for intelligent driving based on a safe envelope boundary, which is fundamentally different from the technical problem addressed by the patent document. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention provides a self-supervised learning evolutionary method and system for intelligent driving based on a safe envelope boundary.

[0008] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: In a first aspect, the present invention provides an intelligent driving self-supervised learning evolutionary method based on a safe envelope boundary, comprising: Based on the dynamic state parameters, a dynamic safety envelope model of the vehicle is constructed using the lateral force. A quasi-Monte Carlo filtering algorithm is used to predict the future motion trajectory of other vehicles to construct a kinematic safety envelope model of the vehicle. A safety reward function is designed based on the distance between the vehicle state and the safety envelope boundary output by the dynamic and kinematic safety envelope models. A comfort reward function is designed using the rate of change of acceleration, and then a comprehensive reward is obtained. Collect driver behavior data and vehicle status, and construct a positive and negative sample dataset based on the comprehensive reward; Based on the deep deterministic policy gradient algorithm, the driving strategy is gradually optimized through the interaction between the agent and the environment; Based on positive and negative sample datasets, a self-supervised learning mechanism is used to correct the deep deterministic policy gradient algorithm and adjust the driving strategy.

[0009] In one embodiment, the step of constructing a dynamic safety envelope model of the vehicle based on lateral force using dynamic state parameters, and using a quasi-Monte Carlo filtering algorithm to predict the future trajectories of other vehicles to construct a kinematic safety envelope model of the vehicle, specifically includes: Vehicle dynamics parameters, including speed, are collected using onboard sensors. acceleration Side slip angle Lateral force Based on vehicle dynamics characteristics, we can obtain ,in, It is the lateral stiffness coefficient; Maximum stable driving speed for: ; Where m is the vehicle mass. It is the coefficient of friction between the tire and the road surface; using dynamic state parameters and m A dynamic safety envelope model is constructed under different working conditions. The dynamic safety envelope model is used to characterize the vehicle's stable driving feasible region that satisfies the tire lateral adhesion constraint, sideslip angle constraint, yaw stability constraint and maximum stable driving speed constraint. The dynamic safety envelope boundary obtained from the dynamic safety envelope model is the boundary of the vehicle's stable driving feasible region. The quasi-Monte Carlo filtering algorithm is used to predict the future vehicle states of other vehicles, and the future trajectory is formed from the continuous vehicle states:

[0010] in, For vehicle kinematics model, To control the input, It is Gaussian noise. This represents the vehicle state at time t. This represents the step size for predicting the motion trajectory; based on the predicted future motion trajectory, a kinematic safety envelope model of the vehicle is constructed.

[0011] In one embodiment, the step of designing a safety reward function based on the distance to the safety envelope boundary output by the vehicle state and dynamic safety envelope model and the kinematic safety envelope model, and designing a comfort reward function using the rate of change of acceleration, thereby obtaining a comprehensive reward, specifically includes: Based on the distance between the vehicle state and the dynamic safety envelope boundary output by the dynamic safety envelope model, and the distance between the vehicle and the kinematic safety envelope boundary output by the kinematic safety envelope model, a safety reward function is designed. : ; in, The distance from the current vehicle state to the boundary of the dynamic safety envelope. This is the distance from the predicted trajectory of this vehicle to the boundary of the kinematic safety envelope; The structured vehicle state at time t is represented. Indicates the vehicle control action at time t; Represents the dynamic safety weight, Indicates kinematic safety weights. This represents the distance normalization scaling parameter. This indicates the penalty when the vehicle's state exceeds any safety envelope boundary; Comfort reward function ; in, The rate of change of acceleration; The maximum permissible rate of change of acceleration; Weighting for comfort within the interior; Comprehensive Rewards , and These are the combined weights of the safety reward function and the comfort reward function, respectively.

[0012] In one embodiment, the collection of driver behavior data and vehicle status, and the construction of a positive and negative sample dataset based on the comprehensive reward, specifically includes: Collect driver driving data and vehicle status to form state-action pairs. ,in, For a moment Vehicle status; The vehicle control actions at time t are extracted from the driver's driving data, including steering wheel angle, accelerator pedal opening, and brake pedal opening; the comprehensive reward is calculated for the state-action pairs at each time point. When the total reward of a state-action pair is greater than a set positive sample threshold, the state-action pair is marked as a positive sample; when the total reward of a state-action pair is less than a set negative sample threshold, the state-action pair is marked as a negative sample, thus obtaining a positive and negative sample dataset. The deep deterministic policy gradient algorithm, through the interaction between the agent and the environment, gradually optimizes the driving strategy, specifically including: Initialize the experience replay pool; Initialize the parameters of the value network, target value network, policy network, and target policy network in the deep deterministic policy gradient algorithm; Get the current vehicle status ,action Rewards and vehicle status at the next moment , forming data pairs And store it in the experience replay pool; In vehicle status Take action below Then, the comprehensive reward was calculated. The rewards received; By sampling from the experience replay pool, the parameters of the value network and policy network are updated, and the driving strategy is gradually optimized.

[0013] In one embodiment, the step of using a self-supervised learning mechanism to correct the deep deterministic policy gradient algorithm and adjust the driving strategy based on positive and negative sample datasets specifically includes: A convolutional neural network is used as a feature encoder to process the input sensor data, which is environmental perception data from cameras, LiDAR, and GPS positioning modules; the sensor data is then processed. Through the convolutional neural network Mapping to a high-dimensional embedding space, outputting embedding features ,in Given the parameters of a convolutional neural network; design a contrastive loss function. The embedded features are optimized by comparing the similarity between positive and negative sample pairs. ; in, For anchor sample embedding features, The embedding features are for positive sample pairs. Embedding features for negative sample pairs; express and The similarity between them is calculated, where N represents the number of negative samples in the same batch; the parameters of the convolutional neural network are updated using the backpropagation algorithm to minimize the contrastive loss function. ; in, For learning rate, To compare the loss functions The gradient; the optimized embedded features Vehicle status The concatenation forms the enhanced state that is input to the deep deterministic policy gradient algorithm. Update the value network and policy network, and adjust the driving strategy.

[0014] Secondly, the present invention provides an intelligent driving self-supervised learning evolutionary system based on a safe envelope boundary, comprising: The comprehensive reward calculation module constructs a dynamic safety envelope model of the vehicle based on the dynamic state parameters and the lateral force. It uses a quasi-Monte Carlo filtering algorithm to predict the future motion trajectory of other vehicles to construct a kinematic safety envelope model of the vehicle. It designs a safety reward function based on the distance between the vehicle state and the safety envelope boundary output by the dynamic and kinematic safety envelope models. It designs a comfort reward function using the rate of change of acceleration, and then obtains the comprehensive reward. The sample construction module collects driver behavior data and vehicle status, and constructs positive and negative sample datasets based on the comprehensive reward. The training module, based on the deep deterministic policy gradient algorithm, gradually optimizes the driving strategy through the interaction between the agent and the environment. The self-supervised learning module, based on positive and negative sample datasets, uses a self-supervised learning mechanism to correct the deep deterministic policy gradient algorithm and adjust the driving strategy.

[0015] In one embodiment, the step of constructing a dynamic safety envelope model of the vehicle based on lateral force using dynamic state parameters, and using a quasi-Monte Carlo filtering algorithm to predict the future trajectories of other vehicles to construct a kinematic safety envelope model of the vehicle, specifically includes: Vehicle dynamics parameters, including speed, are collected using onboard sensors. acceleration Side slip angle Lateral force Based on vehicle dynamics characteristics, we can obtain ,in, It is the lateral stiffness coefficient; Maximum stable driving speed for: ; Where m is the vehicle mass. It is the coefficient of friction between the tire and the road surface; using dynamic state parameters and m A dynamic safety envelope model is constructed under different working conditions. The dynamic safety envelope model is used to characterize the vehicle's stable driving feasible region that satisfies the tire lateral adhesion constraint, sideslip angle constraint, yaw stability constraint and maximum stable driving speed constraint. The dynamic safety envelope boundary obtained from the dynamic safety envelope model is the boundary of the vehicle's stable driving feasible region. The quasi-Monte Carlo filtering algorithm is used to predict the future vehicle states of other vehicles, and the future trajectory is formed from the continuous vehicle states:

[0016] in, For vehicle kinematics model, To control the input, It is Gaussian noise. This represents the vehicle state at time t. This represents the step size for predicting the motion trajectory; based on the predicted future motion trajectory, a kinematic safety envelope model of the vehicle is constructed.

[0017] In one embodiment, the step of designing a safety reward function based on the distance to the safety envelope boundary output by the vehicle state and dynamic safety envelope model and the kinematic safety envelope model, and designing a comfort reward function using the rate of change of acceleration, thereby obtaining a comprehensive reward, specifically includes: Based on the distance between the vehicle state and the dynamic safety envelope boundary output by the dynamic safety envelope model, and the distance between the vehicle and the kinematic safety envelope boundary output by the kinematic safety envelope model, a safety reward function is designed. : ; in, The distance from the current vehicle state to the boundary of the dynamic safety envelope. This is the distance from the predicted trajectory of this vehicle to the boundary of the kinematic safety envelope; The structured vehicle state at time t is represented. Indicates the vehicle control action at time t; Represents the dynamic safety weight, Indicates kinematic safety weights. This represents the distance normalization scaling parameter. This indicates the penalty when the vehicle's state exceeds any safety envelope boundary; Comfort reward function ; in, The rate of change of acceleration; It is the maximum permissible rate of change of acceleration; Weighting for comfort within the interior; Comprehensive Rewards , and These are the combined weights of the safety reward function and the comfort reward function, respectively.

[0018] In one embodiment, the collection of driver behavior data and vehicle status, and the construction of a positive and negative sample dataset based on the comprehensive reward, specifically includes: Collect driver driving data and vehicle status to form state-action pairs. ,in, For a moment Vehicle status; The vehicle control actions at time t are extracted from the driver's driving data, including steering wheel angle, accelerator pedal opening, and brake pedal opening; the comprehensive reward is calculated for the state-action pairs at each time point. When the total reward of a state-action pair is greater than a set positive sample threshold, the state-action pair is marked as a positive sample; when the total reward of a state-action pair is less than a set negative sample threshold, the state-action pair is marked as a negative sample, thus obtaining a positive and negative sample dataset. The deep deterministic policy gradient algorithm, through the interaction between the agent and the environment, gradually optimizes the driving strategy, specifically including: Initialize the experience replay pool; Initialize the parameters of the value network, target value network, policy network, and target policy network in the deep deterministic policy gradient algorithm; Get the current vehicle status ,action Rewards and vehicle status at the next moment , forming data pairs And store it in the experience replay pool; In vehicle status Take action below Then, the comprehensive reward was calculated. The rewards received; By sampling from the experience replay pool, the parameters of the value network and policy network are updated, and the driving strategy is gradually optimized.

[0019] In one embodiment, the step of using a self-supervised learning mechanism to correct the deep deterministic policy gradient algorithm and adjust the driving strategy based on positive and negative sample datasets specifically includes: A convolutional neural network is used as a feature encoder to process the input sensor data, which is environmental perception data from cameras, LiDAR, and GPS positioning modules; the sensor data is then processed. Through the convolutional neural network Mapping to a high-dimensional embedding space, outputting embedding features ,in Given the parameters of a convolutional neural network; design a contrastive loss function. The embedded features are optimized by comparing the similarity between positive and negative sample pairs. ; in, For anchor sample embedding features, The embedding features are for positive sample pairs. Embedding features for negative sample pairs; express and The similarity between them is calculated, where N represents the number of negative samples in the same batch; the parameters of the convolutional neural network are updated using the backpropagation algorithm to minimize the contrastive loss function. ; in, For learning rate, To compare the loss functions The gradient; the optimized embedded features Vehicle status The concatenation forms the enhanced state that is input to the deep deterministic policy gradient algorithm. Update the value network and policy network, and adjust the driving strategy.

[0020] The system and method in this invention correspond to each other; the specific technical solutions applicable to the method are also applicable to the system.

[0021] Compared with the prior art, the beneficial technical effects of the present invention are: 1. This invention, by introducing a dual safety envelope boundary, can precisely control the vehicle's behavior during driving, ensuring that the vehicle always remains within a safe driving area. It monitors the vehicle's status and trajectory in real time, and when the vehicle approaches the safety envelope boundary, the system automatically adjusts its decision-making and control strategies to avoid danger. Compared to traditional methods, this invention has stronger foresight and reliability, effectively reducing the risks of collisions and rollovers, especially in complex environments such as narrow roads and complex intersections.

[0022] 2. This invention utilizes the contrastive learning method within self-supervised learning to automatically extract environmental feature representations from a large amount of unlabeled sensor data, eliminating the need for manual annotation. Through self-supervised learning, the intelligent driving system can better understand and perceive various factors in the dynamic driving environment, including other vehicles, pedestrians, and traffic signs. This enables the system to make more accurate decisions in various driving environments, improving overall driving performance.

[0023] 3. This invention optimizes driving strategies by designing the rate of change of acceleration as a comfort reward function, reducing sudden acceleration or deceleration during vehicle operation, thereby providing a smoother and more comfortable driving experience. Especially in complex traffic situations, the system can ensure safety while smoothing driving behavior as much as possible, improving passenger comfort.

[0024] 4. This invention combines reinforcement learning and self-supervised learning, enabling intelligent driving systems to learn from environmental feedback and progressively optimize their driving strategies. Self-supervised learning provides a deep understanding of the environment, while reinforcement learning continuously adjusts driving behavior through reward mechanisms. This method allows the system to adapt to different driving environments over time, especially in complex or uncertain scenarios, by dynamically adjusting decisions to improve the system's intelligence and efficiency.

[0025] 5. The self-supervised learning method of this invention enables intelligent driving systems to automatically extract features from sensor data without requiring extensive manual annotation. This unsupervised learning approach significantly reduces the cost and workload of data annotation and helps the system adapt more flexibly to new environments or driving scenarios. For example, when encountering a new environment or an unseen driving situation, the system can directly extract and optimize relevant environmental features through comparative learning without re-annotating the data, thereby improving data processing efficiency.

[0026] 6. This invention designs a comprehensive reward system that combines safety and comfort, simultaneously optimizing both the safety and comfort of driving decisions. This method addresses the shortcomings of traditional methods that may only focus on one aspect, providing a more comprehensive evaluation and decision-making standard. For example, during emergency braking, this invention not only considers safety but also assesses the impact of the operation on passenger comfort, minimizing discomfort while ensuring safety, and improving the balance of driving behavior and the driving experience.

[0027] Other beneficial effects of the present invention will be explained in detail through the introduction of specific technical features and technical solutions in specific embodiments. Those skilled in the art should be able to understand the beneficial technical effects brought about by these technical features and technical solutions through the introduction of these technical features and technical solutions. Attached Figure Description

[0028] Figure 1 This is a flowchart of the method of the present invention.

[0029] Figure 2 This is the network structure used in the deep deterministic strategy gradient algorithm of this invention.

[0030] Figure 3 This is a technical roadmap for the present invention. Detailed Implementation

[0031] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.

[0032] like Figure 1 As shown, an intelligent driving self-supervised learning evolution method based on safe envelope boundaries in this invention includes the following steps: S1. Based on the dynamic state parameters, a dynamic safety envelope model of the vehicle is constructed using the lateral force. The quasi-Monte Carlo filtering algorithm is used to predict the future motion trajectory of other vehicles to construct a kinematic safety envelope model of the vehicle. A safety reward function is designed based on the distance between the vehicle state and the safety envelope boundary output by the dynamic and kinematic safety envelope models. A comfort reward function is designed using the rate of change of acceleration, and then a comprehensive reward is obtained. S2, collect driver behavior data and vehicle status, and construct a positive and negative sample dataset based on the comprehensive reward; S3, based on a deep deterministic policy gradient algorithm, gradually optimizes the driving strategy through the interaction between the agent and the environment; S4, based on positive and negative sample datasets, uses a self-supervised learning mechanism to correct the deep deterministic policy gradient algorithm, and adjusts the driving strategy by inputting the state into the deep deterministic policy gradient algorithm.

[0033] The present invention will be described in detail below in several parts.

[0034] 1. Comprehensive rewards.

[0035] Key dynamic parameters of the vehicle, including speed, are collected using onboard sensors. acceleration Side slip angle Lateral force Based on vehicle dynamics characteristics, we can obtain ,in, It is the lateral stiffness coefficient.

[0036] Maximum stable driving speed It can be represented as: ; Where m is the vehicle mass. This refers to the coefficient of friction between the tire and the road surface. A dynamic safety envelope model is constructed under different operating conditions using velocity, acceleration, slip angle, lateral force, lateral stiffness coefficient, vehicle mass, and the coefficient of friction between the tire and the road surface. The dynamic safety envelope model characterizes the feasible region of stable vehicle driving that satisfies constraints on tire lateral adhesion, slip angle, yaw stability, and maximum stable speed. The boundary of the dynamic safety envelope is the boundary of this feasible region of stable vehicle driving.

[0037] The quasi-Monte Carlo filtering algorithm is used to predict the future vehicle states of other vehicles, and the future trajectory is formed from the continuous vehicle states. The prediction process includes the following formulas: ; in, For vehicle kinematics model, To control the input, It is Gaussian noise. This represents the vehicle state at time t. The time step is defined as the state prediction time. Based on the predicted trajectory, a kinematic safety envelope model of the vehicle is constructed to ensure that the vehicle maintains a safe distance from other vehicles.

[0038] Based on the distance between the vehicle state and the dynamic safety envelope boundary output by the dynamic safety envelope model, and the distance between the vehicle and the kinematic safety envelope boundary output by the kinematic safety envelope model, a safety reward function is designed. : ; in, The distance from the current vehicle state to the boundary of the dynamic safety envelope. This is the distance from the predicted trajectory of this vehicle to the boundary of the kinematic safety envelope; The structured vehicle state at time t is represented. Indicates the vehicle control action at time t; This is a penalty term when the vehicle state exceeds any safety envelope boundary. When the predicted state or trajectory of the vehicle approaches any safety envelope boundary, the safety reward is reduced to guide the system to adjust vehicle behavior and move away from the boundary; when the vehicle moves away from the safety envelope boundary, it indicates that the vehicle is within the safe feasible region. The predicted trajectory of the vehicle is obtained by forward rolling of the vehicle kinematic model based on the vehicle state and the vehicle control actions.

[0039] The vehicle's control actions are extracted from the driver's driving data, including steering wheel angle, accelerator pedal opening, brake pedal opening, or their equivalent steering angle, longitudinal acceleration, and braking deceleration.

[0040] The rate of change of acceleration can measure driving comfort. Large changes in acceleration lead to an uncomfortable driving experience and may result in sudden braking or acceleration. Design a comfort reward function: ; in, The rate of change of acceleration; The maximum permissible rate of change of acceleration; The comfort-oriented internal weight is used to penalize driving actions with excessive acceleration changes. Comprehensive Rewards ; and These are the combined weights of the safety reward function and the comfort reward function, respectively. and These are the combined weights of the safety reward function and the comfort reward function, respectively.

[0041] It's important to note here that the weights of the security reward function are typically set. Weights greater than the comfort reward function Because safety usually takes precedence over comfort.

[0042] 2. Construct a dataset of positive and negative samples.

[0043] Using sensors such as radar to collect driver driving data and vehicle status over a period of time, a time series is constructed. ,in, It represents the vehicle state at time t. It refers to the driver's control actions at that moment.

[0044] For each time t, the state-action pair Calculate the overall reward When the combined reward of a state-action pair exceeds a set positive sample threshold, the state-action pair is marked as a positive sample, indicating that the behavior meets safety and comfort requirements. When the combined reward of a state-action pair is less than a set negative sample threshold, the state-action pair is marked as a negative sample, indicating that the behavior does not meet safety and comfort requirements, thus constructing a positive and negative sample dataset. This positive and negative sample dataset is used for constructing positive and negative sample pairs in subsequent self-supervised contrastive learning.

[0045] The driver's driving data includes steering wheel angle, accelerator pedal opening, brake pedal opening, gear shifting status, turn signal status, and driver takeover or intervention behavior; the vehicle status includes vehicle position, speed, acceleration, heading angle, yaw rate, sideslip angle, lateral acceleration, sideslip force, relative position with lane lines, relative distance with surrounding obstacles, and relative speed.

[0046] 3. Train a deep deterministic policy gradient (DDPG) control model.

[0047] like Figure 2 As shown, in this embodiment, the reinforcement learning layer uses the DDPG algorithm, a typical actor-critic type algorithm, whose training process includes: Initialize the network parameters in the DDPG control model, which includes the policy network. Target policy network Value Network and target value network , This represents the vehicle state after fusing the vehicle state with the optimized embedded features. This indicates the vehicle's control actions. Among them... and These are the parameters of the policy network and the target policy network, respectively. and These are the parameters of the value network and the target value network, respectively. The vehicle's state, agent actions, environmental feedback and rewards, and the vehicle's state at the next moment are acquired within a preset time interval using an onboard platform and sensors, forming a data pair. And store it in the experience replay pool, where . The vehicle state at time t is the vehicle state after fusing the vehicle state with the optimized embedded features.

[0048] A training dataset is constructed by sampling from the experience replay pool, and the target action value estimate is calculated based on the target policy network and the target value network. The estimated value of the target action is used to train the current value network; the autonomous vehicle updates the value network Q every M steps, and the value network parameters are updated every C steps. Synchronize to target value network parameters Among them, the estimated value of target action Represented as: ; in, It is in the vehicle status Take action below The reward received later; It is a discount factor used to measure the importance of future rewards; Represents the target value network. Represents the target policy network. This indicates that a target value network and a target policy network are used to estimate the next state action, and the target value network is used to estimate the value of the next state action. When it is in a terminated state, .

[0049] Updating value network parameters by minimizing mean squared error (MSE) The loss function is: ; Where B is the batch size of data sampled from the experience playback pool; It is the estimated value of the target action corresponding to the i-th empirical sample; This indicates the current value network in terms of parameters. Below, regarding the state Take action Estimate the long-term cumulative rewards.

[0050] After updating the value network, stochastic gradient descent is used to update the policy network parameters, so that the policy network outputs driving actions that can obtain higher overall rewards, thereby optimizing the policy.

[0051] This section is responsible for optimizing the vehicle's driving strategy. Overall Rewards As part of experience replay data Value estimate of direct involvement in target actions The calculation shows that the value network learns the long-term cumulative reward determined by both safety and comfort, while the policy network indirectly maximizes the overall reward by maximizing the output of the value network, thereby learning safe and stable driving strategies in the continuous action space.

[0052] 4. Self-monitored learning.

[0053] First, a convolutional neural network (CNN) is used as a feature encoder to process the input sensor observations. Sensor observation This includes raw or pre-processed environmental perception data from cameras, LiDAR, millimeter-wave radar, GPS positioning modules, and inertial measurement units, such as image or video frames, point clouds, obstacle distances, relative speeds, vehicle positioning attitudes, lane markings, and traffic sign information. To avoid symbolic confusion, in this invention... This indicates raw or pre-processed sensor observations. Represents the structured vehicle state. This represents the embedded features after CNN encoding.

[0054] Sensor data By mapping to a high-dimensional embedding space through a convolutional neural network, the output of the CNN is represented as: ; in, The embedding features at time t are represented. This represents the parameters of the CNN network.

[0055] The contrastive loss function is designed to optimize features by comparing the similarity between positive and negative sample pairs. The contrastive loss function is as follows: ; in, For anchor sample embedding features, The embedding features are for positive sample pairs. Embedding features for negative sample pairs; express and The similarity between them.

[0056] Positive sample pairs can be composed of two data augmentation views of the same sample, state-action pairs that are both marked as positive samples in adjacent time periods, or two positive samples under the same safety condition; after inputting the above samples into a convolutional neural network, the corresponding positive sample embedding features are output.

[0057] Negative sample pairs can consist of one positive sample and one negative sample, or two samples with significantly different safety conditions; the negative sample embedding features are output through a convolutional neural network.

[0058] Updating the parameters of a convolutional neural network using the backpropagation algorithm To minimize the contrastive loss function.

[0059] ; in It's the learning rate. It is a contrastive loss function CNN encoding parameters The gradient.

[0060] The optimized embedding features, learned through comparative learning, will serve as the state input for the deep deterministic policy gradient algorithm. It is used for updating the value network and policy network.

[0061] The overall technical approach of this invention is as follows: Figure 3 As shown, this invention extracts and encodes environmental features using a contrastive learning method, and utilizes a convolutional neural network to process data from vehicle sensors to generate high-dimensional embedded feature representations. By employing a contrastive loss function, the feature space is continuously optimized, causing features from similar environments to cluster together, while features from different environments maintain a greater distance. These embedded features, trained through self-supervised learning, are used as state inputs to the DDPG algorithm, helping the agent better understand complex driving environments, thereby improving the accuracy and stability of decision-making. The core objective of this part is to provide efficient environmental perception capabilities, offering more accurate input data for the strategy optimization of intelligent driving systems.

[0062] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0063] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple steps or stages, which are not necessarily completed at the same time, but may be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0064] Based on the description of the above method embodiments, the present invention also provides a system. The system may be a system that uses software (applications), modules, components, servers, clients, etc., using the methods described in the embodiments of this specification, combined with necessary implementation hardware. Since the implementation schemes and methods for solving the problem are similar, the specific system implementations in the embodiments of this specification can be found in the implementations of the foregoing methods, and repeated details will not be described again. Although the system is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0065] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0066] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.

[0067] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A self-supervised learning evolutionary method for intelligent driving based on a safe envelope boundary, characterized in that, include: A dynamic safety envelope model of the vehicle is constructed using lateral force based on dynamic state parameters. A quasi-Monte Carlo filtering algorithm is used to predict the future trajectories of other vehicles to construct a kinematic safety envelope model for the vehicle. A safety reward function is designed based on the distance between the vehicle state and the safety envelope boundaries output by both the dynamic and kinematic safety envelope models. A comfort reward function is designed using the rate of change of acceleration, resulting in a comprehensive reward. The safety reward function is designed based on the distance between the vehicle state and the dynamic safety envelope boundary output by the dynamic safety envelope model, and the distance between the vehicle and the kinematic safety envelope boundary output by the kinematic safety envelope model. : ; The distance from the current vehicle state to the boundary of the dynamic safety envelope. This is the distance from the predicted trajectory of this vehicle to the boundary of the kinematic safety envelope; The structured vehicle state at time t is represented. Indicates the vehicle control action at time t; Represents the dynamic safety weight, Indicates kinematic safety weights. This represents the distance normalization scaling parameter. This represents the penalty term when the vehicle state exceeds any safety envelope boundary; comfort reward function. ; The rate of change of acceleration; The maximum permissible rate of change of acceleration; Weighting of comfort within the interior; overall reward , and These are the combined weights of the safety reward function and the comfort reward function, respectively. Collect driver behavior data and vehicle status, and construct a positive and negative sample dataset based on the comprehensive reward; Based on the deep deterministic policy gradient algorithm, the driving strategy is gradually optimized through the interaction between the agent and the environment; Based on positive and negative sample datasets, a self-supervised learning mechanism is used to correct the deep deterministic policy gradient algorithm and adjust the driving strategy.

2. The self-supervised learning evolutionary method for intelligent driving based on safe envelope boundaries as described in claim 1, characterized in that, The collection of driver behavior data and vehicle status, and the construction of a positive and negative sample dataset based on the comprehensive reward, specifically include: Collect driver driving data and vehicle status to form state-action pairs. ,in, For a moment Vehicle status; The vehicle control actions at time t are extracted from the driver's driving data, including steering wheel angle, accelerator pedal opening, and brake pedal opening; the comprehensive reward is calculated for the state-action pairs at each time point. When the total reward of a state-action pair is greater than a set positive sample threshold, the state-action pair is marked as a positive sample; when the total reward of a state-action pair is less than a set negative sample threshold, the state-action pair is marked as a negative sample, thus obtaining a positive and negative sample dataset. The deep deterministic policy gradient algorithm, through the interaction between the agent and the environment, gradually optimizes the driving strategy, specifically including: Initialize the experience replay pool; Initialize the parameters of the value network, target value network, policy network, and target policy network in the deep deterministic policy gradient algorithm; Get the current vehicle status ,action Rewards and vehicle status at the next moment , forming data pairs And store it in the experience replay pool; In vehicle status Take action below Then, the comprehensive reward was calculated. The rewards received; By sampling from the experience replay pool, the parameters of the value network and policy network are updated, and the driving strategy is gradually optimized.

3. The self-supervised learning evolutionary method for intelligent driving based on safe envelope boundaries according to claim 2, characterized in that, The process of correcting the deep deterministic policy gradient algorithm and adjusting the driving strategy based on positive and negative sample datasets using a self-supervised learning mechanism specifically includes: A convolutional neural network is used as a feature encoder to process the input sensor data, which is environmental perception data from cameras, LiDAR, and GPS positioning modules; the sensor data is then processed. Through the convolutional neural network Mapping to a high-dimensional embedding space, outputting embedding features ,in Given the parameters of a convolutional neural network; design a contrastive loss function. The embedded features are optimized by comparing the similarity between positive and negative sample pairs. ; in, For anchor sample embedding features, The embedding features are for positive sample pairs. Embedding features for negative sample pairs; express and The similarity between them is calculated, where N represents the number of negative samples in the same batch; the parameters of the convolutional neural network are updated using the backpropagation algorithm to minimize the contrastive loss function. ; in, For learning rate, To compare the loss functions The gradient; the optimized embedded features Vehicle status The concatenation forms the enhanced state that is input to the deep deterministic policy gradient algorithm. Update the value network and policy network, and adjust the driving strategy.

4. A self-supervised learning evolutionary system for intelligent driving based on a safe envelope boundary, characterized in that, include: The comprehensive reward calculation module constructs a dynamic safety envelope model of the vehicle based on dynamic state parameters and lateral force. It then uses a quasi-Monte Carlo filtering algorithm to predict the future trajectories of other vehicles to construct a kinematic safety envelope model for the vehicle. A safety reward function is designed based on the distance between the vehicle state and the safety envelope boundaries output by both the dynamic and kinematic safety envelope models. A comfort reward function is designed using the rate of change of acceleration, thus obtaining the comprehensive reward. The safety reward function is designed based on the distance between the vehicle state and the dynamic safety envelope boundary output by the dynamic safety envelope model, as well as the distance between the vehicle and the kinematic safety envelope boundary output by the kinematic safety envelope model. : ; The distance from the current vehicle state to the boundary of the dynamic safety envelope. This is the distance from the predicted trajectory of this vehicle to the boundary of the kinematic safety envelope; The structured vehicle state at time t is represented. Indicates the vehicle control action at time t; Represents the dynamic safety weight, Indicates kinematic safety weights. This represents the distance normalization scaling parameter. This represents the penalty term when the vehicle state exceeds any safety envelope boundary; comfort reward function. ; The rate of change of acceleration; It is the maximum permissible rate of change of acceleration; Weighting of comfort within the interior; overall reward , and These are the combined weights of the safety reward function and the comfort reward function, respectively. The sample construction module collects driver behavior data and vehicle status, and constructs positive and negative sample datasets based on the comprehensive reward. The training module, based on the deep deterministic policy gradient algorithm, gradually optimizes the driving strategy through the interaction between the agent and the environment. The self-supervised learning module, based on positive and negative sample datasets, uses a self-supervised learning mechanism to correct the deep deterministic policy gradient algorithm and adjust the driving strategy.

5. The intelligent driving self-supervised learning evolutionary system based on a safe envelope boundary as described in claim 4, characterized in that, The collection of driver behavior data and vehicle status, and the construction of a positive and negative sample dataset based on the comprehensive reward, specifically include: Collect driver driving data and vehicle status to form state-action pairs. ,in, For a moment Vehicle status; The vehicle control actions at time t are extracted from the driver's driving data, including steering wheel angle, accelerator pedal opening, and brake pedal opening; the comprehensive reward is calculated for the state-action pairs at each time point. When the total reward of a state-action pair is greater than a set positive sample threshold, the state-action pair is marked as a positive sample; when the total reward of a state-action pair is less than a set negative sample threshold, the state-action pair is marked as a negative sample, thus obtaining a positive and negative sample dataset. The deep deterministic policy gradient algorithm, through the interaction between the agent and the environment, gradually optimizes the driving strategy, specifically including: Initialize the experience replay pool; Initialize the parameters of the value network, target value network, policy network, and target policy network in the deep deterministic policy gradient algorithm; Get the current vehicle status ,action Rewards and vehicle status at the next moment , forming data pairs And store it in the experience replay pool; In vehicle status Take action below Then, the comprehensive reward was calculated. The rewards received; By sampling from the experience replay pool, the parameters of the value network and policy network are updated, and the driving strategy is gradually optimized.

6. The intelligent driving self-supervised learning evolutionary system based on safe envelope boundaries according to claim 5, characterized in that, The process of correcting the deep deterministic policy gradient algorithm and adjusting the driving strategy based on positive and negative sample datasets using a self-supervised learning mechanism specifically includes: A convolutional neural network is used as a feature encoder to process the input sensor data, which is environmental perception data from cameras, LiDAR, and GPS positioning modules; the sensor data is then processed. Through the convolutional neural network Mapping to a high-dimensional embedding space, outputting embedding features ,in Given the parameters of a convolutional neural network; design a contrastive loss function. The embedded features are optimized by comparing the similarity between positive and negative sample pairs. ; in, For anchor sample embedding features, The embedding features are for positive sample pairs. Embedding features for negative sample pairs; express and The similarity between them is calculated, where N represents the number of negative samples in the same batch; the parameters of the convolutional neural network are updated using the backpropagation algorithm to minimize the contrastive loss function. ; in, For learning rate, To compare the loss functions The gradient; the optimized embedded features Vehicle status The concatenation forms the enhanced state that is input to the deep deterministic policy gradient algorithm. Update the value network and policy network, and adjust the driving strategy.

Citation Information

Patent Citations

  • Adaptive control method for driving

    CN110745136A

  • Vehicle following speed control method based on deep reinforcement learning

    CN116811882A

  • New energy vehicle operation optimization method and system based on intelligent network connection

    CN119975396A