Intelligent vehicle control method based on deep reinforcement learning and driving style

Through preprocessing of vehicle status and driving style characteristics and multi-objective reward function training, the problem that intelligent driving vehicles are difficult to take into account safety and comfort in complex environments is solved, and personalized intelligent vehicle control is achieved.

CN120245996AInactive Publication Date: 2025-07-04四川吉利学院

Patent Information

Application Number
CN202510735451.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Intelligent driving vehicles are difficult to meet multiple goals such as safety, efficiency, and comfort in complex traffic environments. Traditional methods rely on precise models and lack personalization. End-to-end learning may ignore passenger comfort or personalized needs.

Method used

By preprocessing the vehicle state, environment and driving style features, a multi-objective reward function is constructed, and an intelligent vehicle control network is trained using deep reinforcement learning to embed driving style features to guide decisions.

Benefits of technology

It improves the personalization of the strategy and user acceptance, and achieves balanced decisions in complex environments, taking into account safety, driving style consistency and passenger comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120245996A_ABST
    Figure CN120245996A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent vehicle control method based on deep reinforcement learning and a driving style, and relates to the technical field of vehicle control, and the method comprises the steps: carrying out the preprocessing of the state data, environment data and driving style characteristics of a vehicle, and obtaining a state vector; constructing a reward function of intelligent vehicle control and an initial intelligent vehicle control network; based on the initial control instruction, the state vector and the reward function, training the initial intelligent vehicle control network by using a deep reinforcement learning method to obtain a trained intelligent vehicle control network; and analyzing the collected state vector by using the trained intelligent vehicle control network to obtain a control command and complete intelligent vehicle control. The problem that road traffic safety, vehicle dynamics constraint and driver driving preference are difficult to consider in intelligent vehicle control is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the technical field of vehicle control, and particularly to an intelligent vehicle control method based on deep reinforcement learning and driving style. Background Art

[0002] When an intelligent driving vehicle performs autonomous decision-making and control in a complex traffic environment, it needs to simultaneously meet multiple goals such as safety, efficiency, and comfort. However, traditional control methods (such as Model Predictive Control MPC) rely on accurate vehicle dynamics models and artificially designed cost functions, and it is difficult to adapt to complex and changing environments and personalized driving style preferences; although the end-to-end deep reinforcement learning method can self-learn control strategies by interacting with the environment, if not constrained, it may ignore passenger comfort or personalized needs, and the decision-making process lacks interpretability and personalization. Summary of the Invention

[0003] Aiming at the above deficiencies in the prior art, an intelligent vehicle control method based on deep reinforcement learning and driving style provided by the present invention solves the problem that it is difficult for intelligent vehicle control to take into account road traffic safety, vehicle dynamics constraints, and driver driving preferences.

[0004] To achieve the above invention purpose, the technical solution adopted by the present invention is: an intelligent vehicle control method based on deep reinforcement learning and driving style, including: S1: Preprocess the vehicle's own state data, environmental data, and driving style features to obtain a state vector; S2: Construct a reward function for intelligent vehicle control and an initial intelligent vehicle control network; S3: Based on the initial control command, the state vector, and the reward function, use the deep reinforcement learning method to train the initial intelligent vehicle control network to obtain a trained intelligent vehicle control network; S4: Use the trained intelligent vehicle control network to analyze the collected state vector to obtain a control command, and complete intelligent vehicle control.

[0005] The beneficial effects of the present invention are: an intelligent vehicle control method based on deep reinforcement learning and driving style. (1) By directly embedding driving style features into the state space of reinforcement learning, the intelligent agent's decision-making is not only based on the vehicle and environmental states, but also considers the driver's personalized preferences, improving the personalization degree and user acceptance of the strategy. (2) A multi-objective reward function that comprehensively considers safety, driving style consistency, and passenger comfort is designed to provide a more comprehensive feedback signal for strategy learning, prompting the intelligent agent to make more balanced decisions in a complex traffic environment.

[0006] Further, the S1 includes: Obtain the vehicle's own state data, environmental data, and driving behavior data; Analyze the driving behavior data to obtain driving style characteristics; Normalize, filter noise, and extract features from the vehicle's own state data and environmental data, and combine them with the driving style characteristics to obtain a state vector.

[0007] Furthermore, the reward function includes a safety reward, a driving style consistency reward, a smoothness reward, and a collision penalty. Among them, the expression of the reward function is: ; Among them, represents the reward function, represents the weight coefficient of the safety reward, represents the safety reward, represents time, represents the weight coefficient of the driving style consistency reward, represents the driving style consistency reward, represents the weight coefficient of the smoothness reward, represents the smoothness reward, represents the weight coefficient of the collision penalty, represents the collision penalty; The expression of the safety reward is: ; Among them, represents the safety reward, represents time, represents the safety compliance reward constant, represents the actual distance between the vehicle and the obstacle, represents the preset safety distance threshold, represents the adjustment parameter; The expression of the driving style consistency reward is: ; Among them, represents the driving style consistency reward, represents the matching sensitivity adjustment parameter, represents the mapping function, represents the action vector at time t, represents the target action vector, represents the Euclidean norm; The expression of the smoothness reward is: ; Among them, represents the smoothness reward, denote moment denote the smoothness penalty coefficient denote the action vector at time t denote the action vector at time t - 1 denote the square of the Euclidean norm The expression for the collision penalty is ; where denote the collision penalty denote moment denote the constant penalty coefficient denote the indicator function

[0008] Furthermore, the initial intelligent vehicle control network includes an actor module, configured to simulate vehicle operation based on an initial control instruction and a state vector to obtain a transition sample a critic module, configured to perform value evaluation on the transition sample based on a reward function to obtain a target value

[0009] Furthermore, the S3 includes input the initial control instruction and the state vector into the actor module to simulate vehicle operation, and let the agent interact with the environment repeatedly to sample data to obtain a transition sample input the transition sample into the critic module, and perform value evaluation on the transition sample based on the reward function to obtain the target value output by the critic module train the critic module using the loss function based on the target value output by the critic module to obtain a trained critic module update the parameters of the actor module based on the policy gradient using the trained critic module to obtain a trained intelligent vehicle control network

[0010] Furthermore, the expression for the transition sample is ; ; ; where denote the transition sample denote the complete state vector at time t denote the execution action vector with exploration noise denote the immediate reward denote the complete state vector at time t + 1 denote the action vector at time t represents the exploration noise term, represents the current policy network parameters, represents the target policy network of the complete state vector at time t; The expression of the target value is: ; where, represents the target value, represents the immediate reward of the i-th sample, represents the discount factor, represents the target critic network, represents the target policy network, represents the state vector at the next moment, represents the target policy network parameters, represents the target critic network parameters.

[0011] Furthermore, the expression of the loss function is: ; ; where, represents the mean squared error loss function value, represents the critic network parameters, represents the number of mini-batch samples, represents the target Q value of the i-th sample, represents the current critic network, represents the state vector of the i-th sample, represents the action vector of the i-th sample, represents the loss function with respect to the gradient of represents the learning rate of the critic network.

[0012] Furthermore, the expression of the policy gradient is: ; where, represents the gradient with respect to the policy parameter of represents the policy objective function, represents the number of mini-batch samples, represents the gradient of the Q function with respect to the action, represents the current critic network, represents the state vector of the i-th sample, represents a placeholder, represents the critic network parameters, represents the output action of the current policy network, Represents the actor network parameters.

[0013] Further, the S4 includes: Analyze the collected state vectors using the trained intelligent vehicle control network to obtain the current predicted action sequence; Based on the current predicted action sequence and the real state, construct a prediction deviation. When the prediction deviation is greater than the threshold, generate a local replacement action sequence; otherwise, obtain the control command to complete the intelligent vehicle control.

[0014] Further, the expression of the prediction deviation is: ; where represents the prediction deviation, represents the predicted next state vector, represents the observed next state vector, represents the Euclidean norm; The expression of the local replacement action sequence is: ; where represents the action vector at the k-th step after replacement, represents the relative step index, represents the local planning time domain length.

[0015] By monitoring the deviation between the model prediction and the actual change of the environment, when the deviation exceeds the threshold, trigger local trajectory replanning to improve the robustness and adaptability of the system to the dynamic unknown environment, and ensure that the vehicle can still operate safely in case of emergencies. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] This specification will be further described in the form of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where: Figure 1 is an exemplary flowchart of an intelligent vehicle control method based on deep reinforcement learning and driving style according to some embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0018] Embodiment Figure 1 is an exemplary flowchart of an intelligent vehicle control method based on deep reinforcement learning and driving style as shown in some embodiments of this specification. As Figure 1 shown, the process includes the following steps. In some embodiments, the process can be executed by a processor.

[0019] S1: Preprocess the vehicle's own state data, environmental data, and driving style characteristics to obtain a state vector.

[0020] The vehicle's own state data is data reflecting the vehicle's own dynamic state. For example, the vehicle's own state data can include position, speed, acceleration, and heading angle, etc.

[0021] The environmental data is data reflecting the vehicle's surrounding environmental conditions. For example, the environmental data can include the distance to the nearest obstacle, the degree of deviation from the center of the lane, the state of surrounding vehicles, etc.

[0022] The driving style characteristics are indicators reflecting the driver's personalized driving habits. For example, the driving style characteristics can include the peak value of lateral acceleration, the change rate of steering angle, etc.

[0023] The state vector is a vector used to characterize the state of the vehicle and its surrounding environment. For example, the state vector can be expressed as: ; where represents the state vector, represents the position of the vehicle at the current moment of, represents the longitudinal speed of the vehicle at time of, represents the longitudinal acceleration of the vehicle at time of, represents the heading angle of the vehicle at time of, represents the distance between the vehicle and the nearest obstacle or target vehicle in front, represents the deviation amount between the vehicle center and the center line of the lane, represents the driving style characteristics.

[0024] In some embodiments, the processor can implement S1 based on the following steps: Obtain the vehicle's own state data, environmental data, and driving behavior data; analyze the driving behavior data to obtain driving style characteristics; perform normalization, noise filtering, and feature extraction on the vehicle's own state data and environmental data, and combine them with the driving style characteristics to obtain a state vector.

[0025] The driving behavior data is the command data received by the vehicle during operation. For example, the driving behavior data may include the steering wheel angle, the accelerator and brake pedal positions, and the corresponding time series, etc.

[0026] In some embodiments, the processor can use a convolutional neural network to extract the spatial features of the driving behavior data, use a long short-term memory network to extract the time-dependent features, and then output the driving style features via a fully connected layer. The convolutional neural network and the long short-term memory network are trained using a supervised learning method to classify the driving style into predefined categories (such as steady type, normal type, aggressive type, changeable type, etc.), and each category corresponds to a set of style feature vectors. The recognition module classifies the driving operation data within the latest period of time during operation and outputs the driving style feature vector of the current driver. This feature vector encodes the key patterns of the driver's behavior (such as typical lateral acceleration changes, steering amplitudes, etc.), and is subsequently embedded in the state vector of the reinforcement learning to guide the control decision to approach the driver's style.

[0027] S2: Construct the reward function for intelligent vehicle control and the initial intelligent vehicle control network.

[0028] The reward function is a function used to evaluate the intelligent vehicle control situation. For example, the reward function may include a safety reward, a driving style consistency reward, a smoothness reward, and a collision penalty.

[0029] In some embodiments, the expression of the reward function can be: ; where, represents the reward function, represents the weight coefficient of the safety reward, represents the safety reward, represents time, represents the weight coefficient of the driving style consistency reward, represents the driving style consistency reward, represents the weight coefficient of the smoothness reward, represents the smoothness reward, represents the weight coefficient of the collision penalty, represents the collision penalty.

[0030] In some embodiments, the expression of the safety reward can be: ; where, represents the safety reward, represents time, represents the safety compliance reward constant, Represents the actual distance between the vehicle and the obstacle, Represents the preset safety distance threshold, Represents the adjustment parameter.

[0031] In some embodiments, the expression of the driving style consistency reward can be: ; Wherein, Represents the driving style consistency reward, Represents the matching sensitivity adjustment parameter, Represents the mapping function, Represents the action vector at time t, Represents the target action vector, Represents the Euclidean norm.

[0032] In some embodiments, the expression of the smoothness reward can be: ; Wherein, Represents the smoothness reward, Represents Time, Represents the smoothness penalty coefficient, Represents the action vector at time t, Represents the action vector at time t - 1, Represents the square of the Euclidean norm.

[0033] In some embodiments, the expression of the collision penalty is: ; Wherein, Represents the collision penalty, Represents Time, Represents the constant penalty coefficient, Represents the indicator function.

[0034] The initial intelligent vehicle control network is a neural network used to generate a control instruction to model the vehicle operation according to the control instruction and the state vector.

[0035] In some embodiments, the initial intelligent vehicle control network may include an actor module and a critic module.

[0036] The actor module is used to simulate the vehicle operation based on the initial control instruction and the state vector to obtain a transition sample.

[0037] The transition sample is an interaction sample of the actor module during the simulation of the vehicle operation.

[0038] A critic module, configured to evaluate the value of the transfer sample based on a reward function to obtain a target value.

[0039] The target value is data reflecting the value degree of the simulated vehicle operation for each control instruction.

[0040] S3: Based on the initial control instruction, the state vector, and the reward function, use the deep reinforcement learning method to train the initial intelligent vehicle control network to obtain a trained intelligent vehicle control network.

[0041] In some embodiments, the processor may implement S3 based on the following steps: input the initial control instruction and the state vector into the actor module to simulate vehicle operation, allowing the agent to interact with the environment repeatedly to sample data and obtain transfer samples; input the transfer samples into the critic module, evaluate the value of the transfer samples based on the reward function to obtain the target value output by the critic module; based on the target value output by the critic module, use the loss function to train the critic module to obtain a trained critic module; use the trained critic module to update the parameters of the actor module based on the policy gradient to obtain a trained intelligent vehicle control network.

[0042] In some embodiments, the expression of the transfer sample may be: ; ; ; where represents the transfer sample, represents the complete state vector at time t, represents the execution action vector with exploration noise, represents the immediate reward, represents the complete state vector at time t+1, represents the action vector at time t, represents the exploration noise term, represents the current policy network parameters, represents the target policy network of the complete state vector at time t.

[0043] In some embodiments, the expression of the target value may be: ; where represents the target value, represents the immediate reward of the i-th sample, represents the discount factor, represents the target critic network, represents the target policy network. represents the state vector at the next moment, represents the parameters of the target policy network, represents the parameters of the target critic network.

[0044] The loss function is the loss function used to train the critic module. For example, the loss function can include the mean squared error loss function.

[0045] In some embodiments, the expression of the loss function can be: ; ; where, represents the mean squared error loss function value, represents the critic network parameters, represents the number of mini-batch samples, represents the target Q value of the i-th sample, represents the current critic network, represents the state vector of the i-th sample, represents the action vector of the i-th sample, represents the loss function with respect to gradient, represents the learning rate of the critic network.

[0046] The policy gradient is the data used to control the update speed of the actor module parameters.

[0047] The critic network parameters are a constant.

[0048] In some embodiments, the expression of the policy gradient can be: ; where, represents the gradient with respect to the policy parameters gradient, represents the policy objective function, represents the number of mini-batch samples, represents the gradient of the Q function with respect to the action, represents the current critic network, represents the state vector of the i-th sample, represents a placeholder, represents the critic network parameters, represents the output action of the current policy network, represents the actor network parameters.

[0049] In some embodiments, the processor can, at regular intervals of a certain number of steps, use a smaller update coefficient Perform a sliding average update on the network parameters to obtain the updated critic module parameters and the updated actor module parameters: ; Among them, represents the updated critic module parameters, represents the update coefficient, represents the critic module parameters, represents the updated actor module parameters, represents the actor module parameters.

[0050] S4: Analyze the collected state vectors using the trained intelligent vehicle control network to obtain control commands and complete the intelligent vehicle control.

[0051] The control command is an operation instruction for completing the intelligent vehicle control.

[0052] In some embodiments, the processor can analyze the collected state vectors using the trained intelligent vehicle control network to obtain the current predicted action sequence; Based on the current predicted action sequence and the true state, construct a prediction deviation. When the prediction deviation is greater than the threshold, generate a local replacement action sequence; otherwise, obtain the control command and complete the intelligent vehicle control.

[0053] The current predicted action sequence is a sequence of the next moment states predicted by the current policy.

[0054] The true state is the true next moment state observed by the sensor.

[0055] The prediction deviation is the gap between the current predicted action sequence and the true state.

[0056] In some embodiments, the expression of the prediction deviation can be: ; Among them, represents the prediction deviation, represents the predicted next state vector, represents the observed next state vector, represents the Euclidean norm.

[0057] The local replacement action sequence is a sequence used to replace the local data in the predicted action sequence.

[0058] In some embodiments, the expression of the local replacement action sequence can be: ; Among them, represents the action vector at the k-th step after replacement, represents the relative step index, Indicates the local planning time domain length.

[0059] By monitoring the deviation between the model prediction and the actual environmental changes, when the deviation exceeds the threshold, local trajectory replanning is triggered to improve the system's robustness and adaptability to dynamic unknown environments, ensuring that the vehicle can still operate safely in case of emergencies.

[0060] In some embodiments, the vehicle can be equipped with a variety of sensors and a sufficient computing platform: environmental perception is achieved using sensors such as cameras, lidar (LiDAR), and millimeter-wave radars to obtain information about roads, obstacles, and other traffic participants; information such as position, speed, and acceleration is obtained using GPS, IMU, and in-vehicle CAN bus to acquire the vehicle's own state. A high-performance in-vehicle computing unit (such as a domain controller with a GPU) is used to run deep neural network inferences, including CNN+LSTM style recognition models and Actor-Critic decision networks.

[0061] In some embodiments, the control instructions can be executed through the vehicle drive control interface (steer-by-wire and brake / throttle-by-wire) to complete intelligent vehicle control.

[0062] The system adopts a modular architecture, and each functional module (perception, style recognition, decision control) is independently designed with clear interfaces, facilitating separate verification and optimization, and can be easily integrated into an actual vehicle platform. The training process makes full use of simulator and real vehicle collected data to improve the model's generalization ability.

[0063] In some embodiments of this specification, an intelligent vehicle control method based on deep reinforcement learning and driving style is provided. (1) By directly embedding driving style features into the state space of reinforcement learning, the intelligent agent's decision-making is not only based on the vehicle and environmental states but also takes into account the driver's personalized preferences, improving the personalization level and user acceptance of the policy. (2) A multi-objective reward function that comprehensively considers safety, driving style consistency, and passenger comfort is designed to provide a more comprehensive feedback signal for policy learning, prompting the intelligent agent to make more balanced decisions in complex traffic environments.

Claims

1. An intelligent vehicle control method based on deep reinforcement learning and driving style, characterized in that, Including: S1: Preprocess the vehicle's own state data, environmental data, and driving style characteristics to obtain a state vector; S2: Construct a reward function and an initial intelligent vehicle control network for intelligent vehicle control; S3: Based on the initial control instruction, the state vector, and the reward function, use the deep reinforcement learning method to train the initial intelligent vehicle control network to obtain a trained intelligent vehicle control network; S4: Use the trained intelligent vehicle control network to analyze the collected state vector to obtain a control command and complete intelligent vehicle control.

2. The intelligent vehicle control method based on deep reinforcement learning and driving style according to claim 1, characterized in that The S1 includes: Obtain the vehicle's own state data, environmental data, and driving behavior data; Analyze the driving behavior data to obtain driving style characteristics; Normalize, filter noise, and extract features from the vehicle's own state data and environmental data, and combine them with the driving style characteristics to obtain a state vector.

3. The intelligent vehicle control method based on deep reinforcement learning and driving style according to claim 1, characterized in that The reward function includes a safety reward, a driving style consistency reward, a smoothness reward, and a collision penalty. Among them, the expression of the reward function is: ; Among them, represents the reward function, represents the weight coefficient of the safety reward, represents the safety reward, represents the moment, represents the weight coefficient of the driving style consistency reward, represents the driving style consistency reward, represents the weight coefficient of the smoothness reward, represents the smoothness reward, represents the weight coefficient of the collision penalty, represents the collision penalty; The expression of the safety reward is: ; Among them, represents the safety reward, represents time, represents the safety compliance reward constant, represents the actual distance between the vehicle and the obstacle, represents the preset safety distance threshold, represents the adjustment parameter; The expression of the driving style consistency reward is: ; Among them, represents the driving style consistency reward, represents the matching sensitivity adjustment parameter, represents the mapping function, represents the action vector at time t, represents the target action vector, represents the Euclidean norm; The expression of the smoothness reward is: ; Among them, represents the stability reward, represents the moment, represents the smoothness penalty coefficient, represents the action vector at time t, represents the action vector at time t-1, represents the square of the Euclidean norm; The expression of the collision penalty is: ; Among them, represents the collision penalty, represents time, represents the constant penalty coefficient, represents the indicator function.

4. The intelligent vehicle control method based on deep reinforcement learning and driving style according to claim 1, wherein The initial intelligent vehicle control network includes: An actor module for simulating vehicle operation based on the initial control instruction and the state vector to obtain a transition sample; A critic module for evaluating the value of the transition sample based on the reward function to obtain a target value.

5. The intelligent vehicle control method based on deep reinforcement learning and driving style according to claim 4, characterized in that The S3 includes: Input the initial control instruction and the state vector into the actor module to simulate vehicle operation, and let the agent interact with the environment repeatedly to sample data to obtain a transition sample; Input the transition sample into the critic module, and based on the reward function, evaluate the value of the transition sample to obtain the target value output by the critic module; Based on the target value output by the critic module, use the loss function to train the critic module to obtain a trained critic module; Use the trained critic module to update the parameters of the actor module based on the policy gradient to obtain a trained intelligent vehicle control network.

6. The intelligent vehicle control method based on deep reinforcement learning and driving style according to claim 5, characterized in that, The expression of the transition sample is: ; ; ; Among them, represents the transfer sample, represents the complete state vector at time t, represents the execution action vector with exploration noise, represents the immediate reward, represents the complete state vector at time t + 1, represents the action vector at time t, represents the exploration noise term, represents the current policy network parameters, represents the target policy network of the complete state vector at time t; The expression of the target value is: ; Among them, represents the target value, represents the immediate reward of the i-th sample, represents the discount factor, represents the target critic network, represents the target policy network, represents the state vector at the next moment, represents the parameters of the target policy network, represents the parameters of the target critic network.

7. The intelligent vehicle control method based on deep reinforcement learning and driving style according to claim 5, wherein The expression of the loss function is: ; ; Among them, represents the mean squared error loss function value, represents the critic network parameters, represents the number of mini-batch samples, represents the target Q value of the i-th sample, represents the current critic network, represents the state vector of the i-th sample, represents the action vector of the i-th sample, represents the gradient of the loss function with respect to , represents the learning rate of the critic network.

8. The intelligent vehicle control method based on deep reinforcement learning and driving style according to claim 5, characterized in that The expression of the policy gradient is: ; Among them, represents the gradient of the policy parameter ; represents the policy objective function, represents the number of mini-batch samples, represents the gradient of the Q function with respect to the action, represents the current critic network, represents the state vector of the i-th sample, represents a placeholder, represents the critic network parameters, represents the action output by the current policy network, represents the actor network parameters. It should be noted that there seems to be a semicolon missing in the original Chinese text at the end of the description of . I added it in the translation for better semantic understanding. If this is not allowed according to your strict rules, please adjust accordingly.

9. The intelligent vehicle control method based on deep reinforcement learning and driving style according to claim 1, characterized in that The S4 includes: Use the trained intelligent vehicle control network to analyze the collected state vector to obtain the current predicted action sequence; Based on the current predicted action sequence and the real state, construct a prediction deviation. When the prediction deviation is greater than the threshold, generate a local replacement action sequence; otherwise, obtain a control command and complete intelligent vehicle control.

10. The intelligent vehicle control method based on deep reinforcement learning and driving style according to claim 9, characterized in that, The expression of the prediction deviation is: ; Among them, represents the prediction deviation, represents the predicted next state vector, represents the observed next state vector, represents the Euclidean norm; The expression of the local replacement action sequence is: ; Among them, represents the action vector at the k-th step after replacement, represents the relative step index, represents the local planning time domain length.

Citation Information

Patent Citations

  • Method for planning motion of intelligent vehicle on highway based on multi-scale breadth-first network

    CN116483069A

  • Multi-level human intelligence enhanced automatic driving vehicle decision control method and system

    CN117227758A

  • End-to-end automatic driving lane changing decision-making method based on multi-modal input

    CN118861963A

  • Multi-agent-based multi-lane ramp confluence area vehicle control method and system

    CN119975359A

  • Automatic parking path planning model construction method and device based on deep reinforcement learning, equipment and vehicle

    CN120003469A

Cited By

  • Industrial Internet of Things unmanned vehicle path planning system and method

    CN120806316A