Adaptive Vehicle Extreme Driving Control Method and System Based on Unexpected External Environmental Changes
By adopting an adaptive control method to adapt to changes in the unexpected environment, and combining conservative safety and performance-optimal strategies, a hybrid control strategy is constructed to solve the safety and optimization problems of vehicles under extreme driving conditions when unexpected environmental changes occur, thereby achieving dynamic balance and improving emergency response capabilities.
Patent Information
- Application Number
- CN202310827122.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-06
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-07-06
AI Technical Summary
Existing vehicle extreme driving control methods struggle to achieve a dynamic balance between safety and optimal performance when faced with unexpected changes in the external environment. Traditional control methods sacrifice speed to ensure safety, while intelligent control methods sacrifice performance.
An adaptive control method for unforeseen environmental changes is adopted, which combines a conservative safety control strategy with a performance-optimal extreme driving control strategy. A hybrid control strategy is constructed through reinforcement learning and optimal control algorithms to dynamically balance safety and performance.
When unexpected external environmental changes occur, dynamic balance can be achieved in extreme driving conditions, improving emergency response capabilities and generalization performance, bridging the simulation-reality gap in intelligent algorithms, and enhancing the safety and optimality of extreme driving.
Smart Images

Figure CN116605242B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of active safety technology for automobiles, and in particular to a method and system for adaptive extreme driving control of automobiles in response to unexpected changes in the external environment. Background Technology
[0002] With the popularization and advancement of vehicle intelligence and connectivity, and the increasing demands of people for driving pleasure and experience, a new research field has emerged in recent years: extreme driving. In pursuit of faster driving speeds and a better driving experience, vehicles in extreme driving situations are often subjected to high speeds, sharp turns, and other extreme conditions. The vehicle tires are in a saturated nonlinear state, the vehicle's control margin is reduced, and it is more prone to instability. Therefore, extreme driving control methods are a key focus and challenge in research.
[0003] Currently, the main objectives of vehicle extreme driving control focus on trajectory planning and motion control. Control methods are broadly categorized into model-based traditional control methods and data-driven intelligent control methods. Traditional control methods mainly include sliding mode control, optimal control, and model predictive control, which offer high reliability but rely heavily on accurate vehicle dynamics models. Intelligent control algorithms are data-driven, possessing self-learning and self-determination capabilities without depending on vehicle dynamics; these include deep learning control and reinforcement learning control, among others. Furthermore, when the surrounding environment changes, current control methods primarily employ switching control between traditional and intelligent control approaches.
[0004] However, when unexpected environmental changes occur, the aforementioned control methods struggle to achieve a dynamic balance between safety and optimal performance. Traditional control methods prioritize safety, while intelligent control methods prioritize optimal performance. When unexpected environmental changes occur, switching-based control sacrifices too much optimality to ensure safety, leading to reduced driving performance. Specifically, when unexpected environmental changes occur, such as the sudden appearance of an unexpected obstacle in front of the vehicle or a sudden change in road conditions, intelligent control algorithms alone struggle to achieve stable vehicle control, posing safety risks. Conversely, if only traditional control is used, to ensure vehicle safety in extreme scenarios, traditional control often sacrifices some vehicle speed, increasing lap time. Furthermore, unexpected environmental changes are diverse, requiring vehicles to possess corresponding adaptive capabilities. In conclusion, when unexpected environmental changes occur, current control methods struggle to achieve a dynamic balance between safety and optimal performance during extreme driving conditions, indicating research shortcomings and deficiencies. Summary of the Invention
[0005] The present invention aims to at least partially solve one of the technical problems in the related art.
[0006] To this end, this invention proposes an adaptive vehicle extreme driving control method that adapts to unexpected environmental changes, improves the effectiveness and generalization of vehicle extreme control algorithms, accurately identifies unexpected environmental changes, judges potential risk factors, and formulates control strategies for unexpected environmental changes to enhance adaptability to different scenarios, so as to achieve a dynamic balance between vehicle extreme driving safety and optimal performance.
[0007] Another objective of this invention is to propose an adaptive vehicle extreme driving control system that adapts to unexpected environmental changes.
[0008] To achieve the above objectives, this invention proposes an adaptive vehicle extreme driving control method for unpredictable external environmental changes, comprising:
[0009] Based on vehicle operating environment data, a conservative safety control strategy for online identification of uncertain environmental factors is constructed, and a data-driven, performance-optimal extreme driving control strategy is constructed based on reinforcement learning algorithm and optimal control algorithm.
[0010] Based on the strategy fusion result of the conservative safety control strategy and the performance-optimal extreme driving control strategy, a hybrid control strategy that dynamically balances safety and performance is obtained.
[0011] When the onboard computer determines that the vehicle is in an extreme driving state, it executes the hybrid control strategy and outputs control commands. The actuators respond to the control commands to control the dynamic balance of the vehicle's operation.
[0012] In addition, the adaptive vehicle extreme driving control method for unexpected environmental changes according to the above embodiments of the present invention may also have the following additional technical features:
[0013] Furthermore, in one embodiment of the present invention, the conservative safety control strategy for online identification of uncertain environmental factors based on vehicle operating environment data includes:
[0014] Candidate trajectories are generated based on the vehicle's driving status, and unexpected environmental changes in the track are identified online to obtain environmental change recognition results.
[0015] Calculate the uncertain environmental hazard level corresponding to the candidate trajectory based on the environmental change identification results;
[0016] Based on the comparison results between the uncertain environmental hazard level and the hazard threshold, candidate trajectories that do not exceed the hazard threshold are selected to obtain a conservative safety control strategy and generate conservative safety control input data.
[0017] Furthermore, in one embodiment of the present invention, the construction of a data-driven, performance-optimal extreme driving control strategy based on reinforcement learning algorithms and optimal control algorithms includes:
[0018] Constructing the state space and action space in an assisted reinforcement learning network;
[0019] A simulation environment for establishing a high-precision vehicle dynamics model and a motion-assisted reinforcement learning network based on the state space and action space;
[0020] A reward function is designed based on the optimal performance index of the high-precision vehicle dynamics model, and an expert policy is obtained based on expert examples and prior knowledge. The assisted reinforcement learning network is iteratively trained in the simulation environment based on the reward function and the expert policy to obtain the network parameters.
[0021] In a real-world environment, the onboard computer is used to calculate the network parameters, and the parameter calculation results are used as the optimal performance limit driving control strategy to generate optimal performance input data.
[0022] Furthermore, in one embodiment of the present invention, obtaining a hybrid control strategy that dynamically balances safety and performance based on the strategy fusion result of the conservative safety control strategy and the performance-optimal extreme driving control strategy includes:
[0023] Calculate the optimal control confidence factor based on the analysis results of the distance between the current state and the optimal performance trajectory;
[0024] The conservative safety control confidence factor is calculated based on the analysis results of environmental hazard factors caused by unexpected online changes.
[0025] Based on the performance-optimal control trust factor and the conservative safety control trust factor, a hybrid control factor is calculated and a hybrid control strategy that dynamically balances safety and performance is generated to obtain actual control input data.
[0026] Furthermore, in one embodiment of the present invention, the distance between the current state and the optimal performance trajectory includes the distance between the vehicle and the optimal trajectory in the track tangent direction, the difference between the vehicle's heading angle and the optimal trajectory's heading angle, and the difference between the current vehicle speed and the speed corresponding to the optimal trajectory.
[0027] To achieve the above objectives, another aspect of the present invention proposes an adaptive vehicle extreme driving control system for unforeseen changes in the external environment, comprising:
[0028] The driving control strategy construction module is used to construct a conservative safety control strategy that identifies uncertain environmental factors online based on vehicle operating environment data, and to construct a data-driven, performance-optimal extreme driving control strategy based on reinforcement learning algorithms and optimal control algorithms.
[0029] A hybrid control strategy construction module is used to obtain a hybrid control strategy that dynamically balances safety and performance based on the strategy fusion result of the conservative safety control strategy and the performance-optimal extreme driving control strategy.
[0030] The instruction output state control module is used by the on-board computer to execute the hybrid control strategy and output control instructions when it determines that the vehicle is in an extreme driving state. The actuator responds to the control instructions to control the dynamic balance of the vehicle's operation.
[0031] The adaptive vehicle extreme driving control method and system for unexpected environmental changes in this invention can dynamically balance optimality and safety when unexpected environmental changes occur, thereby improving the emergency response level and generalization performance of the vehicle extreme control algorithm.
[0032] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0033] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0034] Figure 1 This is a flowchart of an adaptive vehicle extreme driving control method according to an embodiment of the present invention;
[0035] Figure 2 This is a flowchart of another adaptive vehicle extreme driving control method according to an embodiment of the present invention;
[0036] Figure 3 This is a schematic diagram of adaptive vehicle extreme driving control based on unexpected environmental changes according to an embodiment of the present invention;
[0037] Figure 4 This is a schematic diagram illustrating unexpected changes in the online environment during extreme driving on a track corner according to an embodiment of the present invention;
[0038] Figure 5 This is a diagram of the assisted reinforcement learning network architecture according to an embodiment of the present invention;
[0039] Figure 6 This is a schematic diagram comparing the control effects of various strategies according to an embodiment of the present invention;
[0040] Figure 7 This is a schematic diagram comparing the control effects of various strategies according to another embodiment of the present invention;
[0041] Figure 8 This is a schematic diagram of the structure of an adaptive vehicle extreme driving control system for unexpected environmental changes according to an embodiment of the present invention. Detailed Implementation
[0042] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0043] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0044] The following description, with reference to the accompanying drawings, describes an adaptive vehicle extreme driving control method and system for unpredictable environmental changes according to embodiments of the present invention.
[0045] Example 1:
[0046] Figure 1 This is a flowchart of an adaptive vehicle extreme driving control method based on unexpected environmental changes, according to an embodiment of the present invention.
[0047] like Figure 1 As shown, the method includes, but is not limited to, the following steps:
[0048] S1. Conservative safety control strategy for online identification of uncertain environmental factors is constructed based on vehicle operating environment data, and data-driven performance-optimal extreme driving control strategy is constructed based on reinforcement learning algorithm and optimal control algorithm.
[0049] S2, based on the strategy fusion result of conservative safety control strategy and performance-optimal extreme driving control strategy, a hybrid control strategy that dynamically balances safety and performance is obtained;
[0050] S3: When the onboard computer determines that the car is in an extreme driving state, it executes a hybrid control strategy and outputs control commands. The actuators respond to the control commands to control the dynamic balance of the car's operation.
[0051] In one embodiment of the present invention, candidate trajectories are generated based on the vehicle's driving state, and unexpected environmental changes in the track are identified online to obtain environmental change identification results. The uncertainty of the environment corresponding to the candidate trajectory is calculated, and the candidate trajectory that does not exceed the danger threshold is selected to obtain a conservative safety control strategy to generate conservative safety control input data.
[0052] In one embodiment of the present invention, a state space and action space under an assisted reinforcement learning network are constructed; a simulation environment for a high-precision vehicle dynamics model and an assisted reinforcement learning network is established; a reward function is designed based on the performance-optimal index of the high-precision vehicle dynamics model; an expert policy is obtained based on expert examples and prior knowledge; the assisted reinforcement learning network is iteratively trained in the simulation environment to obtain network parameters; the network parameters are calculated using an on-board computer in a real environment; and the parameter calculation results are used as the performance-optimal extreme driving control strategy to generate performance-optimal input data.
[0053] In one embodiment of the present invention, a performance-optimal control confidence factor is calculated based on the analysis results of the distance between the current state and the performance-optimal trajectory, a conservative safety control confidence factor is calculated based on the analysis results of environmental hazard factors caused by unexpected online changes, a hybrid control factor is calculated, and a hybrid control strategy that dynamically balances safety and performance is generated to obtain actual control input data.
[0054] In one embodiment of the present invention, the distance between the current state and the optimal performance trajectory includes the distance between the vehicle and the optimal trajectory in the tangential direction of the track, the difference between the vehicle's heading angle and the heading angle of the optimal trajectory, and the difference between the current vehicle speed and the speed corresponding to the optimal trajectory.
[0055] The adaptive vehicle extreme driving control method based on unexpected environmental changes in this invention, from a scenario perspective, can dynamically balance optimality and safety when unexpected environmental changes occur, improving the emergency response level and generalization performance of the vehicle extreme control algorithm. From an algorithmic perspective, this invention proposes a dynamic hybrid control strategy for data-driven intelligent algorithms, expanding the generalization performance of data-driven algorithms and effectively bridging the "simulation-reality gap" of intelligent algorithms.
[0056] Example 2:
[0057] like Figure 2 As shown, the adaptive vehicle extreme driving control method of the present invention for unpredictable environmental changes establishes a conservative safety control strategy that can identify uncertain environmental factors online, constructs a data-driven performance-optimized extreme driving control strategy, and sets up a hybrid control mechanism that dynamically balances safety and performance to obtain the final actual control input. The specific implementation on a typical racetrack is as follows... Figure 3 As shown, the specific steps include:
[0058] Furthermore, a conservative safety control strategy capable of identifying uncertain environmental factors online includes: generating candidate trajectories based on vehicle driving status, identifying unexpected environmental changes on the track online, calculating the degree of uncertainty in the environment experienced by each candidate trajectory, selecting candidate trajectories that do not exceed the danger threshold, and generating a conservative safety control input u using model predictive control or steady-state drift control techniques.safe .
[0059] Furthermore, the data-driven performance-optimal extreme driving strategy refers to a performance-oriented vehicle control strategy that cannot adapt to unforeseen changes in the external environment. This includes, but is not limited to, reinforcement learning algorithms trained in a virtual environment and optimal control algorithms that cannot update in a timely manner the complex constraints brought about by changes in the unforeseen external environment. Specifically, this involves constructing the state space s under the assisted reinforcement learning framework. rl and action space A rl The process involves: establishing a high-precision vehicle dynamics model and a dynamic extreme driving track as the environment for assisted reinforcement learning; designing a reward function based on optimal performance metrics, including immediate and final performance rewards; obtaining an expert policy based on expert examples and prior knowledge; iteratively training the exploration policy under the guidance of the expert policy in a simulation environment, ultimately achieving complete control by the exploration policy; and importing the network parameters into the onboard computer in a real-world environment to generate the optimal input u for the optimal extreme driving policy. opt .
[0060] Furthermore, the hybrid control mechanism that dynamically balances safety and performance includes: analyzing the distance between the current state and the optimal trajectory, and calculating the performance-optimal control confidence factor c. opt Analyze the environmental hazards caused by unexpected online changes and calculate the conservative safety control trust factor c. safe Combined with the performance-optimized control trust factor c opt And conservative security control trust factor c safe Calculate the hybrid control factor τ to obtain the actual control input u. s Specifically:
[0061] Understandably, the distance between the current state and the optimal trajectory needs to be quantitatively considered as state variables that affect performance, including: the distance Δl between the vehicle and the optimal trajectory in the track tangent direction, the difference Δα between the vehicle's heading angle and the heading angle of the optimal trajectory, and the difference Δv between the current vehicle speed and the corresponding speed of the optimal trajectory.
[0062] In one embodiment of the present invention, the performance-optimized control trust factor c opt The value ranges from (0, 1] and should increase as the distance between the current state and the optimal trajectory increases.
[0063]
[0064] In one embodiment of the present invention, the conservative security control trust factor c safe The value range is [0, 1], and it should increase as the environmental risk caused by changes in the external environment increases.
[0065]
[0066] In one embodiment of the present invention, the mixed control factor τ and the actual control input u opt Using a weighted method, the specific calculation is as follows:
[0067]
[0068] u s =τu safe +(1-τ)u opt
[0069] It is known that optimal control strategies also include other performance-oriented vehicle control strategies that cannot adapt to unexpected environmental changes, and are not limited to reinforcement learning algorithms trained in virtual environments and optimal control algorithms that cannot update the complex constraints brought about by unexpected environmental changes in a timely manner.
[0070] For example, this embodiment describes an extreme driving scenario involving high-speed cornering on a racetrack. The sole performance metric is the number of laps the car completes on the racetrack. Two unexpected environmental changes are introduced: accidental vehicle damage on the track and unexpected partial opening of icy or snowy road surfaces on the track. Figure 4 As shown, the assisted reinforcement learning network in this embodiment is as follows: Figure 5 As shown.
[0071] For example, the conservative safety control strategy in this embodiment uses model predictive control. Based on vehicle speed, segmented nodes are generated on the track, and these nodes are selected and connected using Bézier curves to generate a candidate trajectory set Tr.
[0072] Tr={tr1, tr2, tr3,..., tr n-1 tr n}
[0073] For example, vehicle sensing systems such as vehicle radar and cameras can identify unexpected environmental changes online. In this embodiment, when the vehicle is driving on a curve, an unexpected environmental change is detected, which is a vehicle that has been unexpectedly damaged at the curve of the track.
[0074] For example, trajectories that overlap with the envelope of the unexpectedly damaged vehicle are removed from the trajectory set to generate a safe candidate trajectory set Tr. safe The trajectory with the shortest path in the candidate safe trajectory set is selected as the conservative safe trajectory for target tracking. safe .
[0075]
[0076] For example, a two-degree-of-freedom vehicle dynamics model is used to predict and control the conservative safe trajectory t. safe Generate conservative safety control input u safe .
[0077]
[0078] x i+1|t =A t x i|t +B t u i|t
[0079] For example, in this embodiment, the data-driven optimal performance limit driving strategy uses a reinforcement learning method, and the performance metric is the lap time across the track. The state space S includes all the information required for the car's maximum speed limit driving, and the state space in this embodiment is shown below:
[0080] S = [v x v y ,r,γ,s,l,ψ,φ]
[0081] Where, v x v y r and γ represent the longitudinal velocity, lateral velocity, yaw rate, and roll rate of the vehicle in the vehicle coordinate system, respectively; s, l, ψ, and φ represent the vehicle's driving position, lateral offset, yaw rate, and roll rate in the Frenet coordinate system, respectively.
[0082] It is worth noting that designing reinforcement learning algorithms in the Frenet coordinate system will help reduce the sparsity of rewards and improve the training efficiency of the algorithm.
[0083] For example, the vehicle studied in this embodiment is a front-wheel steering four-wheel drive vehicle, and the motion space is constructed as shown below:
[0084] A = [δ] f [T]
[0085] For example, in this embodiment, the high-precision vehicle dynamics model uses a model built into the Carsim software; and the reward function is designed as follows, in conjunction with the target performance:
[0086] R = R i +R t
[0087]
[0088] R t =k t (t e -t)
[0089] Where R i For immediate performance rewards, high speeds at each moment are rewarded; R t As a final performance bonus, the shortest lap time is encouraged. i kt It is a positive number.
[0090] For example, a deep neural network for Critic and Actor, based on a reinforcement learning algorithm, is constructed. In a simulation environment, an expert policy guides the algorithm to iteratively train the Critic and Actor networks to explore strategies, ultimately resulting in complete control by the Actor network. In a real-world environment, the network parameters are imported into an onboard computer and used as the optimal input u to generate the optimal driving strategy. opt .
[0091]
[0092] u~π(a|s i ;θ)
[0093] For example, in this embodiment, the performance-optimized control trust factor c opt The calculation method is as follows:
[0094]
[0095] In the formula, l rl The lateral deviation of the driving position s in the trajectory obtained solely by the optimal control strategy in the absence of unexpected external environmental changes; The longitudinal velocity of the vehicle at position s in the trajectory obtained solely by the optimal control strategy in the absence of unexpected external environmental changes; Let k be the lateral velocity of the vehicle at position s in the trajectory obtained solely by the optimal control strategy in the absence of unexpected external environmental changes. o k o1 k o2 k o3 It is a positive parameter.
[0096] For example, in this embodiment, the conservative security control trust factor c safe The calculation method is as follows:
[0097]
[0098] Where, d UOC α represents the shortest distance from the vehicle's current position to the anticipated changes in the external environment. UOC d0 is the angle between the vehicle's heading and the line connecting the vehicle to the center of the unexpected environmental change; d0 is the distance for online early warning, which is taken as 20m in this embodiment; S UOC k represents the lateral distance occupied by unexpected changes in the external environment. s k s1 It is a positive parameter.
[0099] For example, in this embodiment, the mixed control factor τ is calculated using a weighted method, specifically as follows:
[0100]
[0101] For example, in this embodiment, the final hybrid control input is calculated based on the hybrid control factor to achieve adaptive and dynamic balance between safety and optimal vehicle limit driving control in a track environment, which can adapt to unexpected environmental changes.
[0102] u s =τu safe +(1-τ)u opt
[0103] For example, in this embodiment, the results of using only a conservative safety control strategy, only a data-driven performance-optimal extreme driving strategy, and an adaptive hybrid control strategy are compared as follows: Figure 6 and Figure 7 As shown.
[0104] The adaptive vehicle extreme driving control method according to embodiments of the present invention can dynamically balance optimality and safety when unexpected environmental changes occur, improve the emergency response level and generalization performance of vehicle extreme control algorithms, enhance the backup level of intelligent vehicle extreme control, improve the training efficiency of data-driven performance-driven control strategies, enable them to adapt to large-scale tasks such as track driving, expand the generalization performance of such algorithms, and effectively bridge the "simulation-reality gap" of intelligent algorithms.
[0105] Example 3:
[0106] To achieve the above embodiments, such as Figure 8 As shown, this embodiment also provides an adaptive vehicle extreme driving control system 10 for unexpected environmental changes. The system 10 includes a driving control strategy construction module 100, a hybrid control strategy construction module 200, and a command output state control module 300.
[0107] The driving control strategy construction module 100 is used to construct a conservative safety control strategy that identifies uncertain environmental factors online based on vehicle operating environment data, and to construct a data-driven performance-optimal extreme driving control strategy based on reinforcement learning algorithm and optimal control algorithm.
[0108] The hybrid control strategy construction module 200 is used to obtain a hybrid control strategy that dynamically balances safety and performance based on the strategy fusion result of the conservative safety control strategy and the performance-optimal extreme driving control strategy.
[0109] The instruction output state control module 300 is used by the on-board computer to execute a hybrid control strategy and output control commands when it determines that the vehicle is in an extreme driving state. The actuator responds to the control commands to control the dynamic balance of the vehicle's operation.
[0110] Furthermore, the aforementioned driving control strategy construction module 100 is also used for,
[0111] Candidate trajectories are generated based on the vehicle's driving status, and unexpected environmental changes in the track are identified online to obtain environmental change recognition results.
[0112] Calculate the degree of uncertainty in the environment corresponding to the candidate trajectory based on the environmental change identification results;
[0113] Based on the comparison results between the uncertain environmental hazard level and the hazard threshold, candidate trajectories that do not exceed the hazard threshold are selected to obtain a conservative safety control strategy and generate conservative safety control input data.
[0114] Furthermore, the aforementioned driving control strategy construction module 100 is also used for,
[0115] Constructing the state space and action space in an assisted reinforcement learning network;
[0116] A simulation environment for building a high-precision vehicle dynamics model and an auxiliary reinforcement learning network based on state space and action space;
[0117] A reward function is designed based on the performance optimization index of a high-precision vehicle dynamics model, and an expert policy is obtained based on expert examples and prior knowledge. The assisted reinforcement learning network is iteratively trained in a simulation environment based on the reward function and the expert policy to obtain the network parameters.
[0118] In a real-world environment, onboard computers are used to calculate network parameters, and the results of these parameter calculations are used as the optimal performance limit driving control strategy to generate optimal performance input data.
[0119] Furthermore, the aforementioned hybrid control strategy construction module 200 is also used for,
[0120] Calculate the optimal control confidence factor based on the analysis results of the distance between the current state and the optimal performance trajectory;
[0121] The conservative safety control confidence factor is calculated based on the analysis results of environmental hazard factors caused by unexpected online changes.
[0122] Based on the performance-optimal control trust factor and the conservative safety control trust factor, a hybrid control factor is calculated and a hybrid control strategy that dynamically balances safety and performance is generated to obtain the actual control input data.
[0123] Furthermore, the distance between the current state and the optimal performance trajectory includes the distance between the vehicle and the optimal trajectory in the tangential direction of the track, the difference between the vehicle's heading angle and the heading angle of the optimal trajectory, and the difference between the current vehicle speed and the corresponding speed of the optimal trajectory.
[0124] The adaptive vehicle extreme driving control system based on the present invention can dynamically balance optimality and safety when unexpected environmental changes occur, improve the emergency response level and generalization performance of vehicle extreme control algorithms, enhance the backup level of intelligent vehicle extreme control, improve the training efficiency of data-driven performance-driven control strategies, enable it to adapt to large-scale tasks such as track driving, expand the generalization performance of such algorithms, and effectively bridge the "simulation-reality gap" of intelligent algorithms.
[0125] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0126] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
Claims
1. A method for adaptive vehicle extreme driving control based on unforeseen changes in the external environment, characterized in that, Includes the following steps: Based on vehicle operating environment data, a conservative safety control strategy for online identification of uncertain environmental factors is constructed, and a data-driven, performance-optimal extreme driving control strategy is constructed based on reinforcement learning algorithm and optimal control algorithm. Based on the strategy fusion result of the conservative safety control strategy and the performance-optimal extreme driving control strategy, a hybrid control strategy that dynamically balances safety and performance is obtained. When the onboard computer determines that the vehicle is in an extreme driving state, it executes the hybrid control strategy and outputs control commands. The actuators respond to the control commands to control the dynamic balance of the vehicle's operation. The conservative safety control strategy for online identification of uncertain environmental factors based on vehicle operating environment data includes: Candidate trajectories are generated based on the vehicle's driving status, and unexpected environmental changes in the track are identified online to obtain environmental change recognition results; Calculate the uncertain environmental hazard level corresponding to the candidate trajectory based on the environmental change identification results; Based on the comparison results between the uncertain environmental hazard level and the hazard threshold, candidate trajectories that do not exceed the hazard threshold are selected to obtain a conservative safety control strategy and generate conservative safety control input data; The data-driven, performance-optimal extreme driving control strategy constructed based on reinforcement learning and optimal control algorithms includes: Constructing the state space and action space in an assisted reinforcement learning network; A simulation environment for establishing a high-precision vehicle dynamics model and a motion-assisted reinforcement learning network based on the state space and action space; A reward function is designed based on the optimal performance index of the high-precision vehicle dynamics model, and an expert policy is obtained based on expert examples and prior knowledge. The assisted reinforcement learning network is iteratively trained in the simulation environment based on the reward function and the expert policy to obtain the network parameters. In a real-world environment, the on-board computer is used to calculate the network parameters, and the parameter calculation results are used as the performance-optimal extreme driving control strategy to generate performance-optimal input data. The hybrid control strategy that dynamically balances safety and performance, obtained by fusing the conservative safety control strategy and the performance-optimal extreme driving control strategy, includes: Calculate the optimal control confidence factor based on the analysis results of the distance between the current state and the optimal performance trajectory; The conservative safety control confidence factor is calculated based on the analysis results of environmental hazard factors caused by unexpected online changes. Based on the performance-optimal control trust factor and the conservative safety control trust factor, a hybrid control factor is calculated and a hybrid control strategy that dynamically balances safety and performance is generated to obtain actual control input data.
2. The method according to claim 1, characterized in that, The distance between the current state and the optimal performance trajectory includes the distance between the vehicle and the optimal trajectory in the tangential direction of the track, the difference between the vehicle's heading angle and the heading angle of the optimal trajectory, and the difference between the current vehicle speed and the corresponding speed of the optimal trajectory.
3. A vehicle extreme driving control system that adapts to unforeseen changes in the external environment, characterized in that, include: The driving control strategy construction module is used to construct a conservative safety control strategy that identifies uncertain environmental factors online based on vehicle operating environment data, and to construct a data-driven, performance-optimal extreme driving control strategy based on reinforcement learning algorithms and optimal control algorithms. A hybrid control strategy construction module is used to obtain a hybrid control strategy that dynamically balances safety and performance based on the strategy fusion result of the conservative safety control strategy and the performance-optimal extreme driving control strategy. The instruction output state control module is used by the on-board computer to execute the hybrid control strategy and output control instructions when it determines that the vehicle is in an extreme driving state. The actuator responds to the control instructions to control the dynamic balance of the vehicle's operation. The driving control strategy construction module is also used for, Candidate trajectories are generated based on the vehicle's driving status, and unexpected environmental changes in the track are identified online to obtain environmental change recognition results; Calculate the uncertain environmental hazard level corresponding to the candidate trajectory based on the environmental change identification results; Based on the comparison results between the uncertain environmental hazard level and the hazard threshold, candidate trajectories that do not exceed the hazard threshold are selected to obtain a conservative safety control strategy and generate conservative safety control input data; The driving control strategy construction module is also used for, Constructing the state space and action space in an assisted reinforcement learning network; A simulation environment for establishing a high-precision vehicle dynamics model and a motion-assisted reinforcement learning network based on the state space and action space; A reward function is designed based on the optimal performance index of the high-precision vehicle dynamics model, and an expert policy is obtained based on expert examples and prior knowledge. The assisted reinforcement learning network is iteratively trained in the simulation environment based on the reward function and the expert policy to obtain the network parameters. In a real-world environment, the on-board computer is used to calculate the network parameters, and the parameter calculation results are used as the performance-optimal extreme driving control strategy to generate performance-optimal input data. The hybrid control strategy construction module is also used for, Calculate the optimal control confidence factor based on the analysis results of the distance between the current state and the optimal performance trajectory; The conservative safety control confidence factor is calculated based on the analysis results of environmental hazard factors caused by unexpected online changes. Based on the performance-optimal control trust factor and the conservative safety control trust factor, a hybrid control factor is calculated and a hybrid control strategy that dynamically balances safety and performance is generated to obtain actual control input data.
4. The system according to claim 3, characterized in that, The distance between the current state and the optimal performance trajectory includes the distance between the vehicle and the optimal trajectory in the tangential direction of the track, the difference between the vehicle's heading angle and the heading angle of the optimal trajectory, and the difference between the current vehicle speed and the corresponding speed of the optimal trajectory.
Citation Information
Patent Citations
Vehicle path tracking control method based on hybrid switching of model and reinforcement learning
CN114355897A
System and method for active traction control of a vehicle
US20100161194A1