A training method for a cognitive model of unfamiliar environment states in an online driving scenario
By building the Actor-Critic-Guard network architecture, combining multi-camera acquisition and BEV feature fusion, the problem of long training process and insufficient safety caused by the time-change and high-dimensional characteristics of hybrid vehicles in real driving environments is solved, and the reliability and safety of hybrid vehicle control strategies are improved.
Patent Information
- Application Number
- CN202411509346.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-10-28
AI Technical Summary
In real driving environments, the deep reinforcement learning strategies of hybrid vehicles face the problems of long training processes and insufficient safety and reliability caused by state time-degeneration and high-dimensional characteristics.
Actor-Critic-Guard network architecture is built, and by building a hybrid car model in offline training, multi-camera acquisition and BEV feature fusion, combined with Guard environment cognitive network, we enhance the environmental adaptability and safety of strategy learning, and adopt L2 regular terms to prevent overfitting, so as to perform state cognition and strategy optimization of the online environment.
Improve the reliability and safety of deep reinforcement learning hybrid vehicle control strategies, and enable continuous learning and safety control in complex and changeable real-world environments.
Smart Images

Figure CN119475978B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of new energy vehicles and artificial intelligence, and relates to a method for training a cognitive model of unfamiliar environment states in an online driving scenario. Background Art
[0002] New energy vehicles are regarded as an important measure to promote energy transformation and alleviate the energy crisis. In particular, the development of hybrid electric vehicles (HEVs) is particularly necessary. With the increasing global attention to sustainable development and low - carbon economy, many mainstream and emerging automakers have successively launched pure electric vehicles, hybrid electric vehicles, and fuel cell vehicles. However, hybrid electric vehicles play a key role during the transition period because they can combine the advantages of internal combustion engines and electric drive systems, meeting the requirements of long - distance driving while achieving lower emissions in cities.
[0003] Meanwhile, the development of artificial intelligence technology, especially deep learning and reinforcement learning in machine learning, has gradually attracted the attention of the scientific research community and the industrial community, promoting the development of the AI era. Deep reinforcement learning strategies based on data or sample - driven face three key obstacles: the difference between simulation and reality, credibility, and sample efficiency. In a real - driving environment, the time - varying and high - dimensional characteristics of states make the update process long. Regarding sample efficiency, by timely editing the Markov decision process and connecting key states, the information in the training data becomes more intensive. When facing a real physical system, a dynamic residual model can alleviate the influence of simulation - to - reality dynamics and sensor noise.
[0004] Generally, when deploying a learning - based strategy, it is necessary to implement a backup strategy to ensure safety control. Therefore, some studies have introduced concepts such as uncertainty or confidence to achieve the switching between these two types of strategies. However, a complete solution should include clarifying the applicable scope of data - driven strategies based on historical training samples, evaluating risks or unknown states that occur during testing, and enhancing its applicability through updates in the online environment. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method for training a cognitive model of unfamiliar environment states in an online driving scenario, improving the reliability and safety of the control strategy of a deep reinforcement learning - based hybrid electric vehicle.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for training a cognitive model of unfamiliar environment states in an online driving scenario, specifically including the following steps:
[0008] S1: In the offline training scenario, with adaptive cruise control and lane keeping assist as the primary learning tasks, a hybrid vehicle model is built by loading a three-dimensional driving scenario model. And according to the internal and external parameter matrices of the cameras in the nuScenes dataset, the corresponding RGB camera models are arranged to achieve multi-angle driving image acquisition of the driving position at each moment.
[0009] S2: Before starting the simulation training environment, based on the real-time driving images collected by RGB cameras from six different perspectives, the Camera BEV features of the pure vision feature extraction stream are obtained and combined with the original two-dimensional variable state tensor. Subsequently, using the proximal policy optimization algorithm based on the Actor-Critic architecture as the agent, the concept of "Guard environment cognitive network" is proposed to form an "Actor-Critic-Guard" network architecture that can fully realize the three functions of policy learning, policy evaluation, and environment cognition.
[0010] S3: During the training process of the simulation driving scenario, for the data-driven deep reinforcement learning-based hybrid vehicle control strategy, the traditional Actor-Critic architecture training process continues according to the original plan. At the same time, to effectively train the Guard environment cognitive network, the loss function synchronously includes the loss values and L2 regularization terms regarding Actor, Critic, and Guard. Thus, after the Actor and Critic networks achieve stable fitting to the current training environment, the training environment state features corresponding to the current optimal policy will be recognized.
[0011] S4: After the offline simulation training is completed, a more complex and variable path trajectory similar to the real world is loaded as the test condition, and significant and diverse state deviations are introduced into the test environment, such as: trajectory features, lighting conditions, weather changes, and the state of the host vehicle. Then, the "Guard environment cognitive network" is evaluated in the dynamic driving scenario to determine the applicable scope that needs to be improved and enhanced.
[0012] Furthermore, in step S1, the modeling process of the offline training specifically includes the following steps:
[0013] S11: Based on the training method of the deep reinforcement learning-based control strategy, the state space S, action space A, and reward function R of the adaptive cruise control strategy and the lane keeping assist strategy are set.
[0014] S ACC =(VehSpd,TgtSpd,TrsGear,WhlTrq,CurVat,Slope)
[0015] A ACC =[-1,1] Activation function = tanh
[0016] R ACC = -1 × [abs(A ACC (t) - A Hybrid (t - 1)) + abs(VehSpd(t) - TgtSpd(t))]
[0017] S LKA = (VehSpd, AngDif, YawRat, PreLka, LatOff, CurVat)
[0018] A LKA = [-π / 2, π / 2] Activation function = tanh × π / 2
[0019] R LKA = -1 × [abs(A LKA (t) - A PreLka (t)) + abs(LatOff)]
[0020] Wherein, S ACC is the state of the adaptive cruise control strategy, VehSpd is the current vehicle speed, TgtSpd is the target vehicle speed, TrsGear is the transmission gear, WhlTrq is the torque demand, CurVat is the road curvature, Slope is the road gradient, A ACC is the output action of the deep reinforcement learning, R ACC is the reward of the adaptive cruise control strategy, A Hybrid is the hybrid action that actually controls the vehicle, AngDif is the real-time angle difference, YawRat is the yaw rate, PreLka is the expected steering angle, LatOff is the lateral offset relative to the center lane, S LKA is the state of the lane keeping assist, A LKA is the action of the lane keeping assist, R LKA is the reward of the lane keeping assist, A PreLka is the expected steering action predicted according to the characteristics of the road ahead;
[0021] S12: Load the scenarios for offline training in the autonomous driving 3D simulation software CARLA, and load the Citroen C3 used to represent the vehicle body model. At the same time, build a parallel hybrid vehicle whole vehicle model in the Simulink environment, mainly including a driver model, a controller model (engine, transmission, hybrid control), a power system, and a dynamics model (engine, motor, power battery, transmission, vehicle body, tires, suspension, etc.);
[0022] S13: Arrange RGB cameras at the positions of the six extrinsic parameter matrices corresponding to the Citroen C3 body model to collect real-time scene images of the vehicle driving in a three-dimensional scene, including: front view, front left view, front right view, rear view, rear left view, and rear right view, and obtain images with a size of 1600×900 as perception features.
[0023] Furthermore, step S2 specifically includes the following steps:
[0024] S21: Before starting the simulation training environment, first collect driving images in real time through RGB cameras with six different perspectives; after these images are processed by the BEV Fusion algorithm, BEV features with a dimension of 1×80×128×128 are generated; this process fuses visual information from multiple cameras into a unified bird's-eye view, making subsequent processing more efficient and accurate. To further optimize feature extraction, a convolutional neural network is used to perform dimensionality reduction on S BEV and then separately combine it with variable state tensors such as S ACC and S LKA ; finally, a fully connected network is used to integrate these states to provide multi-modal input information for the training of the deep reinforcement learning-based control strategy;
[0025] S22: The proximal policy optimization, a deep reinforcement learning algorithm based on the Actor-Critic architecture, is used as an agent to learn the energy-saving driving strategy of a hybrid vehicle. The Actor fits the optimal control strategy of the current training environment, and the Critic evaluates the optimality of the current strategy; at the same time, to further enhance the agent's environmental adaptability and recognition ability, actively distinguish strange states that may occur at any time and the dangerous actions that may result from them, and propose the concept of "Guard environmental cognition network"; the core purpose is to identify the state differences caused by the dynamic environment in real time and take timely and safe takeover measures after approaching the boundary of the known space; thus, an "Actor-Critic-Guard" network architecture combining policy learning - policy evaluation - environmental cognition is constructed.
[0026] Furthermore, in step S3, the training steps of the simulation driving scenario are specifically as follows:
[0027] S31: In the training of the data-driven deep reinforcement learning-based hybrid vehicle control strategy, the traditional Actor-Critic architecture training process calculates the update gradient according to the loss function shown in the following formula. Among them, the goal of the Actor is to maximize the expected return, the goal of the Critic is to accurately estimate the value function of the current strategy, and in order to prevent overfitting of the neural network, enhance the generalization of the neural network, and suppress problems such as the plasticity loss of the neural network, an L2 regularization term is added to the loss function of the original value function;
[0028]
[0029] L Critic = E[(V θ (s)-R) 2 +L Critic-L2
[0030] where L Actor is the loss value of the actor network, and L Critic is the loss value of the critic network, π θ (a|s) is the probability distribution of the current policy, is the probability distribution of the old policy, is the advantage estimate, ε is the clipping parameter, V θ (s) is the value function output of the Critic, R is the return value calculated based on the actual reward and advantage calculation, and L Critic-L2 and L Actor-L2 are the L2 regularization terms regarding the weights of the Critic and Actor networks;
[0031] S32: To effectively train the "Guard environmental perception network", the loss function synchronously includes the loss values and L2 regularization terms regarding the Actor, Critic, and Guard. Thus, only after the Actor and Critic networks achieve stable learning of the current training environment, will it start to recognize the state features of the training environment corresponding to the current optimal policy. The training of the "Guard environmental perception network" is mainly based on the following formula L Guard shown below, which specifically includes the comprehensive indicators of the loss functions in multiple aspects:
[0032] L Guard = L Critic + L Actor + L Guard-Label + L Guard-L2
[0033] where L Guard is the loss value of the guardian network, L Guard-Label is the loss of guiding the Guard to learn the environmental perception label, and L Guard-L2 is the L2 regularization term regarding the weights of the Guard network.
[0034] Furthermore, step S4 specifically includes the following steps:
[0035] S41: After the offline simulation training is completed, in order to fully highlight and verify the discrimination effect of the Guard environmental cognition network, a more complex and changeable path trajectory that is close to the real world is loaded as a test condition, and a state deviation with large differences and diversity is introduced to the test environment, such as: trajectory characteristics (more changeable road properties, steeper slopes, sharper curves, etc.), lighting conditions (night, early morning, noon, dusk, etc.), weather changes (sunny, rainy, snowy, foggy, haze, sandstorms, etc.) and vehicle status (higher speed, larger charge and discharge amplitude, more drastic torque and speed change range);
[0036] S42: In the test scenario for dynamic driving environment, verify the evaluation results of the "Guard Environmental Cognitive Network" on the random state tensor, and observe the control effect of the Actor network on the hybrid vehicle; once the real-time state tensor changes outside the range of the historical sample space, promptly start the takeover mode of the safety strategy, and continue to collect unfamiliar samples to complete the continuous learning of the optimal strategy corresponding to the new state in the test scenario. Among them, the most critical is whether the "Guard Environmental Cognitive Network" is accurate and reliable in the dynamic analysis of the current environment tensor, which directly affects the reliability and safety indicators of the deep reinforcement learning hybrid vehicle control strategy.
[0037] The beneficial effect of the present invention is that the "Guard Environmental Cognitive Network" model constructed by the present invention can provide continuous learning of the optimal strategy after inputting more complex and changeable path trajectories that are close to the real world as test conditions and diversified test environments, that is, the present invention greatly improves the reliability and safety of the deep reinforcement learning hybrid vehicle control strategy.
[0038] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0040] Figure 1 It is an overall flow chart of the unfamiliar environment state recognition model training method for online driving scenarios provided by the present invention;
[0041] Figure 2 The Citroen C3 car body model of the unfamiliar environment state cognitive model training method of the present invention;
[0042] Figure 3 is the multi-camera perception image of the training method for the unfamiliar environment state recognition model of the present invention;
[0043] Figure 4 is the Camera BEV feature of the training method for the unfamiliar environment state recognition model of the present invention;
[0044] Figure 5 is the test scenario introducing state differences of the training method for the unfamiliar environment state recognition model of the present invention. Specific Embodiments
[0045] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0046] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0047] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be understood as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0048] Please refer to Figures 1 to 5 , the present invention provides a training method for an unfamiliar environment state recognition model in an online driving scenario, and the process is as Figure 1 shown, and it can be mainly divided into the following stages:
[0049] S1: In the offline training scenario, with adaptive cruise control and lane keeping assist as the primary learning tasks, a hybrid vehicle model is built by loading a three-dimensional driving scenario model, and the corresponding RGB camera models are arranged according to the internal and external parameter matrices of the cameras in the nuScenes dataset, so as to realize the acquisition of multi-angle driving images of the driving position at each moment.
[0050] Step S1 specifically includes the following steps:
[0051] S11: Based on the training method of the deep reinforcement learning-based control strategy, set the state space S, action space A, and reward function R of the adaptive cruise control strategy and the lane keeping assist strategy;
[0052] S ACC =(VehSpd,TgtSpd,TrsGear,WhlTrq,CurVat,Slope)
[0053] A ACC =[-1,1] Activation function = tanh
[0054] R ACC =-1×[abs(A ACC (t)-A Hybrid (t-1))+abs(VehSpd(t)-TgtSpd(t))]
[0055] S LKA =(VehSpd,AngDif,YawRat,PreLka,LatOff,CurVat)
[0056] A LKA =[-π / 2,π / 2] Activation function = tanh×π / 2
[0057] R LKA =-1×[abs(A LKA (t)-A PreLka (t))+abs(LatOff)]
[0058] Among them, S ACC is the state of the adaptive cruise control strategy, VehSpd is the current vehicle speed, TgtSpd is the target vehicle speed, TrsGear is the transmission gear, WhlTrq is the torque demand, CurVat is the road curvature, Slope is the road slope, A ACC is the output action of the deep reinforcement learning, R ACC is the reward of the adaptive cruise control strategy, A Hybridis the true hybrid action that controls the vehicle, AngDif is the real-time angular difference, YawRat is the yaw rate, PreLka is the expected steering angle, LatOff is the lateral offset relative to the center lane, S LKA is the state of lane keeping assist, A LKA is the action of lane keeping assist, R LKA is the reward of lane keeping assist, A PreLka is the expected steering action predicted according to the characteristics of the road ahead;
[0059] S12: Load the scenarios for offline training in the autonomous driving three-dimensional simulation software CARLA, and load the Citroen C3 used to represent the vehicle body model as shown in Figure 2 At the same time, build a parallel hybrid vehicle whole vehicle model in the Simulink environment, mainly including a driver model, a controller model (engine, transmission, hybrid control), a power system, and a dynamics model (engine, motor, power battery, transmission, vehicle body, tires, suspension, etc.).
[0060] S13: Arrange RGB cameras at the positions of the six camera extrinsic parameter matrices corresponding to the Citroen C3 vehicle body model to collect real-time scene images of the vehicle driving in a three-dimensional scene, including: front view, front left view, front right view, rear view, rear left view, and rear right view, and obtain Figure 3 the image with the size of 1600×900 as shown in
[0061] S2: Before starting the simulation training environment, according to the real-time driving images collected by six RGB cameras from different perspectives, obtain the Camera BEV features of the pure vision feature extraction stream, and combine them with the original two-dimensional variable state tensor. Subsequently, use the proximal policy optimization algorithm based on the Actor-Critic architecture as the agent, propose the concept of the Guard environment cognition network, and form the Actor-Critic-Guard network architecture that can fully realize the three functions of policy learning - policy evaluation - environment cognition.
[0062] Step S2 specifically includes the following steps:
[0063] S21: Before starting the simulation training environment, first collect driving images in real time through six RGB cameras from different perspectives. After these images are processed by the BEV Fusion algorithm, generate Figure 4 the BEV features with the dimension of 1×80×128×128 as shown in BEVPerform dimensionality reduction processing and then separately combine with S ACC and S LKA and other variable-type state tensors. Finally, use a fully connected network to integrate these states to provide multi-modal input information for the training of the deep reinforcement learning-based control strategy.
[0064] S22: Deep reinforcement learning algorithm based on the Actor-Critic architecture - Proximal Policy Optimization. As an agent learns the energy-saving driving strategy of a hybrid vehicle, the Actor fits the optimal control strategy of the current training environment, while the Critic evaluates the optimality of the current strategy. At the same time, to further enhance the agent's environmental adaptability and identification ability, actively distinguish unfamiliar states that may occur at any time and the resulting dangerous actions, and propose the concept of "Guard environmental cognition network". The core purpose is to real-time distinguish the state differences caused by the dynamic environment and take timely safety takeover measures after approaching the boundary of the known space. In this regard, an innovative "Actor-Critic-Guard" network architecture combining policy learning - policy evaluation - environmental cognition is constructed.
[0065] S3: During the training process of the simulation driving scenario, for the data-driven deep reinforcement learning-based hybrid vehicle control strategy, the traditional Actor-Critic architecture training process continues according to the original plan. At the same time, to effectively train the Guard environmental cognition network, the loss function synchronously includes the loss values and L2 regularization terms of Actor, Critic, and Guard. Thus, after the Actor and Critic networks achieve stable fitting to the current training environment, the training environment state characteristics corresponding to the current optimal strategy will be recognized.
[0066] Step S3 specifically includes the following steps:
[0067] S31: In the training of the data-driven deep reinforcement learning-based hybrid vehicle control strategy, the traditional Actor-Critic architecture training process calculates and updates the gradient according to the loss function shown in the following formula. Among them, the goal of the Actor is to maximize the expected return, and the goal of the Critic is to accurately estimate the value function of the current strategy. And to prevent overfitting of the neural network, enhance the generalization of the neural network, and suppress the plasticity loss of the neural network and other problems, L2 regularization terms are added to the loss functions of the original value functions.
[0068]
[0069] L Critic = E[(V θ (s)-R) 2 +L Critic-L2
[0070] Among them, L Actor is the loss value of the actor network, and L Critic is the loss value of the critic network. π θ (a|s) is the probability distribution of the current policy, is the probability distribution of the old policy, is the advantage estimation, ε is the clipping parameter, and V θ (s) is the value function output of the Critic. R is the return value calculated based on the actual reward and advantage calculation, and L Critic-L2 and L Actor-L2 are the L2 regularization terms regarding the weights of the Critic and Actor networks.
[0071] S32: To effectively train the "Guard environmental perception network", the loss function synchronously includes the loss values and L2 regularization terms regarding the Actor, Critic, and Guard. Thus, only after the Actor and Critic networks achieve stable learning of the current training environment, will it start to recognize the training environment state features corresponding to the current optimal policy. The training of the "Guard environmental perception network" is mainly based on the following formula L Guard shown below, which specifically includes the comprehensive indicators of loss functions in multiple aspects:
[0072] L Guard = L Critic + L Actor + L Guard-Label + L Guard-L2
[0073] Among them, L Guard is the loss value of the critic network, L Guard-Label is the loss for guiding the Guard to learn the environmental perception label, and L Guard-L2 is the L2 regularization term regarding the weights of the Guard network.
[0074] S4: After the offline simulation training is completed, load a more complex, variable, and approximately real-world path trajectory as the test condition, and introduce significant and diverse state deviations to the test environment, such as: trajectory features, lighting conditions, weather changes, and the state of the ego vehicle. Then, evaluate the Guard environmental perception model in a dynamic driving scenario to determine the applicable scope that needs to be improved and enhanced.
[0075] Step S4 specifically includes the following steps:
[0076] S41: After the offline simulation training is completed, in order to fully highlight and verify the discrimination effect of the Guard environmental perception network, load a more complex, variable, and approximately real-world path trajectory as the test condition, and introduce significant and diverse state deviations to the test environment, such asFigure 5 As shown: trajectory features (more variable road attributes, steeper slopes, sharper curves, etc.), lighting conditions (night, early morning, noon, dusk, etc.), weather changes (sunny, rainy, snowy, foggy, hazy, dusty, etc.) and the state of the host vehicle (higher speed, larger charge and discharge amplitude, wider range of torque and speed changes).
[0077] S42: In the test scenario for a dynamic driving environment, verify the evaluation results of the "Guard environment perception model" for random state tensors, and observe the control effect of the Actor network on the hybrid vehicle. Once the real-time state tensor changes outside the range of the historical sample space, promptly activate the takeover mode of the safety strategy, and continuously collect unfamiliar samples to complete the continuous learning of the optimal strategy corresponding to the new state in the test scenario. Among them, the most crucial thing is whether the dynamic discrimination of the "Guard environment perception network" for the current environment tensor is accurate and reliable, which directly affects the reliability and safety indicators of the control strategy of the deep reinforcement learning-based hybrid vehicle.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A training method for a cognitive model of an unfamiliar environment state in an online driving scenario, characterized in that, The method specifically includes the following steps: S1: In the offline training scenario, taking adaptive cruise control and lane keeping assist as the primary learning tasks, by loading a three-dimensional driving scenario model, building a hybrid vehicle model, and arranging corresponding RGB camera models according to the camera internal and external parameter matrices of the nuScenes dataset, multi-angle driving images of the driving position at each moment are collected; S2: Before starting the simulation training environment, based on the real-time driving images collected from RGB cameras with six different perspectives, obtain the Camera BEV features of the pure vision feature extraction stream and combine them with the original two-dimensional variable state tensor; Subsequently, using the proximal policy optimization algorithm based on the Actor-Critic architecture as the agent, the concept of "Guard environmental perception network" is proposed to form an "Actor-Critic-Guard" network architecture that can fully realize the three functions of policy learning - policy evaluation - environmental perception; Step S2 specifically includes the following steps: S21: Before the simulation training environment is started, first collect driving images in real time through six RGB cameras with different perspectives; after these images are processed by the BEV Fusion algorithm, BEV features with a dimension of 1×80×128×128 are generated; use a convolutional neural network to S BEV perform dimensionality reduction processing, and then separately combine with S ACC and S LKA the variable state tensor; finally, use a fully connected network to integrate these states to provide multi-modal input information for the training of the deep reinforcement learning control strategy; S22: The proximal policy optimization, a deep reinforcement learning algorithm based on the Actor-Critic architecture, is used as the agent to learn the energy-saving driving strategy of the hybrid vehicle. The Actor fits the optimal control strategy of the current training environment, and the Critic evaluates the optimality of the current strategy; At the same time, actively identify the strange states that may occur at any moment and the dangerous actions that may result therefrom, and propose the concept of "Guard environmental perception network"; Real-time distinguish the state differences caused by the dynamic environment, and take timely safety takeover measures after approaching the boundary of the known space; Thus, an "Actor-Critic-Guard" network architecture combining policy learning - policy evaluation - environmental perception is constructed; S3: During the training of the simulation driving scenario, for the data-driven deep reinforcement learning-based hybrid vehicle control strategy, the traditional Actor-Critic architecture training process continues according to the original plan; At the same time, to effectively train the Guard environmental perception network, the loss function synchronously includes the loss values and L2 regularization terms regarding the Actor, Critic, and Guard. Thus, after the Actor and Critic networks stably fit the current training environment, the state features of the training environment corresponding to the current optimal strategy will be recognized; The training steps of the simulation driving scenario are specifically as follows: S31: In the training of the data-driven deep reinforcement learning-based hybrid vehicle control strategy, the traditional Actor-Critic architecture training process calculates and updates the gradient according to the loss function shown in the following formula. Among them, the goal of the Actor is to maximize the expected return, and the goal of the Critic is to accurately estimate the value function of the current strategy. L2 regularization terms are added to the loss functions of the original value functions; Among them, is the loss value of the actor network, is the loss value of the critic network, is the probability distribution of the current policy, is the probability distribution of the old policy, is the advantage estimation, is the clipping parameter, is the value function output of the Critic, R is the return value calculated based on the actual reward and advantage calculation, L Critic-L2 and L Actor-L2 is the L2 regularization term regarding the weights of the Critic and Actor networks; S32: To effectively train the "Guard environmental perception network", the loss function synchronously includes the loss values and L2 regularization terms for the Actor, Critic, and Guard. Thus, only after the Actor and Critic networks achieve stable learning of the current training environment, will it start to recognize the training environment state features corresponding to the current optimal policy. The training of the "Guard environmental perception network" is based on the following formula L Guard as shown, specifically including the comprehensive indicators of the loss function in multiple aspects: Among them, is the loss value of the Guardian network, L Guard-Label is the loss that guides the Guard to learn the environmental cognitive labels, while L Guard-L2 is the L2 regularization term regarding the weights of the Guard network; S4: After the offline simulation training is completed, load a complex, variable, and nearly real-world path trajectory as the test condition, and introduce significant and diverse state deviations to the test environment. Then, evaluate the "Guard environmental perception network" in a dynamic driving scenario to determine the applicable scope that needs to be improved and enhanced.
2. The training method for the unfamiliar environment state recognition model in the online driving scenario according to claim 1, wherein In step S1, the modeling process of the offline training specifically includes the following steps: S11: Set the state space, S S action space, A A and reward function R R of the adaptive cruise control strategy and the lane keeping assist strategy based on the training method of the deep reinforcement learning-based control strategy; Among them, S ACC is the state of the adaptive cruise control strategy, VehSpd is the current vehicle speed, TgtSpd is the target vehicle speed, TrsGear is the transmission gear position, WhlTrq is the torque demand, CurVat is the road curvature, Slope is the road gradient, A ACC is the output action of the deep reinforcement learning, R ACC is the reward of the adaptive cruise control strategy, A Hybrid is the hybrid action that actually controls the vehicle, AngDif is the real-time angular difference, YawRat is the yaw rate, PreLka is the expected steering angle, LatOff is the lateral offset relative to the center lane, S LKA is the state of the lane keeping assist, A LKA is the action of the lane keeping assist, R LKA is the reward of the lane keeping assist, A PreLka is the expected steering action predicted according to the road characteristics ahead; S12: In the autonomous driving three-dimensional simulation software CARLA, load the scenario for offline training, and load the Citroen C3 used to represent the vehicle body model. At the same time, build a parallel hybrid vehicle integrated model in the Simulink environment, including a driver model, a controller model, a power system, and a dynamics model; S13: Arrange RGB cameras at the positions of the six camera extrinsic parameter matrices corresponding to the Citroen C3 vehicle body model to collect real-time scene images of the vehicle driving in a three-dimensional scene, including: front view, front left view, front right view, rear view, rear left view, and rear right view, and obtain images with a size of 1600×900 as perception features.
3. The training method for the unfamiliar environment state recognition model in the online driving scenario according to claim 1, characterized in that, Step S4 specifically includes the following steps: S41: After the offline simulation training is completed, load a complex, variable, and nearly real-world path trajectory as the test condition, and introduce significant and diverse state deviations to the test environment; S42: In the test scenario facing the dynamic driving environment, verify the evaluation results of the "Guard environmental perception network" for the random state tensor, and observe the control effect of the Actor network on the hybrid vehicle; once the real-time state tensor changes outside the range of the historical sample space, immediately activate the takeover mode of the safety strategy, and continuously collect unfamiliar samples to complete the continuous learning of the optimal strategy corresponding to the new state in the test scenario.
Citation Information
Patent Citations
Man-machine co-driving control right decision-making method based on man-vehicle risk state
CN113335291A
Intelligent network connection beyond visual range driving auxiliary system based on 5G position sharing
CN115412883A