Intervention control adaptive learning method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202611273931.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-21
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]目前车辆运动控制系统通常采用基于规则和标定参数的控制策略,即根据车辆横摆角速度误差、侧偏角、轮胎滑移率、侧向加速度等状态参数设定固定的介入门限,当车辆状态超出预设阈值时触发控制系统介入;为了保证车辆安全性,上述控制参数通常由专业试验工程师和标定人员在典型工况下通过大量实车试验确定,并最终固化于介入时机自适应学习器中;然而由于不同驾驶员在驾驶习惯、风险偏好以及车辆操控需求等方面存在显著差异,固定标定参数难以同时满足不同驾驶风格驾驶员的需求
[0016]本发明的介入控制自适应学习方法的有益效果是:根据驾驶员操作信号确定驾驶员激进指数,并利用驾驶员激进指数动态确定介入门限参数和干预强度参数,使车辆运动控制系统能够针对不同驾驶风格采用差异化控制策略;同时结合车辆响应结果和驾驶员反馈信息持续更新驾驶员激进指数与控制参数之间的映射关系,实现车辆运动控制系统介入时机和干预强度的自适应学习,从而在保证车辆稳定性的基础上提高驾驶员对车辆运动控制系统的接受度和主观驾驶满意度。
Smart Images

Figure CN122808760A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle technology, and more specifically, to an intervention control adaptive learning method, device, electronic device, and storage medium. Background Technology
[0002] Vehicle motion control systems are an important component of active safety control in modern intelligent vehicles. They primarily improve vehicle handling stability and driving safety by monitoring and controlling the vehicle's lateral and longitudinal dynamic states in real time and actively intervening when the vehicle experiences unstable states such as understeer, oversteer, or drive wheel slippage.
[0003] Currently, vehicle motion control systems typically employ rule-based and calibration parameter-based control strategies. This involves setting fixed intervention thresholds based on state parameters such as vehicle yaw rate error, sideslip angle, tire slip ratio, and lateral acceleration. When the vehicle's state exceeds these preset thresholds, the control system is triggered to intervene. To ensure vehicle safety, these control parameters are usually determined by professional test engineers and calibration personnel through extensive real-vehicle testing under typical operating conditions and are ultimately embedded in the intervention timing adaptive learner. However, due to significant differences in driving habits, risk preferences, and vehicle handling needs among different drivers, fixed calibration parameters cannot simultaneously meet the needs of drivers with different driving styles. Summary of the Invention
[0004] The problem addressed by this invention is how to achieve adaptive adjustment of the timing and intensity of vehicle motion control intervention.
[0005] To address the aforementioned problems, this invention provides an intervention control adaptive learning method, apparatus, electronic device, and storage medium.
[0006] In a first aspect, the present invention provides an adaptive learning method for interventional control, comprising: The driver aggression index, which characterizes driving style, is determined based on driver operation signals. The intervention threshold parameter and intervention intensity parameter are determined based on the driver aggression index. The intervention condition is determined based on the vehicle stability error and the intervention threshold parameter. When the intervention condition is met, a vehicle motion control command is generated based on the intervention intensity parameter. The mapping relationship between the driver aggression index, the intervention threshold parameter, and the intervention intensity parameter is updated based on the vehicle response results and driver feedback information after the vehicle motion control command is executed.
[0007] Optionally, determining the driver aggression index characterizing driving style based on driver operation signals includes: Extract driving behavior feature parameters from driver operation signals within a preset time window; The driver aggression index is determined based on the driving behavior characteristic parameters.
[0008] Optionally, the extraction of driving behavior feature parameters from driver operation signals within a preset time window includes: The driver's operation characteristics are determined based on the driver's operation signals. The driver's operation characteristics include at least one of the following: steering wheel angular velocity characteristics, brake pedal rate of change characteristics, accelerator pedal rate of change characteristics, and steering wheel angular acceleration characteristics. The driver's operational characteristics are statistically processed to obtain the driving behavior characteristic parameters.
[0009] Optionally, determining the driver aggression index based on the driving behavior characteristic parameters includes: The driving behavior feature parameters are normalized to obtain the corresponding normalized feature parameters; The normalized feature parameters are weighted and fused according to preset weights to obtain the driver's aggressive index.
[0010] Optionally, determining the intervention threshold parameter and intervention intensity parameter based on the driver aggression index includes: The intervention threshold parameters are determined based on the driver aggression index and the standard intervention threshold. The intervention intensity parameter is determined based on the driver aggression index and the steepness coefficient, wherein the steepness coefficient is used to control the steepness of the curve corresponding to the intervention intensity parameter.
[0011] Optionally, updating the mapping relationship between the driver aggression index and the intervention threshold parameter and the intervention intensity parameter based on the vehicle response results and driver feedback information includes: A reward function is constructed based on the vehicle response results and the driver feedback information; The mapping relationship between the driver's aggression index, the intervention threshold parameter, and the intervention intensity parameter is updated based on the reward function.
[0012] Optionally, updating the mapping relationship between the driver aggression index, the intervention threshold parameter, and the intervention intensity parameter based on the reward function includes: A reinforcement learning state space is constructed based on the driver's aggressiveness index, vehicle stability error, vehicle speed, and road adhesion coefficient. Based on the reward function and the reinforcement learning state space, the adjustment actions for the weights of the intervention threshold parameter, the intervention intensity parameter, and the driving behavior feature parameter are determined, and the mapping relationship is updated according to the adjustment actions.
[0013] In a second aspect, the present invention provides an intervention control adaptive learning device, comprising: The first module is used to determine the driver aggression index, which characterizes driving style, based on driver operation signals. The second module is used to determine the intervention threshold parameter and intervention intensity parameter based on the driver aggression index, determine whether the intervention conditions are met based on the vehicle stability error and the intervention threshold parameter, and generate vehicle motion control commands based on the intervention intensity parameter when the intervention conditions are met. The third module is used to update the mapping relationship between the driver aggression index, the intervention threshold parameter, and the intervention intensity parameter based on the vehicle response result and driver feedback information after the vehicle motion control command is executed.
[0014] Thirdly, the present invention provides an electronic device, including a memory and a processor; The memory is used to store computer programs; The processor is configured to implement the intervention control adaptive learning method as described in the first aspect when executing the computer program.
[0015] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the intervention control adaptive learning method as described in the first aspect.
[0016] The beneficial effects of the adaptive learning method for intervention control in this invention are as follows: the driver's aggressive index is determined based on the driver's operation signal, and the intervention threshold parameter and intervention intensity parameter are dynamically determined using the driver's aggressive index, enabling the vehicle motion control system to adopt differentiated control strategies for different driving styles; at the same time, the mapping relationship between the driver's aggressive index and control parameters is continuously updated by combining vehicle response results and driver feedback information, realizing adaptive learning of the intervention timing and intervention intensity of the vehicle motion control system, thereby improving the driver's acceptance of the vehicle motion control system and subjective driving satisfaction while ensuring vehicle stability. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the intervention control adaptive learning method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the process for determining the driver's aggressiveness index according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the process for extracting driving behavior feature parameters according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the specific process of determining the driver's aggressiveness index according to an embodiment of the present invention; Figure 5 This is a schematic diagram illustrating the process of determining the intervention threshold parameter and intervention intensity parameter according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the process for updating the mapping relationship according to an embodiment of the present invention; Figure 7 This is a schematic diagram illustrating the specific process of updating the mapping relationship according to an embodiment of the present invention; Figure 8 This is a system architecture diagram of the intervention control adaptive learning device according to an embodiment of the present invention; Figure 9 This is a system architecture diagram of an electronic device according to an embodiment of the present invention; Figure 10 This is a schematic diagram illustrating the principle of the intervention control adaptive learning method according to an embodiment of the present invention. Detailed Implementation
[0018] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0019] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0020] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0021] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0023] like Figure 1 As shown in the figure, an adaptive learning method for interventional control provided by an embodiment of the present invention includes: S100: Determines the driver aggression index based on driver operation signals to characterize driving style.
[0024] Specifically, in combination Figure 10 As shown, in the DiL test bench (Driver-in-the-Loop simulation test bench), with a sampling period of, for example, T... s =10ms to collect driver operation signals, including steering wheel angular velocity. Brake pedal speed Accelerator pedal change rate and steering wheel angular acceleration (High-frequency impact) The driver aggression index, which characterizes driving style, is determined based on the driver operation signals mentioned above.
[0025] The DiL bench is a testing platform that introduces real drivers into a closed-loop vehicle simulation. It typically includes a driving simulator, a real-time simulation computer, a vehicle dynamics model, and a visual feedback system. The driver performs driving operations through the steering wheel and pedals. The real-time simulation computer calculates the vehicle's operating status based on the driving operations and outputs the vehicle response to the driver through the visual feedback system, thus forming a closed-loop interactive environment between the driver and the vehicle model.
[0026] S200: Determine the intervention threshold parameter and intervention intensity parameter based on the driver aggression index, determine whether the intervention conditions are met based on the vehicle stability error and the intervention threshold parameter, and generate vehicle motion control commands based on the intervention intensity parameter when the intervention conditions are met.
[0027] Specifically, the corresponding intervention threshold parameter and intervention intensity parameter are determined based on the driver's aggression index. Whether the intervention conditions are met is determined based on the vehicle stability error and the intervention threshold parameter. The vehicle stability error can be represented by the yaw rate error, for example, by defining the yaw rate error. ,in, This represents the actual yaw rate, and the reference yaw rate corresponding to the driver's intention. Determined by the vehicle's two-degree-of-freedom model and the driver's steering wheel angle, when intervention conditions are met, vehicle motion control commands are generated based on intervention intensity parameters.
[0028] S300: Update the mapping relationship between the driver aggression index, the intervention threshold parameter, and the intervention intensity parameter based on the vehicle response result and driver feedback information after executing the vehicle motion control command.
[0029] Specifically, after control is completed, vehicle response results and driver feedback information are obtained. The vehicle response results include the change in yaw rate error, path tracking error, and vehicle stability indicators. The driver feedback information includes the driver's subjective score, which can be evaluated in a 1-5 point range. The mapping relationship between the aggressive index and control parameters is updated using the vehicle response results and the driver's subjective score.
[0030] In this embodiment, the driver aggression index is determined based on the driver's operation signal, and the intervention threshold parameter and intervention intensity parameter are dynamically determined using the driver aggression index. This enables the vehicle motion control system to adopt differentiated control strategies for different driving styles. At the same time, the mapping relationship between the driver aggression index and control parameters is continuously updated by combining vehicle response results and driver feedback information. This achieves adaptive learning of the intervention timing and intensity of the vehicle motion control system, thereby improving the driver's acceptance of the vehicle motion control system and subjective driving satisfaction while ensuring vehicle stability.
[0031] Optionally, determining the driver aggression index characterizing driving style based on driver operation signals includes: S110: Extract driving behavior feature parameters from driver operation signals within a preset time window.
[0032] Specifically, in combination Figure 2 As shown, for the steering wheel angular velocity Extract the past =Statistical characteristics within a 3s window: ; in, To the steering wheel angular velocity Corresponding driving behavior characteristic parameters For exponentially decaying weights (the more recent the operation, the higher the weight), N = / T s , This indicates the preset time window length, τ=1.0.
[0033] The same calculation applies to brake pedal speed. The corresponding driving behavior characteristic parameter F β , and the rate of change of the accelerator pedal The corresponding driving behavior characteristic parameter F α and the angular acceleration of the steering wheel Corresponding high-frequency impact factor Among them, percentile 90 This represents the 90th percentile.
[0034] S120: Determine the driver aggression index based on the driving behavior characteristic parameters.
[0035] Specifically, the driving behavior characteristic parameters are normalized, and the driver aggression index is determined based on the normalized driving behavior characteristic parameters and their corresponding weights.
[0036] In this optional embodiment, by extracting driving behavior feature parameters and determining the driver's aggressiveness index within a preset time window, the driver's driving style can be characterized by driving behavior over a continuous time period, reducing the impact of instantaneous operation fluctuations on the driving style recognition results and improving the stability and accuracy of the driver's aggressiveness index.
[0037] Optionally, the extraction of driving behavior feature parameters from driver operation signals within a preset time window includes: S111: Determine driver operation characteristics based on the driver operation signal, wherein the driver operation characteristics include at least one of steering wheel angular velocity characteristics, brake pedal rate of change characteristics, accelerator pedal rate of change characteristics, and steering wheel angular acceleration characteristics.
[0038] Specifically, in combination Figure 3 As shown, referring to the above content, driver operation signals include steering wheel angular velocity. Brake pedal speed Accelerator pedal change rate and steering wheel angular acceleration The driver's operating characteristics can be determined based on the driver's operating signals. The driver's operating characteristics include at least one of the following: steering wheel angular velocity characteristics, brake pedal rate of change characteristics, accelerator pedal rate of change characteristics, and steering wheel angular acceleration characteristics.
[0039] S112: Perform statistical processing on the driver's operation characteristics to obtain the driving behavior characteristic parameters.
[0040] Specifically, the weighted average value and frequency domain energy distribution of each driver's operating characteristics within the time window are calculated to obtain the feature vector as driving behavior feature parameters.
[0041] In this optional embodiment, by extracting at least one of the steering wheel angular velocity features, brake pedal rate of change features, accelerator pedal rate of change features, and steering wheel angular acceleration features, and performing statistical processing on them, the driver's operating habits can be characterized from multiple dimensions such as steering, braking, and power requests. This allows the driving behavior feature parameters to more comprehensively reflect the driver's driving style characteristics and improve the accuracy of driving style recognition.
[0042] Optionally, determining the driver aggression index based on the driving behavior characteristic parameters includes: S121: Normalize the driving behavior feature parameters to obtain the corresponding normalized feature parameters.
[0043] Specifically, in combination Figure 4 As shown, the driving behavior feature parameters are normalized to obtain the corresponding normalized feature parameters: , ; in, The maximum boundary is obtained in advance through statistical analysis of a large amount of driver data. The minimum boundary is obtained in advance through statistical analysis of a large amount of driver data; when k takes different values... hour, These represent driving behavior characteristic parameters. , , , , These represent driving behavior characteristic parameters. , , , The corresponding normalized feature parameters.
[0044] S122: The normalized feature parameters are weighted and fused according to preset weights to obtain the driver's aggressive index.
[0045] Specifically, driver aggression index The final result is a weighted sum: ; In conventional road driving conditions where steering is the primary operation (such as overtaking and lane changing), =0.4, =0.2, =0.2, =0.2, and the sum of the weights is 1.
[0046] Among them, the driver's aggressiveness index is recorded as This is used to characterize the driver's driving style at time t; when time variation is not emphasized, the driver's aggressiveness index can also be simplified as... ; Yaw rate error Similarly.
[0047] In this optional embodiment, by normalizing the driving behavior feature parameters and using preset weights for weighted fusion, the influence of differences in the dimensions and numerical ranges of different feature parameters on the calculation results can be eliminated. At the same time, different weights are used to reflect the degree of contribution of each feature to the driving style, thereby improving the rationality and robustness of the driver aggression index calculation results.
[0048] Optionally, determining the intervention threshold parameter and intervention intensity parameter based on the driver aggression index includes: S210: Determine the intervention threshold parameter based on the driver aggression index and the standard intervention threshold.
[0049] Specifically, in combination Figure 5 As shown, in related technologies, when the yaw rate error | Traditional fixed threshold At this time, VMC (Vehicle Motion Control) intervenes. In this embodiment, a dynamic threshold is used, i.e., the intervention threshold parameter. Represented as: ; in, Indicates the standard intervention threshold (e.g., 0.1 rad / s); in, , representing the radical sensitivity coefficient, ranges from 0.06 to 0.10; As you can see, When the threshold is equal to 1 (indicating extreme radicalism), it increases to... +0.5 Allow for larger errors; When the threshold is 0, the threshold is lowered, allowing for earlier intervention.
[0050] S220: Determine the intervention intensity parameter based on the driver aggression index and the steepness coefficient, wherein the steepness coefficient is used to control the curve steepness corresponding to the intervention intensity parameter.
[0051] Specifically, VMC braking intervention is typically caused by target-added yaw moment. Converted to cylinder pressure slope Traditionally, the calibration is based on a fixed slope. In this embodiment, a slope adjustment factor is defined. : ; Among them, the steepness coefficient λ = 4.0 is used to control the steepness of the S-curve; when When γ = 0, γ ≈ 0.27, the slope of pressure buildup decreases sharply, and gentle intervention is needed; when When γ = 1, γ ≈ 1.73, the slope of the pressure buildup increases, and decisive intervention is warranted.
[0052] The final actual slope of the pressure buildup is used as the intervention intensity parameter. : ; in The nominal slope (e.g., 50 bar / s).
[0053] In this optional embodiment, the intervention threshold parameter is determined based on the driver's aggressiveness index and the standard intervention threshold, and the intervention intensity parameter is determined based on the driver's aggressiveness index and the steepness coefficient, so that the intervention boundary and intervention intensity of the vehicle motion control system can be dynamically adjusted according to the driving style. For drivers with a more aggressive driving style, the timing of control intervention can be appropriately delayed and the abruptness of intervention can be reduced. For drivers with a more conservative driving style, intervention can be initiated earlier and the control effect can be enhanced, thereby taking into account both vehicle safety and driving experience.
[0054] Optionally, updating the mapping relationship between the driver aggression index and the intervention threshold parameter and the intervention intensity parameter based on the vehicle response results and driver feedback information includes: S310: Construct a reward function based on the vehicle response result and the driver feedback information.
[0055] Specifically, in combination Figure 6 As shown, k in the above formula ε ,λ, The parameters are not fixed, but are automatically optimized through reinforcement learning in the DiL virtual environment: (1) State space: s=[ , V x ,μ est ]; (2) Action space: a = [k ε ,λ, , , , The fine-tuning amount, , , , To calculate the driver's aggressiveness index The weights of each item; (3) Reward function: R = R obj -α1* Abruptness of intervention -α2* Path deviation integral; in, It is the driver's aggression index (real-time calculated value, between 0 and 1, where 0 indicates extremely conservative and 1 indicates extremely aggressive). It is the yaw rate error, that is, the difference between the actual yaw rate and the reference yaw rate corresponding to the driver's intention; V x It is the vehicle's longitudinal speed (usually taken as an absolute value); μ est This is an estimated value for the road surface adhesion coefficient, reflecting the maximum usable coefficient of friction between the tire and the road surface (approximately 0.8~1.0 for dry asphalt, 0.2~0.3 for snow, etc.); R obj For the driver's subjective rating, α1 represents the proportional coefficient corresponding to the abruptness of intervention, and α2 represents the proportional coefficient corresponding to the path deviation integral.
[0056] Among them, combined Figure 10 As shown, offline reinforcement learning optimization can learn the optimal mapping relationship between the driver's aggression index and the VMC control parameters, that is, to fine-tune the existing control parameters. For example, in a certain driving situation, PPO slightly increases the intervention threshold while reducing the intervention intensity. In the end, the driver's score improves, the intervention abruptness decreases (for example, sudden intervention of VMC may lead to increased skin conductance and changes in heart rate. These biosignals can be used to determine whether the driver has a significant tension or fright reaction, which can be used as one of the objective representations of intervention abruptness), the path deviation does not increase significantly, and the vehicle remains stable. This can increase the probability of taking the same action again under similar conditions.
[0057] Among them, combined Figure 10 As shown, in the offline reinforcement learning optimization stage, the PPO algorithm is used to optimize the aggressive sensitivity coefficient k. ε Steepness coefficient λ and weights (When i takes the values 1, 2, 3, and 4, the weights are respectively) , , , After optimization, during online operation, the system continuously collects driver input signals such as steering, braking, and acceleration, extracts statistical features from a preset time window, and then normalizes and weights these features (using optimized weights). The driver's aggressiveness index at the current moment is obtained, and then the dynamic intervention threshold is determined based on the driver's aggressiveness index at the current moment and the optimized aggressiveness sensitivity coefficient k. And determine the slope adjustment factor based on the optimized steepness coefficient λ; determine the absolute value of the yaw rate error through vehicle state estimation. When the value exceeds the dynamic intervention threshold, the VMC intervenes and applies torque according to the dynamic slope adjustment factor; otherwise, the VMC does not intervene, maintaining driver control, and finally feeding back the vehicle dynamics response in a closed loop to the driver.
[0058] Among them, the abruptness of intervention refers to the comprehensive characterization index of the subjective discomfort and objective dynamic intensity caused by the sudden change in the longitudinal or lateral dynamic state of the vehicle when switching from manual driving or the original control state to automatic control or assisted control. This index is usually used to describe whether the control intervention is smooth, reflecting the suddenness and impact of the change in the vehicle state when the system performs takeover or correction actions. The larger the value, the more intense the intervention and the less smooth the ride experience.
[0059] The path deviation integral refers to a quantitative indicator of the cumulative deviation of the vehicle's actual driving trajectory from the reference target path over time. This indicator is typically obtained by integrating the lateral deviation or trajectory error over the entire test time interval, reflecting the overall accuracy and stability of the control system during path keeping or trajectory tracking. A larger path deviation integral indicates a longer time or a greater magnitude of deviation from the reference path, resulting in poorer path tracking performance; conversely, a smaller integral indicates better path keeping capability.
[0060] S320: Update the mapping relationship between the driver's aggression index, the intervention threshold parameter, and the intervention intensity parameter based on the reward function.
[0061] Specifically, the PPO algorithm (Proximal Policy Optimization, reinforcement learning algorithm) is used for training, and a mapping model is finally output. This model can be embedded and deployed in the vehicle controller to run in real time on the actual vehicle in the form of a lightweight neural network or lookup table. The model updates the mapping relationship between the driver's aggression index and the intervention threshold parameter and intervention intensity parameter based on the reward function, thereby realizing the VMC intervention timing and intensity control that adapts to the driver's style.
[0062] In this optional embodiment, a reward function is constructed by combining vehicle response results and driver feedback information, and the mapping relationship between the driver's aggressive index and control parameters is updated using the reward function. This allows the control strategy optimization process to not only consider objective indicators such as vehicle stability, but also to incorporate the driver's subjective evaluation results, thereby achieving synergistic optimization of objective performance and subjective experience and improving the personalized adaptability of the control strategy.
[0063] Optionally, updating the mapping relationship between the driver aggression index, the intervention threshold parameter, and the intervention intensity parameter based on the reward function includes: S321: Construct a reinforcement learning state space based on the driver's aggressiveness index, vehicle stability error, vehicle speed, and road adhesion coefficient. Specifically, in combination Figure 7 As shown, in the driver-in-the-loop simulation, the agent of the PPO algorithm uses the current driver's aggressiveness index. Yaw rate error Vehicle speed Vx and road surface adhesion coefficient μ est As a state input, the output performs fine-tuning actions on the VMC intervention threshold, intervention slope, and weighting coefficients.
[0064] S322: Based on the reward function and the reinforcement learning state space, determine the adjustment action for the weights of the intervention threshold parameter, the intervention intensity parameter, and the driving behavior feature parameter, and update the mapping relationship according to the adjustment action.
[0065] Specifically, after each complete driving condition, the subjective score given by the driver (e.g., 1-5 points) is quantified as a reward signal. PPO continuously updates the policy network parameters by maximizing the cumulative reward, and finally trains a policy model that can map the optimal VMC intervention parameters according to the real-time state. The weights of intervention threshold parameters, intervention intensity parameters, and driving behavior characteristic parameters can be adjusted to achieve adaptive VMC intervention timing and intensity control based on the driver's style.
[0066] In this optional embodiment, a reinforcement learning state space is constructed using the driver's aggressiveness index, vehicle stability error, vehicle speed, and road adhesion coefficient. Based on the reward function and the state space, the adjustment actions for the intervention threshold parameter, intervention intensity parameter, and driving behavior feature parameter weights are determined. This enables the learning process to comprehensively consider driver style characteristics, vehicle operating status, and road environment factors, achieving joint optimization of control parameters and feature weights. This improves the accuracy and convergence efficiency of mapping relationship updates and enhances the vehicle motion control system's adaptability to different drivers and different operating conditions.
[0067] During the online operation phase, the trained and solidified driver style mapping model is deployed to the vehicle controller or domain controller. In actual driving, driver operation signals and vehicle status information are collected in real time at a preset sampling period, and the driver aggression index is calculated online. Subsequently, the driver aggression index, along with state variables such as vehicle yaw rate error, vehicle speed, and road adhesion coefficient, are input into the mapping model. The corresponding intervention threshold parameters and intervention intensity parameters are output in real time, and the vehicle stability error is judged based on the dynamic intervention threshold. When the intervention conditions are met, the corresponding vehicle motion control command is generated to correct the vehicle's lateral stability or trajectory deviation. At the same time, by smoothing the intervention intensity parameters and limiting the control output, the continuity and smoothness of the intervention process are ensured. Thus, while ensuring vehicle stability, adaptive and real-time intervention control for different driving styles is achieved.
[0068] like Figure 8As shown, an embodiment of the present invention provides an intervention control adaptive learning device 800, comprising: The first module 810 is used to determine the driver aggression index, which characterizes driving style, based on driver operation signals. The second module 820 is used to determine the intervention threshold parameter and intervention intensity parameter based on the driver aggression index, determine whether the intervention conditions are met based on the vehicle stability error and the intervention threshold parameter, and generate vehicle motion control commands based on the intervention intensity parameter when the intervention conditions are met. The third module 830 is used to update the mapping relationship between the driver aggression index, the intervention threshold parameter, and the intervention intensity parameter based on the vehicle response results and driver feedback information.
[0069] like Figure 9 As shown, an electronic device 900 provided in this embodiment of the invention includes a memory 920 and a processor 910; the memory 920 is used to store a computer program; the processor 910 is used to implement the intervention control adaptive learning method as described above when the computer program is executed.
[0070] Alternatively, an electronic device 900 includes a memory 920 and a processor 910 coupled to the memory 920; the memory 920 is configured to store a computer program; and the processor 910 is configured to perform the following operations when the computer program is executed: The driver aggression index, which characterizes driving style, is determined based on driver operation signals. The intervention threshold parameter and intervention intensity parameter are determined based on the driver aggression index. The intervention condition is determined based on the vehicle stability error and the intervention threshold parameter. When the intervention condition is met, a vehicle motion control command is generated based on the intervention intensity parameter. The mapping relationship between the driver aggression index, the intervention threshold parameter, and the intervention intensity parameter is updated based on the vehicle response results and driver feedback information after the vehicle motion control command is executed.
[0071] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the intervention control adaptive learning method as described above.
[0072] Alternatively, a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the following operations: The driver aggression index, which characterizes driving style, is determined based on driver operation signals. The intervention threshold parameter and intervention intensity parameter are determined based on the driver aggression index. The intervention condition is determined based on the vehicle stability error and the intervention threshold parameter. When the intervention condition is met, a vehicle motion control command is generated based on the intervention intensity parameter. The mapping relationship between the driver aggression index, the intervention threshold parameter, and the intervention intensity parameter is updated based on the vehicle response results and driver feedback information after the vehicle motion control command is executed.
[0073] The present invention will now be described an electronic device 900 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. Electronic device 900 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 900 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0074] Electronic device 900 includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) or a computer program loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0075] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention according to actual needs. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units can be implemented in hardware or as software functional units.
[0076] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. An adaptive learning method for interventional control, characterized in that, include: The driver aggression index, which characterizes driving style, is determined based on driver operation signals. The intervention threshold parameter and intervention intensity parameter are determined based on the driver aggression index. The intervention condition is determined based on the vehicle stability error and the intervention threshold parameter. When the intervention condition is met, a vehicle motion control command is generated based on the intervention intensity parameter. The mapping relationship between the driver aggression index, the intervention threshold parameter, and the intervention intensity parameter is updated based on the vehicle response results and driver feedback information after the vehicle motion control command is executed.
2. The intervention control adaptive learning method according to claim 1, characterized in that, The driver aggression index, which characterizes driving style, is determined based on driver operation signals, including: Extract driving behavior feature parameters from driver operation signals within a preset time window; The driver aggression index is determined based on the driving behavior characteristic parameters.
3. The intervention control adaptive learning method according to claim 2, characterized in that, The extraction of driving behavior feature parameters from driver operation signals within a preset time window includes: The driver's operation characteristics are determined based on the driver's operation signals. The driver's operation characteristics include at least one of the following: steering wheel angular velocity characteristics, brake pedal rate of change characteristics, accelerator pedal rate of change characteristics, and steering wheel angular acceleration characteristics. The driver's operational characteristics are statistically processed to obtain the driving behavior characteristic parameters.
4. The intervention control adaptive learning method according to claim 2, characterized in that, Determining the driver aggression index based on the driving behavior characteristic parameters includes: The driving behavior feature parameters are normalized to obtain the corresponding normalized feature parameters; The normalized feature parameters are weighted and fused according to preset weights to obtain the driver's aggressive index.
5. The intervention control adaptive learning method according to claim 1, characterized in that, The determination of the intervention threshold parameter and intervention intensity parameter based on the driver aggression index includes: The intervention threshold parameters are determined based on the driver aggression index and the standard intervention threshold. The intervention intensity parameter is determined based on the driver aggression index and the steepness coefficient, wherein the steepness coefficient is used to control the steepness of the curve corresponding to the intervention intensity parameter.
6. The intervention control adaptive learning method according to claim 1, characterized in that, The step of updating the mapping relationship between the driver aggression index and the intervention threshold parameter and the intervention intensity parameter based on the vehicle response results and driver feedback information includes: A reward function is constructed based on the vehicle response results and the driver feedback information; The mapping relationship between the driver's aggression index, the intervention threshold parameter, and the intervention intensity parameter is updated based on the reward function.
7. The intervention control adaptive learning method according to claim 6, characterized in that, The step of updating the mapping relationship between the driver's aggression index, the intervention threshold parameter, and the intervention intensity parameter based on the reward function includes: A reinforcement learning state space is constructed based on the driver's aggressiveness index, vehicle stability error, vehicle speed, and road adhesion coefficient. Based on the reward function and the reinforcement learning state space, the adjustment actions for the weights of the intervention threshold parameter, the intervention intensity parameter, and the driving behavior feature parameter are determined, and the mapping relationship is updated according to the adjustment actions.
8. An intervention control adaptive learning device, characterized in that, include: The first module is used to determine the driver aggression index, which characterizes driving style, based on driver operation signals. The second module is used to determine the intervention threshold parameter and intervention intensity parameter based on the driver aggression index, determine whether the intervention conditions are met based on the vehicle stability error and the intervention threshold parameter, and generate vehicle motion control commands based on the intervention intensity parameter when the intervention conditions are met. The third module is used to update the mapping relationship between the driver aggression index, the intervention threshold parameter, and the intervention intensity parameter based on the vehicle response result and driver feedback information after the vehicle motion control command is executed.
9. An electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to implement the intervention control adaptive learning method as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the intervention control adaptive learning method as described in any one of claims 1 to 7.