A hybrid expert following decision method and system driven by bionic memory
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HARBIN INST OF TECH
- Filing Date
- 2025-10-13
- Publication Date
- 2026-08-07
AI Technical Summary
[0002]跟驰决策是纵向控制的核心模块,其根本任务是在保证安全、效率与舒适性之间实现动态平衡,跟驰决策技术的发展经历了从传统模型到现代学习方法的演进,早期的跟驰模型,如安全距离模型、智能驾驶员模型等,主要基于明确的物理规则和数学公式,虽然具有良好的可解释性,但在应对复杂多变的真实交通环境时显得过于僵化;目前,Wen等人利用IQ-Learn 算法结合 SAC 实现了人类驾驶跟驰行为模拟,发现人在跟随自动驾驶车辆时存在显著策略调整,但并未进一步将这些策略差异转化为自动驾驶系统的响应机制;Zhang等人提出了一种双阶段跟驰策略,将安全性作为前置条件嵌入到能耗优化过程中,其有效降低了追尾风险,但在面对紧急情况时缺乏策略重构与动态响应能力;Zhu等人通过轨迹拟合奖励实现了跟驰动作预测,尽管其显著提升了跟驰性能,但在长时间预测中存在误差累积;Bennajeh等人设计了双层模糊决策系统,将“安全优先”内嵌至速度控制模块中,尽管提升了模型的安全性与可解释性,但其面对雨雾、夜间等开放场景时,表现出适应性不足,为提升模型的泛化能力,Yao等人引入了混合专家(Mixture of Experts,MoE)结构,融合智能驾驶员模型(Intelligent Driver Model,IDM)专家与深度强化学习(DeepReinforcement Learning,DRL)策略专家,在高速跟驰任务中提升了 13.75% 的效率;但其专家调度机制仍基于启发式条件,缺乏对感知质量或风险信号的动态响应能力值得注意的是,极端天气场景下的制动控制对跟驰安全构成重大威胁,为此,Liang等人提出动态积累式刹车模型,引入对逼近信息的动态证据累积机制,动态积累式刹车模型展现出类人感知-反应机制在风险控制中的优势,也为构建具备类人情境感知与调节能力的决策结构提供了理论启示
[0055]The beneficial effects of the present invention are as follows: In autonomous driving, especially in complex tasks such as car-following decision-making in open worlds, the present invention addresses the core problems faced by existing technologies by proposing a biomimetic memory-driven hybrid expert car-following decision-making method and system. The specific problems are as follows: (1) Rigid strategy switching mechanism and delayed response: Existing multi-strategy or hybrid expert decision-making methods rely on fixed and heuristic rules (such as simple speed or distance thresholds) for expert selection and switching. This mechanism lacks real-time and refined perception of dynamic environmental risks, resulting in untimely or inappropriate strategy switching in sudden dangerous scenarios, or even behavioral instability; (2) Poor adaptability and insufficient robustness in extreme scenarios: In cases of perception degradation (such as rain, fog, night) In extreme scenarios such as intermittent or obstructed conditions or high uncertainty, existing models lack a robust policy adjustment mechanism. They are unable to adjust decision preferences based on the dynamic changes in perceived quality and risk signals, and cannot achieve adaptive trade-offs among multiple objectives such as safety, efficiency, and comfort. (3) Human-based models are simple and lack human-like risk adjustment capabilities: Existing models of human driving behavior mostly focus on imitating static driving habits, while ignoring the emotional regulation and memory-driven decision-making mechanisms that human drivers possess when facing dynamic risks. For example, humans become tense and conservative when they perceive danger, but more relaxed and efficient on open roads. Existing autonomous driving systems lack this kind of human-like, risk-emotion-based dynamic adjustment capability.
Smart Images

Figure CN121316900B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a hybrid expert car-following decision-making method and system, and more particularly to a biomimetic memory-driven hybrid expert car-following decision-making method and system, belonging to the field of car-following decision-making technology. Background Technology
[0002] Car-following decision-making is the core module of longitudinal control. Its fundamental task is to achieve a dynamic balance between safety, efficiency, and comfort. The development of car-following decision-making technology has evolved from traditional models to modern learning methods. Early car-following models, such as safe distance models and intelligent driver models, were mainly based on explicit physical rules and mathematical formulas. Although they had good interpretability, they were too rigid when dealing with complex and ever-changing real traffic environments. Currently, Wen et al. have used the IQ-Learn algorithm combined with SAC... Simulations of human driving following behavior were achieved, revealing significant strategy adjustments when following autonomous vehicles, but these strategy differences were not further translated into response mechanisms for the autonomous driving system. Zhang et al. proposed a two-stage following strategy, embedding safety as a precondition into the energy consumption optimization process, which effectively reduced the risk of rear-end collisions, but lacked strategy reconstruction and dynamic response capabilities in emergency situations. Zhu et al. achieved following action prediction through trajectory fitting rewards, which significantly improved following performance, but error accumulation occurred in long-term predictions. Bennageh et al. designed a two-layer fuzzy decision system, embedding "safety first" into the speed control module, which improved the model's safety and interpretability, but showed insufficient adaptability in open scenarios such as rain, fog, and nighttime. To improve the model's generalization ability, Yao et al. introduced a Mixture of Experts (MoE) structure, integrating Intelligent Driver Model (IDM) experts and Deep Reinforcement Learning (DRL) policy experts, achieving a 13.75% improvement in high-speed following tasks. While it offers high efficiency, its expert scheduling mechanism is still based on heuristic conditions and lacks the ability to dynamically respond to perceived quality or risk signals. It is worth noting that braking control in extreme weather scenarios poses a significant threat to car-following safety. To address this, Liang et al. proposed a dynamic accumulative braking model, which introduces a dynamic evidence accumulation mechanism for approximation information. The dynamic accumulative braking model demonstrates the advantages of human-like perception-response mechanisms in risk control and provides theoretical inspiration for constructing decision-making structures with human-like situational perception and adjustment capabilities.
[0003] While existing car-following decision-making technologies have made some progress, they still have key shortcomings when dealing with complex open-world scenarios. First, most multi-strategy or hybrid expert-based methods use static rules for strategy switching, lacking dynamic risk-driven mechanisms, leading to lag in response or behavioral instability in highly dynamic environments. Second, under conditions of perception degradation such as rain or fog, or extreme weather, existing models lack robust strategy adjustment mechanisms, making it difficult to maintain safe and stable car-following performance in time-varying environments. Finally, human factors modeling mostly focuses on static driving style classification, lacking simulation of human drivers' ability to adjust their decisions based on dynamic risks and emotional memories, resulting in autonomous driving behavior that is not human-like enough and makes it difficult to make optimal decisions.
[0004] In summary, there is a need for a hybrid expert car-following decision-making method and system that can simulate human emotional feedback, dynamically perceive environmental risks, and adaptively switch to the optimal strategy. Summary of the Invention
[0005] A brief overview of the invention is given below to provide a basic understanding of certain aspects of it. It should be understood that this overview is not an exhaustive summary of the invention. It is not intended to identify key or essential parts of the invention, nor is it intended to limit the scope of the invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description that follows.
[0006] In view of this, in order to solve the problems of lag in strategy switching response or behavioral instability and lack of personalized adjustment mechanism for dynamic risk perception and emotional memory in the traditional car-following decision-making methods and systems in the existing technology in real complex scenarios, the present invention provides a biomimetic memory-driven hybrid expert car-following decision-making method and system.
[0007] Technical solution one is as follows: A biomimetic memory-driven hybrid expert car-following decision-making method, comprising the following steps:
[0008] S1. Construct a hybrid expert library by setting the three DRL agents as different types of experts;
[0009] S2. Based on the hybrid expert database, establish a gating mechanism driven by a biomimetic hippocampus.
[0010] Furthermore, step S1 includes the following steps:
[0011] S11. Set the DDPG strategy approach to a security expert;
[0012] In S11, the training process of the DDPG policy method consists of an Actor-Critic dual network.
[0013] Loss function of Critic network Represented as:
[0014]
[0015]
[0016] in, This indicates the number of training samples in a mini-batch. , They represent the first The state and actions of each sample For the first The next state of a sample. Indicates an immediate reward. As a discount factor, , For state Deterministic policy network, For the goal value, For action value functions;
[0017] The update process of the Actor network is represented as follows:
[0018]
[0019] in, , These represent the policy network parameters of the Critic network and the Actor network, respectively. Representing state Deterministic policy network, Represents the function Regarding the action gradient, Indicates the network parameters of the policy Regarding its output action, that is The gradient;
[0020] The equivalent objective of the DDPG strategy method is to maximize Value, i.e., minimizing the negative The value is used to obtain the loss function of the DDPG strategy method. ;
[0021] Loss function of DDPG strategy method Represented as:
[0022]
[0023] in, This represents the state sampled from the experience replay pool D. The calculated mathematical expectation, Represents the action value function. Indicates the first The state of each sample For state Deterministic policy network;
[0024] S12. Set the SAC strategy approach as an adaptive expert;
[0025] In S12, the objective function of the SAC strategy method is... Represented as:
[0026]
[0027] in, Let represent the expected value of the state-action pair (s, a) sampled from the experience replay pool D. This represents the experience replay pool. Indicates the network parameters of the policy Given the action distribution, This represents the entropy regularization weight coefficient. , Indicates action ,state The entropy value under;
[0028] Entropy regularization weight coefficient Adaptive adjustments are made to obtain the updated entropy regularization weight coefficients. ;
[0029] Updated entropy regularization weight coefficients Represented as:
[0030]
[0031] in, The temperature learning rate represents the entropy. Represents the target entropy. Describe the objective function Regarding the entropy regularization weight coefficient The partial derivatives, Indicates action ,state The entropy value under;
[0032] The goal of the SAC strategy method is to maximize... The sum of the value and the policy entropy is converted into the Actor network loss Actor-Loss in a minimization form, yielding the loss function of the SAC policy method. ;
[0033] Loss function of SAC strategy method Represented as:
[0034]
[0035] in, Indicates action ,state The action value function below;
[0036] S13. Set the PPO strategy approach as an efficiency expert;
[0037] In S13, the loss function of the PPO strategy method Represented as:
[0038]
[0039] Where min represents the minimum value. This represents the probability ratio between the current policy and the old policy. , Indicates the current policy in state Take action below The probability, Indicates the old strategy in state Take action below The probability, Representing state-action pairs The advantage estimate, clip (·) indicates that the policy ratio is clipped within the interval. Functions between This represents the pruning threshold for policy updates. This represents the empirical average calculated from samples at all time steps t in a batch.
[0040] Furthermore, step S2 includes the following steps:
[0041] S21. Perform perceptual category standardization, that is, use a mapping function to standardize the open-world original categories output by the target recognition model into five perceptual semantic categories;
[0042] S22. Combining the five types of perceptual semantics, set a risk estimation distance threshold. The qualitative risk value was calculated. Construct a qualitative risk value classification table and introduce risk-adjusted weight multipliers. This yields the risk weight correction rule, and further, the comprehensive risk level of each objective.
[0043] In step S22, a risk estimation distance threshold is set by referring to the minimum safe distance modeling method and combining the dynamic characteristics of obstacle types and target detection response requirements. Classification of traffic participants Interaction distance By combining these factors, a qualitative risk value classification table is obtained. A risk adjustment weight multiplier is then introduced into this table. Multiplicative adjustment weights are assigned to near-range targets, unknown category targets, and marker targets respectively, resulting in a risk weight correction rule, which further yields the comprehensive risk level of each target. ;
[0044] S23. In terms of overall risk level Introducing risk suppression factors The final emotion score is calculated. To achieve multi-objective risk aggregation and suppression;
[0045] In step S23, a default threshold for the number of risk-concern targets is set. In risk inhibition factors Introducing incremental suppression coefficient Using risk inhibition factors Controlling the nonlinear growth rate of risk values under multi-objective conditions;
[0046] Risk suppression factor Represented as:
[0047]
[0048] in, This indicates the number of targets detected in the current frame;
[0049] Mood Score Represented as:
[0050]
[0051] S24. Based on sentiment score By combining a hybrid expert database, a gating mechanism with a segmented expert activation mechanism is set up to realize the dynamic switching of control strategies between conservative control (DDPG), neutral adaptation (SAC), and aggressive response (PPO) strategies, and establish a gating mechanism to effectively adapt to the strategy requirements under different risk levels.
[0052] In S24, the gate controller Represented as:
[0053] .
[0054] Technical Solution 2: A biomimetic memory-driven hybrid expert car-following decision system, used to execute the biomimetic memory-driven hybrid expert car-following decision method described in Technical Solution 1, including an environmental perception data input module, an emotion score calculation module, and an expert model selection module connected in sequence.
[0055] The beneficial effects of the present invention are as follows: In autonomous driving, especially in complex tasks such as car-following decision-making in open worlds, the present invention addresses the core problems faced by existing technologies by proposing a biomimetic memory-driven hybrid expert car-following decision-making method and system. The specific problems are as follows: (1) Rigid strategy switching mechanism and delayed response: Existing multi-strategy or hybrid expert decision-making methods rely on fixed and heuristic rules (such as simple speed or distance thresholds) for expert selection and switching. This mechanism lacks real-time and refined perception of dynamic environmental risks, resulting in untimely or inappropriate strategy switching in sudden dangerous scenarios, or even behavioral instability; (2) Poor adaptability and insufficient robustness in extreme scenarios: In cases of perception degradation (such as rain, fog, night) In extreme scenarios such as intermittent or obstructed conditions or high uncertainty, existing models lack a robust policy adjustment mechanism. They are unable to adjust decision preferences based on the dynamic changes in perceived quality and risk signals, and cannot achieve adaptive trade-offs among multiple objectives such as safety, efficiency, and comfort. (3) Human-based models are simple and lack human-like risk adjustment capabilities: Existing models of human driving behavior mostly focus on imitating static driving habits, while ignoring the emotional regulation and memory-driven decision-making mechanisms that human drivers possess when facing dynamic risks. For example, humans become tense and conservative when they perceive danger, but more relaxed and efficient on open roads. Existing autonomous driving systems lack this kind of human-like, risk-emotion-based dynamic adjustment capability.
[0056] Unlike traditional models that rely on static rules, this invention's biomimetic gating mechanism can assess the risks of the current driving scenario in real time, ensuring smoother and more timely policy switching and avoiding behavioral instability. In high-risk or perceptually uncertain scenarios, this invention effectively improves system safety through a safety expert driven by emotion scores and DDPG, SAC, and PPO algorithms. Furthermore, based on a multi-expert collaborative optimization architecture, the system can dynamically adjust its strategy under different risk conditions, balancing safety, efficiency, and comfort. Drawing inspiration from the memory and risk assessment mechanisms of the human hippocampus, this invention's decision-making method is closer to human intuition, enhancing… This invention enhances the interpretability and naturalness of human-computer interaction in autonomous driving systems. In summary, by constructing a biomimetic hippocampus-driven gating mechanism, this invention quantifies multi-dimensional perceptual information into a comprehensive "emotional risk score" in real time. Based on this score, it dynamically and intelligently schedules heterogeneous expert strategies focused on safety, adaptation, and efficiency. This invention overcomes the delay and instability of static rule switching, improves decision robustness in uncertain perceptual environments, and achieves more human-like and interpretable dynamic driving behavior adjustment. Thus, it enables multi-objective adaptive optimization of safety, efficiency, and comfort in complex and ever-changing traffic environments. Attached Figure Description
[0057] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:
[0058] Figure 1 A flowchart illustrating a biomimetic memory-driven hybrid expert car-following decision-making method;
[0059] Figure 2 A schematic diagram of the structure of a biomimetic memory-driven hybrid expert car-following decision-making system;
[0060] Figure 3 This is a schematic diagram of an embodiment of a biomimetic memory-driven hybrid expert car-following decision-making method.
[0061] Figure descriptions: 1. Environmental perception data input module; 2. Emotion score calculation module; 3. Expert model selection module. Detailed Implementation
[0062] To make the technical solutions and advantages of the embodiments of the present invention clearer, the exemplary embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0063] Example 1: Reference Figures 1-3 This embodiment details a biomimetic memory-driven hybrid expert car-following decision-making method, specifically including the following steps:
[0064] S1. Construct a hybrid expert library by setting the three DRL agents as different types of experts;
[0065] S2. Based on the hybrid expert database, establish a gating mechanism driven by a biomimetic hippocampus.
[0066] Specifically, the hybrid expert library of the present invention consists of three independent DRL agents that share the same environmental state input, but have significant differences in algorithm structure, training objectives and behavioral characteristics, thus forming a complementary function.
[0067] refer to Figure 3This invention aims to achieve dynamic and intelligent scheduling of different driving strategies by simulating the emotional memory and risk assessment mechanism of the human hippocampus. Its core is a car-following decision system composed of a bionic gating module and a hybrid expert database. The system acts as the decision-making center, and its upstream interface is with the vehicle's perception module. It receives target information (such as category, distance, confidence level, etc.) processed by the perception module and outputs decision commands. After receiving perception information, it makes intelligent decisions and outputs control commands through the internal bionic gating module and hybrid expert database.
[0068] Mixture of Experts (MoE): An ensemble learning neural network architecture that includes a gating network and multiple expert networks. The gating network is responsible for determining which expert network(s) to assign the task to based on the characteristics of the input data, thereby achieving decomposition and efficient processing of complex problems.
[0069] Emotion Score: In this invention, it is used as a quantitative indicator to simulate the degree of tension or risk perception of humans in specific situations. It is calculated based on a comprehensive analysis of environmental perception information (such as obstacle type, distance, and threat level) and is directly used to drive decision-making.
[0070] Gating Mechanism: In this invention, it is a control module that dynamically selects which expert strategy to activate based on "emotional score". It acts like a switch, switching to the expert most suitable for the current scenario according to the level of risk.
[0071] DDPG (Deep Deterministic Policy Gradient): A deep reinforcement learning algorithm for continuous action spaces. Its policy is deterministic, and it has a fast response speed, making it suitable for safety and risk avoidance scenarios that require precise control and rapid reaction.
[0072] SAC (Soft Actor-Critic): A deep reinforcement learning algorithm based on the principle of maximum entropy. While pursuing the maximization of rewards, it also encourages the exploration of policies, so that it can achieve a good balance between safety and efficiency and has strong adaptability.
[0073] PPO (Proximal Policy Optimization): A deep reinforcement learning algorithm based on policy gradients. It ensures training stability by limiting the magnitude of policy updates, and typically exhibits good convergence and high policy execution efficiency.
[0074] Furthermore, step S1 includes the following steps:
[0075] S11. Set the DDPG strategy approach to a security expert;
[0076] In S11, DDPG (Deep Deterministic Policy Gradient) is a deterministic policy method applicable to continuous action spaces. The training process of the DDPG policy method consists of an Actor-Critic dual network.
[0077] Loss function of Critic network Represented as:
[0078]
[0079]
[0080] in, This indicates the number of training samples in a mini-batch. , They represent the first The state and actions of each sample For the first The next state of a sample. Indicates an immediate reward. This is a discount factor used to control the degree of decay in future rewards. , For state Deterministic policy network, For the goal value, For action value functions;
[0081] The update process of the Actor network is represented as follows:
[0082]
[0083] in, , These represent the policy network parameters of the Critic network and the Actor network, respectively. Representing state Deterministic policy network, Represents the function Regarding the action gradient, Indicates the network parameters of the policy Regarding its output action, that is The gradient;
[0084] The equivalent objective of the DDPG strategy method is to maximize Value, i.e., minimizing the negative The value is used to obtain the loss function of the DDPG strategy method. ;
[0085] Loss function of DDPG strategy method Represented as:
[0086]
[0087] in, This represents the state sampled from the experience replay pool D. The calculated mathematical expectation, Represents the action value function. Indicates the first The state of each sample For state Deterministic policy network;
[0088] S12. Set the SAC strategy approach as an adaptive expert;
[0089] In S12, SAC (Soft Actor-Critic) is a policy-determining algorithm based on the maximum entropy principle. It maximizes the entropy of the policy distribution while optimizing the reward to enhance exploratory behavior. The objective function of the SAC policy method is... Represented as:
[0090]
[0091] in, Let represent the expected value of the state-action pair (s, a) sampled from the experience replay pool D. This represents the experience playback pool (experience sampling distribution). Indicates the network parameters of the policy Given the action distribution, This represents the entropy regularization weighting coefficient, which is used to control the exploration intensity. , Indicates action ,state The entropy value under;
[0092] Entropy regularization weight coefficient Adaptive adjustments are made to obtain the updated entropy regularization weight coefficients. ;
[0093] Updated entropy regularization weight coefficients Represented as:
[0094]
[0095] in, The temperature learning rate represents the entropy. This represents the target entropy, which is typically a negative action space dimension. Describe the objective function Regarding the entropy regularization weight coefficient The partial derivative of , which is used to guide . The adaptive adjustment allows the entropy of the policy output to gradually approach the target entropy. This allows for dynamic adjustment of strategies in an exploratory manner. Indicates action ,state The entropy value under;
[0096] The goal of the SAC strategy method is to maximize... The sum of the value and the policy entropy is converted into the Actor network loss Actor-Loss in a minimization form, yielding the loss function of the SAC policy method. ;
[0097] Loss function of SAC strategy method Represented as:
[0098]
[0099] in, Indicates action ,state The action value function below;
[0100] S13. Set the PPO strategy approach as an efficiency expert;
[0101] In S13, Proximal Policy Optimization (PPO) is one of the most widely used policy gradient methods. Its core feature is the introduction of a policy update constraint mechanism to control the magnitude of each policy iteration, thereby improving the stability and robustness of the training process. The loss function of the PPO policy method... Represented as:
[0102]
[0103] Where min represents the minimum value. This represents the probability ratio between the current policy and the old policy. , Indicates the current policy in state Take action below The probability, Indicates the old strategy in state Take action below The probability, Representing state-action pairs The advantage estimate, clip (·) indicates that the policy ratio is clipped within the interval. Functions between This represents the pruning threshold for policy updates, typically set to 0.2. This represents the empirical average calculated from samples at all time steps t in a batch.
[0104] Specifically, since the DDPG strategy method lacks policy entropy adjustment, the training process is more sensitive to environmental feedback, the policy update speed is fast, and it has strong responsiveness. Combined with its good adaptability to continuous control, this invention sets it as a safety expert, which is suitable for scenarios with many vehicles and people, complex scenes, and high-risk potential conflicts, such as collision avoidance at urban intersections, etc., and strengthens its obstacle avoidance ability and safety guarantee role under risk-driven decision-making.
[0105] Given that the SAC strategy method can flexibly adjust the control strategy under different risk levels, it can achieve a dynamic balance between safety and efficiency in medium-risk scenarios such as medium-speed driving, congested lane changing, and urban interaction. Therefore, this invention uses it as an adaptive expert, which is activated when a dynamic trade-off between exploration and stability is required.
[0106] Since the PPO strategy method only allows small updates to the strategy parameters in each round, it has good convergence stability and efficient strategy optimization capabilities. Its output focuses more on improving traffic efficiency. Therefore, this invention uses it as an efficiency expert, which is suitable for low-risk scenarios with sparse traffic flow and clear road rules, in order to achieve fast passage and efficient driving strategies.
[0107] Furthermore, step S2 includes the following steps:
[0108] S21. To facilitate subsequent decision generation and optimization, perception category standardization is performed, that is, the open-world original categories output by the target recognition model are standardized into five perception semantic categories using a mapping function;
[0109] S22. Combining the five types of perceptual semantics, set a risk estimation distance threshold. The qualitative risk value was calculated. Construct a qualitative risk value classification table and introduce risk-adjusted weight multipliers. This yields the risk weight correction rule, and further, the comprehensive risk level of each objective.
[0110] In step S22, a risk estimation distance threshold is set by referring to the minimum safe distance modeling method and combining the dynamic characteristics of obstacle types and target detection response requirements. The distance is 10m, used for pedestrians, obstacles, and motor vehicles, to ensure timely response to potentially high-risk targets and consistency in risk assessment. Traffic sign targets, since they do not participate in dynamic interaction or pose a direct collision risk, are not included in risk estimation based on physical interaction distance. The categories of traffic participants are also considered. Interaction distance By combining these factors, a qualitative risk value classification table is obtained. To further characterize the risk intensity of the target, a risk adjustment weight multiplier is introduced into the qualitative risk value classification table. Multiplicative adjustment weights are assigned to near-range targets, unknown category targets, and marker targets respectively, resulting in a risk weight correction rule, which further yields the comprehensive risk level of each target. It is used to quantify the degree of threat it poses to the driving task in the current scenario;
[0111] S23. In terms of overall risk level Introducing risk suppression factors The final emotion score is calculated. To achieve multi-objective risk aggregation and suppression;
[0112] In step S23, under complex traffic conditions, the perception system may simultaneously detect a large number of interfering targets (such as distant vehicles and non-priority objects). Simply linearly superimposing the risks of all targets may disproportionately amplify the overall risk, thereby misleading strategy selection. Therefore, based on the concept of "Top-K target focusing," a default threshold for the number of risk-focused targets is set. , In risk inhibition factors Introducing incremental suppression coefficient , Using risk inhibition factors Controlling the nonlinear growth rate of risk values under multi-objective conditions;
[0113] Risk suppression factor Represented as:
[0114]
[0115] in, This indicates the number of targets detected in the current frame;
[0116] Mood Score Represented as:
[0117]
[0118] Mood Score The comprehensive risk level under the current perception scenario was comprehensively assessed, which can be regarded as a human-like emotional state indicator. It drives the strategic preference changes of the decision-making system through the hippocampus in a brain-like manner.
[0119] S24. Based on sentiment score By combining a hybrid expert database, a gating mechanism with a segmented expert activation mechanism is set up to realize the dynamic switching of control strategies between conservative control (DDPG), neutral adaptation (SAC), and aggressive response (PPO) strategies, and establish a gating mechanism to effectively adapt to the strategy requirements under different risk levels.
[0120] In S24, the gate controller Represented as:
[0121] .
[0122] Specifically, referring to Table 1, the standardized mapping table of perception categories is obtained through step S21;
[0123] Table 1
[0124]
[0125] Refer to Table 2 for the rules for calculating distance perception risk values;
[0126] Table 2
[0127]
[0128] Referring to Table 3, the risk weight adjustment rules have multiplicative adjustment weights of 1.5, 1.2, and 1.1, respectively.
[0129] Table 3
[0130]
[0131] The gating mechanism in step S24 enables the autonomous driving system to operate under high-risk conditions ( If the value is greater than or equal to 0.7, switch to the conservative "safe" mode, in low-risk ( Switch to aggressive "efficiency" mode when the value is less than 0.4, and switch to medium risk ( When the value is less than 0.7 and greater than 0.4, a flexible "adaptive" mode is adopted, thereby realizing intelligent, dynamic and robust decision response to different driving situations.
[0132] Example 2: Reference Figure 2This embodiment describes a biomimetic memory-driven hybrid expert car-following decision system, used in the biomimetic memory-driven hybrid expert car-following decision method described in Embodiment 1. The system includes an environmental perception data input module 1, an emotion score calculation module 2, and an expert model selection module 3, which are connected in sequence.
[0133] Specifically, the environmental perception data input module 1 inputs the target type, target distance, and confidence target of the perception category into the emotion score calculation module 2. In this embodiment, the environmental perception data input module inputs the scene state into the hybrid expert database, which includes vehicle speed, weather conditions (such as friction coefficient in rainy weather), perception confidence (such as perception target accuracy), etc.
[0134] The emotion score calculation module 2 is used to execute steps S22-S23 to obtain the emotion score. Through the gating device of module 3 Select an expert model and output strategies corresponding to different risk levels.
[0135] Although the invention has been described with reference to a limited number of embodiments, those skilled in the art will understand from the foregoing description that other embodiments are conceivable within the scope of the invention described herein. Furthermore, it should be noted that the language used in this specification has been chosen primarily for readability and instructional purposes, and not for the purpose of interpreting or limiting the subject matter of the invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the invention is illustrative and not restrictive, and the scope of the invention is defined by the appended claims.
Claims
1. A hybrid expert following decision method driven by biomimetic memory, characterized in that, Includes the following steps: S1. Construct a hybrid expert library by setting the three DRL agents as different types of experts; S2. Based on the hybrid expert database, establish a bionic hippocampus-driven gating mechanism, and combine it with the calculated emotion score to output strategies corresponding to different risk levels; S2 includes the following steps: S21. Perform perceptual category standardization, that is, use a mapping function to standardize the open-world original categories output by the target recognition model into five perceptual semantic categories; S22. Combining the five types of perceptual semantics, set a risk estimation distance threshold. The qualitative risk value was calculated. Construct a qualitative risk value classification table and introduce risk-adjusted weight multipliers. This yields the risk weight correction rule, and further, the comprehensive risk level of each objective. In step S22, a risk estimation distance threshold is set by referring to the minimum safe distance modeling method and combining the dynamic characteristics of obstacle types and target detection response requirements. Classification of traffic participants Interaction distance By combining these factors, a qualitative risk value classification table is obtained. A risk adjustment weight multiplier is then introduced into this table. Multiplicative adjustment weights are assigned to near-range targets, unknown category targets, and marker targets respectively, resulting in a risk weight correction rule, which further yields the comprehensive risk level of each target. ; S23. In terms of overall risk level Introducing risk suppression factors The final emotion score is calculated. To achieve multi-objective risk aggregation and suppression; In step S23, a default threshold for the number of risk-concern targets is set. In risk inhibition factors Introducing the incremental suppression coefficient Using risk inhibition factors Controlling the nonlinear growth rate of risk values under multi-objective conditions; Risk suppression factor Represented as: in, This indicates the number of targets detected in the current frame; Mood Score Represented as: S24. Based on sentiment score By combining a hybrid expert database, a gating mechanism with a segmented expert activation mechanism is set up to realize the dynamic switching of control strategies between conservative control (DDPG), neutral adaptation (SAC), and aggressive response (PPO) strategies, and establish a gating mechanism to effectively adapt to the strategy requirements under different risk levels. In S24, the gate controller Represented as: 。 2. The biomimetic memory-driven hybrid expert car-following decision-making method according to claim 1, characterized in that, S1 includes the following steps: S11. Set the DDPG strategy approach to a security expert; In S11, the training process of the DDPG policy method consists of an Actor-Critic dual network. Loss function of Critic network Represented as: in, This indicates the number of training samples in a mini-batch. , They represent the first The state and actions of each sample For the first The next state of a sample. Indicates an immediate reward. As a discount factor, , For state Deterministic policy network, For the goal value, For action value functions; The update process of the Actor network is represented as follows: in, , These represent the policy network parameters of the Critic network and the Actor network, respectively. Representing state Deterministic policy network, Indicates state, Represents the function Regarding the status ,action gradient, Indicates the network parameters of the policy Regarding its output action, that is The gradient; The equivalent objective of the DDPG strategy method is to maximize Value, i.e., minimizing the negative The value is used to obtain the loss function of the DDPG strategy method. ; Loss function of DDPG strategy method Represented as: in, This represents the state sampled from the experience replay pool D. The calculated mathematical expectation, Represents the action value function. Indicates the first The state of each sample For state Deterministic policy network; S12. Set the SAC strategy approach as an adaptive expert; In S12, the objective function of the SAC strategy method is... Represented as: in, Let represent the expected value of the state-action pair (s, a) sampled from the experience replay pool D. This represents the experience replay pool. Indicates the network parameters of the policy Given the action distribution, This represents the entropy regularization weight coefficient. , Indicates action ,state The entropy value under; Entropy regularization weight coefficient Adaptive adjustments are made to obtain the updated entropy regularization weight coefficients. ; Updated entropy regularization weight coefficients Represented as: in, The temperature represents the entropy learning rate. Represents the target entropy. Describe the objective function Regarding the entropy regularization weight coefficient The partial derivatives, Indicates action ,state The entropy value under; The goal of the SAC strategy method is to maximize... The sum of the value and the policy entropy is converted into the Actor network loss Actor-Loss in a minimization form, yielding the loss function of the SAC policy method. ; Loss function of SAC strategy method Represented as: in, Indicates action ,state The action value function below; S13. Set the PPO strategy approach as an efficiency expert; In S13, the loss function of the PPO strategy method Represented as: Where min represents the minimum value. This represents the probability ratio between the current policy and the old policy. , Indicates the current policy in state Take action below The probability, Indicates the old strategy in state Take action below The probability, Representing state-action pairs The advantage estimate, clip (·) indicates that the policy ratio is clipped within the interval. Functions between This represents the pruning threshold for policy updates. This represents the empirical average calculated from samples at all time steps t in a batch.
3. A biomimetic memory-driven hybrid expert car-following decision-making system, characterized in that, The method for performing a biomimetic memory-driven hybrid expert car-following decision-making method according to any one of claims 1-2 includes an environmental perception data input module (1), an emotion score calculation module (2), and an expert model selection module (3) connected in sequence.
Citation Information
Patent Citations
Anthropomorphic automatic driving car-following model based on deep reinforcement learning
CN109733415A
Intelligent automobile human-like car-following behavior control method based on improved deep reinforcement learning
CN115830863A