World model based hybrid expert policy fusion autonomous driving control method
Patent Information
- Application Number
- CN202511671373.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2045-11-14
AI Technical Summary
[0006]本发明的目的在于提出一种基于世界模型推演的混合专家策略融合的自动驾驶控制方法,以解决现有技术在复杂道路环境下存在的策略泛化能力不足、行为控制连贯性差、安全保障机制缺失等问题,实现具备轨迹级风险评估能力、策略动态协同融合能力以及规则与学习策略协作能力的端到端闭环控制系统
[0051]First, this invention proposes a multi-expert strategy structure that integrates long-term, short-term, and rule-based safety strategies, and designs a fusion mechanism based on trajectory-level risk and performance indicators. By introducing a Bayesian soft-selection strategy fusion formula, the system can adaptively allocate the weights of each strategy according to the dynamic performance of the predicted trajectory, achieving a strategy generation method that emphasizes both performance-driven and safety-assured approaches. Especially in dynamic and complex traffic scenarios, the fusion strategy mechanism of this invention improves behavioral continuity and goal achievement rate while also ensuring the smoothness of strategy switching and the robustness of system response.
Smart Images

Figure CN121553176B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation technology, specifically to an autonomous driving control method, system, and storage medium based on a world model and hybrid expert strategy fusion. Background Technology
[0002] With the rapid development of artificial intelligence and automotive technology, autonomous driving systems are evolving from single-sensor control to complex systems that integrate modeling, prediction, and decision-making. The system design goals are gradually shifting from behavioral-responsive control to an intelligent control architecture with environmental reasoning, strategy adaptation, and risk avoidance capabilities. By introducing world modeling mechanisms, multi-strategy fusion mechanisms, and strategy-level risk control measures, the task completion rate and deployment stability of autonomous driving systems in complex environments can be further improved.
[0003] Currently, autonomous driving systems face challenges in real-world road environments, including dense dynamic targets, scattered semantic cues, and high uncertainty in trajectory generation. This places higher demands on the predictive capabilities and response stability of control strategies. Especially under conditions of multi-target coordination, multi-stage task switching, and incomplete observable information, the system not only needs to accurately predict environmental dynamics but also possesses a comprehensive judgment capability regarding trajectory safety and control robustness. Traditional control structures often struggle to support a closed-loop information expression from perception input to action output, and their strategy design lacks flexible adaptability and structural self-evolution capabilities, making it difficult to meet the deployment needs of diverse and complex traffic scenarios.
[0004] Existing technologies primarily employ modular control structures or end-to-end policy learning methods to model autonomous driving behavior. Modular solutions functionally divide perception, decision-making, and control, and execute path generation and action commands through rule systems or planners, offering good engineering interpretability and debugging friendliness, but lacking generalization ability for unknown scenarios. End-to-end learning methods, on the other hand, rely on imitation learning or reinforcement learning, directly learning state-action mapping relationships from raw observation data. They possess a certain degree of adaptability and policy self-learning ability, but are prone to policy degradation, low sample efficiency, and uncontrollable safety issues during long-term task execution. Furthermore, existing end-to-end structures mostly focus on single policy outputs, lacking the ability to model and fuse diverse aspects at the trajectory level, failing to achieve redundant management and risk perception control of policy behavior, and failing to effectively integrate trajectory simulation, risk assessment, and behavior generation into a unified reasoning process.
[0005] Invention CN120356177A discloses an end-to-end autonomous driving trajectory evaluation method based on a world model. It constructs a latent state space model to extrapolate and score multimodal perception inputs, thereby assisting in control path decision-making. Invention CN118928464A proposes a trajectory prediction and control method based on a hybrid expert model. It integrates multiple sub-strategy models to output candidate behavior decisions, improving the diversity and anti-interference capability of control strategies. Invention CN116118767A relates to a safety collaborative control system that operates in parallel with rule-based control and learning strategies. It provides redundant protection paths for learning strategies through a rule controller. However, these methods generally lack the ability to integrate world model extrapolation capabilities, multi-strategy trajectory fusion mechanisms, and rule-based redundant strategy collaboration into a unified framework. They cannot dynamically fuse and adjust based on trajectory-level risks and performance, nor do they establish an end-to-end closed-loop strategy generation process, making it difficult to meet the high requirements of balancing trajectory stability, control coherence, and deployment security. Summary of the Invention
[0006] The purpose of this invention is to propose an autonomous driving control method based on world model deduction and hybrid expert policy fusion, in order to solve the problems of insufficient policy generalization ability, poor behavior control coherence, and lack of safety assurance mechanism in the existing technology in complex road environments, and to realize an end-to-end closed-loop control system with trajectory-level risk assessment capability, dynamic policy collaborative fusion capability, and rule and learning policy collaboration capability.
[0007] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows:
[0008] In a first aspect, the present invention provides an autonomous driving control method based on world model deduction and hybrid expert policy fusion, comprising the following steps:
[0009] Step 1: Acquire multimodal sensor data of the vehicle in the current autonomous driving scenario and generate environmental perception information;
[0010] Step 2: Generate the vehicle's potential state representation by combining the obtained environmental perception information with the vehicle's historical trajectory state through the world model; wherein: the world model is constructed based on a recursive state space model that integrates physical priors;
[0011] Step 3: Based on the obtained potential state representation of the vehicle, generate multiple sets of candidate policy trajectories in the world model using a multi-policy trajectory generation module; the multi-policy trajectory generation module includes long-term policies. Short-term strategy and security strategies Among them: long-term strategy This is used to output long-term task trajectories with task constraints, and it is constructed based on a soft actor-critic Lagrange constraint architecture; short-term strategy Used to generate short-term elite response trajectories, it filters multiple candidate policy trajectories generated by the world model for the current control cycle to select elite trajectories that meet cost constraints, and performs corresponding action extraction and Gaussian distribution fitting to form an instantaneous response policy trajectory with stronger adaptability to the current environment; security strategy It is used to generate deployment safety trajectories. It adopts a rule control method based on intelligent driving model and provides stable yielding actions to ensure deployment safety when other strategies have high-risk prediction results or unstable behavior.
[0012] Step 4: Perform parallel simulation and deduction of the trajectories of each candidate strategy in the world model and calculate the state evolution process, cumulative reward value, cumulative cost value and potential state stability index of each candidate strategy trajectory in the potential space of the world model. At the same time, conduct a unified evaluation of the trajectory quality of each candidate strategy trajectory according to the preset risk criteria.
[0013] Step 5: Construct a strategy fusion module to use the cumulative reward value, cumulative cost value and potential state stability index of each candidate strategy trajectory in the potential space as the fusion basis, dynamically fuse the action output of each candidate strategy trajectory to generate the final action, which is directly applied to the vehicle control interface to achieve end-to-end control execution.
[0014] As a further improvement to the above technical solution, the present invention also includes step 6: collecting environmental perception information, execution actions and trajectory feedback results in the autonomous driving scenario to form a real-world interaction dataset; combining the predictive control variables generated by the world model, and jointly optimizing the model parameters of the two modules that generate the potential state representation of the vehicle and the candidate policy trajectory of the vehicle in the world model through an alternating training mechanism.
[0015] As a further improvement to the above technical solution, the world model described in this invention introduces a priori physical structure and combines it with a simplified vehicle dynamics model to impose structural constraints on acceleration, velocity changes, and trajectory curvature changes during the state transition process. Its physical consistency regularization term is expressed as:
[0016]
[0017] In the formula: , These are the acceleration and steering angle estimated based on prior physical structures, respectively. , The predictive control output of the world model. , These are weighting coefficients, used to penalize predictions and physical model biases, respectively.
[0018] As a further improvement to the above technical solution, the world model described in this invention is trained using a multi-task joint loss function, which is expressed as follows:
[0019]
[0020] In the formula: Indicates the observation reconstruction loss; Indicates the predicted loss from the reward; Indicates the predicted cost loss; This represents the loss of state consistency. Represents a physical consistency regularization term; , , , These represent the loss weights of reward prediction loss, cost prediction loss, state consistency loss, and physical consistency regularization term, respectively.
[0021] As a further improvement to the above technical solution, the long-term strategy described in this invention By introducing the Lagrange constraint mechanism, an optimization objective function with cost constraints is constructed:
[0022]
[0023] in: For long-term strategy In a specific state-action pair The expected value below; For long-term strategy In a specific state-action pair The expected cumulative return; It is an entropy adjustment factor; Indicating long-term strategy The probability of generating action a given state s; For Lagrange multipliers, To preset the cost threshold, For long-term strategy Cost estimation under specific state-action pairs;
[0024] Long-term strategy The policy network parameters are updated using the following objective function:
[0025]
[0026] in: Let KL divergence be the KL divergence. For the normalization term, the Q network and V network use the soft Bellman equation to regress the target value;
[0027] Long-term strategy During training, a hybrid experience pool is constructed by integrating real-world interactive datasets with simulated data generated by world model inference; and in each training cycle, target values are augmented by weighted sampling of samples from different sources and combined with rollout trajectories.
[0028] Long-term strategy The output long-term task trajectory is provided as a reference path to the policy fusion module and is given priority for action generation when the expected cumulative return and cost estimate meet the requirements.
[0029] As a further improvement to the above technical solution, the short-term strategy described in this invention... It has the following characteristics:
[0030] Within each control cycle, the world model is based on the current potential state representation. Parallel simulations were performed on multiple candidate policy trajectories, and the cumulative reward of each candidate policy trajectory was evaluated. With accumulated costs ;
[0031] Short-term strategy Using preset screening criteria, a set of low-risk trajectories is selected from all candidate strategy trajectories. Then, elite trajectories are selected from these to serve as the candidate strategy trajectories for the current control moment. The low-risk trajectory set represents the accumulated cost among the candidate strategy trajectories. The set consisting of all candidate policy trajectories The preset cost threshold is used; the elite trajectory is the candidate strategy trajectory with the highest reward or lowest cost in the set of low-risk trajectories.
[0032] The action sequences in the elite trajectory are used to fit a Gaussian policy distribution. , where: mean With covariance The following KL divergences are obtained by minimizing:
[0033]
[0034] In the formula, Represents an elite trajectory sequence; It is an empirical conditional distribution based on elite trajectory sequences.
[0035] As a further improvement to the above technical solution, the security strategy described in this invention... The IDM model is used to generate longitudinal control actions for the vehicle; the IDM model calculates the current safe acceleration using the following formula:
[0036]
[0037] In the formula: This is the vehicle's maximum acceleration; Current vehicle speed; The desired driving speed; This is the speed control index, used to adjust the degree of nonlinear decay of the speed term; To achieve the desired minimum safe distance; The speed difference with the vehicle in front; This represents the current actual distance between the vehicle and the vehicle in front.
[0038] As a further improvement to the above technical solution, the strategy fusion module of the present invention adopts a hybrid expert structure and uses a long-term strategy. Short-term strategy With security policy As candidate experts, based on the cumulative rewards corresponding to their respective trajectories. Cost estimation Potential state stability index Construct a fusion weight allocation mechanism to output the final action. :
[0039]
[0040] in ; Indicating long-term strategy Output acceleration; Indicating short-term strategy Output acceleration; Indicate security policy Output acceleration; ; Indicating long-term strategy The fusion weights of the output acceleration; Indicating short-term strategy The fusion weights of the output acceleration; Indicate security policy The fusion weights of the output acceleration.
[0041] As a further improvement to the above technical solution, the world model described in this invention introduces a potential state transition detection mechanism to monitor the step KL divergence of the potential state sequence composed of each continuous potential state representation generated in step 2, thereby marking the candidate strategy trajectory with potential transition risk.
[0042] The policy fusion module is also equipped with a security policy fallback mechanism, specifically:
[0043] The policy fusion module, based on the received long-term policy and short-term strategies The output candidate strategy trajectories are identified to determine whether both are candidate strategy trajectories with potential jump risks or whether the cumulative costs corresponding to both exceed the limit. If the identification result shows that both are candidate strategy trajectories with potential jump risks or the judgment result shows that the cumulative costs corresponding to both exceed the limit, the safety policy retreat mechanism is triggered.
[0044] When the security policy yield mechanism is triggered, the weighted fusion control actions are performed as follows: :
[0045]
[0046] in: , for fusion weights; For security strategy Output safety acceleration; Indicates a long-term strategy Short-term strategy As candidate experts, based on the cumulative rewards corresponding to their respective trajectories. Cost estimation Potential state stability index The fusion action output by constructing the fusion weight allocation mechanism is calculated using the following formula:
[0047]
[0048] In the formula: Indicating long-term strategy The fusion weights of the output acceleration; Indicating short-term strategy The fusion weights of the output acceleration; Indicating long-term strategy Output acceleration; Indicating short-term strategy Output acceleration.
[0049] In a second aspect, the present invention also provides a storage medium storing a computer program, which, when executed in a computer, is used to implement the autonomous driving control method based on world model deduction and hybrid expert policy fusion as described in the first aspect.
[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0051] First, this invention proposes a multi-expert strategy structure that integrates long-term, short-term, and rule-based safety strategies, and designs a fusion mechanism based on trajectory-level risk and performance indicators. By introducing a Bayesian soft-selection strategy fusion formula, the system can adaptively allocate the weights of each strategy according to the dynamic performance of the predicted trajectory, achieving a strategy generation method that emphasizes both performance-driven and safety-assured approaches. Especially in dynamic and complex traffic scenarios, the fusion strategy mechanism of this invention improves behavioral continuity and goal achievement rate while also ensuring the smoothness of strategy switching and the robustness of system response.
[0052] Second, this invention integrates a world model deduction mechanism based on the RSSM-PINN structure with multimodal sensing input, enabling high-quality latent state modeling and trajectory prediction while maintaining physical consistency. The constructed RSSM-PINN model encodes information such as images, radar, and vehicle states into unified latent variables, and through a physically guided state transition network and trajectory reconstruction module, achieves continuous modeling of the dynamic environment and multi-step behavior deduction, providing structured, high-fidelity state priors for subsequent policy generation. Compared to traditional model prediction methods based on direct learning of state-action pairs, this structure significantly improves generalization ability and stability under non-ideal observation conditions.
[0053] Third, this invention constructs a closed-loop policy evolution and optimization framework. Relying on the multi-step extrapolation capabilities provided by the world model, it performs fine-tuning of multiple policy candidate trajectories before execution and refits and updates the policy based on the elite trajectory set. By introducing a decoupled reward / cost supervision structure and a short-term policy training mechanism with dynamic resampling, the system can achieve continuous iterative optimization and adaptive evolution of policy behavior during deployment. Furthermore, in terms of safety mechanisms, the system possesses trajectory-level risk monitoring and emergency policy switching capabilities, ensuring rapid degradation to rule-based safety control under high-risk conditions, thus guaranteeing the safety and reliability of the autonomous driving system. Attached Figure Description
[0054] Figure 1 This is a diagram of the architecture of the autonomous driving control system of the present invention;
[0055] Figure 2 This is a schematic diagram of the training process data flow of the present invention. Detailed Implementation
[0056] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0057] This invention discloses an autonomous driving control method based on world model deduction and hybrid expert policy fusion, comprising the following steps:
[0058] Step 1: Acquire multimodal sensor data of the vehicle in the current autonomous driving scenario and generate environmental perception information.
[0059] In autonomous driving scenarios, vehicles are typically equipped with various sensors, such as cameras, LiDAR, millimeter-wave radar, and inertial navigation systems (INS), which can capture environmental information around the vehicle in real time. In this invention, multimodal sensor data is based on multiple types of data information collected by different sensors.
[0060] This invention employs a Visual Transformer (ViT) architecture to construct a perception coding module. Based on the obtained multimodal sensor data, it extracts high-dimensional perception features representing the current environment of the vehicle and outputs environmental perception information for subsequent potential state modeling and policy generation.
[0061] Step 2: Combine the obtained environmental perception information with the vehicle's historical trajectory state using a world model to generate the vehicle's potential state representation.
[0062] The world model employed in this invention is constructed based on a Recursive State Space Model Incorporating Physical Priors (RSSM-PINN), used for parallel inference and safety evaluation of candidate policy trajectories in the potential state space. It includes a state encoder, a state transition network, an observation decoder, a reward decoder, and a cost decoder; wherein:
[0063] The state encoder is used to jointly generate the vehicle's latent state representation by combining the acquired environmental perception information with the vehicle's historical trajectory state. .
[0064] State transition networks are based on given state-action pairs Constructing the posterior distribution It supports multi-step forward sampling prediction of potential state sequences at multiple future time points.
[0065] The observation decoder is used to reconstruct observations to improve the semantic consistency of state embeddings.
[0066] The reward decoder and cost decoder output the predicted reward values in the trajectory, respectively. Compared with cost forecasts This forms a triplet sequence of state-reward-cost.
[0067] The world model incorporates a priori physical structures and, combined with a simplified vehicle dynamics model, imposes structural constraints on acceleration, velocity changes, and trajectory curvature changes during state transitions. Its physical consistency regularization term is expressed as:
[0068]
[0069] In the formula: , These are the acceleration and steering angle estimated based on prior physical structures, respectively. , The predictive control output of the world model. , These are weighting coefficients, used to penalize predictions and physical model biases, respectively.
[0070] The world model is trained using a multi-task joint loss function, which is expressed as follows:
[0071]
[0072] In the formula: Indicates the observation reconstruction loss; Indicates the predicted loss from the reward; Indicates the predicted cost loss; This represents the loss of state consistency. Represents a physical consistency regularization term; , , , These represent the loss weights of reward prediction loss, cost prediction loss, state consistency loss, and physical consistency regularization term, respectively.
[0073] Step 3: Based on the obtained potential state representation of the vehicle, generate multiple sets of policy trajectory candidates through the multi-policy trajectory generation module.
[0074] This invention constructs a multi-policy trajectory generation module (i.e., the aforementioned state transition network) within a world model to generate multiple sets of candidate policy trajectories from the input vehicle's potential state representation. Specifically, the multi-policy trajectory generation module includes long-term policies. Short-term strategy and security strategies ,in:
[0075] Long-term strategy Built upon a soft actor critic-Lagrange constraint architecture, this system outputs long-term task trajectories with task constraints. It includes a policy network, a double Q-value network, and a state-value function network. The policy network generates continuous control actions *x* in the current state *s*, the double Q-value network evaluates the long-term reward of the actions in that state, and the V-network predicts the expected value of the policy in a given state. To achieve the control objective of maximizing expected reward and constraining risk and cost, a Lagrange constraint mechanism is introduced into the long-term policy, constructing an optimization objective function with cost constraints.
[0076]
[0077] in: For long-term strategy In a specific state-action pair The expected value below; For long-term strategy In a specific state-action pair The expected cumulative return; It is an entropy adjustment factor; Indicating long-term strategy The probability of generating action a given state s; For Lagrange multipliers, To preset the cost threshold, For long-term strategy Cost estimation under a specific state-action pair.
[0078] Long-term strategy The policy network parameters are updated using the following objective function:
[0079]
[0080] in: Let KL divergence be the KL divergence. As a normalization term, the Q network and V network use the soft Bellman equation to regress the target value, thereby optimizing the consistency of policy value estimation.
[0081] Long-term strategy During training, a hybrid experience pool is constructed by integrating real-world interactive datasets with simulated data generated by world model extrapolation. In each training cycle, target values are augmented by weighted sampling of samples from different sources and combined with rollout trajectories.
[0082] Long-term strategy The output candidate policy trajectory is provided as a reference path to the policy fusion module and is given priority for action generation when the expected cumulative return and cost estimate meet the requirements.
[0083] Short-term strategy This strategy is used to generate short-term elite response trajectories. It filters elite trajectories that meet cost constraints based on trajectory data generated by a world model, extracts corresponding actions, and fits them to a Gaussian distribution to form an immediate response strategy that is more adaptable to the current environment. Short-term strategy. It has the following characteristics:
[0084] 1) Within each control cycle, the world model is based on the current potential state representation. The system performs parallel simulations of multiple candidate policy trajectories output by the world model and evaluates the cumulative reward of each candidate policy trajectory. With accumulated costs .
[0085] Cumulative Rewards With accumulated costs Calculate according to the following formulas respectively:
[0086]
[0087]
[0088] In the formula, This represents the cumulative reward for each candidate policy trajectory output during control period t. This represents the cumulative cost of each candidate policy trajectory output during control period t.
[0089] 2) Short-term strategy Using preset screening criteria, a set of low-risk trajectories is selected from each candidate strategy trajectory. Then, elite trajectories are selected from these low-risk trajectories to serve as the short-term elite response trajectory for the current control moment. The low-risk trajectory set represents the accumulated cost among the candidate strategy trajectories. The set consisting of all candidate policy trajectories The preset cost threshold is used; the elite trajectory is the candidate strategy trajectory with the highest reward or lowest cost in the set of low-risk trajectories.
[0090] 3) The action sequence in the elite trajectory Used to fit Gaussian policy distribution , where: mean With covariance The following KL divergences are obtained by minimizing:
[0091]
[0092] In the formula, Represents an elite trajectory sequence; It is an empirical conditional distribution based on elite trajectory sequences.
[0093] Security Policy A rule-based control method based on an Intelligent Driving Model (IDM) is used to generate a deployment safety trajectory. The IDM model calculates the current safety acceleration using the following formula:
[0094]
[0095] Where: Where: This is the vehicle's maximum acceleration; Current vehicle speed; The desired driving speed; This is the speed control index, used to adjust the degree of nonlinear decay of the speed term; To achieve the desired minimum safe distance; The speed difference with the vehicle in front; This represents the current actual distance between the vehicle and the vehicle in front.
[0096] Step 4: Perform parallel simulation and deduction of the trajectories of each candidate strategy in the world model, and calculate the state evolution process, cumulative reward value, cumulative cost value and potential state stability index of each candidate strategy trajectory in the potential space of the world model. At the same time, conduct a unified evaluation of the trajectory quality of each candidate strategy trajectory according to the preset risk criteria.
[0097] Specifically, firstly, a multi-step potential state deduction is performed using the state transition network in the world model, based on the current initial state. The potential state sequence corresponding to each trajectory is calculated sequentially, and the dynamic rationality and physical feasibility of the trajectory are evaluated using the physical consistency module. Subsequently, the cumulative reward of each candidate strategy trajectory is calculated based on the reward function and cost function. With accumulated costs Furthermore, a behavior consistency index based on variance or constraint violation frequency is introduced to quantify the stability of the corresponding candidate strategy trajectories over long time scales. The above multi-index evaluation results serve as the core basis for trajectory selection and fusion strategy weight allocation, ensuring that the final output fusion strategy balances task completion efficiency and execution risk control.
[0098] Step 5: Construct a strategy fusion module to use the cumulative reward value, cumulative cost value and potential state stability index of each candidate strategy trajectory in the potential space as the fusion basis, dynamically fuse the action output of each candidate strategy trajectory to generate the final action, which is directly applied to the vehicle control interface to achieve end-to-end control execution.
[0099] The strategy fusion module adopts a hybrid expert structure and focuses on long-term strategies. Short-term strategy With security policy As candidate experts, based on the cumulative rewards corresponding to their respective trajectories. Cost estimation Potential state stability index Construct a fusion weight allocation mechanism to output the final action. :
[0100]
[0101] in ; Indicating long-term strategy Output acceleration; Indicating short-term strategy Output acceleration; Indicate security policy Output acceleration; ; Indicating long-term strategy The fusion weights of the output acceleration; Indicating short-term strategy The fusion weights of the output acceleration; Indicate security policy The fusion weights of the output acceleration.
[0102] As a further optimization of the present invention, the fusion weight of the strategy fusion module... Adopting a risk-sensitive weighting strategy, integrating weights It is obtained through the following calculation method:
[0103]
[0104] in, For the first The cumulative reward for each strategy-generated trajectory; The cumulative cost of the strategy trajectory (such as collision risk, etc.); It is a measure of the state transition of a trajectory in the latent space, used to characterize the stability of the trajectory; , , The fusion coefficient is used to adjust the reward-driven strength, cost-penalty intensity, and trajectory stability weights respectively. Its value can be configured according to the safety requirements and behavioral strategy preferences in different scenarios, or automatically optimized during training through data-driven methods.
[0105] To enhance the modeling capability for risk control, the world model introduces a latent state transition detection mechanism, which detects the latent state sequence composed of consecutive latent state representations generated in step 2. The KL divergence across steps is monitored. If KL fluctuations exceeding a preset threshold occur within several consecutive steps, the strategy trajectory is marked as having potential jump risk. The stability index is defined as follows:
[0106]
[0107] The policy fusion module is equipped with a security policy fallback mechanism, specifically:
[0108] The policy fusion module, based on the received long-term policy and short-term strategies The output candidate strategy trajectories are identified to determine whether both are candidate strategy trajectories with potential jump risks or whether the cumulative costs corresponding to both exceed the limit. If the identification result shows that both are candidate strategy trajectories with potential jump risks or the judgment result shows that the cumulative costs corresponding to both exceed the limit, the safety policy retreat mechanism is triggered.
[0109] When the security policy yield mechanism is triggered, the weighted fusion control actions are performed as follows: :
[0110]
[0111] in: , for fusion weights; For security strategy Output safety acceleration; Indicates a long-term strategy Short-term strategy As candidate experts, based on the cumulative rewards corresponding to their respective trajectories. Cost estimation Potential state stability index The fusion action output by constructing the fusion weight allocation mechanism is calculated using the following formula:
[0112]
[0113] In the formula: Indicating long-term strategy The fusion weights of the output acceleration; Indicating short-term strategy The fusion weights of the output acceleration; Indicating long-term strategy Output acceleration; Indicating short-term strategy Output acceleration.
[0114] End-to-end control execution, where the final control action is directly output by the fusion strategy module and used for vehicle control execution, does not rely on traditional path planning modules or independent controller components, forming a closed-loop adaptive control process from environmental perception to control execution. The end-to-end process includes: first, encoding the multimodal sensor data currently perceived by the vehicle through a visual Transformer structure (perception encoding module) to generate environmental perception information; then, inputting a recursive state space model (world model) based on fused physical priors to output the current potential state representation; next, generating multiple candidate strategy trajectories from long-term, short-term, and safety strategies respectively; then, inputting these candidate strategy trajectories into the fusion strategy module for weighted combination, ultimately outputting a continuous control action signal. This continuous control action signal is directly used to control the vehicle's acceleration, steering angle, or braking system, achieving action-level control without intermediate path planning or trajectory tracking.
[0115] Step 6: Collect environmental perception information, execution actions, and trajectory feedback results in the autonomous driving scenario to form a real-world interaction dataset; combine the predictive control variables generated by the world model, and jointly optimize the model parameters of the two modules in the world model that generate the potential state representation of the vehicle (state modeling network) and the candidate policy trajectory of the vehicle (policy network) through an alternating training mechanism to form a closed-loop control architecture with adaptive evolution capabilities, effectively improving the policy generalization performance and behavioral robustness of the system in complex traffic scenarios.
[0116] The closed-loop control architecture combines driving data from real-world interactions with simulated data generated by the world model, alternating between these data and the joint optimization of the policy network and the world model. This enables continuous evolution and improved generalization capabilities of the control strategy. During training, the long-term policy network and the fusion policy module periodically construct mixed training batches using real-world and simulated data in proportion. This improves the policy's responsiveness to unseen states while maintaining consistency with real-world behavior. The world model undergoes iterative training by incorporating newly collected data, continuously optimizing its state prediction accuracy and risk assessment capabilities. To enhance the stability and deployment consistency of the fusion strategy, the hybrid expert structure employs behavioral distillation during the training phase. Actions generated by the fusion strategy are used as supervisory signals to guide long-term and short-term strategies towards their behavioral distributions, mitigating oscillations and behavioral jumps during multi-strategy fusion.
[0117] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0118] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0119] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0120] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment, causing a series of operational steps to be executed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that run on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0121] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0122] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. An autonomous driving control method based on world model deduction and hybrid expert policy fusion, characterized in that, Includes the following steps: Step 1: Acquire multimodal sensor data of the vehicle in the current autonomous driving scenario and generate environmental perception information; Step 2: Generate the vehicle's potential state representation by combining the obtained environmental perception information with the vehicle's historical trajectory state through the world model; wherein: the world model is constructed based on a recursive state space model that integrates physical priors; Step 3: Based on the obtained potential state representation of the vehicle, generate multiple sets of candidate policy trajectories in the world model using the multi-policy trajectory generation module; the multi-policy trajectory generation module includes long-term policies. Short-term strategy and security strategies Among them: long-term strategy This is used to output long-term task trajectories with task constraints, and it is constructed based on a soft actor-critic Lagrange constraint architecture; short-term strategy Used to generate short-term elite response trajectories, it filters multiple candidate policy trajectories generated by the world model for the current control cycle to select elite trajectories that meet cost constraints, and performs corresponding action extraction and Gaussian distribution fitting to form an instantaneous response policy trajectory with stronger adaptability to the current environment; security strategy It is used to generate deployment safety trajectories. It adopts a rule control method based on intelligent driving model and provides stable yielding actions to ensure deployment safety when other strategies have high-risk prediction results or unstable behavior. Step 4: Perform parallel simulation and deduction of the trajectories of each candidate strategy in the world model and calculate the state evolution process, cumulative reward value, cumulative cost value and potential state stability index of each candidate strategy trajectory in the potential space of the world model. At the same time, conduct a unified evaluation of the trajectory quality of each candidate strategy trajectory according to the preset risk criteria. Step 5: Construct a strategy fusion module to use the cumulative reward value, cumulative cost value and potential state stability index of each candidate strategy trajectory in the potential space as the fusion basis, dynamically fuse the action output of each candidate strategy trajectory to generate the final action, which is directly applied to the vehicle control interface to achieve end-to-end control execution. The strategy fusion module adopts a hybrid expert structure and focuses on long-term strategies. Short-term strategy With security policy As candidate experts, based on the cumulative rewards corresponding to their respective trajectories. Cost estimation Potential state stability index Construct a fusion weight allocation mechanism to output the final action. : ; in ; Indicating long-term strategy Output acceleration; Indicating short-term strategy Output acceleration; Indicate security policy Output acceleration; ; Indicating long-term strategy The fusion weights of the output acceleration; Indicating short-term strategy The fusion weights of the output acceleration; Indicate security policy The fusion weights of the output acceleration; The world model introduces a potential state transition detection mechanism to monitor the step KL divergence of the potential state sequence composed of each continuous potential state representation generated in step 2, thereby marking the candidate policy trajectory with potential transition risk. The policy fusion module is also equipped with a security policy fallback mechanism, specifically: The policy fusion module, based on the received long-term policy and short-term strategies The output candidate strategy trajectories are identified to determine whether both are candidate strategy trajectories with potential jump risks or whether the cumulative costs corresponding to both exceed the limit. If the identification result shows that both are candidate strategy trajectories with potential jump risks or the judgment result shows that the cumulative costs corresponding to both exceed the limit, the safety policy retreat mechanism is triggered. When the security policy yield mechanism is triggered, the weighted fusion control actions are performed as follows: : ; in: , for fusion weights; For security strategy Output safety acceleration; Indicates a long-term strategy Short-term strategy As candidate experts, based on the cumulative rewards corresponding to their respective trajectories. Cost estimation Potential state stability index The fusion action output by constructing the fusion weight allocation mechanism is calculated using the following formula: ; In the formula: Indicating long-term strategy The fusion weights of the output acceleration; Indicating short-term strategy The fusion weights of the output acceleration; Indicating long-term strategy Output acceleration; Indicating short-term strategy Output acceleration.
2. The autonomous driving control method based on world model deduction and hybrid expert policy fusion according to claim 1, characterized in that, It also includes step 6: collecting environmental perception information, execution actions and trajectory feedback results in the autonomous driving scenario to form a real-world interaction dataset; combining the predictive control variables generated by the world model, and jointly optimizing the model parameters of the two modules that generate the potential state representation of the vehicle and the candidate policy trajectory of the vehicle in the world model through an alternating training mechanism.
3. The autonomous driving control method based on world model deduction and hybrid expert policy fusion according to claim 2, characterized in that, The world model incorporates a priori physical structures and, combined with a simplified vehicle dynamics model, imposes structural constraints on acceleration, velocity changes, and trajectory curvature changes during state transitions. Its physical consistency regularization term is expressed as: ; In the formula: , These are the acceleration and steering angle estimated based on prior physical structures, respectively. , The predictive control output of the world model. , These are weighting coefficients, used to penalize predictions and physical model biases, respectively.
4. The autonomous driving control method based on world model deduction and hybrid expert policy fusion according to claim 3, characterized in that, The world model is trained using a multi-task joint loss function, which is expressed as follows: ; In the formula: Indicates the observation reconstruction loss; Indicates the predicted loss from the reward; Indicates the predicted cost loss; This represents the loss of state consistency. Represents a physical consistency regularization term; , , , These represent the loss weights of reward prediction loss, cost prediction loss, state consistency loss, and physical consistency regularization term, respectively.
5. The autonomous driving control method based on world model deduction and hybrid expert policy fusion according to claim 4, characterized in that, Long-term strategy By introducing the Lagrange constraint mechanism, an optimization objective function with cost constraints is constructed: ; in: For long-term strategy In a specific state-action pair The expected value below; For long-term strategy In a specific state-action pair The expected cumulative return; It is an entropy adjustment factor; Indicating long-term strategy The probability of generating action a given state s; For Lagrange multipliers, To preset the cost threshold, For long-term strategy Cost estimation under specific state-action pairs; Long-term strategy The policy network parameters are updated using the following objective function: ; in: Let KL divergence be the KL divergence. For the normalization term, the Q network and V network use the soft Bellman equation to regress the target value; Long-term strategy During training, a hybrid experience pool is constructed by integrating real-world interactive datasets with simulated data generated by world model inference; and in each training cycle, target values are augmented by weighted sampling of samples from different sources and combined with rollout trajectories. Long-term strategy The output long-term task trajectory is provided as a reference path to the policy fusion module and is given priority for action generation when the expected cumulative return and cost estimate meet the requirements.
6. The autonomous driving control method based on world model deduction and hybrid expert policy fusion according to claim 5, characterized in that, Short-term strategy It has the following characteristics: exist Within each control cycle, the world model is based on the current potential state representation. Parallel simulations were performed on multiple candidate policy trajectories, and the cumulative reward of each candidate policy trajectory was evaluated. With accumulated costs ; Short-term strategy Using preset screening criteria, a set of low-risk trajectories is selected from all candidate strategy trajectories. Then, elite trajectories are selected from these to serve as the candidate strategy trajectories for the current control moment. The low-risk trajectory set represents the accumulated cost among the candidate strategy trajectories. The set consisting of all candidate policy trajectories The preset cost threshold is used; the elite trajectory is the candidate strategy trajectory with the highest reward or lowest cost in the set of low-risk trajectories. The action sequences in the elite trajectory are used to fit a Gaussian policy distribution. , where: mean With covariance The following KL divergences are obtained by minimizing: ; In the formula, Represents an elite trajectory sequence; It is an empirical conditional distribution based on elite trajectory sequences.
7. The autonomous driving control method based on world model deduction and hybrid expert policy fusion according to claim 6, characterized in that, Security Policy The IDM model is used to generate longitudinal control actions for the vehicle; the IDM model calculates the current safe acceleration using the following formula: ; In the formula: This is the vehicle's maximum acceleration; Current vehicle speed; The desired driving speed; This is the speed control index, used to adjust the degree of nonlinear decay of the speed term; To achieve the desired minimum safe distance; The speed difference with the vehicle in front; This represents the current actual distance between the vehicle and the vehicle in front.
8. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed in a computer, is used to implement the autonomous driving control method based on world model deduction and hybrid expert policy fusion as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Vehicle control method and device
CN116118767A
End-to-end automatic driving track evaluation method and system based on world model
CN120356177A
Method and device for generating automatic driving decision based on hybrid expert model
CN118928464A
Decision-making method for intelligent predictive driving of networked automobile based on machine learning
CN120197129A