An Autonomous Driving Method Based on Hierarchical World Cognition and Intent Constraint Generation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-08-11
AI Technical Summary
相较于前两类方法,这类方法在一定程度上缓解了语义推理与动作执行之间的脱节问题,但普遍存在以下不足:一是缺乏显式的空间认知机制,难以充分建模复杂交通场景中的三维几何结构和物理边界;二是难以从多源异构信息中稳定提取安全关键目标和关键交互关系,容易受到背景噪声干扰;三是对未来环境演化趋势和潜在交互结果缺乏显式预测能力,在高动态、强交互场景下难以兼顾规划的安全性、前瞻性与鲁棒性
[0112]与现有技术相比,本发明具有的有益效果是:本发明首先通过构建由物理边界认知层和行为意图认知层组成的层级世界认知表征,将环视视觉观测、自车状态、历史轨迹、导航目标和文本驾驶指令进行统一组织,使模型能够分别理解环境边界约束和驾驶任务意图,减少多模态信息直接混合带来的干扰。其次,本发明通过层级抑扰语义推理限制不同认知层之间的无关信息传播,使行为意图认知层能够受控读取与当前决策相关的边界信息,并凝聚得到包含道路结构、交通参与者动态、导航导向和短期驾驶意图的世界演化核,为后续轨迹生成提供稳定条件。最后,本发明基于世界演化核构建意图约束扩散规划器,通过条件中心、条件交叉读取、状态门控
和约束更新量
对去噪轨迹进行逐步调制,并结合安全距离约束、意图对齐约束、规划一致性约束以及动力学投影和平滑校准,生成满足场景边界、驾驶意图和车辆运动约束的连续规划轨迹。与直接轨迹回归或自回归动作生成方法相比,本发明能够减小长时序误差累积,提高复杂交互场景下轨迹生成的连续性、平滑性和可执行性,并改善自动驾驶系统在动态避障、转弯盲区和复杂路口等长尾场景中的规划稳定性与安全性。
Smart Images

Figure CN122390087B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, specifically to an autonomous driving method based on hierarchical world cognition and intention constraint generation. Background Technology
[0002] End-to-end autonomous driving refers to the use of data-driven and differentiable learning models to directly map raw perception information into vehicle control behavior or future trajectories. Due to its advantage of integrated modeling of perception, decision-making, and planning, it has become an important research direction in autonomous driving in recent years. However, the application of existing end-to-end autonomous driving methods in complex real-world traffic environments still faces significant limitations, mainly manifested in the lack of interpretability of the model decision-making process, insufficient generalization ability to complex and long-tail scenarios, and difficulty in maintaining stable and reliable planning performance under rare and dangerous conditions. Therefore, how to improve the model's planning capabilities while enhancing its semantic understanding, spatial cognition, and decision transparency has become a key problem that needs to be addressed in this field.
[0003] With the development of visual language models, autonomous driving research has gradually evolved from the traditional perception-centric end-to-end policy learning approach to an intelligent driving modeling approach with high-level semantic understanding, logical reasoning, and task interaction capabilities. Unlike traditional methods that focus on directly regressing driving actions or trajectories from sensor input, autonomous driving methods based on visual language models place greater emphasis on semantic modeling of complex traffic scenarios, explicit reasoning about driving behavior, and interpretable expression of decision results. By introducing language priors and multimodal alignment mechanisms, these methods have shown great potential in complex scene understanding, navigation command compliance, and high-level behavioral decision-making.
[0004] Existing end-to-end autonomous driving methods based on visual language models can be broadly classified into three categories.
[0005] The first type of approach employs a dual-system collaborative architecture, where the visual language model primarily handles high-level scene understanding and decision-making reasoning. This type of approach typically begins by having the visual language model analyze the traffic environment and output meta-actions, text-based driving commands, or other high-level behavioral intentions. The downstream planning and control model then executes specific trajectory generation and control operations based on these intentions. While this type of approach has made some progress in complex scene understanding, behavioral interpretation, and reasoning transparency, it suffers from a significant separation between high-level semantic decision-making and low-level motion generation. This makes it difficult for the overall system to achieve true end-to-end joint optimization, thus limiting the model's collaborative modeling capabilities in complex driving tasks.
[0006] The second type of approach directly uses the visual language model as the core planning and decision-making unit in the end-to-end driving system. Unlike the dual-system collaborative architecture, this type of approach uses a unified modeling of visual input, text commands, and navigation information to directly utilize the multimodal representation and reasoning capabilities of the visual language model to generate driving plans, behavior sequences, or trajectory expressions. This type of approach has strong advantages in high-level semantic reasoning and result interpretability; however, its outputs are often represented in discrete text, symbolic, or intermediate semantic action forms, making it difficult to achieve a natural connection with the continuous action space required for vehicle control. This results in difficulties in effectively guaranteeing geometric feasibility, dynamic consistency, and trajectory smoothness.
[0007] The third type of approach is an end-to-end driving framework integrating visual language and action, inspired by embodied intelligence and visual-language-action modeling. This type of approach establishes a closer connection between high-level semantic decisions and low-level continuous actions by jointly modeling visual, linguistic, and action representations within a unified architecture and introducing a dedicated action planner or trajectory generator. Compared to the first two types of approaches, this type alleviates the disconnect between semantic reasoning and action execution to some extent, but it generally suffers from the following shortcomings: First, it lacks an explicit spatial cognitive mechanism, making it difficult to fully model the three-dimensional geometric structure and physical boundaries in complex traffic scenarios; second, it struggles to stably extract key safety objectives and key interaction relationships from multi-source heterogeneous information, making it susceptible to background noise interference; third, it lacks explicit predictive ability for future environmental evolution trends and potential interaction outcomes, making it difficult to balance the safety, foresight, and robustness of planning in highly dynamic and highly interactive scenarios. These problems result in existing visual-language-action driving models still exhibiting deficiencies in physical consistency, long-term planning stability, and trajectory generation reliability in complex conditions.
[0008] To address the aforementioned issues, a novel autonomous driving technology solution is needed. This solution should enable the system to not only extract high-quality hierarchical semantic information from multimodal inputs but also to establish a spatial cognitive representation consistent with the real driving environment. Based on this, effective connection between high-level semantic intent and low-level continuous actions can be achieved. Particularly in scenarios involving complex traffic participant interactions, occlusion, blind spot navigation, and dynamic obstacle avoidance, the system needs to simultaneously possess scene geometric understanding, key target recognition, future state prediction, and generative trajectory planning capabilities to meet the comprehensive requirements of autonomous driving for safety, stability, and generalization. Summary of the Invention
[0009] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.
[0010] To solve the above-mentioned technical problems, according to one aspect of the present invention, the present invention provides the following technical solution: an autonomous driving method based on hierarchical world cognition and intention constraint generation, comprising the following steps: S1: Constructing hierarchical world cognition: First, the surround-view visual observation, vehicle status, historical trajectory, navigation target and text driving instructions are uniformly encoded to obtain multimodal lexical representations; then, these lexical representations are organized into two cognitive layers: physical boundary cognitive layer and behavioral intention cognitive layer;
[0011] S2: Perform hierarchical perturbation reasoning and generate world evolution kernel: After obtaining hierarchical world cognition, the information propagation between the two cognitive layers is controlled by hierarchical perturbation mask. The physical boundary cognitive layer mainly maintains a stable expression of the spatial boundary and obstacle state, while the behavioral intention cognitive layer reads the boundary information related to the current driving decision when needed. After controlled reasoning, the world evolution kernel is obtained.
[0012] S3: Execution Intent Constraint Diffusion Planning: The world evolution kernel, navigation priors, behavior-level semantic representation, and motion-level semantic representation are collectively constituted as the condition center. The conditional center is then injected into the diffusion planner: the noisy trajectory is first encoded as trajectory features, the conditional center is read by the trajectory features through cross attention, and then through state gating. Control condition injection strength, forming constraint update quantity During training, safety distance constraints, intent alignment constraints, and planning consistency constraints are introduced; during inference, dynamic projection and smooth calibration are combined to finally output a continuous planned trajectory that satisfies scene boundaries, driving intent, and vehicle motion constraints.
[0013] As a preferred embodiment of the autonomous driving method based on hierarchical world cognition and intent constraint generation described in this invention, the specific steps of S1 are as follows:
[0014] S1.1: Multimodal feature preprocessing:
[0015] Input includes A 6-channel surround view image sequence of frames Navigation target point sequence and text driving instructions Among them, the surround view image sequence is used to characterize the road structure, traffic participant status and spatial boundary information of the environment around the vehicle; the navigation target point sequence is used to characterize the vehicle's target driving direction and path constraints; and the text driving instructions are used to provide high-level task semantic information.
[0016] Text and navigation encoding: text instructions After BPE word segmentation, features are extracted by a text encoder initialized with Qwen2. Navigation target point It is then converted into navigation terms through a three-layer MLP adapter. ;
[0017] Visual representation compression: for surround-view driving observation We employ the AnyRes strategy combined with bilinear interpolation for preprocessing to balance perception resolution and computational cost.
[0018] S1.2: Initialization of Hierarchical World Cognition:
[0019] After completing the multimodal feature preprocessing, two sets of learnable hierarchical world cognition vectors are defined, namely the physical boundary cognition vector. and behavioral intention cognitive vector To enable the two types of cognitive vectors to dynamically and adaptively adjust according to the current driving state and task objectives, hierarchical dynamic routing weights are further generated:
[0020]
[0021] in, Indicates the hierarchical dynamic routing weight; Represents the cognitive vector acting on the physical boundary. Routing weights This represents the action on the cognitive vector of behavioral intention. Routing weights; Represents the route weight mapping matrix. Indicates the vehicle's state characteristics. Indicates the characteristics of historical trajectory, Indicates navigation features, Representing text features;
[0022] Then, a hierarchical world cognition field vector with dynamic prior terms is constructed:
[0023]
[0024] in, Represents the physical boundary perception vector; Represents a cognitive vector of behavioral intention; Indicates the hierarchical position encoding; This represents a dynamic prior code term composed of the vehicle's state and historical trajectory.
[0025] S1.3: State-biased cross-attention mechanism:
[0026] A cross-attention mechanism is introduced, where state bias and target bias work together. The attention score is expressed as follows:
[0027]
[0028] Further, the initialized hierarchical world cognitive representation is obtained:
[0029]
[0030] in, , and These represent the mapping matrices for queries, keys, and values, respectively. This represents the surround-view visual feature lexical units obtained after visual encoder and lexical compression; This represents the feature dimension; the attention is used to make the hierarchical cognitive vectors derived from the features of the visual environment. Read information related to the current state and task objectives. Represents the motion state bias operator; This represents the goal-oriented bias operator.
[0031] As a preferred embodiment of the autonomous driving method based on hierarchical world cognition and intent constraint generation described in this invention, the specific method for visual representation compression in S1.1 is as follows:
[0032] Introducing an entropy-guided cognitive focusing operator Its core logic is to prioritize preserving regions in the feature map that exhibit significant gradient changes and dense semantic information.
[0033]
[0034]
[0035] in, This represents the original visual word sequence output by the visual encoder; Represents the original visual word set or the corresponding feature map index; This represents the set of key visual terms retained after entropy-guided filtering; This is the threshold for the number of key visual words, used to control the number of key visual words retained after compression. This represents the function for calculating the entropy of visual lexical information; This indicates selecting the option with the highest information content. Each word element; It is the original input loop view image sequence; This is a scaling factor used to adjust the spatial scale after pooling; Indicates a visual encoder; These are the key visual feature terms after compression.
[0036] As a preferred embodiment of the autonomous driving method based on hierarchical world cognition and intent constraint generation described in this invention, the specific steps of S2 are as follows:
[0037] S2.1: Design of a hierarchical scrambling mask mechanism:
[0038] Let the i-th and j-th words belong to a certain cognitive layer, then the hierarchical scrambling mask matrix is defined as follows:
[0039]
[0040] in, Indicates the cognitive level to which a word belongs. This indicates the cross-level suppression strength, which is jointly determined by the current word state, the vehicle state, and the behavioral intent cognitive layer features:
[0041]
[0042] in, and They represent the first The word element and the first The current hidden representation of each word element. Indicates the vehicle's motion status. Represents the suppression intensity mapping matrix;
[0043] The restricted attention reasoning form after introducing hierarchical scrambling masks is as follows:
[0044]
[0045] in, These represent the query, key, and value matrices obtained by linear mapping from the input hidden states in the current inference layer, respectively. Represents the hierarchical scrambling mask matrix; This represents the hidden representation after hierarchical perturbation constraints;
[0046] S2.2: Hierarchical Reasoning Expansion and World Evolution Kernel Extraction:
[0047] In this reasoning process, the physical boundary cognition layer mainly completes the aggregation of spatial boundaries and traffic participant states, while the behavioral intention cognition layer, while maintaining the semantic dominance of the task, is controlled to read the key boundary summary in the physical boundary cognition layer, calculates the reasoning result of the physical boundary layer, and is controlled to read the physical boundary summary.
[0048] After completing hierarchical reasoning, behavioral-level semantic representations and motion-level semantic representations are generated:
[0049]
[0050] in, This indicates behavioral-level semantic output; Indicates motion-level semantic output; The aggregated features representing the cognitive layer of behavioral intent after hierarchical perturbation reasoning; For behavior semantic mapping header; For motion trend mapping head;
[0051] To condense the world evolution kernel from the last few hidden states of the main reasoning framework, first calculate the inter-layer aggregation weights:
[0052]
[0053] in, Indicates the main branch of reasoning. The hidden state of the layer; Indicates the total number of layers in the main reasoning structure; Indicates the importance scoring function of the layer; Indicates the first The aggregation weight of the hidden layer states in the condensation of the world evolution nucleus;
[0054] Then calculate the fusion gating between physical boundary information and behavioral intent information:
[0055]
[0056] in, Indicates the first Aggregated representation of hidden states in the physical boundary cognition layer; Indicates the first Aggregated representation of the hidden states of the cognitive layer of behavioral intentions; Represents the fusion gated mapping matrix; Indicates the first The fusion gating coefficient between layer physical boundary information and behavioral intent information;
[0057] Ultimately, the world evolutionary nucleus is defined as:
[0058]
[0059] in, Represents the world's evolutionary nucleus. Indicates the inter-layer aggregate weight. Indicates hierarchical fusion gating, This indicates element-wise multiplication.
[0060] As a preferred embodiment of the autonomous driving method based on hierarchical world cognition and intent constraint generation described in this invention, the calculation method of the reasoning result of the physical boundary layer in S1.2 is as follows: let the query, key, and value corresponding to the physical boundary layer be respectively... The query, key, and value corresponding to the behavioral intent layer are respectively The inference result of the physical boundary layer is then expressed as:
[0061]
[0062] While maintaining its semantic dominance, the behavioral intent cognitive layer also reads the physical boundary summary in a controlled manner.
[0063]
[0064]
[0065] in, Indicates the suppression coefficient. Represents the cross-layer projection operator; Aggregated features representing the cognitive layer of behavioral intention; This represents the injection coefficient mapping matrix.
[0066] As a preferred embodiment of the autonomous driving method based on hierarchical world cognition and intent constraint generation described in this invention, the world evolution kernel comprehensively includes road boundaries, obstacle occupancy, key traffic participant dynamics, navigation guidance, and short-term behavioral intent.
[0067] As a preferred embodiment of the autonomous driving method based on hierarchical world cognition and intent constraint generation described in this invention, the specific method of S3 is as follows:
[0068] First, construct the intention constraint center. ;
[0069] During the training phase, the actual future trajectory After adding noise through forward diffusion, the first... Noisy trajectory at each diffusion time step:
[0070]
[0071] in, Indicates Gaussian noise. Represents the noise scheduling coefficient, noisy trajectory First, the action embedding layer maps the data to trajectory features, and then the time steps are diffused. Simultaneously, after passing through the time-step embedding layer, both are input into the Transformer-based denoising module, and the internal relationships between future trajectory points are modeled through self-attention to obtain the intermediate trajectory representation. ;
[0072] Then As a query, the condition center After conditional embedding, the data is used as the key and value to perform cross-attention reading:
[0073] in, Indicates the first The intermediate trajectory features obtained by noisy trajectories at each diffusion time step after trajectory embedding, time step embedding and self-attention modeling; Represents the cross-attention operator; As a query, As Key and Value; This represents the conditional modulation result read from the conditional center of the trajectory features;
[0074] No. During denoising, each future trajectory point will be determined based on its current state, from... It reads relevant scene boundaries, navigation directions, behavioral intentions, and planning consistency information.
[0075] The following introduces a diffusion-state gating modulation mechanism, which projects the intention constraint center and the world evolution kernel onto the trajectory generation space, and dynamically combines them using gating coefficients:
[0076]
[0077] in, Indicates the first The state gating coefficients at each diffusion time step Represents the state-semantic joint mapping matrix. This represents the activation function, where the gating coefficients determine the respective roles of the condition center and the world evolution kernel in the current denoising stage.
[0078] Projecting the intention constraint center and the world evolution kernel onto the trajectory generation space, and then dynamically combining them using the aforementioned gating coefficients, yields the constraint update quantity:
[0079]
[0080] in, Indicates the first Constraint update amount under each diffusion time step; This represents the projection of the condition center reading result into the trajectory generation space; Represents the projection of the world evolution nucleus into the trajectory generation space;
[0081] Ultimately, the diffusion planner at time step The trajectory update result is represented as follows:
[0082]
[0083] in, Indicates the first Trajectory characteristics modulated at each diffusion time step; This represents a Transformer-based denoising transform operator; Indicates the constraint update amount; the updated Input the action head to predict the noise at the current diffusion time step, and obtain Then, the diffusion planner predicts the noise at the current diffusion time step using the action head;
[0084] During the training phase, a joint optimization objective is constructed. First, the diffuse noise prediction loss is used as the basic reconstruction objective:
[0085]
[0086] in, This represents the actual noise injected during the forward diffusion process. The model represents the condition center. The noise estimate obtained from the prediction under constraints;
[0087] During the training phase, the spread noise prediction loss is used as the basic objective, combined with safety distance constraints, intention alignment constraints, and planning consistency constraints. The total loss function is expressed as:
[0088]
[0089] in, , and These represent the weight coefficients of the safety distance constraint, intention alignment constraint, and planning consistency constraint, respectively.
[0090] During the inference phase, a deterministic inverse diffusion recovery method is used to gradually reconstruct the noisy trajectory. Let the noisy trajectory at the current time step be... First, based on the noise estimation results, the intermediate noise-free trajectory representation is obtained:
[0091]
[0092] The trajectory state of the previous time step is calculated according to the inverse diffusion update rule:
[0093]
[0094] After multiple rounds of reverse diffusion recovery, the final trajectory potential representation is obtained.
[0095] As a preferred embodiment of the autonomous driving method based on hierarchical world cognition and intent constraint generation described in this invention, the condition center It consists of a world evolution kernel, navigation priors, behavior-level semantic representation, and motion-level semantic representation.
[0096]
[0097] in, This indicates the constraint integration process; This represents the global scenario conditions obtained by pooling the world evolution kernel; This represents a behavioral-level semantic representation. Indicates navigation priors, This represents the position code, which is actually used when entering the diffusion planner. First, it goes through a conditional embedding layer, and then through a conditional coding block to complete context modeling. This is used as conditional memory in the diffusion decoder. The conditional information is not just concatenated once at the input, but continuously participates in trajectory feature updates during the denoising process.
[0098] As a preferred embodiment of the autonomous driving method based on hierarchical world cognition and intent constraint generation described in this invention, a safety distance constraint is introduced to ensure trajectory safety. The specific method is as follows:
[0099] in, Indicates the safe distance threshold. Indicates the generated trajectory The distance between the boundary occupied by obstacles as represented by the physical boundary layer;
[0100] An intent-constrained alignment loss is introduced to ensure that the generated trajectory aligns with the high-level behavioral intent in a distributional sense. The specific method is as follows:
[0101]
[0102] in, This represents the distribution of behaviors identified from the generated trajectories. This represents the distribution of target behaviors obtained from behavior-level semantic representation; Indicates KL divergence;
[0103] A planning consistency loss is introduced to ensure that the generated trajectory remains consistent with the navigation target and local motion trends. The specific method is as follows:
[0104]
[0105] in, This represents the trajectory motion feature mapping function, used to map the generated trajectory to the motion semantic space; This represents motion-level semantic representation. This indicates the deviation between the trajectory and the navigation prior.
[0106] As a preferred embodiment of the autonomous driving method based on hierarchical world cognition and intent constraint generation described in this invention, a dynamic projection operator is introduced to ensure that the generated trajectory satisfies the actual motion constraints of the vehicle. Specifically, the estimated trajectory is mapped to the vehicle's executable trajectory space.
[0107]
[0108] in, This represents the trajectory estimate after dynamic projection; This represents the dynamic projection operator, used to correct components in the trajectory that do not satisfy vehicle dynamic constraints based on the current vehicle kinematics information;
[0109] Through the action head The final trajectory generation is completed using the dynamic smoothing calibration operator.
[0110]
[0111] in, This represents the dynamic smoothing calibration operator.
[0112] Compared with existing technologies, the advantages of this invention are as follows: First, by constructing a hierarchical world cognitive representation composed of a physical boundary cognitive layer and a behavioral intention cognitive layer, this invention unifies the surrounding visual observation, vehicle state, historical trajectory, navigation target, and text driving instructions, enabling the model to understand environmental boundary constraints and driving task intentions separately, reducing interference caused by the direct mixing of multimodal information. Second, this invention restricts the propagation of irrelevant information between different cognitive layers through hierarchical perturbation semantic reasoning, allowing the behavioral intention cognitive layer to controllably read boundary information related to the current decision and converge to obtain a world evolution kernel containing road structure, traffic participant dynamics, navigation guidance, and short-term driving intentions, providing stable conditions for subsequent trajectory generation. Finally, this invention constructs an intention constraint diffusion planner based on the world evolution kernel, through condition centers... Conditional cross-reading, state gating and constraint update amount The denoised trajectory is progressively modulated and combined with safety distance constraints, intent alignment constraints, planning consistency constraints, dynamic projection, and smoothing calibration to generate a continuous planned trajectory that satisfies scene boundaries, driving intent, and vehicle motion constraints. Compared with direct trajectory regression or autoregressive action generation methods, this invention can reduce the accumulation of long-term errors, improve the continuity, smoothness, and executability of trajectory generation in complex interactive scenarios, and improve the planning stability and safety of autonomous driving systems in long-tail scenarios such as dynamic obstacle avoidance, blind spots during turns, and complex intersections. Attached Figure Description
[0113] To more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and detailed embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0114] Figure 1 This is an architecture diagram of an autonomous driving method based on hierarchical world cognition and intent constraint generation according to the present invention. Detailed Implementation
[0115] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0116] Secondly, the present invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of the present invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not according to the usual scale. Furthermore, the schematic diagrams are merely examples and should not limit the scope of protection of the present invention. In addition, actual fabrication should include three-dimensional spatial dimensions of length, width, and depth.
[0117] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0118] The invention integrates multi-source heterogeneous information such as surround-view visual observation, vehicle status, historical motion trajectory, navigation target, and text driving instructions to achieve hierarchical cognitive modeling of complex traffic environments, extraction of key driving semantics, and generation of continuous planning trajectories. The core of this invention lies in the following: First, a hierarchical world cognitive representation is constructed in the input stage, organizing surround-view visual observation, vehicle status, historical trajectory, navigation target, and text driving instructions into a physical boundary cognitive layer and a behavioral intention cognitive layer. This enables the system to describe road boundaries, drivable areas, obstacle occupancy relationships, key traffic participant status, navigation target, text semantics, and short-term driving behavior guidance, respectively. Second, a hierarchical scrambling mask is introduced in the semantic reasoning stage to restrict the propagation of irrelevant information between different cognitive layers. This allows physical boundary information and behavioral intention information to interact in a controlled manner while maintaining the stability of their respective expressions, and to coalesce into a world evolution kernel for subsequent planning. Finally, the world evolution kernel, navigation prior, behavioral-level semantic representation, and motion-level semantic representation together constitute an intention constraint center, which is input into the diffusion planner. During the gradual denoising process, continuous action reconstruction is completed, and combined with safety distance constraints, intention alignment constraints, planning consistency constraints, dynamic projection, and smoothing calibration, a continuous planned trajectory that satisfies scene boundary, driving intention, and vehicle motion constraints is output.
[0119] Specifically, this invention provides an autonomous driving method based on hierarchical world cognition and intent constraint generation, which generally includes three main processes: multimodal input encoding, hierarchical world cognition construction, hierarchical perturbation-suppressed semantic reasoning, and intent constraint diffusion planning. For example... Figure 1 As shown, it includes: multimodal input encoding and hierarchical world cognition construction; hierarchical destrained semantic reasoning and world evolution kernel generation; and continuous trajectory generation based on intention-constrained diffusion planning.
[0120] S1: Multimodal Input Encoding and Hierarchical World Cognition Construction: In autonomous driving systems, raw data collected by sensors typically exhibits high dimensionality, significant modal differences, abundant redundant information, and strong perceptual noise. If these raw observations are directly input into subsequent inference and planning modules, the model easily mixes background textures, irrelevant targets, and crucial driving cues, affecting scene understanding and trajectory generation stability. Therefore, this invention first performs unified encoding on multi-source inputs.
[0121] The system first performs unified encoding on surround-view visual observations, vehicle status, historical trajectory, navigation targets, and text-based driving instructions to obtain multimodal lexical representations. These lexical representations are then organized into two cognitive layers: a physical boundary cognitive layer and a behavioral intention cognitive layer. The former primarily describes road boundaries, drivable areas, obstacle occupancy relationships, and the status of key traffic participants; the latter primarily describes navigation targets, textual semantics, and short-term driving behavior guidance. This approach aims to avoid all modal information being initially mixed into a single indistinguishable feature, enabling the model to separately understand "how the environment allows the vehicle to move" and "how the task requires the vehicle to move."
[0122] S1.1: Multimodal feature preprocessing:
[0123] System inputs include A 6-channel surround view image sequence of frames Navigation target point sequence and text driving instructions Among them, the surround view image sequence is used to characterize the road structure, traffic participant status, and spatial boundary information of the vehicle's surrounding environment; the navigation target point sequence is used to characterize the vehicle's target driving direction and path constraints; and the text driving instructions are used to provide high-level task semantic information.
[0124] Text and navigation encoding: text instructions After BPE word segmentation, features are extracted by a text encoder initialized with Qwen2. Navigation destination It is then converted into navigation terms through a three-layer MLP adapter. .
[0125] Visual representation compression: for surround-view driving observation We employ an AnyRes strategy combined with bilinear interpolation for preprocessing to balance perceptual resolution and computational cost. To prevent the computational explosion problem caused by excessive computation when processing large-scale observation features using the full attention mechanism, we introduce an entropy-guided cognitive focusing operator. Its core logic is to prioritize preserving regions in the feature map that exhibit significant gradient changes and dense semantic information.
[0126]
[0127]
[0128] in, This represents the original visual word sequence output by the visual encoder; Represents the original visual word set or the corresponding feature map index; This represents the set of key visual terms retained after entropy-guided filtering; This is the threshold for the number of key visual words, used to control the number of key visual words retained after compression. This represents the function for calculating the entropy of visual lexical information; This indicates selecting the option with the highest information content. Each word element; It is the original input loop view image sequence; This is a scaling factor used to adjust the spatial scale after pooling; Indicates a visual encoder; These are the key visual feature terms after compression.
[0129] S1.2: Initialization of Hierarchical World Cognition:
[0130] After completing the multimodal feature preprocessing, a set of learnable hierarchical world cognition vectors is defined:
[0131]
[0132] in, Indicates the hierarchical dynamic routing weight; Represents the cognitive vector acting on the physical boundary. Routing weights This represents the action on the cognitive vector of behavioral intention. Routing weights; Represents the route weight mapping matrix. Indicates the vehicle's state characteristics. Indicates the characteristics of historical trajectory, Indicates navigation features, Represents text features.
[0133] Then, a hierarchical world cognition vector with dynamic priors is constructed:
[0134]
[0135] in, The physical boundary perception vector is mainly used to characterize the road geometric boundary, drivable area, obstacle occupancy relationship and key traffic participant status. The cognitive vector representing behavioral intent is mainly used to characterize navigation goals, driving behavior orientation, and short-term task intent. Indicates the hierarchical position encoding; This represents a dynamic prior code term composed of the vehicle's state and historical trajectory.
[0136] S1.3: State-biased cross-attention mechanism:
[0137] To better align the hierarchical cognitive initialization results with the current vehicle motion state and driving task, this invention further introduces a cross-attention mechanism that combines state bias and target bias. The attention score is expressed as:
[0138]
[0139] Further, the initialized hierarchical world cognitive representation is obtained:
[0140]
[0141] in, , and These represent the mapping matrices for queries, keys, and values, respectively. This represents the surround-view visual feature lexical units obtained after visual encoder and lexical compression; This represents the feature dimension; the attention is used to make the hierarchical cognitive vectors derived from the features of the visual environment. It reads information related to the current state and task objectives. Represents the motion state bias operator; This represents the target-oriented bias operator. This mechanism enables the model to proactively select more relevant environmental features based on current speed, heading, historical trajectory, and navigation target.
[0142] S2: Hierarchical Suppression Semantic Reasoning and World Evolution Kernel Generation: After obtaining the hierarchical world cognitive representation, the system enters the semantic reasoning stage. Ordinary global self-attention causes all lexical units to interact indiscriminately, easily leading to information such as roadside textures, far-field backgrounds, and irrelevant objects interfering with key driving cues. To address this, this invention introduces a hierarchical suppression mask in the reasoning backbone to control the information propagation between the physical boundary cognitive layer and the behavioral intention cognitive layer.
[0143] The system controls information propagation between the two cognitive layers through hierarchical suppression masks. The physical boundary cognitive layer primarily maintains a stable representation of spatial boundaries and obstacle states, while the behavioral intention cognitive layer reads boundary information relevant to the current driving decision when needed, rather than indiscriminately receiving all environmental features. After controlled reasoning, the system condenses to obtain the world evolution kernel. This world evolution kernel is the core condition for the subsequent planner, containing road structure, traffic participant dynamics, navigation guidance, and short-term driving intentions.
[0144] S2.1: Design of the hierarchical scrambling mask mechanism: Let the i-th word and the j-th word belong to a certain cognitive layer, then the hierarchical scrambling mask matrix is defined as follows:
[0145]
[0146] in, Indicates the cognitive level to which a word belongs. This indicates the cross-level suppression strength, which is jointly determined by the current word state, the vehicle state, and the behavioral intent cognitive layer features:
[0147]
[0148] in, and They represent the first The word element and the first The current hidden representation of each word element. Indicates the vehicle's motion status. This represents the suppression intensity mapping matrix. At this point, cross-layer information propagation is no longer fixed, but changes depending on the current scenario and driving task.
[0149] The restricted attention reasoning form after introducing hierarchical scrambling masks is as follows:
[0150]
[0151] in, These represent the query, key, and value matrices obtained by linear mapping from the input hidden states in the current inference layer, respectively. Represents the hierarchical scrambling mask matrix; This represents the constrained reasoning result obtained under the hierarchical suppression mask. Unlike traditional unified global attention, this result already contains the dynamic suppression relationship between the physical boundary layer and the behavioral intent layer.
[0152] S2.2: Hierarchical Reasoning Unfolding and World Evolution Kernel Extraction: In this reasoning process, the physical boundary cognition layer mainly completes the aggregation of spatial boundaries and traffic participant states, while the behavioral intention cognition layer, while maintaining the task semantics as the main focus, is controlled to read the key boundary summary in the physical boundary cognition layer. Let the query, key, and value corresponding to the physical boundary layer be respectively... The query, key, and value corresponding to the behavioral intent layer are respectively The inference result of the physical boundary layer is then expressed as:
[0153]
[0154] While maintaining its semantic dominance, the behavioral intent cognitive layer also reads the physical boundary summary in a controlled manner.
[0155]
[0156]
[0157] in, Indicates the suppression coefficient. Represents the cross-layer projection operator; Aggregated features representing the cognitive layer of behavioral intention; This represents the injection coefficient mapping matrix. This design allows the behavioral intent cognition layer to receive boundary information relevant to the current decision, but avoids being directly contaminated by all environmental features.
[0158] After completing hierarchical reasoning, the system generates behavior-level semantic representations and motion-level semantic representations:
[0159]
[0160] in, It represents behavioral-level semantic output, used to describe high-level driving behaviors such as lane keeping, deceleration, waiting, steering, lane changing, and detouring; This represents motion-level semantic output, used to characterize semantic representations related to local trajectory trends, obstacle avoidance tendencies, and short-term motion patterns. The aggregated features representing the cognitive layer of behavioral intent after hierarchical perturbation reasoning; For behavior semantic mapping header; For motion trend mapping headers.
[0161] To provide more stable global conditions for the diffusion planner, this invention further condenses the world evolution kernel from the hidden states of the last few layers of the inference backbone. First, the inter-layer aggregation weights are calculated:
[0162]
[0163] in, Indicates the main branch of reasoning. The hidden state of the layer; Indicates the total number of layers in the main reasoning structure; Indicates the importance scoring function of the layer; Indicates the first The aggregation weight of the hidden layer states in the condensation of the world evolution nucleus.
[0164] Then calculate the fusion gating between physical boundary information and behavioral intent information:
[0165]
[0166] in, Indicates the first Aggregated representation of hidden states in the physical boundary cognition layer; Indicates the first Aggregated representation of the hidden states of the cognitive layer of behavioral intentions; Represents the fusion gated mapping matrix; Indicates the first The fusion gating coefficient between layer physical boundary information and behavioral intent information.
[0167] Ultimately, the world evolutionary nucleus is defined as:
[0168]
[0169] in, Represents the world's evolutionary nucleus. Indicates the inter-layer aggregate weight. Indicates hierarchical fusion gating, This indicates element-wise multiplication.
[0170] The world evolution kernel synthesis includes road boundaries, obstacle occupancy, key traffic participant dynamics, navigation guidance, and short-term behavioral intentions. Compared to directly using the last layer of hidden states, this representation can more stably provide scenario and intention conditions for subsequent trajectory generation.
[0171] S3: Continuous Trajectory Generation Based on Intent-Constrained Diffusion Planning: After obtaining the world evolution kernel, the system enters the trajectory generation stage. This invention does not directly regress future trajectory points, but instead uses conditional diffusion denoising to gradually generate continuous trajectories. This method treats future trajectories as results gradually recovered from noise, which can alleviate the problems of averaged trajectories generated by direct regression and long-term time-series error accumulation caused by autoregressive decoding.
[0172] Specifically, the trajectory generation stage does not directly regress future trajectory points, but instead uses a diffusion denoising method to gradually generate the trajectory. The system uses the world evolution kernel, navigation priors, behavior-level semantic representation, and motion-level semantic representation together to form the condition center. The condition center is then injected into the diffusion planner. Specifically, the noisy trajectory is first encoded as trajectory features, the condition center is read by the trajectory features through cross-attention, and then through state gating. Control condition injection strength, forming constraint update quantity This modulates the trajectory generation direction at each denoising stage. During training, safety distance constraints, intent alignment constraints, and planning consistency constraints are introduced; during inference, dynamic projection and smoothing calibration are combined to finally output a continuous planned trajectory that satisfies scene boundaries, driving intent, and vehicle motion constraints.
[0173] First, the system constructs the intention constraint center. This conditional center consists of a world evolution kernel, navigation priors, behavior-level semantic representation, and motion-level semantic representation.
[0174]
[0175] in, This indicates the constraint integration process. This represents the global scenario conditions obtained by pooling the world evolution kernel; This represents a behavioral-level semantic representation. Indicates navigation priors, This indicates the location code. When actually entering the diffusion planner, First, the data passes through a conditional embedding layer (3-layer MLP), then through a conditional coding block to complete context modeling, which is used as conditional memory in the diffusion decoder. The conditional information is not just concatenated once at the input, but continuously participates in trajectory feature updates during the denoising process.
[0176] During the training phase, the actual future trajectory After adding noise through forward diffusion, the first... Noisy trajectory at each diffusion time step:
[0177]
[0178] in, Indicates Gaussian noise. Indicates the noise scheduling coefficient. Noisy trajectory. First, the action embedding layer maps the data to trajectory features, and then the time steps are diffused. Simultaneously, after passing through the time-step embedding layer, both are input into the Transformer-based denoising module, and the internal relationships between future trajectory points are modeled through self-attention to obtain the intermediate trajectory representation. .
[0179] Then As a query, the condition center After conditional embedding, the data is used as the key and value to perform cross-attention reading:
[0180] in, Indicates the first The intermediate trajectory features obtained by noisy trajectories at each diffusion time step after trajectory embedding, time step embedding and self-attention modeling; Represents the cross-attention operator; As a query, As Key and Value; This represents the conditional modulation result read from the conditional center of the trajectory features.
[0181] This is the key to the diffusion architecture. It means that the conditional embedding is not simply added to the trajectory features, but is actively queried by each future trajectory point. In other words, the... During denoising, each future trajectory point will be determined based on its current state, from... It reads the scene boundary, navigation direction, behavioral intent, and motion trend information related to it.
[0182] In the diffusion denoising process, if the conditional information is injected into the trajectory reconstruction module in a static form, it can easily lead to insufficient modulation of the local trajectory generation by the higher-level intent, or inconsistent constraint strength under different driving states. Therefore, this invention introduces a diffusion state-gated modulation mechanism. The intent constraint center and the world evolution kernel are projected onto the trajectory generation space respectively, and dynamically combined by gating coefficients:
[0183]
[0184] in, Indicates the first The state gating coefficients at each diffusion time step Represents the state-semantic joint mapping matrix. This represents the activation function. The gating coefficients here determine the respective roles of the conditional center and the world evolution kernel in the current denoising stage.
[0185] Subsequently, this invention projects the intention constraint center and the world evolution kernel onto the trajectory generation space, and dynamically combines them using the aforementioned gating coefficients to obtain the constraint update amount:
[0186]
[0187] in, Indicates the first Constraint update amount under each diffusion time step; This represents the projection of the condition center reading result into the trajectory generation space; This represents the projection of the world evolution kernel into the trajectory generation space. Through this design, when the trajectory is still in a stage with high noise and unclear overall direction, the model relies more on navigation and behavioral semantics in the condition center; as the trajectory gradually approaches the feasible space, the model makes more use of scene boundaries and dynamic interaction information in the world evolution kernel to refine the local trajectory.
[0188] Ultimately, the diffusion planner at time step The trajectory update result is represented as follows:
[0189]
[0190] in, Indicates the first Trajectory characteristics modulated at each diffusion time step; This represents a Transformer-based denoising transform operator; This indicates the constraint update amount. The updated... Input the action head to predict the noise at the current diffusion time step. After obtaining... Then, the diffusion planner predicts the noise at the current diffusion time step using the action head. To ensure the generated trajectory simultaneously satisfies denoising and reconstruction accuracy, scene boundary safety, and behavioral intent... Figure 1 To ensure consistency, this invention constructs a joint optimization objective during the training phase. First, it uses the diffuse noise prediction loss as the basic reconstruction objective:
[0191]
[0192] in, This represents the actual noise injected during the forward diffusion process. The model represents the condition center. The noise estimate obtained from the prediction under constraints. By minimizing this loss, the model can learn how to progressively recover the target trajectory under multiple constraints.
[0193] Considering that relying solely on noise recovery loss is insufficient to guarantee trajectory safety, this invention introduces a safety distance constraint:
[0194]
[0195] in, Indicates the safe distance threshold. Indicates the generated trajectory The distance between the trajectory and the obstacle-occupied boundary represented by the physical boundary layer. When the trajectory approaches the obstacle boundary, danger zone, or impassable boundary, the penalty increases rapidly to help the model learn to maintain a reasonable distance during training.
[0196] Furthermore, to ensure that the generated trajectory aligns with the high-level behavioral intent in a distributional sense, this invention introduces an intent-constrained alignment loss:
[0197]
[0198] in, This represents the distribution of behaviors identified from the generated trajectories. This represents the distribution of target behaviors obtained from behavior-level semantic representation. This represents the KL divergence.
[0199] To further ensure that the generated trajectory remains consistent with the navigation target and local motion trends, this invention also introduces a planning consistency loss:
[0200]
[0201] in, This represents a trajectory motion feature mapping function, used to map the generated trajectory to a motion semantic space. This represents motion-level semantic representation. This indicates the deviation between the trajectory and the navigation prior. This term is used to constrain the overall direction, local trend, and navigation target of the trajectory to remain consistent.
[0202] During the training phase, this invention uses the spread noise prediction loss as the basic objective, combined with safety distance constraints, intention alignment constraints, and planning consistency constraints. The total loss function can be expressed as:
[0203]
[0204] in, , and These represent the weight coefficients of the safety distance penalty term, the intent constraint alignment term, and the intent guidance domain consistency term, respectively. Through the above design, this invention simultaneously achieves joint optimization of "trajectory denoising and recovery—safety boundary avoidance—behavioral intent constraint—guidance domain constraint maintenance" during the diffusion generation process.
[0205] During the inference phase, this invention preferably employs a deterministic reverse diffusion recovery method to progressively reconstruct the noisy trajectory. Unlike directly using the standard sampling formula, specifically, let the noisy trajectory at the current time step be... First, based on the noise estimation results, the intermediate noise-free trajectory representation is obtained:
[0206]
[0207] To ensure that the generated trajectory meets the actual motion constraints of the vehicle, this invention further introduces a dynamic projection operator to map the estimated trajectory to the vehicle's executable trajectory space:
[0208]
[0209] in, This represents the trajectory estimate after dynamic projection. This represents the dynamic projection operator, which is used to correct components in the trajectory that do not meet vehicle dynamic constraints based on kinematic information such as the current vehicle speed, acceleration, and steering state.
[0210] The trajectory state of the previous time step is calculated according to the inverse diffusion update rule:
[0211]
[0212] After multiple rounds of reverse diffusion recovery, the final potential trajectory representation is obtained. Considering that the trajectory may still exhibit local high-frequency jitter, curvature discontinuity, or short-term control abrupt changes after sampling, this invention further utilizes an action head... The final trajectory generation is completed using the dynamic smoothing calibration operator.
[0213]
[0214] in, This represents the dynamic smoothing calibration operator. After the above three steps, the system outputs the final continuously planned trajectory. This trajectory not only aligns with high-level driving semantics and navigation objectives, but also satisfies scene boundary constraints, obstacle avoidance requirements, and vehicle dynamics constraints, thus achieving a complete closed loop from hierarchical world cognition to intent-constrained generation alignment, and then to continuous action reconstruction.
[0215] Example:
[0216] This invention verifies the disclosed method through simulation. A DWDrive model was built based on the PyTorch framework, and training and testing were conducted using joint training data for question answering and trajectory, extended from the nuScenes dataset. The constructed system mainly includes a multimodal input encoding module, a hierarchical world cognition construction module, a hierarchical perturbation-suppressed semantic reasoning module, and an intent-constrained diffusion action reconstruction module. Specifically, the multimodal input encoding module receives visual observations, text commands, and navigation targets; the hierarchical world cognition construction module forms the physical boundary layer and the behavioral intent layer; the hierarchical perturbation-suppressed semantic reasoning module extracts behavioral-level semantic representations and motion-level semantic representations and condenses the world evolution kernel; and the intent-constrained diffusion action reconstruction module generates future continuous trajectories under the joint constraints of semantic and navigation conditions. The overall training adopts a three-stage setup: the first stage performs navigation alignment pre-training for one epoch, with a total batch size of 64 and a learning rate of... The second phase involves end-to-end driving fine-tuning, training for 3 epochs with a batch size of 16 and a learning rate of [missing information]. The third stage involves generative trajectory optimization, trained for 300 epochs with a batch size of 128 and a learning rate of [missing information]. The diffusion time step was set to 1000, and the inference step number was set to 8. The AdamW optimizer was used during training, with a cosine annealing learning rate strategy and a warmup ratio of 0.03. All training was completed on two NVIDIA A100 GPUs (80GB each).
[0217] This invention proposes an autonomous driving method based on hierarchical world cognition and intent constraint generation. By constructing hierarchical world cognition representations and introducing a hierarchical perturbation semantic reasoning mechanism, the model can gradually establish a stable mapping relationship from multimodal scene understanding to continuous trajectory generation during training. As training iterations proceed, the physical boundary layer's expression of road geometric boundaries, obstacle occupancy relationships, and the dynamics of key traffic participants gradually stabilizes, while the behavioral intent layer's constraint ability on navigation targets and short-term driving intentions gradually strengthens. This demonstrates that the method of this invention can effectively achieve joint optimization between semantic representation learning and diffusion-based trajectory reconstruction. Its advantages are: improving scene understanding through front-end hierarchical cognitive modeling; enhancing the interpretability of high-level decisions through intermediate semantic reasoning; and improving trajectory continuity, smoothness, and dynamic executability through back-end diffusion-based action reconstruction.
[0218] To verify the effectiveness of the method in this application, high-level behavioral decisions and corresponding interpretive texts were evaluated using metrics such as accuracy on a behavioral decision-making task. Experimental results show that the method of this invention achieves superior performance on this task, with a meta-action decision accuracy rate of 92%. These results demonstrate that this invention, through multimodal unified coding and hierarchical cognitive organization, can effectively extract high-level semantic information related to driving decisions and output behavioral-level semantic results with strong consistency and interpretability.
[0219] In motion planning tasks, this invention uses trajectory L2 error and collision rate in the future 1-second, 2-second, and 3-second time domains as core evaluation indicators. Experimental results show that the proposed method achieves superior planning performance on the nuScenes validation set: the L2 errors for 1 second, 2 seconds, and 3 seconds are 0.19m, 0.27m, 0.35m, and 0.27m, respectively, with corresponding collision rates of 0.01%, 0.10%, 0.32%, and 0.14%. These results demonstrate that this invention, by guiding the generation of diffusion trajectories through a world evolution kernel and navigation conditions, can effectively reduce the accumulation of long-term temporal errors and generate smoother, safer continuous trajectories while maintaining high-level semantic consistency.
[0220] Table 1: Comparison of Decision-Making and Programming Performance
[0221]
[0222] In summary, the training setup, quantitative results, ablation experiments, real-time analysis, and generalization tests presented in this application fully demonstrate the feasibility of the proposed autonomous driving method based on hierarchical world cognition and intent constraints. This method can achieve unified modeling from high-level semantic decision-making to low-level continuous trajectory generation under joint input conditions of visual observation, language commands, and navigation targets. It also achieves good overall performance in terms of trajectory accuracy, collision control, planning smoothness, interpretability, and adaptability to complex scenarios, showcasing promising engineering application prospects.
[0223] Although the present invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the disclosed embodiments can be combined with each other in any manner. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. An autonomous driving method based on hierarchical world cognition and intention constraint generation, characterized in that, Includes the following steps: S1: Constructing a hierarchical world cognition: First, the surround-view visual observation, vehicle status, historical trajectory, navigation target and text driving instructions are uniformly encoded to obtain a multimodal lexical representation; Subsequently, these morphemes were organized into two cognitive layers: the physical boundary cognitive layer and the behavioral intention cognitive layer; S2: Perform hierarchical perturbation reasoning and generate world evolution kernel: After obtaining hierarchical world cognition, the information propagation between the two cognitive layers is controlled by hierarchical perturbation mask. The physical boundary cognitive layer mainly maintains a stable expression of the spatial boundary and obstacle state, while the behavioral intention cognitive layer reads the boundary information related to the current driving decision when needed. After controlled reasoning, the world evolution kernel is obtained. The specific steps are as follows: S2.1: Design of a hierarchical scrambling mask mechanism: Let the i-th and j-th words belong to a certain cognitive layer, then the hierarchical scrambling mask matrix is defined as follows: in, Indicates the cognitive level to which a word belongs. This indicates the cross-level suppression strength, which is jointly determined by the current word state, the vehicle state, and the behavioral intent cognitive layer features: in, and They represent the first The word element and the first The current hidden representation of each word element. Indicates the vehicle's motion status. Represents the suppression intensity mapping matrix; The restricted attention reasoning form after introducing hierarchical scrambling masks is as follows: in, These represent the query, key, and value matrices obtained by linear mapping from the input hidden states in the current inference layer, respectively. Represents the hierarchical scrambling mask matrix; This represents the hidden representation after hierarchical perturbation constraints; S2.2: Hierarchical Reasoning Expansion and World Evolution Kernel Extraction: In this reasoning process, the physical boundary cognition layer mainly completes the aggregation of spatial boundaries and traffic participant states, while the behavioral intention cognition layer, while maintaining the semantic dominance of the task, is controlled to read the key boundary summary in the physical boundary cognition layer, calculates the reasoning result of the physical boundary layer, and is controlled to read the physical boundary summary. After completing hierarchical reasoning, behavioral-level semantic representations and motion-level semantic representations are generated: in, This indicates behavioral-level semantic output; Indicates motion-level semantic output; The aggregated features representing the cognitive layer of behavioral intent after hierarchical perturbation reasoning; For behavior semantic mapping header; For motion trend mapping head; To condense the world evolution kernel from the last few hidden states of the main reasoning framework, first calculate the inter-layer aggregation weights: in, Indicates the main branch of reasoning. The hidden state of the layer; Indicates the total number of layers in the main reasoning structure; Indicates the importance scoring function of the layer; Indicates the first The aggregation weight of the hidden layer states in the condensation of the world evolution nucleus; Then calculate the fusion gating between physical boundary information and behavioral intent information: in, Indicates the first Aggregated representation of hidden states in the physical boundary cognition layer; Indicates the first Aggregated representation of the hidden states of the cognitive layer of behavioral intentions; Represents the fusion gated mapping matrix; Indicates the first The fusion gating coefficient between layer physical boundary information and behavioral intent information; Ultimately, the world evolutionary nucleus is defined as: in, Represents the world's evolutionary nucleus. Indicates the inter-layer aggregate weight. Indicates hierarchical fusion gating, This represents element-wise multiplication; S3: Execution Intent Constraint Diffusion Planning: Combining the world evolution kernel, navigation priors, behavioral semantic representation, and motion semantic representation into a condition center. The conditional center is then injected into the diffusion planner: the noisy trajectory is first encoded as trajectory features, the conditional center is read by the trajectory features through cross attention, and then through state gating. Control condition injection strength, forming constraint update quantity During training, safety distance constraints, intent alignment constraints, and planning consistency constraints are introduced; during inference, dynamic projection and smooth calibration are combined to finally output a continuous planned trajectory that satisfies scene boundaries, driving intent, and vehicle motion constraints.
2. The autonomous driving method based on hierarchical world cognition and intent constraint generation according to claim 1, characterized in that, The specific steps of S1 are as follows: S1.1: Multimodal feature preprocessing: Input includes A 6-channel surround view image sequence of frames Navigation target point sequence and text driving instructions Among them, the surround view image sequence is used to characterize the road structure, traffic participant status and spatial boundary information of the environment around the vehicle; the navigation target point sequence is used to characterize the vehicle's target driving direction and path constraints; and the text driving instructions are used to provide high-level task semantic information. Text and navigation encoding: text instructions After BPE word segmentation, features are extracted by a text encoder initialized with Qwen2. Navigation target point It is then converted into navigation terms through a three-layer MLP adapter. ; Visual representation compression: for surround-view driving observation We employ the AnyRes strategy combined with bilinear interpolation for preprocessing to balance perception resolution and computational cost. S1.2: Initialization of Hierarchical World Cognition: After completing the multimodal feature preprocessing, two sets of learnable hierarchical world cognition vectors are defined, namely the physical boundary cognition vector. and behavioral intention cognitive vector To enable the two types of cognitive vectors to dynamically and adaptively adjust according to the current driving state and task objectives, hierarchical dynamic routing weights are further generated: in, Indicates the hierarchical dynamic routing weight; Represents the cognitive vector acting on the physical boundary. Routing weights This represents the action on the cognitive vector of behavioral intention. Routing weights; Represents the route weight mapping matrix. Indicates the vehicle's state characteristics. Indicates the characteristics of historical trajectory, Indicates navigation features, Representing text features; Then, a hierarchical world cognition vector with dynamic priors is constructed: in, Represents the physical boundary perception vector; Represents a cognitive vector of behavioral intention; Indicates the hierarchical position encoding; This represents a dynamic prior code term composed of the vehicle's state and historical trajectory. S1.3: State-biased cross-attention mechanism: A cross-attention mechanism is introduced, where state bias and target bias work together. The attention score is expressed as follows: Further, the initialized hierarchical world cognitive representation is obtained: in, , and These represent the mapping matrices for queries, keys, and values, respectively. This represents the surround-view visual feature lexical units obtained after visual encoder and lexical compression; This represents the feature dimension; the attention is used to make the hierarchical cognitive vectors derived from the features of the visual environment. Read information related to the current state and task objectives. Represents the motion state bias operator; This represents the goal-oriented bias operator.
3. The autonomous driving method based on hierarchical world cognition and intention constraint generation according to claim 2, characterized in that, The specific method for visual representation compression in S1.1 is as follows: Introducing an entropy-guided cognitive focusing operator Its core logic is to prioritize preserving regions in the feature map that exhibit significant gradient changes and dense semantic information. in, This represents the original visual word sequence output by the visual encoder; Represents the original visual word set or the corresponding feature map index; This represents the set of key visual terms retained after entropy-guided filtering; This is the threshold for the number of key visual words, used to control the number of key visual words retained after compression. This represents the function for calculating the entropy of visual lexical information; This indicates selecting the option with the highest information content. Each word element; It is the original input loop view image sequence; This is a scaling factor used to adjust the spatial scale after pooling; Indicates a visual encoder; These are the key visual feature terms after compression.
4. The autonomous driving method based on hierarchical world cognition and intent constraint generation according to claim 1, characterized in that, The calculation method for the inference result of the physical boundary layer in S1.2 is as follows: Let the query, key, and value corresponding to the physical boundary layer be respectively... The query, key, and value corresponding to the behavioral intent layer are respectively The inference result of the physical boundary layer is then expressed as: While maintaining its semantic dominance, the behavioral intent cognitive layer also reads the physical boundary summary in a controlled manner. in, Indicates the suppression coefficient. Represents the cross-layer projection operator; Aggregated features representing the cognitive layer of behavioral intention; This represents the injection coefficient mapping matrix.
5. The autonomous driving method based on hierarchical world cognition and intention constraint generation according to claim 1, characterized in that, The World Evolution Core Synthesis includes road boundaries, obstacle occupancy, key traffic participant dynamics, navigation guidance, and short-term behavioral intentions.
6. The autonomous driving method based on hierarchical world cognition and intent constraint generation according to claim 1, characterized in that, The specific method of S3 is as follows: First, construct the intention constraint center. ; During the training phase, the actual future trajectory After adding noise through forward diffusion, the first... Noisy trajectory at each diffusion time step: in, Indicates Gaussian noise. Represents the noise scheduling coefficient, noisy trajectory First, the action embedding layer maps the data to trajectory features, and then the time steps are diffused. Simultaneously, after passing through a time-step embedding layer, both are input into a Transformer-based denoising module, and the internal relationships between future trajectory points are modeled through self-attention to obtain the intermediate trajectory representation. ; Then As a query, the condition center After conditional embedding, the data is used as the key and value to perform cross-attention reading: in, Indicates the first The intermediate trajectory features obtained by noisy trajectories at each diffusion time step after trajectory embedding, time step embedding and self-attention modeling; Represents the cross-attention operator; As a query, As Key and Value; This represents the conditional modulation result read from the conditional center of the trajectory features; No. During denoising, each future trajectory point will be determined based on its current state, from... It reads relevant scene boundaries, navigation directions, behavioral intentions, and planning consistency information. The following introduces a diffusion-state gating modulation mechanism, which projects the intention constraint center and the world evolution kernel onto the trajectory generation space, and dynamically combines them using gating coefficients: in, Indicates the first The state gating coefficients at each diffusion time step Represents the state-semantic joint mapping matrix. This represents the activation function, where the gating coefficients determine the respective roles of the condition center and the world evolution kernel in the current denoising stage. Projecting the intention constraint center and the world evolution kernel onto the trajectory generation space, and then dynamically combining them using the aforementioned gating coefficients, yields the constraint update quantity: in, Indicates the first Constraint update amount under each diffusion time step; This represents the projection of the condition center reading result into the trajectory generation space; Represents the projection of the world evolution nucleus into the trajectory generation space; Ultimately, the diffusion planner at time step The trajectory update result is represented as follows: in, Indicates the first Trajectory characteristics modulated at each diffusion time step; This represents a Transformer-based denoising transform operator; Indicates the constraint update amount; the updated Input the action head to predict the noise at the current diffusion time step, and obtain Then, the diffusion planner predicts the noise at the current diffusion time step using the action head; During the training phase, a joint optimization objective is constructed. First, the diffuse noise prediction loss is used as the basic reconstruction objective: in, This represents the actual noise injected during the forward diffusion process. The model represents the condition center. The noise estimate obtained from the prediction under constraints; During the training phase, the spread noise prediction loss is used as the basic objective, combined with safety distance constraints, intention alignment constraints, and planning consistency constraints. The total loss function is expressed as: in, , and These represent the weight coefficients of the safety distance constraint, intention alignment constraint, and planning consistency constraint, respectively. During the inference phase, a deterministic inverse diffusion recovery method is used to gradually reconstruct the noisy trajectory. Let the noisy trajectory at the current time step be... First, based on the noise estimation results, the intermediate noise-free trajectory representation is obtained: The trajectory state of the previous time step is calculated according to the inverse diffusion update rule: After multiple rounds of reverse diffusion recovery, the final trajectory potential representation is obtained.
7. The autonomous driving method based on hierarchical world cognition and intent constraint generation according to claim 6, characterized in that, The condition center It consists of a world evolution kernel, navigation priors, behavior-level semantic representation, and motion-level semantic representation. in, This indicates the constraint integration process; This represents the global scenario conditions obtained by pooling the world evolution kernel; This represents a behavioral-level semantic representation. Indicates navigation priors, This represents the position code, which is actually used when entering the diffusion planner. First, it goes through a conditional embedding layer, and then through a conditional coding block to complete context modeling. This is used as conditional memory in the diffusion decoder. The conditional information is not just concatenated once at the input, but continuously participates in trajectory feature updates during the denoising process.
8. The autonomous driving method based on hierarchical world cognition and intent constraint generation according to claim 6, characterized in that, A safety distance constraint is introduced to ensure the safety of the trajectory. The specific method is as follows: in, Indicates the safe distance threshold. Indicates the generated trajectory The distance between the boundary occupied by obstacles as represented by the physical boundary layer; An intent-constrained alignment loss is introduced to ensure that the generated trajectory aligns with the high-level behavioral intent in a distributional sense. The specific method is as follows: in, This represents the distribution of behaviors identified from the generated trajectories. This represents the distribution of target behaviors obtained from behavior-level semantic representation; Indicates KL divergence; A planning consistency loss is introduced to ensure that the generated trajectory remains consistent with the navigation target and local motion trends. The specific method is as follows: in, This represents the trajectory motion feature mapping function, used to map the generated trajectory to the motion semantic space; This represents motion-level semantic representation. This indicates the deviation between the trajectory and the navigation prior.
9. The autonomous driving method based on hierarchical world cognition and intention constraint generation according to claim 6, characterized in that, A dynamic projection operator is introduced to ensure that the generated trajectory satisfies the actual motion constraints of the vehicle. Specifically, the estimated trajectory is mapped to the vehicle's executable trajectory space. in, This represents the trajectory estimate after dynamic projection; This represents the dynamic projection operator, used to correct components in the trajectory that do not satisfy vehicle dynamic constraints based on the current vehicle kinematics information; This represents the noise-free trajectory estimated based on the current noisy trajectory and the predicted noise. Through the action head The final trajectory generation is completed using the dynamic smoothing calibration operator. in, This represents the dynamic smoothing calibration operator.
Citation Information
Patent Citations
Trajectory prediction method and device and vehicle
CN121062757A
Automatic driving track planning method, system and equipment based on generative model and medium
CN121232823A