A visual language action and world model collaborative unified automatic driving method

CN122808775APending Publication Date: 2026-09-25SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610885892.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

语言描述能够与大语言模型自然结合,但其对三维位置、速度、运动轨迹和智能体间动力学关系的物理约束较弱,难以保证推理结果与真实交通场景严格一致

Benefits of technology

[0019]本申请提供的种视觉语言动作和世界模型协同统一的自动驾驶方法,通过离散智能体Token表征交通参与者,并利用四维几何与二维视觉特征进行多模态物理绑定,构建出紧凑且结构化的智能体世界状态;该状态直接输入视觉—语言—动作策略模型,驱动模型围绕关键交通对象进行显式推理与轨迹规划,有效克服了现有方法像素级表征冗余大、规划意图隐式、交互关系建模不足及决策可解释性弱的缺陷,显著提升了复杂动态场景下的计算效率、规划安全性与决策透明度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122808775A_ABST
    Figure CN122808775A_ABST
Patent Text Reader

Abstract

The application provides a visual language action and world model collaborative unified automatic driving method, comprising: receiving multi-view camera images, vehicle state, navigation target and language prompt, and constructing driving context; based on the driving context, initializing discrete agent Token, and using four-dimensional geometric features and two-dimensional visual features to perform multi-modal physical binding on the discrete agent Token to generate agent world state; inputting the agent world state into a visual-language-action strategy model to perform reasoning, action prediction and future trajectory planning. The application effectively overcomes the defects of the existing method, such as large redundancy of pixel-level representation, implicit planning intention, insufficient modeling of interaction relationship and weak decision-making interpretability, and significantly improves the calculation efficiency, planning safety and decision-making transparency in complex dynamic scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and more specifically, to an autonomous driving method that unifies visual language actions and a world model. Background Technology

[0002] In recent years, autonomous driving technology has been gradually evolving from a traditional modular architecture to an end-to-end intelligent driving architecture. Traditional autonomous driving systems typically divide environmental perception, object detection, trajectory prediction, behavior decision-making, and motion planning into multiple independent modules. Each module completes a specific task and then they are cascaded through manually designed interfaces. This approach has a certain degree of interpretability and engineering controllability, but in complex open road environments, it is susceptible to problems such as upstream perception error propagation, inconsistent module objectives, insufficient scene interaction modeling, and limited adaptability to long-tail traffic scenarios.

[0003] With the development of multimodal large models, Vision-Language-Action Models (VLAs) are increasingly being applied to the field of autonomous driving. These models can uniformly process multi-view images, vehicle states, navigation commands, and natural language cues, and generate driving actions or future trajectories based on multimodal understanding results. Compared to traditional end-to-end models, VLAs possess stronger semantic understanding, command following, and reasoning expression capabilities, helping to improve the interpretable decision-making capabilities of autonomous driving systems in complex traffic scenarios. Meanwhile, World Models have also been introduced into autonomous driving tasks. The core idea of ​​World Models is to predict the evolution of future scenarios based on current and historical observations, enabling autonomous driving systems to estimate future states before executing actions, thus avoiding passive reactions based solely on current observations. In recent years, some methods have further combined VLAs with World Models, forming World VLA-based methods. These methods allow the model to incorporate predictions or imaginations of future driving scenarios during multimodal reasoning, generating more reasonable driving behaviors accordingly.

[0004] However, most existing World VLA methods primarily represent the future world state using pixel-level future images, video frames, or dense visual tokens. For example, some methods predict future driving images through autoregressive or diffusion generation, then input the generated results into a vision-language-action model for reasoning and planning. While these methods can generate visually realistic future scenes, the generated images contain a large number of appearance details that are weakly related to driving decisions, such as texture, lighting, background, color, and image continuity. This information consumes a large number of visual tokens and computational resources, but does not necessarily directly serve trajectory planning. For autonomous driving planning, the more critical information is usually not a complete pixel-level future image, but rather the semantic categories, 3D geometric positions, motion trends, future trajectories, and interaction relationships between the vehicle and surrounding traffic participants. For example, in scenarios such as following other vehicles, changing lanes, merging, avoiding pedestrians, and navigating intersections, the vehicle's decisions mainly depend on the dynamic changes and potential conflict relationships of surrounding vehicles and pedestrians. However, in pixel-level future frame representations, the aforementioned planning-related factors are often implicitly and sparsely encoded in a large number of visual tokens. The model needs to re-extract the structured relationships related to decision-making from redundant image information, resulting in low inference efficiency, weak task alignment, and potentially reduced planning reliability in complex scenarios.

[0005] Furthermore, while existing vision-language-action autonomous driving models can output driving actions based on multimodal inputs and outputs, their internal reasoning processes often lack explicit spatiotemporal modeling and agent interaction modeling capabilities. Some models tend to rely on extrapolation of the vehicle's historical state, linguistic priors, or statistical biases in the training data, failing to fully utilize the physical state and dynamic interaction information of surrounding traffic participants. When multiple vehicles, pedestrians, or non-motorized vehicles interact in a scene, the lack of explicit world state representation at the agent level can easily lead to the model's inability to accurately identify key risk objects and explain the basis for its planning decisions. Existing technologies have also attempted to represent the autonomous driving world state using methods such as bird's-eye view, occupancy grid, global latent variables, or linguistic descriptions. Bird's-eye view and occupancy grid can preserve some spatial geometry, but they usually still contain a large amount of low-level spatial details and tend to mix road background, static structure, and dynamic traffic participants in the same representation, which is not conducive to high-level policy models directly performing object-oriented causal reasoning. Although global latent variable representation is relatively compact, it usually compresses multiple traffic participants and their interaction relationships into the overall features, lacking interpretable attribution for specific agents. Language descriptions can be naturally integrated with large language models, but they have weak physical constraints on three-dimensional position, velocity, motion trajectory, and dynamic relationships between agents, making it difficult to guarantee that the inference results are strictly consistent with real traffic scenarios.

[0006] Therefore, existing technologies have at least the following problems: First, the world state representation based on pixel-level future frames contains a large amount of visual redundancy unrelated to planning, resulting in high inference computation overhead; Second, the three-dimensional geometry, motion state, and interaction relationships of traffic participants lack explicit structured expression in dense visual representations; Third, global latent variables or linguistic world states are difficult to simultaneously achieve compactness, physical meaning, and agent-level interpretability; Fourth, existing vision-language-action models may still not fully model the impact of surrounding agents on the vehicle's future trajectory in complex traffic scenarios, thus affecting planning safety and stability.

[0007] A search revealed a Chinese patent application (application number 202411956161.8) that discloses an autonomous driving method based on an end-to-end and multimodal large model. This method encodes multi-source sensor images, vehicle history, navigation, and text information into a global token, inputs it into a large model, and directly outputs the planned trajectory and control commands through parallel decoding, supplemented by policy text interpretation. However, it fails to address the technical bottlenecks of existing end-to-end architectures, such as large redundancy in visual representations, implicit modeling of key traffic participant interactions, lack of explicit causal links in planning decisions, and insufficient safety and interpretability in complex scenarios.

[0008] In summary, the field of autonomous driving urgently needs a world state modeling approach oriented towards planning tasks, enabling the model to explicitly characterize the semantics, three-dimensional geometry, motion trends, and interaction relationships of surrounding traffic participants with low computational complexity, and effectively integrate this structured world state into the reasoning and trajectory planning process of the vision-language-action model, thereby improving the safety, interpretability, and planning performance of end-to-end autonomous driving systems in complex traffic environments. Summary of the Invention

[0009] In view of the deficiencies in the prior art, the purpose of this application is to provide an autonomous driving method that coordinates and unifies visual language actions and world models.

[0010] The first aspect of this application provides an autonomous driving method that coordinates and unifies visual language actions and a world model, comprising: It receives multi-view camera images, vehicle status, navigation targets, and voice prompts to construct driving context information; Initialize the discrete agent token based on the driving context information and generate a query vector; Extract the four-dimensional geometric features and two-dimensional visual features of the multi-view camera images, and perform multimodal physical binding on the query vector; Generate the agent's world state based on the updated agent representation. The agent's world state is input into the visual language action policy model to perform reasoning, action prediction, and future trajectory planning.

[0011] Optionally, the construction of driving context information includes: Encode the multi-view camera images into visual features; The vehicle status, the navigation target, and the language prompt are respectively mapped to a status vector, a target vector, and a text vector; The visual features are then concatenated with each of the vectors to form an initial input sequence, which serves as driving context information.

[0012] Optionally, the step of initializing the discrete agent token and generating a query vector based on the driving context information includes: A fixed number of placeholders are inserted into the driving context as discrete agent tokens, and each discrete agent token corresponds to an addressable agent slot. Each discrete agent token is assigned a learnable embedding vector, which is then transformed into an initial query vector using a linear mapping or a multilayer perceptron. The initial query vector is used to actively retrieve associated traffic participant information to obtain an updated query vector; The intelligent agent slot is dynamically bound to surrounding traffic participants after the retrieval is completed. If no valid traffic participant information is retrieved, the corresponding slot is marked as an empty object.

[0013] Optionally, the step of extracting the four-dimensional geometric features and two-dimensional visual features of the multi-view camera images and performing multimodal physical binding on the query vector includes: Based on the multi-view camera images, four-dimensional geometric features and two-dimensional visual features are obtained; The query vector for each agent is input into the geometric branch and the visual branch respectively; In the geometric branch, the four-dimensional geometric features are mapped to a feature space that is the same as or compatible with the agent query. Spatial location, three-dimensional scale, motion trend and temporal consistency information related to the current agent token are extracted from the four-dimensional geometric features through a cross-attention mechanism to generate a geometric response. In the visual branch, multi-view image features are mapped to the image routing space, and appearance, category, local context and semantic information are extracted from the two-dimensional visual features through a cross-attention mechanism to generate a visual response; Optionally, adaptive gating fusion is also employed in the multimodal physical binding process, specifically as follows: Predict the geometric branch weights and visual branch weights based on the agent's query vector; The geometric response and the visual response are weighted and fused according to the geometric branch weight and the visual branch weight to generate multimodal fusion features; The multimodal fusion features are aggregated with the agent query vector to update the agent query vector.

[0014] Optionally, the geometric branch weights and visual branch weights are dynamically adjusted according to the relative distance between the target and the vehicle.

[0015] Optionally, generating the agent world state based on the bound-updated agent representation includes: The updated agent query vector is projected into the hidden space of the multimodal large model backbone and backfilled into the corresponding discrete agent token position in the input sequence to obtain the complete context sequence; The multimodal large model backbone performs global self-attention computation on the complete context sequence, enabling each discrete agent token to interact across slots with the driving context and other tokens, and outputs global contextualized hidden features. The output vectors located at the token positions of each discrete agent are extracted from the global contextualized hidden features and aggregated into an agent state set in index order.

[0016] Optionally, training may include a physical binding training phase, specifically: Using a dataset labeled with 3D bounding boxes, categories, and future trajectories, prediction results are output through a multimodal physical binding process, including predicted categories, predicted bounding boxes, and predicted trajectories. The prediction results are optimally associated with real traffic participants using a bipartite graph matching algorithm. Based on the association results, the classification loss, bounding box regression loss, and trajectory prediction loss are calculated, and the parameters involved in the multimodal physical binding process are optimized based on the losses.

[0017] Optionally, the training process may also include a reasoning-guided supervised fine-tuning phase, specifically: The bound discrete agent token and the driving context information are input into the multimodal large model; Using expert driving trajectories, high-level action commands, and agent interaction reasoning text as supervision labels, the autoregressive generation loss is calculated. Based on the autoregressive generation loss, the backbone network parameters of the multimodal large model and the policy generation head parameters of the visual language action policy model are optimized, enabling the model to learn explicit driving reasoning based on the agent's world state.

[0018] Optionally, the training process may also include an agent-guided reinforcement learning phase, specifically: Sample multiple candidate planning results for the same driving scenario; The reward score for each candidate planning result is calculated using a comprehensive reward function that includes interactive safety rewards. Based on the difference in reward scores among the candidate planning results, the strategy generation head parameters and planning decision layer parameters of the visual language action strategy model are updated. The interactive safety reward is calculated based on the relative distance between the vehicle's future trajectory and the predicted trajectories of surrounding traffic participants, as well as the collision risk.

[0019] This application provides an autonomous driving method that unifies visual language action and world model. It represents traffic participants through discrete intelligent agent tokens and uses four-dimensional geometry and two-dimensional visual features for multimodal physical binding to construct a compact and structured intelligent agent world state. This state is directly input into the visual-language-action policy model, driving the model to perform explicit reasoning and trajectory planning around key traffic objects. This method effectively overcomes the shortcomings of existing methods, such as large pixel-level representation redundancy, implicit planning intentions, insufficient modeling of interaction relationships, and weak decision interpretability. It significantly improves the computational efficiency, planning security, and decision transparency in complex dynamic scenarios.

[0020] Other technical effects resulting from the additional features will be further illustrated in the corresponding embodiments. Attached Figure Description

[0021] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A flowchart illustrating an autonomous driving method for coordinating and unifying visual language actions and a world model according to an exemplary embodiment; Figure 2 This is a framework diagram illustrating an autonomous driving method that coordinates and unifies visual language actions and a world model according to an exemplary embodiment. Detailed Implementation

[0022] The present application will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application, and these all fall within the protection scope of the present application. Parts not described in detail in the following embodiments can be implemented using existing technology.

[0023] Existing world vision-language-action models rely on pixel-level future frame generation or a large number of dense visual tokens to represent the future world state, resulting in significant visual redundancy, unclear expression of planning-related information, insufficient modeling of agent interactions, and weak interpretability of end-to-end driving decisions. This makes it difficult to achieve structured world state reasoning for agents and safe planning decisions in complex traffic scenarios. To address these issues, this application provides an autonomous driving method that unifies vision, language, actions, and world models to solve the aforementioned problems.

[0024] Reference Figure 1 and Figure 2 As shown in one embodiment of this application, an autonomous driving method that coordinates and unifies visual language actions and a world model is provided, including the following steps: S100 receives multi-view camera images, vehicle status, navigation targets and voice prompts to construct driving context information; S200 initializes the discrete agent token based on driving context information and generates a query vector; S300 extracts four-dimensional geometric features and two-dimensional visual features from multi-view camera images and performs multimodal physical binding on the query vector; S400 generates the agent world state based on the updated agent representation. The S500 inputs the agent's world state into a visual-language-action-policy model to perform reasoning, action prediction, and future trajectory planning.

[0025] The above embodiments enable efficient, interpretable, and safe end-to-end driving decision-making and trajectory planning in complex traffic scenarios. It breaks through the limitations of traditional methods that rely on pixel-level future frames or dense visual tokens to represent environmental states. It innovatively abstracts the driving scenario as a set of discrete intelligent agent tokens and dynamically physically binds them using four-dimensional geometry and two-dimensional visual features, constructing a compact, structured intelligent agent world state. This mechanism allows the vision-language-action policy model to directly perform explicit reasoning and interaction modeling around key traffic participants who truly influence driving decisions. While significantly improving the planning performance and safety in complex scenarios, it effectively avoids the computational bottleneck caused by redundant visual information. The overall solution has a simple architecture and clear task targeting, and can be widely applied to tasks such as multimodal scene understanding, intelligent agent interaction reasoning, trajectory prediction, end-to-end decision planning, and closed-loop control in the field of autonomous driving.

[0026] To obtain accurate driving context, in some specific embodiments of this application, in step S100, receiving multi-view camera images, vehicle status, navigation target, and voice prompts to construct driving context can employ the following steps: S101 encodes multi-view camera images into visual features; Specifically, multi-view camera images are acquired in real time by multiple onboard cameras (front / rear / left / right, etc.).

[0027] S102, map the vehicle status, navigation target and voice prompts into a status vector, a target vector and a text vector respectively; Specifically, the vehicle status, navigation target, and voice prompts are acquired in real time through onboard sensors (such as CAN bus / IMU / GNSS, etc.), onboard navigation system or human-machine interaction terminal, and onboard voice interaction module or preset driving commands, respectively, and together serve as the raw input data for constructing the driving context.

[0028] S103 concatenates the visual features of S101 with the three vectors of S102 to form the initial input sequence.

[0029] The above embodiments construct a unified driving context base through encoding and sequence concatenation, providing an aligned data pool for subsequent multimodal feature retrieval and global interaction modeling.

[0030] To reduce redundant reasoning, in some specific embodiments of this application, S200, initializing the discrete agent token based on driving context information and generating a query vector can be achieved by the following steps: S201 introduces a fixed number of special agent tokens (denoted as Token) into the input sequence of a multimodal large model.<agent_1> to<agent_N> (where N is the maximum number of modelsable units). Specifically, each agent token corresponds to an addressable agent slot, which is used to dynamically bind surrounding traffic participants such as vehicles, pedestrians, and cyclists in the scene; when a token is not matched with an actual traffic participant, it can be marked as an empty object state.

[0031] S202, assign a learnable embedding vector to each token and convert it into an initial agent query via a linear mapping or multilayer perceptron.

[0032] During the initialization phase, the query is used as an unbound potential object slot for subsequent active retrieval. When no actual traffic participant is matched, it is marked as an empty object state.

[0033] The agent query does not correspond to a specific target during the initialization phase. Instead, it serves as an unbound potential object slot, used to proactively retrieve relevant traffic participant information from geometric and visual features. This approach eliminates the need for an independent target query set; instead, it directly utilizes the unique token positions within the multimodal large model to form the agent's world state.

[0034] The above embodiments directly pre-set fixed agent slots in the input sequence, eliminating the need for independent target query lists and significantly reducing memory and computational overhead. These slots serve as probes to be activated, dynamically binding features only when capturing real traffic participants, avoiding redundant reasoning for blank scenarios.

[0035] In order to give the agent token a clear physical meaning, in some specific embodiments of this application, in S300, the four-dimensional geometric features and two-dimensional visual features of the multi-view camera images are extracted, and the query vector is physically bound in a multimodal manner. This can be done through the following steps: S301 acquires four-dimensional geometric features and two-dimensional visual features based on multi-view camera images.

[0036] Specifically, the four-dimensional geometric features come from VGGT-based features processed offline, while the two-dimensional visual features come from the features output by the visual encoder in the multimodal large model.

[0037] In S302, each agent's query is input into the geometric branch and interacts with the four-dimensional geometric features. The geometric branch first maps the four-dimensional geometric features to a feature space that is the same as or compatible with the agent's query, and then extracts spatial location, three-dimensional scale, motion trend, and temporal consistency information related to the current agent's token from the four-dimensional geometric features through a cross-attention mechanism.

[0038] Specifically, the above mapping can be achieved by using an MLP to transform the feature dimension to the same size as the agent query.

[0039] S303, each agent's query is input into the visual branch and interacts with the two-dimensional visual features. The visual branch first maps the multi-view image features to the image routing space (mapped to two-dimensional visual features), and then extracts appearance, category, local context, and semantic information related to the current agent from the two-dimensional visual features through a cross-attention mechanism.

[0040] S304 After extracting information from the geometric and visual branches, the corresponding geometric response, visual response, and original agent query are fused to obtain the updated agent query.

[0041] Specifically, this fusion process can be achieved through weighted fusion, attention fusion, or multilayer perceptrons. By stacking multiple multimodal routing binding layers, the agent token can gradually aggregate information related to its corresponding traffic participants from multi-view images and four-dimensional geometric features, ultimately forming a bound agent token representation.

[0042] In the above embodiments, the geometric branch is mainly used to enhance the agent token's perception of physical spatial structure and dynamic evolution. The visual branch is mainly used to compensate for the lack of information in geometric features regarding target category, visible appearance, and scene semantics.

[0043] Considering that different intelligent agents have varying degrees of dependence on geometric and visual information in different scenarios, some specific embodiments of this application also incorporate geometric gating and visual gating mechanisms. Specifically, For each agent query, its routing weights in the geometric and visual branches are predicted separately, enabling the token to adaptively select more important modal information.

[0044] For example, visual semantics may be more important for distant targets; while for vehicles interacting at close range, three-dimensional geometry and motion state may be more important.

[0045] The gating mechanism in the above embodiments avoids simple splicing or forced competition between geometric features and visual features, enabling each agent token to dynamically fuse multimodal information according to its own state.

[0046] In order to provide a clear object-level interface for subsequent reasoning and trajectory planning, in some specific embodiments of this application, S400, generating the agent's world state based on the updated agent representation can adopt the following steps: S401, the agent query vector obtained by multimodal routing is projected into the hidden space of the multimodal large language model and inserted into the corresponding discrete agent identifier Token position in the input sequence to obtain the model input sequence containing agent tokens.

[0047] Specifically, the multimodal large model is a backbone network based on the Transformer architecture, specifically a Visual-Language-Action (VLA) or Visual-Language Large Language Model (MLLM). Its core components include a visual feature alignment layer, a multi-layer self-attention network, and a feedforward network (FFN). In this step, this model serves as the context encoding backbone, primarily responsible for receiving the fused multimodal input sequence. It achieves cross-modal feature interaction and semantic alignment between the discrete agent token and the driving context, navigation commands, and vehicle state through a global self-attention mechanism.

[0048] For example, pre-trained weights from open-source or commercial VLA / MLLM architectures can be directly loaded as the base (such as LLaVA series, Qwen-VL, DriveVLM, InternVL, etc.). This application only needs to reserve agent token slots in its standard input sequence and use its native context modeling capabilities to complete state solidification without changing the base network structure.

[0049] S402, based on the model input sequence, uses a multimodal large language model to perform contextual modeling of driving context, visual features, navigation information and multiple discrete intelligent agent tokens, so that each discrete intelligent agent token can integrate the semantic identity, three-dimensional geometry, motion state and interaction relationship of its corresponding traffic participant in a unified semantic space, and obtain the contextualized sequence hidden features. S403: Extract the hidden vectors corresponding to the token positions of each discrete agent from the sequence hidden features, and organize them according to the token index order to form a set of agent potential states, which serve as the agent world states for subsequent driving reasoning and trajectory planning.

[0050] In the above embodiments, the hidden states of all agent tokens collectively constitute the agent world state set of the current driving scenario. Compared to pixel-level future images, this set is a more compact, structured, and planning-oriented representation of the world state, providing a clear object-level interface for subsequent reasoning and trajectory planning.

[0051] The above embodiments can focus computational resources and inference attention on the key traffic participants and their dynamic interactions that truly affect autonomous driving planning, instead of redundantly modeling information that is weakly related to planning, such as texture, lighting, background, and image continuity. Compared with existing World VLA methods based on future image generation, global latent variables, or linguistic descriptions, the constructed agent world state simultaneously possesses compactness, physical meaning, interactive perception capabilities, and interpretability (specifically, the decision-making carrier is transformed from abstract latent variables into discrete agent tokens bound to explicit physical attributes; the output results can be accurately traced to the specific agent slot and bound features that triggered the decision, achieving measurable perception, auditable inference, and traceable basis). This reduces the model's reliance on extrapolation of the vehicle's historical state, linguistic priors, and statistical biases of the dataset, making trajectory planning more consistent with the spatial constraints, motion patterns, and interaction logic in real traffic scenarios.

[0052] In one specific embodiment of this application, in step S500, the agent's world state is input into the visual-language-action-policy model to perform reasoning, action prediction, and future trajectory planning, which can be achieved through the following steps: S501 concatenates the agent's world state with driving context information into a policy input sequence, which is then input into the vision-language-action policy model. S502 performs global context modeling through a policy model, explicitly encoding the dynamic interaction relationships between each agent and the vehicle and the environment; S503, based on the interactive context, outputs risk analysis text, high-level action instructions and future trajectory points in an autoregressive manner through the policy generation head to complete end-to-end driving decision-making.

[0053] The visual language action strategy model described in the above embodiment explicitly decouples pixel-level future frame generation from planning-oriented agent interaction inference. It no longer relies on a large number of dense visual tokens to completely reconstruct future images, but instead abstracts the autonomous driving scenario into a set of dynamically bound discrete agent tokens. Each agent token corresponds to the state of a surrounding traffic participant or empty object, representing its semantic category, 3D geometry, motion trend, and interaction context. This allows the model to reason around key objects that truly influence driving decisions, such as vehicles and pedestrians, effectively alleviating the problems of large visual redundancy, implicit planning-related information, and insufficient agent interaction modeling in existing World VLA methods.

[0054] This application is applicable to end-to-end trajectory planning, closed-loop driving control, and complex traffic interaction reasoning tasks for autonomous driving. To achieve a compact, physically meaningful, interactively perceptive, and interpretable visual-language-action model of the autonomous driving agent world, in one specific embodiment of this application, the overall training process can be divided into three stages: physically bound training, reasoning-oriented supervised fine-tuning, and agent-oriented reinforcement learning.

[0055] During the training phase of physical binding, explicit supervision of the agent token is performed using autonomous driving data labeled with 3D bounding boxes, object categories, and future trajectories. The specific process is as follows: First, the multimodal physical binding process is executed: the agent query vector is initialized based on the driving context, and each agent query vector is dynamically fused with four-dimensional geometric features and two-dimensional visual features through cross-attention retrieval of geometric and visual branches to obtain the agent representation after binding update; Subsequently, the class probability (or existence probability), 3D bounding box coordinates, and future trajectory sequence of each agent slot are output in parallel through the agent prediction head. Finally, the prediction results are optimally associated with real traffic participants using the Hungarian bipartite graph matching algorithm. The classification loss, 3D bounding box regression loss, and trajectory prediction loss are calculated, and backpropagation optimization is performed based on the joint loss.

[0056] The core objective of this phase is to enable discrete agent tokens to have verifiable physical anchors, providing a structured and traceable world state basis for subsequent reasoning and planning.

[0057] It should be noted that this stage mainly optimizes the relevant parameters of the multimodal binding links, such as: agent query vector projection layer weights, cross-attention Q / K / V routing matrices for geometric and visual branches, adaptive gating weight prediction network parameters, and multilayer perceptron mapping layer parameters of the agent prediction head.

[0058] In the reasoning-guided supervised fine-tuning phase, expert driving trajectories, high-level actions, and agent-related reasoning annotations are used to train the vision-language-action model to output structured driving responses. The specific process is as follows: The input to the vision-language-action model includes multi-view visual observations, vehicle state, navigation target, language cues (i.e., contextual information), and the bound agent token; the model output includes agent risk analysis, high-level driving actions, and future trajectory points. This stage uses backpropagation through autoregressive cross-entropy loss to primarily optimize the backbone network parameters of the multimodal large model and the policy generation head parameters of the vision-language-action policy model, enabling it to learn how to perform driving reasoning based on the bound agent world state, rather than directly fitting a trajectory from images or vehicle state.

[0059] In the agent-guided reinforcement learning phase, the supervised fine-tuned visual-language-action policy model is used as the initial policy, and training and optimization are performed in the following manner: First, for the same driving scenario input model, multiple sets of candidate planning results (including action sequences and future trajectories) are sampled and generated. Secondly, the candidate results of each group are input into the comprehensive reward function to calculate the quality score. The reward function simultaneously considers the standardization of output format, perceptual consistency, action correctness, trajectory smoothness, driving progress, drivable area constraints, expert imitation degree and interaction safety with surrounding intelligent agents. Among them, the interactive safety reward is quantitatively evaluated based on the minimum distance, collision probability, safety margin and collision time (TTC) between the vehicle's future trajectory and the predicted trajectories of surrounding vehicles, pedestrians and other intelligent agents, and an explicit negative constraint is imposed on high-risk planning. Subsequently, the policy gradient loss is calculated based on the difference in relative reward scores of different candidate results within the same scene, and the policy generation head and planning decision layer parameters of the vision-language-action policy model are updated. Ultimately, through a multi-round sampling-evaluation-update closed loop, the strategy model gradually internalizes the safety boundary and traffic game logic, outputting smoother, safer end-to-end driving decisions that conform to complex interaction patterns.

[0060] This optimization mainly involves optimizing the policy generation head parameters and planning decision layer parameters of the visual language action strategy model.

[0061] This application presents an end-to-end driving decision-making scheme based on agent world state representation, multimodal route binding, and agent-guided policy learning. First, a compact plan-guided world state is constructed using discrete agent tokens, providing a unified intermediate interface for perception, reasoning, and planning. Then, through a multimodal route binding mechanism, each agent token is adaptively fused with four-dimensional geometric features and two-dimensional visual features. The four-dimensional geometric features provide spatial location, temporal consistency, and motion information, while the two-dimensional visual features provide appearance semantics and scene context information, thereby forming a physically meaningful bound agent representation.

[0062] Furthermore, the above embodiments employ a three-stage progressive training strategy: after establishing the perception base through physical binding training, an explicit causal generation chain of "world state, interactive reasoning, and action trajectory" is established through reasoning-guided supervised fine-tuning; finally, based on agent-guided reinforcement learning, the strategy output is optimized using a comprehensive preference function that includes interactive safety rewards, enabling the planning decision to explicitly internalize the dynamic game constraints of surrounding traffic participants. Through the above collaborative mechanism, this scheme constructs a compact, physically verifiable, interactively perceptive, and decision-interpretable agent-world vision-language-action model, effectively overcoming the technical bottlenecks of existing World VLA methods such as pixel-level representation redundancy, implicit interaction relationships, insufficient alignment of planning tasks, and weak decision stability in complex scenarios, achieving efficient and safe end-to-end autonomous driving control.

[0063] The preferred features in the above embodiments can be used individually in any embodiment, or in any combination thereof, provided they do not conflict with each other. Furthermore, parts not described in detail in the embodiments can be implemented using existing technologies.

[0064] The following examples and comparative examples will be used to further illustrate this application in order to better understand the above-mentioned technical solutions. It should be understood that the following are only some examples and are not intended to limit this application.

[0065] End-to-end trajectory planning and closed-loop driving evaluation were performed on the nuScenes public dataset. Table 1 shows the performance comparison results of the embodiments of the present invention and representative methods in the current field under the ST-P3 and UniAD evaluation protocols. Experimental data show that, compared with the comparative schemes, the present application achieves significant optimization in trajectory prediction accuracy (L2 error) and planning safety (collision rate), verifying the effectiveness of the structured world state representation based on discrete agent tokens in complex traffic scenarios.

[0066] Table 1

[0067] As shown in Table 1, the end-to-end trajectory error (L2) and collision risk of the embodiments in this application are significantly lower than those of existing representative vision-language-action models. This performance improvement is mainly due to: This application abandons redundant pixel-level future frame generation and adopts compact discrete agent token explicit modeling of key traffic participants, so that the model planning attention is accurately focused on the core objects that affect decision-making, effectively reducing trajectory extrapolation error. Multimodal adaptive gating binding and global interaction context modeling enhance geometric perception accuracy in close-range game scenarios. Combined with interaction safety reward constraints in the reinforcement learning phase, the collision rate is reduced by orders of magnitude. The experimental results effectively demonstrate the technical advantages of this application in reducing computational redundancy, improving planning interpretability, and enhancing driving safety in complex scenarios, and have clear engineering application value.

[0068] The foregoing has described some specific embodiments of this application. It should be understood that this application is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the substantive content of this application. The above-described preferred features can be used in any combination without conflict.

Claims

1. An autonomous driving method that coordinates and unifies visual language actions and a world model, characterized in that, include: It receives multi-view camera images, vehicle status, navigation targets, and voice prompts to construct driving context information; Initialize the discrete agent token based on the driving context information and generate a query vector; Extract the four-dimensional geometric features and two-dimensional visual features of the multi-view camera images, and perform multimodal physical binding on the query vector; Generate the agent's world state based on the updated agent representation. The agent's world state is input into the visual language action policy model to perform reasoning, action prediction, and future trajectory planning.

2. The method according to claim 1, characterized in that, The construction of driving context information includes: Encode the multi-view camera images into visual features; The vehicle status, the navigation target, and the language prompt are respectively mapped to a status vector, a target vector, and a text vector; The visual features are then concatenated with each of the vectors to form an initial input sequence, which serves as driving context information.

3. The method according to claim 1, characterized in that, The process of initializing the discrete agent token based on the driving context information and generating a query vector includes: A fixed number of placeholders are inserted into the driving context as discrete agent tokens, and each discrete agent token corresponds to an addressable agent slot. Each discrete agent token is assigned a learnable embedding vector, which is then transformed into an initial query vector using a linear mapping or a multilayer perceptron. The initial query vector is used to actively retrieve associated traffic participant information to obtain an updated query vector; The intelligent agent slot is dynamically bound to surrounding traffic participants after the retrieval is completed. If no valid traffic participant information is retrieved, the corresponding slot is marked as an empty object.

4. The method according to claim 1, characterized in that, The step of extracting the four-dimensional geometric features and two-dimensional visual features of the multi-view camera images and performing multimodal physical binding on the query vector includes: Based on the multi-view camera images, four-dimensional geometric features and two-dimensional visual features are obtained; The query vector for each agent is input into the geometric branch and the visual branch respectively; In the geometric branch, the four-dimensional geometric features are mapped to a feature space that is the same as or compatible with the agent query. Spatial location, three-dimensional scale, motion trend and temporal consistency information related to the current agent token are extracted from the four-dimensional geometric features through a cross-attention mechanism to generate a geometric response. In the visual branch, multi-view image features are mapped to the image routing space, and appearance, category, local context and semantic information are extracted from the two-dimensional visual features through a cross-attention mechanism to generate a visual response.

5. The method according to claim 4, characterized in that, The multimodal physical binding also employs adaptive gating fusion, specifically: Predict the geometric branch weights and visual branch weights based on the query vectors; The geometric response and the visual response are weighted and fused according to the geometric branch weight and the visual branch weight to generate multimodal fusion features; The multimodal fusion features are aggregated with the agent query vector to update the agent query vector.

6. The method according to claim 5, characterized in that, The geometric branch weights and visual branch weights are dynamically adjusted according to the relative distance between the target and the vehicle.

7. The method according to claim 5, characterized in that, The generation of the agent world state based on the agent representation after binding update includes: The updated agent query vector is projected into the hidden space of the multimodal large model backbone and backfilled into the corresponding discrete agent token position in the input sequence to obtain the complete context sequence; The multimodal large model backbone performs global self-attention computation on the complete context sequence, enabling each discrete agent token to interact across slots with the driving context and other tokens, and outputs global contextualized hidden features. The output vectors located at the token positions of each discrete agent are extracted from the global contextualized hidden features and aggregated into an agent state set in index order.

8. The method according to claim 1, characterized in that, Training includes a physical binding training phase, specifically: Using a dataset labeled with 3D bounding boxes, categories, and future trajectories, prediction results are output through a multimodal physical binding process, including predicted categories, predicted bounding boxes, and predicted trajectories. The prediction results are optimally associated with real traffic participants using a bipartite graph matching algorithm. Based on the optimal association, the classification loss, bounding box regression loss, and trajectory prediction loss are calculated, and the parameters involved in the multimodal physical binding process are optimized based on the losses.

9. The method according to claim 8, characterized in that, The training also includes a reasoning-oriented supervised fine-tuning phase, specifically: The bound discrete agent token and the driving context information are input into the multimodal large model; Using expert driving trajectories, high-level action commands, and agent interaction reasoning text as supervision labels, the autoregressive generation loss is calculated. Based on the autoregressive generation loss, the backbone network parameters of the multimodal large model and the policy generation head parameters of the visual language action policy model are optimized, enabling the model to learn explicit driving reasoning based on the agent's world state.

10. The method according to claim 9, characterized in that, The training also includes an agent-guided reinforcement learning phase, specifically: Sample multiple candidate planning results for the same driving scenario; The reward score for each candidate planning result is calculated using a comprehensive reward function that includes interactive safety rewards. Based on the difference in reward scores among the candidate planning results, the strategy generation head parameters and planning decision layer parameters of the visual language action strategy model are updated. The interactive safety reward is calculated based on the relative distance between the vehicle's future trajectory and the predicted trajectories of surrounding traffic participants, as well as the collision risk.

Citation Information

Patent Citations

  • Automatic driving method and system based on end-to-end and multi-modal large model

    CN120003527A