End-to-end automatic driving interpretable decision-making system and method based on multi-modal causal atlas

By combining multimodal vision-language models and causal graph construction techniques, the problem of insufficient decision-making transparency in end-to-end autonomous driving systems is solved, enabling highly transparent and credible explainable decisions, supporting user-interactive counterfactual simulations, and improving the robustness and explainability of the system.

CN120951252APending Publication Date: 2025-11-14JIANGSU UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511076429.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing end-to-end autonomous driving systems lack transparency and explainability in the decision-making process, making it difficult to provide robust and explainable decision support in complex scenarios, especially during emergency braking, yaw, or obstacle avoidance maneuvers, where passengers and operations managers find it difficult to understand the basis of the decisions.

Method used

By combining DynRsl-VLM, V3LMA multimodal vision-language models and cross-modal Transformer networks, a causal graph is constructed using the DiffVLA diffusion planning algorithm and the X-Driver model. Combined with the DriveMoE expert network hybrid mechanism, end-to-end interpretable decision-making from multimodal perception to causal analysis is achieved, supporting interactive visual explanation.

Benefits of technology

It significantly improves the decision-making transparency and credibility of autonomous driving systems. Through causal graphs and natural language interpretation, it supports user interaction in adjusting assumptions and enables real-time simulation of causal path changes in counterfactual scenarios, thereby enhancing the interpretability and safety of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951252A_ABST
    Figure CN120951252A_ABST
Patent Text Reader

Abstract

The invention provides an end-to-end automatic driving interpretable decision-making system and method based on a multi-modal causal atlas, and belongs to the technical field of intelligent driving. The method comprises the steps of data input, decision generation, causal analysis, model optimization and interactive interpretation. The objective of the invention is to construct an interpretable end-to-end automatic driving system integrating causal atlas construction, expert network mixing and interactive visualization by combining multi-modal vision-language models such as DynRsl-VLM and V3LMA with an end-to-end network and a DiffVLA diffusion planning algorithm. And the gap of decision transparency and real-time interaction requirements in the prior art is further filled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent driving technology, and in particular to an end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs and a decision system for implementing the decision-making method. Background Technology

[0002] With the rapid development of autonomous driving technology, vehicle architecture has gradually evolved from the traditional perception-planning-control modular model to end-to-end deep learning networks. This end-to-end approach directly maps data from multiple sources of sensors, such as cameras, radar, and LiDAR, to control commands, effectively reducing error accumulation caused by manual information design and intermediate variable transmission between modules. However, existing end-to-end autonomous driving solutions mostly focus on performance indicators (such as trajectory tracking accuracy and collision rate reduction), paying insufficient attention to the transparency and interpretability of the decision-making process. When the system performs emergency braking, yaw, or obstacle avoidance maneuvers, passengers and operations management personnel find it difficult to intuitively understand the decision-making basis from the "black box" model, thus restricting the safety and reliability of the system when deployed on a large scale in real-world scenarios.

[0003] On the other hand, multimodal perception technology and visual-language models (VLM / VLA) have received considerable attention in recent years. Researchers have attempted to jointly encode multi-source data such as images, point clouds, depth maps, and high-resolution maps with pre-trained language cues, utilizing cross-modal attention mechanisms to obtain more semantic scene understanding. However, most existing works only apply VLM or VLA to perception or detection tasks, and exploration of its real-time fusion of language cues and direct generation of control commands in end-to-end driving frameworks remains insufficient. Furthermore, only a few studies have introduced causal inference (such as structural equation modeling) into the control process to capture the interactions between traffic participants, resulting in a lack of robustness and interpretability in complex urban scenarios or long-tailed extreme events.

[0004] It should be noted that although some studies have attempted to analyze the decision-making process of deep models through adversarial intervention or attention visualization, most existing methods can only provide partial static explanations. They cannot combine control command outputs with multi-step reasoning processes to generate a complete visualized causal graph, nor can they output interactive "counterfactual" simulation results in complex scenarios. Therefore, current technology has not yet formed a holistic solution that can organically integrate multimodal perception, end-to-end decision-making, causal inference, and visual interaction, thereby building a highly transparent and reliable interpretable framework for autonomous driving systems while ensuring real-time performance and accuracy. Summary of the Invention

[0005] In view of this, to address the technical problem that existing technologies have not yet formed a holistic solution that can organically integrate multimodal perception, end-to-end decision-making, causal inference, and visual interaction, thereby constructing a highly transparent and reliable interpretable framework for autonomous driving systems while ensuring real-time performance and accuracy, this invention provides an end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs. This method aims to combine multimodal vision-language models such as DynRsl-VLM and V3LMA with end-to-end networks and the DiffVLA diffusion planning algorithm to construct an interpretable end-to-end autonomous driving system that integrates causal graph construction, expert network hybridization, and interactive visualization, further filling the gap in existing technologies regarding decision transparency and real-time interaction requirements.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] On the one hand, an end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs includes the following steps:

[0008] Step S1: Data Input

[0009] By utilizing the DynRsl-VLM dynamic resolution visual-language model and the V3LMA 3D perception enhancement language model, multi-source temporal information from camera, radar, stereo camera and map data is uniformly encoded and generated to generate real-time linguistic scene descriptions.

[0010] Step S2, Decision Generation

[0011] In the cross-modal Transformer network, the real-time linguistic scene description generated in step S1 is used as input. Combined with the DiffVLA differentiable diffusion mechanism, an end-to-end mapping from perceptual features to continuous control commands is realized, while attention weights and diffusion noise weights are recorded.

[0012] Step S3, Causal Analysis

[0013] Using the perceptual features and continuous control commands in step S2 as input, a causal graph is constructed through the X-Driver model. The causal weights of nodes and edges are estimated through the variational causal method, and the critical paths that affect decision-making are extracted to establish an interpretable causal structure for decision-making.

[0014] Step S4: Model Optimization

[0015] Based on the causal graph generated by S3, a hybrid mechanism of progressive course learning and DriveMoE expert network is adopted to optimize causal weight estimation and critical path selection, and output the optimized critical causal path.

[0016] Step S5, Interactive Explanation

[0017] By integrating the diffusion noise impact recorded in S2 and the key causal path optimized in S4, a visual graph and natural language explanation are generated. Users can interactively adjust the assumptions and simulate the changes in causal paths under counterfactual scenarios in real time, ultimately achieving transparent verification of the decision-making process.

[0018] On the other hand, the present invention also provides a system for executing the above-mentioned end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs, including a multimodal perception and preprocessing module, a causal reasoning and decision extraction module, and an interpretable and visual interaction module;

[0019] The multimodal sensing and preprocessing module includes:

[0020] Data Acquisition and Alignment: Integrate heterogeneous time-series data from multiple sources such as cameras, radar, stereo cameras, and high-definition maps, and achieve data synchronization through a unified timestamp and spatial coordinate system;

[0021] Standardization and cleaning: Convert data such as images, point clouds, and depth maps into a fixed format, and clean and reduce noise, removing outliers.

[0022] Feature fusion encoding: Using DynRsl-VLM and V3LMA, multi-scale feature extraction and cross-modal fusion are performed on the preprocessed data to generate a unified feature tensor containing two-dimensional vision, three-dimensional geometry and language prompts, providing input for subsequent decision-making;

[0023] The causal reasoning and decision extraction module includes:

[0024] End-to-end decision generation: Based on cross-modal fusion features, the cross-modal Transformer and DiffVLA diffusion planning algorithm are used to directly map the data into continuous control commands. Attention weights and diffusion noise weights are recorded for subsequent interpretation.

[0025] Causal graph construction: Using the X-Driver model, traffic participants are treated as nodes, and the causal weights between nodes are estimated through structural equation modeling to construct a dynamic causal graph; by comparing real-world scenarios with adversarial intervention scenarios, key causal paths influencing decision-making are screened.

[0026] Model optimization: The CurricuVLM course learning strategy is adopted to dynamically adjust the weight of training samples according to the complexity of the scene. Combined with the DriveMoE expert network hybrid mechanism, the optimal expert network is selected according to the real-time scene features to improve the robustness of causal path extraction.

[0027] The interpretable and visually interactive module includes:

[0028] Visual presentation: The causal graph, critical path, and diffuse noise weights are transformed into intuitive charts, and the causal relationships are displayed through the node-edge topology, dynamically rendering vehicle trajectories and collision probabilities;

[0029] Natural Language Interpretation: Automatically generates textual explanations based on key causal paths to help users understand the basis for their decisions;

[0030] Counterfactual simulation: Supports user interaction and intervention, reconstructs causal graphs in real time and updates decision results, demonstrating the impact of environmental changes on system decisions.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] (1) The multimodal feature fusion method based on DynRsl-VLM and V3LMA proposed in this invention encodes two-dimensional images, radar point clouds, stereo depth maps and map elements in a unified manner and generates real-time language descriptions. Based on the alignment of multi-source data, it deeply mines the semantic association between visual information and three-dimensional data, which significantly improves the semantic interpretability of perception results.

[0033] (2) The interpretable end-to-end decision generation method proposed in this invention integrates SOLVE cross-modal Transformer and DiffVLA diffusion planning. By retaining attention weights and noise contribution weights in the diffusion iteration process in the end-to-end network, the continuous control instruction reasoning process is visualized. Furthermore, attention distribution and noise influence weights are applied to decision interpretation, which significantly enhances the interpretability of the decision process.

[0034] (3) This invention adopts a method that combines X-Driver causal inference with DriveMoE expert hybrid mechanism to construct an optimized causal subgraph based on multimodal perception and end-to-end decision-making. It extracts key causal paths that affect decision-making through positive / negative example comparison learning and supports interactive counterfactual simulation, which significantly improves the accuracy of causal relationship extraction and interactive visualization capabilities, thereby enhancing the transparency and credibility of system decision-making. Attached Figure Description

[0035] Figure 1 This is a general block diagram of the interpretable end-to-end autonomous driving decision-making system described in this invention.

[0036] Figure 2 This is a schematic diagram of the interpretable end-to-end autonomous driving method described in this invention.

[0037] Figure 3 This is a schematic diagram of the cross-modal feature fusion and decision network architecture described in this invention.

[0038] Figure 4 This is a schematic diagram of the causal graph construction and expert hybrid module described in this invention.

[0039] Figure 5 This is a schematic diagram of the diffusion visualization and interactive interface described in this invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0041] The end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs described in this invention can be divided into five sequentially executed sub-steps, with the output of the previous step serving as the input for the next step, such as... Figure 2 As shown, the processing flow is as follows:

[0042] An end-to-end interpretable decision-making method for autonomous driving based on multimodal causal graphs includes the following steps:

[0043] Step S1: Data Input

[0044] By utilizing the DynRsl-VLM dynamic resolution visual-language model and the V3LMA 3D perception-enhanced language model, multi-source temporal information from cameras, radar, stereo cameras, and maps is uniformly encoded (preferably uniformly encoded in a multi-scale feature pyramid) to generate real-time linguistic scene descriptions. Step S1 is the multimodal data acquisition and preprocessing step, which involves acquiring data from cameras, radar, stereo cameras, and high-definition maps, and performing alignment, format standardization, filtering, and noise reduction to output feature sequences that meet the model input requirements.

[0045] Step S2, Decision Generation

[0046] In the cross-modal Transformer network, the real-time linguistic scene description generated in step S1 is used as input. Combined with the DiffVLA differentiable diffusion mechanism, an end-to-end mapping from perceptual features to continuous control commands is achieved, while attention weights and diffusion noise weights are recorded. Step S2 is the cross-modal feature extraction and end-to-end decision generation. After fusing the preprocessed features with the DynRsl-VLM and V3LMA encoders, the results are fed into the SOLVE / DiffVLA network, which simultaneously outputs continuous control commands and records diffusion noise weights.

[0047] Step S3, Causal Analysis

[0048] Using the perceptual features and continuous control commands from step S2 as input, a causal graph is constructed through the X-Driver model. The causal weights of nodes and edges are estimated using a variational causal method, and critical paths influencing decisions are extracted to establish an interpretable causal structure for the decisions. Step S3 is the causal graph construction and preliminary path selection—a structural equation network is constructed based on cross-modal fusion features and decision results, generating positive / negative causal subgraphs and selecting candidate critical causal paths through path scoring.

[0049] Step S4: Model Optimization

[0050] Based on the causal graph generated in S3, a hybrid mechanism of progressive course learning and DriveMoE expert network is adopted to optimize causal weight estimation and critical path selection, and output the optimized critical causal path. Step S4 is the optimization of expert hybridization and course learning. CurricuVLM courses are scheduled according to the difficulty of the scenario, and the DriveMoE gating mechanism is used to dynamically select different experts to extract and weightedly fuse the final critical path.

[0051] Step S5, Interactive Explanation

[0052] Integrating the diffusion noise impact recorded in S2 and the key causal path optimized in S4, a visual graph and natural language explanation are generated. This allows users to interactively adjust the assumptions and simulate causal path changes in counterfactual scenarios in real time, ultimately achieving transparent verification of the decision-making process. Step S5 is

[0089] Diffusion Visualization and Interactive Interface—using the noise contribution weights recorded during the DiffVLA diffusion iteration process and the extracted key causal paths, a visual graph and natural language explanation are generated, supporting user counterfactual interaction.

[0053] In this invention, step S1 can specifically be:

[0054] Step S11: Spatiotemporal alignment of multi-source data

[0055] Align camera image frames, radar point cloud data, stereo camera depth maps, and high-definition map data with a unified timestamp and spatial coordinate system to obtain a synchronized sample set under the same reference system.

[0056] Step S12: Data format standardization

[0057] The format of the processed raw data is standardized, the camera images are converted into a fixed resolution pixel matrix, the radar point cloud is represented by dense rasterization, the stereo depth map is also mapped to a depth pixel matrix, and the map information is extracted into vectorized road elements.

[0058] Step S13: Cleaning and Noise Reduction Treatment

[0059] The standardized data was cleaned and denoised in the following way: noise point clouds with abrupt changes in continuous frames and non-continuous image frames were removed; Gaussian filtering was used to smooth the image noise; radius neighborhood filtering was used to remove isolated points in the point cloud data; and bilateral filtering was applied to the stereo depth map to reduce the impact of measurement errors on subsequent perception.

[0060] Step S14: Construction of multimodal feature sequences

[0061] Based on the multimodal features collected by the aforementioned high-precision sensor at each time step t, a sequence {I} is constructed. t ,P t D t M t}, where I t ∈R H×W×3 Represents the image pixel tensor. N represents the radar point cloud p The three-dimensional coordinates of point D t ∈R H×W Represents the depth pixel matrix, Represents vectorized map elements;

[0062] Step S15: Unified Feature Tensor Fusion

[0063] Constructing a unified feature tensor:

[0064] X τ =Concat(Enc img (I τ ),Enc pc (P τ ),Enc dep (D τ ),Enc map (M τ )),

[0065] Among them, Enc img (·), Enc pc (·), Enc dep (·) and Enc map (·) represent the pre-trained image, point cloud, depth and map encoding network, respectively, to obtain the fused features. This serves as the input for subsequent DynRsl-VLM and V3LMA.

[0066] Wherein, feature dimension d f The feature sequence length L is determined by the sum of the output dimensions of each encoder. f The number of time steps after fusion is determined; the parameters of each encoder are initialized and fine-tuned during the training phase based on the pre-trained weights of DynRsl-VLM and V3LMA.

[0067] In this invention, step S2 can specifically be:

[0068] Step S21: Multimodal feature separation and encoding

[0069] by As inputs to DynRsl-VLM and V3LMA, two-dimensional language features are obtained respectively. and three-dimensional language features Where d m The multimodal feature dimension designed jointly for DynRsl-VLM and V3LMA. Here... This is the "unified feature tensor" obtained at each time step t, following step S15.

[0070] This means collecting these fused features from S15 in chronological order to form a time series of length T: (X1, X2, X3...X...) t T is the preset historical time window size, corresponding to the high-dimensional features after the fusion of all modes at time t.

[0071] In step S21, the entire sequence of length T is used as the input to the network, allowing the model to not only see the environmental information of the current frame, but also capture the temporal context of the past (T-1) frames, so as to learn the temporal dependencies in dynamic scenes.

[0072] Specifically, to The inputs to the DynRsl-VLM and V3LMA encoders can perform cross-temporal and cross-modal feature fusion within them, generating a multimodal representation that includes both historical trajectory information and real-time semantic cues. This provides a complete spatiotemporal context for the subsequent Transformer attention network to extract spatial dependencies and generate control commands.

[0073] Step S22: Cross-modal feature fusion and location encoding

[0074] Will and By splicing along the channel direction, cross-modal fusion features are obtained. Add learnable positional codes after splicing.

[0075] Step S23, Cross-modal Transformer encoding

[0076] In a cross-modal Transformer, let the query matrix... Key matrix Value matrix in For learnable weights, dk Indicates the dimension of the key / query vector in the attention header;

[0077] Cross-modal feature A is calculated using a multi-head self-attention mechanism. t The formula for calculating single-head attention is:

[0078]

[0079] Where Softmax(·) represents row-wise normalization. is a scaling factor used to stabilize the gradient, indicating that matrix K is transposed (i.e., rows and columns are interchanged). It is the transpose of K, used to calculate the similarity between the query and the key in the attention mechanism.

[0080] Multi-head attention is achieved by concatenating h heads after parallel computation:

[0081] A t =Concat(head1,…,head) h W O ,head i =Attention(Q) i ,K i V i ),

[0082] in, To output the mapping matrix, Q i ,K i V i These are slices of Q, K, and V at the i-th head, respectively;

[0083] A t The final feature is obtained after passing through several layers of feedforward fully connected networks and LayerNorm.

[0084] Let H t After regression head MLP ctrl Predictive continuous control command u t =[δ t ,a t ,b t ], where δ t For steering angle (unit: radians), a t For acceleration (unit: m / s²), b t The braking intensity is expressed in relative terms; the output delay of this process can be controlled within 60ms.

[0085] In this invention, the construction of the causal graph and the extraction of critical paths in step S3 include:

[0086] Step S31: Filtering of traffic entity nodes

[0087] based on The multimodal perception features and decision-making results in the process select all traffic participants (such as pedestrians, vehicles, and traffic lights) in the scene as the causal graph node set V within the same time t. t ={v1,…,v N}; This refers to the pair of "multimodal fusion features H" output at each time step t in step S2. t "and control command u" t The time series set of "".

[0088] Step S32: Estimation of Causal Structure Equation and Weight Matrix

[0089] Using the X-Driver model for each pair of nodes (v i ,v j Establish structural equations:

[0090]

[0091] Where, x i Represents node v i Extracted multimodal feature values ​​(H can be taken as H) t (the row vector of the corresponding entity in the middle), a ij Indicates from node v i to v j causal weight, ε j Noise term;

[0092] The least squares estimation or regularized regression method is used to evaluate {a}. ij Estimate the causal weight matrix C to construct the causal weight matrix. t ∈R N×N Simultaneously, according to the current control command u t The dependencies between the features of each node form a complete causal graph G. t =(V t E t ), where edge set E t ={(v i ,v j )|[C t ] ij ≠0};

[0093] Step S33: Comparison of positive examples and counterfactual trajectories

[0094] Input the real driving trajectory and the counterfactual trajectory under "adversarial intervention" (such as manually deleting a pedestrian or changing a traffic light status) into the same model, and calculate the positive example graph. With negative example diagram

[0095] right and By comparing the differences between node weights and edge weights, a set of critical causal paths P is defined. t ={p k}, where each path Satisfying path gain score:

[0096]

[0097] in, and These are the causal weights of the corresponding edges in the positive and negative example graphs, respectively. The path with the higher score is the key causal path that influences the decision.

[0098] In this invention, the specific process of combining causal model training with expert network in step S4 is as follows:

[0099] Step S41: Measure the complexity of the scene

[0100] Within the CurricuVLM framework, training samples are categorized into gradient difficulty levels based on the complexity of different driving scenarios. Assuming a scenario *s* contains N(s) traffic entities and D(s) dynamic obstacles, these can be combined using empirical weights α and β to form a difficulty score.

[0101]

[0102] Where, N max With D max The maximum number of vehicles and the maximum proportion of dynamic obstacles in the dataset are used to normalize N(s) and D(s), respectively. φ(s) describes the relative degree of the scene from "simple" to "complex". φ(s)∈[0,1] α and β are empirical weight hyperparameters that measure the relative influence of vehicle density and dynamic obstacle proportion on the scene complexity score, respectively, and satisfy α+β=1;

[0103] Step S42, Gradient Sample Training

[0104] The training samples are graded according to the difficulty score φ(s), and the sampling weights w are dynamically adjusted during the training process. t (s):

[0105]

[0106] Where t is the current iteration round number, T is the total number of iteration rounds, and Φ max To maximize the possible difficulty score, the model is trained in the order of "easy → medium → difficult" so that the causal subgraph extractor has a stronger generalization ability for complex scenes.

[0107] Step S43, DriveMoE Expert Network Fusion

[0108] Construct a DriveMoE expert network, and at the same time t, fuse the features H output by DynRsl-VLM and V3LMA. t Input gating network:

[0109] g t =Softmax(W g H t +b g )∈R E ,

[0110] in, With b g ∈R E Here, E represents the gating weights and biases, and E represents the number of expert networks; the gating output g is... t =[g t,1 ,…,g t,E ] represents the weight of each expert, satisfying

[0111] Expert sub-networks i Based on H t Independent estimation of causal graph G t,i and critical path P t,i The final set of causal paths P t Obtained by weighted fusion:

[0112]

[0113] Of these, only the score s(p)×g is retained. t,i Paths greater than a threshold τ are filtered out to remove weak causal relationships.

[0114] In this invention, step S5 specifically includes:

[0115] Step S51: Record the path weights in the diffusion planning.

[0116] In the t-th iteration of the DiffVLA diffusion programming network, for the current sample x t Through noise prediction network ε θ (x t ,t) Calculate the predicted noise and use the diffusion sampling formula:

[0117]

[0118] Where, α t To preset the noise attenuation coefficient, σ t For random noise intensity, zt ~N(0,I);

[0119] By recording each step ε θ (x t The contribution weight of ,t) to the final control command u0 provides data basis for generating "diffusion path visualization".

[0120] Step S52: Automatic generation of readable text explanation

[0121] Based on the key causal path set P extracted in steps S3 and S4 t Automatically generate readable text explanation E:

[0122]

[0123] The template function Template(·) is based on the node sequence. With corresponding causal weights Automatically fill in the language description, such as "detected". The causal path has a higher weight, so the corresponding braking strategy is selected.

[0124] Step S53: Visualization and Interaction of Cause-and-Effect Subgraphs

[0125] Load the causal subgraph G at the current moment in the user interface. t With P t Each path is labeled with its color depth and width, representing the size of s(p); when the user clicks any node v with the mouse or touch... i In real time, it displays the upstream and downstream nodes associated with it and their corresponding weights;

[0126] Step S54, Counterfactual Intervention Simulation

[0127] When the user selects a certain node v i Perform counterfactual intervention (e.g., "assuming the pedestrian does not exist"), and then re-run the structural equation network from step S3 to apply the intervened features. Alternatively, you can change the language description of the traffic light status and re-enter it to obtain a counterfactual cause-effect graph. And update the critical path To present "how the system will adjust its decisions under this assumption";

[0128] Step S55: Synchronous Display of Multimodal Results

[0129] The above textual explanations and visualization results are packaged and synchronized to the user's terminal, presenting the system's causal inferences and decision-making basis in the current and counterfactual scenarios in the form of pictures and text.

[0130] This invention, through the organic combination of the above steps, realizes end-to-end autonomous driving decision-making from multimodal perception to interpretable causal graphs and visualization of diffusion planning actions. It can also support counterfactual simulation on a real-time interactive interface, providing highly transparent decision-making and interpretation capabilities for autonomous driving systems.

[0131] Figure 3 An exemplary cross-modal feature fusion and decision network is shown, which is used to generate end-to-end control commands and extract interpretable information from features after multimodal perception.

[0132] Specifically, refer to Figure 3 The cross-modal feature fusion and decision network architecture shown will convert the original feature tensor output by the multi-source preprocessing module into a multi-source preprocessing module. The input is fed to the DynRsl-VLM and processed in parallel with the V3LMA encoder.

[0133] DynRsl-VLM first performs multi-scale pyramid encoding on the features of the 2D image and radar point cloud, and then fuses predefined linguistic cues (such as "pedestrian is veering off the lane") with visual features through cross-modal attention to output 2D-linguistic features. Meanwhile, V3LMA performs 3D context encoding on stereo depth maps and high-resolution map elements, and embeds semantic cues such as "road boundaries" and "speed limit signs" into depth features through a structured cross-modal fusion layer, outputting 3D-linguistic features. Then and Cross-modal fusion features are obtained by concatenating them along the channel dimension. And add learnable positional coding This constitutes the input to the Transformer.

[0134] In the Transformer encoder, let

[0135]

[0136] in,

[0137] Each attention should be focused on:

[0138]

[0139] Calculate the attention weights for each, and then obtain them through matrix concatenation and linear mapping.

[0140]

[0141] in,

[0142] After completing the multi-head self-attention, At It is fed into a two-layer feedforward network, in the form of H ' t =ReLU(A t W1+b1)W2+b2, and after residual connection and LayerNorm normalization, the final high-dimensional cross-modal representation is obtained.

[0143] In order to generate control commands, H t By converging along the time dimension (e.g., taking the first vector or global average), a fixed-length vector is obtained. And regress the continuous control vector through the multilayer perceptron control head. Where δ t For the steering angle (in radians), a t Let b be the acceleration (m / s²). t This represents the braking intensity (relatively normalized value).

[0144] Besides returning to u t In addition, this module also runs in parallel with the DiffVLA diffusion planning network: using random noise x T Initially, the iterative denoising process is as follows:

[0145]

[0146] Where, α t This is the noise attenuation coefficient. ε θ (x t ,t) is the result of H t and current noise x t The estimated value σ obtained by the input noise prediction network t z is a scalar measure of noise variance. t ~N(0,I). Throughout the entire diffusion iteration, ε θ (x t The ,t) will be retained to measure the contribution of each step to the final generated instructions, providing key information for subsequent visualization.

[0147] The module shown not only completes end-to-end cross-modal perception and control generation, but also outputs attention weights and diffusion noise prediction information to the interpretable level.

[0148] The X-Driver causal inference and DriveMoE expert hybrid mechanism can further transform the output of the cross-modal decision-making stage into interpretable causal relationships.

[0149] Specifically, refer to Figure 4 The causal graph construction and expert hybrid module shown in the figure, at each time t, the system first performs H t With control command u tThe node feature set {x1,…,x} mapped to traffic participants N},in Let i represent the multimodal features of i entities. Based on this, a structural equation model is adopted:

[0150]

[0151] For each pair of nodes v i →v j Perform causal weighting a ij The estimate, of which W u To control the context mapping matrix, ε j This represents the noise term. The causal weight matrix A, which imparts sparsity, can be obtained through least squares or Lasso regression. t =[a ij The system constructs "positive example diagrams" in parallel. (Using real features and instructions) and "negative example diagrams" (Re-estimate after intervention on the target node), and define the path score based on the difference in weights between positive and negative examples:

[0152]

[0153] in, The higher the path score, the more significant the impact of that path on the current decision. The system uses this score to select a set of candidate critical paths {P}. t,i}

[0154] To improve the robustness of causal extraction in long-tail extreme scenarios, this invention employs the CurricuVLM course learning strategy during the offline training phase, calculating a difficulty score for each training session s:

[0155]

[0156] Where N(s) is the number of entities in the scene, D(s) is the proportion of dynamic obstacles, α and β are empirical weights, and N max D max These are the maximum values ​​in the dataset;

[0157] The sample weights are dynamically assigned based on the current iteration round t and the total number of rounds T:

[0158]

[0159] Gradually switch from simple to complex scenarios to ensure that each expert network can fully learn its corresponding scenario type.

[0160] During the online inference phase, the system will use H t Flatten the input and enter the DriveMoE gating network:

[0161] g t =Softmax(W g Flatten(H t )+b g )∈R E ,

[0162] in, b g ∈R E E represents the number of experts.

[0163] Gated output {g t,1 ,…,g t,E} represents the relative contribution of different experts in the current scenario, and each expert subnetwork is based on H t Independent estimation of candidate path P t,i And with:

[0164]

[0165] The final critical path with a weighted score exceeding the threshold τ is selected.

[0166] Figure 4 The causal subgraph G output by the module shown t With critical path P t , combined Figure 3 The noise prediction information recorded during the DiffVLA iteration process generates rich visualization maps and linked explanatory texts for the interpretable and visually interactive modules.

[0167] Figure 5 This is a schematic diagram of the multimodal interpretation visualization and interactive interface described in this invention.

[0168] The diagram contains three main functional modules: the top left corner is the "Diffusion Process Visualization" area, which intuitively displays the iterative denoising process from the noise state xT to the generated state x0. The weight parameters of the noise prediction network are marked next to each step, reflecting the degree of contribution of different steps to the final result.

[0169] The bottom left corner is the "Causal Graph Display" area, which uses a node-edge topology to present key causal relationships in driving scenarios. Circular nodes label traffic participants such as pedestrians and vehicles, and the connecting edges are labeled with unequal causal weight values.

[0170] The main area on the right is the "Interactive Driving Scene" area, which includes: a 3D bird's-eye view of the road scene, overlaid rendering of vehicle trajectory lines (the color depth indicates the probability of collision), a right-side explanation panel that dynamically displays the explanatory text of the currently selected causal path, allows users to select any traffic participant by touch, and the system will highlight its relevant causal links, and a counterfactual reasoning button at the bottom, which, when clicked, can instantly generate and compare the changes in the causal graph before and after intervention.

[0171] The interface uses different colors (such as red for high risk and green for safety) and dynamic prompts to intuitively present the decision-making basis of the autonomous driving system.

[0172] The present invention also provides a system for executing the above-described end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs, including a multimodal perception and preprocessing module, a causal reasoning and decision extraction module, and an interpretable and visual interaction module;

[0173] The multimodal sensing and preprocessing module includes:

[0174] Data Acquisition and Alignment: Integrate heterogeneous time-series data from multiple sources such as cameras, radar, stereo cameras, and high-definition maps, and achieve data synchronization through a unified timestamp and spatial coordinate system;

[0175] Standardization and cleaning: Convert data such as images, point clouds, and depth maps into a fixed format, and clean and reduce noise, removing outliers.

[0176] Feature fusion encoding: Using DynRsl-VLM and V3LMA, multi-scale feature extraction and cross-modal fusion are performed on the preprocessed data to generate a unified feature tensor containing two-dimensional vision, three-dimensional geometry and language prompts, providing input for subsequent decision-making;

[0177] The causal reasoning and decision extraction module includes:

[0178] End-to-end decision generation: Based on cross-modal fusion features, the cross-modal Transformer and DiffVLA diffusion planning algorithm are used to directly map the data into continuous control commands. Attention weights and diffusion noise weights are recorded for subsequent interpretation.

[0179] Causal graph construction: Using the X-Driver model, traffic participants are treated as nodes, and the causal weights between nodes are estimated through structural equation modeling to construct a dynamic causal graph; by comparing real-world scenarios with adversarial intervention scenarios, key causal paths influencing decision-making are screened.

[0180] Model optimization: The CurricuVLM course learning strategy is adopted to dynamically adjust the weight of training samples according to the complexity of the scene. Combined with the DriveMoE expert network hybrid mechanism, the optimal expert network is selected according to the real-time scene features to improve the robustness of causal path extraction.

[0181] The interpretable and visually interactive module includes:

[0182] Visual presentation: The causal graph, critical path, and diffuse noise weights are transformed into intuitive charts, and the causal relationships are displayed through the node-edge topology, dynamically rendering vehicle trajectories and collision probabilities;

[0183] Natural Language Interpretation: Automatically generates textual explanations based on key causal paths to help users understand the basis for their decisions;

[0184] Counterfactual simulation: Supports user interaction and intervention, reconstructs causal graphs in real time and updates decision results, demonstrating the impact of environmental changes on system decisions.

[0185] like Figure 1 As shown, the present invention provides an exemplary end-to-end autonomous driving interpretable decision system and method based on multimodal causal graphs, including a multimodal perception and preprocessing module, a causal reasoning and decision extraction module, and an interpretable and visual interaction module.

[0186] The multimodal perception and preprocessing module is responsible for the comprehensive collection and unified processing of environmental information by the system. This involves real-time acquisition of temporal data of the vehicle's surroundings from heterogeneous sensor sources such as front-mounted cameras, LiDAR, stereo depth cameras, and high-definition maps. Under a unified spatiotemporal reference, this data is accurately registered and standardized in format. After filtering, noise reduction, and anomaly removal, the data is cleaned and aligned. The module then sends the fused features to the cross-modal coding subsystem to generate a unified representation that includes 2D vision, 3D geometry, and predefined language prompts, providing a basis for the next step of decision-making.

[0187] The causal reasoning and decision extraction module takes the preprocessed multimodal features as input and performs cross-modal fusion and action prediction in an end-to-end network. Visual and verbal cues are fused as learnable tokens to generate a continuous framework of control commands such as steering, acceleration, and braking. While obtaining continuous control outputs, the module retains the attention weights and activation features of the intermediate layers of the network, providing a basis for interpretability analysis. In parallel, a real-time causal graph is constructed: the system maps traffic participants (pedestrians, other vehicles, traffic lights, etc.) in the current frame to graph nodes, estimates the causal weights between nodes using a structural equation model, and filters out the critical paths influencing the current decision using a "positive example-negative example" comparison. Combining the CurricuVLM curriculum learning strategy and the DriveMoE expert hybrid mechanism, the system can dynamically allocate expert networks under different scene complexities to precisely extract the most representative causal relationships.

[0188] The module outputs two results: one is real-time control commands for the vehicle to execute, and the other is several key causal paths for the interpretable and visual interactive modules to use.

[0189] The Explainable and Visual Interaction module is primarily responsible for transforming the internal processes of causal reasoning and diffusion denoising into an intuitive and interactive interface. It maps key causal paths to a visual graph and generates corresponding natural language annotations, enabling users to intuitively understand why the system will perform a certain action at this moment. During the denoising iteration process of the diffusion planning network, it saves the impact of each step of noise prediction and language prompts on the final control. When a user makes a hypothesis about a certain environmental factor through interactive means such as clicking or touching (e.g., "assuming the pedestrian has not appeared"), the system will immediately reconstruct the counterfactual causal graph and refresh the visual interface, showing "how the causal path will change if the environment changes, and what different decisions the system will make."

[0190] Through this two-way interaction, users can comprehensively examine system decisions at both the spatial and semantic levels, ensuring the transparency of decisions and significantly increasing the trust of passengers or operators in the internal logic of the autonomous driving system.

[0191] The above description is merely a preferred embodiment of the present invention; however, the scope of protection of the present invention is not limited thereto; any equivalent substitutions or modifications made by those skilled in the art within the technical scope disclosed in the present invention, based on the technical solution and its improved concept, should be covered within the scope of protection of the present invention.

Claims

1. An end-to-end explainable decision-making method for autonomous driving based on multimodal causal graphs, characterized in that, Includes the following steps: Step S1: Data Input By utilizing the DynRsl-VLM dynamic resolution visual-language model and the V3LMA 3D perception enhancement language model, multi-source temporal information from camera, radar, stereo camera and map data is uniformly encoded and generated to generate real-time linguistic scene descriptions. Step S2, Decision Generation In the cross-modal Transformer network, the real-time linguistic scene description generated in step S1 is used as input. Combined with the DiffVLA differentiable diffusion mechanism, an end-to-end mapping from perceptual features to continuous control commands is realized, while attention weights and diffusion noise weights are recorded. Step S3, Causal Analysis Using the perceptual features and continuous control commands in step S2 as input, a causal graph is constructed through the X-Driver model. The causal weights of nodes and edges are estimated through the variational causal method, and the critical paths that affect decision-making are extracted to establish an interpretable causal structure for decision-making. Step S4: Model Optimization Based on the causal graph generated by S3, a hybrid mechanism of progressive course learning and DriveMoE expert network is adopted to optimize causal weight estimation and critical path selection, and output the optimized critical causal path. Step S5, Interactive Explanation By integrating the diffusion noise impact recorded in S2 and the key causal path optimized in S4, a visual graph and natural language explanation are generated. Users can interactively adjust the assumptions and simulate the changes in causal paths under counterfactual scenarios in real time, ultimately achieving transparent verification of the decision-making process.

2. The end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs according to claim 1, characterized in that, In step S1, the unified encoding of multi-source time series information is specifically as follows: Step S11: Spatiotemporal alignment of multi-source data Align camera image frames, radar point cloud data, stereo camera depth maps, and high-definition map data with a unified timestamp and spatial coordinate system to obtain a synchronized sample set under the same reference system. Step S12: Data format standardization The format of the processed raw data is standardized, the camera images are converted into a fixed resolution pixel matrix, the radar point cloud is represented by dense rasterization, the stereo depth map is also mapped to a depth pixel matrix, and the map information is extracted into vectorized road elements. Step S13: Cleaning and Noise Reduction Treatment The standardized data is then cleaned and noise-reduced. Step S14: Construction of multimodal feature sequences After integrating the processed data at time step t, a time-series feature sequence {I} is constructed. t ,P t D t M t }, where I t ∈R H×W×3 Represents the image pixel tensor. N represents the radar point cloud p The three-dimensional coordinates of point D t ∈R H×W M represents the depth pixel matrix. t ∈R Lm Represents vectorized map elements; Step S15: Unified Feature Tensor Fusion Through the pre-trained encoder Enc img (·), Enc pc (·), Enc dep (·) and Enc map (·) Encode the features of the image, point cloud, depth map, and map respectively, and concatenate the encoding results into a unified feature tensor to obtain the fused features. As input for subsequent DynRsl-VLM and V3LMA; X τ =Concat(Enc img (I τ ),Enc pc (P τ ),Enc dep (D τ ),Enc map (M τ )), Wherein, feature dimension d f The feature sequence length L is determined by the sum of the output dimensions of each encoder. f The number of time steps after fusion is determined; the parameters of each encoder are initialized and fine-tuned during the training phase based on the pre-trained weights of DynRsl-VLM and V3LMA.

3. The end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs according to claim 2, characterized in that, In step S13, cleaning and noise reduction are performed as follows: noise point clouds with abrupt changes in continuous frames and non-continuous image frames are removed; Gaussian filtering is used to smooth image noise; radius neighborhood filtering is used to remove isolated points in point cloud data; and bilateral filtering is applied to the stereo depth map to reduce the impact of measurement errors on subsequent perception.

4. The end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs according to claim 1, characterized in that, In step S2, the end-to-end mapping from perceived features to continuous control commands includes: Step S21: Multimodal feature separation and encoding by As inputs to DynRsl-VLM and V3LMA, two-dimensional language features are obtained respectively. and three-dimensional language features Where d m Multimodal feature dimensions jointly designed for DynRsl-VLM and V3LMA; Step S22: Cross-modal feature fusion and location encoding Will and By splicing along the channel direction, cross-modal fusion features are obtained. Add learnable positional codes after splicing. Step S23, Cross-modal Transformer encoding In a cross-modal Transformer, let the query matrix... Key matrix Value matrix in For learnable weights, d k This represents the dimension of the key / query vector in the attention head; the formula for calculating single-head attention is: Where Softmax(·) represents row-wise normalization, which is used to assign feature weights. This is a scaling factor used to stabilize the gradient; Cross-modal feature A is calculated using a multi-head self-attention mechanism. t Multi-head attention is achieved by concatenating h heads after parallel computation: A t =Concat(head1,…,head h )W O ,head i =Attention(Q i ,K i ,V i ), in, To output the mapping matrix, Q i ,K i V i These are slices of Q, K, and V at the i-th head, respectively; A t The final features are obtained after passing through several layers of feedforward fully connected networks and LayerNorm. Let H t After regression head MLP ctrl Predictive continuous control command u t =[δ t ,a t ,b t ], where δ t For steering angle (unit: radians), a t For acceleration (unit: meters per second), b t Braking intensity (unit: relative value).

5. The end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs according to claim 1, characterized in that, Step S3, the construction of the causal graph and the extraction of critical paths, includes: Step S31: Filtering of traffic entity nodes based on The multimodal perception features and decision results are used to select all traffic participants in the scene as the causal graph node set V within the same time t. t ={v1,…,v N }; Step S32: Estimation of Causal Structure Equation and Weight Matrix Using the X-Driver model for each pair of nodes (v i ,v j Establish structural equations: Where, x i Represents node v i Extracted multimodal feature values, a ij Indicates from node v i to v j causal weight, ε j Noise term; The least squares estimation or regularized regression method is used to evaluate {a}. ij Estimate the causal weight matrix C to construct the causal weight matrix. t ∈R N×N Simultaneously, according to the current control command u t The dependencies between the features of each node form a complete causal graph G. t =(V t E t ), where edge set E t ={(v i ,v j )|[C t ] ij ≠0}; Step S33: Comparison of positive examples and counterfactual trajectories Input the real driving trajectory and the counterfactual trajectory under "adversarial intervention" into the same model respectively, and calculate the positive example diagram. With negative example diagram right and By comparing the differences between node weights and edge weights, a set of critical causal paths P is defined. t ={p k }, where each path Satisfying path gain score: in and These are the causal weights of the corresponding edges in the positive and negative example graphs, respectively. The path with the higher score is the key causal path that influences the decision.

6. The end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs according to claim 1, characterized in that, The specific process of step S4 is as follows: Step S41: Measure the complexity of the scene Within the CurricuVLM framework, training samples are divided into gradient difficulty levels based on the complexity of different driving scenarios. Assuming that a scenario s contains N(s) traffic entities and D(s) dynamic obstacles, the difficulty score can be synthesized using empirical weights α and β. Where, N max With D max These represent the maximum number of vehicles and the maximum proportion of dynamic obstacles in the dataset, respectively. φ(s) describes the relative degree of the scene from "simple" to "complex", φ(s)∈[0,1]. α and β are empirical weight hyperparameters, which measure the relative influence of vehicle density and the proportion of dynamic obstacles on the scene complexity score, respectively, and satisfy α+β=1. Step S42, Gradient Sample Training The training samples are graded according to the difficulty score φ(s), and the sampling weights w are dynamically adjusted during the training process. t (s): Where t is the current iteration round number, T is the total number of iteration rounds, and Φ max To maximize the possible difficulty score, the model is trained in the order of "easy → medium → difficult" so that the causal subgraph extractor has a stronger generalization ability for complex scenes. Step S43, DriveMoE Expert Network Fusion Construct a DriveMoE expert network, and at the same time t, fuse the features H output by DynRsl-VLM and V3LMA. t Input gating network: g t =Softmax(W g H t +b g )∈R E , in, With b g ∈R E Here, E represents the gating weights and biases, and E represents the number of expert networks; the gating output g is... t =[g t,1 ,…,g t,E ] represents the weight of each expert, satisfying Expert sub-networks i Based on H t Independent estimation of causal graph G t,i and critical path P t,i The final set of causal paths P t Obtained by weighted fusion: Of these, only the score s(p)×g is retained. t,i Paths greater than a threshold τ are filtered out to remove weak causal relationships.

7. The end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs according to claim 1, characterized in that, Step S5 is as follows: Step S51: Record the path weights in the diffusion planning. In the t-th iteration of the DiffVLA diffusion programming network, for the current sample x t Through noise prediction network ε θ (x t ,t) Calculate the predicted noise and use the diffusion sampling formula: Where, α t To preset the noise attenuation coefficient, σ t For random noise intensity, z t ~N(0,I); By recording each step ε θ (x t The contribution weight of ,t) to the final control command u0 provides data basis for generating "diffusion path visualization"; Step S52: Automatic generation of readable text explanation Based on the key causal path set P extracted in steps S3 and S4 t Automatically generate readable text explanation E: The template function Template(·) is based on the node sequence. With corresponding causal weights Automatically fill in the language description; Step S53: Visualization and Interaction of Cause-and-Effect Subgraphs Load the current causal subgraph into the user interface and label each path with color depth and width to represent its size; when the user clicks any node with the mouse or touch, display the upstream and downstream nodes associated with it and their corresponding weights in real time. Step S54, Counterfactual Intervention Simulation When a user chooses to perform counterfactual intervention on a certain node, the structural equation network of step S3 is rerun, and the features after intervention or the linguistic description of the change in traffic light status is re-inputted to obtain a counterfactual causal graph and update the critical path to present "how the system will adjust its decision under this assumption". Step S55: Synchronous Display of Multimodal Results The above textual explanations and visualization results are packaged and synchronized to the user's terminal, presenting the system's causal inferences and decision-making basis in the current and counterfactual scenarios in the form of pictures and text.

8. A system for executing the end-to-end autonomous driving interpretable decision-making method based on multimodal causal graphs according to any one of claims 1-7, characterized in that, It includes a multimodal perception and preprocessing module, a causal reasoning and decision extraction module, and an interpretable and visual interaction module; The multimodal sensing and preprocessing module includes: Data Acquisition and Alignment: Integrate heterogeneous time-series data from multiple sources such as cameras, radar, stereo cameras, and high-definition maps, and achieve data synchronization through a unified timestamp and spatial coordinate system; Standardization and cleaning: Convert data such as images, point clouds, and depth maps into a fixed format, and clean and reduce noise, removing outliers. Feature fusion encoding: Using DynRsl-VLM and V3LMA, multi-scale feature extraction and cross-modal fusion are performed on the preprocessed data to generate a unified feature tensor containing two-dimensional vision, three-dimensional geometry and language prompts, providing input for subsequent decision-making; The causal reasoning and decision extraction module includes: End-to-end decision generation: Based on cross-modal fusion features, the cross-modal Transformer and DiffVLA diffusion planning algorithm are used to directly map the data into continuous control commands. Attention weights and diffusion noise weights are recorded for subsequent interpretation. Causal graph construction: Using the X-Driver model, traffic participants are treated as nodes, and the causal weights between nodes are estimated through structural equation modeling to construct a dynamic causal graph; by comparing real-world scenarios with adversarial intervention scenarios, key causal paths influencing decision-making are screened. Model optimization: The CurricuVLM course learning strategy is adopted to dynamically adjust the weight of training samples according to the complexity of the scene. Combined with the DriveMoE expert network hybrid mechanism, the optimal expert network is selected according to the real-time scene features to improve the robustness of causal path extraction. The interpretable and visually interactive module includes: Visual presentation: The causal graph, critical path, and diffuse noise weights are transformed into intuitive charts, and the causal relationships are displayed through the node-edge topology, dynamically rendering vehicle trajectories and collision probabilities; Natural Language Interpretation: Automatically generates textual explanations based on key causal paths to help users understand the basis for their decisions; Counterfactual simulation: Supports user interaction and intervention, reconstructs causal graphs in real time and updates decision results, demonstrating the impact of environmental changes on system decisions.

Citation Information

Cited By

  • Pedestrian intention recognition method and system based on hierarchical spatio-temporal causal diagram and state diffusion

    CN121708571A

  • Causal reasoning method and device for automatic driving decision

    CN122153457A

  • Automatic driving data set construction system and method

    CN122217287A