An Interpretable Hierarchical Reasoning Multimodal Method and System for Highway Scenarios

By introducing a hierarchical reasoning structure and causal activation mechanism into the multimodal model, the problems of insufficient reasoning depth and weak interpretability in highway scenarios are solved, achieving efficient and transparent multimodal reasoning and improving the model's application capability and credibility in complex traffic scenarios.

CN120822624BActive Publication Date: 2025-12-02SHANDONG HI SPEED GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511332313.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-02
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Existing multimodal large models suffer from insufficient inference depth, severe contextual interference, and weak interpretability in highway scenarios, making it difficult to meet the requirements of high-level abstract understanding and low-latency response.

Method used

We employ an interpretable hierarchical reasoning multimodal approach, which performs global semantic understanding and long-term reasoning in the high-level part, and fine-grained computation in the low-level part. We combine this with a causal activation mechanism for visual interpretation to construct a deep reasoning path.

Benefits of technology

It significantly improves the model's application capability and credibility in complex traffic scenarios, enhances the model's robustness and transparency, and supports decision tracking, fault tracing, and safety verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822624B_ABST
    Figure CN120822624B_ABST
Patent Text Reader

Abstract

This application belongs to the field of electronic digital data processing technology, specifically relating to an interpretable hierarchical reasoning multimodal method and system for highway scenarios. The method includes: acquiring multimodal data and preprocessing the multimodal data; performing feature encoding and semantic fusion on the preprocessed multimodal data to form unified features; using the unified features for hierarchical recursive collaborative reasoning, including obtaining the final semantic state through global semantic understanding and long-term reasoning in the high-level part; updating data, performing fine-grained calculations, and supplementary inferences through the low-level part; generating natural language answers through an encoder based on the outputs of the high-level and low-level parts, using the final semantic state as a condition; calculating the contribution ratio of the natural language answers by calculating the interpretation distribution of the high-level and low-level parts; and performing interpretability analysis on the natural language answers, performing causal activation mapping calculations on each generated result to obtain a visual interpretation graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic digital data processing technology, and specifically relates to an interpretable hierarchical reasoning multimodal method and system for highway scenarios. Background Technology

[0002] With the widespread application of multimodal large models in complex intelligent systems such as autonomous driving, they have demonstrated powerful capabilities in perception, understanding, and decision-making. However, these models typically employ a "single-layer, shallow" inference structure, making it difficult to meet the dual requirements of high-level abstract understanding and low-latency response in highway scenarios. On the one hand, current mainstream large models largely rely on a sequential language reasoning approach similar to thought chains (CoT). While this provides some interpretability in task decomposition, it lacks structured modeling capabilities and suffers from weaknesses such as fragile task decomposition, unstable inference paths, and long response times. On the other hand, existing multimodal models rely heavily on traditional visual activation methods for interpretability analysis, failing to fully consider contextual interference between generative token data, leading to biased understanding of model behavior. Therefore, there is an urgent need for a multimodal model architecture for complex highway traffic scenarios, possessing hierarchical reasoning capabilities and an explicit interpretability mechanism at the token data level, to balance accuracy, efficiency, and controllability. Summary of the Invention

[0003] This application is proposed based on the aforementioned needs of the prior art. The technical problem to be solved by this application is to provide an interpretable hierarchical reasoning multimodal method and system for highway scenarios to improve the accuracy and efficiency of understanding complex traffic scenarios.

[0004] To address the above problems, the technical solution provided in this application includes:

[0005] This paper presents an interpretable hierarchical reasoning multimodal method for highway scenarios, comprising: acquiring multimodal data and preprocessing the multimodal data; performing feature encoding and semantic fusion on the preprocessed multimodal data to form unified features; using the unified features for hierarchical recursive collaborative reasoning, including obtaining the final semantic state through global semantic understanding and long-term reasoning in the high-level part; updating the data, performing fine-grained calculations and supplementary inferences in the low-level part; generating natural language responses using an encoder based on the final semantic state; and performing interpretability analysis on the natural language responses, and then performing causal activation mapping calculations on each generated result to obtain a visual interpretation graph.

[0006] Preferably, the acquisition of multimodal data includes acquiring multimodal input data from vehicle-side and road-side sources in a highway scenario. The vehicle-side data includes forward-facing camera images, navigation text, historical trajectories, and vehicle status information, including speed, acceleration, and lane change commands. The road-side data includes roadside units, supplementary image information from monitoring cameras, target detection results, road surface status data, and cross-lane environmental perception information, wherein the road surface status data includes lane markings, traffic congestion, and construction warnings.

[0007] Preferably, the preprocessing of the multimodal data includes time synchronization, coordinate unification and format standardization of all multimodal data.

[0008] Preferably, the step of performing feature encoding and semantic fusion on the preprocessed multimodal data to form a unified feature includes: the preprocessed multimodal data includes image data, text data, and perceptual information data, wherein the image data is used to extract global and local spatial semantics through a visual large language model and to form first data through an image encoder; the text data is converted into context-aware vectors by a pre-trained language model and to form second data through a text encoder; the perceptual information data is used to generate semantic embeddings through a structural encoder and to form third data through a structural data embedder; and the first data, second data, and third data are fused through a cross-modal attention mechanism to form a unified feature.

[0009] Preferably, the unified feature is represented by hierarchical recursive collaborative reasoning as follows: , ,in, This represents the output state of the higher-level component at time t. This indicates the processing procedure at the lower level. This represents the output state of the higher-level component at time t-1. The input is sensed and encoded at time t. This represents the output state of the lower layer at time t. This indicates the processing procedure at the lower level. This represents the output state of the lower layer at time t-1.

[0010] Preferably, the answer generation module uses the final high-level state of the hierarchical reasoning module as a condition and employs a decoder to generate a natural language answer in the following form: ,in, The output is a natural language response. This is a sequence-to-sequence modeling decoding function used to map an abstract semantic vector space into an output sequence that conforms to natural language rules. This represents the output status at the end of the higher-level section. Q represents the output state at the end of the lower-level part, where Q is the question.

[0011] Preferably, after the natural language response undergoes interpretability analysis, a causal activation mapping calculation is performed on each generated result to obtain a visual interpretation diagram. This includes visual analysis of each credential data in the natural language response generation process and refined separation of activation paths between visual and linguistic elements, represented as follows: ,in, The first one obtained after causal activation mapping and filtering is... An interpretable activation graph of each token, reflecting its causal contribution to natural language response generation. This is the activation graph of the original i-th token data. Its context token credential data set, Let be the interference estimation function. For noise reduction filter, This is the scaling factor.

[0012] Preferably, after the natural language response undergoes interpretability analysis, causal activation mapping is calculated for each generated result to obtain a visual interpretation map. This includes identifying early voucher data interference, restoring the true correspondence between the current voucher data and the input image, performing denoising processing, and outputting a visual interpretation map corresponding to each voucher data.

[0013] A hierarchical reasoning multimodal system for highway scenarios is also provided, comprising: an acquisition and preprocessing module for acquiring and preprocessing multimodal data; a fusion module for feature encoding and semantic fusion of the preprocessed multimodal data to form unified features; a hierarchical recursive collaborative reasoning module for using the unified features through hierarchical recursive collaborative reasoning, including obtaining the final semantic state through global semantic understanding and long-term reasoning at a high level, and updating, fine-grained calculation, and supplementary inference through a low level; a natural language answer generation module for generating natural language answers using an encoder based on the final semantic state; and a visualization explanation graph output module for performing causal activation mapping calculation on each generated result after interpretability analysis of the natural language answers to obtain a visualization explanation graph.

[0014] Compared with existing technologies, this application proposes an interpretable hierarchical reasoning multimodal large-scale model structure to address the problems of insufficient reasoning depth, severe contextual interference, and weak interpretability in the complex traffic environment of highways. This structure integrates a "high-low layer cascaded" reasoning architecture with a token data-level explicit visualization interpretation mechanism. The model consists of two collaborative recursive modules: a high-level module responsible for abstracting and modeling global traffic intentions and strategy planning in highway scenarios, possessing the ability to perform low-frequency, slow updates and long-term information integration; and a low-level module providing rapid response and fine-grained execution for detailed tasks such as vehicle control and obstacle avoidance, achieving high-frequency, precise control. The two modules iteratively collaborate through temporal cascading to construct a deep reasoning path. Simultaneously, a TAM-based causal activation mechanism is introduced to remove contextual interference and enhance the saliency of the visualization mapping of generated token data, thereby achieving end-to-end causal interpretable modeling, suitable for decision tracking, fault tracing, and security verification in complex scenarios.

[0015] This invention significantly enhances the application capability and reliability of multimodal large models in real-world complex scenarios such as highways. On one hand, the hierarchical inference structure effectively improves the model's computational depth and task adaptability, enabling it to flexibly handle high-order planning tasks such as high-speed merging, lane changing, and strategic overtaking, while also ensuring low-latency execution control, thus strengthening the model's robustness and generalization ability in dynamic environments. On the other hand, the TAM mechanism enables data-level visualization and causal interference removal of token credentials, significantly improving the transparency of the model's internal inference path. This supports process-level verification and behavior-level auditing of the model's output, providing technical support for building a reliable and controllable intelligent driving system. Furthermore, this method exhibits good compatibility and scalability with existing mainstream multimodal models, and is expected to be widely applied as a general-purpose interpretable module in intelligent transportation, vehicle AI, and other embodied intelligence systems. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings.

[0017] Figure 1 This is a flowchart illustrating the steps of an interpretable hierarchical reasoning multimodal method for highway scenarios in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of the data flow of an interpretable hierarchical reasoning multimodal method for highway scenarios in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] In the description of the embodiments of this application, it should be noted that, unless otherwise explicitly specified and limited, the term "connected" should be interpreted broadly. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0021] Throughout the text, the terms “top,” “bottom,” “above,” “below,” and “on top” refer to the relative positions of components of the device, such as the relative positions of the top and bottom substrates within the device. It is understood that the device is multifunctional and independent of its spatial orientation.

[0022] To facilitate understanding of the embodiments of this application, the following will provide further explanation and description with reference to the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of this application.

[0023] Example 1

[0024] This embodiment provides an interpretable hierarchical reasoning multimodal method for highway scenarios, such as... Figures 1-2 As shown.

[0025] The interpretable hierarchical reasoning multimodal method for highway scenarios includes:

[0026] Acquire multimodal data and preprocess the multimodal data.

[0027] Acquire multimodal input data from both vehicle and roadside units in highway scenarios. Vehicle-side data primarily includes images from forward-facing cameras, navigation text, historical trajectories, and vehicle status information (such as speed, acceleration, and lane change commands). Roadside data includes roadside units (RSUs), supplementary images from surveillance cameras, object detection results, road surface status data (lane lines, traffic congestion, construction warnings), and cross-lane environmental perception information.

[0028] All multimodal data undergoes time synchronization, coordinate unification, and format standardization.

[0029] The preprocessed multimodal data is then subjected to feature encoding and semantic fusion to form a unified feature.

[0030] Multimodal data, including image data, text data, and perceptual data, is input into the encoding module, which comprises multiple feature encoders. Different types of multimodal data are input into their respective feature encoders. Specifically, image data (e.g., images from vehicle cameras and roadside equipment) is processed by a visual Transformer to extract global and local spatial semantics; text data (e.g., navigation instructions and traffic signs) is transformed into context-aware vectors by a pre-trained language model; and map and structured perceptual information (e.g., lane structure, obstacle locations, and traffic flow status) are processed by a structure encoder to generate semantic embeddings. Furthermore, to achieve modality alignment and semantic fusion, the encoding module employs an image encoder (ViT), a text encoder (LLaMA), and a structured data embedder to encode different types of data into unified high-dimensional representation vectors, which are then fused using a cross-modal attention mechanism. Additionally, to adapt to the complex spatial semantic structures in highway scenarios, the module introduces spatial location embedding and semantic label guidance mechanisms during the fusion stage, thereby improving the sensitivity of the encoded features to key elements such as lane topology and obstacle locations.

[0031] The modal features output from each feature encoder are aligned with the semantic fusion module through a cross-modal attention mechanism to form a shared representation vector as a unified feature. This unified feature not only includes the vehicle's own local perspective but also incorporates extended environmental information provided by roadside equipment, providing a more complete and visually advantageous input foundation for subsequent hierarchical inference.

[0032] These unified features are then passed as input to subsequent hierarchical reasoning modules.

[0033] The above-mentioned parts perform feature extraction and alignment on the input multimodal information such as images, text, maps, and sensor data to form a unified semantic embedding representation.

[0034] The unified feature is achieved through hierarchical recursive collaborative reasoning, including obtaining the final semantic state through global semantic understanding and long-term reasoning in the high-level part, and updating data, fine-grained calculation and supplementary inference in the low-level part.

[0035] The unified feature captures abstract semantics, scene intent, and long-term dependency information at a slower frequency through a high-level component, constructing a global reasoning trajectory; and processes local target behavior, dynamic interactions, and short-term dependency details at a faster frequency through a low-level component. The reasoning process unfolds recursively, gradually evolving intermediate semantic states to achieve hierarchical reasoning construction from perception to understanding. Coupling is achieved through a unified state evolution equation, thus simultaneously completing reasoning and causal interpretability output in highway scenarios. This mechanism can effectively simulate the "strategic-tactical" cognitive mechanism in human driving.

[0036] The above process simulates the human reasoning process of "abstract planning and detailed implementation" in traffic decision-making. Specifically, the high-level part is responsible for global semantic understanding and long-term reasoning; the low-level part is responsible for rapid updates, fine-grained calculations, and supplementary inferences. At time... High-level reasoning state Updating via a recursive function is represented as follows:

[0037]

[0038] This represents the output state of the higher-level component at time t. This indicates the processing procedure at the lower level. This represents the output state of the higher-level component at time t-1. The input is the perceived encoding at time t. This inference state is used to model global semantics, such as lane narrowing, traffic accidents, or merging intentions.

[0039] At any moment Low-level reasoning state Then, updates are performed under high-level semantic constraints:

[0040]

[0041] This represents the output state of the lower layer at time t. This indicates the processing procedure at the lower level. This represents the output state of the lower layer at time t-1. The input is sensed and encoded at time t. This represents the output state of the high-level component at time t. This low-level inference state is used to capture fine-grained dynamic information such as changes in adjacent vehicle trajectories, local deceleration behavior, or construction sign detection. Through the above process, a hierarchical inference mechanism for collaborative modeling of "high-level semantics - low-level details" is formed.

[0042] This architecture decouples high-frequency detailed reasoning from low-frequency global planning. The higher-level part controls the global semantic trajectory and task intent, determining whether it is a merging scenario, whether a lane change is needed, etc., while the lower-level part performs specific calculations based on this, including real-time calculations for understanding the current lane status and semantic inference of obstacles. This module does not rely on chained prompt templates and has end-to-end learning and generalization capabilities.

[0043] The above part is based on a high-low recursive structure to carry out multi-stage reasoning. The high-level part is used to model global traffic intentions, complex logical relationships and long-term contextual dependencies, while the low-level part focuses on fine-grained semantic inference and spatiotemporal consistency processing.

[0044] Conditioned by the final semantic state, a natural language answer is generated by an encoder based on the outputs of the high-level and low-level parts. The contribution percentage to the natural language answer is calculated by computing the interpretation distribution of the high-level and low-level parts.

[0045] Once the hierarchical reasoning is complete, the final semantic state generated by the high-level part and the output of the low-level part will be used to drive the language generation module to answer questions posed by the user or triggered by the system. The language generation module generates token data based on the current reasoning state, supporting single-round or multi-round question-and-answer sessions, task explanations, behavioral descriptions, and other highway-related natural language outputs.

[0046] Specifically, the answer generation module uses the final high-level state of the hierarchical reasoning module as a condition and employs a decoder to generate a natural language answer, represented as follows:

[0047]

[0048] in, The output is a natural language response. For sequence-to-sequence modeling decoding functions, This represents the output status at the end of the higher-level section. Q represents the output state at the end of the lower-level part, where Q is the question.

[0049] This module uses the autoregressive language generation model LLaMA, supporting various output types such as question-and-answer generation, traffic scenario descriptions, and abnormal behavior explanations. It is particularly well-suited for the diverse questions users might ask in highway scenarios, such as "Why is the vehicle ahead slowing down?", "Is it permissible to change lanes?", or "Is there any dangerous driving behavior currently occurring?". By combining the semantic state of the inference module, it achieves context-consistent and semantically accurate natural language responses.

[0050] High-level explanations highlight the causal origins of global semantics, such as "the construction area was identified by the RSU camera," while low-level explanations annotate the contribution sources of local dynamics, such as "the deceleration of the vehicle ahead was provided by the onboard forward image token." In real-world highway scenarios, when a user asks "Is there a risk ahead?", the system answers "Construction ahead has caused lane narrowing, posing a potential risk," simultaneously providing both a high-level explanation path (relying on the RSU's perception of construction signs) and a low-level explanation path (relying on the deceleration features of the vehicle ahead). This achieves integrated reasoning and explanation, ensuring the transparency and traceability of the question-and-answer results.

[0051] The contribution percentage for natural language responses is calculated by interpreting the distribution of high-level and low-level components, as shown below:

[0052]

[0053] in Quantitative indicators of causal contribution Attention weights, measuring the input Regarding the significance of the current state, the difference term measures the causal contribution of the input to the state, where... The system obtains the corresponding high-level interpretation result immediately upon generating the inference state. Compared with low-level interpretation results This enables full-chain tracing of causal paths.

[0054] The above process generates responses in natural language based on the reasoning results, covering accurate answers to various questions such as object states, driving intentions, and abnormal behaviors in highway scenarios.

[0055] After performing interpretability analysis on the natural language responses, causal activation mapping is calculated for each generated result to obtain a visual explanatory graph.

[0056] After outputting a natural language response, the system automatically initiates the interpretability analysis module to perform causal activation mapping calculations on each generated token data. This mechanism eliminates interference from earlier token data, restores the true correspondence between the current token data and the input image, and performs noise reduction using a Rank Gaussian Filter. Finally, the system outputs a visual explanation map corresponding to each response token data, achieving explicit explanation at the token data level—location-based and traceable—providing technical support for model behavior transparency, reliability assessment, and scenario auditing. This enables explicit explanation generation at the token data level and contextual causal analysis.

[0057] To enhance the controllability and credibility of large-scale models in real-world traffic systems, this module introduces a TAM mechanism to perform visual analysis of each output token data during the response generation process. This mechanism uses causal reasoning and contextual interference modeling to finely separate the activation paths between vision and language. Formalized as follows:

[0058]

[0059] in, The first one obtained after causal activation mapping and filtering is... An interpretable visual activation graph for each token. This is the activation graph of the original i-th token data. Its context token credential data set, Let be the interference estimation function. For noise reduction filter, The scaling factor is used. This module supports generating data-level visualizations of token credentials, which can intuitively display the visual origin, contextual dependencies, and suspicious activations of each word in a response. This provides strong support for system debugging, fault tracing, and model credibility verification, and is particularly suitable for highway scenarios with regulatory sensitivity or high reliability requirements.

[0060] The above process is based on TAM's causal reasoning mechanism. It performs visual semantic alignment and context redundancy removal interpretation on each output token, realizing the visualization of explicit reasoning paths at the token level and in a hierarchical manner, thereby enhancing the transparency and credibility of the model's behavior.

[0061] In summary, this technical solution employs hierarchical recursive collaborative reasoning, drawing inspiration from the human brain's abstract-detail collaborative mechanism across multiple time scales. The high-level components handle planning and reasoning for slow decisions regarding intent, structure, and strategy in traffic scenarios, while the low-level components handle rapid responses to lane changes and obstacle avoidance. This structure significantly improves the model's response efficiency and robustness in the high-dynamic scenarios of highways, helping to overcome the bottleneck of limited reasoning capabilities in traditional single-layer Transformer models. Furthermore, this structure does not rely on large-scale pre-training or external chained hints, offering advantages such as lightweight design, high efficiency, and ease of deployment, making it particularly suitable for in-vehicle systems with limited computing resources or high real-time requirements. Additionally, by introducing a TAM (token activation map) mechanism, combined with a causal reasoning model, redundancy is removed from the activation paths used to generate token data, enabling clear localization and explicit explanation of specific reasoning processes. This mechanism supports fine-grained diagnosis of complex behaviors such as "abnormal reasoning" and "target misidentification" in highway scenarios, providing strong support for model credibility verification, system security assessment, and decision controllability. Furthermore, the proposed explanation mechanism is compatible with existing mainstream multimodal models and has good scalability, making it widely applicable to various practical scenarios such as autonomous driving and intelligent traffic management.

[0062] Example 2

[0063] This embodiment provides an interpretable hierarchical reasoning multimodal system for highway scenarios, including:

[0064] The acquisition and preprocessing module acquires multimodal data and preprocesses the multimodal data.

[0065] The fusion module performs feature encoding and semantic fusion on the preprocessed multimodal data to form unified features.

[0066] The hierarchical recursive collaborative reasoning module unifies features through hierarchical recursive collaborative reasoning, including obtaining the final semantic state through global semantic understanding and long-term reasoning in the high-level part; and updating data, fine-grained calculation and supplementary inference in the low-level part.

[0067] The natural language response generation module uses an encoder to generate natural language responses based on the final semantic state.

[0068] The visualization and explanation output module performs interpretability analysis on the natural language responses and calculates causal activation mapping for each generated result to obtain a visualization and explanation graph.

[0069] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. An interpretable hierarchical reasoning multimodal method for highway scenarios, characterized in that, include: Acquire multimodal data and preprocess the multimodal data. The preprocessed multimodal data includes image data, text data, and perceptual information data. The preprocessed multimodal data is then subjected to feature encoding and semantic fusion to form unified features; The unified feature is derived through hierarchical recursive collaborative reasoning, including obtaining the final semantic state through global semantic understanding and long-term reasoning at the high level; and updating data, performing fine-grained calculations, and supplementary inferences at the low level. This hierarchical recursive collaborative reasoning of the unified feature is represented as follows: in, This represents the output state of the higher-level component at time t. This indicates the processing procedure at the higher levels. This represents the output state of the higher-level component at time t-1. The input is sensed and encoded at time t. This represents the output state of the lower layer at time t. This indicates the processing procedure at the lower level. This represents the output state of the lower layer at time t-1; Conditioned by the final semantic state, the encoder generates a natural language response based on the outputs of the high-level and low-level parts. The contribution percentage to the natural language response is calculated by evaluating the interpretation distributions of the high-level and low-level parts. The response generation module, conditional by the final high-level state of the hierarchical reasoning module, uses a decoder to generate a natural language response in the following form: ,in, The output is a natural language response. This is a sequence-to-sequence modeling decoding function used to map an abstract semantic vector space into an output sequence that conforms to natural language rules. This represents the output status at the end of the higher-level section. Q represents the output state at the end of the lower-level part, where Q is the question. After interpretability analysis of the natural language responses, causal activation mapping is calculated for each generated result to obtain a visual interpretation diagram. This includes visual analysis of each credential data in the natural language response generation process and fine-grained separation of activation paths between visual and language data, represented as follows: in, The first one obtained after causal activation mapping and filtering is... An interpretable visualization of the activation graph of a token credential data point reflects its causal contribution to natural language response generation. This is the activation graph of the original i-th token data. Its context token credential data set, Let be the interference estimation function. For noise reduction filters, This is the scaling factor.

2. The interpretable hierarchical reasoning multimodal method for highway scenarios according to claim 1, characterized in that, The acquisition of multimodal data includes, The system acquires multimodal input data from both vehicle-side and roadside units in a highway scenario. The vehicle-side data includes forward-facing camera images, navigation text, historical trajectories, and vehicle status information, including speed, acceleration, and lane change commands. The roadside data includes roadside units, supplementary image information from monitoring cameras, target detection results, road surface status data, and cross-lane environmental perception information, including lane markings, traffic congestion, and construction warnings.

3. The interpretable hierarchical reasoning multimodal method for highway scenarios according to claim 1, characterized in that, The preprocessing of multimodal data includes time synchronization, coordinate unification, and format standardization of all multimodal data.

4. The interpretable hierarchical reasoning multimodal method for highway scenarios according to claim 1, characterized in that, The step of performing feature encoding and semantic fusion on the preprocessed multimodal data to form unified features includes: The image data is used to extract global and local spatial semantics through a visual big language model and then to form the first data through an image encoder; Text data is transformed into context-aware vectors by a pre-trained language model and then used to form second data through a text encoder. Perceptual information data is encoded into semantic embeddings by a structural encoder, and then formed into third data by a structural data embedder. The first, second, and third data are fused together using a cross-modal attention mechanism to form a unified feature.

5. The interpretable hierarchical reasoning multimodal method for highway scenarios according to claim 1, characterized in that, After interpretability analysis of the natural language response, causal activation mapping calculation is performed on each generated result to obtain a visual interpretation map. This includes identifying interference from early voucher data, restoring the true correspondence between the current voucher data and the input image, performing noise reduction processing, and outputting a visual interpretation map corresponding to each voucher data.

6. An interpretable hierarchical reasoning multimodal system for highway scenarios, characterized in that, include: The acquisition and preprocessing module acquires multimodal data and preprocesses the multimodal data. The preprocessed multimodal data includes image data, text data, and perceptual information data. The fusion module performs feature encoding and semantic fusion on the preprocessed multimodal data to form unified features; The hierarchical recursive collaborative reasoning module unifies features through hierarchical recursive collaborative reasoning, including obtaining the final semantic state through global semantic understanding and long-term reasoning in the high-level part; and updating data, fine-grained calculation, and supplementary inference in the low-level part. The unified features, through hierarchical recursive collaborative reasoning, are represented as follows: in, This represents the output state of the higher-level component at time t. This indicates the processing procedure at the higher levels. This represents the output state of the higher-level component at time t-1. The input is sensed and encoded at time t. This represents the output state of the lower layer at time t. This indicates the processing procedure at the lower level. This represents the output state of the lower layer at time t-1; The natural language response generation module, conditioned on the final semantic state, uses an encoder to generate natural language responses and calculates the contribution percentage to the natural language response by evaluating the interpretation distribution of the high-level and low-level parts. The response generation module, conditioned on the final high-level state of the hierarchical reasoning module, uses a decoder to generate natural language responses in the following form: ,in, The output is a natural language response. This is a sequence-to-sequence modeling decoding function used to map an abstract semantic vector space into an output sequence that conforms to natural language rules. This represents the output status at the end of the higher-level section. Q represents the output state at the end of the lower-level part, where Q is the question. The visualization and explanation output module performs interpretability analysis on the natural language responses, calculates causal activation mappings for each generated result, and obtains a visualization and explanation diagram. This includes visual analysis of each piece of evidence data in the natural language response generation process and fine-grained separation of activation paths between visual and linguistic elements, represented as follows: in, The first one obtained after causal activation mapping and filtering is... An interpretable visualization of the activation graph of a token credential data point reflects its causal contribution to natural language response generation. This is the activation graph of the original i-th token data. Its context token credential data set, Let be the interference estimation function. For noise reduction filters, This is the scaling factor.

Citation Information

Patent Citations

  • Visual question-answering method and system based on multi-modal hierarchical structure representation and alignment

    CN117540810A

  • Interpretable visual positioning method based on multi-modal semantics

    CN117911786A