Cockpit application dynamic scheduling and scene adaptation system based on multi-modal intent fusion

By acquiring multimodal interaction data in the smart cockpit, performing semantic understanding and hierarchical arbitration, generating structured intent units, and performing spatiotemporal linkage matching, the problem of semantic information loss under multimodal conflict is solved, and efficient interaction in the smart cockpit is achieved.

CN122264468APending Publication Date: 2026-06-23WUHAN CHELING ZHILIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing technologies typically treat multimodal input conflicts as a single-choice problem, resulting in the complete discarding of semantic information in suppressed or abandoned modalities, which fails to meet the user's complex intent requirements.

Method used

By acquiring multimodal interaction data within the cockpit, semantic understanding is performed to generate structured intent units, the degree of conflict is quantified and graded arbitration is conducted, an execution instruction set and an intent cache queue to be activated are generated, and the suppressed intent units are activated by spatiotemporal linkage matching.

Benefits of technology

It achieves the preservation and delayed utilization of semantic information in multimodal conflict scenarios, improves the interactive experience of the smart cockpit, and ensures the timely execution of the user's main intent and the comprehensive response to complex intents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122264468A_ABST
    Figure CN122264468A_ABST
Patent Text Reader

Abstract

The application discloses a cockpit application dynamic scheduling and scene adaptation system based on multi-modal intention fusion, belongs to the technical field of intelligent chips, machine learning and dynamic control, and solves the problem that semantic information is completely discarded in a multi-modal conflict scene in a traditional scheme by introducing a dynamic scheduling and delay activation mechanism of multi-modal intention. The application can intelligently identify and cache the suppressed secondary intention, and reactivate the intention when the vehicle drives to a specific space-time condition. Therefore, not only is the timely execution of the user's main intention ensured, but also a comprehensive response to the user's composite intention is realized, the intelligent level of intelligent cockpit interaction and the user experience are improved, and the loss of semantic information is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of intelligent chips, machine learning and dynamic control technology, and in particular to a cockpit application dynamic scheduling and scenario adaptation system based on multimodal intent fusion. Background Technology

[0002] Intelligent cockpit multimodal interaction technology integrates multiple input modalities such as voice, gestures, and gaze to provide users with a natural and efficient human-computer interaction experience. Typical solutions include a dynamic allocation system based on priority scores, which selects the optimal execution modality and suppresses other modalities by calculating the real-time priority scores of each modality; or a system that uses semantic conflict coefficients to determine the modality with the highest confidence when the conflict coefficient is greater than a threshold, and performs weighted fusion of modalities when the conflict coefficient is less than or equal to the threshold.

[0003] However, when dealing with multimodal input conflicts, related technologies treat conflict handling as a "selection" problem—either selecting the modality with the highest priority for execution or selecting the modality with the highest confidence for output. This approach results in the complete discarding of semantic information in suppressed or abandoned modalities, failing to meet the user's complex intent needs. For example, when a driver's voice command is "navigate to the company" while their gaze is fixed on a roadside restaurant for an extended period, the technology only executes the navigation command, permanently discarding the "roadside attention" intent inherent in the gaze modality, and failing to proactively provide reminders as the vehicle approaches the restaurant.

[0004] Therefore, how to preserve and delay the use of semantic information in multimodal conflict scenarios has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0005] This application provides a cockpit application dynamic scheduling and scenario adaptation system based on multimodal intent fusion, the technical solution of which is as follows: On the one hand, a cockpit application dynamic scheduling and scenario adaptation system based on multimodal intent fusion is provided. The system includes a processor and a memory, and the processor is configured to perform the following steps: Acquire multimodal interaction data collected in the cockpit, perform semantic understanding on the multimodal interaction data, and generate multiple structured intent units; The conflict level of the multiple structured intent units and the cockpit scene knowledge graph is quantified to obtain conflict level scores and scene semantic association information; The multiple structured intent units, the conflict degree score, and the scene semantic association information are subjected to hierarchical arbitration to generate a cockpit application execution instruction set and an intent cache queue to be activated. The hierarchical arbitration determines the main graph unit and the suppressed intent unit according to the conflict degree score, adds the instruction corresponding to the main graph unit to the execution instruction set, and adds the suppressed intent unit and its associated spatiotemporal triggering condition to the intent cache queue to be activated. Based on the intent cache queue to be activated, the scene semantic association information, and the real-time collected vehicle location information, the suppressed intent units in the intent cache queue to be activated are spatiotemporally linked for matching, and activation instructions are generated and added to the execution instruction set. Attached Figure Description

[0006] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0007] Figure 1 This is a flowchart of a method for dynamic scheduling and scenario adaptation of cockpit applications based on multimodal intent fusion, provided in an embodiment of this application. Figure 2 This is a flowchart of another method for dynamic scheduling and scenario adaptation of cockpit applications based on multimodal intent fusion provided in this application embodiment; Figure 3 This is a flowchart of another method for dynamic scheduling and scenario adaptation of cockpit applications based on multimodal intent fusion provided in the embodiments of this application; Figure 4 This is a flowchart of another method for dynamic scheduling and scenario adaptation of cockpit applications based on multimodal intent fusion provided in the embodiments of this application; Figure 5 This is a flowchart of another method for dynamic scheduling and scenario adaptation of cockpit applications based on multimodal intent fusion provided in the embodiments of this application. Detailed Implementation

[0008] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0009] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.

[0010] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0011] In related technologies, intelligent cockpit multimodal interaction technologies typically treat multimodal input conflicts as a single "choice" problem, executing only the modal command with the highest priority or confidence. This approach results in the complete discarding of semantic information contained in suppressed or abandoned modalities, failing to meet the complex and compound intentions of users. For example, a driver's secondary intentions (such as attention to waypoints) are ignored because they conflict with the primary intentions and cannot be utilized in subsequent scenarios, thus limiting the intelligence level of the interaction system.

[0012] To address this, this application proposes a cockpit application dynamic scheduling and scenario adaptation system based on multimodal intent fusion. The system includes a processor and memory; see [link to relevant documentation]. Figure 1 The processor is configured to perform the following steps: 101. Acquire multimodal interaction data collected in the cockpit, perform semantic understanding on the multimodal interaction data, and generate multiple structured intent units.

[0013] 102. Quantify the degree of conflict for multiple structured intent units and cockpit scene knowledge graphs to obtain conflict degree scores and scene semantic association information.

[0014] 103. Perform hierarchical arbitration on multiple structured intent units, conflict degree scores and scene semantic association information to generate a cockpit application execution instruction set and an intent cache queue to be activated. The hierarchical arbitration determines the main graph unit and the suppressed intent unit according to the conflict degree score, adds the instruction corresponding to the main graph unit to the execution instruction set, and adds the suppressed intent unit and its associated spatiotemporal triggering conditions to the intent cache queue to be activated.

[0015] 104. Based on the intent cache queue to be activated, scene semantic association information, and real-time vehicle location information, perform spatiotemporal linkage matching on the suppressed intent units in the intent cache queue to be activated, generate activation instructions, and add them to the execution instruction set.

[0016] Thus, this application achieves the preservation and delayed utilization of semantic information in multimodal conflict scenarios, thereby improving the interactive experience of the smart cockpit.

[0017] For ease of understanding, the following explains some key terms in this embodiment: Multimodal interaction data refers to the raw data collected from user interactions with the vehicle in a smart cockpit environment through various sensors, such as voice, gaze, gestures, touch, vehicle dynamics, and biometric data. This data collectively reflects the user's intent and the current state of the cockpit environment.

[0018] A structured intent unit refers to user intent information represented in a unified and standardized format after semantic understanding of the raw multimodal interaction data. Each structured intent unit typically includes an intent description, source modality, and initial confidence level, facilitating subsequent conflict analysis and scheduling.

[0019] The cockpit scene knowledge graph is a structured knowledge base containing entities, concepts, and their relationships within the cockpit. It represents various elements in the cockpit environment (such as devices, functions, locations, and user states) and their semantic, temporal, and command conflict relationships through nodes and edges, providing rich contextual information for quantifying intent conflicts and adapting to different scenarios.

[0020] The conflict score is a numerical metric that quantifies the degree of conflict between multiple structured intent units. This score takes into account both semantic contradictions and operational mutual exclusions between intents, and is used to assess the potential conflicts or risks that may arise when different intents are executed simultaneously.

[0021] Scene semantic association information refers to information extracted from the cockpit scene knowledge graph that describes the semantic relationships between structured intent units or between intent units and the cockpit environment. This can include device associations, operation timing associations, and instruction mutual exclusion relationships, providing a basis for intent arbitration and trigger condition generation.

[0022] Hierarchical arbitration refers to a decision-making process that categorizes intent processing into different levels based on the degree of conflict between intents. For intents with low conflict levels, semantic fusion can be performed. For intents with high conflict levels, priority determination is required to identify the intent graph unit and the suppressed intent unit.

[0023] The cockpit application execution command set is a collection of cockpit application control commands generated and prepared for immediate execution by the system based on the arbitration result. These commands directly drive various functions and applications within the cockpit.

[0024] The pending intent cache queue is a queue used to temporarily store intent units that were suppressed during the hierarchical arbitration process. These suppressed intent units are not discarded, but are associated with specific spatiotemporal triggering conditions, waiting to be reactivated at an appropriate time in the future.

[0025] A concept unit refers to the intent unit that the system determines to be executed with priority during the hierarchical arbitration process. Its corresponding instruction will be added to the cockpit application execution instruction set.

[0026] Suppressed intent units refer to intent units that are temporarily not executed during the hierarchical arbitration process due to conflicts with intent graph units or lower priority. These intent units will be cached and activated again when their spatiotemporal triggering conditions are met.

[0027] Spatiotemporal triggering conditions refer to the conditions associated with a suppressed intent unit that determine when and where the intent unit can be activated. These typically include a trigger location and a trigger time window, ensuring that the suppressed intent is reused in the most relevant context.

[0028] Spatiotemporal linkage matching refers to the process of matching suppressed intent units in the intent cache queue with real-time collected vehicle location information, current status information of drivers and passengers, and scene semantic association information. When the matching degree reaches a certain threshold, the suppressed intent unit will be activated.

[0029] The technical solution of this application will be described in detail below with reference to specific embodiments.

[0030] The system acquires multimodal interaction data collected within the cockpit and performs semantic understanding on this data to generate multiple structured intent units. Specifically, multimodal interaction data can be directly collected by various sensors within the cockpit, such as microphones collecting voice data, cameras collecting gaze and gesture data, and touchscreens collecting touch data. This raw data is input into the semantic parsing module. This module analyzes the data for each modality independently, identifying the user intent contained within. For example, voice data is processed using speech recognition and natural language understanding technologies to extract keywords and intent types of user commands. Gesture data is processed using image recognition technology to identify the meaning of user gestures. After completing semantic understanding, the identified intent information is encapsulated in a specific format to form structured intent units. Each intent unit can include an intent description, the source modality, and an initial confidence level. For example, when a user says "turn on the air conditioning," an intent unit can be generated, representing the intent to "turn on the air conditioning," originating from the voice modality. When the user simultaneously points to the car window, another intent unit can be generated, representing the intent to "point to the car window," originating from the gesture modality.

[0031] The conflict level of multiple generated structured intent units and the cockpit scene knowledge graph is quantified to obtain conflict level scores and scene semantic association information. The cockpit scene knowledge graph can be pre-built, storing the relationships between various devices, functions, locations, and user behavior patterns within the cockpit. When quantifying conflict, the target objects and operations involved in each intent unit can be analyzed and compared with relevant information in the knowledge graph. For example, when one intent is "open the window" and another is "close the window," a direct operational conflict between these two intents can be identified. Based on a set of rules, it can be assessed whether there are semantic contradictions or operational mutual exclusions between these intents. For example, the rules can define "navigation" and "play music" as low conflict, while "open the window" and "close the window" as high conflict. In this way, a conflict level score can be calculated for each pair or group of intent units, reflecting the intensity of conflict that may arise when they are executed simultaneously. At the same time, the association information between intent units or between intent units and the cockpit environment can be extracted from the knowledge graph. For example, whether two intents point to the same device or whether they have a temporal sequence, this information is used as scene semantic association information.

[0032] A hierarchical arbitration process is performed on multiple structured intent units, conflict severity scores, and scene semantic association information to generate a cockpit application execution instruction set and a queue of intents to be activated. The hierarchical arbitration process determines how to handle these intents based on their conflict severity scores. For example, when the conflict severity score is low, these intents can be considered to be executed in parallel or combined. When the conflict severity score is high, the system needs to make more complex decisions. Specifically, based on the arbitration strategy, one conflicting intent unit can be selected as the main graph unit, and the remaining intent units can be identified as suppressed intent units. For example, the intent with the highest initial confidence can be simply selected as the main graph unit, or the intent with the highest safety level can be selected. The cockpit application control instructions corresponding to the main graph unit are immediately added to the cockpit application execution instruction set, awaiting execution. Suppressed intent units are not executed immediately but are cached by the system. To ensure that these suppressed intent units can be reused at an appropriate time in the future, a spatiotemporal trigger condition is generated and associated with each suppressed intent unit. This condition can be a geographical location, such as "when the vehicle arrives at a specific location," or a time period, such as "within the next five minutes." These suppressed intent units and their associated spatiotemporal triggering conditions are stored in the intent cache queue to be activated.

[0033] Based on the pending intent cache queue, scene semantic association information, and real-time vehicle location information, the system performs spatiotemporal matching on suppressed intent units in the pending intent cache queue, generates activation commands, and adds them to the execution command set. The system continuously monitors the vehicle's real-time location information, such as the current geographic coordinates obtained via GPS. Simultaneously, it retrieves suppressed intent units and their associated spatiotemporal trigger conditions from the pending intent cache queue. The real-time vehicle location is compared with the trigger location of the suppressed intent unit. For example, if the trigger condition for a suppressed intent is "when the vehicle approaches a restaurant," the system determines whether the current vehicle location has entered the specific area of ​​that restaurant. When the real-time location information meets the spatiotemporal trigger condition of the suppressed intent unit, the system considers the suppressed intent unit to have met the activation condition. At this point, a corresponding activation command is generated, such as "remind the user to pay attention to the restaurant," and this activation command is added to the cockpit application execution command set, allowing previously suppressed intents to be executed at an appropriate time, achieving delayed utilization of intents.

[0034] This application addresses the problem of complete loss of semantic information in multimodal conflict scenarios in traditional solutions by introducing a dynamic scheduling and delayed activation mechanism for multimodal intents. It intelligently identifies and caches suppressed secondary intents, reactivating them when the vehicle reaches specific spatiotemporal conditions. This not only ensures the timely execution of the user's primary intents but also achieves a comprehensive response to the user's complex intents, improving the intelligence level and user experience of the smart cockpit interaction while avoiding the loss of semantic information.

[0035] In some of the embodiments described above in this application, it is proposed to acquire multimodal interaction data and perform semantic understanding to generate structured intent units to support subsequent conflict quantification and arbitration. However, in this process, if the data acquisition is incomplete or the semantic parsing is unstructured, the intent information may be incomplete, and the details of multimodal intents cannot be effectively captured, thereby affecting conflict handling and delayed use of intents.

[0036] To address this, this application further proposes a method for acquiring multimodal interaction data collected within the cockpit, performing semantic understanding on this multimodal interaction data, and generating multiple structured intent units. See [link to relevant documentation]. Figure 2 ,include: 201. Collect raw multimodal data streams from within the cockpit, including voice data, gaze data, gesture data, touch data, vehicle dynamics data, biometric data, and external environment data.

[0037] 202. Perform feature extraction and semantic parsing on the original multimodal data stream to obtain the intent feature vector set corresponding to each modality.

[0038] 203. Based on the intent feature vector set and intent structuring rules, instantiate and generate the multiple structured intent units, wherein each structured intent unit includes intent type, target object, modality source, original confidence, timestamp, security level and semantic vector.

[0039] For example, raw multimodal data streams are collected within the cockpit, including voice data, gaze data, gesture data, touch data, vehicle dynamics data, biometric data, and external environmental data. Raw multimodal data streams refer to the unprocessed collection of various data types acquired in real-time or near real-time from the intelligent cockpit environment. Their purpose is to comprehensively capture the interactive behaviors and physiological states of the driver and passengers, as well as information about vehicle operation and the external environment, providing rich and comprehensive raw input for subsequent intent recognition. For instance, this can be achieved through various sensors integrated within the cockpit, such as microphone arrays for voice data, infrared cameras or eye trackers for gaze data, depth sensors or inertial measurement units (IMUs) for gesture data, touchscreen sensors for touch data, vehicle buses (such as the CAN bus) for vehicle dynamics data (such as vehicle speed and steering angle), biosensors (such as heart rate monitors and EEG caps) for biometric data, and external cameras, radar, GPS, etc., for external environmental data. In addition, it can also integrate with in-vehicle infotainment systems, driver assistance systems (ADAS), and vehicle networking platforms through data interfaces to acquire and aggregate data streams from different subsystems in real time, forming a unified raw multimodal data stream.

[0040] Feature extraction and semantic parsing are performed on the original multimodal data stream to obtain the intent feature vector set corresponding to each modality. This step aims to transform the original, heterogeneous multimodal data into a unified, computable intent feature vector representation, thereby eliminating the heterogeneity between modalities and extracting semantic information highly relevant to the user's intent. This is crucial for subsequent intent structuring and conflict resolution. For example, specialized feature extraction algorithms can be used for different modalities. For speech data, acoustic models (such as MFCC and FBank features) combined with ASR (Automatic Speech Recognition) technology can be used to extract textual semantic features. For gaze data, eye-tracking algorithms can be used to extract features such as gaze point and gaze duration. For gesture data, skeleton extraction and posture recognition algorithms can be used to extract gesture type and trajectory features. Then, deep learning models (such as Transformer and LSTM) are used to perform semantic parsing on these features to generate intent feature vectors for each modality. Alternatively, an end-to-end multimodal fusion model can be used, which directly inputs the raw multimodal data into a unified neural network architecture. This network contains a feature extraction layer and a semantic parsing layer, which automatically learns and outputs the intent feature vector set corresponding to each modality.

[0041] Based on the intent feature vector set and intent structuring rules, multiple structured intent units are instantiated. Each structured intent unit includes intent type, target object, modality source, original confidence level, timestamp, security level, and semantic vector. This step transforms the abstract intent feature vector into standardized, machine-understandable structured intent units, providing a unified data format and rich information dimensions for subsequent intent conflict quantification, arbitration, and delayed activation. Through structuring, various aspects of user intent can be clearly expressed, such as intent category, target object, source, reliability, occurrence time, importance, and semantic connotation. For example, a set of intent structuring rules can be predefined, including the definition and value range of fields such as intent type (e.g., navigation, entertainment, air conditioning control), target object (e.g., destination, song, temperature), modality source (e.g., voice, gesture), and security level (e.g., high, medium, low). Then, using an intent classifier and entity recognition model, the intent type and target object are parsed from the intent feature vector set. These are then combined with modality recognition results, model output confidence, data acquisition time, and a pre-defined security level mapping table to populate the fields of the structured intent unit. Semantic vectors can be generated by using a pre-trained word embedding model or sentence embedding model to transform textual information such as intent type and target object into high-dimensional vector representations. Alternatively, template matching or knowledge graph mapping methods can be used to input the intent feature vector set into the intent recognition module, identify the initial intent concept, and populate the fields of the structured intent unit based on a pre-defined intent template and knowledge graph query.

[0042] Through the aforementioned technical solution, this application ensures comprehensive perception of the driver's and passengers' intentions by collecting raw multimodal data streams within the cockpit, including voice data, gaze data, gesture data, touch data, vehicle dynamic data, biometric data, and external environmental data. This avoids the problem of incomplete intention understanding caused by the omission of single-modal information. Feature extraction and semantic parsing are performed on this raw multimodal data stream to obtain intention feature vector sets corresponding to each modality. This transforms the heterogeneous raw data into a unified vector representation rich in semantic information, laying the foundation for subsequent intention processing. Based on this intention feature vector set and intention structuring rules, multiple structured intention units are instantiated. Each structured intention unit includes intention type, target object, modality source, original confidence level, timestamp, security level, and semantic vector. This structured representation not only clearly defines the various attributes of the intention, such as its type, target object, source, reliability, occurrence time, importance, and deep semantics, but also provides a standardized, rich, and easily processed information carrier for subsequent conflict quantification, graded arbitration, and spatiotemporal linkage matching of suppressed intention units. In this way, even in multimodal conflict scenarios, suppressed intent information can be completely preserved in a structured form, carrying the necessary spatiotemporal triggering conditions. This provides a solid data foundation for realizing delayed use of intent and meeting users' complex intent needs, solving the problem of intent information loss in traditional solutions.

[0043] In some of the solutions described above in this application, feature extraction and semantic parsing of multimodal data are proposed to generate structured intent units. However, in this process, the extracted features may be inaccurate due to the potential for temporal asynchrony or reliability issues caused by the driving context when using voice data, gaze data, and gesture data. This affects the quality of the generated intent units.

[0044] To address this, this application further proposes feature extraction and semantic parsing of the original multimodal data stream to obtain intent feature vector sets corresponding to each modality. Specifically, this includes: temporally aligning the speech data, gaze data, and gesture data to obtain a spatiotemporally synchronized multimodal feature sequence. Based on this spatiotemporally synchronized multimodal feature sequence, a cross-modal attention mechanism is used to extract and fuse semantic features, resulting in speech intent feature vectors, gaze intent feature vectors, and gesture intent feature vectors. Based on the vehicle dynamic data and the biometric data, dynamic weight adjustments are made to the speech intent feature vector, gaze intent feature vector, and gesture intent feature vector, respectively, to obtain the intent feature vector sets corresponding to each modality.

[0045] "Temporal alignment" refers to synchronizing data streams from different modalities (such as voice, gaze, and gesture) along the time dimension to ensure that these data have a consistent reference point in time. Its purpose is to eliminate data asynchrony caused by differences in sensor acquisition frequencies, transmission delays, or processing speeds, providing an accurate and unified time benchmark for subsequent multimodal fusion, thereby avoiding semantic misjudgments caused by temporal misalignment. For example, precise timestamps can be assigned to each modal data stream, and before data fusion, techniques such as interpolation, resampling, or sliding windows can be used to map all modal data onto a unified time axis, ensuring that at any given time point, each modal data corresponds accurately. Alternatively, Dynamic Time Warping (DTW) algorithms can be used. These algorithms can handle the non-linear scaling of different modal data sequences along the time axis, finding the optimal matching path between them, thus achieving more flexible and robust temporal alignment.

[0046] Cross-modal attention is a deep learning technique whose core idea is to allow models to selectively focus on and integrate information from other modalities that is most relevant to the current task when processing information from a specific modality. This helps models capture deep semantic connections and complementarities between different modalities, thereby more effectively extracting semantically rich features during the fusion process and improving the accuracy of understanding user intent. For example, an attention network based on the Transformer architecture can be constructed, where the feature sequence of each modality serves as the query, and the feature sequences of other modalities serve as the key and value. Attention weights are then calculated to aggregate information from different modalities, generating a fused modality-specific feature vector. Alternatively, a multi-head attention mechanism can be used, allowing the model to learn different cross-modal association patterns from different representation subspaces. The outputs of these patterns are then concatenated or weighted to obtain more comprehensive fused semantic features.

[0047] "Dynamic weight correction" refers to assigning different importance or reliability weights to the intent feature vectors of different modalities based on real-time changes in the external environment (such as vehicle dynamic data) and internal states (such as biometric data). Its purpose is to adapt to complex driving environments and user states, reduce the impact of unreliable modalities, and increase the contribution of reliable modalities, thereby improving the accuracy and robustness of intent recognition. For example, a predictive model can be used, taking vehicle dynamic data (such as vehicle speed, acceleration, steering angle, etc.) and biometric data (such as driver heart rate, eye movement trajectory, facial expressions, etc.) as input, and outputting dynamically corrected weights for voice, gaze, and gesture intent feature vectors. In another implementation, a rule-based weight correction strategy can be established, pre-setting default weights for each modality under different driving scenarios (such as congestion, highways, parking) and driver / passenger states (such as fatigue, distraction, focus), and triggering corresponding weight adjustment rules based on real-time data.

[0048] Through the above technical solutions, this application can solve the problems of temporal asynchrony and reliability fluctuations in the multimodal feature extraction process, and improve the generation quality of intent units. For example, by temporally aligning voice data, gaze data, and gesture data, precise synchronization of different modal data in time is ensured, eliminating potential semantic biases caused by acquisition or processing delays, and laying a solid foundation for subsequent feature fusion. Employing a cross-modal attention mechanism enables the intelligent capture and integration of complementary information between modalities when extracting fused semantic features, thereby generating more representative and discriminative voice intent feature vectors, gaze intent feature vectors, and gesture intent feature vectors, avoiding information isolation and one-sided understanding. Furthermore, dynamic weight correction of these intent feature vectors based on vehicle dynamic data and biometric data allows for adaptive adjustment of the reliability weights of each modality according to real-time driving situations (such as vehicle driving status) and the state of the driver and passengers (such as focus level and fatigue). For example, when the driver is highly focused or the vehicle is driving smoothly, the weights of each modality may be relatively balanced. When the driver is distracted or the vehicle is in complex operating conditions, the weight of susceptible modalities (such as gestures) can be reduced while the weight of relatively stable modalities (such as voice) can be increased. This effectively suppresses the negative impact of unreliable input on intent recognition, ensuring that the system can extract accurate and robust intent feature vector sets in various complex and dynamic scenarios. These high-quality intent feature vector sets will directly improve the instantiation accuracy of subsequent structured intent units, providing more reliable input for the dynamic scheduling and scenario adaptation system of cockpit applications, thereby achieving a more accurate and intelligent cockpit interaction experience.

[0049] In some of the solutions mentioned above in this application, dynamic weight correction of multimodal intent feature vectors based on vehicle dynamic data and biometric data is proposed to improve the reliability of intent feature vectors. However, in this process, since the impact of the urgency of the driving situation and the driver's concentration level on the reliability of intent is not fully considered, the intent feature vector may not accurately reflect the real intent in subsequent conflict handling, thereby erroneously discarding valuable semantic information and increasing the risk that the user's complex intent needs will not be met.

[0050] To address this, this application proposes a method for dynamically weighting and correcting the voice intent feature vector, gaze intent feature vector, and gesture intent feature vector based on vehicle dynamic data and biometric data, thereby obtaining the intent feature vector set corresponding to each modality. Specifically, the method includes: extracting driving context features based on the vehicle dynamic data, whereby these features characterize the urgency of the current driving state; determining a reliability correction factor for each modality based on the biometric data, whereby this reliability correction factor is positively correlated with the driver's level of focus; fusing the driving context features with the reliability correction factor for each modality to obtain dynamic correction weights corresponding to each modality; and applying these dynamic correction weights to the voice intent feature vector, gaze intent feature vector, and gesture intent feature vector, respectively, to obtain the intent feature vector set corresponding to each modality.

[0051] Based on the vehicle's dynamic data, driving situation features are extracted to characterize the urgency of the current driving state. The purpose of these features is to quantify the impact of the current driving environment on the driver's cognitive load and operational precision, thus providing crucial information for assessing the reliability of subsequent intentions. For example, in high-urgency situations such as high-speed driving, traffic congestion, or emergency braking, the driver's attention may be highly focused on the driving task, potentially interfering with or reducing the reliability of other modal intention expressions. In one implementation, parameters such as vehicle speed, acceleration, braking intensity, steering angle, and distance to the vehicle in front can be analyzed from the vehicle's dynamic data. Combined with a pre-defined driving situation model, the current driving state can be categorized into different levels, such as "low urgency," "medium urgency," or "high urgency," and each level can be assigned a corresponding numerical or vector representation as a driving situation feature. In another implementation, vehicle sensor data can be used to acquire information about the surrounding environment. Combined with high-precision map data, current road type, traffic flow, and weather conditions can be identified. A deep learning model can then extract feature vectors reflecting the urgency of the driving from these multi-source data sources.

[0052] Based on this biometric data, a reliability correction factor is determined for each modality. This reliability correction factor is positively correlated with the driver's level of focus. The reliability correction factor is a numerical value used to quantify the reliability of the driver's expressed intentions in a specific modality. Its purpose is to assess the impact of factors such as the driver's current mental state and cognitive load on the input quality of different interaction modalities, thereby dynamically adjusting the weight of each modality's intention. For example, when the driver's focus is high, the accuracy of their voice commands or gestures is usually higher, and the correction factor should be larger. Conversely, when the driver is distracted or fatigued, the reliability of their interaction intentions may decrease, and the correction factor should be smaller. In one implementation, physiological indicators such as eye movement data, heart rate variability, and electroencephalograms in the biometric data can be analyzed, combined with a machine learning model, to assess the driver's level of focus in real time and map it to a value between 0 and 1 as the reliability correction factor. In another implementation, facial expression recognition, head posture estimation and other technologies can be used in combination with the interaction history data of drivers and passengers to determine whether drivers and passengers are in a state of fatigue, distraction, or emotional fluctuation, and reliability correction factors for each modality can be preset or dynamically calculated based on these states.

[0053] The driving scenario features are fused with the reliability correction factors for each modality to obtain dynamic correction weights for each modality. The fusion of driving scenario features and reliability correction factors aims to comprehensively consider both the external environment and the internal state of the occupants, generating dynamic correction weights that fully reflect the current credibility of each modality's intent. These weights are directly used to adjust the original intent feature vector, ensuring more accurate understanding and processing of multimodal intents across different driving scenarios and occupant states. One implementation method uses a weighted summation approach, for example, dynamic correction weight = α × driving scenario feature value + β × reliability correction factor value, where α and β are preset weight coefficients that can be optimized through experimentation or machine learning methods. Another implementation method uses a neural network model for fusion, taking the driving scenario feature vector and the reliability correction factors for each modality as input, and learning and outputting the dynamic correction weights for each modality through a small fully connected neural network or attention mechanism network.

[0054] The dynamic correction weights are applied to the voice intent feature vector, the gaze intent feature vector, and the gesture intent feature vector, respectively, to obtain the intent feature vector set corresponding to each modality. Applying the calculated dynamic correction weights to the original intent feature vector is a key step in achieving dynamic adjustment of intent reliability. Through this "application" mechanism, the expressive strength of high-reliability intents can be effectively enhanced while weakening the influence of low-reliability intents. This allows the corrected intent feature vector set to more accurately reflect the driver's true intent, providing a more reliable input for subsequent intent arbitration and fusion. In one implementation, vector multiplication or element-wise multiplication can be used; for example, the corrected intent feature vector = dynamic correction weight × original intent feature vector. In another implementation, a gating mechanism can be used to use the dynamic correction weights as a gating signal to control the information flow of the original intent feature vector, thereby achieving dynamic adjustment of the intent feature vector.

[0055] The aforementioned technical solution identifies the urgency of the current driving state and adjusts the reliability assessment of intentions accordingly, preventing the erroneous suppression of important intentions due to the driver's focus on driving. Simultaneously, based on biometric data, it can assess the driver's and passengers' level of concentration in real time and generate reliability correction factors for each modality, effectively filtering out "noise" intentions caused by poor driver / passenger condition and improving the accuracy of intention recognition. By fusing driving context features with reliability correction factors, dynamic correction weights for each modality are generated and applied to the voice intention feature vector, gaze intention feature vector, and gesture intention feature vector. This process comprehensively considers the dual impact of external environment and internal state on intention reliability, achieving adaptive weight adjustment. Through this dynamic correction, the representation of intention feature vectors can be more accurately optimized, strengthening high-reliability, high-priority intentions while appropriately suppressing low-reliability, low-priority intentions. This ensures that the true intentions of drivers and passengers can be more accurately captured and preserved during subsequent semantic understanding and intention arbitration, reducing the risk of valuable semantic information being erroneously discarded. By dynamically adjusting the weights, a more reliable and accurate set of intent feature vectors can be generated. This provides high-quality input for subsequent intent conflict quantification, hierarchical arbitration, and spatiotemporal linkage matching, improving the ability to understand the complex intents of drivers and passengers and the accuracy of scenario adaptation, thereby better meeting user needs.

[0056] In some of the embodiments described above in this application, a structured intent unit is proposed to represent user intent based on the intent feature vector set. However, in its implementation, the intent unit may lack sufficient semantic depth and security considerations, which may lead to the inability to accurately evaluate the correlation and priority between intents during subsequent conflict quantification, thereby affecting the accuracy of conflict arbitration and the integrity of intent information.

[0057] In response, this application further proposes a method for instantiating and generating multiple structured intent units based on the intent feature vector set and intent structuring rules, specifically including: For each modal intent feature vector in the intent feature vector set, the corresponding intent type and target object are obtained by parsing based on the intent classifier.

[0058] Based on the intent type and target object obtained from the parsing, the semantic vector is generated through a semantic embedding model.

[0059] The security level is determined based on the intent type, the modality source, and the security level mapping table.

[0060] Based on the original confidence level and security level in the intent feature vector set, the original confidence level field is initialized to generate the corresponding structured intent unit, thus obtaining the multiple structured intent units.

[0061] For each modal intent feature vector in the intent feature vector set, the corresponding intent type and target object are obtained through parsing using an intent classifier. This step aims to extract the intent type and specific operation object with clear semantics from the abstract intent feature vectors. The intent type is a classification of the user's intent, such as "play music," "navigate," or "adjust the temperature." The target object is the specific entity that the intent is applied to, such as "Jay Chou's song," "company," or "22 degrees." The intent classifier can be implemented in various ways. For example, it can be a pre-trained deep learning model, such as a classifier based on the Transformer architecture, which can receive intent feature vectors as input and output the intent type and its corresponding confidence score, while extracting the target object from the text or semantic information associated with the feature vectors using Named Entity Recognition (NER) technology. Another implementation method is to use a combination of rule-based and machine learning approaches, identifying the intent type and target object through pre-defined keyword matching and syntactic analysis rules, and supplementing it with traditional machine learning models such as Support Vector Machines (SVM) or Random Forests for auxiliary classification and entity extraction.

[0062] Based on the parsed intent type and target object, a semantic vector is generated using a semantic embedding model. The purpose of this step is to transform the discrete intent type and target object into a continuous, dense vector representation, facilitating subsequent semantic similarity calculation and quantitative analysis. Semantic embedding models can capture the semantic relationships between words or phrases. For example, word embedding models such as Word2Vec and GloVe can be used to convert the intent type and target object into word vectors, and then averaged, concatenated, or more complex pooling operations can be used to obtain the semantic vector of the intent. Furthermore, sentence embedding models such as Sentence-BERT or Universal Sentence Encoder can be used to directly map the phrase or sentence composed of the intent type and target object into a high-dimensional semantic space, generating a semantic vector with richer semantic information.

[0063] Based on the intent type, modal source, and safety level mapping table, the safety level is determined. This step aims to assign a safety priority to each intent unit to address potential safety risks in the cockpit environment. The safety level mapping table is a predefined set of rules that maps specific intent types and modal source combinations to different safety levels. For example, when the intent type is "emergency braking," the safety level may be set to the highest level regardless of whether the modal source is voice or touch. Conversely, when the intent type is "play music" and the modal source is voice, the safety level may be lower. Furthermore, a dynamic evaluation mechanism can be employed to adjust the safety level in real time, taking into account the current driving context (e.g., high-speed driving, complex road conditions) and the physiological state of the occupants (e.g., fatigue, distraction), ensuring the safest decisions are made in different situations.

[0064] Based on the original confidence level and security level in the intent feature vector set, the original confidence level field is initialized to generate the corresponding structured intent unit. This step integrates the original intent recognition confidence level and security level to form a more comprehensive confidence assessment. One implementation is to perform a weighted sum or product operation on the original confidence level and security level. For example, initialize confidence level = original confidence level × (1 + security level weight factor), where the security level weight factor is set according to the security level. The higher the security level, the larger the weight factor, thereby improving the confidence level of high-security-level intents. Another implementation is to modify the original confidence level according to the security level. For example, when the security level reaches the preset highest level, even if the original confidence level is slightly low, it can be raised to a higher baseline value to ensure the effectiveness of the key intent. The information such as the intent type, target object, modality source, original confidence level, timestamp, security level, and semantic vector obtained from the above analysis and determination is integrated to form a complete structured intent unit.

[0065] Through the above technical solutions, this application can solve the problems of insufficient semantic depth and lack of security considerations in intent representation. For example, the intent classifier accurately parses the intent type and target object, laying the foundation for subsequent processing. The introduction of the semantic embedding model enables each structured intent unit to have a high-dimensional semantic vector, greatly enhancing the semantic depth of the intent, thereby more accurately capturing the correlation between intents and providing a precise quantitative basis for subsequent conflict degree quantification (such as semantic conflict component calculation). At the same time, the security level is determined based on the intent type, modality source, and security level mapping table, ensuring that intents with high security levels are given priority in the subsequent hierarchical arbitration process, improving the security and reliability of the system. In addition, the confidence field is initialized by fusing the original confidence level with the security level, so that the generated structured intent unit not only reflects the accuracy of intent recognition but also incorporates security considerations, reducing the negative impact of low-reliability or low-security intents on system decision-making. These enriched and optimized structured intent units provide high-quality, high-dimensional input for the cockpit application dynamic scheduling and scenario adaptation system. This enables more accurate assessment of the correlation and priority between intents when dealing with multimodal conflicts, avoiding the problems of inaccurate conflict arbitration and loss of intent information caused by insufficient information in traditional solutions. As a result, it can better support the identification and delayed utilization of composite intents, improving user experience and system intelligence.

[0066] In some of the embodiments described above in this application, a conflict degree quantification method is proposed to quantify the degree of conflict between intent units. However, in its implementation, due to the lack of detailed analysis of semantic association and temporal security factors, the quantification result may be inaccurate, which may lead to the inability to effectively distinguish between fusionable intents and intents that need to be suppressed during arbitration, thereby failing to make full use of the semantic information of suppressed intents.

[0067] To address this, this application further proposes quantifying the conflict level of the multiple structured intent units and the cockpit scene knowledge graph, obtaining conflict level scores and scene semantic association information. See [link to relevant documentation]. Figure 3 The quantification process specifically includes: 301. Map the multiple structured intent units to the cockpit scene knowledge graph to obtain the corresponding node of each structured intent unit in the cockpit scene knowledge graph, and extract the path association, temporal association and instruction conflict relationship between the corresponding nodes from the cockpit scene knowledge graph.

[0068] 302. Based on the node distance between the corresponding nodes, the semantic vector distance between the multiple structured intent units, and the path association relationship, the semantic conflict component is calculated.

[0069] 303. Based on the differences in time urgency, security level, and instruction conflict relationship among the multiple structured intent units, the timing security conflict component is calculated.

[0070] 304. The semantic conflict component and the temporal security conflict component are weighted and fused to obtain the conflict degree score, and the path association, the temporal association, and the instruction conflict are used as the semantic association information of the scene.

[0071] This process involves mapping multiple structured intent units to a cockpit scene knowledge graph, obtaining the corresponding nodes for each structured intent unit within the knowledge graph, and extracting path relationships, temporal relationships, and command conflict relationships between these corresponding nodes. The aim is to associate abstract structured intent units generated from multimodal interaction data with the actual semantic background of the cockpit environment. By mapping intent units to specific entity or concept nodes in the cockpit scene knowledge graph, rich contextual information can be provided for the intent units, enabling a more accurate understanding of their potential meanings and their relationships with other intents or cockpit elements. Simultaneously, the path relationships, temporal relationships, and command conflict relationships extracted from the knowledge graph can reveal potential connections or conflicts between intent units in space, time, and behavior execution, providing a multi-dimensional basis for subsequent conflict quantification. In practice, semantic matching methods can be employed, such as using the intent type, target object, and semantic vector of the intent unit to perform entity linking or concept matching in the cockpit scene knowledge graph to find the most relevant nodes. If multiple candidate nodes exist, the best corresponding node can be determined by calculating the similarity between the semantic vector of the intent unit and the semantic embedding vector of the candidate node. Alternatively, this can be achieved through a predefined rule base and ontology mapping mechanism, pre-defining the corresponding entity or relation type in the knowledge graph for each intent type and target object. When generating structured intent units, mapping is performed directly according to these rules. Relationships in the knowledge graph can then be extracted using graph traversal algorithms or pre-defined query statements.

[0072] Based on the node distance between corresponding nodes, the semantic vector distance between multiple structured intent units, and the path association, a semantic conflict component is calculated. This component quantifies the degree of semantic conflict between different structured intent units. It comprehensively considers the degree of association between intent units in the knowledge graph (node ​​distance), the semantic similarity of the intent units themselves (semantic vector distance), and whether there is a clear path association between them. Through this multi-dimensional consideration, the semantic compatibility or contradiction of intent units can be evaluated more precisely. For example, node distance can be calculated using shortest path algorithms in the knowledge graph (such as Dijkstra's algorithm or BFS) to reflect the semantic distance between entities. Semantic vector distance can be measured using cosine similarity or Euclidean distance. The path association can be a Boolean value indicating whether a predefined semantic path exists. The semantic conflict component can be designed as a weighted sum or product of these factors; for example, the larger the node distance, the smaller the semantic vector distance (higher similarity), and the weaker the path association, the higher the semantic conflict component. Another approach is to utilize graph embedding technology, embedding nodes and relationships from the knowledge graph into a low-dimensional vector space, and then directly calculating the distance between the corresponding node embedding vectors as the node distance. Semantic vector distance is calculated directly using the semantic vectors of the intent unit. Path association can act as a correction factor, reducing conflict components when strong association paths exist, and increasing them when they don't.

[0073] Based on the differences in time urgency, safety level, and command conflict relationships among multiple structured intent units, a temporal safety conflict component is calculated. This component aims to assess the urgency of intent unit execution, its impact on cockpit safety, and the existence of direct command contradictions. In a cockpit environment, some intents may have strict time window requirements or directly affect driving safety, thus requiring separate assessment of these non-semantic level conflicts. Time urgency differences can be calculated by comparing the timestamps or preset execution deadlines of intent units; the greater the difference, the higher the urgency conflict. Safety level differences are quantified according to the safety level of the intent units (e.g., high, medium, low); the greater the level difference, the higher the safety conflict. Command conflict relationships can be a predefined conflict matrix indicating which commands are mutually exclusive. The temporal safety conflict component can integrate these factors; for example, when a direct command conflict exists, the component is directly set to its maximum value. Otherwise, it is calculated based on a weighted sum of time urgency and safety level differences. In addition, machine learning models can be used to predict temporal security conflict components. Input features include timestamps, security levels, and instruction types. The model learns conflict patterns from historical data to output component values. Instruction conflict relationships can serve as an important feature or constraint of the model.

[0074] The semantic conflict component and the temporal safety conflict component are weighted and fused to obtain the conflict severity score. The path association, temporal association, and instruction conflict relationship are then used as the semantic association information for the scene. This step is the core of conflict quantification. By comprehensively considering semantic and temporal safety conflicts, a comprehensive conflict severity score is generated. This score can more accurately reflect the potential degree of contradiction between multiple intent units, providing a decision-making basis for subsequent intent arbitration and scheduling. Simultaneously, outputting the original association relationship as scene semantic association information preserves this detailed contextual information for more refined judgment and utilization in subsequent intent fusion, arbitration, and activation of suppressed intents. The weighted fusion can use a simple linear weighted summation: Conflict Severity Score = W_Semantic × Semantic Conflict Component + W_Temporal Safety × Temporal Safety Conflict Component, where W_Semantic and W_Temporal Safety are preset weights reflecting the importance of semantic and temporal safety conflicts in the overall conflict assessment. The scene semantic association information can be directly packaged into a data structure or object for transmission. Nonlinear fusion methods can also be employed, such as decision trees or neural network models, which take semantic conflict components and temporal security conflict components as input and output conflict severity scores. This approach can learn more complex conflict patterns. Scene semantic association information can then be stored and transmitted in the form of graph structures or relational databases.

[0075] Through the above technical solutions, this application can improve the accuracy and comprehensiveness of conflict degree quantification by introducing a refined mapping and component calculation mechanism, thereby solving the arbitration decision problem caused by inaccurate quantification. For example, mapping multiple structured intent units to a cockpit scene knowledge graph to obtain corresponding nodes can associate intent units with entities in the knowledge graph, providing a semantic context basis and avoiding isolated intent analysis. Extracting path associations, temporal associations, and instruction conflict relationships between corresponding nodes from the cockpit scene knowledge graph captures the potential connections and conflict points between intents, providing multi-dimensional basis for quantification and preventing the omission of key association information. The semantic conflict component is calculated based on the node distance between corresponding nodes, the semantic vector distance between multiple structured intent units, and the path association. Combining node distance to reflect the strength of association between entities, semantic vector distance to reflect intent similarity, and path association to indicate the existence of associated paths, these factors ensure that the semantic conflict component accurately reflects the degree of conflict at the semantic level, avoiding the bias caused by a single distance indicator. Based on the differences in time urgency, safety level, and command conflict relationships among multiple structured intent units, a temporal safety conflict component is calculated. The time urgency difference captures temporal urgency, the safety level difference highlights safety priority, and the command conflict relationship directly identifies the existence of conflict, ensuring the component comprehensively covers both temporal and safety dimensions and preventing the omission of key safety factors in driving scenarios. The semantic conflict component and the temporal safety conflict component are weighted and fused to obtain a conflict degree score. The fusion process balances semantic and temporal safety factors, generating a comprehensive quantitative result to support more reasonable arbitration decisions. Simultaneously, path association relationships, temporal association relationships, and command conflict relationships are output as scene semantic association information. These relationships are retained to provide rich context for subsequent hierarchical arbitration, facilitating the identification of fusionable intents or the setting of precise trigger conditions. This ensures that the semantic information of suppressed intents is fully utilized, thus solving the problem of completely discarding the semantic information of suppressed intents in existing technologies and achieving semantic information preservation and delayed utilization in multimodal conflict scenarios.

[0076] In some of the embodiments described above in this application, a method is proposed to map structured intent units to a cockpit scene knowledge graph to obtain corresponding nodes for conflict quantification. However, in its implementation, when the intent type and target object of the intent unit do not have a directly matching entity in the knowledge graph, the corresponding node may not be found, causing the intent unit to be ignored or mishandled, affecting the accuracy of conflict quantification and the integrity of intent processing.

[0077] To address this, this application further proposes mapping multiple structured intent units to a cockpit scene knowledge graph to obtain the corresponding node for each structured intent unit in the cockpit scene knowledge graph. This process includes: performing entity matching in the cockpit scene knowledge graph based on the intent type and target object of each structured intent unit to obtain an initial candidate node set. If the initial candidate node set contains multiple candidate nodes, the candidate node with the highest similarity is selected as the corresponding node based on the similarity between the semantic vector of each structured intent unit and the semantic embedding vectors of the multiple candidate nodes. If the initial candidate node set is empty, similar node retrieval is performed in the cockpit scene knowledge graph based on the semantic vector of each structured intent unit, and the retrieved similar nodes are used as the corresponding node.

[0078] For example, in the step of "based on the intent type and target object of each structured intent unit, perform entity matching in the cockpit scene knowledge graph to obtain an initial candidate node set," the intent type and target object are the core semantic attributes of the structured intent unit, used for preliminary entity localization in the cockpit scene knowledge graph. Entity matching aims to utilize these core attributes to find directly corresponding entities or concepts in the knowledge graph. For instance, exact matching can be used, directly using the intent type and target object as query conditions to search in the entity index of the knowledge graph. Alternatively, fuzzy matching or pattern matching can be used, allowing for a certain degree of semantic generalization or using regular expressions for matching to handle the diversity of user expressions.

[0079] In the step of "selecting the candidate node with the highest similarity as the corresponding node based on the similarity between the semantic vector of each structured intent unit and the semantic embedding vectors of the multiple candidate nodes when the initial candidate node set contains multiple candidate nodes," a deeper semantic analysis is needed to determine the most accurate match when the initial matching result is ambiguous, i.e., when there are multiple potential candidate nodes. The semantic vector of the structured intent unit and the semantic embedding vectors of each candidate node in the knowledge graph provide the representation of intent and entity in the semantic space. By calculating the similarity between these vectors, the degree of semantic association between them can be quantified. For example, the cosine similarity between the semantic vector of the structured intent unit and the semantic embedding vectors of each candidate node can be calculated, and the candidate node with the largest cosine value can be selected. Alternatively, the Euclidean distance between them can be calculated, and the candidate node with the smallest distance can be selected.

[0080] In the step of "when the initial candidate node set is empty, similar node retrieval is performed in the cockpit scene knowledge graph based on the semantic vectors of each structured intent unit, and the retrieved similar nodes are taken as the corresponding nodes," if direct entity matching fails to find any candidate nodes, it indicates that there may be no perfectly matching entities in the knowledge graph, but there may be semantically highly related entities. In this case, similar node retrieval aims to find the semantically closest nodes in the entire cockpit scene knowledge graph through semantic vectors to avoid the loss of intent information. For example, all entity nodes in the cockpit scene knowledge graph can be pre-converted into semantic embedding vectors and stored in a vector database. Then, the semantic vectors of the structured intent units can be used as query vectors to perform nearest neighbor search in the vector database to retrieve the semantically most similar nodes. Alternatively, graph embedding techniques can be used to embed the nodes in the knowledge graph into a low-dimensional vector space, and then the similarity between the semantic vectors of the structured intent units and these graph embedding vectors can be calculated, selecting the node with the highest similarity.

[0081] Through the above technical solution, this application provides a robust and multi-layered mapping mechanism to ensure that each structured intent unit can find a corresponding node in the cockpit scene knowledge graph. This mechanism performs initial entity matching based on intent type and target object, utilizing the core attributes of the intent for localization. When multiple candidate nodes exist, precise filtering is performed using semantic vector similarity, resolving matching ambiguity and avoiding incorrect mapping. Even without a direct match, similar nodes can be retrieved using semantic vectors, ensuring that all intent information is incorporated into the knowledge graph framework. This solves the problem that when the intent type and target object of an intent unit do not have a directly matching entity in the knowledge graph, a corresponding node may not be found, leading to the intent unit being ignored or incorrectly processed. This improves the accuracy of conflict level quantification and the completeness of intent processing, providing reliable and comprehensive basic data for subsequent conflict level score calculation, hierarchical arbitration, and spatiotemporal linkage matching, thereby enhancing the overall intelligence level and user experience of the cockpit application dynamic scheduling and scene adaptation system.

[0082] In some of the embodiments described above in this application, a method for calculating semantic conflict components to quantify the degree of conflict between intent units is proposed. However, if path association is not considered during its implementation, the conflict quantification may be inaccurate, and the semantic association between intent units may not be effectively distinguished, thereby overestimating the degree of conflict or ignoring potential associations, which may affect the accuracy of subsequent arbitration decisions.

[0083] To address this, this application further proposes a method for calculating semantic conflict components, including: calculating semantic conflict components based on node distances between corresponding nodes, semantic vector distances between multiple structured intent units, and path associations. For example, a first sub-component is calculated based on the node distances between corresponding nodes, and this first sub-component is positively correlated with the node distance. When quantifying the conflict degree of multiple structured intent units, these structured intent units need to be mapped to a cockpit scene knowledge graph to obtain the corresponding nodes of each structured intent unit in the knowledge graph. These corresponding nodes represent specific entities or concepts of the intent in the cockpit scene. Node distance refers to the shortest path length or semantic distance between two corresponding nodes connected by an edge in the cockpit scene knowledge graph. This distance reflects the degree of association or spatial proximity of the entities or concepts pointed to by the two intents in the knowledge graph. For example, node distance can be obtained by calculating the number of edges contained in the shortest path between two nodes using graph traversal algorithms (such as breadth-first search or depth-first search). Another approach is to characterize node distance by calculating the distance between these node embedding vectors (e.g., Euclidean or cosine distance) if the nodes in the knowledge graph have embedding vectors. The first subcomponent, calculated based on the node distance between corresponding nodes, quantifies the degree of conflict between intent units at the knowledge graph level. This subcomponent is positively correlated with node distance; that is, the larger the node distance, the larger the value of the first subcomponent, indicating a weaker association or greater spatial distance between the entities or concepts associated with the intent in the knowledge graph, and thus a higher potential degree of conflict. For example, the first subcomponent can be simply set as a linear function of node distance, such as C1 = k1 × node distance, where k1 is a positive coefficient. Alternatively, a non-linear function can be used, such as C1 = f(node ​​distance), where f is a monotonically increasing function, such as a logarithmic or exponential function, to more precisely reflect the impact of distance on conflict.

[0084] Based on the semantic vector distances between multiple structured intent units, a second sub-component is calculated, which is positively correlated with the semantic vector distances. Each structured intent unit contains a semantic vector, generated through a semantic embedding model, used to represent the deep semantic information of the intent unit. The semantic vector distance between multiple structured intent units refers to the distance between the semantic vectors of these intent units, directly reflecting the degree of similarity or difference in semantic content between different intent units. For example, the semantic vector distance can be obtained by calculating the cosine distance (1-cosine similarity) between two semantic vectors; a larger cosine distance indicates a greater semantic difference. Euclidean distance can also be used to measure distance in vector space; a larger distance also indicates a greater semantic difference. The second sub-component, calculated based on the semantic vector distances between multiple structured intent units, is used to quantify the degree of conflict between intent units at the semantic content level. This sub-component is positively correlated with the semantic vector distances; that is, the larger the semantic vector distances, the larger the value of the second sub-component, indicating greater semantic differences between intent units and thus a higher potential degree of conflict. For example, the second subcomponent can be set as a linear function of the semantic vector distance, such as C2 = k2 × semantic vector distance, where k2 is a positive coefficient. Alternatively, a non-linear function can be used, such as C2 = g(semantic vector distance), where g is a monotonically increasing function to more flexibly capture the impact of semantic differences on conflict.

[0085] When the path association indicates that a path relationship exists between corresponding nodes, the weighted sum of the first and second sub-components is multiplied by a preset path association discount coefficient to obtain the semantic conflict component. The path association is extracted from the cockpit scene knowledge graph and is used to indicate whether a predefined, meaningful path exists between two corresponding nodes. This path may represent a logical, functional, or spatiotemporal connection between intents. For example, if one intent is "navigate to the company" and another is "play music," there may be no direct path relationship between them. However, if one intent is "navigate to the company" and another is "pass by a gas station," there may be a "waypoint" path relationship. Path associations can be identified by predefining specific types of edges or path patterns in the knowledge graph. For example, relation types such as "contains," "adjacent," and "pass by" can be defined; when these relations exist between two nodes, a path relationship is considered to exist. The preset path association discount coefficient is a value between 0 and 1, used to correct the calculation results of the semantic conflict component when a path relationship exists. Its function is to reduce the conflict score between intent units that have related paths in the knowledge graph, reflecting that they are not completely conflicting, but rather there may be some possibility of some kind of synergy or sequential execution. For example, the discount factor can be a fixed value, such as 0.5 or 0.8, determined through experience or system tuning. It can also be dynamically adjusted according to the type or strength of the path association; for example, a smaller discount factor corresponds to a strong association, and a larger discount factor corresponds to a weak association.

[0086] If the path association indicates that there is no path association between the corresponding nodes, the weighted sum of the first sub-component and the second sub-component is taken as the semantic conflict component. This means that when no predefined path association is detected between intent units in the knowledge graph, their semantic conflict components will not be discounted to ensure a true reflection of their conflict level.

[0087] Through the above technical solution, this application comprehensively considers the entity association (node ​​distance) of intent units in the knowledge graph, the semantic differences of the intents themselves (semantic vector distance), and whether there are predefined path associations between them when calculating semantic conflict components. For example, by introducing a first sub-component and a second sub-component, the degree of conflict of intents is quantified at the entity level and the semantic level, respectively. More importantly, when a path association is detected between intent units, a preset path association discount coefficient can be applied to effectively reduce its semantic conflict component, thereby avoiding unnecessary overestimation of conflict for intents with potential associations or collaborations. Conversely, for intents without path associations, no discount is applied, ensuring a true reflection of their conflict degree. This refined conflict quantification method enables more accurate identification of real conflicts and potential collaborations between intents, providing a more reliable and refined basis for subsequent hierarchical arbitration, thereby improving the accuracy and intelligence level of dynamic scheduling and scenario adaptation in cockpit applications, avoiding misjudgments or omissions caused by inaccurate conflict quantification, and ultimately improving the user experience.

[0088] In some of the embodiments described above in this application, a calculation of timing security conflict components is proposed to quantify the degree of conflict. However, in this process, due to the lack of specific calculation rules for differences in time urgency, security level, and command conflict relationships, it may be impossible to accurately quantify the conflict when there is a direct command conflict, thereby affecting the accuracy of subsequent arbitration.

[0089] To address this, this application further proposes a method for calculating a timing security conflict component based on differences in time urgency, security levels, and command conflict relationships among multiple structured intent units. Specifically, this includes: calculating a time conflict sub-component based on the time urgency differences among multiple structured intent units, where the time conflict sub-component is positively correlated with the time urgency difference; calculating a security conflict sub-component based on the security level differences among multiple structured intent units, where the security conflict sub-component is positively correlated with the security level differences; when the command conflict relationship indicates a direct command conflict among multiple structured intent units, setting the weighted sum of the time conflict sub-component and the security conflict sub-component to a preset maximum value to obtain the timing security conflict component; and when the command conflict relationship indicates no direct command conflict among multiple structured intent units, using the weighted sum of the time conflict sub-component and the security conflict sub-component as the timing security conflict component.

[0090] For example, when calculating time conflict subcomponents, it's necessary to obtain the time urgency difference between multiple structured intent units. This time urgency difference reflects the degree of urgency or the gap in expected completion time between different intent units. For instance, one intent might require "execution immediately," while another might allow "execution within the next five minutes," creating a time urgency difference. This time conflict subcomponent aims to quantify this temporal inconsistency. One implementation is to assign a time urgency score to each structured intent unit, which can be determined based on the semantic content of the intent (such as "immediately," "as soon as possible," "later") or the difference between its associated timestamp and the current time, and then calculate the absolute difference between these scores as the time urgency difference. Another implementation is to define an expected execution time window for each intent unit and calculate the time urgency difference based on the overlap or interval between these time windows, mapping it to a time conflict subcomponent. This subcomponent is positively correlated with the degree of difference; that is, the greater the difference, the higher the subcomponent value.

[0091] When calculating the safety conflict subcomponent, it is necessary to obtain the safety level differences between multiple structured intent units. This safety level difference characterizes the degree of safety risk or impact on occupant safety that different intent units may pose during execution. For example, one intent might involve vehicle control (such as "emergency braking"), which has a high safety level. Another intent might only involve entertainment system adjustments (such as "playing music"), which has a low safety level. This safety conflict subcomponent is used to quantify this incompatibility at the safety level. One implementation is to pre-define a safety level for each structured intent unit (e.g., a value from 1 to 5, with higher values ​​indicating higher safety levels), and then calculate the absolute difference between these safety levels as the safety level difference. Another implementation is to determine the safety level based on the intent type and target object by consulting a predefined safety risk assessment table, and then calculate the safety level difference based on this. This subcomponent is positively correlated with the safety level difference; that is, the greater the difference, the higher the subcomponent value.

[0092] When a command conflict relationship indicates a direct command conflict among multiple structured intent units, the weighted sum of the temporal conflict subcomponent and the security conflict subcomponent is set to a preset maximum value. This command conflict relationship is predefined in the cockpit scenario knowledge graph and is used to identify fundamental, irreconcilable contradictions between two or more intent units, such as the simultaneous occurrence of "open the window" and "close the window." When the system detects such a direct command conflict, it means that these intents cannot be executed simultaneously or compatibly within a short period, requiring explicit arbitration. Setting the weighted sum to a preset maximum value aims to highlight the severity of this conflict, ensuring that it is given the highest priority in subsequent hierarchical arbitration, thereby preventing the system from getting bogged down in the execution of contradictory commands. This preset maximum value can be a pre-set fixed value, such as 1.0 or 100, to ensure it has the absolute highest weight among all conflict scores.

[0093] When there is no direct command conflict between multiple structured intent units, indicating a conflict relationship, the weighted sum of the time conflict subcomponent and the safety conflict subcomponent is taken as the temporal safety conflict component. When there is no direct, fundamental command conflict between intent units, the degree of conflict can be quantified by comprehensively considering differences in time urgency and safety level. This weighted sum allows the system to assign different importance weights to time conflicts and safety conflicts according to actual needs. For example, in some driving scenarios, time urgency may be more critical, while in others, safety considerations take precedence. By adjusting the weights, the impact of different conflict factors on the overall conflict degree can be flexibly reflected, resulting in a more refined and reasonable temporal safety conflict component.

[0094] Through the above technical solutions, this application addresses potential accuracy issues in conflict quantification by refining the specific methods for calculating timing-based security conflict components. This ensures that the degree of conflict is accurately reflected in direct command conflict scenarios, thereby supporting the accuracy of subsequent arbitration. Specifically, time conflict sub-components are calculated based on differences in time urgency, quantifying the degree of temporal conflict and preventing the neglect of time-critical intentions. Security conflict sub-components are calculated based on differences in security levels, quantifying security-related conflicts and ensuring that high-security-level intentions are prioritized. When direct command conflicts exist, the weighted sum of the sub-components is set to a preset maximum value, highlighting the severity of the conflict and preventing the arbitration process from underestimating critical conflicts. When no direct command conflicts exist, the weighted sum of the sub-components and the security conflict sub-components is used as the component, maintaining computational flexibility to adapt to different conflict scenarios. These features collectively ensure the accuracy of conflict quantification, providing a reliable basis for subsequent graded arbitration, thereby improving the intelligence level and user experience of the cockpit application dynamic scheduling and scenario adaptation system.

[0095] In some of the solutions mentioned above in this application, a hierarchical arbitration is proposed to handle intent conflicts and generate execution instruction sets and cache queues based on conflict degree scores. However, in this process, how to effectively divide intent units into different levels based on conflict scores, perform fusion processing on low-conflict intents, arbitrate high-conflict intents, and generate precise spatiotemporal triggering conditions for suppressed intents in order to avoid the complete discarding of semantic information and achieve delayed activation is a technical problem.

[0096] To address this, this application proposes a cockpit application dynamic scheduling and scenario adaptation system based on multimodal intent fusion. Its processor is configured to perform hierarchical arbitration on multiple structured intent units, conflict severity scores, and scenario semantic association information, generating a cockpit application execution instruction set and a queue of intents to be activated. For example, see... Figure 4 The tiered arbitration includes the following steps: 401. Based on the preset threshold range of the conflict level score, multiple structured intent units are divided into a first-level intent unit set and a second-level intent unit set.

[0097] 402. Perform semantic fusion on the structured intent units in the first-level intent unit set to generate composite intent units, and add the cockpit application control commands corresponding to the composite intent units to the cockpit application execution command set.

[0098] 403. Arbitrate the structured intent units in the second-level intent unit set to determine the main idea unit and the suppressed intent unit, and add the cockpit application control commands corresponding to the main idea unit to the cockpit application execution command set.

[0099] 404. Based on the suppressed intent unit and scene semantic association information, generate the spatiotemporal triggering conditions corresponding to the suppressed intent unit, and store the suppressed intent unit and spatiotemporal triggering conditions together to obtain the intent cache queue to be activated.

[0100] The conflict severity score quantifies the severity of semantic and temporal security conflicts among multiple structured intent units. This score is calculated by comprehensively considering factors such as semantic similarity, time urgency, security level differences, and command conflict relationships between intent units. For example, a higher score indicates a more severe conflict between intent units, requiring arbitration. A lower score indicates a smaller conflict, potentially allowing for fusion. The preset threshold range is a pre-defined series of numerical ranges used to classify conflict severity scores, thus dividing structured intent units into different processing levels. For example, a fusion threshold can be set; intent units below this threshold are considered fusionable, while those above or equal to this threshold require arbitration. This threshold range can be dynamically adjusted or set based on actual application scenarios, user experience requirements, or security strategies, or through expert experience. The first-level intent unit set contains structured intent units with lower conflict severity scores. These intent units are considered semantically compatible or complementary, suitable for fusion processing to form more complex composite intents. For example, when a user gives the voice command "open the car window" while simultaneously focusing their gaze on the window control button, these two intents might be classified as first-level. The second-level intent unit set contains structured intent units with high conflict scores or direct conflicts. These intent units have semantic or temporal security conflicts that are not suitable for direct fusion and require an arbitration mechanism to determine which intent unit should be executed first and which should be suppressed. For example, when a user's voice command "play music" is accompanied by a gesture command "turn off the speakers," these two intents may be classified into the second level.

[0101] Semantic fusion refers to the process of integrating multiple semantically related or complementary structured intent units into a higher-level, more comprehensive composite intent unit. Its purpose is to retain and utilize all relevant intent information, avoiding information loss caused by executing a single intent. For example, this can be achieved by constructing an intent graph, analyzing the semantic relationships between intent units, and then using graph neural networks for feature aggregation, or by combining multiple simple intents into a complex intent through rule engines and template matching. A composite intent unit is a new intent representation formed by semantic fusion of multiple first-level intent units. It typically contains an intent graph and one or more supplementary intents, and may include the spatiotemporal triggering conditions and fusion types corresponding to these supplementary intents. For example, "Navigate to the company and remind me to buy coffee when passing Starbucks" is a composite intent unit. Cockpit application control commands are specific operational instructions generated by the system based on the parsing of structured intent units or composite intent units, used to directly control various applications or functions within the cockpit. Examples include "Play music," "Adjust the air conditioning temperature," and "Turn on navigation." The cockpit application execution instruction set is a set of instructions that contains the cockpit application control commands that need to be executed immediately. These instructions can come from the idea diagram unit or from the composite idea unit.

[0102] Arbitration refers to the process of selecting one conflicting structured intent unit as the main graph unit for execution from multiple conflicting structured intent units, based on preset strategies (such as security level, confidence level, user preferences, etc.), while treating the other conflicting intent units as suppressed intent units. For example, a decision tree model based on priority ranking or a machine learning model can be used to decide on conflicting intents. A main graph unit is a structured intent unit selected for priority execution during the arbitration process. It represents the user intent that the system currently considers most important or urgent. A suppressed intent unit is a structured intent unit that is not selected for immediate execution during the arbitration process, but whose semantic information is retained and cached, waiting to be activated again when specific conditions are met in the future. Scene semantic association information is extracted from the cockpit scene knowledge graph and is used to describe the relationships between structured intent units and between intent units and cockpit environment entities, including path associations, temporal associations, and command conflict relationships. For example, it can indicate whether two intents are spatially related or whether there is a temporal order. Spatiotemporal triggering conditions are conditions set for suppressed intent units to activate the intent under specific future time or spatial conditions. It typically includes the trigger location (such as geographic coordinates or POI points) and the trigger time window (such as a specific time period or after an event occurs). For example, a suppressed intent will only be reconsidered and activated when a vehicle arrives at a specific location or within a certain time period. The pending intent cache queue is a queue that stores suppressed intent units and their associated spatiotemporal trigger conditions. The intent units in this queue are continuously monitored, and when their spatiotemporal trigger conditions are met, they are reactivated and added to the execution instruction set.

[0103] Through the above technical solution, this application introduces a hierarchical arbitration mechanism, solving the problem of complete discarding of semantic information in multimodal intent conflict processing and enabling delayed utilization of suppressed intents. For example, by dividing intent units into a first-level intent unit set and a second-level intent unit set based on a preset threshold range of conflict severity scores, fine-grained management of intent conflicts is achieved. For first-level intent units with low conflict severity, the system performs semantic fusion to generate composite intent units, thereby enabling simultaneous responses to multiple compatible user intents, avoiding the loss of semantic information caused by executing a single intent, and improving the richness and coherence of the user experience. For example, when a user's voice command "open the window" is simultaneously focused on the window control button, these two intents can be fused to ensure the window is opened, and may be fine-tuned based on the focus of the gaze, rather than simply selecting one. For second-level intent units with high conflict severity, the system arbitrates to determine the main intent unit and the suppressed intent unit. This processing method ensures that key intents are executed immediately without simply discarding the semantic information of suppressed intents. Conversely, based on the suppressed intent unit and the semantic association information of the scene, precise spatiotemporal triggering conditions are generated for it, and these conditions are stored in the cache queue of intents to be activated. For example, when the driver's voice command is "navigate to the company" while their gaze is fixed on a roadside restaurant for an extended period, "navigate to the company" will be executed immediately as the main intent, while the "roadside attention" intent will be treated as a suppressed intent, and a spatiotemporal triggering condition of "when the vehicle approaches the restaurant" will be generated for it. This mechanism ensures that the suppressed intent can be reactivated at an appropriate time and place in the future, thereby making full use of all intent information, greatly improving the intelligence of the cockpit application and the responsiveness of user intents, and avoiding the drawback of the permanent discarding of the "roadside attention" intent in traditional solutions.

[0104] In some of the embodiments described above in this application, an intention unit set is proposed to be divided based on the degree of conflict score for hierarchical arbitration. However, in its implementation, there is a lack of a clear threshold to distinguish between semantically fusionable intentions and intentions that need to be arbitrated, resulting in inaccurate intention division and failure to effectively preserve and delay the use of semantic information.

[0105] To address this, this application further proposes dividing multiple structured intent units into a first-level intent unit set and a second-level intent unit set based on a preset threshold range for the conflict level score. Specifically, this includes: classifying structured intent units with conflict level scores less than a preset fusion threshold into the first-level intent unit set, where the preset fusion threshold is used to distinguish semantically fusionable intents from those requiring arbitration; and classifying structured intent units with conflict level scores greater than or equal to the preset fusion threshold into the second-level intent unit set.

[0106] The conflict severity score is a numerical value obtained by quantifying the conflict severity of multiple structured intent units and the cockpit scene knowledge graph. It quantifies the semantic and temporal conflict severity between different intent units. This score is the basis for classifying intent units, and its value directly reflects the severity of the conflict between intents. The preset threshold range refers to one or more numerical ranges set for the conflict severity score, used to classify intent units. This range provides an objective dividing standard, ensuring the consistency and quantifiability of intent classification. For example, a normalized range of 0 to 1 can be set, containing one or more specific threshold points. The preset fusion threshold is a key critical value within the preset threshold range, specifically used to distinguish between semantically fusionable intents and intents requiring arbitration. This threshold can be a fixed value, such as 0.3 or 0.5, or a parameter dynamically adjusted based on historical data or expert experience. Its setting aims to ensure that low-conflict intents can be effectively fused, while high-conflict intents enter the arbitration process. The first-level intent unit set includes structured intent units whose conflict severity scores are lower than the preset fusion threshold. These intent units are determined by the system to be semantically compatible or complementary, suitable for semantic fusion processing to form more comprehensive composite instructions that better reflect the user's true intentions. The second-level intent unit set includes structured intent units whose conflict scores are greater than or equal to a preset fusion threshold. These intent units are determined by the system to have high conflict and require an arbitration mechanism to determine their priority and execution method to avoid instruction conflicts or improper operations. Distinguishing between semantically fusionable intents and intents requiring arbitration is the core function of the preset fusion threshold. Through clear threshold definition, it is possible to avoid the erroneous fusion of inherently conflicting intents, while ensuring the effective integration of synergistic intents, thereby maximizing the preservation and utilization of semantic information in multimodal interactions.

[0107] Through the above technical solution, this application provides a clear and objective basis for the classification of intent units by introducing a preset fusion threshold. For example, structured intent units with conflict scores less than the preset fusion threshold are classified into the first-level intent unit set, allowing these semantically fusionable intents to enter the subsequent semantic fusion processing flow. This avoids the loss of semantic information due to improper conflict handling and achieves a comprehensive understanding and preservation of the user's complex intents. Simultaneously, structured intent units with conflict scores greater than or equal to the preset fusion threshold are classified into the second-level intent unit set, ensuring that intents with high conflict levels are accurately identified and sent to a dedicated arbitration mechanism, avoiding system chaos or erroneous execution that may be caused by improper fusion. This refined hierarchical processing based on clear thresholds improves the accuracy and robustness of cockpit application dynamic scheduling and scenario adaptation in handling multimodal intent conflicts, enabling more intelligent responses to user needs and laying a solid foundation for the delayed activation of suppressed intents. This solves the technical problems of inaccurate intent classification and the inability to effectively preserve and utilize semantic information in a delayed manner.

[0108] In some of the solutions mentioned above in this application, semantic fusion is proposed to generate composite intent units by performing semantic fusion on structured intent units in the first-level intent unit set. However, in this process, semantic fusion may fail to fully capture the deep semantic associations and entity relationships between structured intent units, resulting in the generated composite intent units being inaccurate or incomplete and unable to effectively integrate relevant information in the cockpit scene knowledge graph.

[0109] To address this, this application further proposes semantic fusion of structured intent units in the first-level intent unit set to generate composite intent units. Specifically, this includes the following steps: First, each structured intent unit in the first-level intent unit set is used as a node, and entities associated with each structured intent unit in the cockpit scene knowledge graph are used as additional nodes to construct a fusion graph network. Second, based on the semantic relationships between nodes in the fusion graph network, information propagation and aggregation are performed through a graph neural network to obtain the fusion representation vector of each structured intent unit. Third, based on the fusion representation vector of each structured intent unit, the composite intent unit is generated. This composite intent unit includes a main graph, a supplementary intent list, a fusion type, and the spatiotemporal triggering conditions corresponding to the supplementary intent list.

[0110] For example, constructing a fusion graph network aims to unify the representation of intent units to be fused and related background knowledge for deeper semantic analysis. In one implementation, each structured intent unit (containing intent type, target object, semantic vector, etc.) can be directly treated as a node in the graph, with its attributes serving as the intent unit's features. Simultaneously, entities related to the target object and intent type of these intent units (e.g., specific locations, devices, people, behaviors, etc.) are retrieved from the cockpit scene knowledge graph and added as additional nodes in the graph. Edges between nodes can represent predefined semantic relationships, such as "contains," "located in," "acts on," etc. In another implementation, the semantic vectors of structured intent units can be used as node features, while entities in the knowledge graph are used as additional node features through their embedding vectors. The graph can be constructed using an adjacency matrix or adjacency list, and the edge weights can be determined based on the semantic similarity or predefined relationship strength between entities.

[0111] Based on the semantic relationships between nodes in this fusion graph network, information propagation and aggregation are performed using graph neural networks to obtain the fusion representation vector of each structured intent unit. Graph Neural Networks (GNNs) can utilize graph structures for information interaction and feature learning, thereby generating a richer and more context-rich fusion representation for each structured intent unit. For example, Graph Convolutional Networks (GCNs) can be used, which updates the representation of the current node by aggregating information from neighboring nodes, thus capturing the local graph structure and node features. In each propagation layer, each node receives feature information from its neighboring nodes and performs a nonlinear transformation based on its own features to generate a new node representation. After multiple layers of GCN propagation, each node of a structured intent unit will contain its own information and its contextual information in the fusion graph network, forming a fusion representation vector. Alternatively, Graph Attention Networks (GATs) can be used, which, by introducing an attention mechanism, allows nodes to assign different weights when aggregating neighbor information, thus capturing important semantic relationships more flexibly. For example, for an intent unit node, it may pay more attention to knowledge graph entity nodes directly related to it and assign them higher attention weights. By using multi-layered GAT, a more discriminative and semantically rich fusion representation vector can be obtained.

[0112] Based on the fusion representation vectors of each structured intent unit, a composite intent unit is generated. This composite intent unit includes a main intent graph, a list of supplementary intents, a fusion type, and spatiotemporal triggering conditions corresponding to the list of supplementary intents. This represents the output of semantic fusion, transforming the fusion representations obtained through deep learning into an executable and understandable composite intent structure for subsequent cockpit application scheduling. For example, a decoder module can be designed that receives the fusion representation vectors of all structured intent units as input. A classifier or clustering algorithm identifies the most core intent from these fusion representations as the main intent graph. Then, other intents are identified as supplementary intents, and the fusion type is determined based on their semantic relationship with the main intent graph (e.g., whether it is a modification, limitation, or parallel relationship). For each supplementary intent, the corresponding spatiotemporal triggering conditions are generated by combining its target object and location information in the cockpit scene knowledge graph. Alternatively, a sequence generation model can be used, taking the fusion representation vectors as conditional inputs to directly generate the components of the composite intent unit. For example, a Transformer-based decoder is used to predict the idea graph, and then, based on the idea graph and the fused representation, a list of supplementary intentions, fusion types, and spatiotemporal triggering conditions for each supplementary intention are generated progressively. The generation of spatiotemporal triggering conditions can combine a pre-trained geographic information model and a time prediction model.

[0113] Through the above technical solution, this application achieves more accurate and comprehensive semantic fusion by constructing a fusion graph network and applying graph neural network technology. By using structured intent units as nodes and introducing related entities from the cockpit scene knowledge graph as additional nodes, a fusion graph network is constructed. This captures deep semantic connections and entity relationships between structured intent units, avoiding the semantic information loss caused by simply fusion ignoring external knowledge. Based on the semantic connections between nodes in this network, information propagation and aggregation are performed through graph neural networks, effectively integrating the semantic information between nodes and achieving dynamic fusion of intent features, thus obtaining more accurate fusion representation vectors. Based on these fusion representation vectors, composite intent units are generated, including a main intent graph, a supplementary intent list, a fusion type, and the spatiotemporal triggering conditions corresponding to the supplementary intent list. This ensures the integrity and executability of the composite intent, solves the problem of incomplete intent unit generation, better meets the user's composite intent needs, and avoids the loss of semantic information.

[0114] In response, this application further proposes a specific method for generating composite intent units based on the fusion representation vectors of each structured intent unit. This method addresses the challenges of accurately decoding the fusion representation vectors to identify intent types, reasonably distinguishing between main intent graphs and supplementary intents, and setting effective spatiotemporal triggering conditions to preserve the ambiguity of suppressed intent information during the generation of composite intent units for fusion of semantic intents. This avoids confusion in fusion logic and prevents supplementary intents from being activated at appropriate times and spaces, thus failing to meet the user's composite needs.

[0115] The method for generating composite intent units proposed in this application includes the following steps: The fusion representation vectors of each structured intent unit are decoded to obtain the corresponding fusion intent type. These fusion representation vectors are obtained after information propagation and aggregation through a graph neural network, containing deep semantic information about each structured intent unit and its associated entities within the fusion graph network. Decoding these fusion representation vectors aims to transform these abstract, high-dimensional vector representations into human-understandable and system-operable intent types. For example, the decoding process can be implemented in several ways. One approach is to use a pre-trained neural network decoder, typically composed of one or more fully connected layers, which takes the fusion representation vector as input and outputs a probability distribution representing the likelihood of the vector corresponding to different intent types, selecting the intent type with the highest probability as the fusion intent type. Another approach is to use a rule-based pattern matching method, pre-defining mapping rules between the fusion representation vector and specific intent types. When the fusion representation vector conforms to a certain rule, it is decoded into the corresponding intent type.

[0116] Based on the types of fusion intents and the semantic relationships between nodes in the fusion graph network, one main graph unit and at least one supplementary intent unit are determined from each structured intent unit. This step aims to identify which of the multiple intent units is the user's most critical intent requiring immediate response (the main graph), and which are secondary intents that can be delayed (supplementary intents). Methods for determining the main graph unit and supplementary intent units can include: one approach is based on preset priority rules, such as assigning different priorities to different fusion intent types according to factors like security level, time urgency, and user historical preferences, identifying the highest-priority intent unit as the main intent unit and the rest as supplementary intent units. Another approach is to utilize a machine learning classifier, using the fusion intent type and semantic relationships (such as node distance and path association) as feature inputs to train a classification model to determine whether each intent unit is a main graph or a supplementary intent.

[0117] The fusion type is determined based on the semantic association between the main idea unit and the supplementary intention unit. The fusion type describes the logical relationship between the main idea and the supplementary intention, such as whether it is a "modifying relationship" (e.g., "play music" + "turn down volume"), a "parallel relationship" (e.g., "navigation" + "find nearby restaurants"), or a "causal relationship". Methods for determining the fusion type can include: one approach is to predefine the fusion type, for example, by searching for the corresponding fusion type in a pre-defined fusion type table based on the combination of intention types of the main idea and the supplementary intention. Another approach is to use the semantic association strength and type for judgment; for example, if the semantic association path between the main idea and the supplementary intention in the fusion graph network is short and indicates a modifying relationship, then it is determined to be a modifying fusion.

[0118] Based on the target object corresponding to the supplementary intent unit and the location entity relationships in the cockpit scene knowledge graph, spatiotemporal triggering conditions are generated for the supplementary intent unit. Spatiotemporal triggering conditions are crucial to ensuring that the supplementary intent is activated at the appropriate time and place. Methods for generating spatiotemporal triggering conditions can include: one approach is to utilize the geographic coordinates of the target object and navigation path information in the cockpit scene knowledge graph. For example, if the target object of the supplementary intent is a "roadside restaurant," the expected time when the vehicle will pass near the restaurant can be calculated based on the restaurant's geographic location and the vehicle's current navigation path, thus generating a triggering condition that includes a geographic range and a time window. Another approach is based on a distance threshold between the target object and the vehicle's current location; for example, triggering when the distance between the vehicle and the target object is less than a preset threshold.

[0119] The main idea unit, supplementary intent units, fusion type, and spatiotemporal triggering conditions are combined to generate a composite intent unit. This step integrates all the key information identified and generated above into a unified, structured data entity, facilitating subsequent storage, management, and execution by the system. The combination method can adopt a structured data format, such as JSON, XML, or object-oriented data structures, which includes detailed information about the main idea and a list of supplementary intents (each supplementary intent contains its own information, fusion type, and corresponding spatiotemporal triggering conditions).

[0120] Through the above technical solution, this application can ensure that the main graph and supplementary intent are accurately identified and effective triggering conditions are set during the multimodal intent fusion process, thereby solving the problem of semantic information being discarded in traditional solutions. For example, by decoding the fusion representation vector of each structured intent unit, abstract fusion semantic information can be transformed into specific and operable fusion intent types, laying the foundation for subsequent intent differentiation. Combining the fusion intent type and the semantic association between nodes in the fusion graph network, the main graph unit and at least one supplementary intent unit can be intelligently divided, allowing the main graph to be executed first, while the supplementary intent is effectively retained. Based on the semantic association between the main graph unit and the supplementary intent unit, the fusion type can be accurately determined, ensuring that the fusion logic conforms to the actual context semantics. At the same time, by utilizing the target object corresponding to the supplementary intent unit and the location entity association in the cockpit scene knowledge graph, precise spatiotemporal triggering conditions can be generated, ensuring that the retained supplementary intent can be activated in a timely manner when the vehicle travels to a specific location or meets specific time conditions, avoiding permanent information loss. Combining the main intent unit, supplementary intent unit, fusion type, and spatiotemporal triggering conditions into a complete composite intent unit not only facilitates unified system management and execution, but also enables a comprehensive understanding and dynamic response to the user's composite intent, thereby improving the intelligence level of human-computer interaction and user experience in the smart cockpit.

[0121] In some of the embodiments described above in this application, an arbitration of structured intent units in the second-level intent unit set is proposed to determine the main graph unit and the suppressed intent unit. However, in its implementation, when the confidence levels of multiple intent units are similar, simple arbitration based solely on the security level or the original confidence level may lead to decision bias, making it impossible to optimize the selection using the user's historical preferences, thereby affecting the subsequent activation accuracy of the suppressed intent unit.

[0122] In response, this application further proposes to arbitrate the structured intent units in the second-level intent unit set to determine the idea graph units and the suppressed intent units, including: Based on the security level of each structured intent unit in the second-level intent unit set, a security priority arbitration is performed, and the structured intent unit with the highest security level is selected as the candidate intent graph unit.

[0123] Based on the semantic association information of the scene and the candidate idea graph unit, the confidence of the remaining structured intention units in the second-level intention unit set is arbitrated to obtain the dynamic confidence of each remaining structured intention unit.

[0124] In the case of a structured intent unit where the difference between the dynamic confidence level and the original confidence level of the candidate intent unit is less than a preset deviation threshold, the intent unit is determined from the candidate intent unit and the structured intent unit based on the user's historical preference data, and the remaining structured intent units are used as the suppressed intent unit.

[0125] In the absence of a structured intention unit whose dynamic confidence level differs from the original confidence level of the candidate intention unit by less than the preset deviation threshold, the candidate intention unit is used as the intention unit, and other structured intention units in the second-level intention unit set are used as the suppressed intention units.

[0126] Safety level is an indicator that measures the potential impact of an intent unit's execution on the safety of occupants or the vehicle. It can be preset or dynamically evaluated based on intent type (e.g., navigation commands typically have a higher safety level, while entertainment commands may have a lower), target object (e.g., targets related to vehicle control have a higher safety level), and current driving context (e.g., the safety level of all intents may be dynamically increased during high-speed driving). For example, safety level can be defined as a discrete value (e.g., 1-5, with level 1 being the highest), or as a continuous numerical representation of its potential risk. Safety priority arbitration refers to the mechanism that prioritizes the intent unit with the highest safety level when multiple intent units conflict. Its role is to ensure that the cockpit system's decisions always prioritize the safety of occupants and the vehicle. One implementation is to directly compare the safety levels of each intent unit and select the one with the highest value. Another implementation is to assign different weights to different safety levels, giving higher-safety-level intents a higher priority during comprehensive evaluation. Candidate intent units are the intent units initially determined as having the highest priority after safety priority arbitration. They serve as the benchmark for subsequent arbitration processes, representing the safest intent direction at present.

[0127] Scene semantic association information refers to data extracted from the cockpit scene knowledge graph that describes the relationships between different intent units or between intent units and the cockpit environment. This can include path associations, temporal associations, and command conflict relationships. This information provides richer context during arbitration, for example, indicating whether two intents are spatially related, have a temporal sequence, or have direct command conflicts. Confidence arbitration refers to the process of using the confidence level of intent units to assist in decision-making, based on consideration of security levels. Its purpose is to select the most likely solution that matches the user's true intent by assessing the reliability of intents when security levels are similar or security priorities have been determined. One implementation method is to modify the original confidence level based on preset rules or models and combined with scene semantic association information. Another implementation method is to use a machine learning model, inputting intent features and scene information, and directly outputting the modified confidence level. Dynamic confidence level refers to the confidence level obtained after modifying the original confidence level of the intent unit by combining scene semantic association information during the confidence arbitration process. Compared to raw confidence, dynamic confidence better reflects the actual reliability and rationality of an intention in the current cockpit scenario. For example, when an intention is highly relevant to the current scenario, its dynamic confidence may be increased; conversely, it may be decreased.

[0128] The preset deviation threshold is a value used to determine whether the difference between the dynamic confidence level and the original confidence level of two intent units is small enough. Its function is to define the range of "similar confidence levels." When the difference is less than this threshold, it indicates that a clear decision cannot be made based solely on confidence level, and a deeper level of judgment is needed. This threshold can be set through expert experience or dynamically adjusted through historical data analysis and optimization algorithms. User historical preference data refers to the information recorded by the system regarding the intent units selected by drivers and passengers in similar conflict scenarios in the past. This data reflects the user's personalized decision-making habits and preferences. For example, it can be stored as a rule such as "when intent A and intent B conflict, the user usually chooses A," or represented by a more complex model (such as a recommendation model based on user behavior sequences). Its function is to provide personalized decision-making basis when confidence levels are similar, making the system behavior more in line with user expectations. The intent graph unit refers to the intent unit that the system determines needs to be executed immediately after multi-level arbitration. It is a major component of the current cockpit application's execution instruction set. Suppressed intent units refer to intent units that are not selected as primary intent units during the arbitration process, but whose semantic information is not discarded but is temporarily cached and will be activated again when subsequent spatiotemporal conditions are met.

[0129] Through the above technical solution, this application introduces a more refined and intelligent arbitration mechanism when handling multimodal intent conflicts, solving the decision-making bias problem that may arise from simple arbitration when the confidence levels of multiple intent units are similar. By prioritizing safety arbitration based on safety levels, it ensures that the cockpit prioritizes responses to the intents with the greatest impact on the safety of passengers and the vehicle under any circumstances, laying a safe foundation for subsequent decisions. Introducing scene semantic association information to arbitrate the confidence levels of remaining intent units generates dynamic confidence levels, making the reliability assessment of intents no longer static but incorporating contextual information from the current cockpit environment, thereby improving the accuracy of confidence assessment. More importantly, when the difference between the dynamic confidence level and the original confidence level of the candidate intent graph unit is less than a preset deviation threshold, decisions can be made based on the user's historical preference data. This greatly enhances the personalization of the arbitration results and user satisfaction, avoiding decisions that contradict user habits in ambiguous situations. This multi-level, dynamic arbitration strategy, which incorporates user preferences, not only ensures the accuracy of the main idea graph unit but also allows the suppressed intention unit to be reasonably cached, providing accurate input for subsequent spatiotemporal linkage matching and activation. This avoids the problem of semantic information being completely discarded in traditional solutions, and achieves a more comprehensive understanding and response to the user's complex intentions.

[0130] In some of the embodiments described above in this application, confidence-based arbitration is proposed for selecting intention graph units. However, in this process, relying solely on the original confidence may ignore the semantic relationships between intentions, resulting in inaccurate arbitration results and failure to utilize information about suppressed intentions.

[0131] To address this, this application further proposes a step of arbitrating the structured intent units in the second-level intent unit set to determine the idea graph units and suppressed intent units. Specifically, based on scene semantic association information and candidate idea graph units, confidence arbitration is performed on the remaining structured intent units in the second-level intent unit set to obtain the dynamic confidence of each remaining structured intent unit. This includes: obtaining the original confidence of each remaining structured intent unit and the scene semantic association information between each remaining structured intent unit and the candidate idea graph unit; determining the semantic association correction factor corresponding to each remaining structured intent unit based on the scene semantic association information between each remaining structured intent unit and the candidate idea graph unit; this semantic association correction factor is negatively correlated with the node distance in the scene semantic association information and positively correlated with the path association relationship; and the product of the original confidence and the semantic association correction factor is used as the dynamic confidence of each remaining structured intent unit.

[0132] This involves acquiring the initial confidence level of each remaining structured intent unit and the scene semantic association information between each remaining structured intent unit and candidate intent graph units. The initial confidence level refers to the initial credibility or probability value of each structured intent unit when it is instantiated after semantic understanding of the multimodal interaction data. This value reflects a preliminary assessment of the accuracy of the intent unit's identification, providing a basic, uncontextualized reliability basis for subsequent intent arbitration. It can be obtained by using the probability values ​​synchronously output by the intent classifier when parsing intent type and target object, or by assigning initial values ​​based on preset rules combined with factors such as the clarity and completeness of the modality source. The scene semantic association information refers to structured data extracted from the cockpit scene knowledge graph, used to describe the relationships between different structured intent units or between intent units and cockpit scene entities. According to the above implementation method, this information specifically includes path association relationships, temporal association relationships, and instruction conflict relationships. Its function is to provide contextual connections between intent units, helping the system understand the mutual influence and potential conflicts of intent units in a specific cockpit scene. The acquisition method can be either by querying the predefined relationships between nodes corresponding to intent units in the cockpit scene knowledge graph, or by dynamically generating the knowledge through a graph inference algorithm.

[0133] Based on the scene semantic association information between each remaining structured intent unit and candidate idea graph unit, a semantic association correction factor is determined for each remaining structured intent unit. This semantic association correction factor is negatively correlated with the node distance in the scene semantic association information and positively correlated with the path association relationship. This semantic association correction factor is a numerical value used to adjust the original confidence of the structured intent unit; it quantifies the semantic association strength and consistency between the intent unit and the candidate idea graph unit. The role of this factor is to incorporate the contextual relationship between intent units into their confidence evaluation, allowing semantically more relevant and consistent intents to receive higher weights, and vice versa. The node distance refers to the length of the shortest path between the nodes corresponding to two structured intent units in the cockpit scene knowledge graph. The smaller the node distance, the closer the semantic association between the two intent units. The calculation method can be to find the shortest path in the knowledge graph using breadth-first search (BFS) or depth-first search (DFS) algorithms, and use the number of edges or the sum of edge weights on the path as the distance. The path association refers to whether a predefined, semantically specific connection path exists between nodes corresponding to two structured intent units in the cockpit scene knowledge graph. For example, one intent is "play music" and another is "adjust volume," and there might be a "control" path association between them. This relationship can be predefined in the knowledge graph or dynamically identified based on the graph structure. When a path association exists, it usually means that the two intents are functionally or logically mutually supportive or complementary.

[0134] The product of the original confidence score and the semantic association correction factor is used as the dynamic confidence score for each remaining structured intent unit. This dynamic confidence score is a new confidence value obtained by correcting the original confidence score after considering the scene semantic association between the intent unit and candidate intent graph units. This value more accurately reflects the actual credibility and importance of the intent unit in the current multi-intent conflict scenario. Its purpose is to provide a more comprehensive and context-aware evaluation basis for subsequent arbitration decisions, avoiding misjudgments that may occur based solely on the original confidence score. The calculation method directly uses multiplication, that is, multiplying the original confidence score by the semantic association correction factor to obtain a value that comprehensively considers the initial recognition accuracy and semantic association strength.

[0135] By introducing a dynamic confidence mechanism, this application addresses the problem that relying solely on the original confidence level might overlook the semantic connections between intents, leading to inaccurate arbitration results. For example, it obtains the original confidence level of each remaining structured intent unit and its scene semantic connection information with candidate intent graph units. The original confidence level provides preliminary reliability for intent recognition, while the scene semantic connection information reveals deeper relationships between intent units. Based on this scene semantic connection information, a semantic connection correction factor is determined for each remaining structured intent unit. This correction factor is cleverly negatively correlated with the node distance in the knowledge graph; that is, the closer the intents are semantically, the larger the correction factor. Simultaneously, it is positively correlated with path association; that is, intents with functional or logical connections also have larger correction factors. This design allows the correction factor to accurately quantify the semantic coordination or conflict degree between intents. Multiplying the original confidence level by this semantic connection correction factor yields the dynamic confidence level of each remaining structured intent unit. In this way, dynamic confidence not only reflects the accuracy of intent recognition itself, but also incorporates its contextual association with other intents in the current cockpit scenario. This allows the arbitration process to more intelligently and accurately assess the priority of each intent unit. This avoids simply discarding the semantic information of suppressed intents, but instead provides a more refined basis for subsequent intent activation and scenario adaptation by dynamically adjusting its confidence, thereby improving the intelligence level of dynamic scheduling and scenario adaptation in cockpit applications and enhancing the user experience.

[0136] In some of the solutions mentioned above in this application, the idea graph unit is determined based on the difference in confidence level and the difference in security level. However, in the implementation process, when the difference in confidence level is small, relying solely on these factors may not fully take into account the user's historical preferences, resulting in arbitration results that are not personalized or accurate enough.

[0137] To address this, this application further proposes determining idea graph units from candidate idea graph units and structured intention units based on user historical preference data. Specifically, this involves: matching the combination of intention units (consisting of candidate idea graph units and structured intention units) with user historical preference data to obtain historical adoption records corresponding to these combinations. If historical adoption records exist, the intention units adopted by the user in these records are used as idea graph units. If historical adoption records do not exist, the idea graph units are determined based on the differences in initial confidence and security levels between candidate idea graph units and structured intention units.

[0138] User historical preference data refers to a collection of data accumulated by the system over a long period, reflecting users' actual choices and behaviors in specific multimodal intent conflict scenarios. This data typically includes the correspondence between combinations of intent units in historical conflict scenarios and the intent units adopted by the user. For example, when a conflict is identified between the voice intent of "navigate to the company" and the visual intent of "focus on roadside restaurants," if the user chooses "navigate to the company," this record will be stored. User historical preference data can be obtained through explicit confirmation by the user after system arbitration (such as clicking a confirmation button) or implicit behavior (such as subsequent operations related to a certain intent). Figure 1 This data can be collected through methods such as [unclear - possibly "interaction" or "collection"], or by using machine learning to model user behavior patterns in similar conflict scenarios to predict user preferences. This data can be stored in local storage or on a cloud server and associated with the user ID.

[0139] The combination of candidate idea graph units and structured intention units defines the set of two or more intention units that need to be arbitrated. In the second-level intention unit set of hierarchical arbitration, after security priority arbitration, there will be one candidate idea graph unit and one or more remaining structured intention units that conflict with that candidate idea graph unit. Here, "intention unit combination" specifically refers to the pairing of a candidate idea graph unit with one of the remaining structured intention units. The currently pending candidate idea graph unit is paired with each of the conflicting structured intention units to form the intention unit combination to be queried.

[0140] Matching within user historical preference data aims to identify historical decision patterns that are similar to or identical to the current intent unit combination from massive amounts of historical data. This matching can employ exact matching, which searches for historical records that are completely identical to the current intent unit combination (including key information such as intent type and target object). Alternatively, it can use fuzzy matching or semantic similarity-based matching, such as using the semantic vectors of intent units to calculate similarity and identify semantically close historical conflict scenarios.

[0141] Historical adoption records refer to historical entries in a user's historical preference data that successfully match the current intent unit combination. Each record typically includes the conflicting intent unit combination, the intent unit the user chose to execute, and possible contextual information (such as time, location, vehicle status, etc.), providing personalized reference based on the user's past behavior for the current arbitration decision. The intent unit adopted by the user refers to the intent unit that the user chose to execute when faced with a specific intent conflict in the historical adoption record. It directly reflects the user's true preference in similar situations and is the core basis for personalized arbitration. The system directly extracts the intent units marked as "adopted by the user" from the matched historical adoption records.

[0142] The original confidence level difference and security level difference are the default or backup mechanisms for arbitration, used when historical preference data cannot provide clear guidance. The original confidence level difference refers to the numerical difference between the original confidence levels of the two intent units, reflecting the certainty of intent identification. The security level difference refers to the numerical difference between the security levels of the two intent units, reflecting the potential security risks of executing the intent. These differences serve as a fallback mechanism, ensuring that even in the absence of historical preference data, the system can still make reasonable arbitration based on the accuracy of intent identification and potential risks. In practice, the difference is calculated by comparing the original confidence level and security level values ​​of the two intent units, and a decision is made according to preset rules.

[0143] By incorporating historical user preference data, this application can more accurately identify intention graph units, solving the problem that relying solely on confidence and security levels may lead to insufficiently personalized or accurate arbitration results when confidence differences are small. For example, based on the current combination of intent units to be arbitrated, matching is performed in historical user preference data to obtain historical adoption records related to the user's past behavior patterns. When historical adoption records exist, the system directly uses the intent units adopted by the user in similar conflict scenarios as intention graph units. This ensures that arbitration decisions fully respect the user's personalized preferences, improves the consistency and satisfaction of the user experience, and avoids situations where system decisions do not match user habits. Even if historical adoption records are not available, the system can revert to an arbitration mechanism based on the original confidence and security level differences, ensuring the robustness and security of the decision and maintaining reliable operation under various conditions. This arbitration strategy, which combines historical preferences with real-time judgment, enables the cockpit application's dynamic scheduling and scenario adaptation to not only make safe and reasonable decisions when handling multimodal intent conflicts, but also provide highly personalized and intelligent services, thereby better meeting the user's complex intent needs and preventing valuable intent information from being simply discarded.

[0144] In some of the solutions mentioned above in this application, a method is proposed to determine the idea graph unit based on the difference in original confidence level and the difference in security level when user historical preference data is missing. However, in this process, there is uncertainty about how to accurately weigh the security level and confidence level to make a reliable decision, which may lead to subjective or inconsistent intention choices and affect the accuracy and security of cockpit application scheduling.

[0145] To address this, this application proposes a scheme for determining a concept graph unit, specifically including: determining the concept graph unit based on the difference in original confidence level and the difference in security level between the candidate concept graph unit and the structured intent unit, including: if the difference in security level is greater than a preset security threshold, selecting the intent unit with the higher security level as the concept graph unit; if the difference in security level is less than or equal to the preset security threshold, selecting the intent unit with the higher original confidence level as the concept graph unit.

[0146] The safety level difference refers to the numerical difference in safety levels associated with two intent units to be arbitrated. This difference is a key indicator for assessing the degree of conflict in safety priorities between intent units, and its magnitude directly reflects the difference in safety risks that different intent executions may bring. For example, safety levels can be quantified as integer values ​​from 1 to 5, where 5 represents the highest safety level and 1 represents the lowest safety level; the safety level difference is then the absolute value of the difference between these two integer values. Alternatively, safety levels can also be continuous values ​​calculated based on a risk assessment model; in this case, the safety level difference is the difference between these continuous values. The preset safety threshold is a pre-set numerical limit used to determine whether the safety level difference is sufficient to decide whether to prioritize safety factors in arbitration. This threshold can be configured according to the safety standards of the cockpit application, industry regulations, or the risk tolerance of actual driving scenarios. For example, a fixed value, such as "2," can be set, indicating that when the safety levels of two intent units differ by more than 2 levels, safety factors have absolute priority. Alternatively, the threshold can be dynamically adjusted; for example, in high-speed driving or complex road conditions, the preset safety threshold can be increased to ensure that decisions are made under stricter safety standards. A higher safety level intent unit refers to the intent unit assigned a higher safety level value among two compared intent units. The safety level is typically associated with the intent type, the target object, and its potential impact on occupant or vehicle safety. For example, the safety level of the "emergency braking" intent is much higher than that of the "play music" intent. In arbitration, identifying a higher safety level intent unit means prioritizing the intent with the least or most beneficial impact on occupant and vehicle safety. The difference in original confidence scores refers to the numerical difference in the original confidence scores associated with two intent units to be arbitrated. Original confidence is a quantitative assessment of the accuracy of intent recognition during the semantic understanding stage, reflecting the degree of confidence in the authenticity of the intent. The larger the difference, the greater the uncertainty regarding the reliability of the two intents. A higher original confidence level intent unit refers to the intent unit assigned a higher original confidence value among two compared intent units. Original confidence is typically a probability value between 0 and 1; a higher value indicates a more accurate and reliable recognition of the intent. When there are differences in security levels, the intention unit with higher original confidence is selected to prioritize the execution of instructions that the system believes are most likely to accurately reflect the user's true intentions, thereby improving the accuracy of cockpit application scheduling.

[0147] Through the above technical solution, this application provides an objective and robust intent arbitration mechanism for scenarios where user historical preference data is lacking. This mechanism introduces a preset safety threshold, clarifying the priority of safety level and initial confidence level in the decision-making process. When the safety level difference between two intent units (i.e., greater than the preset safety threshold) is greater than the preset safety threshold, the system prioritizes the intent unit with the higher safety level as the main graph unit. This ensures that when potential safety risks exist, cockpit application scheduling prioritizes the safety of passengers and the vehicle, avoiding safety hazards caused by misjudgment or suboptimal selection. Conversely, when the safety level difference is not greater than or equal to the preset safety threshold, the system selects the intent unit with the higher initial confidence level as the main graph unit. This allows for the priority execution of user intents with higher recognition accuracy when safety risks are comparable, thereby improving the accuracy of cockpit application scheduling and user experience. This hierarchical decision-making logic solves the problem of balancing safety and accuracy in intent selection when historical data is unavailable, avoiding subjective judgment and inconsistent intent selection, making cockpit application scheduling more intelligent and reliable.

[0148] In some of the solutions mentioned above in this application, spatiotemporal triggering conditions are proposed to delay the activation of suppressed intent units. However, in this process, if the path association and timestamp in the target object and scene semantic association information of the suppressed intent unit are not accurately utilized, the determination of the trigger position and time window may be inaccurate or inefficient, and the real-time vehicle position and driver and passenger status may not be effectively matched, thus causing invalid activation or missing the activation opportunity.

[0149] To address this, this application further proposes generating spatiotemporal triggering conditions corresponding to suppressed intention units based on the suppressed intention unit and scene semantic association information. This process includes: determining the triggering location corresponding to the suppressed intention unit based on the target object of the suppressed intention unit and the path association relationship in the scene semantic association information; the triggering location being a mapping point of the target object on a preset navigation path or the geographical location of the target object; determining the triggering time window corresponding to the suppressed intention unit based on the timestamp of the suppressed intention unit and the temporal association relationship in the scene semantic association information; and combining the triggering location with the triggering time window to obtain the spatiotemporal triggering conditions corresponding to the suppressed intention unit.

[0150] For example, the target object of the suppressed intent unit refers to the entity that the suppressed intent unit points to or is associated with, such as a specific location (e.g., a restaurant, gas station), an in-vehicle device (e.g., a window, air conditioner), or a driver or passenger. This target object provides a specific spatial reference for determining the trigger location. In practical applications, the target object can be a predefined point of interest (POI), whose geographic coordinates are stored in a map database. Alternatively, the target object can also be an identifiable physical entity within the cabin, identified and located by sensors within the cabin (e.g., cameras).

[0151] The path associations in the semantic association information of this scenario refer to the spatial or logical connections between different entities (including the intended target object) in the cockpit scene knowledge graph, especially connections related to the vehicle's driving path. These relationships are used to help determine the trigger location, particularly in navigation scenarios, by associating the target object with the vehicle's driving path. These path associations can be represented as edges between nodes in the knowledge graph, such as "located on...path" or "passing near...," and include attributes such as distance and direction. Alternatively, these path associations can be dynamically generated by analyzing the topological relationship between the target object and the current vehicle navigation path, for example, by determining whether the target object is within the buffer zone of the currently planned path.

[0152] When determining the trigger location for a suppressed intent unit, this location refers to the geographic point or region used to activate the suppressed intent unit in subsequent spatiotemporal linkage matching. This aims to precisely define the spatial conditions for intent activation, avoiding premature or late activation. When the target object is near a preset navigation path, the nearest projection point of the target object on the navigation path can be used as the trigger location. This can be achieved by calculating the vertical distance from the target object's geographic coordinates to the navigation path segment and the projection point. Alternatively, when the target object itself has explicit geographic coordinates (such as a POI) and does not need to be associated with a specific navigation path, these geographic coordinates can be directly used as the trigger location.

[0153] The timestamp of the suppressed intent unit refers to the system time record when the original multimodal interaction data was collected and a structured intent unit was generated. This timestamp provides the original time information of the intent's occurrence, serving as a benchmark for determining the trigger time window. The timestamp can be the system's unified UTC time or local time, accurate to the millisecond level. It can also include the start and end times of the intent's duration, forming a time period.

[0154] The temporal relationships in the semantic association information of this scenario refer to the temporal order or temporal dependency between different entities or events in the cockpit scenario knowledge graph. This relationship is used to correct or expand the trigger time determined based on timestamps, making it more consistent with the dynamic changes of the actual scenario. This temporal relationship can be represented as "before," "after," or "occurring simultaneously" relationships between event nodes in the knowledge graph, and may include time interval attributes. Alternatively, this temporal relationship can be defined by analyzing historical data or preset rules, such as "5 minutes before arriving at a certain location" or "immediately after an event occurs."

[0155] When determining the trigger time window corresponding to the suppressed intent unit, this window refers to the time period used to activate the suppressed intent unit in subsequent spatiotemporal linkage matching. Its purpose is to precisely define the time conditions for intent activation and avoid activation at inappropriate times. Based on the original timestamp and temporal correlation, a relative time window can be calculated, such as "original timestamp + preset delay time" or "original timestamp + dynamic delay calculated based on temporal correlation." This trigger time window can be an absolute time period (e.g., "14:30-14:45") or a relative time period (e.g., "within a 5-minute drive from the trigger location").

[0156] The trigger location is combined with the trigger time window to obtain the spatiotemporal trigger condition corresponding to the suppressed intent unit. This combination integrates the trigger conditions in the spatial and temporal dimensions into a unified activation condition, forming a complete activation rule that can be used for subsequent matching, ensuring that the suppressed intent unit is activated "when and where". This spatiotemporal trigger condition can be represented as a data structure containing the geographic coordinates of the trigger location, the start and end times of the trigger time window, and possible trigger states (such as vehicle speed, occupant status). Alternatively, the spatiotemporal trigger condition can be a logical expression, such as "when the vehicle position is within the [distance threshold] of [trigger location] AND the current time is within the [trigger time window]".

[0157] Through the above technical solution, this application can accurately generate the spatiotemporal triggering conditions of suppressed intent units, solving the problem of inaccurate or inability to activate suppressed intent units in traditional solutions. For example, by utilizing the path association relationship between the target object of the suppressed intent unit and the scene semantic association information, the intent can be accurately associated with the actual geographic space, such as mapping the target object to a preset navigation path or directly using its geographic location, thereby ensuring the accuracy of the trigger location. At the same time, by combining the timestamp of the intent occurrence and the temporal association relationship in the scene semantic association information, a reasonable time window can be dynamically determined, avoiding the limitations of fixed-time triggering. This close integration of spatial and temporal dimensions enables the subsequent activation process to more accurately match real-time vehicle location information and driver / passenger status information, improving the activation success rate of suppressed intent units and user experience, avoiding invalid activation or missed activation opportunities, thereby achieving effective maintenance and delayed utilization of the user's complex intents.

[0158] In some of the solutions mentioned above in this application, a spatiotemporal linkage matching based on the intent cache queue to be activated is proposed to activate the suppressed intent unit. However, in this process, the matching mechanism may not be able to effectively combine real-time vehicle location and driver and passenger status information, resulting in inaccurate activation timing or missed opportunities, and cannot ensure that the suppressed intent is accurately activated under appropriate spatiotemporal conditions.

[0159] In response, this application further proposes a method for performing spatiotemporal linkage matching on the suppressed intent units in the intent cache queue to be activated, based on the semantic association information of the scene and the real-time collected vehicle location information, to generate activation instructions and add them to the execution instruction set.

[0160] For example, see Figure 5 The method includes: 501. Obtain the spatiotemporal triggering conditions corresponding to each suppressed intent unit in the intent cache queue to be activated.

[0161] This step aims to provide a clear triggering basis for subsequent spatiotemporal linkage matching. Spatiotemporal trigger conditions are pre-defined for each suppressed intent unit, indicating when and where that intent unit can be reactivated. For example, the spatiotemporal trigger conditions associated with the suppressed intent unit can be read directly from the pending intent cache queue. Alternatively, the corresponding spatiotemporal trigger conditions can be dynamically retrieved from a separate trigger condition database or configuration table by querying the metadata or identifier associated with the suppressed intent unit.

[0162] 502. Based on the real-time collected vehicle location information and the trigger position in each spatiotemporal trigger condition, determine the spatial matching degree of each suppressed intent unit.

[0163] This step assesses the degree to which the vehicle's current location matches the preset trigger location of the suppressed intent unit, quantifying the extent to which the vehicle is geographically close to or has reached the trigger location. For example, this can be achieved by calculating the Euclidean distance or navigation path distance between the vehicle's current GPS coordinates and the trigger location's GPS coordinates; a smaller distance indicates a higher spatial match. Alternatively, it can be determined whether the vehicle's current location has entered a preset geofence area centered on the trigger location, and quantified based on the depth of entry into the area.

[0164] 503. Based on the current state information of the driver and passengers and the triggering state in each spatiotemporal triggering condition, determine the state matching degree of each suppressed intention unit.

[0165] This step aims to assess whether the current physiological or psychological state of the driver or passenger meets the activation conditions of the suppressed intention unit, reflecting the appropriateness of the driver or passenger when a specific intention is activated. For example, it can be done by analyzing the driver's or passenger's facial expressions, heart rate, eye movements, and other biometric data, combined with preset trigger states (e.g., non-fatigue, non-distraction, pleasure, etc.), to determine the degree of consistency between the current state and the trigger state. Alternatively, it can be done by combining driver and passenger behavior data (e.g., posture, tone of voice) collected by in-vehicle sensors (e.g., cameras, microphones), using machine learning models to identify the driver's or passenger's current emotion or level of focus, and comparing it with the trigger state.

[0166] 504. The spatial matching degree and the state matching degree are weighted and fused to obtain the spatiotemporal linkage matching degree of each suppressed intention unit.

[0167] This step integrates matching information from both spatial and state dimensions to form a unified and more comprehensive evaluation metric, avoiding the limitations of single-dimensional judgment. For example, a linear weighted summation method can be used, where the spatiotemporal linkage matching degree equals the spatial matching degree multiplied by a preset spatial weight plus the state matching degree multiplied by a preset state weight. The weights can be adjusted according to the importance of the actual application scenario. Alternatively, a nonlinear fusion model can be used, such as one based on a neural network or fuzzy logic system, taking the spatial and state matching degrees as input and outputting the spatiotemporal linkage matching degree to capture more complex relationships.

[0168] 505. Add the cockpit application control commands corresponding to the suppressed intent units whose spatiotemporal linkage matching degree exceeds the preset activation threshold as activation commands to the execution command set.

[0169] This step is the decision-making stage, where the decision to activate the suppressed intent unit is based on a comprehensive assessment of the spatiotemporal linkage matching degree. Only when the matching degree reaches a sufficiently high level is the intent unit considered safe and effective for activation. For example, a simple numerical comparison can be used, comparing the spatiotemporal linkage matching degree with a fixed preset activation threshold; if it exceeds the threshold, an activation command is generated. Alternatively, a dynamic threshold adjustment mechanism can be employed, adjusting the preset activation threshold in real time based on factors such as the current driving situation, vehicle speed, and occupant safety level to adapt to different activation needs and safety considerations.

[0170] Through the above technical solution, this application can obtain the spatiotemporal triggering conditions corresponding to each suppressed intention unit, providing a clear activation standard for subsequent matching. Based on the real-time collected vehicle location information and the triggering position in each spatiotemporal triggering condition, the spatial matching degree of each suppressed intention unit can be accurately determined, ensuring that the activation of the intention is closely integrated with the actual driving environment of the vehicle. At the same time, based on the current state information of the driver and passengers and the triggering state in each spatiotemporal triggering condition, the suitability of the driver and passengers when triggering a specific intention can be evaluated, avoiding activation of the intention in an unsuitable state such as driver and passengers being fatigued or distracted, thereby improving the safety of the interaction and the user experience. By weightedly fusing the spatial matching degree and the state matching degree, the spatiotemporal linkage matching degree of each suppressed intention unit is obtained. This application comprehensively considers the two key dimensions of space and driver and passenger state, avoiding the bias that may be caused by a single dimension judgment, making the activation decision more comprehensive and accurate. The cockpit application control commands corresponding to the suppressed intent units whose spatiotemporal linkage matching degree exceeds the preset activation threshold are used as activation commands and added to the execution command set. This ensures that the suppressed intent is activated only when both the space and the status of the driver and passengers meet the conditions, preventing accidental activation or waste of resources. It achieves accurate and timely activation of suppressed intent, thereby solving the problems of inaccurate activation of suppressed intent units, inaccurate activation timing, or missed opportunities in traditional solutions, and improving the intelligence of the intelligent cockpit system and user satisfaction.

[0171] In some of the embodiments described above in this application, a method is proposed to determine the spatial matching degree based on vehicle location information and trigger location to evaluate the activation conditions of the suppressed intention unit. However, in this process, considering only spatial distance may lead to inaccurate matching and cannot effectively combine the vehicle driving direction for accurate evaluation, thereby affecting the activation timing and reliability of the suppressed intention.

[0172] In response, this application further proposes a method for determining the spatial matching degree of each suppressed intent unit based on real-time collected vehicle location information and the trigger positions in various spatiotemporal trigger conditions. The specific steps include: Based on the vehicle's current geographic coordinates and the target geographic coordinates corresponding to the trigger location, the spatial distance is calculated. Spatial distance is a quantitative indicator measuring the physical proximity between the vehicle's current location and the target trigger location. Its purpose is to provide a basic measure of positional difference to determine whether the vehicle is sufficiently close to the target trigger point. In practical applications, Euclidean distance can be used, which calculates the straight-line distance on a two-dimensional plane using the Pythagorean theorem based on the difference in latitude and longitude between two geographic coordinate points. Alternatively, great-circle distance, such as the Haversine formula or Vincenty formula, can be used to calculate the shortest distance between two points on the Earth's surface, taking into account the Earth's curvature. This is more accurate for long distances or high-precision requirements.

[0173] Based on the vehicle's current driving direction and the azimuth angle of the target's geographical coordinates relative to the current geographical coordinates, a direction consistency coefficient is calculated. This coefficient assesses the degree of matching between the vehicle's current driving direction and the azimuth angle of the target trigger position relative to the vehicle's current position. Its purpose is to introduce a directional factor, preventing erroneous activation of intents when the vehicle's orientation is opposite to or deviates from the target position, thereby improving the accuracy and rationality of the matching. Specifically, the direction consistency coefficient can be obtained by calculating the cosine of the angle between the vehicle's current driving direction vector and the vector pointing from the current position to the target position. A smaller angle and a larger cosine value indicate greater direction consistency. Alternatively, a piecewise function can be defined to classify the direction consistency coefficient into different levels based on the angle difference between the two directions. For example, an angle difference within a certain threshold indicates high consistency, while a difference exceeding this threshold indicates low or inconsistent consistency.

[0174] The spatial distance and the directional consistency coefficient are fused to obtain the spatial matching degree. The purpose of fusing spatial distance and the directional consistency coefficient is to comprehensively consider both positional proximity and directional matching degree, thus obtaining a more comprehensive and accurate spatial matching degree. This fusion avoids the limitations of single-dimensional evaluation, ensuring that suppressed intent units are only activated when the vehicle is both close to and facing the target. One fusion method is a weighted average method, where weights are assigned to both the spatial distance (or its normalized value) and the directional consistency coefficient, and their weighted sum is used as the spatial matching degree. The weight settings can be adjusted based on actual application scenarios and experience. Another fusion method is a product method, where the spatial distance (or its normalized value) is directly multiplied by the directional consistency coefficient to obtain the spatial matching degree. This method reduces the matching degree if either factor fails to meet the requirements.

[0175] By introducing a directional consistency coefficient and fusing it with spatial distance, this application can more accurately assess the spatial activation conditions of suppressed intent units. For example, calculating spatial distance provides a quantitative basis for positional proximity, ensuring that activation judgments are based on actual physical distance. Simultaneously, calculating the directional consistency coefficient effectively compensates for the shortcomings of considering only distance, preventing false activation of intents when the vehicle is traveling in the opposite direction or deviating from the target direction. Fusing these two factors ensures that spatial matching not only reflects the proximity of the vehicle to the target location but also comprehensively considers whether the vehicle's driving direction aligns with the target orientation. This comprehensive evaluation mechanism improves the accuracy and reliability of spatial matching, ensuring that suppressed intent units are activated at the most appropriate time, avoiding misjudgments or omissions caused by single-dimensional evaluation, and thus enhancing the intelligence level of dynamic scheduling and scenario adaptation in cockpit applications and the user experience.

[0176] In some of the embodiments described above in this application, a method is proposed to determine the state matching degree based on the current state information of the driver and passengers and the trigger state to calculate the spatiotemporal linkage matching degree. However, in this process, relying solely on a simple comparison between the state information and the trigger state may lead to inaccurate state matching degree evaluation, and may fail to identify the adaptability of the real-time state type of the driver and passengers to the preset trigger state, thereby affecting the accuracy and timing appropriateness of the activation of the suppressed intention unit.

[0177] To address this, this application further proposes a method for determining the state matching degree of each suppressed intention unit based on the current state information of the driver / passenger and the trigger states in various spatiotemporal trigger conditions. Specifically, this method includes: acquiring biometric data and facial expression data from the current state information of the driver / passenger; determining the current state type of the driver / passenger based on the biometric data and facial expression data, which includes fatigue, distraction, and pleasure; parsing a preset trigger state type from the spatiotemporal trigger conditions; setting the state matching degree to a first value when the current state type matches the preset trigger state type; and setting the state matching degree to a second value, which is less than the first value, when the current state type does not match the preset trigger state type.

[0178] For example, acquiring biometric and facial expression data from the current state information of drivers and passengers aims to capture their real-time physiological and emotional states. Biometric data can include, but is not limited to, physiological indicators such as heart rate, respiratory rate, skin conductance, and electroencephalogram (EEG). This data can be collected in real time through contact sensors integrated into the seats and steering wheel, or wearable devices (such as smart bracelets and smartwatches), or through non-contact sensors (such as millimeter-wave radar and infrared sensors) monitoring physiological signals such as subtle breathing and heartbeat movements. Facial expression data is primarily acquired by capturing facial images or video streams of drivers and passengers through in-cabin cameras. Computer vision technologies, such as deep learning models, are used to detect facial landmarks and identify Action Units (AUs) to quantify the facial features of drivers and passengers. Furthermore, eye-tracking technology can be combined to analyze the driver's blink frequency, pupil size, and eyelid closure, serving as a supplement to the facial expression data.

[0179] Based on the acquired biometric and facial expression data, the system further determines the current state type of the driver and passengers. This current state type can include fatigue, distraction, and pleasure. Various techniques can be used to determine the current state type. For example, a multimodal fusion machine learning model can be constructed, using biometric and facial expression data as input features, and classifying the state through a pre-trained classifier (such as support vector machine, neural network, random forest, etc.). This model, by learning from a large amount of labeled data under different states, can identify states such as fatigue, distraction, or pleasure for the driver and passengers. Another method is based on rule engines and threshold judgments. For example, when the system detects that the driver or passenger blinks excessively frequently, yawns frequently, or has abnormal heart rate fluctuations, it can determine a fatigue state. When the gaze is deviated from the road for a long time, the head posture is abnormal, or the facial expression shows confusion or anxiety, it can be judged as a distraction state. When the facial expression shows positive emotional characteristics such as upturned corners of the mouth and smile lines at the corners of the eyes, it can be judged as a pleasure state. These rules and thresholds can be set and optimized based on expert experience and actual driving data.

[0180] Simultaneously, the system parses the preset trigger state type from the spatiotemporal trigger conditions corresponding to the suppressed intent unit. These spatiotemporal trigger conditions are generated for the suppressed intent unit during the hierarchical arbitration phase and are used to define its spatiotemporal context of activation. The preset trigger state type is a component of these spatiotemporal trigger conditions, specifying under what driver / passenger states the suppressed intent unit is suitable for activation. The parsing process typically involves reading structured data (such as JSON or XML formats) and directly extracting the preset state type identifier from predefined fields.

[0181] After obtaining the current state type and the preset trigger state type, the system performs a matching judgment. If the current state type and the preset trigger state type are completely identical, they are considered a match, and the state matching degree is set to a preset first value. This first value is usually a high value, such as 1.0, indicating a high match, which is beneficial for subsequent activation decisions. Conversely, if the current state type and the preset trigger state type are inconsistent, they are considered a mismatch, and the state matching degree is set to a preset second value. This second value is less than the first value, such as 0.0 or a small positive number, indicating a low matching degree, thereby suppressing inappropriate activation.

[0182] Through the above technical solution, this application can meticulously evaluate the compatibility between the real-time state of the driver and passengers and the state required for the activation of the suppressed intention unit. By acquiring biometric data and facial expression data, it can capture the physiological and emotional changes of the driver and passengers in real time and objectively, providing rich and reliable evidence for state recognition. Based on this multimodal data, it can accurately identify various current state types of the driver and passengers, such as fatigue, distraction, or pleasure, thereby avoiding the inaccuracies caused by relying solely on a single modality or simple rule judgment. The identified current state type is precisely matched with the preset trigger state type parsed from the spatiotemporal trigger conditions, and different state matching degree values ​​are assigned according to the matching results, so that the activation decision of the suppressed intention unit can fully consider the actual state of the driver and passengers. When the state of the driver and passengers is highly consistent with the preset conditions, the state matching degree is high, which helps to improve the spatiotemporal linkage matching degree and prompts the suppressed intention unit to be activated at the most appropriate time. Conversely, when the driver and passengers' status does not meet the activation conditions, the status matching degree is low, which avoids activating the intention in inappropriate situations, thereby improving the intelligent level of dynamic scheduling and scenario adaptation of the cockpit application and the comfort and safety of the user experience.

[0183] The following example will provide a more detailed explanation of the above technical solution: In a typical driving scenario, User A is driving to work. At this time, the in-cabin system, through the coordinated operation of the processor and memory, performs the following operations: The system acquires multimodal interaction data collected within the cockpit. For example, user A issues the voice command "Navigate to the company," while simultaneously maintaining a prolonged gaze on a "coffee shop" sign on the roadside and lightly touching the air conditioning control panel to lower the temperature. This raw multimodal data stream, including voice data, gaze data, and touch data, is collected by the system. Feature extraction and semantic parsing are performed on this data. For instance, the voice and gaze data are temporally aligned, and a cross-modal attention mechanism is used to extract and fuse semantic features. Simultaneously, the system dynamically adjusts the weights of the intent feature vectors by combining vehicle dynamic data (such as vehicle speed and direction of travel) and user A's biometric data (such as heart rate and pupil state) to improve the accuracy and reliability of intent recognition. Based on these intent feature vector sets and pre-defined intent structuring rules, the system instantiates and generates multiple structured intent units. For example, the generated intent unit 1 is: "Navigate to the company" (modal source: speech, original confidence: high, security level: medium), intent unit 2 is: "Follow the coffee shop" (modal source: gaze, original confidence: medium, security level: low), and intent unit 3 is: "Lower the air conditioning temperature" (modal source: touch, original confidence: high, security level: high). Each intent unit includes intent type, target object, modal source, original confidence, timestamp, security level, and semantic vector.

[0184] The conflict level of these structured intent units and the cockpit scene knowledge graph is quantified. Intent units are mapped to the cockpit scene knowledge graph, and path associations, temporal associations, and command conflict relationships are extracted between them. For example, "navigate to the company" and "follow the coffee shop" may have a potential conflict because the coffee shop may not be on the navigation path to the company, or it may require temporarily deviating from the main path. "Lower the air conditioning temperature" has a lower conflict level with the former two. Based on the semantic vector distance, node distance, time urgency difference, safety level difference, and command conflict relationship between intent units, semantic conflict components and temporal safety conflict components are calculated and weighted and fused to obtain a conflict level score between each intent unit. Simultaneously, path associations, temporal associations, and command conflict relationships are used as scene semantic association information.

[0185] The system performs hierarchical arbitration on these structured intent units, conflict scores, and scene semantic association information to generate a cockpit application execution instruction set and a cache queue of intents to be activated. Based on a preset threshold range for the conflict score, the system divides intent units into a first-level intent unit set and a second-level intent unit set. For example, intent unit 3, "lower air conditioning temperature," has a low conflict score and is assigned to the first-level intent unit set. Semantic fusion is performed to generate a composite intent unit, and its corresponding cockpit application control instruction, "Execute: Lower Air Conditioning Temperature," is added to the cockpit application execution instruction set. Intent units 1, "Navigate to Company," and 2, "Follow Coffee Shop," have high conflict scores and are assigned to the second-level intent unit set. Arbitration is then performed on them, prioritizing safety based on security level. Since "Navigate to Company" typically has a higher security level, it is initially selected as a candidate intent unit. Then, the system combines scene semantic association information and user historical preference data for confidence arbitration. Assuming user A's historical preference data shows that in similar scenarios, user A usually prioritizes completing the navigation task but will also subsequently follow up on points of interest along the way. Therefore, "Navigate to the company" is identified as the primary intent unit, and its corresponding instruction "Execute: Start navigation to the company" is added to the cockpit application's execution instruction set. "Follow the coffee shop" is identified as the suppressed intent unit. Unlike existing technologies that directly discard suppressed intents, this approach generates corresponding spatiotemporal triggering conditions based on the suppressed intent unit "Follow the coffee shop" and its contextual semantic association information. For example, the trigger location is determined to be the geographical location of the "coffee shop," and the trigger time window is determined to be "when the vehicle approaches the coffee shop." This suppressed intent unit and its associated spatiotemporal triggering conditions are stored, forming an entry in the pending intent cache queue.

[0186] Based on the pending intent cache queue, scene semantic association information, and real-time vehicle location information, the system performs spatiotemporal linkage matching on suppressed intent units in the cache queue, generates activation commands, and adds them to the execution command set. While the vehicle is in motion, the system continuously acquires real-time vehicle location information. As the vehicle approaches the "coffee shop," it retrieves the spatiotemporal trigger conditions corresponding to the "follow coffee shop" intent from the pending intent cache queue. The system calculates the spatial matching degree based on the vehicle's current location information and the trigger location (coffee shop's geographical location), and determines the state matching degree based on the driver's current state information (e.g., not fatigued, not distracted) and the trigger state (e.g., user not in an emergency driving state). The spatial matching degree and state matching degree are weighted and fused to obtain the spatiotemporal linkage matching degree. When this matching degree exceeds a preset activation threshold, the system generates an activation command, "Reminding user A that there is a coffee shop you have previously followed ahead; do you need navigation to it?" and adds it to the cockpit application execution command set.

[0187] Through the above process, this system avoids the problem in existing technologies where suppressed semantic information is completely discarded due to conflict handling. For example, when user A's voice command "navigate to the company" conflicts with the visual command "focus on the coffee shop," this system does not simply choose to execute the navigation command and abandon the coffee shop intention. Instead, it caches the "focus on the coffee shop" intention and its spatiotemporal triggering conditions, and reactivates it at an appropriate time (when the vehicle approaches the coffee shop and the user's state allows it). This achieves the preservation and delayed utilization of semantic information in multimodal conflict scenarios, improving user experience and system intelligence.

[0188] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0189] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0190] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A cockpit application dynamic scheduling and scenario adaptation system based on multimodal intent fusion, characterized in that, The system includes a processor and memory, the processor being configured to perform the following steps: Acquire multimodal interaction data collected in the cockpit, perform semantic understanding on the multimodal interaction data, and generate multiple structured intent units; The conflict level of the multiple structured intent units and the cockpit scene knowledge graph is quantified to obtain conflict level scores and scene semantic association information; The multiple structured intent units, the conflict degree score, and the scene semantic association information are subjected to hierarchical arbitration to generate a cockpit application execution instruction set and an intent cache queue to be activated. The hierarchical arbitration determines the main graph unit and the suppressed intent unit according to the conflict degree score, adds the instruction corresponding to the main graph unit to the execution instruction set, and adds the suppressed intent unit and its associated spatiotemporal triggering condition to the intent cache queue to be activated. Based on the intent cache queue to be activated, the scene semantic association information, and the real-time collected vehicle location information, the suppressed intent units in the intent cache queue to be activated are spatiotemporally linked for matching, and activation instructions are generated and added to the execution instruction set.

2. The system according to claim 1, characterized in that, The process involves acquiring multimodal interaction data collected within the cockpit, performing semantic understanding on the multimodal interaction data, and generating multiple structured intent units, including: The raw multimodal data stream inside the cockpit is collected, including voice data, gaze data, gesture data, touch data, vehicle dynamics data, biometric data, and external environment data. Feature extraction and semantic parsing are performed on the original multimodal data stream to obtain the intent feature vector set corresponding to each modality; Based on the intent feature vector set and intent structuring rules, the plurality of structured intent units are instantiated and generated, wherein each structured intent unit includes intent type, target object, modality source, original confidence level, timestamp, security level, and semantic vector.

3. The system according to claim 1, characterized in that, The process of quantifying the conflict level of the multiple structured intent units and the cockpit scene knowledge graph to obtain conflict level scores and scene semantic association information includes: The multiple structured intent units are mapped to the cockpit scene knowledge graph to obtain the corresponding node of each structured intent unit in the cockpit scene knowledge graph, and the path association, temporal association and instruction conflict relationship between the corresponding nodes are extracted from the cockpit scene knowledge graph. Based on the node distance between the corresponding nodes, the semantic vector distance between the multiple structured intent units, and the path association relationship, the semantic conflict component is calculated; Based on the differences in time urgency, security level, and instruction conflict relationships among the multiple structured intent units, a timing security conflict component is calculated. The semantic conflict component and the temporal security conflict component are weighted and fused to obtain the conflict degree score, and the path association relationship, the temporal association relationship and the instruction conflict relationship are used as the scene semantic association information.

4. The system according to claim 3, characterized in that, The step of mapping the plurality of structured intent units to the cockpit scene knowledge graph to obtain the corresponding node of each structured intent unit in the cockpit scene knowledge graph includes: Based on the intent type and target object of each structured intent unit, entity matching is performed in the cockpit scene knowledge graph to obtain an initial candidate node set; When the initial candidate node set contains multiple candidate nodes, the candidate node with the highest similarity is selected as the corresponding node based on the similarity between the semantic vector of each structured intent unit and the semantic embedding vector of the multiple candidate nodes. If the initial candidate node set is empty, similar nodes are retrieved in the cockpit scene knowledge graph based on the semantic vector of each structured intent unit, and the retrieved similar nodes are used as the corresponding nodes.

5. The system according to claim 1, characterized in that, The step of hierarchically arbitrating the multiple structured intent units, the conflict degree scores, and the scene semantic association information to generate a cockpit application execution instruction set and a queue of intents to be activated includes: Based on the preset threshold range where the conflict level score is located, the multiple structured intent units are divided into a first-level intent unit set and a second-level intent unit set; Semantic fusion is performed on the structured intent units in the first-level intent unit set to generate composite intent units, and the cockpit application control commands corresponding to the composite intent units are added to the cockpit application execution command set. Arbitrate the structured intent units in the second-level intent unit set to determine the main idea unit and the suppressed intent unit, and add the cockpit application control command corresponding to the main idea unit to the cockpit application execution command set; Based on the suppressed intent unit and the scene semantic association information, a spatiotemporal triggering condition corresponding to the suppressed intent unit is generated, and the suppressed intent unit and the spatiotemporal triggering condition are associated and stored to obtain the intent cache queue to be activated.

6. The system according to claim 5, characterized in that, The arbitration of structured intent units in the second-level intent unit set to determine the main idea unit and the suppressed intent unit includes: Based on the security level of each structured intent unit in the second-level intent unit set, a security priority arbitration is performed, and the structured intent unit with the highest security level is selected as the candidate intent graph unit. Based on the scene semantic association information and the candidate idea graph units, the remaining structured intention units in the second-level intention unit set are arbitrated with confidence to obtain the dynamic confidence of each remaining structured intention unit. In the case of a structured intent unit where the difference between the dynamic confidence level and the original confidence level of the candidate intent graph unit is less than a preset deviation threshold, the intent graph unit is determined from the candidate intent graph unit and the structured intent unit based on the user's historical preference data, and the remaining structured intent units are used as the suppressed intent units. In the absence of a structured intent unit whose dynamic confidence level differs from the original confidence level of the candidate intent graph unit by less than the preset deviation threshold, the candidate intent graph unit is used as the intent graph unit, and other structured intent units in the second-level intent unit set are used as the suppressed intent units.

7. The system according to claim 6, characterized in that, The step of arbitrating the confidence of the remaining structured intent units in the second-level intent unit set based on the scene semantic association information and the candidate intent graph units, to obtain the dynamic confidence of each remaining structured intent unit, includes: Obtain the original confidence level of each remaining structured intent unit and the scene semantic association information between each remaining structured intent unit and the candidate idea graph unit; Based on the scene semantic association information between each remaining structured intent unit and the candidate intent graph unit, a semantic association correction factor corresponding to each remaining structured intent unit is determined. The semantic association correction factor is negatively correlated with the node distance in the scene semantic association information and positively correlated with the path association relationship. The product of the original confidence level and the semantic association correction factor is used as the dynamic confidence level of each remaining structured intent unit.

8. The system according to claim 6, characterized in that, The user historical preference data includes the correspondence between combinations of intent units in historical conflict scenarios and the intent units ultimately adopted by the user. The step of determining the intent graph unit from the candidate intent graph units and the structured intent units based on the user historical preference data includes: Based on the combination of the candidate idea graph unit and the structured intention unit, the user's historical preference data is matched to obtain the historical adoption record corresponding to the combination of intention units; If the historical adoption record exists, the intention unit of the user's final adoption in the historical adoption record shall be used as the intention graph unit; If the historical adoption record does not exist, the idea graph unit is determined based on the difference in original confidence level and the difference in security level between the candidate idea graph unit and the structured idea unit.

9. The system according to claim 1, characterized in that, The process involves performing spatiotemporal matching on suppressed intent units in the intent cache queue based on the intent to be activated, the scene semantic association information, and the real-time collected vehicle location information, generating activation instructions and adding them to the execution instruction set, including: Obtain the spatiotemporal triggering conditions corresponding to each suppressed intent unit in the intent cache queue to be activated; Based on the real-time collected vehicle location information and the trigger position in each spatiotemporal trigger condition, the spatial matching degree of each suppressed intent unit is determined. Based on the current state information of the driver and passengers and the triggering state in each spatiotemporal triggering condition, the state matching degree of each suppressed intention unit is determined. The spatial matching degree and the state matching degree are weighted and fused to obtain the spatiotemporal linkage matching degree of each suppressed intention unit; The cockpit application control commands corresponding to the suppressed intent units whose spatiotemporal linkage matching degree exceeds the preset activation threshold are added to the execution command set as activation commands.

10. The system according to claim 9, characterized in that, The determination of the state matching degree of each suppressed intention unit based on the current state information of the driver and passengers and the trigger state in each spatiotemporal trigger condition includes: Obtain biometric data and facial expression data from the current status information of the driver and passengers; Based on the biometric data and the facial expression data, the current state type of the driver and passengers is determined, including fatigue, distraction, and pleasure. The preset trigger state type is obtained by parsing the spatiotemporal trigger conditions; If the current state type matches the preset trigger state type, the state matching degree is set to a first value; If the current state type does not match the preset trigger state type, the state matching degree is set to a second value, which is less than the first value.