Intelligent cockpit vehicle cloud collaborative intelligent decision-making system and method based on multi-modal perception
The intelligent cockpit vehicle-cloud collaborative intelligent decision-making system with multimodal perception solves the problems of insufficient computing power on the vehicle side and large processing latency on the cloud side, realizes efficient collaboration between the vehicle side and the cloud side, and improves the real-time performance and personalized service capabilities of the intelligent cockpit system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA YOUKE COMM TECH
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-28
AI Technical Summary
In existing in-vehicle intelligent cockpit systems, insufficient vehicle-side computing power, large cloud processing latency, and high communication coupling result in insufficient response speed and stability, making it impossible to achieve efficient personalized services and real-time requirements.
The intelligent cockpit vehicle-cloud collaborative intelligent decision-making system adopts multimodal perception, including a vehicle-side intelligent perception layer, a collaborative communication layer, and a cloud-side intelligent decision-making layer. Through layered processing, separation of data flow and control flow, asynchronous object storage and uploading, and real-time communication protocols, it achieves efficient collaborative decision-making between the vehicle and the cloud.
It achieves efficient collaboration between the vehicle and the cloud, ensuring real-time and personalized decision-making, improving the human-computer interaction experience and system reliability, and ensuring timely response to safety incidents and continuity of information exchange.
Smart Images

Figure CN121929079A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle-mounted intelligent cockpit technology, and in particular to an intelligent cockpit vehicle-cloud collaborative intelligent decision-making system and method based on multimodal perception. Background Technology
[0002] With the continuous enrichment of in-vehicle sensors and smart cockpit functions, vehicles possess multimodal capabilities such as voice interaction, visual perception, and environmental monitoring. However, existing solutions suffer from the following problems: on the one hand, on-vehicle computing power is limited, making it unable to handle large-scale model calculations for extended periods; on the other hand, pure cloud-based models are limited by network latency and bandwidth, resulting in insufficient response speed and stability. Furthermore, traditional data transmission methods often couple large files with control flow, increasing system burden and making it difficult to guarantee real-time performance. Therefore, there is an urgent need for an intelligent decision-making architecture that can work collaboratively between the vehicle and the cloud, ensuring both real-time performance and handling complex tasks, to improve the human-machine interaction experience and system reliability of the in-vehicle cockpit. Summary of the Invention
[0003] This invention proposes an intelligent cockpit vehicle-cloud collaborative intelligent decision-making system and method based on multimodal perception, which solves the problems of insufficient computing power on the vehicle side, large processing latency in the cloud, high communication coupling, and lack of personalized services. It can achieve efficient collaboration, rapid response, and personalized decision-making between the vehicle side and the cloud.
[0004] The present invention adopts the following technical solution.
[0005] A multimodal perception-based intelligent cockpit vehicle-cloud collaborative intelligent decision-making system, the decision-making system comprising:
[0006] The vehicle-side intelligent perception layer is used to execute computer vision models through pluggable devices on the vehicle side, collect multimodal data from driver monitoring, occupant monitoring, cameras, and radar, and generate event data through layered processing.
[0007] The intelligent event triggering module is used to determine events based on detection results and context information, and to encapsulate triggering events in a structured manner and set priorities.
[0008] The collaborative communication layer is used to upload data using a method that separates object storage direct transmission from metadata API, and to provide feedback on task status and results through the SSE real-time communication protocol.
[0009] The cloud-based intelligent decision-making layer receives the event data, selects a fast link, a hybrid link, or a complex link for processing based on the task complexity, and generates safety prompts, in-cabin interactive responses, or personalized recommendation results by combining user profiles and external services.
[0010] The vehicle-side intelligent perception layer is based on a pluggable computer vision model architecture and combines multimodal sensors located inside and outside the vehicle cabin to perceive the environment inside and outside the cabin.
[0011] The vehicle-side intelligent perception layer is used to detect driver fatigue, distraction, emotions and gestures, as well as to identify occupant status and behavior, thereby providing semantic input for cabin services and human-machine interaction. It aligns multimodal data through a unified timestamp and coordinate system, and adopts a hierarchical processing mechanism: L1 basic detection, L2 reporting and decision-making, L3 fine recognition and L4 semantic analysis, to achieve a step-by-step process from low-latency detection to high-level semantic understanding.
[0012] Multimodal sensors inside and outside the vehicle include Driver Monitoring System (DMS), Occupant Monitoring System (OMS), in-cabin microphone array, forward-facing camera, and millimeter-wave radar;
[0013] The vehicle-side intelligent perception layer is based on an intelligent event triggering mechanism, which identifies key event types inside and outside the cabin through rule engine and context information analysis.
[0014] Contextual information includes vehicle operating status, geographical location, time, driver status, user profile, and weather conditions;
[0015] Key event types include safety events (such as driver fatigue, obstacles ahead) and interactive events (such as occupants issuing voice commands, the system detecting entertainment needs);
[0016] The intelligent event triggering mechanism classifies different events into risk levels through a security assessment module. The priority scoring model used for classification sorts events based on risk, time sensitivity, context, and personalized needs, ensuring that security events and high-value interactions are processed first. The data encapsulation module packages events into standardized data packets for further processing in the cloud.
[0017] The collaborative communication layer adopts a "data flow and control flow separation" architecture. Specifically, large-volume media data (images, videos, etc.) are directly transmitted through object storage, supporting segmentation and breakpoint resumption; control messages (object_key, hash, context, idempotent key) are uploaded through metadata API to avoid redundant calculations.
[0018] The collaborative communication layer provides a real-time event stream based on SSE (Server-Sent Events) for continuous feedback on task progress (such as task_created, response_chunk, task_completed), and supports reconnection after disconnection and incremental resending in weak network environments. This design ensures both real-time response to security events and continuity of information and service interaction within the cockpit.
[0019] The cloud-based intelligent decision-making layer includes dual input interfaces: a multimodal direct input interface to process real-time safety events reported by the vehicle to ensure low-latency response; and a central routing input interface to handle complex cross-domain, multi-step tasks to support the orchestration and fusion of complex tasks.
[0020] The cloud-based intelligent decision-making layer uses a cloud-based routing mechanism to score tasks based on a complexity assessment model and selects between fast, hybrid, or complex paths. In some embodiments, the complexity score combines the number of tool calls, multimodal dependencies, and real-time requirements, and can dynamically update thresholds based on historical execution results. The cloud-based multimodal analysis engine calls large-scale models, combining time, geographic location, and user profiles to perform semantic reasoning and personalized recommendations, generating optimized decision results.
[0021] The decision-making system includes a closed-loop execution and correction mechanism. Specifically, in some embodiments, when network degradation or high system load causes complex link delays, the system automatically falls back to the fastest link to return the minimum usable result (such as a basic reminder or brief answer), and reissues the corrected result after the link is restored. The corrected result includes a version number and timestamp, and the cockpit interface displays "Updated / Corrected," ensuring consistency in user perception. Through this closed-loop mechanism of real-time response—supplementary reasoning—consistency correction, both timely handling of security incidents and smooth cockpit interaction are ensured.
[0022] The cockpit domain controller includes at least one processor, memory, and peripheral interfaces, and communicates with other ECUs in the vehicle via in-vehicle Ethernet and / or CAN bus;
[0023] The Driver Monitoring System (DMS) is deployed above the steering wheel to capture the driver's facial expression; the Occupant Monitoring System (OMS) is deployed above the center console to detect occupant expressions and postures; and the cabin microphone array is installed at the four corners of the roof to collect voice and ambient sound.
[0024] A forward-facing camera and a millimeter-wave radar are installed at the front of the vehicle to collect information about the external environment.
[0025] The sensors in the vehicle-side intelligent perception layer also include LiDAR, which is installed on the roof or bumper to enhance environmental modeling capabilities.
[0026] The input data from the multimodal sensors to the decision-making system is integrated and processed by the cockpit domain controller to provide data support to the decision-making system.
[0027] Multimodal sensors are connected to the cockpit domain controller via MIPI-CSI / FPD-Link, I²S / TDM, and Ethernet interfaces; raw / feature / context data are aligned with a unified timestamp and coordinate system before being input into the hierarchical perception and event triggering process of this invention.
[0028] The collaborative communication layer adopts an "asynchronous object storage upload + metadata API notification" working mode to separate data flow from control flow. All data packets uploaded by the collaborative communication layer include a data type tag (data_type) to distinguish between raw media data (raw_media), feature vector data (feature_vector), and context information (context_meta). The addition of the data type tag is completed in the data encapsulation stage on the vehicle side. Specifically, before uploading to object storage, the system prepares structured metadata notifications (e.g., JSON or Protobuf format). During this process, a field named "data_type" is explicitly set in the metadata according to the data source and processing level.
[0029] After receiving the metadata API notification, the cloud selects different processing paths based on the data type marker; for example, it prioritizes parsing feature vectors for quick determination, or delays calling media data for deep inference, thereby improving the efficiency and accuracy of data processing.
[0030] The asynchronous object storage upload process includes hardware calls and asynchronous execution. Specifically, the asynchronous object storage upload is executed by a dedicated communication module (such as a 5G / V2X module) within the vehicle's cockpit domain controller. When the vehicle's intelligent perception layer generates a large amount of media data (such as a video recording from a front-facing camera), the cockpit domain controller's main processor instruction communication module initializes this upload task, so that the task is managed by an independent background upload service process. This process is responsible for data fragmentation, packaging, and establishing a connection with cloud object storage (such as AWS S3) via the TCP / IP protocol stack and transmitting the data. By placing the upload task in the background, it is ensured that the process will not block or preempt computing resources used for high-priority tasks such as in-cabin human-machine interaction and instrument display, thereby ensuring the system's real-time response capability.
[0031] The intelligent cockpit vehicle-cloud collaborative intelligent decision-making method uses the intelligent cockpit vehicle-cloud collaborative intelligent decision-making system based on multimodal perception described above. In this decision-making method, the processing results in the cloud are sent to the vehicle's cockpit domain controller through the collaborative communication layer. The cockpit domain controller then distributes the results to the relevant ECUs and cockpit interaction modules of the entire vehicle, thereby achieving unified output of safety prompts, vehicle control, and personalized services throughout the entire vehicle.
[0032] In the decision-making method, the vehicle-side intelligent perception layer includes a layered processing flow of basic detection, reporting decision, fine recognition, and semantic analysis;
[0033] The intelligent event triggering module includes a rule engine, a context analysis module, a security assessment module, and a data encapsulation module;
[0034] The context analysis module confirms the triggering conditions by combining vehicle operating status, geographical location, time, driver status, passenger status, and environmental conditions;
[0035] The security assessment module outputs high, medium, and low risk levels and sorts the events using a priority scoring model.
[0036] The working process of the collaborative communication layer includes uploading credential generation, direct transmission of object storage, and metadata API notification, wherein the metadata API carries object_key, verification hash, context information, and idempotent key;
[0037] The collaborative communication layer pushes task statuses such as task_created, response_chunk, and task_completed through SSE (Server-Sent Events) real-time event streams, supporting reconnection after disconnection and incremental resending;
[0038] The cloud-based intelligent decision-making layer performs data integrity verification and access control during the preprocessing stage, and calculates task scores based on a complexity evaluation model.
[0039] The complexity assessment model includes tool call factor, multimodal processing factor, personalized requirement factor and real-time correction factor, and selects fast link, hybrid link or complex link according to task score;
[0040] The output of the cloud-based intelligent decision-making layer includes perception-based results, semantic-based results, decision-based results, and task-based results.
[0041] In the aforementioned decision-making method, at the vehicle-side intelligent perception layer, the system adopts a modular, pluggable computer vision (CV) model architecture to achieve flexible expansion and dynamic management of different visual perception functions.
[0042] like Figure 2 As shown in Module 1 of the pluggable CV model management architecture diagram, in the pluggable computer vision model architecture, each model is provided as an independent plug-in package. The plug-in package typically contains model files (such as ONNX, TensorFlowLite, PyTorch Mobile formats), configuration files (config.json), interface definition files, and preprocessing and postprocessing logic scripts.
[0043] like Figure 2 As shown in Module 2 of the pluggable CV model management architecture diagram, the decision method introduces a complete model management framework into the decision system. The plug-in package is registered through the plug-in registry, enabling the decision system to identify its metadata at runtime, including version number, function category and dependency relationship.
[0044] The model management framework also includes a lifecycle management module for loading, initializing, running, and unloading plugins, enabling the decision system to monitor their status during runtime; a resource scheduler for allocating computing resources between GPUs / CPUs, controlling memory usage, and ensuring optimal performance when multiple plugins are running in parallel; and a performance monitoring module for real-time statistics of inference time and resource utilization, and for generating error logs and performance reports for diagnosis and optimization.
[0045] like Figure 2 As shown in module 3 of the pluggable CV model management architecture diagram, the decision-making method adopts the following methods in terms of security and version control: the decision-making system performs digital signature verification and permission checks on the plugin package to ensure its source is trustworthy and the code is secure; the decision-making system runs the plugin in a sandbox environment to avoid malicious plugin behavior from affecting the main system; the decision-making system supports compatibility checks and rollback operations for the plugin through the version management module, so that when an anomaly occurs in the new version, it can be restored to the stable version; the update mechanism of the decision-making system is based on OTA differential update to push updates in batches without interrupting services, and automatically performs verification after the update is completed.
[0046] like Figure 2 As shown in module 4 of the pluggable CV model management architecture diagram, when the decision method is running in the decision system, the environment layer model of the decision system is executed through a unified runtime (such as ONNX Runtime, TensorFlow Lite or PyTorchMobile), and interacts with other components of the vehicle system through the interface adaptation layer to achieve unified conversion of input and output data formats. An event bus is used for inter-plugin communication and message distribution, and a task scheduler is used to prioritize and manage different tasks to ensure that critical tasks are executed first in high-load scenarios.
[0047] like Figure 2 As shown in module 5 of the pluggable CV model management architecture diagram, the decision-making method supports integration with external management platforms through a decision-making system. The cloud-based management platform is used to uniformly release, update, and analyze the performance of plugins. Developer tools provide local debugging, log viewing, and performance testing capabilities. The integration interface of the decision-making system ensures that the plugins can work collaboratively with other services or hardware abstraction layers of the vehicle system, achieving integrated closed-loop management from development, deployment, to operation.
[0048] like Figure 3 The vehicle-side layered perception flowchart shows that the decision-making method includes the vehicle-side layered perception process, which includes the following steps:
[0049] Step S1: The vehicle first timestamps and aligns the multimodal data from the camera, millimeter-wave radar, lidar and DMS / OMS with the coordinate system, and then processes them step by step according to the L1-L4 hierarchical perception process.
[0050] Step S2: The L1 basic detection layer outputs the basic results of the bounding boxes, categories, and confidence scores of objects / lane lines / traffic signs;
[0051] Step S3: The L2 reporting decision layer, based on L1, combines vehicle operating status (vehicle speed, steering angle, etc.) to execute rule judgment, decide whether to generate a reporting event, and set priority and initial route suggestion;
[0052] In step S4, the L3 fine recognition layer performs higher-precision recognition of key targets (such as gestures, facial expressions, and fine-grained object classification) to enhance interaction and scene understanding.
[0053] In step S5, the L4 semantic analysis layer integrates the L1–L3 results with the context of vehicle status, geographical location, and historical trajectory to construct a scene graph and perform intent reasoning to generate semantically complex events.
[0054] Through the above layering, the system provides high-level semantics while ensuring low latency, providing high-quality input for subsequent event triggering and cloud-based collaborative decision-making.
[0055] like Figure 4 As shown in the event triggering mechanism flowchart, in the event triggering mechanism of the vehicle-side intelligent perception layer, the system performs routine detection and recognition of sensor data, and also realizes real-time response to key driving scenarios through the multimodal event triggering mechanism. This mechanism is completed collaboratively by the rule engine, context analysis module, safety assessment module and data encapsulation module to ensure that the event has high priority and reliability throughout the entire process of triggering, judging and reporting.
[0056] Step S6: Using a rule engine module based on a predefined rule base and a dynamic update strategy, the detection results of CV model, radar, and DMS / OMS multi-source inputs are quickly compared; for example, when the system detects a traffic sign ahead, if the rule base specifies "trigger parsing when the vehicle approaches a critical road section", a trigger signal will be generated immediately.
[0057] The rule base not only stores commonly used rules locally, but also supports the dynamic distribution of new rules from the cloud via OTA to adapt to complex and ever-changing road environments.
[0058] Step S7: Through the context analysis module, the detection results are combined with the vehicle's operating status and environmental factors to reconfirm the triggering conditions. Specifically, this includes: geographical location analysis (determining whether the vehicle has entered a high-risk area or special road), time factor consideration (nighttime, rain, snow, and other special periods), and user status assessment (such as driver distraction or fatigue). Context enhancement helps the decision-making system avoid false alarms from a single sensor and improves the accuracy of event triggering.
[0059] Step S8: After the incident is initially confirmed, the security assessment module further determines its risk level and urgency, and selects the corresponding security strategy according to different levels.
[0060] Step S9: The data encapsulation module structures and packages the information related to the triggering event. This process includes: keyframe extraction (selecting the most representative frame based on the importance score of the image frame), context data encapsulation (such as GPS, speed, timestamp, navigation status), integration of preliminary analysis results (including confidence, detection box, and classification results output by the CV model), and addition of triggering cause markers and priority information.
[0061] Step S10: The packaged event packet is uploaded to the cloud through the collaborative communication layer. High-priority events are reported immediately through a low-latency link, while medium- and low-priority events are transmitted with delay through a regular link. The events are then parsed and differentiated in the cloud.
[0062] Through the layered filtering and enhancement of the above-mentioned triggering mechanism, a complete closed loop of "detection-judgment-evaluation-encapsulation-reporting" is achieved at the vehicle end, providing a fast and reliable response in safety-critical scenarios, and ensuring that the event data uploaded to the cloud has structured, context-complete, and traceable characteristics, thereby improving the accuracy and efficiency of subsequent cloud-based decision analysis.
[0063] In step S8, the priority assessment method adopts a weighted scoring model, which is expressed by the following formula:
[0064] P=α⋅R+β⋅T+γ⋅C+δ⋅E−ε⋅L;
[0065] Where R represents risk level, T represents time sensitivity, C represents contextual factors, E represents event category weight, L represents system load status, and α, β, γ, δ, and ε are the weight coefficients of each dimension; the risk level R is divided into three levels: high, medium, and low and mapped to normalized interval values.
[0066] High-risk events correspond to R=1.0, medium-risk events correspond to R=0.6, low-risk events correspond to R=0.3, and non-safe events correspond to R=0.1;
[0067] The time sensitivity T is assigned a value based on the task's real-time requirements: strong real-time events correspond to T=1.0, medium real-time events correspond to T=0.5, and weak real-time events correspond to T=0.2.
[0068] The context C is determined based on external conditions. Events triggered at night, in rainy or snowy weather, or at complex intersections correspond to C=0.8, events triggered in normal environments correspond to C=0.5, and events triggered in low-risk environments correspond to C=0.2.
[0069] The category weight E is preset according to the event category, with core security events corresponding to E=1.0, minor security events corresponding to E=0.6, and non-security auxiliary events corresponding to E=0.3;
[0070] The system load state L is determined based on the vehicle-side computing and communication resource status. High load corresponds to L=0.7, medium load corresponds to L=0.4, and low load corresponds to L=0.1.
[0071] The weighting coefficients are set to: α=0.3, β=0.2, γ=0.2, δ=0.2, ε=0.1;
[0072] After calculating the priority score P, the system classifies it into three levels: when P ≥ 0.8, the event is judged as high priority and reported immediately via a low-latency link; when 0.4 ≤ P < 0.8, the event is judged as medium priority and reported via a regular link; when P < 0.4, the event is judged as low priority and can be reported after delay or batch aggregation. For example, when a pedestrian is detected crossing the road ahead, the safety assessment module outputs a high risk level, R = 1.0; since this event requires an immediate response, the time sensitivity T = 1.0; triggered at night or in rainy weather, the context environment is assigned a value of C = 0.8; this event category belongs to the core security category, E = 1.0; under normal system resource conditions, L = 0.2. With a calculated priority score P = 0.92, it is judged as a high-priority event and will be immediately uploaded via a low-latency link.
[0073] like Figure 6 As shown in the flowchart of the cloud-based intelligent decision-making layer, in step S10, the cloud sets up dual input interfaces to handle two types of tasks: one is multimodal direct input from the vehicle, used to handle cockpit perception events that are safety-related or have high real-time requirements; the other is central routing input that is transferred after central orchestration, used to handle complex cockpit intelligent tasks that require cross-service orchestration, personalization, and context integration, including the following process:
[0074] In process S11, both types of interfaces uniformly perform data integrity verification, format standardization, and security checks at the entry point, including event packet structure verification, media reference validity verification, task idempotent key identification, and authentication based on service access control lists (ACLs). For tasks containing media references, the preprocessing layer only parses metadata and reference identifiers, and large files are retrieved as needed in subsequent links according to the references, avoiding entry point blockage.
[0075] Process S12, Task Complexity Scoring and Intelligent Routing: To ensure the intelligent cockpit's response efficiency and deep reasoning capabilities under various tasks, a task complexity scoring module is set up in the cloud. This module generates a complexity score for each task through quantitative evaluation of multiple dimensions and selects the appropriate processing link accordingly.
[0076] The rating factors include:
[0077] Tool Invocation Complexity: The number of external services or tools required for the task to be invoked, such as maps, weather, POI, or user profiling services. Each additional external invocation increases the complexity score proportionally, reflecting potential latency and resource consumption in the process.
[0078] Multimodal data processing requirements: If the task involves multimodal fusion processing such as vision, audio, vehicle status, and historical trajectory, a certain weight will be added to the baseline score; single-modal tasks will score lower.
[0079] Intensity of personalized needs: If the task requires personalized reasoning based on user profiles and historical behavioral habits, the score will be increased based on the degree of dependence on personalized parameters; for tasks that only rely on common rules, the score will be lower.
[0080] Real-time requirements: If the task belongs to a strong real-time scenario of security incident or dialogue interaction, it should be marked as "time sensitive" in the complexity score. Even if the task itself calls many services, the weight should be adjusted to prioritize low-latency links.
[0081] The scoring mechanism uses a weighted cumulative method to calculate the task complexity score; the number of tool calls is weighted at 0.3, multimodal processing at 0.4, personalized needs at 0.2, and real-time marking is adjusted by a factor of 0.1 in the score.
[0082] The tool call complexity factor adopts a binary design. When the task involves at least one external tool call, the factor takes a value of 1.0; when the task does not involve any external tool calls, the factor takes a value of 0. In this way, the complexity score is avoided from exceeding the normalized range of 0 to 1 due to the cumulative number of tool calls, thereby ensuring the uniformity and controllability of the scoring model.
[0083] Routing strategy: When the score is below the first threshold θ1 (0.3), the task enters the fast link, where the lightweight engine performs basic analysis and response; when the score is between θ1 and θ2 (0.3–0.7), the task enters the hybrid link, which first returns preliminary results quickly and then calls the deep analysis system to supplement the results; when the score is above θ2 (0.7), the task enters the complex link, where deep reasoning and planning are performed by multimodal large models and cross-service orchestration.
[0084] In some embodiments, the complexity score threshold is not fixed, but dynamically adjusted based on system resource status (such as GPU utilization and network congestion) and historical task performance. For example, θ1 can be appropriately increased under high load to reduce the proportion of tasks entering complex links and ensure the overall stability of the cockpit response.
[0085] Process S13, Multimodal Deep Analysis Layer: In the fast link, the system prioritizes returning the minimum result set required for cockpit safety / interaction (such as structured prompts, basic explanations, and executable commands), ensuring that the interaction latency is in the sub-second to hundreds of millisecond range; when necessary, it immediately pulls keyframes or small-volume features for lightweight inference based on references; in the hybrid link, the system first sends back fast results to ensure response continuity, and then completes supplementary inference on the slow system or multimodal model (such as higher confidence recognition, regulatory / POI matching, and profile fusion), and writes back to the cockpit with result correction events; in the complex link, the decision system performs deep semantic understanding and planning based on the multimodal large model and cross-service orchestration, supports long temporal context and personalized synthetic output, and attaches executable templates and fallback strategies to sensitive commands (such as cockpit settings, media and navigation switching) during the generation stage;
[0086] In process S14, result generation and degradation / correction involves integrating the rapid results, supplementary inference, and complex inference results and returning them to the vehicle in real time. When there is link delay or insufficient resources, the process reverts to the rapid link to return the "minimum usable result," which is then supplemented and corrected later. In some embodiments, the result generation of the cloud-based intelligent decision-making layer includes the following types:
[0087] Perception-related results: These confirm or correct perception events reported by the vehicle and output the results of target detection and recognition. For example, the cloud verifies the category and confidence level of obstacles ahead, or performs secondary confirmation on traffic sign recognition results, thereby improving detection accuracy.
[0088] Semantic results: Combining contextual information and multimodal input, outputting results that reflect user intent and scene understanding; for example, generating an intent judgment of "navigate to the nearest gas station" based on driver voice commands and gesture recognition, or increasing the risk level in scenarios such as "night + rainy day + winding road section", achieving semantic-level comprehensive reasoning.
[0089] Decision-making results: Generate driving strategies and safety decision outputs; for example, the cloud provides driving change suggestions such as "decelerate", "change lanes" or "maintain distance" based on the fusion results, and selects the corresponding safety strategy (immediate reporting, delayed reporting or local prompt only).
[0090] Task-based results: Generate complex task outputs for cross-service requests or personalized needs; for example, when a user requests "plan a trip from Shanghai to Hangzhou", the cloud integrates map, weather and POI services to generate a complete trip planning result; or in recommendation tasks, combine user profiles and historical preferences to output personalized recommendations after sorting.
[0091] In process S15, the vehicle first applies for an upload permit from the cloud authentication service;
[0092] In process S16, a pre-signed URL and object_key are generated and returned in the cloud;
[0093] In process S17, the vehicle transmits large media files such as images / videos directly to the object storage in a segmented / resumable manner. The storage side returns the ETag / segment list after completion.
[0094] In process S18, the vehicle then reports a business request through the metadata API, which includes an object reference object_key, a verification hash (such as SHA-256) and context metadata, as well as an idempotency key idempotency_key with a data type tag data_type;
[0095] In process S19, the business hub receives and schedules tasks. The running service of the business hub selects different processing paths based on the data type markers as needed: when it is feature data, it prioritizes rapid parsing for immediate judgment of security-related events; when it is raw media data, it calls object storage on demand for deep inference for fine-grained analysis of complex tasks or personalized recommendations; when it is context information, it is used for task enhancement and result supplementation. The cloud records the content hash and metadata digest for each object_key, performing rapid consistency verification before business access. When duplicate reporting (network retries or cockpit playback) occurs, idempotent keys ensure that only one business side effect occurs, and subsequent requests return the status and result reference of the first task. This mode significantly reduces the bandwidth and computational pressure on the API surface, facilitating the introduction of CDN / edge acceleration and multi-region disaster recovery.
[0096] In process S20, when using real-time communication (SSE event stream), to ensure observability and fine-grained progress feedback on the cockpit side, the collaborative communication layer provides SSE (Server-Sent Events) real-time event streams. The vehicle subscribes to the event channel with a task identifier, and the cloud pushes task_created, plan_generated, tool_call_started, response_chunk, and task_completed events in stages. For network jitter, the system uses Last-Event-ID and heartbeat to maintain connection continuity and uses exponential backoff for reconnection and speed limiting. When the vehicle goes offline and comes back online, the event channel incrementally resends events based on Last-Event-ID to ensure the cockpit has complete awareness of the task status. If network degradation or excessive system load occurs during task execution, the business hub can also trigger degradation and correction mechanisms, so that the task returns the "minimum available result" in the SSE push first, and then sends the corrected result after the link is restored, thereby maintaining the consistency of the user experience and results on the vehicle side.
[0097] In step S10, the cloud operation includes degradation and correction, specifically: when network degradation or slow system load increases, causing complex / hybrid links to fail to complete within the specified time limit, the decision system's router automatically triggers a degradation strategy: prioritizing fallback to a fast link to return the "minimum available result," while using cached results to maintain user experience, and recording the task idempotency key and context so that execution can continue after the link recovers and the corrected result is sent out; for supplementary results that may cause user misunderstanding, version numbers and timestamps are used to mark them, and consistency correction is performed on the cockpit UI side (displaying "updated" or "corrected" prompts), thus forming a closed loop of real-time response—supplementary reasoning—consistency correction;
[0098] In step S10, after receiving the metadata API notification, the API gateway of the cloud-based business hub immediately parses its payload (such as a JSON message), reads the value of the "data_type" field, and based on this value, the cloud-based intelligent routing and scheduling module will execute the preset differentiated processing logic:
[0099] a) When data_type is "feature_vector", the routing module determines that this is a high-time-sensitivity task and immediately forwards the request to the low-latency real-time analysis service. This service can perform fast calculations (such as fatigue driving judgment) directly based on the feature vectors carried in the metadata, without waiting for the download of large media files, thus achieving a response time in seconds.
[0100] b) When data_type is "raw_media", the routing module pushes the task to a message queue (such as Kafka or RabbitMQ). The background deep learning inference cluster (usually equipped with GPU / NPU) acts as a consumer to asynchronously retrieve the task from the queue, pulls the corresponding media file from the object storage according to the object_key in the metadata, and performs complex and time-consuming analysis (such as complete semantic understanding of traffic scenarios).
[0101] c) When data_type is "context_meta", the request may be distributed to the data management service to update event logs or enrich user profiles without triggering large-scale computation.
[0102] This invention relates to the field of intelligent cockpit technology, and discloses an intelligent cockpit vehicle-cloud collaborative intelligent decision-making system based on multimodal perception and its implementation method. This system enables rapid response to safety events (such as driver fatigue detection alerts and external abnormal obstacle recognition) and personalized service output (such as voice-recommended restaurants and content recommendations) in intelligent cockpit scenarios, thereby improving the human-machine interaction experience and the continuity of in-cabin services. The system includes a vehicle-side intelligent perception layer, a collaborative communication layer, and a cloud-based intelligent decision-making layer. The vehicle-side intelligent perception layer, through a pluggable visual model architecture, combined with driver monitoring (DMS), occupant monitoring (OMS), voice and environmental sensors, achieves multimodal perception of driver status, occupant behavior, and the external environment, and employs a layered processing mechanism to complete step-by-step reasoning from basic detection to semantic analysis. The collaborative communication layer adopts an architecture that separates object storage direct transmission from metadata API, and combines it with SSE (Server-Sent Events) real-time event streams to achieve efficient uploading of large volumes of data and low-latency transmission of small volumes of control information, ensuring observable task progress and the continuity of cockpit interaction. The cloud-based intelligent decision-making layer selects fast, hybrid, or complex links based on a complexity assessment model. It combines user profiles with external services (maps, weather, POIs, etc.) to generate semantic decisions and personalized recommendations, and maintains reliable response in weak network environments through degradation and correction mechanisms. Compared to existing technologies, this invention achieves efficient processing of multimodal data, personalized service recommendations, and low-latency security event response in cockpit scenarios, improving user experience and system stability.
[0103] The purpose of this invention is to provide a vehicle-cloud collaborative intelligent decision-making system and its implementation method based on multimodal perception for use in intelligent cockpit environments, in order to solve the problems of insufficient vehicle-side computing power, large cloud processing latency, high communication coupling, and lack of personalized services in existing systems. The results of cloud processing are sent to the vehicle's cockpit domain controller via a collaborative communication layer, and then further distributed by the cockpit domain controller to relevant ECUs and cockpit interaction modules throughout the vehicle. This enables unified output of safety prompts, vehicle control, and personalized services across the entire vehicle, achieving efficient collaboration, rapid response, and personalized decision-making between the vehicle and the cloud. Attached Figure Description
[0104] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0105] Appendix Figure 1 This is a schematic diagram of the intelligent cockpit vehicle-cloud collaborative system architecture provided in an embodiment of the present invention (showing the overall architectural relationship between the vehicle-side intelligent perception layer, collaborative communication layer, cloud-based intelligent decision-making layer, and external system integration module).
[0106] Appendix Figure 2 A schematic diagram of the pluggable CV model management architecture provided in the embodiments of the present invention (showing the relationship and operation mechanism between the model plugin package, development toolchain, plugin library and model management framework).
[0107] Appendix Figure 3 The vehicle-side layered perception flowchart provided in this embodiment of the invention (shows the layered processing flow of basic detection, reporting decision, fine recognition and semantic analysis after the multimodal sensor input undergoes spatiotemporal alignment and fusion).
[0108] Appendix Figure 4 This is a schematic diagram of the event triggering mechanism provided in an embodiment of the present invention (showing the collaborative work of the rule engine, context analysis, security assessment and data encapsulation modules to realize event identification, priority determination and reporting).
[0109] Appendix Figure 5 This is a schematic diagram of the collaborative communication layer process provided in an embodiment of the present invention (showing the process of obtaining uploaded credentials on the vehicle side, direct transmission of object storage, metadata API notification, business scheduling, and SSE real-time event stream push).
[0110] Appendix Figure 6 This is a schematic diagram of the cloud-based intelligent decision-making layer process provided in an embodiment of the present invention (showing the process by which task input, after preprocessing and complexity assessment, enters the fast link, hybrid link, or complex link for processing, and generates corresponding results to be returned to the vehicle).
[0111] Appendix Figure 7This is a schematic diagram of the external system integration framework provided for an embodiment of the present invention (showing the integration relationship between external resources such as map services, weather services, POI services, user data centers, log and monitoring systems, and object storage and the cloud decision layer). Detailed Implementation
[0112] As shown in the figure, the intelligent cockpit vehicle-cloud collaborative intelligent decision-making system based on multimodal perception includes:
[0113] The vehicle-side intelligent perception layer is used to execute computer vision models through pluggable devices on the vehicle side, collect multimodal data from driver monitoring, occupant monitoring, cameras, and radar, and generate event data through layered processing.
[0114] The intelligent event triggering module is used to determine events based on detection results and context information, and to encapsulate triggering events in a structured manner and set priorities.
[0115] The collaborative communication layer is used to upload data using a method that separates object storage direct transmission from metadata API, and to provide feedback on task status and results through the SSE real-time communication protocol.
[0116] The cloud-based intelligent decision-making layer receives the event data, selects a fast link, a hybrid link, or a complex link for processing based on the task complexity, and generates safety prompts, in-cabin interactive responses, or personalized recommendation results by combining user profiles and external services.
[0117] The vehicle-side intelligent perception layer is based on a pluggable computer vision model architecture and combines multimodal sensors located inside and outside the vehicle cabin to perceive the environment inside and outside the cabin.
[0118] The vehicle-side intelligent perception layer is used to detect driver fatigue, distraction, emotions and gestures, as well as to identify occupant status and behavior, thereby providing semantic input for cabin services and human-machine interaction. It aligns multimodal data through a unified timestamp and coordinate system, and adopts a hierarchical processing mechanism: L1 basic detection, L2 reporting and decision-making, L3 fine recognition and L4 semantic analysis, to achieve a step-by-step process from low-latency detection to high-level semantic understanding.
[0119] Multimodal sensors inside and outside the vehicle include Driver Monitoring System (DMS), Occupant Monitoring System (OMS), in-cabin microphone array, forward-facing camera, and millimeter-wave radar;
[0120] The vehicle-side intelligent perception layer is based on an intelligent event triggering mechanism, which identifies key event types inside and outside the cabin through rule engine and context information analysis.
[0121] Contextual information includes vehicle operating status, geographical location, time, driver status, user profile, and weather conditions;
[0122] Key event types include safety events (such as driver fatigue, obstacles ahead) and interactive events (such as occupants issuing voice commands, the system detecting entertainment needs);
[0123] The intelligent event triggering mechanism classifies different events into risk levels through a security assessment module. The priority scoring model used for classification sorts events based on risk, time sensitivity, context, and personalized needs, ensuring that security events and high-value interactions are processed first. The data encapsulation module packages events into standardized data packets for further processing in the cloud.
[0124] The collaborative communication layer adopts a "data flow and control flow separation" architecture. Specifically, large-volume media data (images, videos, etc.) are directly transmitted through object storage, supporting segmentation and breakpoint resumption; control messages (object_key, hash, context, idempotent key) are uploaded through metadata API to avoid redundant calculations.
[0125] The collaborative communication layer provides a real-time event stream based on SSE (Server-Sent Events) for continuous feedback on task progress (such as task_created, response_chunk, task_completed), and supports reconnection after disconnection and incremental resending in weak network environments. This design ensures both real-time response to security events and continuity of information and service interaction within the cockpit.
[0126] The cloud-based intelligent decision-making layer includes dual input interfaces: a multimodal direct input interface to process real-time safety events reported by the vehicle to ensure low-latency response; and a central routing input interface to handle complex cross-domain, multi-step tasks to support the orchestration and fusion of complex tasks.
[0127] The cloud-based intelligent decision-making layer uses a cloud-based routing mechanism to score tasks based on a complexity assessment model and selects between fast, hybrid, or complex paths. In some embodiments, the complexity score combines the number of tool calls, multimodal dependencies, and real-time requirements, and can dynamically update thresholds based on historical execution results. The cloud-based multimodal analysis engine calls large-scale models, combining time, geographic location, and user profiles to perform semantic reasoning and personalized recommendations, generating optimized decision results.
[0128] The decision-making system includes a closed-loop execution and correction mechanism. Specifically, in some embodiments, when network degradation or high system load causes complex link delays, the system automatically falls back to the fastest link to return the minimum usable result (such as a basic reminder or brief answer), and reissues the corrected result after the link is restored. The corrected result includes a version number and timestamp, and the cockpit interface displays "Updated / Corrected," ensuring consistency in user perception. Through this closed-loop mechanism of real-time response—supplementary reasoning—consistency correction, both timely handling of security incidents and smooth cockpit interaction are ensured.
[0129] The cockpit domain controller includes at least one processor, memory, and peripheral interfaces, and communicates with other ECUs in the vehicle via in-vehicle Ethernet and / or CAN bus;
[0130] The Driver Monitoring System (DMS) is deployed above the steering wheel to capture the driver's facial expression; the Occupant Monitoring System (OMS) is deployed above the center console to detect occupant expressions and postures; and the cabin microphone array is installed at the four corners of the roof to collect voice and ambient sound.
[0131] A forward-facing camera and a millimeter-wave radar are installed at the front of the vehicle to collect information about the external environment.
[0132] The sensors in the vehicle-side intelligent perception layer also include LiDAR, which is installed on the roof or bumper to enhance environmental modeling capabilities.
[0133] The input data from the multimodal sensors to the decision-making system is integrated and processed by the cockpit domain controller to provide data support to the decision-making system.
[0134] Multimodal sensors are connected to the cockpit domain controller via MIPI-CSI / FPD-Link, I²S / TDM, and Ethernet interfaces; raw / feature / context data are aligned with a unified timestamp and coordinate system before being input into the hierarchical perception and event triggering process of this invention.
[0135] The collaborative communication layer adopts an "asynchronous object storage upload + metadata API notification" working mode to separate data flow from control flow. All data packets uploaded by the collaborative communication layer include a data type tag (data_type) to distinguish between raw media data (raw_media), feature vector data (feature_vector), and context information (context_meta). The addition of the data type tag is completed in the data encapsulation stage on the vehicle side. Specifically, before uploading to object storage, the system prepares structured metadata notifications (e.g., JSON or Protobuf format). During this process, a field named "data_type" is explicitly set in the metadata according to the data source and processing level.
[0136] After receiving the metadata API notification, the cloud selects different processing paths based on the data type marker; for example, it prioritizes parsing feature vectors for quick determination, or delays calling media data for deep inference, thereby improving the efficiency and accuracy of data processing.
[0137] The asynchronous object storage upload process includes hardware calls and asynchronous execution. Specifically, the asynchronous object storage upload is executed by a dedicated communication module (such as a 5G / V2X module) within the vehicle's cockpit domain controller. When the vehicle's intelligent perception layer generates a large amount of media data (such as a video recording from a front-facing camera), the cockpit domain controller's main processor instruction communication module initializes this upload task, so that the task is managed by an independent background upload service process. This process is responsible for data fragmentation, packaging, and establishing a connection with cloud object storage (such as AWS S3) via the TCP / IP protocol stack and transmitting the data. By placing the upload task in the background, it is ensured that the process will not block or preempt computing resources used for high-priority tasks such as in-cabin human-machine interaction and instrument display, thereby ensuring the system's real-time response capability.
[0138] The intelligent cockpit vehicle-cloud collaborative intelligent decision-making method uses the intelligent cockpit vehicle-cloud collaborative intelligent decision-making system based on multimodal perception described above. In this decision-making method, the processing results in the cloud are sent to the vehicle's cockpit domain controller through the collaborative communication layer. The cockpit domain controller then distributes the results to the relevant ECUs and cockpit interaction modules of the entire vehicle, thereby achieving unified output of safety prompts, vehicle control, and personalized services throughout the entire vehicle.
[0139] In the decision-making method, the vehicle-side intelligent perception layer includes a layered processing flow of basic detection, reporting decision, fine recognition, and semantic analysis;
[0140] The intelligent event triggering module includes a rule engine, a context analysis module, a security assessment module, and a data encapsulation module;
[0141] The context analysis module confirms the triggering conditions by combining vehicle operating status, geographical location, time, driver status, passenger status, and environmental conditions;
[0142] The security assessment module outputs high, medium, and low risk levels and sorts the events using a priority scoring model.
[0143] The working process of the collaborative communication layer includes uploading credential generation, direct transmission of object storage, and metadata API notification, wherein the metadata API carries object_key, verification hash, context information, and idempotent key;
[0144] The collaborative communication layer pushes task statuses such as task_created, response_chunk, and task_completed through SSE (Server-Sent Events) real-time event streams, supporting reconnection after disconnection and incremental resending;
[0145] The cloud-based intelligent decision-making layer performs data integrity verification and access control during the preprocessing stage, and calculates task scores based on a complexity evaluation model.
[0146] The complexity assessment model includes tool call factor, multimodal processing factor, personalized requirement factor and real-time correction factor, and selects fast link, hybrid link or complex link according to task score;
[0147] The output of the cloud-based intelligent decision-making layer includes perception-based results, semantic-based results, decision-based results, and task-based results.
[0148] In the aforementioned decision-making method, at the vehicle-side intelligent perception layer, the system adopts a modular, pluggable computer vision (CV) model architecture to achieve flexible expansion and dynamic management of different visual perception functions.
[0149] like Figure 2 As shown in Module 1 of the pluggable CV model management architecture diagram, in the pluggable computer vision model architecture, each model is provided as an independent plug-in package. The plug-in package typically contains model files (such as ONNX, TensorFlowLite, PyTorch Mobile formats), configuration files (config.json), interface definition files, and preprocessing and postprocessing logic scripts.
[0150] like Figure 2 As shown in Module 2 of the pluggable CV model management architecture diagram, the decision method introduces a complete model management framework into the decision system. The plug-in package is registered through the plug-in registry, enabling the decision system to identify its metadata at runtime, including version number, function category and dependency relationship.
[0151] The model management framework also includes a lifecycle management module for loading, initializing, running, and unloading plugins, enabling the decision system to monitor their status during runtime; a resource scheduler for allocating computing resources between GPUs / CPUs, controlling memory usage, and ensuring optimal performance when multiple plugins are running in parallel; and a performance monitoring module for real-time statistics of inference time and resource utilization, and for generating error logs and performance reports for diagnosis and optimization.
[0152] like Figure 2As shown in module 3 of the pluggable CV model management architecture diagram, the decision-making method adopts the following methods in terms of security and version control: the decision-making system performs digital signature verification and permission checks on the plugin package to ensure its source is trustworthy and the code is secure; the decision-making system runs the plugin in a sandbox environment to avoid malicious plugin behavior from affecting the main system; the decision-making system supports compatibility checks and rollback operations for the plugin through the version management module, so that when an anomaly occurs in the new version, it can be restored to the stable version; the update mechanism of the decision-making system is based on OTA differential update to push updates in batches without interrupting services, and automatically performs verification after the update is completed.
[0153] like Figure 2 As shown in module 4 of the pluggable CV model management architecture diagram, when the decision method is running in the decision system, the environment layer model of the decision system is executed through a unified runtime (such as ONNX Runtime, TensorFlow Lite or PyTorchMobile), and interacts with other components of the vehicle system through the interface adaptation layer to achieve unified conversion of input and output data formats. An event bus is used for inter-plugin communication and message distribution, and a task scheduler is used to prioritize and manage different tasks to ensure that critical tasks are executed first in high-load scenarios.
[0154] like Figure 2 As shown in module 5 of the pluggable CV model management architecture diagram, the decision-making method supports integration with external management platforms through a decision-making system. The cloud-based management platform is used to uniformly release, update, and analyze the performance of plugins. Developer tools provide local debugging, log viewing, and performance testing capabilities. The integration interface of the decision-making system ensures that the plugins can work collaboratively with other services or hardware abstraction layers of the vehicle system, achieving integrated closed-loop management from development, deployment, to operation.
[0155] like Figure 3 The vehicle-side layered perception flowchart shows that the decision-making method includes the vehicle-side layered perception process, which includes the following steps:
[0156] Step S1: The vehicle first timestamps and aligns the multimodal data from the camera, millimeter-wave radar, lidar and DMS / OMS with the coordinate system, and then processes them step by step according to the L1-L4 hierarchical perception process.
[0157] Step S2: The L1 basic detection layer outputs the basic results of the bounding boxes, categories, and confidence scores of objects / lane lines / traffic signs;
[0158] Step S3: The L2 reporting decision layer, based on L1, combines vehicle operating status (vehicle speed, steering angle, etc.) to execute rule judgment, decide whether to generate a reporting event, and set priority and initial route suggestion;
[0159] In step S4, the L3 fine recognition layer performs higher-precision recognition of key targets (such as gestures, facial expressions, and fine-grained object classification) to enhance interaction and scene understanding.
[0160] In step S5, the L4 semantic analysis layer integrates the L1–L3 results with the context of vehicle status, geographical location, and historical trajectory to construct a scene graph and perform intent reasoning to generate semantically complex events.
[0161] Through the above layering, the system provides high-level semantics while ensuring low latency, providing high-quality input for subsequent event triggering and cloud-based collaborative decision-making.
[0162] like Figure 4 As shown in the event triggering mechanism flowchart, in the event triggering mechanism of the vehicle-side intelligent perception layer, the system performs routine detection and recognition of sensor data, and also realizes real-time response to key driving scenarios through the multimodal event triggering mechanism. This mechanism is completed collaboratively by the rule engine, context analysis module, safety assessment module and data encapsulation module to ensure that the event has high priority and reliability throughout the entire process of triggering, judging and reporting.
[0163] Step S6: Using a rule engine module based on a predefined rule base and a dynamic update strategy, the detection results of CV model, radar, and DMS / OMS multi-source inputs are quickly compared; for example, when the system detects a traffic sign ahead, if the rule base specifies "trigger parsing when the vehicle approaches a critical road section", a trigger signal will be generated immediately.
[0164] The rule base not only stores commonly used rules locally, but also supports the dynamic distribution of new rules from the cloud via OTA to adapt to complex and ever-changing road environments.
[0165] Step S7: Through the context analysis module, the detection results are combined with the vehicle's operating status and environmental factors to reconfirm the triggering conditions. Specifically, this includes: geographical location analysis (determining whether the vehicle has entered a high-risk area or special road), time factor consideration (nighttime, rain, snow, and other special periods), and user status assessment (such as driver distraction or fatigue). Context enhancement helps the decision-making system avoid false alarms from a single sensor and improves the accuracy of event triggering.
[0166] Step S8: After the incident is initially confirmed, the security assessment module further determines its risk level and urgency, and selects the corresponding security strategy according to different levels.
[0167] Step S9: The data encapsulation module structures and packages the information related to the triggering event. This process includes: keyframe extraction (selecting the most representative frame based on the importance score of the image frame), context data encapsulation (such as GPS, speed, timestamp, navigation status), integration of preliminary analysis results (including confidence, detection box, and classification results output by the CV model), and addition of triggering cause markers and priority information.
[0168] Step S10: The packaged event packet is uploaded to the cloud through the collaborative communication layer. High-priority events are reported immediately through a low-latency link, while medium- and low-priority events are transmitted with delay through a regular link. The events are then parsed and differentiated in the cloud.
[0169] Through the layered filtering and enhancement of the above-mentioned triggering mechanism, a complete closed loop of "detection-judgment-evaluation-encapsulation-reporting" is achieved at the vehicle end, providing a fast and reliable response in safety-critical scenarios, and ensuring that the event data uploaded to the cloud has structured, context-complete, and traceable characteristics, thereby improving the accuracy and efficiency of subsequent cloud-based decision analysis.
[0170] In step S8, the priority assessment method adopts a weighted scoring model, which is expressed by the following formula:
[0171] P=α⋅R+β⋅T+γ⋅C+δ⋅E−ε⋅L;
[0172] Where R represents risk level, T represents time sensitivity, C represents contextual factors, E represents event category weight, L represents system load status, and α, β, γ, δ, and ε are the weight coefficients of each dimension; the risk level R is divided into three levels: high, medium, and low and mapped to normalized interval values.
[0173] High-risk events correspond to R=1.0, medium-risk events correspond to R=0.6, low-risk events correspond to R=0.3, and non-safe events correspond to R=0.1;
[0174] The time sensitivity T is assigned a value based on the task's real-time requirements: strong real-time events correspond to T=1.0, medium real-time events correspond to T=0.5, and weak real-time events correspond to T=0.2.
[0175] The context C is determined based on external conditions. Events triggered at night, in rainy or snowy weather, or at complex intersections correspond to C=0.8, events triggered in normal environments correspond to C=0.5, and events triggered in low-risk environments correspond to C=0.2.
[0176] The category weight E is preset according to the event category, with core security events corresponding to E=1.0, minor security events corresponding to E=0.6, and non-security auxiliary events corresponding to E=0.3;
[0177] The system load state L is determined based on the vehicle-side computing and communication resource status. High load corresponds to L=0.7, medium load corresponds to L=0.4, and low load corresponds to L=0.1.
[0178] The weighting coefficients are set to: α=0.3, β=0.2, γ=0.2, δ=0.2, ε=0.1;
[0179] After calculating the priority score P, the system classifies it into three levels: when P ≥ 0.8, the event is judged as high priority and reported immediately via a low-latency link; when 0.4 ≤ P < 0.8, the event is judged as medium priority and reported via a regular link; when P < 0.4, the event is judged as low priority and can be reported after delay or batch aggregation. For example, when a pedestrian is detected crossing the road ahead, the safety assessment module outputs a high risk level, R = 1.0; since this event requires an immediate response, the time sensitivity T = 1.0; triggered at night or in rainy weather, the context environment is assigned a value of C = 0.8; this event category belongs to the core security category, E = 1.0; under normal system resource conditions, L = 0.2. With a calculated priority score P = 0.92, it is judged as a high-priority event and will be immediately uploaded via a low-latency link.
[0180] like Figure 6 As shown in the flowchart of the cloud-based intelligent decision-making layer, in step S10, the cloud sets up dual input interfaces to handle two types of tasks: one is multimodal direct input from the vehicle, used to handle cockpit perception events that are safety-related or have high real-time requirements; the other is central routing input that is transferred after central orchestration, used to handle complex cockpit intelligent tasks that require cross-service orchestration, personalization, and context integration, including the following process:
[0181] In process S11, both types of interfaces uniformly perform data integrity verification, format standardization, and security checks at the entry point, including event packet structure verification, media reference validity verification, task idempotent key identification, and authentication based on service access control lists (ACLs). For tasks containing media references, the preprocessing layer only parses metadata and reference identifiers, and large files are retrieved as needed in subsequent links according to the references, avoiding entry point blockage.
[0182] Process S12, Task Complexity Scoring and Intelligent Routing: To ensure the intelligent cockpit's response efficiency and deep reasoning capabilities under various tasks, a task complexity scoring module is set up in the cloud. This module generates a complexity score for each task through quantitative evaluation of multiple dimensions and selects the appropriate processing link accordingly.
[0183] The rating factors include:
[0184] Tool Invocation Complexity: The number of external services or tools required for the task to be invoked, such as maps, weather, POI, or user profiling services. Each additional external invocation increases the complexity score proportionally, reflecting potential latency and resource consumption in the process.
[0185] Multimodal data processing requirements: If the task involves multimodal fusion processing such as vision, audio, vehicle status, and historical trajectory, a certain weight will be added to the baseline score; single-modal tasks will score lower.
[0186] Intensity of personalized needs: If the task requires personalized reasoning based on user profiles and historical behavioral habits, the score will be increased based on the degree of dependence on personalized parameters; for tasks that only rely on common rules, the score will be lower.
[0187] Real-time requirements: If the task belongs to a strong real-time scenario of security incident or dialogue interaction, it should be marked as "time sensitive" in the complexity score. Even if the task itself calls many services, the weight should be adjusted to prioritize low-latency links.
[0188] The scoring mechanism uses a weighted cumulative method to calculate the task complexity score; the number of tool calls is weighted at 0.3, multimodal processing at 0.4, personalized needs at 0.2, and real-time marking is adjusted by a factor of 0.1 in the score.
[0189] The tool call complexity factor adopts a binary design. When the task involves at least one external tool call, the factor takes a value of 1.0; when the task does not involve any external tool calls, the factor takes a value of 0. In this way, the complexity score is avoided from exceeding the normalized range of 0 to 1 due to the cumulative number of tool calls, thereby ensuring the uniformity and controllability of the scoring model.
[0190] Routing strategy: When the score is below the first threshold θ1 (0.3), the task enters the fast link, where the lightweight engine performs basic analysis and response; when the score is between θ1 and θ2 (0.3–0.7), the task enters the hybrid link, which first returns preliminary results quickly and then calls the deep analysis system to supplement the results; when the score is above θ2 (0.7), the task enters the complex link, where deep reasoning and planning are performed by multimodal large models and cross-service orchestration.
[0191] In some embodiments, the complexity score threshold is not fixed, but dynamically adjusted based on system resource status (such as GPU utilization and network congestion) and historical task performance. For example, θ1 can be appropriately increased under high load to reduce the proportion of tasks entering complex links and ensure the overall stability of the cockpit response.
[0192] Process S13, Multimodal Deep Analysis Layer: In the fast link, the system prioritizes returning the minimum result set required for cockpit safety / interaction (such as structured prompts, basic explanations, and executable commands), ensuring that the interaction latency is in the sub-second to hundreds of millisecond range; when necessary, it immediately pulls keyframes or small-volume features for lightweight inference based on references; in the hybrid link, the system first sends back fast results to ensure response continuity, and then completes supplementary inference on the slow system or multimodal model (such as higher confidence recognition, regulatory / POI matching, and profile fusion), and writes back to the cockpit with result correction events; in the complex link, the decision system performs deep semantic understanding and planning based on the multimodal large model and cross-service orchestration, supports long temporal context and personalized synthetic output, and attaches executable templates and fallback strategies to sensitive commands (such as cockpit settings, media and navigation switching) during the generation stage;
[0193] In process S14, result generation and degradation / correction involves integrating the rapid results, supplementary inference, and complex inference results and returning them to the vehicle in real time. When there is link delay or insufficient resources, the process reverts to the rapid link to return the "minimum usable result," which is then supplemented and corrected later. In some embodiments, the result generation of the cloud-based intelligent decision-making layer includes the following types:
[0194] Perception-related results: These confirm or correct perception events reported by the vehicle and output the results of target detection and recognition. For example, the cloud verifies the category and confidence level of obstacles ahead, or performs secondary confirmation on traffic sign recognition results, thereby improving detection accuracy.
[0195] Semantic results: Combining contextual information and multimodal input, outputting results that reflect user intent and scene understanding; for example, generating an intent judgment of "navigate to the nearest gas station" based on driver voice commands and gesture recognition, or increasing the risk level in scenarios such as "night + rainy day + winding road section", achieving semantic-level comprehensive reasoning.
[0196] Decision-making results: Generate driving strategies and safety decision outputs; for example, the cloud provides driving change suggestions such as "decelerate", "change lanes" or "maintain distance" based on the fusion results, and selects the corresponding safety strategy (immediate reporting, delayed reporting or local prompt only).
[0197] Task-based results: Generate complex task outputs for cross-service requests or personalized needs; for example, when a user requests "plan a trip from Shanghai to Hangzhou", the cloud integrates map, weather and POI services to generate a complete trip planning result; or in recommendation tasks, combine user profiles and historical preferences to output personalized recommendations after sorting.
[0198] In process S15, the vehicle first applies for an upload permit from the cloud authentication service;
[0199] In process S16, a pre-signed URL and object_key are generated and returned in the cloud;
[0200] In process S17, the vehicle transmits large media files such as images / videos directly to the object storage in a segmented / resumable manner. The storage side returns the ETag / segment list after completion.
[0201] In process S18, the vehicle then reports a business request through the metadata API, which includes an object reference object_key, a verification hash (such as SHA-256) and context metadata, as well as an idempotency key idempotency_key with a data type tag data_type;
[0202] In process S19, the business hub receives and schedules tasks. The running service of the business hub selects different processing paths based on the data type markers as needed: when it is feature data, it prioritizes rapid parsing for immediate judgment of security-related events; when it is raw media data, it calls object storage on demand for deep inference for fine-grained analysis of complex tasks or personalized recommendations; when it is context information, it is used for task enhancement and result supplementation. The cloud records the content hash and metadata digest for each object_key, performing rapid consistency verification before business access. When duplicate reporting (network retries or cockpit playback) occurs, idempotent keys ensure that only one business side effect occurs, and subsequent requests return the status and result reference of the first task. This mode significantly reduces the bandwidth and computational pressure on the API surface, facilitating the introduction of CDN / edge acceleration and multi-region disaster recovery.
[0203] In process S20, when using real-time communication (SSE event stream), to ensure observability and fine-grained progress feedback on the cockpit side, the collaborative communication layer provides SSE (Server-Sent Events) real-time event streams. The vehicle subscribes to the event channel with a task identifier, and the cloud pushes task_created, plan_generated, tool_call_started, response_chunk, and task_completed events in stages. For network jitter, the system uses Last-Event-ID and heartbeat to maintain connection continuity and uses exponential backoff for reconnection and speed limiting. When the vehicle goes offline and comes back online, the event channel incrementally resends events based on Last-Event-ID to ensure the cockpit has complete awareness of the task status. If network degradation or excessive system load occurs during task execution, the business hub can also trigger degradation and correction mechanisms, so that the task returns the "minimum available result" in the SSE push first, and then sends the corrected result after the link is restored, thereby maintaining the consistency of the user experience and results on the vehicle side.
[0204] In step S10, the cloud operation includes degradation and correction, specifically: when network degradation or slow system load increases, causing complex / hybrid links to fail to complete within the specified time limit, the decision system's router automatically triggers a degradation strategy: prioritizing fallback to a fast link to return the "minimum available result," while using cached results to maintain user experience, and recording the task idempotency key and context so that execution can continue after the link recovers and the corrected result is sent out; for supplementary results that may cause user misunderstanding, version numbers and timestamps are used to mark them, and consistency correction is performed on the cockpit UI side (displaying "updated" or "corrected" prompts), thus forming a closed loop of real-time response—supplementary reasoning—consistency correction;
[0205] In step S10, after receiving the metadata API notification, the API gateway of the cloud-based business hub immediately parses its payload (such as a JSON message), reads the value of the "data_type" field, and based on this value, the cloud-based intelligent routing and scheduling module will execute the preset differentiated processing logic:
[0206] a) When data_type is "feature_vector", the routing module determines that this is a high-time-sensitivity task and immediately forwards the request to the low-latency real-time analysis service. This service can perform fast calculations (such as fatigue driving judgment) directly based on the feature vectors carried in the metadata, without waiting for the download of large media files, thus achieving a response time in seconds.
[0207] b) When data_type is "raw_media", the routing module pushes the task to a message queue (such as Kafka or RabbitMQ). The background deep learning inference cluster (usually equipped with GPU / NPU) acts as a consumer to asynchronously retrieve the task from the queue, pulls the corresponding media file from the object storage according to the object_key in the metadata, and performs complex and time-consuming analysis (such as complete semantic understanding of traffic scenarios).
[0208] c) When data_type is "context_meta", the request may be distributed to the data management service to update event logs or enrich user profiles without triggering large-scale computation.
[0209] Example:
[0210] In this example, when the decision-making system is integrated with external systems, such as Figure 7 In the external system integration framework shown,
[0211] Module 6 serves as the horizontal foundation for the object storage / logging and monitoring system: in addition to storing media files, the object storage can also store feature files and intermediate results to support secondary inference; the logging / monitoring system provides indicator dashboards and alarms, and provides closed-loop evidence in the adaptive tuning of scoring thresholds and link hysteresis. Through the collaboration of the above external systems, this example achieves stable service capabilities of low-latency response + deep inference + personalized correction in the intelligent cockpit scenario.
[0212] Module 7 is the user data center, used for cockpit personalization data, including: user profiles and historical behaviors (frequently visited locations, preferred media, interaction habits) are subject to hierarchical authorization and anonymization strategies, and are only authorized to participate in inference in personalized tasks; when the profile is unavailable or authorization is revoked, the system degrades to general output according to the default security policy.
[0213] Module 8 is the external vertical domain service module, establishing a unified, observable, and rate-limitable access between the cloud and external vertical domain services: the map service provides road features, speed limits / road segment attributes, and route constraints; the weather service provides micro-weather queries for routes and time windows; and the POI service provides structured attributes such as available facilities along the route (charging stations, service areas, restaurants, etc.) and their opening hours. All of these services are uniformly managed through a service catalog and call gateway. The request chain records latency, error codes, and retries before and after the call, used for adaptive optimization of route scoring.
Claims
1. A multimodal perception-based intelligent cockpit vehicle-cloud collaborative intelligent decision-making system, characterized in that: The decision-making system includes: The vehicle-side intelligent perception layer is used to execute computer vision models through pluggable devices on the vehicle side, collect multimodal data from driver monitoring, occupant monitoring, cameras, and radar, and generate event data through layered processing. The intelligent event triggering module is used to determine events based on detection results and context information, and to encapsulate the triggering events in a structured manner and set priorities. The collaborative communication layer is used to upload data using a method that separates object storage direct transmission from metadata API, and to provide feedback on task status and results through the SSE real-time communication protocol. The cloud-based intelligent decision-making layer receives the event data, selects a fast link, a hybrid link, or a complex link for processing based on the task complexity, and generates safety prompts, in-cabin interactive responses, or personalized recommendation results by combining user profiles and external services.
2. The intelligent cockpit vehicle-cloud collaborative intelligent decision-making system based on multimodal perception according to claim 1, characterized in that: The vehicle-side intelligent perception layer is based on a pluggable computer vision model architecture and combines multimodal sensors located inside and outside the vehicle cabin to perceive the environment inside and outside the cabin. The vehicle-side intelligent perception layer is used to detect driver fatigue, distraction, emotions and gestures, as well as to identify occupant status and behavior, thereby providing semantic input for cabin services and human-machine interaction. It aligns multimodal data through a unified timestamp and coordinate system, and adopts a hierarchical processing mechanism: L1 basic detection, L2 reporting and decision-making, L3 fine recognition and L4 semantic analysis, to achieve a step-by-step process from low-latency detection to high-level semantic understanding. Multimodal sensors inside and outside the vehicle include Driver Monitoring System (DMS), Occupant Monitoring System (OMS), in-cabin microphone array, forward-facing camera, and millimeter-wave radar; The vehicle-side intelligent perception layer is based on an intelligent event triggering mechanism, which identifies key event types inside and outside the cabin through rule engine and context information analysis. Contextual information includes vehicle operating status, geographical location, time, driver status, user profile, and weather conditions; Key incident types include both security incidents and interaction incidents; The intelligent event triggering mechanism classifies different events into risk levels through a security assessment module. The priority scoring model used for classification sorts events based on risk, time sensitivity, context, and personalized needs, ensuring that security events and high-value interactions are processed first. The data encapsulation module packages events into standardized data packets for further processing in the cloud. The collaborative communication layer adopts a "data flow and control flow separation" architecture, specifically: large volumes of media data are directly transmitted through object storage, supporting segmentation and breakpoint resumption; control messages are uploaded through metadata API to avoid redundant calculations. The collaborative communication layer provides a real-time event stream based on SSE for continuous feedback on task progress and supports reconnection after disconnection and incremental retransmission in weak network environments. The cloud-based intelligent decision-making layer includes dual input interfaces: a multimodal direct input interface to process real-time safety events reported by the vehicle to ensure low-latency response; The central routing input interface is used to handle complex cross-domain, multi-step tasks to support the orchestration and fusion of complex tasks; The cloud-based intelligent decision-making layer uses a cloud-based routing mechanism to score tasks based on a complexity assessment model and selects between fast links, hybrid modes, or complex links. The decision-making system includes a closed-loop execution and correction mechanism, specifically: when network degradation or high system load causes complex link delays, the system automatically falls back to the fast link to return the minimum usable result, and resends the corrected result after the link is restored; The revised version includes a version number and timestamp to ensure consistency for users.
3. The intelligent cockpit vehicle-cloud collaborative intelligent decision-making system based on multimodal perception according to claim 2, characterized in that: The cockpit domain controller includes at least one processor, memory, and peripheral interfaces, and communicates with other ECUs in the vehicle via in-vehicle Ethernet and / or CAN bus; The Driver Monitoring System (DMS) is deployed above the steering wheel to capture the driver's facial expressions. The Occupant Monitoring System (OMS) is deployed above the center console to detect occupant expressions and postures. The in-cabin microphone array is installed at the four corners of the roof to collect voice and ambient sound; A forward-facing camera and a millimeter-wave radar are installed at the front of the vehicle to collect information about the external environment. The sensors in the vehicle-side intelligent perception layer also include LiDAR, which is installed on the roof or bumper to enhance environmental modeling capabilities. The input data from the multimodal sensors to the decision-making system is integrated and processed by the cockpit domain controller to provide data support to the decision-making system. Multimodal sensors are connected to the cockpit domain controller via MIPI-CSI / FPD-Link, I²S / TDM, and Ethernet interfaces; raw / feature / context data are aligned with a unified timestamp and coordinate system before being input into the hierarchical perception and event triggering process of this invention.
4. The intelligent cockpit vehicle-cloud collaborative intelligent decision-making system based on multimodal perception according to claim 3, characterized in that: The collaborative communication layer adopts an "asynchronous object storage upload + metadata API notification" working mode to achieve separation of data flow and control flow. All data packets uploaded by the collaborative communication layer include a data type tag (data_type) to distinguish between raw media data (raw_media), feature vector data (feature_vector), and context information (context_meta). The addition of the data type tag is completed in the data encapsulation stage on the vehicle side. Specifically, before uploading to object storage, the system prepares a structured metadata notification and explicitly sets a field named "data_type" in the metadata according to the source and processing level of the data. After receiving the metadata API notification, the cloud selects different processing paths based on the data type tag. The asynchronous object storage upload process includes hardware calls and asynchronous execution. Specifically, the asynchronous object storage upload is executed by a dedicated communication module within the vehicle's cockpit domain controller. When the vehicle's intelligent perception layer generates a large volume of media data, the cockpit domain controller's main processor instruction communication module initializes this upload task, so that the task is managed by an independent background upload service process. This process is responsible for data fragmentation, packaging, and establishing a connection with the cloud object storage through the TCP / IP protocol stack to transmit data.
5. An intelligent cockpit vehicle-cloud collaborative intelligent decision-making method, using the intelligent cockpit vehicle-cloud collaborative intelligent decision-making system based on multimodal perception as described in claim 3, characterized in that: In the aforementioned decision-making method, the processing results in the cloud are sent to the vehicle's cockpit domain controller through the collaborative communication layer. The cockpit domain controller then distributes these results to the relevant ECUs and cockpit interaction modules throughout the vehicle, thereby enabling unified output of safety prompts, vehicle control, and personalized services across the entire vehicle.
6. The intelligent cockpit vehicle-cloud collaborative intelligent decision-making method according to claim 5, characterized in that: In the decision-making method, the vehicle-side intelligent perception layer includes a layered processing flow of basic detection, reporting decision, fine recognition, and semantic analysis; The intelligent event triggering module includes a rule engine, a context analysis module, a security assessment module, and a data encapsulation module; The context analysis module confirms the triggering conditions by combining vehicle operating status, geographical location, time, driver status, passenger status, and environmental conditions; The security assessment module outputs high, medium, and low risk levels and sorts the events using a priority scoring model. The working process of the collaborative communication layer includes uploading credential generation, direct transmission of object storage, and metadata API notification, wherein the metadata API carries object_key, verification hash, context information, and idempotent key; The collaborative communication layer pushes the task status of task_created, response_chunk, and task_completed through SSE real-time event stream, and supports disconnection reconnection and incremental resending. The cloud-based intelligent decision-making layer performs data integrity verification and access control during the preprocessing stage, and calculates task scores based on a complexity evaluation model. The complexity assessment model includes tool call factor, multimodal processing factor, personalized requirement factor and real-time correction factor, and selects fast link, hybrid link or complex link according to task score; The output of the cloud-based intelligent decision-making layer includes perception-based results, semantic-based results, decision-based results, and task-based results.
7. The intelligent cockpit vehicle-cloud collaborative intelligent decision-making method according to claim 6, characterized in that: In the aforementioned decision-making method, at the vehicle-side intelligent perception layer, the system adopts a modular, pluggable computer vision model architecture to achieve flexible expansion and dynamic management of different visual perception functions. In the pluggable computer vision model architecture, each model is provided as an independent plug-in package, which contains model files, configuration files, interface definition files, and preprocessing and postprocessing logic scripts. The decision-making method introduces a complete model management framework into the decision-making system. Plugin packages are registered through a plugin registry, enabling the decision-making system to identify their metadata at runtime, including version number, function category, and dependencies. The model management framework also includes a lifecycle management module, which is used to load, initialize, run and unload plugins, enabling the decision system to monitor their status during runtime. The resource scheduler is used to allocate computing resources between GPUs / CPUs, control memory usage, and ensure optimal performance when multiple plugins are running in parallel. The model management framework also includes a performance monitoring module for real-time statistics of inference time and resource utilization, and generates error logs and performance reports for diagnosis and optimization. The decision-making approach for security and version control specifically employs the following method: digital signature verification and permission checks are performed on the plugin package through a decision-making system to ensure its trustworthy origin and code security. The decision-making system runs plugins in a sandbox environment to prevent malicious plugin behavior from affecting the main system. The decision-making system supports compatibility checks and rollback operations for plugins through the version management module, so that when an anomaly occurs in the new version, it can be restored to a stable version. The decision-making system's update mechanism is based on OTA differential updates to push updates in batches without interrupting services, and automatically performs verification after the update is completed. During the runtime of the decision system, the environment layer model of the decision system is executed in a unified manner at runtime and interacts with other components of the vehicle system through the interface adaptation layer to achieve unified conversion of input and output data formats. An event bus is used for inter-plugin communication and message distribution, and a task scheduler is used to prioritize and manage different tasks to ensure that critical tasks are executed first under high load scenarios. The decision-making approach supports integration with external management platforms through a decision-making system, uses a cloud-based management platform to uniformly release plugin versions, push updates, and perform performance analysis; and provides local debugging, log viewing, and performance testing capabilities through developer tools. By using the integration interface of the decision system, the plugin can be made to work together with other services or hardware abstraction layers of the vehicle system, thus achieving integrated closed-loop management from development, deployment to operation.
8. The intelligent cockpit vehicle-cloud collaborative intelligent decision-making method according to claim 6, characterized in that: The decision-making method includes a layered perception process on the vehicle side, comprising the following steps: Step S1: The vehicle first timestamps and aligns the multimodal data from the camera, millimeter-wave radar, lidar and DMS / OMS with the coordinate system, and then processes them step by step according to the L1-L4 hierarchical perception process. Step S2: The L1 basic detection layer outputs the basic results of the bounding boxes, categories, and confidence scores of objects / lane lines / traffic signs; Step S3: The L2 reporting decision layer, based on L1 and combined with the vehicle operating status execution rules, decides whether to generate a reporting event and sets the priority and initial route suggestion; Step S4: The L3 fine recognition layer performs higher-precision recognition of key targets, enhancing interaction and scene understanding; In step S5, the L4 semantic analysis layer integrates the L1–L3 results with the context of vehicle status, geographical location, and historical trajectory to construct a scene graph and perform intent reasoning to generate semantically complex events. In the event triggering mechanism of the vehicle-side intelligent perception layer, the system performs routine detection and recognition of sensor data, and also realizes real-time response to key driving scenarios through a multimodal event triggering mechanism. This mechanism is completed collaboratively by the rule engine, context analysis module, safety assessment module and data encapsulation module to ensure that events have high priority and reliability throughout the entire process of triggering, judging and reporting. Step S6: Use a rule engine module based on a predefined rule base and dynamic update strategy to quickly compare the detection results of CV model, radar, and DMS / OMS multi-source inputs; The rule base not only stores commonly used rules locally, but also supports the dynamic delivery of new rules from the cloud via OTA. Step S7: Through the context analysis module, the detection results are combined with the vehicle's operating status and environmental factors to reconfirm the triggering conditions. Specifically, this includes: geographical location analysis, time factor consideration, and user status assessment. Context enhancement helps the decision-making system avoid false alarms from a single sensor and improves the accuracy of event triggering. Step S8: After the incident is initially confirmed, the security assessment module further determines its risk level and urgency, and selects the corresponding security strategy according to different levels. Step S9: The data encapsulation module performs structured packaging of information related to the triggering event; this process includes: keyframe extraction, context data encapsulation, integration of preliminary analysis results, and addition of triggering cause markers and priority information; Step S10: The packaged event packet is uploaded to the cloud through the collaborative communication layer. High-priority events are reported immediately through a low-latency link, while medium- and low-priority events are transmitted with delay through a regular link. The events are then parsed and differentiated in the cloud. Through the layered filtering and enhancement of the above-mentioned triggering mechanism, a complete closed loop of "detection-judgment-evaluation-encapsulation-reporting" is achieved at the vehicle end, providing a fast and reliable response in safety-critical scenarios, and ensuring that the event data uploaded to the cloud has structured, context-complete, and traceable characteristics, thereby improving the accuracy and efficiency of subsequent cloud-based decision analysis.
9. The intelligent cockpit vehicle-cloud collaborative intelligent decision-making method according to claim 8, characterized in that: In step S8, the priority assessment method adopts a weighted scoring model, which is expressed by the following formula: P=α⋅R+β⋅T+γ⋅C+δ⋅E−ε⋅L; Where R represents risk level, T represents time sensitivity, C represents contextual factors, E represents event category weight, L represents system load status, and α, β, γ, δ, and ε are the weight coefficients of each dimension; the risk level R is divided into three levels: high, medium, and low and mapped to normalized interval values. High-risk events correspond to R=1.0, medium-risk events correspond to R=0.6, low-risk events correspond to R=0.3, and non-safe events correspond to R=0.1; The time sensitivity T is assigned a value based on the task's real-time requirements; The context C is determined based on external conditions; Category weight E is preset based on the event category; The system load status L is determined based on the vehicle-side computing and communication resource status; After calculating the priority score P, the system divides it into three levels: when P≥0.8, the event is judged as high priority and reported immediately through a low-latency link; when 0.4≤P<0.8, the event is judged as medium priority and reported through a regular link; when P<0.4, the event is judged as low priority and can be reported after delay or batch aggregation.
10. The intelligent cockpit vehicle-cloud collaborative intelligent decision-making method according to claim 8, characterized in that: As shown in the flowchart of the cloud-based intelligent decision-making layer in Figure 6, in step S10, the cloud sets up dual input interfaces to handle two types of tasks respectively: one is multimodal direct input from the vehicle, which is used to handle cockpit perception events that are related to safety or have high real-time requirements. The second is the central routing input, which is transferred after central orchestration. This is used to handle complex cockpit intelligence tasks that require cross-service orchestration, personalization, and context integration, including the following processes: In process S11, both types of interfaces uniformly perform data integrity verification, format standardization, and security checks at the entry point, including event packet structure verification, media reference validity verification, task idempotent key identification, and authentication based on service access control lists (ACLs). For tasks containing media references, the preprocessing layer only parses metadata and reference identifiers, and large files are retrieved as needed in subsequent links according to the references, avoiding entry point blockage. Process S12, Task Complexity Scoring and Intelligent Routing: The task complexity scoring module is set up in the cloud. This module generates a complexity score for each task through quantitative evaluation of multiple dimensions and selects the appropriate processing link accordingly. The rating factors include: Tool call complexity: The number of external services or tools that the task needs to call. For each additional external call, the complexity score increases proportionally, reflecting the potential latency and resource consumption in the chain. Multimodal data processing requirements: If the task involves multimodal fusion processing, a certain weight will be added to the baseline score; single-modal tasks will receive lower scores. Personalization requirement intensity: If the task requires personalized reasoning based on user profiles and historical behavioral habits, the score will be increased based on the degree of dependence on personalized parameters. Real-time requirements: If the task belongs to a strong real-time scenario of security incident or dialogue interaction, then mark it as "time sensitive" in the complexity score, and select low-latency links by adjusting the weights. The scoring mechanism uses a weighted cumulative method to calculate the task complexity score; The tool call complexity factor adopts a binary design. When the task involves at least one external tool call, the factor takes a value of 1.0; when the task does not involve any external tool call, the factor takes a value of 0. Routing strategy: When the score is below the first threshold θ1, the task enters the fast link, where the lightweight engine performs basic analysis and response; when the score is between θ1 and θ2, the task enters the hybrid link, which first returns preliminary results quickly and then calls the deep analysis system to supplement the results; when the score is above θ2, the task enters the complex link, where multimodal large models and cross-service orchestration perform deep reasoning and planning. Process S13, Multimodal Deep Analysis Layer: In the fast link, the system prioritizes returning the minimum result set required for cockpit safety / interaction, ensuring that the interaction latency is in the sub-second to hundreds of millisecond range; when necessary, keyframes or small features are pulled in real time by reference for lightweight inference; in the hybrid link, the system first sends back fast results to ensure response continuity, then completes supplementary inference on the slow system or multimodal model, and writes back to the cockpit with result correction events; in the complex link, the decision system performs deep semantic understanding and planning based on the multimodal large model and cross-service orchestration, supports long temporal context and personalized synthetic output, and attaches executable templates and fallback strategies to sensitive instructions during the generation phase; In process S14, result generation and degradation / correction are performed by integrating the fast results, supplementary inference, and complex inference results and returning them to the vehicle in real time. When there is a link delay or insufficient resources, the process falls back to the fast link to return the "minimum available result," which is then supplemented and corrected later. The result generation of the cloud-based intelligent decision-making layer includes the following types: Perception-related results: Confirm or correct the perception events reported by the vehicle, and output the results of target detection and recognition; Semantic results: Combining contextual information and multimodal input, outputting results reflecting user intent and scene understanding; Decision-related results: Generate driving strategy and safety decision outputs; Task-related results: Generate complex task outputs for cross-service requests or personalized needs; In process S15, the vehicle first applies for an upload permit from the cloud authentication service; In process S16, a pre-signed URL and object_key are generated and returned in the cloud; In process S17, the vehicle transmits large media files such as images / videos directly to the object storage in a segmented / resumable manner. The storage side returns the ETag / segment list after completion. In process S18, the vehicle then reports a business request through the metadata API, which includes an object reference object_key, a verification hash and context metadata, as well as an idempotency key idempotency_key with a data type tag data_type. In process S19, the business hub receives and schedules tasks. The running service of the business hub selects different processing paths based on the data type marker when needed: when it is feature data, it prioritizes rapid parsing for immediate judgment of security-related events; when it is raw media data, it calls object storage on demand for deep inference for fine-grained analysis of complex tasks or personalized recommendations; when it is context information, it is used for task enhancement and result supplementation. The cloud records the content hash and metadata summary for each object_key, performing rapid consistency verification before business access. When duplicate reporting occurs, idempotent keys ensure that only one business side effect occurs, and subsequent requests return the status and result reference of the first task. This mode reduces bandwidth and computational pressure on the API surface, facilitating the introduction of CDN / edge acceleration and multi-region disaster recovery. In process S20, when using real-time communication, to ensure observability and fine-grained progress feedback on the cockpit side, the collaborative communication layer provides SSE real-time event streams. The vehicle subscribes to the event channel with a task identifier, and the cloud pushes task_created, plan_generated, tool_call_started, response_chunk, and task_completed events in stages. For network jitter, the system uses Last-Event-ID and heartbeat to maintain connection continuity and uses exponential backoff for reconnection and speed limiting. When the vehicle goes offline and comes back online, the event channel incrementally resends events based on Last-Event-ID to ensure the cockpit has complete awareness of the task status. If network degradation or excessive system load occurs during task execution, the business hub can also trigger degradation and correction mechanisms, so that the task returns the "minimum available result" in the SSE push first, and then sends the corrected result after the link is restored, thereby maintaining the consistency of the user experience and results on the vehicle side. In step S10, the cloud operation includes degradation and correction, specifically: when network degradation or slow system load increases, causing complex / hybrid links to fail to complete within the specified time limit, the decision system's router automatically triggers a degradation strategy: prioritizing fallback to a fast link to return the "minimum available result," while using cached results to maintain user experience, and recording the task idempotency key and context so that execution can continue after the link recovers and the corrected result is issued; for supplementary results that may cause user misunderstanding, version numbers and timestamps are used to mark them, and consistency correction is performed on the cockpit UI side (displaying "updated" or "corrected" prompts), thereby forming a closed loop of real-time response—supplementary reasoning—consistency correction; In step S10, after receiving the metadata API notification, the API gateway of the cloud-based business hub immediately parses its payload and reads the value of the "data_type" field. Based on this value, the cloud-based intelligent routing and scheduling module will execute the preset differentiated processing logic.