Vehicle active service system based on end-side multi-modal large model and dynamic reasoning method thereof
Through the end-side multimodal large model and MCP dynamic inference controller, the network delay and privacy leakage problems of the vehicle system are solved, real-time and personalized vehicle proactive service decisions are realized, and the system's response speed and decision accuracy are improved.
Patent Information
- Application Number
- CN202511205449.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing in-vehicle active service systems rely on cloud processing, resulting in high network latency, slow response time, high risk of privacy leakage, and lack of dynamic reasoning mechanisms, making them unable to meet real-time and personalized service requirements.
A large multimodal model on the end is adopted, combined with multimodal data collection, cross-modal attention fusion and MCP dynamic inference controller to realize proactive service decision-making in vehicle scenarios. The model is compressed to the vehicle end through knowledge distillation and quantification technology, and supplementary data is dynamically requested for multi-step reasoning.
It realizes real-time response and intelligent decision-making of the vehicle-mounted active service system, improves reasoning flexibility and environmental adaptability, reduces computing load and energy consumption, and ensures local processing and privacy security of user data.
Smart Images

Figure CN120724397A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle networking technology, and in particular to a vehicle active service system based on a terminal-side multimodal large model and a dynamic reasoning method thereof. Background Art
[0002] With the rapid development of intelligent vehicles, in-vehicle proactive service systems have become an important technology for improving the driving experience. Existing in-vehicle systems mainly rely on the driver's active operation to trigger corresponding functions, but still have many shortcomings in understanding driving scenarios and providing personalized services.
[0003] Traditional in-vehicle service systems typically use large cloud-based models for data processing and decision-making reasoning, which has the following technical flaws: First, cloud-based reliance leads to high network latency, with response times typically exceeding 500-1000 milliseconds, making it difficult to meet the real-time requirements of in-vehicle scenarios; second, user data needs to be uploaded to the cloud for processing, posing a risk of privacy leakage; third, the recognition accuracy of single-modality sensors is limited, especially in low-light environments, where traditional RGB cameras cannot effectively identify complex driving scenarios; finally, existing systems lack a dynamic reasoning mechanism and are unable to dynamically obtain the required parameters based on the intermediate results of the reasoning process, resulting in inaccurate and inflexible reasoning decisions.
[0004] The invention patent with the Chinese publication number CN115175138B discloses a method and an in-vehicle system for providing online services for in-vehicle applications based on user characteristics. The patent detects the characteristic data of each user currently served by multiple second in-vehicle devices through a first in-vehicle device, determines the user's in-vehicle application usage preference information based on the characteristic data, obtains the corresponding online access token and shares it with the second in-vehicle device. When the target in-vehicle device receives the user's online access instruction, it obtains the corresponding token from the shared access token and provides online services to the user based on the access instruction and token. This technical solution mainly solves the problem that users can enjoy personalized services without registering a personal in-vehicle account. However, its service triggering still relies on the user's active operation rather than the system's active perception, and lacks the ability to understand the environment. Summary of the Invention
[0005] The technical solution of the present invention is implemented as follows: The present invention provides a vehicle active service system based on a large multi-modal model on the terminal side, comprising: A multimodal data acquisition module is used to collect multimodal perception data of the vehicle; wherein the multimodal perception data includes image data, in-vehicle audio signal data and vehicle driving status parameter data; A multimodal large model processing unit is used to extract features from multimodal perception data to obtain features of different modalities. The multimodal large model is compressed using knowledge distillation and quantization and deployed on the vehicle side. The cross-modal attention fusion module is used to convert feature data of different modalities into query, key, and value triplets, calculate cross-modal similarity weights through a bidirectional attention mechanism, and fuse the query, key, value triplets and cross-modal similarity weights to obtain fused features; The MCP dynamic inference controller is used to make a preliminary judgment on the vehicle scene based on the fusion features. When the preliminary judgment result needs further confirmation, it generates supplementary data collection instructions to control the multimodal data collection module to collect specific parameter data. The supplementary data is combined with the fusion features to perform multi-step reasoning to determine the current vehicle scene type, and the user profile and historical behavior data are combined to generate proactive service decisions; The active service execution module is used to receive active service decisions, generate vehicle-mounted equipment control instructions and execute active service operations.
[0006] Based on the above technical solution, preferably, the image data includes RGB image data and infrared image data, and the multimodal data acquisition module further includes: A brightness scoring unit is used to convert the RGB image data into a color space and calculate a brightness score value. When the brightness score value is lower than a preset threshold, the visible light camera is switched to the infrared camera for image acquisition. The infrared vision enhancement unit is used to collect infrared thermal radiation image data using an infrared camera, and perform thermal radiation feature extraction and dynamic exposure compensation on the infrared thermal radiation image to obtain infrared image data.
[0007] Based on the above technical solution, preferably, the multimodal large model processing unit adopts a teacher-student network architecture to perform knowledge distillation, specifically including: In the output layer distillation, the output probability distributions of the teacher model and the student model are smoothed, and the difference between the smoothed probability distributions of the teacher model and the student model is calculated using the KL divergence loss function. The KL divergence loss is weightedly combined with the cross entropy loss of the student model on the true label to form a joint loss function for optimization training; In the intermediate feature distillation, the intermediate feature map sizes of the teacher model and the student model are adjusted through a 1×1 convolution adaptation layer, and the mean square error loss function is used to align the intermediate feature representations of the teacher model and the student model; The 8-bit integer quantization technology is used to quantize the weight parameters and activation values of the student model. The model parameter storage accuracy is reduced by channel-by-channel quantization, and the large multimodal model is compressed to less than 500MB.
[0008] Based on the above technical solution, preferably, the step of converting the feature data of different modalities into a query, key, and value triplet specifically includes: A visual transformer encoder is used to extract visual feature vectors from image data, an audio embedding feature vector is extracted from audio signal data, and a state feature vector is generated by feature encoding the vehicle driving state parameter data; The eigenvectors of different modes are projected and transformed into query vector, key vector and value vector respectively through linear transformation matrix; The feature dimensions of the query vector, key vector, and value vector are uniformly set to the same feature space dimension to form a triplet representation of uniform dimension.
[0009] Based on the above technical solution, preferably, the calculation method of the bidirectional attention mechanism specifically includes: The forward attention weight matrix is obtained by performing a dot product operation on the query vector and the key vector and applying softmax normalization based on the scaling factor. The inverse attention weight matrix is obtained by performing a transposed dot product operation on the key vector and the query vector and applying softmax normalization based on the scaling factor. The forward attention weight matrix and the value vector are weightedly summed to obtain the forward attention output, and the reverse attention weight matrix and the query vector are weighted summed to obtain the reverse attention output. The weighted average calculation is performed based on the forward attention output and the reverse attention output, and after residual connection and layer normalization processing, the fusion feature is output.
[0010] Based on the above technical solution, preferably, the MCP dynamic inference controller includes a dynamic inference client and multiple multi-parameter servers, and the processing steps of the MCP dynamic inference controller specifically include: The dynamic reasoning client communicates with the multi-parameter server through a multi-step conditional reasoning protocol. When it starts, it automatically identifies all available multi-parameter servers and obtains the data collection capability description information provided by each server. Based on the fusion features, a preset scene classification model is used to perform a preliminary classification judgment of the vehicle scene, and a preliminary judgment result including a scene type identifier and a confidence value is generated; When the confidence value of the preliminary judgment result is lower than the set threshold, the dynamic reasoning client sends a supplementary data collection instruction to the corresponding multi-parameter server according to the scene type identifier.
[0011] Based on the above technical solutions, preferably, the method of generating proactive service decisions by combining user profiles and historical behavior data specifically includes: extracting physiological and behavioral characteristic parameters of the driver as first inference parameters using a driving state detection algorithm, wherein the driving state detection algorithm includes at least one of a fatigue detection algorithm based on facial key point detection, a steering wheel grip detection algorithm, and an attention distraction detection algorithm; When the first inference parameter indicates an abnormal driving state, the dynamic inference client requests the multi-parameter server to collect the corresponding verification parameter as the second inference parameter, and then performs feature splicing on the first and second inference parameters and inputs them into the multi-step inference network for scene confirmation; User personal preference features are extracted from the user portrait database, and historical behavior pattern features are extracted from the database. The confirmed scenario type is weighted and fused with the user personal preference features and historical behavior pattern features to generate proactive service decisions, where proactive service decisions include service type, service parameters, and execution timing.
[0012] On the basis of the above technical solution, preferably, the active service execution module includes: A decision parsing unit is used to parse the service type, service parameters, and execution timing information in the active service decision, match the corresponding vehicle device control protocol according to the service type, and generate device control parameters; an instruction generating unit, configured to generate control instructions according to the device control parameters, wherein the control instructions include seat vibration control instructions, air conditioning temperature adjustment control instructions, navigation rest area search control instructions, and audio volume adjustment control instructions; The execution monitoring unit is used to monitor the execution status of the control instructions of the vehicle-mounted equipment through the equipment execution interface, record the control instruction issuance time, execution completion time and execution result status, and generate execution feedback information including execution success rate and response time.
[0013] More preferably, the active service execution module further includes a user interaction and feedback processing mechanism: Before executing a proactive service operation, a proactive service semantic prompt message is generated through speech synthesis technology to explain the detected driving scenario and the service measures to be taken to the user. At the same time, a proactive service confirmation card with confirmation and cancellation options is displayed on the in-vehicle display interface; Receive feedback information input by the user through voice commands or touch operations, execute the vehicle-mounted device control instructions when confirmation feedback is received, terminate the current service execution process when cancellation feedback is received, and update the user's feedback information as personalized preference data to the database.
[0014] The present invention also provides a dynamic reasoning method for vehicle proactive services based on a large, on-device multimodal model, which is applied to the vehicle proactive service system based on the large, on-device multimodal model described above, and includes the following steps: S1. Collecting multimodal perception data of the vehicle; wherein the multimodal perception data includes image data, in-vehicle audio signal data, and vehicle driving state parameter data; S2. Feature extraction is performed on multimodal perception data to obtain features of different modalities. The multimodal large model is compressed using knowledge distillation and quantization and deployed on the vehicle side. S3. Convert the feature data of different modalities into query, key, and value triplets, calculate the cross-modal similarity weight through a bidirectional attention mechanism, and fuse them based on the query, key, and value triplets and the cross-modal similarity weight to obtain the fused features; S4. Perform a preliminary judgment on the vehicle scenario based on the fused features. When the preliminary judgment requires further confirmation, generate a supplementary data collection instruction to control the multimodal data collection module to collect specific parameter data, combine the supplementary data with the fused features to perform multi-step reasoning, determine the current vehicle scenario type, and generate proactive service decisions based on user profiles and historical behavior data; S5. Receive active service decisions, generate vehicle-mounted equipment control instructions, and execute active service operations.
[0015] The vehicle active service system based on the terminal-side multimodal large model and its dynamic reasoning method have the following advantages over the prior art: (1) Through the end-side multimodal large model and MCP dynamic reasoning protocol, the real-time response and intelligent decision-making capabilities of the vehicle-mounted active service system are realized. Through multimodal feature fusion and cross-modal attention mechanism, the system transforms from passive response to active service mode, significantly improving reasoning flexibility and environmental adaptability. (2) By using knowledge distillation technology to compress large-scale multimodal models into lightweight versions suitable for on-board deployment, and through the teacher-student network architecture and intermediate feature alignment mechanism, the model storage requirements and computational complexity are greatly reduced while maintaining inference accuracy, eliminating dependence on network connections and achieving millisecond-level response speeds; (3) The MCP protocol adopts a multi-step reasoning network architecture, which can decide whether to supplement the second reasoning parameter based on the confidence assessment of the first reasoning parameter, realizing on-demand allocation of computing resources. While ensuring the accuracy of decision-making, it significantly reduces the computing load and energy consumption of the end-side equipment, and improves the overall operating efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 This is an architecture diagram of the vehicle active service system based on the end-side multimodal large model of the present invention; Figure 2 This is a hardware deployment diagram of the vehicle active service system based on the end-side multimodal large model of the present invention; Figure 3 This is a flow chart of the MCP protocol of the vehicle active service system based on the end-side multimodal large model of the present invention; Figure 4 Schematic diagram of the MCP protocol of the vehicle active service system based on the end-side multimodal large model of the present invention; Figure 5 This is a schematic diagram of the key points of the eyes in the fatigue detection algorithm of the vehicle active service system based on the end-side multimodal large model of the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0019] like Figure 1 As shown, the present invention provides a vehicle active service system based on a large multi-modal model on the terminal side, including: A multimodal data acquisition module is used to collect multimodal perception data of the vehicle; wherein the multimodal perception data includes image data, in-vehicle audio signal data and vehicle driving status parameter data; A multimodal large model processing unit is used to extract features from multimodal perception data to obtain features of different modalities. The multimodal large model is compressed using knowledge distillation and quantization and deployed on the vehicle side. The cross-modal attention fusion module is used to convert feature data of different modalities into query, key, and value triplets, calculate cross-modal similarity weights through a bidirectional attention mechanism, and fuse the query, key, value triplets and cross-modal similarity weights to obtain fused features; The MCP dynamic inference controller is used to make a preliminary judgment on the vehicle scene based on the fusion features. When the preliminary judgment result needs further confirmation, it generates supplementary data collection instructions to control the multimodal data collection module to collect specific parameter data. The supplementary data is combined with the fusion features to perform multi-step reasoning to determine the current vehicle scene type, and the user profile and historical behavior data are combined to generate proactive service decisions; The active service execution module is used to receive active service decisions, generate vehicle-mounted equipment control instructions and execute active service operations.
[0020] This invention realizes the real-time response and intelligent decision-making capabilities of the in-vehicle active service system through the end-side multimodal large model and the MCP dynamic reasoning protocol. Through multimodal feature fusion and cross-modal attention mechanism, it transforms from a passive response to an active service mode, significantly improving reasoning flexibility and environmental adaptability, and providing a more reliable technical solution for in-vehicle intelligent interaction.
[0021] like Figure 2 As shown in the figure, the vehicle active service system includes data acquisition components such as RGB cameras, infrared cameras, voice sensors and on-board SDKs. It intelligently switches visual sensors through the brightness scoring unit, extracts multimodal features through ASR voice recognition and VLM visual language model, and performs preliminary scene analysis through the active service inference engine; MCP protocol MCPServer (multi-parameter server) 1, 2, 3 and other parameter servers are centrally coordinated and managed through MCP Client (dynamic inference client), and the inference agent multi-step inference module dynamically requests additional parameters based on the preliminary analysis results and performs multi-step conditional inference; the active service execution module includes an active service engine and an execution engine, and controls specific on-board hardware to perform active service operations through on-board device interfaces such as RYT SDK1, SDK2, SDK3. The entire architecture realizes modular collaborative work of data acquisition, inference decision-making and service execution.
[0022] In one embodiment of the present invention, the image data includes RGB image data and infrared image data, and the multimodal data acquisition module further includes: A brightness scoring unit is used to convert the RGB image data into a color space and calculate a brightness score value. When the brightness score value is lower than a preset threshold, the visible light camera is switched to the infrared camera for image acquisition. The infrared vision enhancement unit is used to collect infrared thermal radiation image data using an infrared camera, and perform thermal radiation feature extraction and dynamic exposure compensation on the infrared thermal radiation image to obtain infrared image data.
[0023] As you can understand, RGB image data is captured using a high-definition RGB camera installed near the rearview mirror. This camera primarily captures visual information such as the driver's facial expressions, eye state, and head posture. Infrared thermal radiation image data is captured using an 850nm infrared array camera, which forms images by detecting infrared radiation emitted by objects of different temperatures. A thermal radiation feature extraction algorithm analyzes the radiation intensity distribution in different regions of the infrared image to identify key features such as the facial contour and eye area. A dynamic exposure compensation algorithm automatically adjusts the camera's exposure parameters based on the ambient temperature and the thermal radiation intensity of the target object, ensuring appropriate contrast and clarity in the infrared image. The processed infrared image data has the same resolution and frame rate as the visible light image, ensuring consistent data format.
[0024] Specifically, the low-light recognition rate of the present invention is compared with the traditional cloud solution and the local fixed model. The data is shown in Table 1: Table 1 As can be seen from Table 1, this solution significantly improves the dark light recognition rate and image detection accuracy through the brightness scoring unit and the infrared vision enhancement unit.
[0025] In one embodiment of the present invention, the brightness scoring algorithm is implemented using a V channel detection method based on the HSV color space, specifically including: Convert the image captured by the RGB camera from the RGB color space to the HSV format. The V channel in the HSV color space directly represents the brightness value. The V channel data is extracted and two algorithms are used to calculate the brightness score in parallel: the average method calculates the average value of all pixels in the V channel. When the average value is lower than the 30% threshold, the environment is judged to be too dark; the proportion method counts the proportion of pixels with V values less than 20% of the total pixels. When this proportion exceeds 50%, it is judged to be a dark environment. When both algorithms meet the dark light judgment conditions at the same time, they automatically switch from the RGB camera to the infrared camera for image acquisition.
[0026] In one embodiment of the present invention, in-vehicle audio signal data is collected by an in-vehicle microphone array, including sound information such as the driver's voice, in-vehicle conversations, and ambient noise. This in-vehicle audio signal data is preprocessed using noise reduction, echo cancellation, and voice enhancement. The noise reduction algorithm uses adaptive filtering technology to effectively suppress background interference such as engine noise and road noise. The voice enhancement algorithm uses spectral analysis and time-domain processing to highlight the human voice frequency band and improve the clarity of the voice signal.
[0027] In one embodiment of the present invention, vehicle status sensors are used to collect vehicle operating status parameter data, including key parameters such as vehicle speed, steering wheel angle, brake status, turn signal status, gear position information, etc. These sensors are connected to the vehicle's electronic control unit via the onboard CAN bus, enabling real-time acquisition of various vehicle operating status information.
[0028] In one embodiment of the present invention, the multimodal large model processing unit adopts a sub-modal encoder design, using specialized feature extraction networks for different types of input data. For image data, a visual transformer encoder (ViTEncoder) is used to extract visual feature vectors. It can process both RGB and infrared image data. The ViT encoder divides the input image into fixed-size image blocks, each of which is converted into an initial embedding vector via a linear projection layer. Hierarchical visual features are then extracted through a multi-layer Transformer structure. For audio signal data, an audio feature extraction network is used to generate audio embedding feature vectors. The audio processing process includes pre-emphasis, time domain segmentation and smoothing, and fast Fourier transform (FFT). Spectral and temporal features are extracted through a deep convolutional neural network. The audio feature vectors encode multi-dimensional information such as speech content, speaker characteristics, and emotional state. Vehicle driving state parameter data is processed through a specialized state feature encoder. This encoder uses a multi-layer perceptron structure to uniformly encode discrete and continuous parameters such as vehicle speed, steering wheel angle, brake status, and turn signal status into a state feature vector. The dimensions of the state feature vector are consistent with those of the visual and audio feature vectors.
[0029] Specifically, the multimodal large model processing unit adopts a teacher-student network architecture to perform knowledge distillation, specifically including: In the output layer distillation, the output probability distributions of the teacher model and the student model are smoothed, and the difference between the smoothed probability distributions of the teacher model and the student model is calculated using the KL divergence loss function. The KL divergence loss is weightedly combined with the cross entropy loss of the student model on the true label to form a joint loss function for optimization training; In the intermediate feature distillation, the intermediate feature map sizes of the teacher model and the student model are adjusted through a 1×1 convolution adaptation layer, and the mean square error loss function is used to align the intermediate feature representations of the teacher model and the student model; The 8-bit integer quantization technology is used to quantize the weight parameters and activation values of the student model. The model parameter storage accuracy is reduced by channel-by-channel quantization, and the large multimodal model is compressed to less than 500MB.
[0030] Furthermore, the quantization process uses post-training quantization (PTQ), enabling model compression without retraining. Weight quantization employs a channel-by-channel approach, calculating independent quantization parameters for each output channel of each convolutional or fully connected layer. Activation quantization employs a dynamic quantization strategy, dynamically calculating quantization parameters during inference based on the actual distribution of input data. For each activation tensor, the quantization parameters are updated using a moving average to adapt to changes in the input data distribution.
[0031] Specifically, quantization uses two strategies: post-training quantization (PTQ) and quantization-aware training (QAT). PTQ requires no retraining and is suitable for deployment. The specific process includes preparing the original FP32 model, collecting 100-1000 representative calibration data samples, applying quantization algorithms such as MinMax, Histogram, or MSE approximation, and exporting the model in INT8 / INT4 format. Weight quantization uses a channel-by-channel approach, calculating independent quantization parameters for each output channel of each convolutional or fully connected layer. Activation quantization uses a dynamic quantization strategy, dynamically calculating quantization parameters based on the actual distribution of the input data and updating them using a moving average to adapt to changes in the input data distribution.
[0032] This invention uses knowledge distillation technology to compress large-scale multimodal models into lightweight versions suitable for on-board deployment. Through the teacher-student network architecture and intermediate feature alignment mechanism, it greatly reduces the model storage requirements and computational complexity while maintaining inference accuracy, eliminates dependence on network connections, achieves millisecond-level response speeds, and ensures local processing of user data, effectively solving the technical problems of network latency and privacy security in cloud solutions.
[0033] In one embodiment of the present invention, the cross-modal attention fusion module receives visual feature vectors, audio embedding feature vectors, and state feature vectors from the multimodal large model processing unit. The cross-modal attention fusion module performs dimensionality unification and numerical normalization on the input feature vectors. The dimensions of all feature vectors are uniformly set to 512, represented by 32-bit floating-point numbers. Features of different original dimensions are mapped to a unified feature space through a linear projection layer. Feature vector normalization uses the L2 normalization method.
[0034] Specifically, the conversion of feature data of different modalities into query, key, and value triples includes: A visual transformer encoder is used to extract visual feature vectors from image data, an audio embedding feature vector is extracted from audio signal data, and a state feature vector is generated by feature encoding the vehicle driving state parameter data; The eigenvectors of different modes are projected and transformed into query vector Q, key vector K and value vector V respectively through linear transformation matrix; The feature dimensions of the query vector Q, key vector K, and value vector V are uniformly set to the same feature space dimension to form a triplet representation of uniform dimension.
[0035] In one embodiment of the present invention, the calculation method of the bidirectional attention mechanism specifically includes: The forward attention weight matrix is obtained by performing a dot product operation on the query vector Q and the key vector K and applying softmax normalization based on the scaling factor. The calculation formula is: in, represents the forward attention weight matrix, represents the softmax normalization function, represents the query vector, represents the key vector, represents transpose, Indicates the feature space dimension, which is 512 in the present invention. represents the scaling factor; The inverse attention weight matrix is obtained by performing a transposed dot product operation on the key vector K and the query vector Q, and applying softmax normalization based on the scaling factor. The calculation formula is: in, represents the inverse attention weight matrix; The forward attention weight matrix and the value vector are weightedly summed to obtain the forward attention output, and the reverse attention weight matrix and the query vector are weighted summed to obtain the reverse attention output. The weighted average calculation is performed based on the forward attention output and the reverse attention output, and after residual connection and layer normalization processing, the fusion feature is output.
[0036] In one embodiment of the present invention, the MCP dynamic inference controller includes a dynamic inference client (MCPClient) and multiple multi-parameter servers (MCPServer). The processing steps of the MCP dynamic inference controller specifically include: The dynamic reasoning client communicates with the multi-parameter server through a multi-step conditional reasoning protocol. When it starts, it automatically identifies all available multi-parameter servers and obtains the data collection capability description information provided by each server. Based on the fusion features, a preset scene classification model is used to perform a preliminary classification judgment of the vehicle scene, and a preliminary judgment result including a scene type identifier and a confidence value is generated; When the confidence value of the preliminary judgment result is lower than the set threshold, the dynamic reasoning client sends a supplementary data collection instruction to the corresponding multi-parameter server according to the scene type identifier.
[0037] As you can understand, the dynamic inference client serves as the core coordination unit, responsible for interaction with the onboard large model, inference logic control, and service scheduling. Multi-parameter servers are responsible for different types of data collection and preprocessing tasks, including user information management servers, vehicle status query servers, environment perception servers, and multimedia control servers.
[0038] The MCP protocol implementation uses a three-step architecture deployment: the first step is to build an MCP Client module as a bridge connecting the large model and in-vehicle services. It has the ability to communicate with the in-vehicle MCP Server through the standard MCP protocol, automatically discover available MCP Servers at startup, and pull and cache service description information. The second step is to develop multiple MCP Server modules, including a user information management server, a vehicle condition query server, an environmental perception server, and a multimedia control server. Each server provides an independent functional module and is registered with the MCP Client. The third step is to preset the available MCP Server information in the MCP Client configuration file to support hot loading and dynamic discovery mechanisms. The large model calls the MCP Client through Function Calling to obtain a service list, selects the appropriate service based on the service description, and executes the specific call.
[0039] Specifically, the scenario classification model utilizes a Transformer-based multi-classification network structure, capable of identifying various driving scenarios, including fatigue driving, distracted driving, environmental anomalies, and user needs. Based on the scenario type identifier and confidence distribution, the dynamic inference client intelligently selects the supplementary parameter types to be collected and sends supplementary data collection instructions to the corresponding multi-parameter server. The data structure of the supplementary data collection instruction includes information such as the request type, parameter specifications, collection accuracy, and time window.
[0040] like Figure 3 As shown in the figure, after a user submits a request, the big model requests a service description from the MCP Client through the Function Call mechanism. The MCP Client then obtains service description information from MCP Servers A and B in turn and returns a service list. The big model selects the appropriate service based on the service description (such as Server A's query service). The MCP Client forwards the request to the corresponding MCP Server for specific operations. The service execution results are returned to the big model and the user through the same path. The entire process embodies the MCP protocol's standardized interaction mechanism for service discovery, dynamic invocation, and result return, ensuring efficient coordination between the device-side big model and the in-vehicle service module.
[0041] The present invention uses a step-by-step conditional reasoning mechanism to dynamically request the required sensor data based on the preliminary reasoning results, avoiding the computational redundancy of full data collection and processing in traditional solutions. The MCP protocol adopts a multi-step reasoning network architecture, which can determine whether to supplement the second reasoning parameter based on the confidence assessment of the first reasoning parameter, realizing on-demand allocation of computing resources. While ensuring decision-making accuracy, it significantly reduces the computing load and energy consumption of the end-side device, and improves the overall operating efficiency of the system.
[0042] In one embodiment of the present invention, the process of generating a proactive service decision by combining user profiles and historical behavior data specifically includes: extracting physiological and behavioral characteristic parameters of the driver as first inference parameters using a driving state detection algorithm, wherein the driving state detection algorithm includes at least one of a fatigue detection algorithm based on facial key point detection, a steering wheel grip detection algorithm, and an attention distraction detection algorithm; When the first inference parameter indicates an abnormal driving state, the dynamic inference client requests the multi-parameter server to collect the corresponding verification parameter as the second inference parameter, and then performs feature splicing on the first and second inference parameters and inputs them into the multi-step inference network for scene confirmation; User personal preference features are extracted from the user portrait database, and historical behavior pattern features are extracted from the database. The confirmed scenario type is weighted and fused with the user's personal preference features and historical behavior pattern features to generate proactive service decisions. The proactive service decisions include service type, service parameters, and execution sequence. Service parameters define the specific configuration of the service, such as vibration intensity, temperature setting, search range, etc. The execution sequence specifies the execution order and time interval of the services to ensure the coordination of multiple service operations.
[0043] It can be understood that the output of the MCP dynamic reasoning controller is an active service decision containing complete decision information. The decision is encoded in JSON format and contains information such as service identification, parameter configuration, execution timing, priority, etc., providing clear execution guidance for the active service execution module.
[0044] like Figure 4 As shown in the figure, after the inference engine initiates the initial request, the MCP controller uses the first parameter set P1 for preliminary analysis. When the preliminary results require further confirmation, the system triggers the sensor cluster to collect supplementary data and obtain the second parameter set P2. The MCP controller integrates and analyzes these multi-step parameters to generate the final service instruction output. This entire process embodies the core concept of dynamic reasoning, which is to dynamically obtain the required parameters based on the intermediate results of the reasoning process, thus implementing an adaptive decision-making mechanism.
[0045] In one embodiment of the present invention, the multi-step reasoning network refers to the core network architecture for implementing step-by-step conditional reasoning in the MCP (Multi-step Conditional Protocol) dynamic reasoning controller. Specifically, The multi-step inference network utilizes a Transformer-based multi-classification network architecture. It receives as input a concatenated feature vector of a first inference parameter (such as the EAR value) and a second inference parameter (such as steering wheel grip data or vehicle lateral displacement). Using a multi-layer attention mechanism, the network gradually analyzes the correlations between these parameters and outputs an inference result that includes a scene confirmation result and a confidence score. The core advantage of this network lies in its ability to dynamically request and integrate new parameter information based on intermediate results of the inference process, enabling an adaptive inference decision-making process.
[0046] like Figure 5 As shown, specifically, taking fatigue driving detection as an example, the face key point detection algorithm is used to extract the driver's eye key points P1-P6, and the eye aspect ratio is calculated. The calculation formula is: in, represents the eye aspect ratio, Indicates the coordinates of the first key point of the eye, Indicates the coordinates of the second key point of the eye. Indicates the coordinates of the third key point of the eye, Indicates the coordinates of the fourth key point of the eye. Indicates the coordinates of the fifth key point of the eye. Indicates the coordinates of the sixth key point of the eye, represents the Euclidean distance.
[0047] The EAR value is approximately 0.25-0.35 in the normal eye-open state, and close to 0 in the eye-closed state. When the EAR parameter is continuously lower than the threshold of 0.2 for more than 2 seconds, it is determined that there may be a fatigue state. At this time, the dynamic reasoning client requests the multi-parameter server to collect steering wheel grip data, vehicle lateral displacement, vehicle speed change rate, etc. as the second reasoning parameter, connect the first reasoning parameter (EAR value) and the second reasoning parameter by vector, and input the spliced feature vector into the multi-step reasoning network for scene confirmation. The network output includes the scene confirmation result and confidence score.
[0048] In one embodiment of the present invention, the active service execution module includes: A decision parsing unit is used to parse the service type, service parameters, and execution timing information in the active service decision, match the corresponding vehicle device control protocol according to the service type, and generate device control parameters; an instruction generating unit, configured to generate control instructions according to the device control parameters, wherein the control instructions include seat vibration control instructions, air conditioning temperature adjustment control instructions, navigation rest area search control instructions, and audio volume adjustment control instructions; The execution monitoring unit is used to monitor the execution status of the control instructions of the vehicle-mounted equipment through the equipment execution interface, record the control instruction issuance time, execution completion time and execution result status, and generate execution feedback information including execution success rate and response time.
[0049] It can be understood that the decision parsing unit is responsible for comprehensively parsing and processing the active service decisions from the MCP dynamic inference controller, receiving three key information dimensions: service type, service parameters and execution timing. The service type is identified by a predefined enumeration value, including standard service types such as SEAT_VIBRATION (seat vibration), AC_TEMPERATURE (air conditioning temperature adjustment), NAVIGATION_SEARCH (navigation rest area search), and AUDIO_VOLUME (audio volume adjustment). The service parameter parsing adopts a key-value pair mapping mechanism to convert the abstract service configuration into specific device control parameters. The execution timing parsing uses a time series analysis algorithm to convert the timing information in the decision into an accurate execution plan.
[0050] Specifically, for the seat vibration service, parameters such as vibration intensity level, vibration mode and duration are analyzed; for the air conditioning temperature adjustment service, parameters such as target temperature value, wind speed level and air outlet mode are analyzed.
[0051] In one embodiment of the present invention, the generation process of the seat vibration control instruction includes three steps: device address configuration, vibration parameter setting and execution triggering. The device address adopts the CAN bus standard address format, and the vibration parameters control the motor speed and vibration intensity through PWM signals. The air conditioning temperature adjustment control instruction is implemented through the HVAC control protocol, including functions such as temperature setting, wind speed adjustment and air outlet control. The navigation rest area search control instruction is implemented through the vehicle navigation system API interface, including functions such as search range setting, POI type screening and route planning. The audio volume adjustment control instruction is implemented through the audio system control interface, supporting the adjustment of the main volume, partition volume and sound effect mode.
[0052] In one embodiment of the present invention, the execution monitoring unit utilizes an asynchronous monitoring mechanism, assigning a separate monitoring thread to each control instruction to prevent execution anomalies on a single device from impacting overall system performance. The monitoring thread obtains execution status information through the device feedback interface, including instruction receipt confirmation, execution progress updates, and completion status notifications. Execution status monitoring utilizes a finite state machine model, defining standard states such as PENDING, EXECUTING, COMPLETED, FAILED, and TIMEOUT. State transitions are triggered by time events, device feedback events, and exception events.
[0053] In one embodiment of the present invention, the active service execution module further includes a user interaction and feedback processing mechanism: Before executing a proactive service operation, a proactive service semantic prompt message is generated through speech synthesis technology to explain the detected driving scenario and the service measures to be taken to the user. At the same time, a proactive service confirmation card with confirmation and cancellation options is displayed on the in-vehicle display interface; Receive feedback information input by the user through voice commands or touch operations, execute the vehicle-mounted device control instructions when confirmation feedback is received, terminate the current service execution process when cancellation feedback is received, and update the user's feedback information as personalized preference data to the database.
[0054] As you can understand, user interaction is enabled by the in-vehicle speech synthesis system, which uses natural language generation technology to convert detected driving scenarios and upcoming service actions into easy-to-understand voice prompts. This speech synthesis utilizes a parameterized text-to-speech model, supporting dynamic adjustment of intonation, speaking rate, and timbre to suit different user preferences and driving environments. Card display adheres to in-vehicle HMI design specifications to ensure readability and operational safety while driving. Card display duration utilizes a dynamic adjustment mechanism, automatically adjusting the display duration based on the complexity of the driving environment.
[0055] The feedback processing mechanism adopts a state machine design and executes corresponding processing procedures according to different feedback types of users. When confirmation feedback is received, the preset vehicle equipment control instructions are immediately executed; when cancellation feedback is received, the current service execution process is terminated and the user's rejection preference is recorded; when later feedback is received, the service request is added to the delay queue and the user is prompted again after the preset time.
[0056] The updated personalized preference data is stored in the database to provide personalized reference for subsequent proactive service decisions. The database adopts NoSQL architecture to support flexible expansion and efficient query of user portraits. The storage format of personalized data includes fields such as user ID, scenario type, service type, preference weight, update time, etc. to ensure data integrity and traceability.
[0057] The present invention also provides a dynamic reasoning method for vehicle proactive services based on a large, on-device multimodal model, which is applied to the vehicle proactive service system based on the large, on-device multimodal model described above, and includes the following steps: S1. Collecting multimodal perception data of the vehicle; wherein the multimodal perception data includes image data, in-vehicle audio signal data, and vehicle driving state parameter data; S2. Feature extraction is performed on multimodal perception data to obtain features of different modalities. The multimodal large model is compressed using knowledge distillation and quantization and deployed on the vehicle side. S3. Convert the feature data of different modalities into query, key, and value triplets, calculate the cross-modal similarity weight through a bidirectional attention mechanism, and fuse them based on the query, key, and value triplets and the cross-modal similarity weight to obtain the fused features; S4. Perform a preliminary judgment on the vehicle scenario based on the fused features. When the preliminary judgment requires further confirmation, generate a supplementary data collection instruction to control the multimodal data collection module to collect specific parameter data, combine the supplementary data with the fused features to perform multi-step reasoning, determine the current vehicle scenario type, and generate proactive service decisions based on user profiles and historical behavior data; S5. Receive active service decisions, generate vehicle-mounted equipment control instructions, and execute active service operations.
[0058] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A vehicle active service system based on a large, multimodal model on the device side, characterized by: include: Multimodal data acquisition module, used to collect multimodal perception data of the vehicle; The multimodal perception data includes image data, in-vehicle audio signal data and vehicle driving status parameter data; A multimodal large model processing unit is used to extract features from multimodal perception data to obtain features of different modalities; The multimodal large model is compressed using knowledge distillation and quantization and deployed on the vehicle side. The cross-modal attention fusion module is used to convert feature data of different modalities into query, key, and value triplets, calculate cross-modal similarity weights through a bidirectional attention mechanism, and fuse the query, key, value triplets and cross-modal similarity weights to obtain fused features; The MCP dynamic inference controller is used to make a preliminary judgment on the vehicle scene based on the fusion features. When the preliminary judgment result needs further confirmation, it generates supplementary data collection instructions to control the multimodal data collection module to collect specific parameter data. The supplementary data is combined with the fusion features to perform multi-step reasoning to determine the current vehicle scene type, and the user profile and historical behavior data are combined to generate proactive service decisions; The active service execution module is used to receive active service decisions, generate vehicle-mounted equipment control instructions and execute active service operations.
2. The vehicle active service system based on a large, multimodal model on the device side according to claim 1, characterized in that: The image data includes RGB image data and infrared image data, and the multimodal data acquisition module further includes: A brightness scoring unit is used to convert the RGB image data into a color space and calculate a brightness score value. When the brightness score value is lower than a preset threshold, the visible light camera is switched to the infrared camera for image acquisition. The infrared vision enhancement unit is used to collect infrared thermal radiation image data using an infrared camera, and perform thermal radiation feature extraction and dynamic exposure compensation on the infrared thermal radiation image to obtain infrared image data.
3. The vehicle active service system based on a large, multi-modal model on the device side according to claim 1, characterized in that: The multimodal large model processing unit adopts a teacher-student network architecture to perform knowledge distillation, specifically including: In the output layer distillation, the output probability distributions of the teacher model and the student model are smoothed, and the difference between the smoothed probability distributions of the teacher model and the student model is calculated using the KL divergence loss function. The KL divergence loss is weightedly combined with the cross entropy loss of the student model on the true label to form a joint loss function for optimization training; In the intermediate feature distillation, the intermediate feature map sizes of the teacher model and the student model are adjusted through a 1×1 convolution adaptation layer, and the mean square error loss function is used to align the intermediate feature representations of the teacher model and the student model; The 8-bit integer quantization technology is used to quantize the weight parameters and activation values of the student model. The model parameter storage accuracy is reduced by channel-by-channel quantization, and the large multimodal model is compressed to less than 500MB.
4. The vehicle active service system based on a large, multi-modal model on the device side according to claim 1, characterized in that: The conversion of feature data of different modalities into query, key, and value triples specifically includes: A visual transformer encoder is used to extract visual feature vectors from image data, an audio embedding feature vector is extracted from audio signal data, and a state feature vector is generated by feature encoding the vehicle driving state parameter data; The eigenvectors of different modes are projected and transformed into query vector, key vector and value vector respectively through linear transformation matrix; The feature dimensions of the query vector, key vector, and value vector are uniformly set to the same feature space dimension to form a triplet representation of uniform dimension.
5. The vehicle active service system based on a large, multi-modal model on the device side according to claim 4, characterized in that: The calculation method of the bidirectional attention mechanism specifically includes: The forward attention weight matrix is obtained by performing a dot product operation on the query vector and the key vector and applying softmax normalization based on the scaling factor. The inverse attention weight matrix is obtained by performing a transposed dot product operation on the key vector and the query vector and applying softmax normalization based on the scaling factor. The forward attention weight matrix and the value vector are weightedly summed to obtain the forward attention output, and the reverse attention weight matrix and the query vector are weighted summed to obtain the reverse attention output. The weighted average calculation is performed based on the forward attention output and the reverse attention output, and after residual connection and layer normalization processing, the fusion feature is output.
6. The vehicle active service system based on a large, multi-modal model on the device side according to claim 1, characterized in that: The MCP dynamic inference controller includes a dynamic inference client and multiple multi-parameter servers. The processing steps of the MCP dynamic inference controller specifically include: The dynamic reasoning client communicates with the multi-parameter server through a multi-step conditional reasoning protocol. When it starts, it automatically identifies all available multi-parameter servers and obtains the data collection capability description information provided by each server. Based on the fusion features, a preset scene classification model is used to perform a preliminary classification judgment of the vehicle scene, and a preliminary judgment result including a scene type identifier and a confidence value is generated; When the confidence value of the preliminary judgment result is lower than the set threshold, the dynamic reasoning client sends a supplementary data collection instruction to the corresponding multi-parameter server according to the scene type identifier.
7. The vehicle active service system based on a large, multi-modal model on the device side according to claim 6, characterized in that: The generation of proactive service decisions by combining user profiles and historical behavior data specifically includes: extracting physiological and behavioral characteristic parameters of the driver as first inference parameters using a driving state detection algorithm, wherein the driving state detection algorithm includes at least one of a fatigue detection algorithm based on facial key point detection, a steering wheel grip detection algorithm, and an attention distraction detection algorithm; When the first inference parameter indicates an abnormal driving state, the dynamic inference client requests the multi-parameter server to collect the corresponding verification parameter as the second inference parameter, and then performs feature splicing on the first and second inference parameters and inputs them into the multi-step inference network for scene confirmation; User personal preference features are extracted from the user portrait database, and historical behavior pattern features are extracted from the database. The confirmed scenario type is weighted and fused with the user personal preference features and historical behavior pattern features to generate proactive service decisions, where proactive service decisions include service type, service parameters, and execution timing.
8. The vehicle active service system based on a large, multi-modal model on the device side according to claim 1, characterized in that: The active service execution module includes: A decision parsing unit is used to parse the service type, service parameters, and execution timing information in the active service decision, match the corresponding vehicle device control protocol according to the service type, and generate device control parameters; an instruction generating unit, configured to generate control instructions according to the device control parameters, wherein the control instructions include seat vibration control instructions, air conditioning temperature adjustment control instructions, navigation rest area search control instructions, and audio volume adjustment control instructions; The execution monitoring unit is used to monitor the execution status of the control instructions of the vehicle-mounted equipment through the equipment execution interface, record the control instruction issuance time, execution completion time and execution result status, and generate execution feedback information including execution success rate and response time.
9. The vehicle active service system based on a large, multi-modal model on the device side according to claim 8, characterized in that: The active service execution module also includes a user interaction and feedback processing mechanism: Before executing a proactive service operation, a proactive service semantic prompt message is generated through speech synthesis technology to explain the detected driving scenario and the service measures to be taken to the user. At the same time, a proactive service confirmation card with confirmation and cancellation options is displayed on the in-vehicle display interface; Receive feedback information input by the user through voice commands or touch operations, execute the vehicle-mounted device control instructions when confirmation feedback is received, terminate the current service execution process when cancellation feedback is received, and update the user's feedback information as personalized preference data to the database.
10. A dynamic reasoning method for vehicle proactive services based on a large, multimodal model on the device side, characterized by: The vehicle active service system based on the terminal-side multimodal large model as described in any one of claims 1 to 9 comprises the following steps: S1. Collecting multimodal perception data of the vehicle; wherein the multimodal perception data includes image data, in-vehicle audio signal data, and vehicle driving state parameter data; S2. Feature extraction is performed on multimodal perception data to obtain features of different modalities. The multimodal large model is compressed using knowledge distillation and quantization and deployed on the vehicle side. S3. Convert the feature data of different modalities into query, key, and value triplets, calculate the cross-modal similarity weight through a bidirectional attention mechanism, and fuse them based on the query, key, and value triplets and the cross-modal similarity weight to obtain the fused features; S4. Perform a preliminary judgment on the vehicle scenario based on the fused features. When the preliminary judgment requires further confirmation, generate a supplementary data collection instruction to control the multimodal data collection module to collect specific parameter data, combine the supplementary data with the fused features to perform multi-step reasoning, determine the current vehicle scenario type, and generate proactive service decisions based on user profiles and historical behavior data; S5. Receive active service decisions, generate vehicle-mounted equipment control instructions, and execute active service operations.
Citation Information
Patent Citations
Vehicle-mounted intelligent scene system and method
CN114943031A
Pavement type identification method and device based on bidirectional time sequence feature adaptive fusion
CN119919902A
Multi-modal data fusion method and system based on large model agent
CN120354357A
Intelligent automobile control system and method based on artificial intelligence
CN120482081A
Image recognition and analysis system based on AI
CN120495674A
Cited By
AI multi-mode voice interaction method based on vehicle-mounted intelligent terminal and electronic equipment
CN121171226A
Method and device for determining interactive interface, terminal equipment and storage medium
CN121523575A
Space-time sensing subway passenger flow prediction method based on large model and multi-source information fusion
CN121543042A
Space-time perception subway passenger flow prediction method based on large model and multi-source information fusion
CN121543042B
End-side multi-mode fusion analysis and active service interaction method and system
CN121580285A