Vehicle control methods, devices, vehicle control application systems and their testing methods

CN121019203BActive Publication Date: 2026-09-01CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511298835.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2026-09-01
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

[0004]然而,由于多源数据往往参数量庞大,在模型处理过程中对计算资源要求极高

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121019203B_ABST
    Figure CN121019203B_ABST
Patent Text Reader

Abstract

This invention relates to a vehicle control method, device, vehicle control application system, and testing method thereof. The method includes: acquiring multi-source data information, including user voice data and / or vehicle environment data, user profile data, and the current load level of the vehicle; constructing multimodal prompt words based on a pre-set intelligent model, according to the user voice data and / or vehicle environment data, and the load level, where the load level indicates the length of the multimodal prompt words; obtaining the vehicle control intent based on the multimodal prompt words; and generating vehicle control information based on the vehicle control intent and user profile data. This method can effectively improve the accuracy of vehicle control intent recognition and decision-making, provide users with more personalized vehicle control suggestions, and effectively solve the problem of delayed response due to limited onboard computing power, thus bringing a better user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle technology, and specifically to a vehicle control method, device, vehicle control application system, and testing method thereof. Background Technology

[0002] With the continuous development of automotive intelligence, intelligent cockpits have gradually become an important part of modern automobiles. As one of the core interaction methods of intelligent cockpits, in-vehicle voice systems have evolved from recognizing simple commands to becoming multifunctional intelligent assistants.

[0003] Current in-vehicle voice systems primarily rely on preset command sets and limited semantic understanding capabilities. While users can achieve basic vehicle control through voice, these systems only understand semantics based on the user's voice, which can easily lead to inaccurate intent recognition. Some related technologies use deep learning models to learn from multi-source data (such as voice data and user data) and attempt to fuse multi-source data to improve the accuracy of intent recognition.

[0004] However, multi-source data often involves a large number of parameters, placing extremely high demands on computing resources during model processing. In contrast, the strict constraints of onboard computing power and energy consumption make it difficult to support efficient real-time inference of the model, resulting in significant system response latency and increased power consumption, severely limiting the performance of the voice system in a real-world vehicle environment. Summary of the Invention

[0005] The present invention provides a vehicle control method, apparatus, vehicle control application system, and testing method thereof to at least solve one of the aforementioned technical problems.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] According to a first aspect of the present invention, a vehicle control method is provided, comprising:

[0008] Acquire multi-source data information, including user voice data and / or vehicle environment data, user profile data, and the current load level of the vehicle;

[0009] Based on a pre-set intelligent model, multimodal prompt words are constructed according to the user voice data and / or the vehicle environment data and the load level. The load level is used to indicate the length information of the multimodal prompt words. The vehicle control intention is obtained according to the multimodal prompt words, and vehicle control information is generated according to the vehicle control intention and the user profile data.

[0010] The vehicle control intent includes user vehicle control intent and / or scenario vehicle control intent. The intelligent model is a model trained based on historical multi-source data information and used to generate vehicle control information. The vehicle control information is used to instruct control operations on at least one actuator of the vehicle.

[0011] In one implementation, the step of constructing multimodal prompt words based on the user voice data and / or the vehicle environment data, and the load level includes:

[0012] Semantic text features are extracted from the user voice data, and environmental perception features and vehicle status features are extracted from the vehicle environment data.

[0013] Based on the semantic text features, environmental perception features, and vehicle state features, an initial multimodal prompt word is constructed;

[0014] Based on the current load level of the vehicle, the initial multimodal prompt words are processed to obtain multimodal prompt words.

[0015] In one implementation, the step of processing the initial multimodal cue word according to the current load level of the vehicle to obtain a multimodal cue word includes:

[0016] Based on the length information indicated by the load level, the number of semantic text features, environmental perception features, and vehicle status features in the initial multimodal prompt is adjusted to obtain the adjusted feature information;

[0017] The multimodal prompt words are obtained based on the adjusted feature information.

[0018] In one embodiment, the vehicle control intent carries an intent classification for the vehicle control intent, the intent classification being used to indicate the strength of the intent of the vehicle control intent.

[0019] In one implementation, generating vehicle control information based on the vehicle control intent and the user profile data includes:

[0020] User biometrics are extracted from the user profile data, and the user biometrics include at least one of age features, action features, and location features;

[0021] Based on the vehicle control intent, the user's biometrics, and the predefined vehicle control rule knowledge graph in the intelligent model, a structured instruction sequence for vehicle control is constructed, the structured instruction sequence including control instructions for multiple actuators;

[0022] The vehicle control information is generated based on the structured instruction sequence.

[0023] In one implementation, it further includes:

[0024] Determine whether the intelligent model generates the vehicle control information within a predefined execution time window;

[0025] If the vehicle control information is not generated within the preset time period, a predefined vehicle control baseline strategy is triggered, and the resources of the intelligent model are released; wherein, the vehicle control baseline strategy is used to query and execute basic vehicle control operations corresponding to the user voice data and / or vehicle environment data.

[0026] In one embodiment, the environmental data includes environmental perception data and vehicle status data, wherein the environmental perception data includes weather data and road data; the acquisition of multi-source data information includes:

[0027] User voice data is collected through the vehicle's voice acquisition device;

[0028] Environmental perception data is acquired through the cloud or the vehicle's environmental perception sensors, and vehicle status data is acquired through the vehicle's domain controller.

[0029] User profile data is collected using the vehicle's image acquisition device.

[0030] In one implementation, after generating vehicle control information based on the vehicle control intent and the user profile data, the method further includes:

[0031] The information generation link corresponding to the vehicle control information is displayed. The information generation link includes at least the input information corresponding to the multi-source data information and the output information corresponding to the vehicle control information.

[0032] In response to the confirmation information for the information generation link, the vehicle control operation is performed based on the vehicle control information.

[0033] In one implementation, it further includes:

[0034] According to predefined instruction recognition rules, valid and invalid control instructions in the vehicle control information are identified, and invalid control instructions are filtered out.

[0035] The step of performing vehicle control operations based on the vehicle control information includes:

[0036] According to the predefined instruction priority rules, the effective control instructions are time-arranged to obtain the timing arrangement result;

[0037] Based on the timing arrangement result, the vehicle control operations are executed sequentially according to the effective control commands.

[0038] According to a second aspect of the present invention, a vehicle control device is provided, comprising:

[0039] The data acquisition module is used to acquire multi-source data information, including user voice data and / or vehicle environment data, user profile data, and the current load level of the vehicle.

[0040] The model processing module is used to construct multimodal prompt words based on a pre-set intelligent model, according to the user voice data and / or the vehicle environment data, and the load level, wherein the load level is used to indicate the length information of the multimodal prompt words, and to obtain vehicle control intentions based on the multimodal prompt words, and to generate vehicle control information based on the vehicle control intentions and the user profile data; wherein the vehicle control intentions include user vehicle control intentions and / or scenario vehicle control intentions, the intelligent model is a model trained based on historical multi-source data information and used to generate vehicle control information, and the vehicle control information is used to instruct control operations on at least one actuator of the vehicle.

[0041] According to a third aspect of the present invention, a vehicle control application system is provided, comprising:

[0042] At least one processor; and

[0043] A memory communicatively connected to the at least one processor; wherein,

[0044] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the vehicle control method provided in any of the first aspects above.

[0045] According to a fourth aspect of the present invention, a testing method for a vehicle control application system is provided, comprising:

[0046] A test set for testing the vehicle control application system is constructed based on multi-source data information with different state values.

[0047] The test set is input into the vehicle control application system through a test application to perform segmented testing on the vehicle control application system, and obtain segmented test information for each segmented test; wherein, the segmented test includes at least a first segmented test on the construction of multimodal prompt words, and a second segmented test on vehicle control information and its execution results;

[0048] Based on the segmented test information, the test results of the vehicle control application system are obtained.

[0049] In one implementation, inputting the test set into the vehicle control application system via a test application includes:

[0050] The test set is converted into a JSON file format, and the converted test set is input into the vehicle control application system through the test application.

[0051] In one implementation, the segmented testing of the vehicle control application system includes:

[0052] Obtain log data from the vehicle control application system regarding the test set. The log data includes a first log for constructing the multimodal prompt words and a second log for executing the vehicle control information and obtaining the execution result.

[0053] The first segmented test is executed based on the first log and the predefined prompt word evaluation rules; and the second segmented test is executed based on the second log and the predefined execution result evaluation rules.

[0054] According to a fifth aspect of the present invention, a testing apparatus for a vehicle control application system is provided, comprising:

[0055] The test set construction module is used to construct a test set for testing the vehicle control application system based on multi-source data information with different state values.

[0056] The segmented testing module is used to input the test set into the vehicle control application system through the test application, perform segmented testing on the vehicle control application system, and obtain segmented test information for each segmented test; wherein, the segmented test includes at least a first segmented test on the construction of multimodal prompt words and a second segmented test on vehicle control information and its execution results;

[0057] The test result acquisition module is used to acquire the test results of the vehicle control application system based on the segmented test information.

[0058] According to a sixth aspect of the present invention, a vehicle is provided, the vehicle including the vehicle control application system provided in the third aspect above.

[0059] The technical solution provided by this invention acquires multi-source data information about the vehicle and then uses an intelligent model to identify vehicle control intentions. These intentions include both user and environmental intents. Compared to methods that only use voice data to identify vehicle control intentions, this approach integrates the vehicle environment, resulting in more accurate intention recognition. Furthermore, it combines vehicle control intentions and user profile data to generate vehicle control information, taking into account individual user differences during the information generation process. Moreover, considering the limited computing resources on the vehicle side, this embodiment uses an intelligent model that combines user voice, vehicle environment, and vehicle load level to construct multimodal prompts. After obtaining the vehicle control intention, it integrates user profile data to generate vehicle control information. Compared to merging all data within the model, this significantly reduces the amount of data processing during model training and inference, decreases the consumption of computing power and memory, makes the intelligent model more adaptable to the vehicle hardware environment, improves operating efficiency and response speed, and effectively reduces the processing latency from user-input voice data to vehicle control information, providing a better user experience. In addition, after generating vehicle control information, the visualization of the process allows users to intuitively understand the generation process of vehicle control information, thereby enhancing user interactivity. Attached Figure Description

[0060] Figure 1 A possible system architecture diagram provided for an embodiment of the present invention;

[0061] Figure 2a This is one of the flowcharts illustrating the vehicle control method provided in an embodiment of the present invention;

[0062] Figure 2b This is a second schematic flowchart of the vehicle control method provided in an embodiment of the present invention;

[0063] Figure 3 A schematic diagram of the HMI interface provided in an embodiment of the present invention;

[0064] Figure 4a This is one of the flowcharts illustrating the load classification method provided in an embodiment of the present invention;

[0065] Figure 4b The second schematic flowchart of the load classification method provided in the embodiment of the present invention;

[0066] Figure 5 This is a flowchart illustrating the prompt construction method provided in an embodiment of the present invention.

[0067] Figure 6 A flowchart illustrating an exemplary vehicle control method provided in an embodiment of the present invention;

[0068] Figure 7 This is an example diagram of control commands in an embodiment of the present invention;

[0069] Figure 8a This is one of the structural schematic diagrams of the vehicle control device provided in the embodiments of the present invention;

[0070] Figure 8b This is a second schematic diagram of the structure of the vehicle control device provided in an embodiment of the present invention;

[0071] Figure 9 This is a schematic diagram of the structure of a vehicle control application system provided in an embodiment of the present invention;

[0072] Figure 10 A flowchart illustrating a testing method for a vehicle control application system provided in an embodiment of the present invention;

[0073] Figure 11 This is a schematic diagram of the structure of a test device for a vehicle control application system provided in an embodiment of the present invention;

[0074] Figure 12 This is a structural schematic diagram of a vehicle provided in an embodiment of the present invention. Detailed Implementation

[0075] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0076] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0077] Current in-vehicle voice systems primarily rely on preset instruction sets and limited semantic understanding capabilities for speech recognition, resulting in low intent recognition accuracy and poor interactivity. By integrating a smart cockpit with a Large Language Model (LLM), combining the natural language processing and deep semantic understanding capabilities of the LLM model, higher-precision intent recognition and more natural voice interaction can be achieved, enabling vehicle control such as air conditioning and window control. However, these technologies mainly analyze user voice data, lacking adaptability to the vehicle environment and individual flexibility, making it difficult to fully understand the user's true intent and provide personalized vehicle control suggestions. Furthermore, the large number of parameters in multi-source data often places high demands on computing resources. Strict constraints on in-vehicle computing power and energy consumption make it difficult to support efficient real-time inference of the model, leading to significant system response latency and increased power consumption, severely limiting the performance of the voice system in a real-world vehicle environment.

[0078] In view of this, embodiments of the present invention provide a vehicle control method, device, vehicle control application system, and testing method thereof. This method acquires multi-source data information, including user voice data and / or vehicle environment data, and user profile data. Based on a pre-set intelligent model, it processes the user voice data and / or vehicle environment data to obtain vehicle control intent, and generates vehicle control information based on the vehicle control intent and user profile data. The vehicle control intent includes user vehicle control intent and / or scenario vehicle control intent. The intelligent model is trained based on historical multi-source data information and used to generate the vehicle control information. The vehicle control information is used to instruct control operations on at least one actuator of the vehicle. In this process, by acquiring multi-source data information about the vehicle and then using an intelligent model to identify the vehicle control intent, which includes both user and environmental vehicle control intent, compared to methods that only use voice data to identify vehicle control intent, this method incorporates the vehicle environment, making the intent recognition result more accurate. Furthermore, by combining vehicle control intent and user profile data to generate vehicle control information, individual user differences are considered during the vehicle control information generation process. Furthermore, this embodiment addresses the limited computing resources on the vehicle side. The intelligent model combines user voice, vehicle environment, and vehicle load level to construct multimodal prompts. After obtaining the vehicle control intent, it integrates user profile data to generate vehicle control information. Compared to merging all data within the model, this significantly reduces the amount of data processing during model training and inference, minimizing the consumption of computing power and memory. This makes the intelligent model more adaptable to the vehicle hardware environment, improving operational efficiency and response speed, and effectively reducing the processing latency from user-input voice data to vehicle control information, resulting in a better user experience. In addition, after generating vehicle control information, a visual link display method allows users to intuitively understand the generation process of vehicle control information, enhancing user interactivity.

[0079] For example, Figure 1 is a possible system architecture provided by the embodiment of the present invention, as shown in Figure 1 , it may include an adaptive vehicle control application (app) 110, a voice app 120, a vehicle control app 130, an Operating System Layer (OS) 140 and a cloud 150. Wherein,

[0080] the adaptive vehicle control app 110 can serve as the core functional layer of the system architecture, and the method of the embodiment of the present invention can be applied in the adaptive vehicle control app 110. The adaptive vehicle control app 110 may include a semantic transceiver module 111, a load grading module 112, a semantic discrimination module 113, a Prompt construction module 114, an instruction-to-semantic module 115, a streaming parsing module 116, a model inference module 117, a perception module 118, and so on. Wherein, the semantic transceiving module 111 is connected to the voice app 120, and is configured to receive user voice data, convert the user voice data into semantic text, and may also feed back feedback information about vehicle control information to the user; the load grading module 112 is configured to identify the current load level of the vehicle; the semantic discrimination module 113 is configured to parse the meaning corresponding to the semantic text; the Prompt construction module 114 is configured to construct multimodal prompts for model inference; the instruction-to-semantic module 115 is configured to convert a control instruction into semantic text (for example, a natural language understanding (NLU) parsing engine can be used to convert the instruction text into the semantics corresponding to a vehicle control application programming interface (API), so as to implement linkage control of multiple devices such as air conditioners, seats, and lights); the streaming parsing module 116 is configured to process the inference result of the model, that is, vehicle control information, which can be a structured instruction sequence including multiple control instructions, and parse the structured instruction sequence into individual control instructions by means of streaming parsing; the model inference module 117 is configured to perform the inference process of vehicle control intentions in combination with prompts, and generate vehicle control information by combining the vehicle control intentions and user portrait data; the perception module 118 is configured to obtain vehicle internal and external environment data (which can include user portrait data), can receive weather data from the weather service of the cloud 150 or other vehicle environment data, and perform perception and conversion processing on the received data. Specifically, the vehicle internal and external environment data can be obtained from the OS 140. Specifically, the OS 140 can include a Car Service, a GPS-Service, an Integrated Monitoring System (IMS) / Occupant Monitoring System (OMS), and the perception module 118 can obtain vehicle environment data (vehicle state data, road data), user portrait data, etc. from the OS 140.

[0081] Voice app 120, serving as the interactive entry layer of this system architecture, transmits user voice data (or voice commands) to adaptive vehicle control app 110 via inter-process communication (IPC). Alternatively, it receives vehicle control information from adaptive vehicle control app 110 and transmits it to vehicle control app 130 (in some embodiments, vehicle control information from adaptive vehicle control app 110 can also be directly transmitted to vehicle control app 130). Vehicle control app 130 can act as an intermediary, connecting downwards to OS layer services to transmit vehicle control information to CarService, enabling vehicle control functions such as air conditioning, speed, gear shift, seat heating, seat ventilation, windows, etc. Correspondingly, the OS layer can transmit this data to adaptive vehicle control app 110 as in-vehicle and external environmental perception data, realizing a complete closed loop of "perception (including voice input) - decision-making - execution - feedback" for voice and / or environmental vehicle control. This allows the vehicle to dynamically adjust its control logic based on the real-time environment, improving interactive intelligence and vehicle control adaptability.

[0082] based on Figure 1 The following example illustrates a possible system architecture, applied to the vehicle control method described above. Figure 1 The adaptive vehicle control app 110 will be used as an example for detailed explanation. The technical solution of the present invention and how it solves the aforementioned technical problems will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that these specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0083] First, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0084] Figure 2a and Figure 2b This is a flowchart illustrating a vehicle control method proposed in an embodiment of the present invention, as shown below. Figure 2a As shown, the method includes steps S201-S204.

[0085] Step S201: Obtain multi-source data information, which includes user voice data and / or vehicle environment data, user profile data, and the vehicle's current load level. In this embodiment, multi-source data information refers to vehicle-related multi-source data information. Optionally, environmental data may include environmental perception data and vehicle status data. Environmental perception data includes weather data and road data, while vehicle status data may include vehicle speed status, air conditioning status, window status, etc. User profile data may be basic user data, such as one or more of the following: user age, user location, user posture, and user habits.

[0086] In one optional implementation, the acquisition of the aforementioned multi-source data information can also be achieved by: acquiring user voice data through the vehicle's voice acquisition device; acquiring environmental perception data through the cloud and / or the vehicle's environmental perception sensors; and acquiring vehicle status data through the vehicle's domain controller. Finally, user profile data can be acquired through the vehicle's image acquisition device.

[0087] For example, the voice acquisition device can be an in-vehicle microphone. The environmental perception module can be a sensor device that perceives the environment, such as a temperature sensor (e.g., an in-vehicle temperature sensor to obtain the in-vehicle temperature), a camera (e.g., an in-vehicle forward-facing 800W streaming media camera to capture images outside the cabin to obtain road data and other vehicle environmental perception data), or it can obtain weather data (such as outside temperature data) from cloud-based weather services and / or road data from cloud-based navigation services. The image acquisition device for collecting user profile data can be an in-vehicle monitoring camera.

[0088] Furthermore, vehicle status data and user profile data can also be obtained through OS140 running on the vehicle domain controller. For example, vehicle status information such as air conditioning, vehicle speed, gear position, seat heating, seat ventilation, and windows can be obtained from CarService. User profile data or other in-vehicle and out-of-vehicle monitoring data can also be obtained through IMS / OMS to obtain corresponding vehicle environmental data. For example, when not communicating with the cloud, weather data and road data can also be obtained through IMS / OMS, and so on.

[0089] In some embodiments, to improve the real-time processing efficiency of the model, the process of acquiring multi-source data information can be designed, which may include outlier filtering, data transformation, and data updating (acquisition). For example, a three-layer data governance mechanism can be adopted: first, a data link monitoring framework is established, and the most recent valid data is automatically injected as a default value for scenarios such as abnormal weather services, and outliers such as sudden speed jumps are filtered through threshold verification. Secondly, semantic mapping rules are designed to transform the raw signals into semantic descriptions that the interactive scenario can understand (e.g., mapping real-time vehicle speed to graded states such as "medium-high speed driving," "(0,20] low speed," "(20,40] medium-low speed," "(40,60] medium speed," "(60,80] medium-high speed," "greater than 80, high speed," etc.). Graded update strategies can be executed based on data characteristics. For example, high-frequency data such as vehicle speed / GPS can be acquired with appropriate delays and filtered using change thresholds; medium-frequency state changes such as gear shifting can be updated using event-driven methods; and low-frequency data such as weather information can be updated through a combination of timed polling and change detection. In this way, dynamic perception of multi-source data information can be achieved, providing data support for subsequent model processing.

[0090] Step S202: Based on a pre-set intelligent model, construct multimodal prompt words according to user voice data and / or vehicle environment data, and load level. The load level is used to indicate the length information of the multimodal prompt words. Obtain vehicle control intent based on the multimodal prompt words, and generate vehicle control information based on vehicle control intent and user profile data. The vehicle control intent includes user vehicle control intent and / or scenario vehicle control intent. The intelligent model is a model trained based on historical multi-source data information and used to generate vehicle control information. The vehicle control information is used to instruct control operations on at least one actuator of the vehicle.

[0091] In related technologies, some intelligent models exist that infer user intent by fusing multiple types of data. This typically involves inputting all types of data into the model for training and directly combining all data during inference. This embodiment addresses the limited computing resources on the vehicle side by combining the generation of multimodal prompts based on load level (different load level indicators construct multimodal prompts matching the corresponding load level; for high load levels, shorter multimodal prompts are generated to reduce computational overhead, while lower load levels can retain longer multimodal prompts to provide richer contextual information), and the phased processing of multi-source data (first, multimodal prompts are constructed using user voice data and / or vehicle environment data, along with load level, to generate vehicle control intent; then, user profile data is fused to generate vehicle control information). Compared to the method of merging all data, this significantly reduces the amount of data processing during model training and inference, reduces the consumption of computing power and memory, makes the intelligent model more adaptable to the vehicle hardware environment, improves operating efficiency and response speed, and effectively reduces the processing latency from user-input voice data to vehicle control information.

[0092] For example, by inputting multi-source data into a pre-built intelligent model, the intelligent model constructs multimodal prompt words based on real-time user voice data, vehicle environment data, and dynamic load levels. During the construction of the multimodal prompt words, the load level serves as a key control parameter, indicating the length of the multimodal prompt words. After constructing the multimodal prompt words, the intelligent model performs a first-stage inference based on them, outputting the user's vehicle control intent (e.g., adjusting the air conditioning temperature) and / or the scenario's vehicle control intent (e.g., automatically closing windows in rainy weather). Subsequently, combined with user profile data, the model performs a second-stage processing to generate the final vehicle control information (e.g., setting the air conditioning to 24°C and closing the windows).

[0093] As a further example, the intelligent model can employ a lightweight, two-stage cascaded neural network architecture to decouple the core processing layer, thereby enabling step-by-step and efficient processing of information from multiple sources. For instance, in the first-stage inference process, multimodal fusion encoding in the model can simultaneously extract and encode features from the input user voice data and vehicle environment data, resulting in a rich set of initial feature vectors. The prompt word construction module in the model (which can be an adaptation layer with gating or attention mechanisms, where load level serves as the key control signal, determining the behavior of its internal learnable selector) constructs multimodal prompt words modulated by the load level, which are then fed into the intent decoder (such as a fully connected network). The decoder parses and infers from these prompt words, outputting the vehicle control intent. This module can generate vehicle control information using a simple mapping table (e.g., pre-establishing an intent-parameter mapping table so that when the vehicle control intent is to adjust the air conditioning temperature and the user profile is identified as age A, the target temperature is directly mapped to 24℃). Alternatively, it can convert the vehicle control intent information into feature vectors, simultaneously retrieve user profile data from storage, and use a feature cross-network (e.g., using a simple attention fusion layer) to align and fuse the intent features and user profile features (e.g., calculating the attention weights between intent features and profile features, dynamically weighting and fusing them, and then inferring the final temperature setting of 25℃ and fan speed level 2 through the network). This vehicle control information can be a structured sequence of instructions that can directly control the execution of each actuator. Through this phased processing, most of the complex multimodal fusion and intent inference computations are limited to the first stage, and the computational load is dynamically adjusted according to the load level. The second stage only requires a lightweight feature fusion, which greatly reduces the peak computational overhead and memory usage of a single inference, better adapting to the environment of limited vehicle resources and solving the problem of limited onboard computing power.

[0094] In this embodiment, an intelligent model (such as a large language model LLM, which has multi-turn context awareness and dialogue coherence understanding capabilities) trained based on historical multi-source data (such as historical user voice data, vehicle environment data, user profile data, and vehicle control information corresponding to actuator control records) can be used. Optionally, considering that the vehicle control information corresponding to the user's historical data (i.e., the vehicle control information of the user's historical decisions) may not conform to the system's functional logic, safety specifications, or personalized adaptation requirements (e.g., historical vehicle control information has operational redundancy, does not match the actuator parameters of the current vehicle model, or has not incorporated the latest scenario safety rules, etc.), this historical multi-source data information can be other sample data, such as historical user voice data, vehicle environment data, user profile data, and vehicle control information after manual fine-tuning.

[0095] During training, the output of vehicle control intentions can be trained first using user voice data and vehicle environment data, and then the output of vehicle control information can be trained using vehicle control intentions and user profile data. In other words, the model training phase can be divided into two stages. The first stage focuses on extracting vehicle control intentions, inputting only historical user voice data (e.g., "It's a little hot") and vehicle environment data (e.g., vehicle speed and road type at the corresponding time). Through training, the model learns the intention reasoning ability under the "voice-environment" association. For example, from "It's a little hot," it outputs the intention to "lower the car's interior temperature." Secondly, the second stage, based on the vehicle control intentions output in the first stage, introduces user profile data (e.g., the speaker is an elderly person) for secondary training, focusing on learning the vehicle control information generation logic under "intention-profile" matching. For example, for the intention to "lower the car's interior temperature," combined with the elderly user profile and environmental data, it outputs specific vehicle control information such as "set the air conditioner to 24°C and turn on the internal circulation." During training, the model can be fine-tuned using historical multi-source data. For example, if the historical multi-source data shows that the user's vehicle control information in a high-speed scenario is that the windows are fully open, which does not comply with safety regulations, the model can be made more adaptable by processing the training data.

[0096] During model application, the vehicle inputs real-time user voice data (such as commands like "turn on the air conditioner" or "lower the windows") and / or vehicle environmental data (such as interior and exterior temperatures, vehicle speed, road data, and weather conditions) and user profile data (such as elderly / children) into the intelligent model. The intelligent model performs semantic understanding, environmental feature extraction, and intent reasoning on the input data to generate vehicle control intents that include user vehicle control intents (i.e., control needs directly expressed by the user through voice, such as adjusting the seat angle) and / or scenario-based vehicle control intents (i.e., implicit control needs derived from vehicle environmental data, such as the unmentioned intent to turn on the air conditioner in a high-temperature environment). Next, the intelligent model further integrates user profile data (such as user age, frequently used seat positions, etc.) to personalize the vehicle control intents, ultimately generating vehicle control information. This information can explicitly instruct specific control operations on at least one actuator (or execution device, such as the air conditioner, windows, or seats) of the vehicle to achieve an efficient vehicle control process.

[0097] It should be understood that this embodiment is only used as an example to illustrate the concept of multi-source data information simultaneously including user voice data and vehicle environment data. In some embodiments, only voice data and user profile data may be considered, or when the user does not provide voice input, vehicle environment data may be used to identify the scene vehicle control intent, and vehicle control information may be generated using the scene vehicle control intent and user profile data. For example, when the user does not provide voice input and a sudden drop in ambient temperature is detected, the intelligent model can identify the scene vehicle control intent (such as raising the temperature) based on the vehicle environment data, and recommend personalized vehicle control information to the user using the scene vehicle control intent and user profile data, such as raising the air conditioning temperature, closing the windows, etc.

[0098] To further enhance the user interaction experience, such as Figure 2b As shown, after generating vehicle control information based on vehicle control intent and user profile data in step S202, the following steps S203 and S204 are also included.

[0099] Step S203: Display the information generation link corresponding to the vehicle control information. The information generation link includes at least the input information corresponding to the multi-source data information and the output information corresponding to the vehicle control information.

[0100] After generating vehicle control information, the system displays the corresponding information generation chain. This information generation chain can include the entire process information from input to output, or it can display content including input information, output information, and other key node information. Based on this information generation chain, the user can further decide whether to proceed with the subsequent vehicle control process.

[0101] For example, the multimodal reasoning logic process of vehicle control information can be transformed into a dynamic decision-making flowchart and rendered in real time on the in-vehicle interactive interface (HMI). This enables generative interactive visualization of "user needs → feature analysis → service triggering," making the model decision-making process transparent and enhancing the confidence of adaptive vehicle control results. Users can better understand the model decision-making process based on this, thereby making more informed decisions to confirm or adjust vehicle control information. This makes the vehicle control information ultimately used to execute vehicle control more in line with the user's personal needs and preferences, thereby improving the user experience.

[0102] As a further example, taking the intelligent model using an LLM model (which can output vehicle control information in natural language) as an example, the in-vehicle interactive interface (HMI) is as follows: Figure 3As shown, when a user inputs the voice data "It's so hot," the system, upon receiving this voice data, uses an LLM model combined with vehicle environmental data and user profiles to make vehicle control decisions (i.e., generate vehicle control information). The decision result, such as "Turn on the air conditioning and seat ventilation," is fed back to the user in natural language, such as "Are you feeling hot? Let me turn on the air conditioning and seat ventilation for you based on the environment." When displayed on the HMI interface, the HMI can include controls (such as buttons) - "View Execution Decision." By triggering this control, the user can view the information generation chain of this decision, including the reasoning chain from input information to output information. Further exemplifying this, the information generation chain can include user input, displaying the user's voice data, such as "It's so hot"; dynamic feature display, such as the in-car temperature of 28°C, elderly passenger in the front seat, air quality, etc.; and decision labels, which can be displayed as safety, comfort, and performance, representing the effect and user experience achieved by the model's output vehicle control services for different scenarios. If the user's vehicle control intention is in a scenario affecting driving, the model will provide a conservative service based on the "safety" label when outputting vehicle control. When the vehicle control intent is in a daily scenario and energy consumption is sufficient, the model will provide more comprehensive services based on the "comfort" label when outputting vehicle control information; control logic, such as compressing control commands into verb-object phrases within 8 bytes (such as "temperature gradient compensation" and "energy consumption dynamic balance"), shows users the main reasons for the current vehicle control device's operation after inference by the model, and the brief introduction makes it easy for users to directly understand its control logic; output results, displaying vehicle control information, such as air conditioning cooling; temperature 26 degrees; fan speed 2, etc.

[0103] Step S204: In response to the confirmation information for the information generation link, perform vehicle control operations based on the vehicle control information.

[0104] Understandably, in response to, is used to indicate the conditions or states on which the operation being performed depends. When the conditions or states on which it depends are met, one or more operations being performed can be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations being performed are executed.

[0105] In some embodiments, if a user, after reviewing the information generation process, deems the generated vehicle control information unsuitable for their needs, they can modify the information. This can be done, for example, by issuing a voice command or by inputting modifications into the HMI interface. The system then executes vehicle control based on the modified information. Optionally, executing vehicle control operations based on the vehicle control information can be achieved by sending the information to an actuator for execution, or by sending it to a vehicle control app via a voice app, thus enabling the actuator to execute the operation. This embodiment does not specifically limit the method of executing vehicle control operations.

[0106] In one implementation, the steps described above, which involve constructing multimodal prompts based on a pre-set intelligent model, user voice data, / or vehicle environment data, and load level, may include the following steps:

[0107] Semantic text features are extracted from user voice data, and environmental perception features and vehicle status features are extracted from vehicle environment data.

[0108] Based on semantic text features, environmental perception features, and vehicle status features, initial multimodal prompt words are constructed;

[0109] Based on the vehicle's current load level, the initial multimodal cue words are processed to obtain multimodal cue words.

[0110] For example, semantic information can be extracted from user voice data, such as by converting speech to text (speech recognition) and analyzing the text to extract key semantic features. Natural Language Processing (NLP) techniques can be used to extract semantic text features. Environmental perception features (such as weather and road features) and vehicle status features (such as vehicle speed and air conditioning status) can be extracted from vehicle environmental data using computer vision techniques. Next, by integrating the semantic text features extracted from the voice data with features extracted from the environmental and vehicle status data, a multimodal prompt is constructed. This prompt integrates the user's (semantic) intent, the current environmental conditions, and the vehicle status, providing a more comprehensive information input for the intelligent model. The intelligent model utilizes its trained knowledge and reasoning abilities to analyze and process this multimodal prompt, considering not only the intent corresponding to the user's voice itself but also integrating environmental information. This allows for more accurate identification of vehicle control intent, and the vehicle control information generated based on this intent can effectively improve the environmental adaptability of voice-controlled vehicles.

[0111] Considering the vehicle's processing of user voice data and other data (such as receiving user voice data), Figure 1 During the process of receiving signals via IPC communication and model inference tasks, a significant amount of computing resources (such as CPU computing power, NPU (AI computing power), and memory (bandwidth / throughput)) are required. Insufficient computing power can easily lead to high latency in the voice-controlled vehicle process, resulting in the vehicle's inability to respond promptly and impacting the user experience. In one implementation, the vehicle's current load information can also be obtained, and the vehicle's current load level can be determined based on this information. The aforementioned multimodal prompt word construction process incorporates the vehicle's current load level when constructing multimodal prompt words.

[0112] For example, the current load level of the vehicle can be obtained based on the vehicle's current load information, which may refer to the hardware resource usage allocated to the adaptive vehicle control app (such as CPU utilization, NPU computing power utilization, and memory bandwidth utilization), directly affecting data processing latency. Data processing latency can be determined by the first latency from when the user issues a voice command (i.e., user voice data input) to when the system (i.e., the adaptive vehicle control app) receives the user's voice data, and the second latency from the start time of model inference to the end time of model inference. For example, if the first delay is lower than a first preset threshold (e.g., 50ms) and the second delay is lower than a second preset threshold (e.g., 100ms), the load level is determined as the first load level Q1; if the first delay is greater than or equal to the first preset threshold and less than the third preset threshold (e.g., 100ms), and the second delay is greater than or equal to the first preset threshold and less than or equal to the fourth preset threshold (e.g., 200ms), the load level is determined as the second load level Q2; if the first delay is greater than or equal to the third preset threshold and less than the fifth preset threshold (e.g., 20ms), or the second delay is greater than or equal to the fourth preset threshold and less than the sixth preset threshold (e.g., 300ms), the load level is determined as the third load level Q3; if the first delay is greater than or equal to the fifth preset threshold, or the second delay is greater than or equal to the sixth preset threshold, the load level is determined as the fourth load level Q4. Wherein, load level Q4 > Q3 > Q2 > Q1. Further, the above steps, based on the vehicle's current load level, process the initial multimodal prompt words to obtain multimodal prompt words, which may include the following steps:

[0113] Based on the length information of the load level indication, the number of semantic text features, environmental perception features, and vehicle status features in the initial multimodal prompt words are adjusted to obtain the adjusted feature information;

[0114] Based on the adjusted feature information, multimodal cue words are obtained.

[0115] For different load levels, the initial multimodal prompts are simplified to shorten the prompts used for model inference, thereby reducing the computational load and time of model inference, adapting to the current load capacity, and improving response efficiency. Specifically, the prompt processing method combined with the load level can be achieved by adjusting the length information of the multimodal prompts and performing feature filtering based on this length information. Optionally, it can be combined with... Figure 4a and Figure 4bAs shown, the user issues a voice command "XXX". The voice app collects the user's voice data and records a timestamp ts1. The adaptive vehicle control app calls the user's voice data via IPC and records a timestamp ts2 to calculate the first latency (i.e., data transmission time): ts2 - ts1. Combined with the time from the start of the last model inference ts3 to the end of the last model inference ts4, the second latency (i.e., model inference time): ts4 - ts3 is calculated. Using ts2 - ts1 and ts4 - ts3, load grading is performed. The processing methods for different load levels can be exemplified as follows:

[0116] Q1 (Low Load) can be used without simplification, retaining 100% of the length of the initial multimodal cue words (such as including all features) to ensure the accuracy of model inference in low load scenarios.

[0117] Q2 (Light Load): Simplify to 60% of the length of the initial prompt, remove 40% of the features, such as removing redundant modifiers (e.g., retain the core instruction "turn on the air conditioner" and remove modal particles and repetitive expressions in the speech), retain key environmental and state features, and control the model to output no more than four instructions.

[0118] Q3 (Medium Load): Simplify to 20% of the length of the initial prompt, retain core semantic instructions and necessary environmental parameters (such as "turn on the air conditioner" + "in-car temperature 28℃", omitting secondary features such as "current road speed limit"), execute more simplified model inference, and control the model output instructions to within 2.

[0119] Q4 (Heavy Load): The initial prompt can be completely reduced (i.e., no prompt is built). The system returns the semantics corresponding to the original speech data (e.g., if the user says "turn on the air conditioner," simply send the command "turn on the air conditioner"), freeing up model resources and reducing intermediate computation steps. Furthermore, the environmental perception module's environmental perception and the prompt building module's prompt word construction can be stopped, further optimizing system resources.

[0120] In one example, combining Figure 5As shown, the prompt construction process can include three stages from left to right: data acquisition, data perception, and prompt construction. In the data acquisition stage, raw data is collected. Vehicle status data such as vehicle speed, gear position, air conditioning, seat heating, seat ventilation, and windows are obtained through CarService. Map navigation data can be obtained from navigation apps, weather data from cloud-based weather services, user voice data from voice apps, and other relevant data from IMS / OMS. In the data perception stage, the raw data is processed to convert it into a data format that the model can handle. Specifically, this can include monitoring and converting vehicle speed data, gear position data, seat heating data, seat ventilation data, air conditioning data, window data, navigation data, image data, speaker location data, and weather and air quality data, resulting in processed multi-source data information. For example, vehicle speed data monitoring and conversion can involve converting the raw signal from the vehicle speed sensor into a "value + state" (such as 60km / h, constant speed) that the model can understand; image data monitoring and conversion → processing camera images to recognize user actions and cabin environment (such as recognizing "user raising hand" and "whether the window is closed"); weather and air quality acquisition and conversion → connecting to weather services to convert "cloudy" and "PM2.5 value" into standardized environmental parameters, and so on. During the prompt construction process, by receiving multi-source data information that has undergone perception processing, features can be extracted from the multi-source data information, such as semantic text features, environmental perception features, and vehicle state features, and combined with the load classification results, multimodal prompt words that conform to the corresponding load level can be constructed to balance model accuracy and latency.

[0121] By adopting the above-mentioned hierarchical simplification strategy, and dynamically adjusting the information density of prompt words, the model prioritizes inference speed when the load increases and takes into account inference accuracy when the load is sufficient, thus effectively achieving a balance between model processing efficiency and output accuracy.

[0122] Optionally, the vehicle control intent may carry an intent hierarchy for the vehicle control intent, which is used to indicate the strength of the intent of the vehicle control intent.

[0123] For example, intelligent models can determine whether a user has a clear intention to adjust their speech when expressing certain situational intents based on different modifiers. Alternatively, for voice commands without modifiers, vehicle environment data can be combined to determine the intent level. For example, if the user's command is "cold," different intent levels can be assigned to different in-car temperatures. Optionally, intent levels can include weakness, strength, and extreme levels, thereby helping the model provide more scenario-based vehicle control services. By incorporating modifier intent intensity mapping samples (e.g., "a little cold" corresponds to weakness, "too cold" to strength, and "unbearably cold" to extreme) and environmental parameter-intent intensity association samples (e.g., when a user says "cold," an in-car temperature of 15°C corresponds to strength, 10°C to extreme, and 20°C to weakness), the model learns the logic for determining the strength of intent under different input conditions. For vehicle control intents under different intent levels, it generates corresponding vehicle control information (e.g., in the air conditioning control scenario, the air conditioning adjustment value increases sequentially from weak to extreme for different intent levels).

[0124] In one implementation, the above-mentioned generation of vehicle control information based on vehicle control intent and user profile data may include the following steps:

[0125] Extract user biometrics from user profile data. User biometrics include at least one of age characteristics, action characteristics, and location characteristics.

[0126] Based on vehicle control intent, user biometrics, and a knowledge graph of predefined vehicle control rules in the intelligent model, a structured instruction sequence for vehicle control is constructed. The structured instruction sequence includes control instructions for multiple actuators.

[0127] Vehicle control information is generated based on the structured instruction sequence.

[0128] For example, user profile data can be user image data. During feature extraction, convolutional neural networks, pose estimation algorithms, and other methods can be used to extract relevant user biometric features. For instance, for age features, deep learning models (such as convolutional neural networks) can be used to analyze facial features to estimate the user's age. Alternatively, pose estimation algorithms can be used to identify the user's body posture and movement features. Furthermore, background information in the image can be combined to extract the user's location features.

[0129] After extracting the user's biometrics, guided by the vehicle control intent, and combined with the extracted user biometrics, the system calls upon a predefined vehicle control rule knowledge graph in the intelligent model (those skilled in the art can set this vehicle control rule knowledge graph based on a large amount of historical data or experience, which may include rules such as actuator association, scene constraints, and preference mapping). The system filters and matches logic and sorts out priorities, constructing a structured instruction sequence containing multiple actuator control instructions (which may contain multiple control instructions, such as closing the driver's side window, setting the air conditioner to 26 degrees Celsius, and setting the fan speed to level 3). The system can also convert each control instruction into a standardized control signal according to the actuator type, and can mark the execution priority for each execution (e.g., safety actuators (such as transmission actuators that affect vehicle speed) > comfort actuators (such as air conditioners or other actuators that affect temperature)). The system generates instruction information that can directly drive the vehicle actuators to perform actions, thereby achieving intelligent vehicle control.

[0130] In some embodiments, to further improve the prediction accuracy of the intelligent model, during the training process, the constraints of equipment control rules under different working conditions can be further considered. By establishing a dynamic mapping relationship between the characteristic parameters of the control equipment (or actuator or execution device) and the model output, the accuracy of model decision-making and scenario adaptability can be improved. In this embodiment, the characteristic changes of various execution devices are used as key training constraints to construct a closed-loop feedback mechanism of equipment characteristics-control output. This can effectively prevent the model from generating arbitrary control commands that are detached from physical constraints. The following uses the air conditioner, seat, and window as examples in Table 1 to illustrate this in detail:

[0131] Table 1:

[0132]

[0133] In some embodiments, considering the different vehicle control intentions identified, the vehicle control information inferred by the model may include control instructions for multiple execution devices. To ensure that each execution device executes in an orderly manner, or to optimize execution (omitting unnecessary execution processes), restrictions and priority judgments can be set for different execution devices during model training for different vehicle control intentions, thereby improving user experience and reducing energy consumption. Specifically, there are different levels of rule restrictions on the execution devices output by the model, which can be divided into three levels: devices that must be executed according to the vehicle control intention, devices that are executed based on environmental decisions, and devices that are not executed unless necessary (similar to labeling each control instruction in the control information to distinguish the level of each execution device). Examples of device restrictions and priorities in different vehicle control intentions will be given below. For example, if the vehicle control intention is to cool down: air conditioning control must be executed, and the air conditioning control preset can be determined by combining the execution logic in Table 1 above; whether the seat is controlled can be determined based on temperature characteristics; whether the windows are controlled can be determined based on characteristics such as vehicle speed, temperature, and weather. Heating Intent: Air conditioning control must be executed. The control threshold for air conditioning is determined according to the execution logic in Table 1 above. Whether the seats are controlled depends on temperature characteristics. Whether the windows are controlled depends on vehicle speed, temperature, weather, and other characteristics. Purification and Ventilation Intent: For the intention of a stuffy, poorly ventilated interior: Whether the windows are controlled depends on vehicle speed, temperature, weather, and other characteristics. If the windows do not meet the requirements, air conditioning control will be executed. For the intention of an unpleasant odor in the car: Whether the windows are controlled depends on vehicle speed, temperature, weather, and other characteristics. If the windows do not meet the requirements, air conditioning control will be executed. For the intention of air purification: Window control (closed) must be executed. Air conditioning control (air purification, recirculation) must be executed. For the intention of rest: For the intention to sleep or rest, seat control is prioritized, with the front passenger seat controlled based on passenger characteristics. Whether the air conditioning is controlled depends on temperature characteristics. Whether the windows are controlled depends on vehicle speed, temperature, weather, and other characteristics. For the intention of relaxation: Seat massage must be executed. Whether the seat position is controlled depends on vehicle speed, passenger characteristics, etc.. Whether the air conditioning is controlled depends on temperature characteristics. Whether the windows are controlled depends on vehicle speed, temperature, weather, and other characteristics.

[0134] In some embodiments, further considering the inference latency issue of the intelligent model, a dynamic latency compensation mechanism can be constructed to deploy a dual-channel collaborative strategy between the vehicle control app and the intelligent model: A timeout circuit breaker strategy is established, by setting an execution time window (e.g., 1000ms, which can be adjusted by those skilled in the art according to hardware performance or scenario requirements). When the intelligent model does not output vehicle control information within this execution time window (or the output vehicle control information does not contain valid control commands), the vehicle control baseline strategy is automatically triggered, and model resources are released. Secondly, a command validity detection module is constructed. When the intelligent model outputs valid control commands in real time, a dynamic latency refresh algorithm continuously optimizes the timing arrangement of the voice execution queue to ensure the atomic execution of streaming commands and dynamic adaptation of system resources. Simultaneously, the large model output paradigm undergoes a key upgrade and transformation, evolving from raw text generation to structured service call commands. Through standardized API interfaces, it connects to the vehicle control bus, achieving end-to-end digital mapping from natural language understanding to onboard actuator actions, realizing the real-time response capability and operational robustness of the vehicle control system.

[0135] Alternatively, the method may also include the following steps:

[0136] Determine whether the intelligent model generates vehicle control information within a predefined execution time window;

[0137] If no vehicle control information is generated within the preset time period, the predefined vehicle control baseline strategy is triggered, and the resources of the intelligent model are released. The vehicle control baseline strategy is used to query and execute basic vehicle control operations corresponding to user voice data and / or vehicle environment data.

[0138] For example, in the interaction logic between the vehicle control app and the intelligent model, a predefined execution time window is used as the maximum tolerable time limit for the model to generate valid vehicle control information. By continuously monitoring the output status of the intelligent model, it is determined whether it generates vehicle control information within the aforementioned time window. If the intelligent model does not generate vehicle control information within the time window, or the generated information does not contain valid instructions (such as invalid text output due to model illusion), the vehicle control baseline strategy is automatically triggered. This strategy pre-stores basic control logic corresponding to user voice data (such as "turn on the air conditioner"), and / or vehicle environment data (such as in-vehicle temperature 30℃) and user profile (such as elderly) (such as directly executing "air conditioner cooling mode, target temperature 26℃"), so that the user's core needs can still be responded to. At the same time, the system releases the hardware resources (such as NPU computing power and memory) occupied by the intelligent model, avoiding the accumulation of load caused by resource waste.

[0139] In one implementation, the method further includes: identifying valid control commands and invalid control commands in the vehicle control information according to predefined command recognition rules, and filtering out invalid control commands.

[0140] The steps for executing vehicle control operations based on vehicle control information are as follows: according to predefined instruction priority rules, the effective control instructions are time-sequentially arranged to obtain the timing arrangement result; based on the timing arrangement result, the vehicle control operations are executed sequentially according to the effective control instructions.

[0141] For example, those skilled in the art can set instruction recognition rules based on actual applications or experience (such as "instructions containing actuator ID, control parameters, and operation type are valid control instructions", for example, "air conditioning actuator: cooling, 26°C"; instructions without a clear execution object or parameters are invalid instructions). Based on these instruction recognition rules, the vehicle control information output by the intelligent model can be filtered, invalid instructions can be removed, and valid control instructions that can be executed can be retained, so as to facilitate the efficient execution of vehicle control information.

[0142] Next, based on predefined instruction priority rules (e.g., safety actuator instructions > comfort actuator instructions, such as "brake instruction" taking precedence over "air conditioning adjustment instruction"), valid control instructions are sequentially ordered to generate a timing arrangement result containing the execution order and time interval (e.g., "execute window closing instruction first, then air conditioning purification instruction 1 second later"). According to the timing arrangement result, valid control instructions are sequentially sent to the vehicle control bus to drive the corresponding actuators. During this process, the dynamic delay refresh algorithm adjusts the instruction execution interval in real time based on system resource load (e.g., CPU utilization, bus bandwidth), ensuring that the atomic execution of streaming instructions (one vehicle control instruction is executed before the next one starts) adapts to the resource state.

[0143] Optionally, considering that the generated control information is streaming data composed of multiple control commands, to facilitate efficient execution of the control commands, a streaming text segmenter can be used to segment it into separate multi-control command texts in real time. The command texts can then be converted into vehicle control API commands using an NLU parsing engine. Furthermore, corresponding voice control protocol data can be sent to the voice app using a customized protocol, thereby achieving coordinated control of multiple devices such as air conditioning, seats, and lights. Figure 6As shown, when a user issues a voice command, such as saying "It's a bit hot," the user's voice data is distributed to the intelligent model, and a countdown begins. If the countdown is 0, the vehicle control baseline strategy is used to implement the vehicle control process (based on the vehicle control baseline strategy, the corresponding basic vehicle control operation is queried and sent to the vehicle control app, and the corresponding status is adjusted through CarService, such as windows, air conditioning, seat ventilation, etc.). When the countdown has not expired (>0), the control process is executed based on the output of the intelligent model. Specifically, voice data is first received and combined with weather service (weather, air quality), navigation app (road attributes, latitude and longitude data), virtual car (gear, speed, air conditioning mode, temperature, air volume, in-vehicle and out-of-vehicle temperature, seat ventilation, seat massage, window opening), and load level data to construct (multimodal) prompts. After constructing multimodal prompts, model inference is performed to identify vehicle control intentions. Vehicle control information is generated by combining user profile data and vehicle control intentions. This information is streaming text corresponding to a structured instruction sequence. A streaming text segmenter (based on temporal arrangement) performs streaming parsing, segmenting the text into individual control instructions and counting them (n-1, where n represents the total number of control instructions). For example, during streaming segmentation, semicolons can be used for segmentation, and periods are used as terminators to complete the instruction sequence loop. Real-time monitoring is performed based on an incremental processing mode. When a semicolon or period is detected, the integrity of a single control text is triggered, achieving synchronous operation of real-time instruction stream parsing and execution triggering. To facilitate efficient execution of control instructions, a corresponding state identifier (index) can be assigned to each control instruction while it is being stream-segmented. The state identifier can contain a triple state identifier (0-start, 1-continue, 2-terminate, corresponding to the first instruction obtained after segmentation, the intermediate instruction obtained after segmentation, and the last instruction obtained after segmentation, respectively). The control instructions after streaming parsing and assigning state identifiers are as follows: Figure 7 As shown. Next, the NLU parsing engine sequentially converts the control commands into vehicle control API semantics, carrying the corresponding index, and transmits them to the vehicle control app to execute the control flow. This can effectively avoid execution disorder caused by the parsing of multiple commands mixed together. At the same time, it allows the vehicle control terminal to call resources according to the status identifier (such as calling more resources for the last command), avoiding resource idleness or overload, thereby reducing response latency.

[0144] To verify the actual performance of the embodiments of the present invention under limited computing resources on the vehicle side, benchmark tests were conducted against mainstream industry solutions, as shown in Table 2:

[0145] Table 2

[0146]

[0147] It should be understood that “4x LPDDR5x 10667Mbps + 4-Lane UFS 4.1” in Table 2 refers to Low Power Double Data Rate 5x (LPDDR5x) memory with a 64-bit bit width and a total bandwidth of up to 5.3GB / s.

[0148] As shown in Table 2 above, even though the vehicle's basic computing power is inferior (CPU clock speed 2.9GHz vs 4.2GHz, NPU computing power 35TOPS vs 100TOPS), it still surpasses the competition in key performance indicators: the prefill speed is 215 tokens / s faster, and the decoding speed is also faster, achieving a lead of 11.4 tokens / s. The test results fully demonstrate that this solution, through its multimodal prompt word construction and phased fusion with vehicle control intent architecture, significantly reduces data processing complexity, achieves superior performance in resource-constrained environments, and effectively solves the vehicle-side computing power bottleneck problem.

[0149] In summary, this invention provides a vehicle control solution adapted to smart cockpits. By understanding the user's natural contextual voice commands or recognizing external scenes in real time, it proactively provides a comprehensive suite of vehicle control services, offering users an intelligent and differentiated driving experience. The voice-based vehicle control solutions described in this invention cover various high-frequency scenarios, from temperature control and ventilation / purification to proactive services in abnormal weather and parking scenarios, achieving dynamic and differentiated services and providing users with a better experience.

[0150] Figure 8a This is a schematic diagram of the structure of a vehicle control device provided in an embodiment of the present invention, as shown below. Figure 8a As shown, the device 800 includes a data acquisition module 801 and a model processing module 802. Wherein,

[0151] Data acquisition module 801 is used to acquire multi-source data information, including user voice data and / or vehicle environment data, user profile data, and the current load level of the vehicle;

[0152] The model processing module 802 is used to construct multimodal prompt words based on a preset intelligent model, user voice data and / or vehicle environment data, and load level. The load level is used to indicate the length information of the multimodal prompt words. The module also obtains the vehicle control intent based on the multimodal prompt words and generates vehicle control information based on the vehicle control intent and user profile data.

[0153] Among them, vehicle control intent includes user vehicle control intent and / or scenario vehicle control intent, intelligent model is a model trained based on historical multi-source data information and used to generate vehicle control information, and vehicle control information is used to instruct control operations on at least one actuator of the vehicle.

[0154] Furthermore, such as Figure 8b As shown, it may also include a display module 803 and a vehicle control module 804:

[0155] Display module 803 is used to display the information generation link corresponding to vehicle control information. The information generation link includes at least the input information corresponding to multi-source data information and the output information corresponding to vehicle control information.

[0156] The control module 804 is used to execute vehicle control operations based on vehicle control information in response to confirmation information for the information generation link.

[0157] In one embodiment, the model processing module 802 includes:

[0158] The first feature extraction unit is used to extract semantic text features from user voice data and environmental perception features and vehicle status features from vehicle environment data.

[0159] The cue word construction unit is used to construct initial multimodal cue words based on semantic text features, environmental perception features, and vehicle state features; and,

[0160] Based on the vehicle's current load level, the initial multimodal cue words are processed to obtain multimodal cue words.

[0161] In one implementation, the prompt word construction unit is specifically used for:

[0162] Based on the length information of the load level indication, the number of semantic text features, environmental perception features, and vehicle status features in the initial multimodal prompt words are adjusted to obtain the adjusted feature information;

[0163] Based on the adjusted feature information, multimodal cue words are obtained.

[0164] In one implementation, the vehicle control intent carries an intent hierarchy for the vehicle control intent, the intent hierarchy being used to indicate the strength of the intent of the vehicle control intent.

[0165] In one embodiment, the model processing module 802 includes:

[0166] The second feature extraction unit is used to extract user biometric features from user profile data. User biometric features include at least one of age features, action features, and location features.

[0167] The instruction construction unit is used to construct a structured instruction sequence for vehicle control based on vehicle control intent, user biometrics, and a predefined vehicle control rule knowledge graph in the intelligent model. The structured instruction sequence includes control instructions for multiple actuators; and to generate vehicle control information based on the structured instruction sequence.

[0168] In one implementation, it further includes:

[0169] The judgment module is used to determine whether the intelligent model generates vehicle control information within a predefined execution time window;

[0170] The triggering module is used to trigger a predefined vehicle control baseline strategy and release the resources of the intelligent model if no vehicle control information is generated within a preset time period. The vehicle control baseline strategy is used to query and execute basic vehicle control operations corresponding to user voice data and / or vehicle environment data.

[0171] In one implementation, it further includes:

[0172] The instruction recognition module is used to identify valid and invalid control instructions in vehicle control information according to predefined instruction recognition rules, and to filter out invalid control instructions.

[0173] Vehicle control module 804 includes:

[0174] The timing orchestration unit is used to orchestrate the timing of valid control instructions according to predefined instruction priority rules, and obtain the timing orchestration result;

[0175] The control unit is used to execute vehicle control operations sequentially according to valid control commands based on the timing arrangement results.

[0176] In one embodiment, environmental data includes environmental perception data and vehicle status data, wherein the environmental perception data includes weather data and road data; the data acquisition module 801 includes:

[0177] The first acquisition unit is used to acquire user voice data through the vehicle's voice acquisition device;

[0178] The second acquisition unit is used to acquire environmental perception data through the cloud and / or the vehicle's environmental perception sensors, and to acquire vehicle status data through the vehicle's domain controller.

[0179] The third acquisition unit is used to acquire user profile data through the vehicle's image acquisition device.

[0180] Figure 9 This is a vehicle control application system provided in an embodiment of the present invention, such as... Figure 9 As shown, the vehicle control application system 900 may include:

[0181] At least one processor 901; and

[0182] Memory 902 is communicatively connected to at least one processor; wherein,

[0183] The memory 902 stores instructions that can be executed by at least one processor, which enables the at least one processor to perform the vehicle control method described above.

[0184] Figure 10 This invention provides a testing method for a vehicle control application system. Optionally, the vehicle control application system can be... Figure 1 Adaptive vehicle control apps, such as Figure 10 As shown, the method includes steps S1001-S1003.

[0185] Step S1001: Construct a test set for testing the vehicle control application system based on multi-source data information with different state values.

[0186] In this embodiment, the state value, i.e. the data value, such as different state values ​​of temperature, are different temperature values.

[0187] Step S1002: Input the test set into the vehicle control application system through the test application, perform segmented testing on the vehicle control application system, and obtain segmented test information for each segmented test; wherein, the segmented test includes at least a first segmented test on the construction of multimodal prompt words, and a second segmented test on vehicle control information and its execution results.

[0188] In this embodiment, considering that the vehicle control application system involves multiple modules connected in series from data reception to model processing and vehicle control information execution, and that the data flow link is relatively long, the upstream scenario is difficult to realistically fulfill the requirements of the long link, leading to difficulties in guaranteeing the accuracy of the test results. This embodiment performs the test of the vehicle control application system through segmented testing, which includes verifying the accuracy of multi-source data feature extraction and constructing multimodal prompt words, and testing the execution results after multimodal prompt word inference.

[0189] Specifically, the test suite can be converted into a JSON file format, and then input into the vehicle control application system through the testing application. JSON (JavaScript Object Notation) is a lightweight data-interchange format that uses a hierarchical tree structure of key-value pairs and supports data types such as strings, numbers, booleans, arrays, objects, and null. Since the test suite involves a large amount of data, converting it to JSON format facilitates the vehicle's understanding (avoiding confusion between different data types) and simulation processing to obtain the response results for the test suite.

[0190] Optionally, the above steps involve segmented testing of the vehicle control application system, including:

[0191] Obtain log data from the vehicle control application system regarding the test set. The log data includes a first log of constructing multimodal prompts and a second log of executing vehicle control information and obtaining execution results.

[0192] The first segment test is executed based on the first log and the predefined prompt word evaluation rules; and the second segment test is executed based on the second log and the predefined execution result evaluation rules.

[0193] Specifically, the log data of the vehicle control application system regarding the test set, i.e., the log data of the response results to the test set, can be divided into two logs. The first log can record all data from the input to the prompt construction process, and the second log can record all data from the construction of the prompt to the execution result of the vehicle control based on the vehicle control information. Optionally, predefined prompt evaluation rules can evaluate the prompt construction process from multiple dimensions (such as whether the prompt is missing, whether the prompt is accurate, the latency of user voice data reception, etc.), and predefined execution result evaluation rules can evaluate the execution result from multiple dimensions (such as whether the control commands in the generated vehicle control information are effective, whether the timing arrangement of the commands is reasonable, whether the execution is based on the control commands, etc.). These prompt evaluation rules or execution result evaluation rules can be evaluation rules given by a large amount of user test data, or they can be evaluation rules provided by cloud-based large models (such as DeepSeek or ChatGPT-4o and other existing multimodal models). This embodiment does not make any special limitations on this.

[0194] In this process, by injecting JSON files into the vehicle's infotainment system, the system can better understand the data, build user scenarios to verify differentiated scenario output results, conduct segmented testing based on the prompt, and utilize large cloud models to evaluate the rationality of adaptive vehicle control output results, which can effectively improve the testing effect.

[0195] Step S1003: Obtain the test results of the vehicle control application system based on the segmented test information.

[0196] In this embodiment, by summarizing the segmented test information from the first segment test (accuracy of prompt words) and the second segment test (validity of execution results), and combining it with the rationality evaluation of the cloud-based large model, a complete test result for the vehicle control application system is formed. If both segmented tests meet the evaluation rules, the system performance is deemed qualified. If any segmented test fails (e.g., missing prompt words or incorrect execution instructions), the problem in the corresponding link (e.g., the feature extraction module or the actuator communication module) is located, providing a basis for system optimization.

[0197] In summary, the technical solutions provided by the embodiments of the present invention have at least the following beneficial effects:

[0198] 1) Personalized Intelligent Vehicle Control: During model processing, user voice data and vehicle environment data are combined to recognize vehicle control intentions. Dynamic user profiles, including age, actions, and habits, are integrated to generate vehicle control information, reducing interaction steps in complex scenarios and ultimately outputting accurate, differentiated vehicle control commands (e.g., "seat ventilation + air conditioning cooling" linkage control). 2) Visualized Decision Display: The model's reasoning logic is transformed into a visual flowchart on the in-vehicle interface (presenting the information chain of "user needs → feature analysis → service output"), supporting user tracking of the decision-making process, improving the confidence of vehicle control results, and significantly enhancing the explicitness of execution. 3) Flexible Performance Adjustment: Model performance is dynamically and hierarchically controlled based on CPU / NPU load (e.g., automatically compressing the number of commands and simplifying the reasoning process under high load), reducing system resource consumption by 35% and keeping complex command processing latency within 200ms; Ecosystem Linkage Response: Through streaming command parsing, cabin devices such as air conditioning, lights, and seats can be linked, supporting "one-sentence multi-control" interaction, improving the efficiency of multi-device collaboration in the cabin by 60%, and reducing the response time for complex needs (e.g., "cooling + ventilation + seat adjustment") to 1 second. 4) Segmented testing and adaptation: By injecting structured simulation data into the system, a realistic driving environment is simulated. Combined with a multi-dimensional feature verification system under real vehicle conditions (covering three dimensions: weather simulation, vehicle state simulation, and human behavior modeling), segmented testing is conducted on multimodal prompts and execution results, forming a closed-loop verification process of "data link → semantic conversion → execution feedback". Based on this testing scheme, the workload of manual review can be reduced, and the verification efficiency of large model products can be significantly improved.

[0199] Figure 11 This is a schematic diagram of the structure of a test device for a vehicle control application system provided in an embodiment of the present invention, as shown below. Figure 11 As shown, the device includes a test set construction module 1101, a segmented testing module 1102, and a test result acquisition module, wherein...

[0200] The test set construction module 1101 is used to construct a test set for testing vehicle control application systems based on multi-source data information with different state values.

[0201] The segmented testing module 1102 is used to input the test set into the vehicle control application system through the test application, perform segmented testing on the vehicle control application system, and obtain segmented test information for each segmented test; wherein, the segmented test includes at least a first segmented test on the construction of multimodal prompt words and a second segmented test on vehicle control information and its execution results;

[0202] The test result acquisition module 1103 is used to acquire the test results of the vehicle control application system based on the segmented test information.

[0203] In one embodiment, the segmented testing module 1102 includes:

[0204] The format conversion unit is used to convert the test set into JSON file format, and the converted test set is input into the vehicle control application system by the test application.

[0205] In one embodiment, the segmented testing module 1102 includes:

[0206] The log construction unit is used to obtain log data of the vehicle control application system about the test set. The log data includes the first log of constructing multimodal prompt words and the second log of executing vehicle control information and obtaining the execution result.

[0207] The test unit performs a first segment test based on the first log and predefined prompt word evaluation rules; and performs a second segment test based on the second log and predefined execution result evaluation rules.

[0208] Figure 12 This is a schematic diagram of the structure of a vehicle provided by the present invention. The vehicle 1200 includes... Figure 9 The corresponding vehicle control application system is 900.

[0209] The present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the above-described method.

[0210] The present invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the method described above.

[0211] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0212] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0213] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0214] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0215] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0216] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0217] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0218] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

[0219] The above embodiments are merely preferred embodiments provided to fully illustrate the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are all within the scope of protection of the present invention.

Claims

1. A vehicle control method, characterized in that, include: Acquire multi-source data information, including user voice data and / or vehicle environment data, user profile data, and the current load level of the vehicle; Based on a pre-set intelligent model, multimodal prompt words are constructed according to the user voice data and / or the vehicle environment data and the load level. The load level is used to indicate the length information of the multimodal prompt words. The vehicle control intention is obtained according to the multimodal prompt words, and vehicle control information is generated according to the vehicle control intention and the user profile data. The vehicle control intent includes user vehicle control intent and / or scenario vehicle control intent. The intelligent model is a model trained based on historical multi-source data information and used to generate vehicle control information. The vehicle control information is used to instruct control operations on at least one actuator of the vehicle. The step of constructing multimodal prompt words based on the user's voice data and / or the vehicle's environmental data, and the load level includes: Based on the semantic text features extracted from the user voice data, the environmental perception features extracted from the vehicle environment data, and the vehicle state features, an initial multimodal prompt word is constructed. Based on the length information of the load level indication of the vehicle, the number of semantic text features, environmental perception features, and vehicle status features in the initial multimodal prompt are adjusted to obtain the adjusted feature information; The multimodal prompt words are obtained based on the adjusted feature information.

2. The method according to claim 1, characterized in that, The vehicle control intent carries an intent classification for the vehicle control intent, and the intent classification is used to indicate the strength of the intent of the vehicle control intent.

3. The method according to claim 1 or 2, characterized in that, The step of generating vehicle control information based on the vehicle control intent and the user profile data includes: User biometrics are extracted from the user profile data, and the user biometrics include at least one of age features, action features, and location features; Based on the vehicle control intent, the user's biometrics, and the predefined vehicle control rule knowledge graph in the intelligent model, a structured instruction sequence for vehicle control is constructed, the structured instruction sequence including control instructions for multiple actuators; The vehicle control information is generated based on the structured instruction sequence.

4. The method according to claim 1 or 2, characterized in that, Also includes: Determine whether the intelligent model generates the vehicle control information within a predefined execution time window; If the vehicle control information is not generated within the preset time period, a predefined vehicle control baseline strategy is triggered, and the resources of the intelligent model are released; wherein, the vehicle control baseline strategy is used to query and execute basic vehicle control operations corresponding to the user voice data and / or vehicle environment data.

5. The method according to claim 1 or 2, characterized in that, The environmental data includes environmental perception data and vehicle status data, wherein the environmental perception data includes weather data and road data; the acquisition of multi-source data information includes: User voice data is collected through the vehicle's voice acquisition device; Environmental perception data is acquired through the cloud or the vehicle's environmental perception sensors, and vehicle status data is acquired through the vehicle's domain controller. User profile data is collected using the vehicle's image acquisition device.

6. The method according to claim 1 or 2, characterized in that, After generating vehicle control information based on the vehicle control intent and the user profile data, the process also includes: The information generation link corresponding to the vehicle control information is displayed. The information generation link includes at least the input information corresponding to the multi-source data information and the output information corresponding to the vehicle control information. In response to the confirmation information for the information generation link, the vehicle control operation is performed based on the vehicle control information.

7. The method according to claim 6, characterized in that, Also includes: According to predefined instruction recognition rules, valid and invalid control instructions in the vehicle control information are identified, and invalid control instructions are filtered out. The step of performing vehicle control operations based on the vehicle control information includes: According to the predefined instruction priority rules, the effective control instructions are time-arranged to obtain the timing arrangement result; Based on the timing arrangement result, the vehicle control operations are executed sequentially according to the effective control commands.

8. A vehicle control device, characterized in that, include: The data acquisition module is used to acquire multi-source data information, including user voice data and / or vehicle environment data, user profile data, and the current load level of the vehicle. The model processing module is used to construct multimodal prompt words based on a pre-set intelligent model, according to the user voice data and / or the vehicle environment data, and the load level, wherein the load level is used to indicate the length information of the multimodal prompt words, and to obtain vehicle control intentions based on the multimodal prompt words, and to generate vehicle control information based on the vehicle control intentions and the user profile data; wherein the vehicle control intentions include user vehicle control intentions and / or scenario vehicle control intentions, the intelligent model is a model trained based on historical multi-source data information and used to generate vehicle control information, and the vehicle control information is used to instruct control operations on at least one actuator of the vehicle; The step of constructing multimodal prompt words based on the user's voice data and / or the vehicle's environmental data, and the load level includes: Based on the semantic text features extracted from the user voice data, the environmental perception features extracted from the vehicle environment data, and the vehicle state features, an initial multimodal prompt word is constructed. Based on the length information of the load level indication of the vehicle, the number of semantic text features, environmental perception features, and vehicle status features in the initial multimodal prompt are adjusted to obtain the adjusted feature information; The multimodal prompt words are obtained based on the adjusted feature information.

9. A vehicle control application system, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the vehicle control method according to any one of claims 1-7.

10. A test method for the vehicle control application system of claim 9, characterized in that, include: A test set for testing the vehicle control application system is constructed based on multi-source data information with different state values. The test set is input into the vehicle control application system through a test application to perform segmented testing on the vehicle control application system, and obtain segmented test information for each segmented test; wherein, the segmented test includes at least a first segmented test on the construction of multimodal prompt words, and a second segmented test on vehicle control information and its execution results; Based on the segmented test information, the test results of the vehicle control application system are obtained.

11. The method according to claim 10, characterized in that, The step of inputting the test set into the vehicle control application system through a test application includes: The test set is converted into a JSON file format, and the converted test set is input into the vehicle control application system through the test application.

12. The method according to claim 10 or 11, characterized in that, The segmented testing of the vehicle control application system includes: Obtain log data from the vehicle control application system regarding the test set. The log data includes a first log for constructing the multimodal prompt words and a second log for executing the vehicle control information and obtaining the execution result. The first segmented test is executed based on the first log and the predefined prompt word evaluation rules; and the second segmented test is executed based on the second log and the predefined execution result evaluation rules.

13. A testing apparatus for the vehicle control application system of claim 9, characterized in that, include: The test set construction module is used to construct a test set for testing the vehicle control application system based on multi-source data information with different state values. The segmented testing module is used to input the test set into the vehicle control application system through the test application, perform segmented testing on the vehicle control application system, and obtain segmented test information for each segmented test; wherein, the segmented test includes at least a first segmented test on the construction of multimodal prompt words and a second segmented test on vehicle control information and its execution results; The test result acquisition module is used to acquire the test results of the vehicle control application system based on the segmented test information.

14. A vehicle, characterized in that, The vehicle includes the vehicle control application system as described in claim 9.

Citation Information

Patent Citations

  • Multi-modal interaction method for vehicle, storage medium and electronic equipment

    CN111966320A

  • Vehicle multi-dimensional control method and device, electronic equipment and storage medium

    CN118182502A