Vehicle control command determination method, device and vehicle

By acquiring multimodal state label data and using scenario orchestration models to identify user intent, the accuracy and reliability issues of traditional in-vehicle voice recognition systems in complex environments are resolved, efficient and accurate vehicle control command conversion is achieved, and the user experience is improved.

CN120080867BActive Publication Date: 2025-09-23CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510565898.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-23
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Traditional in-vehicle voice recognition systems have difficulty accurately converting users' desired in-vehicle operations in complex in-vehicle environments. They are affected by noise interference, multi-tasking requirements and poor scene adaptability, resulting in low user control efficiency.

Method used

By acquiring multimodal state label data, using the scenario orchestration model to identify user intent and generate vehicle control commands, combined with multimodal data fusion and large model fine-tuning, the accuracy and reliability of command conversion are ensured.

Benefits of technology

It improves the processing capability of voice commands in the vehicle environment, enhances the user experience and vehicle intelligence level, and ensures the accuracy and reliability of command conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120080867B_ABST
    Figure CN120080867B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, device, and vehicle for determining vehicle control instructions, and relates to the field of intelligent driving technology. The method for determining vehicle control instructions includes: obtaining first multimodal state label data about an in-vehicle scene from an in-vehicle terminal; performing intent recognition on user interaction information in the first multimodal state label data to obtain user intent information; determining whether the user intent information includes an intent for vehicle control; if the user intent information includes an intent for vehicle control, determining a first vehicle control instruction based on the user intent information, the first multimodal state label data, and a preset scene arrangement model, and sending the instruction to the in-vehicle terminal to control the vehicle. The embodiment of the present application can deeply analyze the user's control intention for the vehicle contained in the user interaction information based on the multimodal state label data in the in-vehicle scene, improve the ability to understand the user's intent, and output vehicle control instructions that better meet the user's needs, thereby enhancing the user's driving experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent driving technology, and in particular to a method and device for determining vehicle control instructions, and a vehicle. Background Art

[0002] With the continuous development of intelligent driving technology, vehicles have become more than just means of transportation; they have also become terminals for information processing and intelligent decision-making. In vehicle scenarios, users interact with onboard systems through voice commands, and speech recognition and natural language processing technologies play a key role in this process.

[0003] Traditional speech recognition systems typically rely on user command data (such as voice commands) and face many challenges in the changing in-vehicle environment, such as noise interference and multi-tasking requirements. It is difficult to accurately convert command data into the user's desired in-vehicle operations, reducing user control efficiency. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present application provides a vehicle control instruction determination method, device and vehicle.

[0005] In a first aspect, the present application provides a method for determining a vehicle control instruction, which is applied to the cloud. The method includes:

[0006] Acquire first multimodal state tag data about an in-vehicle scene from the vehicle computer, where the first multimodal state tag data is determined based on multimodal state data collected by multiple sensors on the vehicle;

[0007] Performing intent recognition on the user interaction information in the first multimodal state tag data to obtain user intent information;

[0008] determining whether the user intention information includes an intention for vehicle control;

[0009] If the user intention information includes an intention for vehicle control, a first vehicle control instruction is determined based on the user intention information, the first multimodal state tag data and a preset scenario arrangement model, and is sent to the vehicle terminal to control the vehicle.

[0010] Optionally, the user interaction information includes: a user voice signal;

[0011] Performing intent recognition on the user interaction information in the first multimodal state tag data to obtain user intent information includes:

[0012] Performing intent recognition on the user voice signal to obtain objective intent information and emotional intent information;

[0013] The objective intention information and the emotional intention information are determined as the user intention information.

[0014] Optionally, determining whether the user intention information includes an intention for vehicle control includes:

[0015] determining whether the user intention information includes a preset keyword associated with the vehicle;

[0016] If the user intention information includes a preset keyword associated with the vehicle, it is determined that the user intention information includes an intention for vehicle control.

[0017] Optionally, the scene arrangement model is obtained by fine-tuning a pre-trained large model, and the fine-tuning training method of the pre-trained large model includes:

[0018] Obtaining model fine-tuning prompt words, user interaction information from multiple vehicle terminals, second multimodal state label data about multiple vehicle scenarios, and second vehicle control instructions labeled for the multiple vehicle scenarios;

[0019] The pre-trained large model is fine-tuned based on the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, multiple second vehicle control instructions and a low-rank adaptation algorithm to obtain the scene arrangement model.

[0020] Optionally, fine-tuning the pre-trained large model based on the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, multiple second vehicle control instructions, and a low-rank adaptation algorithm to obtain the scene arrangement model includes:

[0021] Obtaining a weight matrix including a plurality of first model parameters in a pre-trained large model, wherein the number of the first model parameters in the weight matrix is ​​less than the number of remaining model parameters in the pre-trained large model, and the remaining model parameters are model parameters in the pre-trained large model other than the first model parameters in the weight matrix;

[0022] Decomposing the weight matrix into a first low-rank matrix and a second low-rank matrix;

[0023] Training the first low-rank matrix and the second low-rank matrix using the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, and multiple second vehicle control instructions to obtain a first fine-tuning parameter matrix and a second fine-tuning parameter matrix;

[0024] The scene arrangement model is obtained by replacing the corresponding first model parameters in the pre-trained large model with the second model parameters in the first fine-tuning parameter matrix and the second fine-tuning parameter matrix.

[0025] Optionally, after obtaining user interaction information from multiple vehicle terminals, second multimodal state label data about the vehicle scene, and second vehicle control instructions labeled for the multiple vehicle scenes, the method further includes:

[0026] Get the verification prompt word;

[0027] Inputting the verification prompt word, any one of the user interaction information, the second multimodal state label data, and the second vehicle control instruction corresponding to the second multimodal state label data into a preset large model, so that the preset large model performs rationality verification on the second vehicle control instruction and obtains a rationality verification result;

[0028] If the rationality verification result meets the verification pass conditions, the step of fine-tuning the pre-trained large model based on the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, multiple second vehicle control instructions and low-rank adaptation algorithm is executed to obtain the scene arrangement model.

[0029] Optionally, determining a first vehicle control instruction based on the user intention information, the first multimodal state tag data, and a preset scenario arrangement model includes:

[0030] Inputting the user intent information and the first multimodal state label data into the scenario choreography model, so that the scenario choreography model determines a set of candidate control instructions based on the user intent information, and determines a candidate control instruction from the set of candidate control instructions based on the user intent information and the first multimodal state label data;

[0031] The candidate control instruction is converted according to an instruction conversion rule to obtain the first vehicle control instruction.

[0032] Optionally, the scenario arrangement model determines a set of candidate control instructions based on the user intent information, and determines a first vehicle control instruction from the set of candidate control instructions based on the user intent information and the first multimodal state tag data, including:

[0033] The scenario arrangement model determines the candidate control instruction set based on the objective intention information in the user intention information and the first multimodal state label data;

[0034] Based on the emotional intention information in the user intention information and the first multimodal status label data, determine the first vehicle control instruction whose priority and instruction execution method match the emotional intention information and the first multimodal status label data in the candidate control instruction set.

[0035] Optionally, converting the candidate control instruction according to an instruction conversion rule to obtain the first vehicle control instruction includes:

[0036] Performing synonym correction on the candidate control instruction to obtain a corrected candidate control instruction;

[0037] The modified candidate control instruction is converted according to a standard instruction format to obtain the first vehicle control instruction.

[0038] In a second aspect, the present application provides a method for determining a vehicle control instruction, which is applied to a vehicle computer, and the method includes:

[0039] Acquire multimodal status data collected by multiple sensors installed on the vehicle;

[0040] Performing multimodal preprocessing on the multimodal state data to obtain first multimodal state label data, and sending the first multimodal state label data to the cloud;

[0041] receiving a first vehicle control instruction from the cloud;

[0042] The first vehicle control instruction is executed.

[0043] Optionally, the multimodal status data includes: image data and / or video data collected by a camera, sound data collected by a microphone, and perception data collected by a vehicle operating status and environment perception sensor;

[0044] Performing multimodal preprocessing on the multimodal state data to obtain first multimodal state label data includes:

[0045] Identify the image data and / or video data to obtain identification text information, and generate first tag information based on the identification text information;

[0046] locating a sound source of the sound data to obtain sound source information, and converting the sound source information into second label information;

[0047] Converting the perception data into third label information;

[0048] A combination of the first tag information, the second tag information, and the third tag information is determined as the first multimodal state tag data.

[0049] Optionally, after executing the first vehicle control instruction, the method further includes:

[0050] Obtaining an instruction execution result corresponding to the first vehicle control instruction;

[0051] Send the instruction execution result to the user.

[0052] In a third aspect, the present application provides a vehicle control instruction determination device, which is applied to the cloud, and the device includes:

[0053] a first acquisition module, configured to acquire first multimodal state tag data about an in-vehicle scene from an in-vehicle computer, wherein the first multimodal state tag data is determined based on multimodal state data collected by a plurality of sensors on the vehicle;

[0054] an intention recognition module, configured to perform intention recognition on the user interaction information in the first multimodal state tag data to obtain user intention information;

[0055] A first determining module, configured to determine whether the user intention information includes an intention for vehicle control;

[0056] The second determination module is used to determine a first vehicle control instruction based on the user intention information, the first multimodal state tag data and a preset scenario arrangement model if the user intention information includes an intention for vehicle control, and send it to the vehicle terminal to control the vehicle.

[0057] In a fourth aspect, the present application provides a vehicle control instruction determination device, which is applied to a vehicle computer, and the device includes:

[0058] A second acquisition module is used to acquire multimodal state data collected by multiple sensors installed on the vehicle;

[0059] a preprocessing module, configured to perform multimodal preprocessing on the multimodal state data to obtain first multimodal state label data, and send the first multimodal state label data to the cloud;

[0060] A receiving module, configured to receive a first vehicle control instruction from the cloud;

[0061] An execution module is used to execute the first vehicle control instruction.

[0062] In a fifth aspect, the present application provides a vehicle, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0063] Memory for storing computer programs;

[0064] The processor is configured to implement any of the vehicle control instruction determination methods described in the first aspect or any of the vehicle control instruction determination methods described in the second aspect when executing a program stored in the memory.

[0065] Beneficial effects of the present invention: The embodiment of the present application can determine user intention information based on user interaction information in the first multimodal state label data by obtaining the first multimodal state label data of the in-vehicle scene. When the user intention information includes intention for vehicle control, the scene arrangement model can be used to generate a first vehicle control instruction based on the user intention information and the first multimodal state label data, thereby realizing in-depth analysis of the user's control intention for the vehicle contained in the user interaction information based on the multimodal state label data in the in-vehicle scene, improving the ability to understand user intentions, and outputting vehicle control instructions that better meet user needs, thereby enhancing the user's driving experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0067] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0068] Figure 1 A flow chart of a method for determining a vehicle control instruction provided in an embodiment of the present application;

[0069] Figure 2 for Figure 1 Flowchart of step S104;

[0070] Figure 3 for Figure 2 Flowchart of step S201;

[0071] Figure 4 for Figure 2 Flowchart of step S202;

[0072] Figure 5 A flowchart of a fine-tuning training method for a pre-trained large model provided in an embodiment of the present application;

[0073] Figure 6 Another flowchart of a fine-tuning training method for a pre-trained large model provided in an embodiment of the present application;

[0074] Figure 7 for Figure 5 Flowchart of step S502;

[0075] Figure 8 A flowchart of another method for determining vehicle control instructions provided in an embodiment of the present application;

[0076] Figure 9 A structural diagram of a vehicle control instruction determination device provided in an embodiment of the present application;

[0077] Figure 10 A structural diagram of another vehicle control instruction determination device provided in an embodiment of the present application;

[0078] Figure 11 A structural diagram of a vehicle provided in an embodiment of the present application. DETAILED DESCRIPTION

[0079] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0080] Existing in-vehicle voice command recognition and processing systems still have many shortcomings and challenges in practical applications, specifically in the following aspects:

[0081] 1. Single-modality reliance: Traditional in-vehicle speech recognition systems rely primarily on single-modality speech data, limiting their processing capabilities for complex in-vehicle scenarios. Single-modality systems are susceptible to noise, fuzzy speech, and external interference, resulting in low speech recognition accuracy and failing to meet high user expectations.

[0082] 2. Inadequate multitasking capabilities: In a car, user voice commands often involve multitasking, such as controlling navigation, music playback, and the air conditioning system simultaneously. Existing technologies exhibit limitations in understanding and executing these complex commands, making it difficult to provide a smooth user experience.

[0083] 3. Poor scenario adaptability: The in-vehicle environment is dynamic, and the understanding requirements for voice commands vary in different driving scenarios. Existing voice recognition systems lack sufficient adaptability to complex and changing in-vehicle scenarios, resulting in suboptimal voice command processing.

[0084] 4. Inadequate data fusion: Although some advanced systems have begun experimenting with multimodal data fusion, efficient integration of multimodal data remains a challenge in practical applications. Existing methods often fail to fully leverage the complementary strengths of each modality when integrating multimodal information such as voice, image, and sensor data, limiting overall system performance.

[0085] 5. Issues with command conversion accuracy and reliability: When processing and converting voice commands, large models may return inaccurate or erroneous commands due to data noise, model bias, and other factors. This can lead to serious operational errors in vehicle scenarios. Therefore, existing technologies lack an effective verification system to validate and correct the commands returned by large models and ensure the accuracy and reliability of command conversion.

[0086] Due to the above-mentioned deficiencies, traditional in-vehicle voice command recognition and processing systems have difficulty in accurately converting command data into the in-vehicle operations desired by users, reducing the efficiency of user control of the vehicle. To this end, the embodiments of the present application provide a method, device, and vehicle for determining vehicle control commands, which can convert comprehensive voice commands involving complex scene arrangement requirements into vehicle control commands in an in-vehicle environment through the fusion of multimodal data and fine-tuning of a large model. For example, when a user issues a comprehensive voice command intended to simultaneously control the air conditioning, navigation, and entertainment systems, the vehicle control command is automatically determined based on the comprehensive voice command. The vehicle control command can simultaneously control the air conditioning, navigation, and entertainment systems, thereby improving the system's voice command processing capabilities in complex in-vehicle scenarios. At the same time, a verification system is introduced to verify and correct the commands returned by the large model, ensuring the accuracy and reliability of the command conversion, improving the user experience, and realizing efficient and accurate conversion of voice commands in complex in-vehicle scenarios, thereby improving the intelligence level of the in-vehicle system and the user experience.

[0087] The embodiment of the present application provides a method for determining a vehicle control instruction, which is applied to the cloud. Figure 1 As shown, the method includes:

[0088] Step S101, obtaining first multimodal state tag data about an in-vehicle scene from an in-vehicle terminal;

[0089] In an embodiment of the present application, the first multimodal state label data is determined based on multimodal state data collected by multiple sensors on the vehicle, and the multimodal state data includes: image data and / or video data collected by a camera, sound data collected by a microphone, and perception data collected by a vehicle operation status and environmental perception sensor, etc. Furthermore, the multimodal state data includes: sound data, environmental state data, vehicle body component state data, character posture data, etc.

[0090] Step S102: performing intent recognition on the user interaction information in the first multimodal state tag data to obtain user intent information;

[0091] In the embodiment of the present application, user interaction information includes: user voice signals input by the user through voice, user text information input by the user on the control panel, user action information of user-specified actions collected by the camera, etc. In actual applications, users can also input user interaction information through other methods, which are not limited here.

[0092] In one embodiment of the present application, the user interaction information includes: a user voice signal; step S102 performs intent recognition on the user interaction information in the first multimodal state tag data to obtain user intent information, including:

[0093] Intention recognition is performed on the user voice signal to obtain objective intention information and emotional intention information; and the objective intention information and the emotional intention information are determined as the user intention information.

[0094] When the user's voice signal is used for intention recognition, feature extraction can be performed. The system extracts key features from the voice signal to obtain emotional intention information. Emotional intention information includes features such as pitch, speaking speed, and emotion. These features help to accurately capture the driver's emotions and intentions. At the same time, the user's voice signal needs to be processed. Advanced voice processing algorithms are used to clean and enhance the original voice signal, and then the objective intention information in the voice signal is identified to ensure the accuracy and reliability of the analysis.

[0095] Step S103, determining whether the user intention information includes an intention for vehicle control;

[0096] In actual applications, users in a vehicle may make sounds when talking to each other, talking to others through mobile terminals such as mobile phones, listening to the radio, playing music, etc. Therefore, the user interaction information in the multimodal state label data uploaded by the vehicle may contain intentions other than vehicle control. Therefore, it is necessary to further determine whether the user intention information contains intentions for vehicle control.

[0097] In one embodiment of the present application, step S103 determines whether the user intention information includes an intention for vehicle control, including:

[0098] Determine whether the user intention information includes preset keywords associated with the vehicle; if the user intention information includes preset keywords associated with the vehicle, determine that the user intention information includes intention for vehicle control.

[0099] In this embodiment, the preset keywords may include, for example: the name of the vehicle-side interactive assistant (such as: Xiao An), and / or the actions that need to be controlled to be performed by the vehicle (such as: open, close, windows, doors, play, music, increase, decrease, temperature, etc.).

[0100] Step S104: If the user intention information includes an intention for vehicle control, a first vehicle control instruction is determined based on the user intention information, the first multimodal state tag data, and a preset scenario arrangement model, and sent to the vehicle terminal to control the vehicle.

[0101] In this step, if the user intention information includes intention for vehicle control, the user intention information and the first multimodal state label data can be input into the scenario orchestration model so that the scenario orchestration model outputs a first vehicle control instruction, and then the first vehicle control instruction is sent to the vehicle-mounted terminal so that the vehicle-mounted terminal controls the vehicle to execute the first vehicle control instruction.

[0102] If the user intention information does not include an intention for vehicle control, the process ends.

[0103] In one embodiment of the present application, step S104 determines a first vehicle control instruction based on the user intention information, the first multimodal state tag data and a preset scenario arrangement model, such as Figure 2 As shown, including:

[0104] Step S201: inputting the user intent information and the first multimodal state tag data into the scenario choreography model, so that the scenario choreography model determines a set of candidate control instructions based on the user intent information, and determines a candidate control instruction from the set of candidate control instructions based on the user intent information and the first multimodal state tag data;

[0105] The scenario arrangement model not only relies on voice data, but also combines other modal information, such as the driver's facial expressions and gestures. By analyzing the multimodal state label data, the model can determine a set of candidate control instructions, and determine the candidate control instructions in the candidate control instruction set based on the user intention information and the first multimodal state label data.

[0106] In one embodiment of the present application, the scenario arrangement model in step S201 determines a set of candidate control instructions based on the user intention information, and determines a first vehicle control instruction from the set of candidate control instructions based on the user intention information and the first multimodal state tag data, such as Figure 3 As shown, including:

[0107] Step S301: The scenario arrangement model determines the candidate control instruction set based on the objective intention information in the user intention information and the first multimodal state label data;

[0108] In this step, the scenario arrangement model determines the candidate control instruction set based on the textual content without emotional connotation in the user intention information, i.e., the objective intention information and the first multimodal state label data. The candidate control instruction set may include multiple candidate control instructions with different priorities and instruction execution methods.

[0109] Step S302: Based on the emotional intention information in the user intention information and the first multimodal state label data, determine the first vehicle control instruction whose priority and instruction execution method match the emotional intention information and the first multimodal state label data in the candidate control instruction set.

[0110] The scenario orchestration model, based on emotional intent information and the first multimodal state label data, can accurately determine the user's current state. For example, through facial expressions and gestures, the model can identify whether the driver is nervous, relaxed, or focused. Based on the driver's state, the system can dynamically adjust the priority and execution of instructions. For example, if the driver is detected to be nervous, the system may prioritize soothing music playback or autonomous driving functions to reduce driver stress.

[0111] Step S202: convert the candidate control instruction according to an instruction conversion rule to obtain the first vehicle control instruction.

[0112] In this embodiment, the system converts the recognized voice commands through a set of command conversion rules to ensure accurate execution of in-vehicle operations.

[0113] In one embodiment of the present application, step S202 converts the candidate control instruction according to the instruction conversion rule to obtain the first vehicle control instruction, such as Figure 4 As shown, including:

[0114] Step S401, performing synonym correction on the candidate control instruction to obtain a corrected candidate control instruction;

[0115] First, to optimize the matching process, the cloud-based system can handle subtle differences in voice commands. Whether the driver uses expressions such as "Open the driver's door" or "Open the main driver's door," the system will recognize and match the standard command "DriverDoorOpen." Through this correction mechanism, the system can effectively handle a variety of different voice expressions, ensuring accurate and consistent operation.

[0116] Step S402: convert the corrected candidate control instruction according to a standard instruction format to obtain the first vehicle control instruction.

[0117] The system is also configured with a standard instruction format to ensure consistency and accuracy of operations. For example:

[0118] Example 1: The standard command for opening the driver's door is DriverDoorOpen.

[0119] Example 2: The standard instruction for playing music is PlayMusic.

[0120] In actual use cases, the driver's voice commands may differ from predefined standard commands. The scenario orchestration model may return commands such as "Driver, open the door." The system uses a command matching mechanism to compare and match these voice commands with predefined standard commands. For example, "Driver, open the door" is matched to the standard command DriverDoorOpen.

[0121] For example, when the user says "Driver, please remember to bring your phone when you get off the car," the system generates corresponding trigger, state, and execution commands based on the multimodal state tag data:

[0122] Trigger data: The main driving door is opened

[0123] Status data: Driver is occupied

[0124] Execution data: Voice broadcast, remember to bring your mobile phone

[0125] Text matching logic: Matches words or phrases in the trigger data, status data, and execution data with strings in the standard dictionary. For example, if the trigger data "driver's door" matches the "driver's door" in the standard dictionary, that is, "driver's door" = "driver's door". If the match is successful, the first successfully matched data (the first data refers to the data in the trigger data, status data, and execution data that successfully matches the standard dictionary) is converted into text to obtain instructions that can be used by the vehicle computer. If the match fails, the second unmatched data (the second data refers to the data in the trigger data, status data, and execution data that failed to match the standard dictionary) is processed through synonym matching logic.

[0126] Synonym matching logic: A synonym library can be pre-configured, containing multiple standard words and synonyms corresponding to each standard word. The second data is matched against the multiple synonyms in the synonym library. If a match is successful, the synonyms in the second data are corrected to the standard words corresponding to the synonyms. For example, "driver's door" = "main driver's door", "driver's side door" = "main driver's door". The successfully matched third data is then converted into text to obtain instructions that can be used by the vehicle computer. If the match fails, the unmatched fourth data is processed through text vector matching logic.

[0127] Text vector matching logic: The fourth data, such as the word "main driver's door," is converted into a word vector. The word vector is matched with a standard word vector library to return a similarity matching result. If the similarity matching result indicates a successful text vector match, the successfully matched fifth data is converted into text to obtain instructions that can be used by the vehicle computer.

[0128] Vehicle computer commands: Convert text into commands that can be used by the vehicle computer.

[0129] For example: "DriverDoorOpen" = "DriverDoorOpen", "DriverDoorClose" = "DriverDoorClose".

[0130] Finally, if the text vector matching fails: the matching failure indicates that the vehicle computer does not support the current command and will not issue the command.

[0131] Through the above steps, the cloud can achieve accurate recognition and conversion of voice commands, greatly improving the user experience of in-car voice interaction. It not only ensures the accurate execution of commands, but also can flexibly respond to different voice expressions, thereby providing more intelligent and humane services.

[0132] After generating a first vehicle control command and sending it to the vehicle computer to control the vehicle, the system can verify the accuracy of the generated first vehicle control command and the executed action. Furthermore, this ensures that there are no omissions or errors in the command conversion process. For example, after receiving the "open sunroof" command, the system actually executes the corresponding action. Furthermore, by simulating different usage scenarios, the system's response speed and accuracy to commands can be comprehensively tested.

[0133] The embodiment of the present application can determine user intention information based on user interaction information in the first multimodal state tag data by obtaining the first multimodal state tag data of the in-vehicle scene. When the user intention information includes intention for vehicle control, the scene arrangement model can be used to generate a first vehicle control instruction based on the user intention information and the first multimodal state tag data, thereby realizing in-depth analysis of the user's control intention for the vehicle contained in the user interaction information based on the multimodal state tag data in the in-vehicle scene, improving the ability to understand user intentions, and outputting vehicle control instructions that better meet user needs, thereby enhancing the user's driving experience.

[0134] In another embodiment of the present application, the scene arrangement model is obtained by fine-tuning the pre-trained large model, and the fine-tuning training method of the pre-trained large model is as follows: Figure 5 As shown, including:

[0135] Step S501: obtaining model fine-tuning prompt words, user interaction information from multiple vehicle terminals, second multimodal state label data for multiple vehicle scenarios, and second vehicle control instructions labeled for the multiple vehicle scenarios respectively;

[0136] In an embodiment of the present application, the second multimodal state label data is also obtained after multimodal preprocessing on the vehicle side. Each group of second multimodal state label data is determined based on the multimodal state data collected by multiple sensors on each vehicle. The multimodal state data includes: image data and / or video data collected by the camera, sound data collected by the microphone, and perception data collected by the vehicle operation status and environmental perception sensors, etc. Furthermore, the multimodal state data includes: sound data, environmental status data, vehicle body component status data, character posture data, etc.

[0137] The second vehicle control instructions are annotated by the user or staff for the in-vehicle scene of the second multimodal state label data. As an example, after the user or staff inputs user interaction information, if the vehicle-side cannot make a response that meets the user's expectations, the second vehicle control instructions manually input by the user can be obtained for fine-tuning training of the pre-trained large model.

[0138] User interaction information from multiple vehicle terminals, second multimodal state label data about the vehicle scene, and second vehicle control instructions labeled for multiple vehicle scenes can use the following JSON format to express the input and output relationship of multimodal data:

[0139] {

[0140] "instruction": "prompt model fine-tuning prompt words";

[0141] "input": "Text spoken by the user, multimodal data labels",

[0142] "output": "{

[0143] Trigger name, trigger condition, trigger parameters;

[0144] Status: [{status name, status condition, status parameters}]

[0145] Execution: [execution object, execution condition, execution parameters]

[0146] }"}

[0147] Example

[0148] For example, when the user says "I open the sunroof and air conditioning when I get in the car", the system generates the corresponding state and execution command based on the multimodal state label data:

[0149] Trigger condition: driver opens the door, trigger parameter: open the door

[0150] Multimodal state labels: Person: Hostess; Position: Co-pilot

[0151] State generation: State name: character, state condition: equal (default is the same as the actual situation), state parameter: hostess.

[0152] State name: Position, State condition: Equal, State parameter: Co-pilot.

[0153] To ensure the accuracy and reliability of the system, the quality of the training data must be rigorously tested and verified. For example, prompt word engineering can be used to verify the logical consistency of the training data. Furthermore, by setting a series of predefined prompt words and scenarios, the system's response in different situations can be tested. The system's generated actions can be checked to ensure they are consistent with expectations, such as correctly opening the sunroof and air conditioning when the user detects "hostess." By continuously adjusting prompt words and training data, the system can ensure that it responds logically in a variety of complex scenarios.

[0154] The scenario orchestration model, trained using the previously verified and validated training data, can identify different multimodal tags (such as visual and audio) and generate and execute corresponding actions to enhance the user's driving experience. For example, when the system detects a "hostess" sitting in the passenger seat, it automatically generates the corresponding status conditions and opens the sunroof and air conditioning to ensure a comfortable environment and meet the user's personalized needs. This intelligent operation, based on deep learning and analysis of multimodal data, ensures attentive service in various driving scenarios.

[0155] In one embodiment of the present application, after obtaining user interaction information from multiple vehicle terminals, second multimodal state label data about the vehicle scene, and second vehicle control instructions respectively labeled for multiple vehicle scenes in step S501, the method is as follows: Figure 6 As shown, it also includes:

[0156] Step S601, obtaining a verification prompt word;

[0157] As an example, a verification prompt word is as follows:

[0158] Background information:

[0159] You are a car user experience judge.

[0160] Task Description

[0161] Based on the request made by the user and the actual actions performed by the vehicle, determine whether it is reasonable and make a judgment.

[0162] Input format:

[0163] {

[0164] Input: user interaction information and multimodal state label data input by the user;

[0165] output: the actual action performed by the vehicle;

[0166] }

[0167] Output format:

[0168] {

[0169] Judgement: Reasonable / Unreasonable

[0170] }

[0171] Sample output:

[0172] {

[0173] Verdict: Reasonable

[0174] }

[0175] Step S602: Inputting the verification prompt word, any user interaction information, the second multimodal state tag data, and the second vehicle control instruction corresponding to the second multimodal state tag data into a preset large model, so that the preset large model performs rationality verification on the second vehicle control instruction and obtains a rationality verification result;

[0176] In this step, the verification prompt word, one set of user interaction information, the second multimodal state label data and the second vehicle control instruction corresponding to the second multimodal state label data can be input into the preset large model, so that the preset large model verifies whether the second vehicle control instruction is reasonable and obtains a rationality verification result.

[0177] Step S603: If the rationality verification result meets the verification pass conditions, the pre-trained large model is fine-tuned based on the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, multiple second vehicle control instructions and low-rank adaptation algorithm to obtain the scene arrangement model.

[0178] Step S502: Fine-tune the pre-trained large model based on the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, multiple second vehicle control instructions and a low-rank adaptation algorithm to obtain the scene arrangement model.

[0179] In the embodiment of the present application, the fine-tuning process of the pre-trained large model mainly adopts the low-rank adaptation algorithm (LoRA) method, which introduces a small number of trainable parameters to achieve efficient adjustment of the large model to adapt to specific task requirements.

[0180] As an example, the prompt words for a pre-trained large model during fine-tuning training are as follows:

[0181] Background information:

[0182] You are a car assistant, and users may mention various needs during driving, such as adjusting the air conditioning, navigation, making and receiving calls, etc. You need to understand the user's intentions and translate them into specific scenarios and parameters.

[0183] Task Description:

[0184] Based on the user's description, create a reasonable car scenario and give it a name of no more than 10 characters. You need to extract information such as the triggering event, the preceding state, and the executed action from the user's description.

[0185] Parameter condition list:

[0186] Composite interval, not equal to, equal to, greater than or equal to, greater than, less than or equal to, less than

[0187] Input format:

[0188] User description. For example: "Initiate navigation and the estimated arrival distance is 2 kilometers, remind me to slow down."

[0189] Output format:

[0190] {

[0191] "Scene Name": "Scene Name",

[0192] "Trigger Name": "Trigger Event",

[0193] "Trigger Parameters": "Trigger Parameters",

[0194] "Trigger Condition": "Trigger Condition",

[0195] "Trigger parameter unit": "Trigger parameter unit",

[0196] "Pre-state array": [

[0197] {

[0198] "Status Name": "Status Condition",

[0199] "Status Parameters": "Status Parameters",

[0200] "Status parameter unit": "Status parameter unit",

[0201] "Status Condition": "Status Condition"

[0202] }

[0203] ],

[0204] "Execution action array": [

[0205] {

[0206] "Execution Name": "Execute Action",

[0207] "Execution parameters": "Execution parameters",

[0208] "Execution Condition": "Execution Condition",

[0209] "Execution parameter unit": "Execution parameter unit"

[0210] } ]

[0212] }

[0213] Example:

[0214] enter:

[0215] "Initiate navigation and estimate the arrival distance to 2 kilometers and remind me to slow down"

[0216] Output:

[0217] {

[0218] "Scene Name": "Navigation Reminder",

[0219] "Trigger Name": "Map",

[0220] "Trigger Parameters": "Initiate Navigation",

[0221] "Trigger Condition": "Equal",

[0222] "Trigger parameter unit": "",

[0223] "Pre-state array": [

[0224] {

[0225] "Status Name": "Distance",

[0226] "Status Parameter": "2",

[0227] "Status parameter unit": "km",

[0228] "Status Condition": "Equal"

[0229] }

[0230] ],

[0231] "Execution action array": [

[0232] {

[0233] "Execution Name": "Reminder",

[0234] "Execution parameters": "Vehicle slow down",

[0235] "Execution Condition": "Equal",

[0236] "Execution parameter unit": ""

[0237] } ]

[0239] }

[0240] In one embodiment of the present application, step S502 fine-tunes the pre-trained large model based on the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, multiple second vehicle control instructions and low-rank adaptation algorithm to obtain the scene arrangement model, such as Figure 7 As shown, including:

[0241] Step S701, obtaining a weight matrix including a plurality of first model parameters in a pre-trained large model, wherein the number of first model parameters in the weight matrix is ​​less than the number of remaining model parameters in the pre-trained large model, and the remaining model parameters are model parameters in the pre-trained large model other than the first model parameters in the weight matrix;

[0242] Step S702, decomposing the weight matrix into a first low-rank matrix and a second low-rank matrix;

[0243] Step S703: training the first low-rank matrix and the second low-rank matrix using the model fine-tuning prompt words, the plurality of user interaction information, the plurality of second multimodal state label data, and the plurality of second vehicle control instructions to obtain a first fine-tuning parameter matrix and a second fine-tuning parameter matrix;

[0244] Step S704: Use the second model parameters in the first fine-tuning parameter matrix and the second fine-tuning parameter matrix to replace the corresponding first model parameters in the pre-trained large model to obtain the scene arrangement model.

[0245] In the embodiment of the present application, the main goal of the fine-tuning training process of the pre-trained model is to reduce the number of parameters that need to be trained during the fine-tuning process, thereby reducing computing costs and storage requirements. In the embodiment of the present application, the weights of the pre-trained model are fine-tuned by introducing a low-rank matrix, rather than directly adjusting all weight parameters. The fine-tuning training process is as follows:

[0246] Pre-trained model freezing: During fine-tuning, most of the weight parameters of the pre-trained model remain unchanged (frozen), and only a small number of parameters are adjusted. These small number of parameters are represented by low-rank matrices;

[0247] Low-rank matrix: Assuming that some weight matrix in the pre-trained model is (W), LoRA decomposes it into two low-rank matrices (A) and (B), where the ranks of (A) and (B) are low, so the number of parameters that need to be trained is greatly reduced.

[0248] Fine-tuning process: During fine-tuning, only the low-rank matrices (A) and (B) are trained, while the other weights of the pre-trained model remain unchanged. This significantly reduces computing resource consumption while ensuring the fine-tuning effect.

[0249] The scene orchestration model of this application receives voice, image and sensor data simultaneously at the input layer of the model, enabling it to understand complex information in the vehicle scene across modalities; by adjusting the model parameters, it performs well in multi-task processing, especially in the accurate recognition of voice commands and scene understanding; the fine-tuned model can more effectively process and fuse information from different modalities, significantly improving its performance in complex vehicle environments.

[0250] In another embodiment of the present application, a vehicle control instruction determination method is also provided, which is applied to the vehicle terminal, such as Figure 8 As shown, the method includes:

[0251] Step S801, acquiring multimodal status data collected by multiple sensors installed on the vehicle;

[0252] In the embodiment of the present application, the multimodal status data includes: image data and / or video data collected by the camera, sound data collected by the microphone, and perception data collected by the vehicle operation status and environment perception sensor.

[0253] The system can collect multimodal status data in the vehicle environment in real time through on-board sensors, cameras, microphones and other devices. These data include but are not limited to: voice data: collecting the driver's voice commands through the microphone; image and video data: capturing the interior environment, the driver's facial expressions and gestures and other information through the in-vehicle camera; sensor data: obtaining speed, acceleration, steering wheel angle, ambient temperature and other information through various sensors in the vehicle.

[0254] Step S802: performing multimodal preprocessing on the multimodal state data to obtain first multimodal state label data, and sending the first multimodal state label data to the cloud;

[0255] In one embodiment of the present application, step S802 performs multimodal preprocessing on the multimodal state data to obtain first multimodal state label data, including: identifying the image data and / or video data to obtain identification text information, and generating first label information based on the identification text information; locating the sound source of the sound data to obtain sound source information, and converting the sound source information into second label information; converting the perception data into third label information; and determining the combination of the first label information, the second label information and the third label information as the first multimodal state label data.

[0256] In this implementation, multimodal state data can be preprocessed (e.g., noise filtering, image enhancement, and signal calibration) to ensure data accuracy and consistency. To enhance the data's application value while protecting user privacy, the processed data is converted into labeled data. For example, passenger identification: The system records relevant image information when recognizing the hostess boarding the vehicle and generates a "hostess" label through image matching during subsequent rides. Voice source location: The system can identify the source direction of voice commands (e.g., driver's seat, front passenger seat, or rear seat) and generate corresponding labels. Environmental data tags: Data such as current vehicle speed and interior temperature are converted into tagged information for further analysis and response.

[0257] The system is also capable of managing multi-user scenarios. When multiple passengers are in the vehicle at the same time, the system uses voice and image recognition technology to distinguish between different users and provide personalized services based on their identities. For example, after recognizing the voice commands of users in different seats, the system will adjust settings such as audio volume, playback content, or seat heating to meet the needs of each passenger.

[0258] Step S803, receiving a first vehicle control instruction from the cloud;

[0259] Step S804: execute the first vehicle control instruction.

[0260] In the embodiment of the present application, the verified instructions will be sent to the various subsystems of the vehicle system for execution to ensure accurate and efficient operation. The specific execution process is as follows:

[0261] 1. System: Executes path planning and navigation instructions to ensure that the vehicle follows the optimal path and provides accurate navigation services.

[0262] 2. Air Conditioning System: Adjusts the interior temperature and air speed based on commands to provide a comfortable interior environment. For example, if a "turn up the temperature" command is received, the system automatically raises the interior temperature.

[0263] 3. Entertainment System: Play music, adjust volume, or switch media based on your commands. Whether selecting a specific song or adjusting the volume, the system responds quickly and executes the corresponding action.

[0264] 4. Follow-up Conversations: The system supports dynamic adjustments and further actions based on subsequent conversations. For example, users can use voice commands at any time to request that the system adjust current settings or submit new requests. The system will respond and execute the new instructions promptly, ensuring a continuously optimized user experience.

[0265] In one embodiment of the present application, after executing the first vehicle control instruction, the method further includes:

[0266] Obtaining an instruction execution result corresponding to the first vehicle control instruction; and sending the instruction execution result to a user.

[0267] During command execution, the system monitors the results of each subsystem in real time and provides timely feedback to the user. For example, after a navigation command is executed, the system notifies the user of the updated route; after the air conditioning temperature is adjusted, the system displays the current temperature setting. This real-time feedback mechanism ensures that users are informed of the results of their operations, improving the transparency and satisfaction of the interactive experience.

[0268] Through the above steps, the in-vehicle system of this embodiment can achieve accurate execution and real-time feedback of voice commands, providing users with an intelligent and personalized in-vehicle experience.

[0269] The present application also provides the following overall optimization directions:

[0270] 1. Response speed: Optimize the data transmission and instruction processing speed between modules to ensure that the system's response speed meets real-time requirements.

[0271] 2. User interaction: By analyzing user usage habits and preferences, we continuously optimize the system's interactive interface and feedback mechanism to improve user experience.

[0272] Through these implementations, the present invention can effectively improve the efficiency of voice command conversion and execution in vehicle-mounted scenarios, significantly improving the user's driving experience.

[0273] In another embodiment of the present application, a vehicle control instruction determination device is provided, which is applied to the cloud. Figure 9 As shown, the device includes:

[0274] A first acquisition module 11 is configured to acquire first multimodal state tag data about an in-vehicle scene from an in-vehicle terminal, wherein the first multimodal state tag data is determined based on multimodal state data collected by a plurality of sensors on the vehicle;

[0275] an intention recognition module 12, configured to perform intention recognition on the user interaction information in the first multimodal state tag data to obtain user intention information;

[0276] A first determining module 13 is configured to determine whether the user intention information includes an intention for vehicle control;

[0277] The second determination module 14 is used to determine a first vehicle control instruction based on the user intention information, the first multimodal state tag data and a preset scenario arrangement model if the user intention information includes an intention for vehicle control, and send it to the vehicle terminal to control the vehicle.

[0278] In another embodiment of the present application, a vehicle control instruction determination device is provided, which is applied to the vehicle terminal, such as Figure 10 As shown, the device includes:

[0279] The second acquisition module 21 is used to acquire multimodal state data collected by multiple sensors installed on the vehicle;

[0280] A preprocessing module 22 is configured to perform multimodal preprocessing on the multimodal state data to obtain first multimodal state label data, and send the first multimodal state label data to the cloud;

[0281] A receiving module 23 is configured to receive a first vehicle control instruction from the cloud;

[0282] The execution module 24 is configured to execute the first vehicle control instruction.

[0283] In yet another embodiment of the present application, a vehicle is provided, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0284] Memory for storing computer programs;

[0285] The processor is used to implement any of the aforementioned vehicle control instruction determination methods applied to the cloud or any of the aforementioned vehicle control instruction determination methods applied to the vehicle terminal when executing the program stored in the memory.

[0286] In the vehicle provided by an embodiment of the present invention, the processor obtains first multimodal status label data of the in-vehicle scene by executing a program stored in the memory, and can determine user intention information based on user interaction information in the first multimodal status label data. When the user intention information includes an intention for vehicle control, the scene arrangement model can be used to generate a first vehicle control instruction based on the user intention information and the first multimodal status label data, thereby realizing in-depth analysis of the user's control intention for the vehicle contained in the user interaction information based on the multimodal status label data in the in-vehicle scene, improving the ability to understand the user's intention, and outputting vehicle control instructions that better meet user needs, thereby enhancing the user's driving experience.

[0287] The communication bus 1140 mentioned in the above vehicle can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The communication bus 1140 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0288] The communication interface 1120 is used for communication between the vehicle and other devices.

[0289] The memory 1130 may include a random access memory (RAM) or a non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.

[0290] The above-mentioned processor 1110 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0291] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0292] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for determining a vehicle control instruction, characterized in that: Applied to the cloud, the method includes: Acquire first multimodal state tag data about an in-vehicle scene from the vehicle computer, where the first multimodal state tag data is determined based on multimodal state data collected by multiple sensors on the vehicle; Performing intent recognition on the user interaction information in the first multimodal state tag data to obtain user intent information; determining whether the user intention information includes an intention for vehicle control; If the user intention information includes an intention for vehicle control, a first vehicle control instruction is determined based on the user intention information, the first multimodal state label data and a preset scenario orchestration model, and sent to the vehicle terminal to control the vehicle; the scenario orchestration model determines a set of candidate control instructions based on the objective intention information in the user intention information and the first multimodal state label data; based on the emotional intention information in the user intention information and the first multimodal state label data, a first vehicle control instruction whose priority and instruction execution method match the emotional intention information and the first multimodal state label data is determined in the candidate control instruction set.

2. The vehicle control instruction determination method according to claim 1, characterized in that: The user interaction information includes: user voice signal; Performing intent recognition on the user interaction information in the first multimodal state tag data to obtain user intent information includes: Performing intent recognition on the user voice signal to obtain objective intent information and emotional intent information; The objective intention information and the emotional intention information are determined as the user intention information.

3. The vehicle control instruction determination method according to claim 1, characterized in that: Determining whether the user intention information includes an intention for vehicle control includes: determining whether the user intention information includes a preset keyword associated with the vehicle; If the user intention information includes a preset keyword associated with the vehicle, it is determined that the user intention information includes an intention for vehicle control.

4. The vehicle control instruction determination method according to claim 1, characterized in that: The scene arrangement model is obtained by fine-tuning the pre-trained large model. The fine-tuning training method of the pre-trained large model includes: Obtaining model fine-tuning prompt words, user interaction information from multiple vehicle terminals, second multimodal state label data about multiple vehicle scenarios, and second vehicle control instructions labeled for the multiple vehicle scenarios; The pre-trained large model is fine-tuned based on the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, multiple second vehicle control instructions and a low-rank adaptation algorithm to obtain the scene arrangement model.

5. The vehicle control instruction determination method according to claim 4, characterized in that: Fine-tuning the pre-trained large model based on the model fine-tuning prompt words, multiple second multimodal state label data, multiple user interaction information, multiple second vehicle control instructions, and a low-rank adaptation algorithm to obtain the scene arrangement model, including: Obtaining a weight matrix including a plurality of first model parameters in a pre-trained large model, wherein the number of the first model parameters in the weight matrix is ​​less than the number of remaining model parameters in the pre-trained large model, and the remaining model parameters are model parameters in the pre-trained large model other than the first model parameters in the weight matrix; Decomposing the weight matrix into a first low-rank matrix and a second low-rank matrix; Training the first low-rank matrix and the second low-rank matrix using the model fine-tuning prompt words, the plurality of user interaction information, the plurality of second multimodal state label data, and the plurality of second vehicle control instructions to obtain a first fine-tuning parameter matrix and a second fine-tuning parameter matrix; The scene arrangement model is obtained by replacing the corresponding first model parameters in the pre-trained large model with the second model parameters in the first fine-tuning parameter matrix and the second fine-tuning parameter matrix.

6. The vehicle control instruction determination method according to claim 4, characterized in that: After obtaining user interaction information from multiple vehicle terminals, second multimodal state label data about the vehicle scene, and second vehicle control instructions labeled for the multiple vehicle scenes, the method further includes: Get the verification prompt word; Inputting the verification prompt word, any one of the user interaction information, the second multimodal state label data, and the second vehicle control instruction corresponding to the second multimodal state label data into a preset large model, so that the preset large model performs rationality verification on the second vehicle control instruction and obtains a rationality verification result; If the rationality verification result meets the verification pass conditions, the step of fine-tuning the pre-trained large model based on the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, multiple second vehicle control instructions and low-rank adaptation algorithm is executed to obtain the scene arrangement model.

7. The vehicle control instruction determination method according to claim 1, characterized in that: Determining a first vehicle control instruction based on the user intention information, the first multimodal state tag data, and a preset scenario arrangement model includes: Inputting the user intent information and the first multimodal state label data into the scenario choreography model, so that the scenario choreography model determines a set of candidate control instructions based on the user intent information, and determines a candidate control instruction from the set of candidate control instructions based on the user intent information and the first multimodal state label data; The candidate control instruction is converted according to an instruction conversion rule to obtain the first vehicle control instruction.

8. The vehicle control instruction determination method according to claim 7, characterized in that: Converting the candidate control instruction according to an instruction conversion rule to obtain the first vehicle control instruction includes: Performing synonym correction on the candidate control instruction to obtain a corrected candidate control instruction; The modified candidate control instruction is converted according to a standard instruction format to obtain the first vehicle control instruction.

9. A method for determining a vehicle control instruction according to any one of claims 1 to 8, characterized in that: Applied to the vehicle terminal, the method includes: Acquire multimodal status data collected by multiple sensors installed on the vehicle; Performing multimodal preprocessing on the multimodal state data to obtain first multimodal state label data, and sending the first multimodal state label data to the cloud; receiving a first vehicle control instruction from the cloud; The first vehicle control instruction is executed.

10. The vehicle control instruction determination method according to claim 9, characterized in that: The multimodal status data includes: image data and / or video data collected by a camera, sound data collected by a microphone, and perception data collected by a vehicle operating status and environment perception sensor; Performing multimodal preprocessing on the multimodal state data to obtain first multimodal state label data includes: Identify the image data and / or video data to obtain identification text information, and generate first tag information based on the identification text information; locating a sound source of the sound data to obtain sound source information, and converting the sound source information into second label information; Converting the perception data into third label information; A combination of the first tag information, the second tag information, and the third tag information is determined as the first multimodal state tag data.

11. The vehicle control instruction determination method according to claim 10, characterized in that: After executing the first vehicle control instruction, the method further includes: Obtaining an instruction execution result corresponding to the first vehicle control instruction; Send the instruction execution result to the user.

12. A vehicle control instruction determination device, characterized in that: Applied to the cloud, the device includes: a first acquisition module, configured to acquire first multimodal state tag data about an in-vehicle scene from an in-vehicle computer, wherein the first multimodal state tag data is determined based on multimodal state data collected by a plurality of sensors on the vehicle; an intention recognition module, configured to perform intention recognition on the user interaction information in the first multimodal state tag data to obtain user intention information; A first determining module, configured to determine whether the user intention information includes an intention for vehicle control; The second determination module is used to determine a first vehicle control instruction based on the user intention information, the first multimodal state label data and a preset scenario arrangement model if the user intention information includes an intention for vehicle control, and send it to the vehicle terminal to control the vehicle; the scenario arrangement model determines a set of candidate control instructions based on the objective intention information in the user intention information and the first multimodal state label data; based on the emotional intention information in the user intention information and the first multimodal state label data, determine the first vehicle control instruction whose priority and instruction execution method match the emotional intention information and the first multimodal state label data in the candidate control instruction set.

13. A vehicle control instruction determination device according to claim 12, characterized in that: Applied to the vehicle terminal, the device includes: A second acquisition module is used to acquire multimodal state data collected by multiple sensors installed on the vehicle; a preprocessing module, configured to perform multimodal preprocessing on the multimodal state data to obtain first multimodal state label data, and send the first multimodal state label data to the cloud; A receiving module, configured to receive a first vehicle control instruction from the cloud; An execution module is used to execute the first vehicle control instruction.

14. A vehicle, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; The processor is configured to implement the vehicle control instruction determination method described in any one of claims 1 to 8 or the vehicle control instruction determination method described in any one of claims 9 to 11 when executing the program stored in the memory.

Citation Information

Patent Citations

  • Multi-modal data collaborative man-machine interaction method and system and vehicle-mounted multimedia device

    CN111737670A

  • Vehicle interaction method and device, electronic equipment, storage medium and vehicle

    CN117235320A

  • Terminal equipment and corpus data generation method

    CN118349844A

  • Vehicle control method, device, equipment and medium

    CN119811396A

  • Human-computer interaction method, device, equipment and medium

    CN120071971A