Vehicle control instruction determination method and device and vehicle
By obtaining multimodal state label data of on-board scenarios and identifying user intentions, and using scene orchestration models to generate vehicle control instructions, the problem of traditional on-board voice recognition systems accurately converting instructions in complex on-board environments is solved, and the user experience is improved.
Patent Information
- Application Number
- CN202510565898.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
In the changing vehicle environment, traditional vehicle voice recognition systems are difficult to accurately convert command data into the on-board operation that users expect, resulting in low efficiency of user vehicle control.
By obtaining multimodal state label data of on-board scenarios, identifying user intention information, and using the scene orchestration model to generate vehicle control instructions, realizing a deep understanding and accurate conversion of user intentions.
It improves the ability to understand user intentions, outputs vehicle control instructions that are more in line with user needs, and improves the user's driving experience.
Smart Images

Figure CN120080867A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent driving technology, and in particular, to a method, an apparatus, and a vehicle for determining a vehicle control instruction. Background Art
[0002] With the continuous development of intelligent driving technology, vehicles have not only become means of transportation but also terminals for information processing and intelligent decision-making. In the in-vehicle scenario, users interact with the in-vehicle system through voice commands, and voice recognition and natural language processing technologies play a key role in this process.
[0003] Traditional voice recognition systems usually rely on user command data (such as voice commands) and face many challenges in the changing in-vehicle environment, such as noise interference and multi-tasking requirements. It is difficult to accurately convert the command data into the in-vehicle operations expected by users, reducing the efficiency of users controlling the vehicle. Summary of the Invention
[0004] To solve the above technical problems or at least partially solve the above technical problems, the present application provides a method, an apparatus, and a vehicle for determining a vehicle control instruction.
[0005] In a first aspect, the present application provides a method for determining a vehicle control instruction, which is applied to the cloud, and the method includes: Obtaining first multi-modal state label data regarding an in-vehicle scenario from an in-vehicle terminal, where the first multi-modal state label data is determined according to multi-modal state data collected by multiple sensors on the vehicle; Performing intent recognition on user interaction information in the first multi-modal state label data to obtain user intent information; Determining whether the user intent information includes an intent for vehicle control; If the user intent information includes an intent for vehicle control, determining a first vehicle control instruction based on the user intent information, the first multi-modal state label data, and a preset scenario orchestration model, and sending the first vehicle control instruction to the in-vehicle terminal to control the vehicle.
[0006] Optionally, the user interaction information includes: a user voice signal; Performing intent recognition on user interaction information in the first multi-modal state label data to obtain user intent information, including: Performing intent recognition on the user voice signal to obtain objective intent information and emotional intent information; Determining the objective intent information and the emotional intent information as the user intent information.
[0007] Optionally, determining whether the user intent information includes an intent for vehicle control includes: Determine whether the user intention information contains a preset keyword associated with a vehicle; If the user intention information contains a preset keyword associated with a vehicle, determine that the user intention information contains an intention for vehicle control.
[0008] Optionally, the scenario orchestration model is obtained by fine-tuning a pre-trained large model, and the fine-tuning training method of the pre-trained large model includes: Obtain model fine-tuning prompt words, user interaction information from multiple in-vehicle terminals, second multi-modal status label data for multiple in-vehicle scenarios, and second vehicle control commands respectively annotated for multiple in-vehicle scenarios; Fine-tune the pre-trained large model based on the model fine-tuning prompt words, multiple pieces of the user interaction information, multiple pieces of the second multi-modal status label data, multiple pieces of the second vehicle control commands, and the low-rank adaptation algorithm to obtain the scenario orchestration model.
[0009] Optionally, fine-tuning the pre-trained large model based on the model fine-tuning prompt words, multiple user interaction information, multiple pieces of the second multi-modal status label data, multiple pieces of the second vehicle control commands, and the low-rank adaptation algorithm to obtain the scenario orchestration model includes: Obtain a weight matrix containing multiple first model parameters in the pre-trained large model, where the number of first model parameters in the weight matrix is less than the number of remaining model parameters in the pre-trained large model, and the remaining model parameters are the model parameters in the pre-trained large model except the first model parameters in the weight matrix; Decompose the weight matrix into a first low-rank matrix and a second low-rank matrix; Train the first low-rank matrix and the second low-rank matrix using the model fine-tuning prompt words, multiple user interaction information, multiple pieces of the second multi-modal status label data, and multiple pieces of the second vehicle control commands to obtain a first fine-tuning parameter matrix and a second fine-tuning parameter matrix; Replace the corresponding first model parameters in the pre-trained large model with the second model parameters in the first fine-tuning parameter matrix and the second fine-tuning parameter matrix to obtain the scenario orchestration model.
[0010] Optionally, after obtaining the user interaction information from multiple in-vehicle terminals, the second multi-modal status label data for in-vehicle scenarios, and the second vehicle control commands respectively annotated for multiple in-vehicle scenarios, the method further includes: Obtain a verification prompt word; Input the verification prompt word, any one of the user interaction information, the second multi-modal status label data, and the second vehicle control instruction corresponding to the second multi-modal status label data into a preset large model, so that the preset large model verifies the rationality of the second vehicle control instruction to obtain a rationality verification result; If the rationality verification result meets the verification passing condition, execute the step of fine-tuning and training the pre-trained large model based on the model fine-tuning prompt word, multiple user interaction information, multiple pieces of the second multi-modal status label data, multiple second vehicle control instructions, and the low-rank adaptation algorithm to obtain the scenario orchestration model.
[0011] Optionally, determining the first vehicle control instruction based on the user intention information, the first multi-modal status label data, and a preset scenario orchestration model includes: Input the user intention information and the first multi-modal status label data into the scenario orchestration model, so that the scenario orchestration model determines a candidate control instruction set based on the user intention information, and determines a candidate control instruction in the candidate control instruction set based on the user intention information and the first multi-modal status label data; Convert the candidate control instruction according to the instruction conversion rule to obtain the first vehicle control instruction.
[0012] Optionally, the scenario orchestration model determines a candidate control instruction set based on the user intention information, and determines the first vehicle control instruction in the candidate control instruction set based on the user intention information and the first multi-modal status label data, including: The scenario orchestration model determines the candidate control instruction set based on the objective intention information in the user intention information and the first multi-modal status label data; Based on the emotional intention information in the user intention information and the first multi-modal status label data in the candidate control instruction set, determine the first vehicle control instruction whose priority and instruction execution method match the emotional intention information and the first multi-modal status label data.
[0013] Optionally, converting the candidate control instruction according to the instruction conversion rule to obtain the first vehicle control instruction includes: Perform synonym correction on the candidate control instruction to obtain a corrected candidate control instruction; Convert the corrected candidate control instruction according to the standard instruction format to obtain the first vehicle control instruction.
[0014] In a second aspect, the present application provides a method for determining a vehicle control instruction, which is applied to a vehicle head unit, and the method includes: Obtain multi-modal status data collected by multiple sensors installed on the vehicle; Perform multi-modal preprocessing on the multi-modal status data to obtain first multi-modal status label data, and send the first multi-modal status label data to the cloud; Receive a first vehicle control instruction from the cloud; Execute the first vehicle control instruction.
[0015] Optionally, the multi-modal status data includes: image data and / or video data collected by a camera, sound data collected by a microphone, and perception data collected by a vehicle operation status and environment perception sensor; Performing multi-modal preprocessing on the multi-modal status data to obtain first multi-modal status label data includes: Identify the image data and / or video data to obtain recognized text information, and generate first label information based on the recognized text information; Locate the sound source of the sound data to obtain sound source information, and convert the sound source information into second label information; Convert the perception data into third label information; Determine the combination of the first label information, the second label information, and the third label information as the first multi-modal status label data.
[0016] Optionally, after executing the first vehicle control instruction, the method further includes: Obtain the instruction execution result corresponding to the first vehicle control instruction; Send the instruction execution result to the user.
[0017] In a third aspect, the present application provides a vehicle control instruction determination device, which is applied to the cloud, and the device includes: A first acquisition module, configured to acquire first multi-modal status label data about an in-vehicle scene from a vehicle terminal, where the first multi-modal status label data is determined according to multi-modal status data collected by multiple sensors on the vehicle; An intention recognition module, configured to perform intention recognition on user interaction information in the first multi-modal status label data to obtain user intention information; A first determination module, configured to determine whether the user intention information includes an intention for vehicle control; A second determination module, configured to, if the user intention information includes an intention for vehicle control, determine a first vehicle control instruction based on the user intention information, the first multi-modal status label data, and a preset scenario arrangement model, and send it to the vehicle terminal to control the vehicle.
[0018] Fourth aspect, the present application provides a vehicle control instruction determination device, which is applied to the in-vehicle computer terminal. The device includes: A second acquisition module, configured to acquire multi-modal state data collected by a plurality of sensors arranged on the vehicle; A preprocessing module, configured to perform multi-modal preprocessing on the multi-modal state data to obtain first multi-modal state label data, and send the first multi-modal state label data to the cloud; A receiving module, configured to receive a first vehicle control instruction from the cloud; An execution module, configured to execute the first vehicle control instruction.
[0019] Fifth aspect, the present application provides a vehicle, which includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; The processor, when executing the program stored on the memory, implements the vehicle control instruction determination method according to any one of the first aspect or the vehicle control instruction determination method according to any one of the second aspect.
[0020] Advantageous effects of the present invention: By acquiring the first multi-modal state label data of the in-vehicle scenario in the embodiments of the present application, the user intention information can be determined based on the user interaction information in the first multi-modal state label data. When the user intention information includes an intention for vehicle control, the scenario orchestration model can be used to generate a first vehicle control instruction according to the user intention information and the first multi-modal state label data, so as to realize the in-depth analysis of the user's control intention for the vehicle included in the multi-modal state label data in the in-vehicle scenario, improve the ability to understand the user's intention, and output a vehicle control instruction that better meets the user's needs, so as to enhance the user's driving experience. Description of the Drawings
[0021] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present invention, and are used together with the description to explain the principles of the present invention.
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a flowchart of a vehicle control instruction determination method provided by an embodiment of the present application; Figure 2 is Figure 1Flowchart of step S104 in Figure 3 is Figure 2 Flowchart of step S201 in Figure 4 is Figure 2 Flowchart of step S202 in Figure 5 Flowchart of a fine-tuning training method for a pre-trained large model provided by an embodiment of the present application; Figure 6 Another flowchart of a fine-tuning training method for a pre-trained large model provided by an embodiment of the present application; Figure 7 is Figure 5 Flowchart of step S502 in Figure 8 Flowchart of another vehicle control instruction determination method provided by an embodiment of the present application; Figure 9 Structure diagram of a vehicle control instruction determination device provided by an embodiment of the present application; Figure 10 Structure diagram of another vehicle control instruction determination device provided by an embodiment of the present application; Figure 11 Structure diagram of a vehicle provided by an embodiment of the present application. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0025] Existing in-vehicle voice command recognition and processing systems still have many deficiencies and challenges in practical applications, which are specifically manifested in the following aspects: 1. Single-modal dependence: Traditional in-vehicle voice recognition systems mainly rely on single-modal voice data, and their processing capabilities are limited for complex in-vehicle scenarios. Single-modal systems are easily affected by noise, voice ambiguity, and external interference, resulting in low voice recognition accuracy and unable to meet the high expectations of users.
[0026] 2. Insufficient multi-task processing ability: In the in-vehicle environment, users' voice commands often involve multi-task processing, such as simultaneously controlling navigation, music playback, and the air conditioning system. Existing technologies show certain limitations in understanding and executing these complex commands, making it difficult to provide a smooth user experience.
[0027] 3. Poor scene adaptability: The in-vehicle environment has the characteristic of dynamic changes, and the understanding requirements for voice commands are different in different driving scenarios. Existing voice recognition systems lack sufficient adaptability when facing complex and changeable in-vehicle scenarios, resulting in unsatisfactory processing effects of voice commands.
[0028] 4. Insufficient data fusion: Although some advanced systems have begun to attempt multi-modal data fusion, in practical applications, the efficient fusion of multi-modal data remains a difficult problem. Existing methods often fail to fully utilize the complementary advantages of each modal data when integrating multi-modal information such as voice, image, and sensor data, resulting in limited overall system performance.
[0029] 5. Accuracy and reliability issues of instruction conversion: When processing and converting voice commands, large models may return inaccurate or incorrect commands due to reasons such as data noise and model bias. This may lead to serious operation errors in the in-vehicle scenario. Therefore, the existing technology lacks an effective verification system to verify and correct the commands returned by the large model to ensure the accuracy and reliability of instruction conversion.
[0030] Due to the above deficiencies, traditional in-vehicle voice command recognition and processing systems are difficult to accurately convert command data into the in-vehicle operations expected by users, reducing the user's vehicle control efficiency. For this reason, the embodiments of the present application provide a method, device, and vehicle for determining vehicle control commands, which can, in the in-vehicle environment, convert comprehensive voice commands involving complex scene choreography requirements into vehicle control commands through the fusion of multi-modal data and fine-tuning of large models. For example, when the user issues a comprehensive voice command intended to control the air conditioner, navigation, and entertainment systems simultaneously, the vehicle control command can be automatically determined according to the comprehensive voice command, and this vehicle control command can control the air conditioner, navigation, and entertainment systems simultaneously, improving the voice command processing ability of the system in complex in-vehicle scenarios. At the same time, a verification system is introduced to verify and correct the commands returned by the large model to ensure the accuracy and reliability of instruction conversion, improve the user experience, achieve efficient and accurate conversion of voice commands in complex in-vehicle scenarios, and improve the intelligent level and user experience of the in-vehicle system.
[0031] The embodiments of the present application provide a method for determining vehicle control commands, which is applied to the cloud, as Figure 1 shown, and the method includes: Step S101, obtaining first multi-modal status tag data about the in-vehicle scene from the vehicle head unit; In the embodiments of the present application, the first multi-modal status tag data is determined according to the multi-modal status data collected by multiple sensors on the vehicle. The multi-modal status data includes: image data and / or video data collected by a camera, sound data collected by a microphone, and perception data collected by vehicle running status and environment perception sensors, etc. Further, the multi-modal status data includes: sound data, environmental status data, vehicle body component status data, posture data of people, and so on.
[0032] Step S102: Perform intention recognition on the user interaction information in the first multi-modal status tag data to obtain user intention information. In the embodiments of the present application, the user interaction information includes: user voice signals input by the user through voice, user text information input by the user on the control panel, user action information of user-specified actions collected by the camera, etc. In practical applications, the user can also input user interaction information in other ways, which is not limited here.
[0033] In an implementation manner of the present application, the user interaction information includes: user voice signals; Step S102 performs intention recognition on the user interaction information in the first multi-modal status tag data to obtain user intention information, including: Perform intention recognition on the user voice signal to obtain objective intention information and emotional intention information; determine the objective intention information and the emotional intention information as the user intention information.
[0034] When performing intention recognition on the user voice signal, feature extraction can be performed. The system extracts key features from the voice signal to obtain emotional intention information. The emotional intention information includes features such as pitch, speech rate, and emotion. These features help to accurately capture the driver's emotions and intentions; at the same time, it is also necessary to perform signal processing on the user voice signal, use advanced speech processing algorithms to clean and enhance the original voice signal, and then recognize the objective intention information in the voice signal to ensure the accuracy and reliability of the analysis.
[0035] Step S103: Determine whether the user intention information includes an intention for vehicle control. In practical applications, users in the vehicle may talk to each other, talk to others through mobile terminals such as mobile phones, or listen to the radio, play music, etc., which will all generate sounds. Therefore, the user interaction information in the multi-modal status tag data uploaded by the in-vehicle device may not be an intention for vehicle control. Therefore, it is necessary to further determine whether the user intention information includes an intention for vehicle control.
[0036] In an implementation manner of the present application, Step S103 determines whether the user intention information includes an intention for vehicle control, including: Determine whether the preset keywords associated with the vehicle are included in the user intention information; if the preset keywords associated with the vehicle are included in the user intention information, determine that the user intention information includes the intention for vehicle control.
[0037] In this embodiment, the preset keywords may include, for example: the name of the in-vehicle terminal interaction assistant (e.g., Xiao An), and / or, the actions that need to be controlled for the vehicle to execute (e.g., open, close, window, door, play, music, turn up, turn down, temperature, etc.).
[0038] Step S104, if the user intention information includes the intention for vehicle control, determine the first vehicle control instruction based on the user intention information, the first multi-modal status label data, and the preset scenario orchestration model, and send it to the in-vehicle terminal to control the vehicle.
[0039] In this step, if the user intention information includes the intention for vehicle control, the user intention information and the first multi-modal status label data can be input into the scenario orchestration model, so that the scenario orchestration model outputs the first vehicle control instruction, and then the first vehicle control instruction is sent to the in-vehicle terminal, so that the in-vehicle terminal controls the vehicle to execute the first vehicle control instruction.
[0040] If the user intention information does not include the intention for vehicle control, end the process.
[0041] In an embodiment of the present application, step S104 determines the first vehicle control instruction based on the user intention information, the first multi-modal status label data, and the preset scenario orchestration model, as Figure 2 shown, including: Step S201, input the user intention information and the first multi-modal status label data into the scenario orchestration model, so that the scenario orchestration model determines a candidate control instruction set based on the user intention information, and determines a candidate control instruction in the candidate control instruction set based on the user intention information and the first multi-modal status label data; The scenario orchestration model not only depends on voice data, but also combines other modal information, such as the driver's facial expressions, gestures, etc. By analyzing the multi-modal status label data, the model can determine the candidate control instruction set and determine the candidate control instruction in the candidate control instruction set based on the user intention information and the first multi-modal status label data.
[0042] In an embodiment of the present application, in step S201, the scenario orchestration model determines the candidate control instruction set based on the user intention information, and determines the first vehicle control instruction in the candidate control instruction set based on the user intention information and the first multi-modal status label data, asFigure 3 As shown in the figure, it includes: Step S301, the scenario orchestration model determines the candidate control instruction set based on the objective intention information in the user intention information and the first multi-modal status tag data; In this step, the scenario orchestration model determines the candidate control instruction set based on the literal content without emotional color, that is, the objective intention information in the user intention information and the first multi-modal status tag data. The candidate control instruction set can contain multiple candidate control instructions with different priorities and instruction execution methods.
[0043] Step S302, based on the emotional intention information in the user intention information and the first multi-modal status tag data, in the candidate control instruction set, determine the first vehicle control instruction whose priority and instruction execution method match the emotional intention information and the first multi-modal status tag data.
[0044] Based on the emotional intention information and the first multi-modal status tag data, the scenario orchestration model can accurately judge the user's current state. For example, through facial expressions and gestures, the model can identify whether the driver is nervous, relaxed or focused. Furthermore, based on the driver's state, the system can dynamically adjust the priority and execution method of the instruction. For example, when it is detected that the driver is nervous, the system may give priority to playing soothing music or enabling the autopilot function to relieve the driver's stress.
[0045] Step S202, convert the candidate control instruction according to the instruction conversion rule to obtain the first vehicle control instruction.
[0046] In this embodiment, the system converts the recognized voice instruction through a set of instruction conversion rules to ensure the accurate execution of in-vehicle operations.
[0047] In an implementation manner of the present application, step S202 converts the candidate control instruction according to the instruction conversion rule to obtain the first vehicle control instruction, as Figure 4 shown, including: Step S401, perform synonym correction on the candidate control instruction to obtain the corrected candidate control instruction; First, in order to optimize the matching process, the cloud system can handle the subtle differences in voice instructions. Whether the driver uses expressions such as "open the driver's door" or "open the main driver's door", the system can recognize and match the standard instruction DriverDoorOpen. Through this correction mechanism, the system can effectively handle various different voice expressions to ensure the accuracy and consistency of operations.
[0048] Step S402: Convert the corrected candidate control instruction according to the standard instruction format to obtain the first vehicle control instruction.
[0049] The system is also configured with a standard instruction format to ensure the consistency and accuracy of operations. For example: Example 1: The standard instruction for the driver to open the door is DriverDoorOpen.
[0050] Example 2: The standard instruction for playing music is PlayMusic.
[0051] In actual usage scenarios, the driver's voice expressions may be different from the predefined standard instructions. The scenario orchestration model may return instructions such as "Driver opens the door"; the system compares and matches these voice instructions with the predefined standard instructions through an instruction matching mechanism. For example, "Driver opens the door" is matched to the standard instruction DriverDoorOpen.
[0052] For example: When the user says "Remind me to take my phone when getting out of the driver's seat", the system generates corresponding trigger, status, and execution commands based on the multi-modal status tag data: Trigger data: The driver's door opens Status data: There is someone in the driver's seat Execution data: Voice broadcast to remind to take the phone Text matching logic: Match the words or short phrases in the trigger data, status data, and execution data with the strings in the standard dictionary respectively. For example, the trigger data "Driver's door" is matched to "Driver's door" in the standard dictionary, that is, "Driver's door" = "Driver's door". If the match is successful, the successfully matched data is not corrected, and the first successfully matched data (the first data refers to the data in the trigger data, status data, and execution data that is successfully matched with the standard dictionary) is text-converted to obtain an instruction available for the in-vehicle computer; if the match fails, the second failed-matched data (the second data refers to the data in the trigger data, status data, and execution data that fails to match the standard dictionary) is processed by the synonym matching logic; Synonym matching logic: A synonym library can be pre-configured. The synonym library contains multiple standard words and their corresponding synonyms. The second data is matched with multiple synonyms in the synonym library respectively. If the match is successful, the synonyms in the second data are corrected to the standard words corresponding to the synonyms. For example, "Driver's door" = "Driver's door", "Driver side door" = "Driver's door", and then the third successfully matched data is text-converted to obtain an instruction available for the in-vehicle computer. If the match fails, the fourth failed-matched data is processed by the text vector matching logic; Text vector matching logic: Convert the fourth data, such as the word "driver's door", into a word vector. By matching the word vector with the standard word vector library, a similarity matching result is returned. If the similarity matching result indicates successful text vector matching, the fifth data that matches successfully is text-converted to obtain an instruction available for the in-vehicle unit; In-vehicle unit instruction: Text is converted into an instruction available for the in-vehicle unit.
[0053] For example: "Open the driver's door" = "DriverDoorOpen", "Close the driver's door" = "DriverDoorClose".
[0054] Finally, if the text vector matching fails: It is considered that the in-vehicle unit does not support the current instruction, and the instruction is not sent.
[0055] Through the above steps, the cloud can achieve accurate recognition and conversion of voice instructions, greatly improving the user experience of in-vehicle voice interaction. It not only ensures the accurate execution of instructions but also can flexibly handle different voice expressions, thus providing more intelligent and user-friendly services.
[0056] After generating the first vehicle control instruction and sending it to the in-vehicle unit terminal to control the vehicle, it is also possible to verify whether the generated first vehicle control instruction and the executed action are accurate and error-free. Further, it can ensure that there are no omissions or errors during the instruction conversion process. For example, when the system receives the instruction "Open the sunroof", it actually executes the corresponding action; it can also comprehensively detect the response speed and accuracy of the system to instructions by simulating different usage scenarios.
[0057] In the embodiment of the present application, by obtaining the first multi-modal state label data of the vehicle-mounted scenario, the user intention information can be determined based on the user interaction information in the first multi-modal state label data. When the user intention information includes the intention for vehicle control, the scenario orchestration model can generate the first vehicle control instruction according to the user intention information and the first multi-modal state label data, realizing in-depth analysis of the user's control intention for the vehicle included in the user interaction information based on the multi-modal state label data in the vehicle-mounted scenario, improving the ability to understand the user's intention, and outputting a vehicle control instruction that better meets the user's needs to enhance the user's driving experience.
[0058] In another embodiment of the present application, the scenario orchestration model is obtained by fine-tuning a pre-trained large model. The fine-tuning training method of the pre-trained large model is as Figure 5 shown and includes: Step S501, obtain model fine-tuning prompt words, user interaction information from multiple in-vehicle unit terminals, second multi-modal state label data regarding multiple vehicle-mounted scenarios, and second vehicle control instructions respectively annotated for multiple vehicle-mounted scenarios; In the embodiments of the present application, the second multi-modal status tag data is also obtained after multi-modal preprocessing on the in-vehicle device side. Each group of second multi-modal status tag data is determined according to the multi-modal status data collected by multiple sensors on each vehicle. The multi-modal status data includes: image data and / or video data collected by a camera, sound data collected by a microphone, and perception data collected by a vehicle operation status and environment perception sensor, etc. Further, the multi-modal status data includes: sound data, environmental status data, vehicle body component status data, human posture data, and so on.
[0059] The second vehicle control instruction is marked for the in-vehicle scenario of the second multi-modal status tag data by the user or the staff. As an example, after the user or the staff inputs user interaction information, if the in-vehicle device side cannot make a response that meets the user's expectations, the second vehicle control instruction manually input by the user can be obtained to be used for the fine-tuning training of the pre-trained large model.
[0060] The user interaction information from multiple in-vehicle device sides, the second multi-modal status tag data regarding the in-vehicle scenario, and the second vehicle control instructions marked for multiple in-vehicle scenarios respectively can adopt the following JSON format to express the input and output relationship of the multi-modal data: { "instruction": "prompt model fine-tuning prompt"; "input": "Text spoken by the user, multi-modal data label", "output": "{ Trigger name, trigger condition, trigger parameter; status: [{status name, status condition, status parameter}] execute: [execution object, execution condition, execution parameter] }"}} Example For example, when the user says "Open the sunroof and air conditioner when I get in the car", the system will generate corresponding status and execution commands according to the multi-modal status tag data: Trigger condition: The driver's door is opened, trigger parameter: Door opened Multi-modal status tag: Person: Female owner; Location: Passenger seat Status generation: Status name: Person, status condition: Equal to (default is the same as the actual situation), status parameter: Female owner.
[0061] Status name: Location, status condition: Equal to, status parameter: Passenger seat.
[0062] To ensure the accuracy and reliability of the system, the quality of the training data needs to be strictly inspected and verified. For example, it can be checked through prompt engineering whether the training data is logical. Further, by setting a series of predefined prompts and scenarios, the system's responses in different situations can be tested; it can be checked whether the actions generated by the system are consistent with the expectations. For example, when "the hostess" is detected, whether the system correctly performs the actions of opening the sunroof and air conditioner; by continuously adjusting the prompts and training data, it can be ensured that the system can make logical responses in various complex scenarios.
[0063] The scenario orchestration model trained with the training data that has passed the above inspections and verifications can, by recognizing different multi-modal tags (such as vision, sound, etc.), enable the system to generate and execute corresponding actions to enhance the user's driving experience. For example, when the system detects that "the hostess" is sitting in the passenger seat, it will automatically generate corresponding status conditions and perform the operations of opening the sunroof and air conditioner to ensure a comfortable environment and meet the user's personalized needs. Such intelligent operations are based on the deep learning and analysis of multi-modal data to ensure that considerate services can be provided in various driving scenarios.
[0064] In an implementation manner of the present application, after obtaining user interaction information from multiple in-vehicle terminals, second multi-modal status tag data regarding in-vehicle scenarios, and second vehicle control instructions respectively annotated for multiple in-vehicle scenarios in step S501, the method is as Figure 6 shown and further includes: Step S601, obtaining verification prompts; As an example, a kind of verification prompt is as follows: Background information: You are an automotive user experience judge.
[0065] Task description Based on the requests put forward by the user and the actions actually performed by the vehicle, judge whether they are reasonable and make a judgment.
[0066] Input format: { input: User input user interaction information and multi-modal status tag data; output: Actions actually performed by the vehicle; } Output format: { Judgment: Reasonable / Unreasonable } Example output: { Judgment: Reasonable } Step S602: Input the verification prompt word, any one of the user interaction information, the second multi-modal status label data, and the second vehicle control instruction corresponding to the second multi-modal status label data into a preset large model, so that the preset large model verifies the rationality of the second vehicle control instruction and obtains a rationality verification result. In this step, the verification prompt word, one set of user interaction information, the second multi-modal status label data, and the second vehicle control instruction corresponding to the second multi-modal status label data can be input into the preset large model, so that the preset large model verifies whether the second vehicle control instruction is reasonable and obtains a rationality verification result.
[0067] Step S603: If the rationality verification result meets the verification passing condition, perform the step of fine-tuning and training the pre-trained large model based on the model fine-tuning prompt word, multiple pieces of user interaction information, multiple pieces of the second multi-modal status label data, multiple second vehicle control instructions, and the low-rank adaptation algorithm to obtain the scenario orchestration model.
[0068] Step S502: Fine-tune and train the pre-trained large model based on the model fine-tuning prompt word, multiple pieces of user interaction information, multiple pieces of the second multi-modal status label data, multiple second vehicle control instructions, and the low-rank adaptation algorithm to obtain the scenario orchestration model.
[0069] In the embodiments of the present application, the fine-tuning process of the pre-trained large model mainly adopts the low-rank adaptation algorithm (Low-Rank Adaptation, LoRA) method. By introducing a small number of trainable parameters, the efficient adjustment of the large model is realized to meet the specific task requirements.
[0070] As an example, the prompt words for a pre-trained large model during fine-tuning training are as follows: Background information: You are a car assistant. Users may mention various requirements during driving, such as adjusting the air conditioner, navigation, making or answering calls, etc. You need to understand the user's intention and convert it into specific scenarios and parameters.
[0071] Task description: According to the user's description, orchestrate a reasonable car scenario and summarize a scenario name with no more than 10 characters. You need to extract information such as trigger events, preconditions, and execution actions from the user's description.
[0072] Parameter condition list: Compound interval, not equal to, equal to, greater than or equal to, greater than, less than or equal to, less than Input format: Description of the user. For example: "Initiate navigation and remind me to slow down the vehicle when the estimated arrival distance is 2 kilometers" Output format: { "Scenario Name": "Scenario Name", "Trigger Name": "Trigger Event", "Trigger Parameter": "Trigger Parameter", "Trigger Condition": "Trigger Condition", "Unit of Trigger Parameter": "Trigger Parameter Unit", "Array of Pre - status": { "Status Name": "Status Condition", "Status Parameter": "Status Parameter", "Unit of Status Parameter": "Status Parameter Unit", "Status Condition": "Status Condition" } , "Array of Execution Actions": { "Execution Name": "Execution Action", "Execution Parameter": "Execution Parameter", "Execution Condition": "Execution Condition", "Unit of Execution Parameter": "Execution Parameter Unit" } } Example: Input: "Initiate navigation and remind me to slow down the vehicle when the estimated arrival distance is 2 kilometers" Output: { "Scenario Name": "Navigation Reminder", "Trigger Name": "Map", "Trigger Parameter": "Initiate navigation", "Trigger Condition": "Equal to", "Unit of Trigger Parameter": "", "Array of Pre - status": { "Status Name": "Distance", "Status Parameter": "2", "Unit of Status Parameter": "kilometer", "Status Condition": "Equal to" } , "Execution action array": { "Execution name": "Reminder", "Execution parameter": "Vehicle slow down", "Execution condition": "Equal to", "Execution parameter unit": "" } } In an implementation manner of the present application, step S502 fine-tunes and trains a pre-trained large model based on the model fine-tuning prompt words, the multiple user interaction information, the multiple second multi-modal status label data, the multiple second vehicle control instructions, and the low-rank adaptation algorithm to obtain the scenario choreography model. As Figure 7 shown, it includes: Step S701: Obtain a weight matrix containing multiple first model parameters in the pre-trained large model. The number of first model parameters in the weight matrix is less than the number of remaining model parameters in the pre-trained large model. The remaining model parameters are the model parameters in the pre-trained large model except for the first model parameters in the weight matrix; Step S702: Decompose the weight matrix into a first low-rank matrix and a second low-rank matrix; Step S703: Use the model fine-tuning prompt words, the multiple user interaction information, the multiple second multi-modal status label data, and the multiple second vehicle control instructions to train the first low-rank matrix and the second low-rank matrix to obtain a first fine-tuning parameter matrix and a second fine-tuning parameter matrix; Step S704: Use the second model parameters in the first fine-tuning parameter matrix and the second fine-tuning parameter matrix to replace the corresponding first model parameters in the pre-trained large model to obtain the scenario choreography model.
[0073] In the embodiment of the present application, the main objective of the fine-tuning training process of the pre-trained model is to reduce the number of parameters that need to be trained during the fine-tuning process, thereby reducing the computational cost and storage requirements. In the embodiment of the present application, a low-rank matrix is introduced to fine-tune the weights of the pre-trained model instead of directly adjusting all weight parameters. The fine-tuning training process is as follows: Freezing of the pre-trained model: During the fine-tuning process, most of the weight parameters of the pre-trained model remain unchanged (frozen), and only a small number of parameters are adjusted. These small number of parameters are represented by low-rank matrices; Low-rank matrix: Assume that some weight matrices in the pre-trained model are (W), and LoRA decomposes it into two low-rank matrices (A) and (B). Among them, the ranks of (A) and (B) are relatively low, so the number of parameters that need to be trained is greatly reduced.
[0074] Fine-tuning process: During the fine-tuning process, only the low-rank matrices (A) and (B) are trained, while the other weights of the pre-trained model remain unchanged. This can significantly reduce the consumption of computing resources while ensuring the fine-tuning effect.
[0075] The scenario arrangement model of this application can cross-modally understand complex information in the vehicle-mounted scenario by simultaneously receiving voice, image, and sensor data at the input layer of the model; by adjusting the model parameters, it performs excellently in multi-task processing, especially in the accurate recognition of voice commands and scenario understanding; the fine-tuned model can more effectively process and fuse information from different modalities, significantly improving its performance in complex vehicle-mounted environments.
[0076] In another embodiment of this application, a method for determining vehicle control instructions is also provided, which is applied to the in-vehicle computer, as Figure 8 shown, the method includes: Step S801, obtain multi-modal status data collected by multiple sensors arranged on the vehicle; In the embodiments of this application, the multi-modal status data includes: image data and / or video data collected by a camera, sound data collected by a microphone, and perception data collected by a vehicle operation status and environment perception sensor. The system can collect multi-modal status data in the vehicle-mounted environment in real time through devices such as vehicle-mounted sensors, cameras, and microphones. These data include but are not limited to: voice data: collect the driver's voice commands through a microphone; image and video data: capture information such as the in-vehicle environment, the driver's facial expressions, and gestures through an in-vehicle camera; sensor data: obtain information such as speed, acceleration, steering wheel angle, and environmental temperature through various sensors of the vehicle.
[0077] Step S802, perform multi-modal preprocessing on the multi-modal status data to obtain first multi-modal status label data, and send the first multi-modal status label data to the cloud; In an implementation manner of this application, step S802 performs multi-modal preprocessing on the multi-modal status data to obtain first multi-modal status label data, including: identifying the image data and / or video data to obtain recognition text information, generating first label information based on the recognition text information; locating the sound source of the sound data to obtain sound source information, and converting the sound source information into second label information; converting the perception data into third label information; determining the combination of the first label information, the second label information, and the third label information as the first multi-modal status label data.
[0078] In this embodiment, multi-modal state data can be pre-processed (such as noise filtering, image enhancement, and signal calibration) to ensure the accuracy and consistency of the data. To enhance the application value of the data while protecting user privacy, the processed data will be converted into labeled data. For example: Passenger identification: When the system identifies the hostess getting into the car, it records relevant image information, and during her subsequent ride, it generates the label "hostess" through image matching; Voice source localization: The system can identify the source direction of voice commands (such as the driver's seat, passenger seat, or rear row) and generate corresponding labels; Environmental data labeling: Data such as the current vehicle speed and the temperature inside the car will be converted into label information for further analysis and response.
[0079] The system can also manage multi-user scenarios. When multiple passengers are in the car at the same time, the system uses voice and image recognition technologies to distinguish different users and provides personalized services according to their identities. For example, after recognizing the voice commands of users in different seats, the system will adjust settings such as the audio volume, playing content, or seat heating respectively to meet the needs of each passenger.
[0080] Step S803: Receive a first vehicle control instruction from the cloud. Step S804: Execute the first vehicle control instruction.
[0081] In the embodiment of the present application, the verified instruction will be sent to each subsystem of the vehicle system for execution to ensure precise and efficient operation. The specific execution process is as follows: 1. System: Execute path planning and navigation instructions to ensure that the vehicle travels along the optimal path and provides accurate navigation services.
[0082] 2. Air conditioning system: Adjust the temperature and wind speed inside the car according to the instruction to provide a comfortable in-vehicle environment. For example, when receiving the instruction "raise the temperature", the system will automatically increase the temperature inside the car.
[0083] 3. Entertainment system: Play music, adjust the volume, or switch media according to the instruction. Whether it is selecting a specific song to play or adjusting the volume, the system can quickly respond and perform the corresponding operation.
[0084] 4. Subsequent conversation: The system supports dynamic adjustment and further operations based on subsequent conversations. For example, users can at any time request the system to adjust the current settings or propose new requirements through voice commands, and the system will respond in a timely manner and execute the new instruction to ensure a continuously optimized user experience.
[0085] In an embodiment of the present application, after executing the first vehicle control instruction, the method further includes: Obtain the instruction execution result corresponding to the first vehicle control instruction; send the instruction execution result to the user.
[0086] During the execution of instructions, the system will monitor the instruction execution results of each subsystem in real time and provide the feedback information to the user in a timely manner. For example, after a navigation instruction is executed, the system will notify the user that the route has been updated; after the air conditioner temperature adjustment is completed, the system will display the current temperature setting. This real-time feedback mechanism ensures that users can understand the operation results in a timely manner, improving the transparency and satisfaction of the interaction experience.
[0087] Through the above steps, the vehicle-mounted system of this embodiment can achieve precise execution and real-time feedback of voice instructions, providing users with an intelligent and personalized in-vehicle experience.
[0088] The embodiments of this application also provide the following overall optimization directions: 1. Response speed: Optimize the data transmission and instruction processing speeds between modules to ensure that the system's response speed meets real-time requirements.
[0089] 2. User interaction: Continuously optimize the system's interaction interface and feedback mechanism by analyzing users' usage habits and preferences to improve the user experience.
[0090] Through these implementation manners, the present invention can effectively improve the voice instruction conversion and execution efficiency in the vehicle-mounted scenario, significantly improving the driving experience of users.
[0091] In another embodiment of this application, a vehicle control instruction determination device is also provided, which is applied to the cloud, as Figure 9 shown. The device includes: A first acquisition module 11, configured to acquire first multi-modal state label data regarding a vehicle-mounted scenario from a vehicle head unit, where the first multi-modal state label data is determined according to multi-modal state data collected by multiple sensors on the vehicle; An intention recognition module 12, configured to perform intention recognition on the user interaction information in the first multi-modal state label data to obtain user intention information; A first determination module 13, configured to determine whether the user intention information includes an intention for vehicle control; A second determination module 14, configured to, if the user intention information includes an intention for vehicle control, determine a first vehicle control instruction based on the user intention information, the first multi-modal state label data, and a preset scenario arrangement model, and send it to the vehicle head unit to control the vehicle.
[0092] In another embodiment of this application, a vehicle control instruction determination device is also provided, which is applied to the vehicle head unit, as Figure 10 shown. The device includes: The second acquisition module 21 is configured to acquire multimodal status data collected by a plurality of sensors disposed on the vehicle; The preprocessing module 22 is configured to perform multimodal preprocessing on the multimodal status data to obtain first multimodal status label data, and send the first multimodal status label data to the cloud; The receiving module 23 is configured to receive a first vehicle control instruction from the cloud; The execution module 24 is configured to execute the first vehicle control instruction.
[0093] In another embodiment of the present application, a vehicle is further provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; When the processor is configured to execute the program stored on the memory, it implements any one of the foregoing vehicle control instruction determination methods applied to the cloud or any one of the foregoing vehicle control instruction determination methods applied to the in-vehicle device.
[0094] For the vehicle provided in the embodiment of the present invention, the processor can determine user intention information based on the user interaction information in the first multimodal status label data by executing the program stored on the memory and obtaining the first multimodal status label data of the in-vehicle scenario. When the user intention information includes an intention for vehicle control, the scenario orchestration model can be used to generate a first vehicle control instruction according to the user intention information and the first multimodal status label data, so as to realize in-depth analysis of the user's control intention for the vehicle included in the multimodal status label data in the in-vehicle scenario, improve the ability to understand the user's intention, output a vehicle control instruction that better meets the user's needs, and enhance the user's driving experience.
[0095] The communication bus 1140 mentioned in the above vehicle may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 1140 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 11 only a thick line is used to represent it here, but it does not mean that there is only one bus or one type of bus.
[0096] The communication interface 1120 is used for communication between the above vehicle and other devices.
[0097] The memory 1130 may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0098] The aforementioned processor 1110 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0099] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0100] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for determining a vehicle control instruction, characterized in that: Applied to the cloud, the method includes: Acquire first multimodal state tag data about a vehicle-mounted scene from a vehicle machine end, wherein the first multimodal state tag data is determined based on multimodal state data collected by a plurality of sensors on the vehicle; Performing intent recognition on the user interaction information in the first multimodal state label data to obtain user intent information; determining whether the user intention information includes an intention for vehicle control; If the user intention information includes an intention for vehicle control, a first vehicle control instruction is determined based on the user intention information, the first multimodal state tag data and a preset scenario arrangement model, and is sent to the vehicle machine to control the vehicle.
2. The vehicle control instruction determination method according to claim 1, characterized in that: The user interaction information includes: user voice signal; Performing intent recognition on the user interaction information in the first multimodal state label data to obtain user intent information includes: Performing intent recognition on the user voice signal to obtain objective intent information and emotional intent information; The objective intention information and the emotional intention information are determined as the user intention information.
3. The vehicle control instruction determination method according to claim 1, characterized in that: Determining whether the user intention information includes an intention for vehicle control includes: Determining whether the user intention information includes a preset keyword associated with the vehicle; If the user intention information includes a preset keyword associated with the vehicle, it is determined that the user intention information includes an intention for vehicle control.
4. The vehicle control instruction determination method according to claim 1, characterized in that: The scene arrangement model is obtained by fine-tuning the pre-trained large model, and the fine-tuning training method of the pre-trained large model includes: Obtaining model fine-tuning prompt words, user interaction information from multiple vehicle terminals, second multimodal state label data about multiple vehicle-mounted scenes, and second vehicle control instructions respectively labeled for the multiple vehicle-mounted scenes; Based on the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, multiple second vehicle control instructions and a low-rank adaptation algorithm, the pre-trained large model is fine-tuned to obtain the scene arrangement model.
5. The vehicle control instruction determination method according to claim 4, characterized in that: The pre-trained large model is fine-tuned based on the model fine-tuning prompt word, a plurality of the second multimodal state label data, a plurality of the user interaction information, a plurality of the second vehicle control instructions and a low-rank adaptation algorithm to obtain the scene arrangement model, including: Obtain a weight matrix including a plurality of first model parameters in a pre-trained large model, wherein the number of first model parameters in the weight matrix is less than the number of remaining model parameters in the pre-trained large model, and the remaining model parameters are model parameters in the pre-trained large model other than the first model parameters in the weight matrix; Decomposing the weight matrix into a first low-rank matrix and a second low-rank matrix; Using the model fine-tuning prompt words, the plurality of user interaction information, the plurality of the second multimodal state label data, and the plurality of the second vehicle control instructions to train the first low-rank matrix and the second low-rank matrix to obtain a first fine-tuning parameter matrix and a second fine-tuning parameter matrix; The scene arrangement model is obtained by replacing the corresponding first model parameters in the pre-trained large model with the second model parameters in the first fine-tuning parameter matrix and the second fine-tuning parameter matrix.
6. The vehicle control instruction determination method according to claim 4, characterized in that: After obtaining user interaction information from multiple vehicle terminals, second multimodal state label data about vehicle scenes, and second vehicle control instructions respectively labeled for multiple vehicle scenes, the method further includes: Get the verification prompt word; Inputting the verification prompt word, any of the user interaction information, the second multimodal state label data, and the second vehicle control instruction corresponding to the second multimodal state label data into a preset large model, so that the preset large model performs rationality verification on the second vehicle control instruction to obtain a rationality verification result; If the rationality verification result meets the verification pass conditions, the step of fine-tuning the pre-trained large model based on the model fine-tuning prompt words, multiple user interaction information, multiple second multimodal state label data, multiple second vehicle control instructions and low-rank adaptation algorithm is executed to obtain the scene arrangement model.
7. The vehicle control instruction determination method according to claim 1, characterized in that: Determining a first vehicle control instruction based on the user intention information, the first multimodal state label data, and a preset scenario arrangement model includes: Inputting the user intention information and the first multimodal state label data into the scenario choreography model, so that the scenario choreography model determines a candidate control instruction set based on the user intention information, and determines a candidate control instruction in the candidate control instruction set based on the user intention information and the first multimodal state label data; The candidate control instruction is converted according to an instruction conversion rule to obtain the first vehicle control instruction.
8. The vehicle control instruction determination method according to claim 7, characterized in that: The scenario arrangement model determines a candidate control instruction set based on the user intention information, and determines a first vehicle control instruction from the candidate control instruction set based on the user intention information and the first multimodal state tag data, including: The scenario arrangement model determines the candidate control instruction set based on the objective intention information in the user intention information and the first multimodal state label data; Based on the emotional intention information in the user intention information and the first multimodal status label data, determine the first vehicle control instruction whose priority and instruction execution method match the emotional intention information and the first multimodal status label data in the candidate control instruction set.
9. The vehicle control instruction determination method according to claim 7, characterized in that: Converting the candidate control instruction according to an instruction conversion rule to obtain the first vehicle control instruction includes: Performing synonym correction on the candidate control instruction to obtain a corrected candidate control instruction; The modified candidate control instruction is converted according to a standard instruction format to obtain the first vehicle control instruction.
10. A method for determining a vehicle control instruction, characterized in that: Applied to the vehicle terminal, the method includes: Acquire multimodal status data collected by multiple sensors installed on the vehicle; Performing multimodal preprocessing on the multimodal state data to obtain first multimodal state label data, and sending the first multimodal state label data to the cloud; Receiving a first vehicle control instruction from the cloud; The first vehicle control instruction is executed.
11. The vehicle control instruction determination method according to claim 10, characterized in that: The multimodal status data includes: image data and / or video data collected by a camera, sound data collected by a microphone, and perception data collected by a vehicle operating status and environment perception sensor; Performing multimodal preprocessing on the multimodal state data to obtain first multimodal state label data includes: Identify the image data and / or video data to obtain identification text information, and generate first label information based on the identification text information; Positioning a sound source of the sound data to obtain sound source information, and converting the sound source information into second label information; Converting the sensed data into third label information; A combination of the first tag information, the second tag information, and the third tag information is determined as the first multimodal state tag data.
12. The vehicle control instruction determination method according to claim 11, characterized in that: After executing the first vehicle control instruction, the method further includes: Obtaining an instruction execution result corresponding to the first vehicle control instruction; The instruction execution result is sent to the user.
13. A vehicle control instruction determination device, characterized in that: Applied to the cloud, the device comprises: A first acquisition module, used to acquire first multimodal state label data about a vehicle-mounted scene from a vehicle-mounted terminal, wherein the first multimodal state label data is determined based on multimodal state data collected by a plurality of sensors on the vehicle; an intention recognition module, configured to perform intention recognition on the user interaction information in the first multimodal state label data to obtain user intention information; A first determination module, configured to determine whether the user intention information includes an intention for vehicle control; The second determination module is used to determine the first vehicle control instruction based on the user intention information, the first multimodal state label data and the preset scenario arrangement model if the user intention information includes the intention for vehicle control, and send it to the vehicle terminal to control the vehicle.
14. A vehicle control instruction determination device, characterized in that: Applied to the vehicle terminal, the device comprises: A second acquisition module is used to acquire multimodal state data collected by multiple sensors installed on the vehicle; A preprocessing module, configured to perform multimodal preprocessing on the multimodal state data to obtain first multimodal state label data, and send the first multimodal state label data to the cloud; A receiving module, used for receiving a first vehicle control instruction from the cloud; An execution module is used to execute the first vehicle control instruction.
15. A vehicle, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the vehicle control instruction determination method described in any one of claims 1 to 9 or the vehicle control instruction determination method described in any one of claims 10 to 12 when executing the program stored in the memory.
Citation Information
Patent Citations
Multi-modal data collaborative man-machine interaction method and system and vehicle-mounted multimedia device
CN111737670A
Vehicle interaction method and device, electronic equipment, storage medium and vehicle
CN117235320A
Terminal equipment and corpus data generation method
CN118349844A
Intelligent interaction method and system applied to programs, electronic equipment and storage medium
CN118551846A
Vehicle control method, device, equipment and medium
CN119811396A
Cited By
Vehicle-mounted application control method and device, vehicle, medium and product
CN120697783A
A method, device, vehicle, medium, and product for vehicle-mounted applications.
CN120697783B
Vehicle control method, intelligent cabin and vehicle
CN120986324A
Vehicle direction control method, controller and automobile
CN121019609A