Voice interaction system for vehicle

By generating vehicle scene combination commands through in-vehicle voice adaptation units and cloud-based large models, the problem of users having difficulty efficiently adjusting multiple components in the car is solved, enabling the rapid generation of comfortable driving scenarios and improving the user experience.

CN120895036APending Publication Date: 2025-11-04FAW VOLKSWAGEN AUTOMOTIVE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511197755.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Users find it difficult to efficiently adjust various components in a vehicle to achieve a comfortable driving experience, and existing technologies cannot effectively utilize AI technology to improve the user experience.

Method used

The system uses an in-vehicle voice adaptation unit to receive voice commands and convert them into text commands. It then generates scene combination commands through a cloud-based large model to control in-vehicle applications and controllers. The in-vehicle scenario mode terminal generates and displays scene information and executes user commands.

Benefits of technology

It enables the rapid generation of in-vehicle scenes through voice input, reducing the time required for various automotive hardware and software controls and improving the user's driving experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895036A_ABST
    Figure CN120895036A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction system for a vehicle, and the system comprises a vehicle-mounted voice adaptation unit which is used for receiving a voice instruction inputted by a user, and converting the voice instruction into a text instruction; the cloud receives the character instruction through the voice adaptation unit and generates a scene combination instruction according to the character instruction by adopting a preset large model, and the scene combination instruction is configured to control the operation of the vehicle-mounted application and / or the vehicle-mounted controller and output the scene combination instruction to the voice adaptation unit; the vehicle-mounted contextual model terminal receives the scene combination instruction through a voice adaptation unit and generates scene information and inquiry information, and the vehicle-mounted contextual model terminal is configured to receive a scene opening instruction based on the inquiry information and input by a user; and in response to the scene opening instruction, connecting the corresponding vehicle-mounted application and / or the vehicle-mounted controller to output a scene combination instruction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cloud-based vehicle scene mode generation, in particular to a voice interaction system for a vehicle. BACKGROUND

[0002] With the continuous innovation of AI technology, AI empowers various fields to provide users with better experiences. In the current vehicle environment, users cannot efficiently adjust a large number of components in the vehicle to achieve a more comfortable effect in a short time because of the large number of software and hardware in the vehicle. By empowering the entire vehicle with AI, AI can better understand the user's voice input scene, generate a combination control scheme for the vehicle software and hardware through a large model algorithm, and control the corresponding components in the vehicle to achieve the user's desired scene, thereby greatly reducing the time for controlling multiple vehicle software and hardware and improving the user's vehicle experience. SUMMARY

[0003] To solve at least one aspect of the above technical problems, embodiments of the present application provide a voice interaction system for a vehicle, comprising:

[0004] a vehicle voice adaptation unit configured to receive a voice instruction input by a user and convert the voice instruction into a text instruction;

[0005] a cloud configured to receive the text instruction from the voice adaptation unit and generate a scene combination instruction based on the text instruction using a preset large model, the scene combination instruction being configured to control the operation of a vehicle application and / or a vehicle controller, and output the scene combination instruction to the voice adaptation unit;

[0006] a vehicle scene mode terminal configured to receive the scene combination instruction from the voice adaptation unit and generate scene information and inquiry information, the vehicle scene mode terminal being configured to receive a scene start instruction input by the user based on the inquiry information and connect the corresponding vehicle application and / or vehicle controller in response to the scene start instruction to output the scene combination instruction.

[0007] Preferably, the vehicle application includes a map application, a voice application, a navigation application, a music application, a maintenance application, a charging application, and a refueling application.

[0008] Preferably, the vehicle controller includes an air conditioner controller, a seat controller, an atmosphere lamp controller, a window controller, a door controller, a sunroof controller, and a steering wheel heater controller.

[0009] Preferably, the vehicle scene mode terminal comprises:

[0010] a listening unit configured to listen to the current vehicle state in real time and receive the scene combination instruction from the vehicle voice adaptation unit;

[0011] The mode trigger management unit is used to assess the scenario triggering conditions based on the vehicle status and output the scenario triggering assessment results.

[0012] The mode-to-interface unit is used to convert scene trigger evaluation result information and scene combination instructions;

[0013] The mode management unit generates scene modes based on scene combination instructions. A scene mode includes mode triggering conditions and scene combination instructions.

[0014] The pattern storage unit is used to store scene pattern information.

[0015] Preferably, the in-vehicle scenario mode terminal also includes an interface display unit, which receives scenario trigger evaluation result information and scenario combination instructions through the mode-to-interface unit to output and display them.

[0016] Preferably, the pattern storage unit includes a pattern data warehouse and a pattern memory, wherein the pattern data warehouse is used to cache scene pattern information and the pattern memory is used to store scene pattern information.

[0017] Preferably, the in-vehicle scenario mode terminal further includes:

[0018] The user interface is used to receive user input commands to activate the scene.

[0019] The execution interface logic unit responds to the scene start command and outputs the trigger condition and scene combination command according to the scene mode;

[0020] The task generation unit and the response and mode trigger management unit output the qualified scenario trigger evaluation results to generate task sequence information;

[0021] The task management unit generates a task distribution sequence based on the task sequence information.

[0022] The task distribution unit connects the in-vehicle application and / or the in-vehicle controller to execute the scenario mode according to the task distribution sequence.

[0023] Preferably, the task management unit receives the execution result through the vehicle controller, and the mode trigger management unit receives the execution result through the task management unit and outputs the execution result through the execution interface logic unit.

[0024] The voice interaction system for vehicles in this invention has the following technical effects: To input the user's desired in-vehicle scenario into a microphone via voice, the microphone transmits the voice information to the infotainment system host. The infotainment system host can connect to a cloud backend via a gateway. The cloud backend parses the user's information and generates a vehicle scenario solution, which is then provided to the infotainment system host (including vehicle scenario information from multiple apps and vehicle control units). The infotainment system host distributes the generated vehicle scenario information to the corresponding infotainment system apps and vehicle control units, achieving AI-powered intelligent scenario generation through multi-operation joint control of the entire vehicle. By introducing large-scale AI model capabilities into the cloud, it can intelligently analyze user needs and combine them with the vehicle's own capabilities to provide users with vehicle scenario-related services. Attached Figure Description

[0025] To better understand the above and other objects, features, advantages, and functions of the present invention, reference can be made to the embodiments shown in the accompanying drawings. The same reference numerals in the drawings refer to the same parts. Those skilled in the art should understand that the drawings are intended to schematically illustrate preferred embodiments of the invention and are not intended to limit the scope of the invention; the parts in the drawings are not drawn to scale.

[0026] Figure 1 A schematic diagram illustrating an application scenario of a voice interaction system for a vehicle according to an embodiment of the present invention is shown.

[0027] Figure 2 A schematic diagram of the scene mode generation process for a voice interaction system for a vehicle according to an embodiment of the present invention is shown.

[0028] Figure 3 A schematic diagram of the scene mode execution flow of a voice interaction system for a vehicle according to an embodiment of the present invention is shown. Detailed Implementation

[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0030] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0031] To at least partially address one or more of the aforementioned problems and other potential issues, embodiments of this disclosure propose an embodiment of the present invention providing a voice interaction system for vehicles, comprising: an in-vehicle voice adaptation unit, an in-vehicle scenario mode terminal, and a cloud. The in-vehicle voice adaptation unit and the in-vehicle scenario mode terminal are mounted on the infotainment system host (HU) of the vehicle's infotainment system. The HU connects to the cloud via a gateway using a network communication protocol (e.g., HTTP / MQTT). The cloud includes an AIBE (AI-based education) cloud AI engine, which provides software infrastructure for training and inference of pre-set large models.

[0032] The in-vehicle voice adaptation unit is used to receive voice commands input by the user and convert the voice commands into text commands.

[0033] Specifically, such as Figure 1 As shown, the user inputs voice commands through the vehicle's microphone, and the vehicle's voice adaptation unit converts the voice commands into text.

[0034] The cloud receives text commands through the voice adaptation unit and generates scene combination commands based on the text commands using a preset large model. The scene combination commands are configured to control the operation of in-vehicle applications and / or in-vehicle controllers, and the scene combination commands are output to the voice adaptation unit.

[0035] Specifically, the pre-built cloud-based model possesses deep semantic understanding capabilities, enabling it to understand spoken language, separate multiple intents, and generate corresponding scenario-based instruction combinations based on the user's intents. For example, the model can recognize the user's voice and decompose it into specific control tasks. It can also invoke multiple intelligent agent modules to handle different control tasks. For instance, by parsing the user's voice input, it can generate multiple user intents and their corresponding skills, and then invoke the appropriate intelligent agent modules to complete the tasks. The scenario-based instruction combinations generated in the cloud are output to the vehicle via the voice adaptation unit.

[0036] The cloud-based pre-defined model determines whether a text command is a pre-defined command. If the text command is not a pre-defined command, it generates a scene combination command based on the text command. If the text command is a pre-defined command, it directly calls the scene combination command corresponding to the pre-defined command and outputs it to the vehicle's infotainment system. Pre-defined commands include in-vehicle commands and user-stored commands.

[0037] The in-vehicle scenario mode terminal receives scenario combination commands and generates scenario information and query information through the voice adaptation unit. The in-vehicle scenario mode terminal is configured to receive scenario activation commands based on query information input by the user, and in response to the scenario activation commands, connect to the corresponding in-vehicle application and / or in-vehicle controller to output scenario combination commands.

[0038] Specifically, the vehicle scene mode terminal is a vehicle scene mode application (Vehicle Scene Model APP). It receives scene combination commands output from the cloud through a voice adapter unit (SDS Adapter) and generates scene information and query information. The query information is used to output to the user to prompt the user whether the scene is activated. When the user inputs a scene activation command, the vehicle scene mode terminal connects to the corresponding vehicle application and / or vehicle controller to output scene combination commands.

[0039] The in-vehicle scenario mode terminal breaks down scenario information and distributes it to software apps such as Navi and Media on the vehicle's infotainment system, passing vehicle control functions to the vehicle control interface (AndroidCarAPI / CnsExtAPI). The vehicle control interface then transmits this information to specific vehicle control actuators via the Car Control Service, through methods such as CAN / Enternet.

[0040] In some embodiments, in-vehicle applications include map applications, voice applications, navigation applications, music applications, maintenance applications, charging applications, and refueling applications.

[0041] Specifically, the scenario combination instructions include application instructions, which are used to control the corresponding in-vehicle application software functions.

[0042] In some embodiments, the vehicle controller includes an air conditioning controller, a seat controller, an ambient lighting controller, a window controller, a door controller, a sunroof controller, and a steering wheel heater controller.

[0043] Specifically, the scenario combination instructions include vehicle control instructions, which are used to control the corresponding vehicle controller to perform corresponding operations, such as temperature adjustment, seat adjustment, and window adjustment.

[0044] Example 1

[0045] like Figure 1As shown, the user inputs a voice command through the microphone: "It's cold today, can you help me warm up?"

[0046] The user's speech is converted into text by the Speech Adaptor Unit (SDS Adapter layer), and the text is then transmitted to the FTB cloud.

[0047] The cloud parses the received user commands (text commands), identifies the command as a large model command (the voice is not a preset voice command, but can be considered a large model command), and transmits the command to the AIBE preset large model backend for parsing. After parsing, the preset large model algorithm provides a solution that conforms to the vehicle scenario, such as turning on the air conditioner, setting the air conditioner temperature to 30℃, closing the windows, closing the sunroof, turning on the seat heating, turning on the steering wheel heating, and other scenario combination commands.

[0048] The AIBE large model backend generates combined instructions and transmits them to the vehicle's infotainment system (HU) through the gateway. The HU then transmits the information to the in-vehicle scene model APP through the SDS Adapter layer, generating the corresponding scene and asking the user whether to activate the scene.

[0049] Upon receiving a user's command to activate a scene, the in-vehicle scene mode terminal will send the corresponding software functions to music, navigation, and other software apps. The corresponding vehicle control functions will be transmitted to the relevant ECUs such as air conditioning control, seat control, and ambient lighting control via the vehicle control service (Android Car API / Cns Ext API) through the CAN bus or Ethernet. Ultimately, this will enable one-click execution of the combined commands.

[0050] In some embodiments, the in-vehicle scenario mode terminal includes: a monitoring unit for real-time monitoring of the current vehicle status and receiving scenario combination instructions through an in-vehicle voice adaptation unit; a mode trigger management unit for evaluating scenario trigger conditions based on the vehicle status and outputting scenario trigger evaluation results; a mode-to-interface unit for converting scenario trigger evaluation result information and scenario combination instructions; a mode management unit for generating scenario modes based on scenario combination instructions, wherein the scenario mode includes mode trigger conditions and scenario combination instructions; and a mode storage unit for storing scenario mode information.

[0051] Specifically, the in-vehicle scenario mode terminal creates modes by combining a listening unit, a mode trigger management unit, a mode to interface unit, a mode management unit, and a mode storage unit with the cloud.

[0052] like Figure 2As shown, the user inputs "create a suitable rest mode" via voice command. The user's words will enter the SDS Adapter voice adaptation layer, and then be parsed by the cloud's large model to generate scene modes (including scene combination commands) and returned to the voice adaptation layer.

[0053] The scene information received by the in-vehicle voice adaptation unit is monitored by the monitoring unit, which monitors the current vehicle status in real time, such as vehicle speed and gear position.

[0054] The monitoring unit distributes vehicle status and scene modes to the Trigger Manager.

[0055] The mode trigger management unit determines whether the current vehicle state meets the trigger conditions of the mode. For example, it is not suitable to generate a mode when the gear is in D and there is speed. It also arranges the order of mode execution, such as closing the windows first, then turning on the air conditioner and folding the seat, and passes the result of whether the trigger is met to the mode to interface unit (AI ViewModel).

[0056] The mode-to-interface unit displays the triggered scene mode information on the interface and passes the information to the AIFragment interface presentation layer for display. The AI ​​View Model mode-to-interface layer also passes the conditions triggered by the mode and related parameters to the mode management unit (Model Factory) for mode management, and the generated mode is passed to the mode storage unit for storage.

[0057] In some embodiments, the in-vehicle scenario mode terminal further includes an interface display unit, which receives scenario trigger evaluation result information and scenario combination instructions through the mode-to-interface unit to output and display them.

[0058] Specifically, the interface display unit (AI Fragment) is used to receive and display scene trigger evaluation result information and scene combination instruction information.

[0059] In some embodiments, the pattern storage unit includes a pattern data warehouse and a pattern memory, wherein the pattern data warehouse is used to cache scene pattern information and the pattern memory is used to store scene pattern information.

[0060] Specifically, the Mode Repo caches patterns for easy current access, while the Mode Memory (DB) persists patterns for future use by users.

[0061] In some embodiments, the in-vehicle scenario mode terminal further includes: a user interface for receiving a scenario activation command input by a user; an execution interface logic unit that, in response to the scenario activation command, outputs triggering conditions and scenario combination commands according to the scenario mode; a task generation unit that, in response to the scenario trigger evaluation result that meets the conditions output by the mode trigger management unit, generates task sequence information; a task management unit that generates a task distribution sequence according to the task sequence information; and a task distribution unit that, according to the task distribution sequence, connects to the in-vehicle application and / or the in-vehicle controller to execute the scenario mode. In some embodiments, the task management unit receives the execution result through the in-vehicle controller, and the mode trigger management unit receives the execution result through the task management unit and outputs the execution result through the execution interface logic unit.

[0062] Specifically, such as Figure 3 The diagram shows the sequence of events for mode execution. For example, based on the previously created rest mode, the user activates the rest mode by clicking the user interface (Fragment) and inputting a scene activation command. The mode is triggered to the execution interface logic unit (View Model). The execution interface logic unit then passes the trigger conditions and the action to be executed (scene combination command) to the mode trigger management unit (Trigger Manager). The mode trigger management unit determines whether execution is possible based on the trigger conditions. If the conditions are met, it generates a task sequence (Task) and sends it to the task generation unit (Task Factory). The task generation unit passes the arranged task execution order to the task management unit (TaskManager). The task management unit arranges the task execution order, such as turning on the air conditioner first, then adjusting the fan speed, etc. The task management unit then passes the arranged order to the task dispatch unit (Task Dispatcher). The task dispatch unit distributes the executed operations to the specific Car Controll mpl vehicle control execution unit, thereby executing the corresponding vehicle control operation. The CarControl1mpl vehicle control execution unit returns the scenario mode execution result to the task management unit. The task management unit then sends the execution result to the mode trigger management unit, and finally, the execution interface logic is sent to the interface for display. The monitoring unit continuously monitors vehicle signals and sends the detected signals back to the mode trigger management unit to determine whether the scenario meets the trigger conditions.

[0063] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand this document.

Claims

1. A voice interaction system for vehicles, characterized in that, include: The in-vehicle voice adapter unit is used to receive voice commands input by the user and convert the voice commands into text commands; In the cloud, text commands are received through the voice adaptation unit, and scene combination commands are generated based on the text commands using a preset large model. The scene combination commands are configured to control the operation of the vehicle application and / or the vehicle controller, and the scene combination commands are output to the voice adaptation unit. The in-vehicle scenario mode terminal receives scenario combination commands and generates scenario information and query information through a voice adaptation unit. The in-vehicle scenario mode terminal is configured to receive scenario activation commands based on query information input by the user, and in response to the scenario activation commands, connect to the corresponding in-vehicle application and / or in-vehicle controller to output scenario combination commands.

2. The system according to claim 1, characterized in that, In-vehicle applications include map applications, voice applications, navigation applications, music applications, maintenance applications, charging applications, and refueling applications.

3. The system according to claim 1, characterized in that, The vehicle controllers include air conditioning controllers, seat controllers, ambient lighting controllers, window controllers, door controllers, sunroof controllers, and steering wheel heater controllers.

4. The system according to claim 1, characterized in that, The in-vehicle scenario mode terminal includes: The monitoring unit is used to monitor the current vehicle status in real time and receive scene combination commands through the in-vehicle voice adaptation unit. The mode trigger management unit is used to assess the scenario triggering conditions based on the vehicle status and output the scenario triggering assessment results. The mode-to-interface unit is used to convert scene trigger evaluation result information and scene combination instructions; The mode management unit generates scene modes based on scene combination instructions. A scene mode includes mode triggering conditions and scene combination instructions. The pattern storage unit is used to store scene pattern information.

5. The system according to claim 4, characterized in that, The in-vehicle scenario mode terminal also includes an interface display unit. The interface display unit receives scenario trigger evaluation result information and scenario information through the mode-to-interface unit and outputs the display.

6. The system according to claim 5, characterized in that, The pattern storage unit includes a pattern data warehouse and a pattern memory. The pattern data warehouse is used to cache scene pattern information, and the pattern memory is used to store scene pattern information.

7. The system according to claim 6, characterized in that, The in-vehicle scenario mode terminal also includes: The user interface is used to receive user input commands to activate the scene. The execution interface logic unit responds to the scene start command and outputs the trigger condition and scene combination command according to the scene mode; The task generation unit and the response and mode trigger management unit output the qualified scenario trigger evaluation results to generate task sequence information; The task management unit generates a task distribution sequence based on the task sequence information. The task distribution unit connects the in-vehicle application and / or the in-vehicle controller to execute the scenario mode according to the task distribution sequence.

8. The system according to claim 7, characterized in that, The task management unit receives the execution results through the vehicle controller, and the mode trigger management unit receives the execution results through the task management unit and outputs the execution results through the execution interface logic unit.

Citation Information

Patent Citations

  • Vehicle-mounted voice instruction recommendation method and device and model training method

    CN115547302A

  • Question and answer method, device and system and vehicle

    CN120179762A

  • Scene generation method and device, equipment and storage medium

    CN120510842A

  • Gaze detection using one or more neural networks

    US20210056306A1

  • Automobile control method and apparatus, computer device, and storage medium

    WO2023024680A1