Method for operating a voice assistant system in a vehicle
By spatially directing acoustic responses based on speaker position, the method enhances user engagement and reduces distraction in vehicle voice assistant systems, offering an immersive interaction experience.
Patent Information
- Application Number
- DE102024003531
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-12-24
- Estimated Expiration
- 2044-10-25
AI Technical Summary
Existing voice assistant systems in vehicles lack the ability to provide immersive and localized audio responses, leading to potential user distraction and a less engaging experience.
A method that determines the position of a speaker within the vehicle and spatially directs acoustic responses based on spoken content, using audio animations and three-dimensional audio output to create an immersive interaction experience.
Enhances user engagement by providing localized and moving audio outputs, reducing distraction and creating a more interactive and realistic interaction with the voice assistant system.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for operating a voice assistant system in a vehicle.
[0002] From WO 2017 / 139533 A1, a system for controlling multiple entertainment systems and speakers via voice commands is known. The system receives voice commands and identifies speakers that play audio output in the vicinity of the commands. The system generates a voice output and sends it to the speakers along with a command to reduce the volume of the audio output during playback. Furthermore, the system sends a command to the speakers while transmitting the voice output to another device for playback. The system reduces the audio output generated by the speakers and plays the voice output through the input device.
[0003] Furthermore, DE 10 2016 114 413 A1 describes a device for generating object-dependent audio data and a method for generating object-dependent audio data in a vehicle interior. The method provides that information and warning signals originating from driver assistance systems and their sensors are processed by an object-based audio generator, which exchanges data with a database of audio data, into object-based sounds and reproduced through loudspeakers located in the vehicle interior.
[0004] Furthermore, DE 10 2020 003 922 A1 describes a voice control system in a motor vehicle with seats arranged in several rows, each equipped with microphones for receiving voice commands from occupants. The voice control system comprises an evaluation unit designed to analyze a voice command transmitted by the microphones and determine data for generating a voice output. The system also includes an output unit that generates a voice output from the data received by the evaluation unit and transmits it to the loudspeakers assigned to the seats in response to a voice command. The system further includes a separate output unit for generating a voice output directed at a seat in a rear row, distinct from the one used for generating a voice output directed at a seat in a front row.
[0005] US 2020 / 0111489A1 describes a device with - a microphone that picks up audio signals in a vehicle cabin, - a loudspeaker that outputs audio signals into the vehicle cabin, - an interpreter who interprets the meaning of the audio signals picked up by the microphone, - a display provided in the vehicle cabin and - a controller that displays an agent image in the form of an address to an inmate in an area of the display and causes the speaker to emit audio signals by which the agent image speaks to an inmate, wherein the controller changes a face direction of the agent image to a direction different from the direction of the inmate who is a conversation target, in the event that an utterance in respect of face direction is interpreted by the interpreter after the agent image has been displayed on the display.
[0006] The invention is based on the objective of providing a method for operating a voice assistant system in a vehicle and a voice assistant system.
[0007] The problem is solved according to the invention by a method which has the features specified in claim 1.
[0008] Advantageous embodiments of the invention are the subject of the dependent claims.
[0009] In the inventive method for operating a voice assistant system in a vehicle, the position of a speaker in the vehicle is determined based on detected acoustic signals, and the speaker receives an acoustic response in reaction to their spoken content. If spatial audio output is available, an acoustic response from the voice assistant system, related to the speaker's content and / or occupant-related, is spatially directed towards a vehicle area and / or an occupant, wherein an audio signal to be output as a response from the voice assistant system is linked to position information of the speaker and / or a vehicle area.
[0010] By applying this method, a vehicle occupant, especially a speaker, can experience an immersive exchange of information between themselves and the voice assistant system. For example, the method makes it possible for a voice assistant to move audibly through the vehicle or to provide the speaker with targeted, acoustically localizable directions to objects within the vehicle.
[0011] Localization of speech output can also be used to support the speaker, so that the speaker is less distracted in comparison.
[0012] Furthermore, audio animations corresponding to generated text-to-speech prompts are created for the output of the audio signal as audio data, according to a predefined duration. This gives the speaker, i.e., the user of the voice assistant system, the feeling of speaking to a real person as a voice assistant who answers their questions and / or provides information on desired topics. Moreover, such audio animation, particularly depending on the spoken content of the speaker, can create the impression that a voice assistant is in the vehicle and, for example, moves away from the user to open a car window.
[0013] In particular, spatial audio makes it possible to animate sound sources in three-dimensional space. The location of such a sound source at a specific time can be represented by corresponding coordinates.
[0014] Audio animation, therefore, refers to a description of movement and / or loudness and / or other changing properties of the sound source in space over a defined period of time.
[0015] Furthermore, for at least individual text-to-speech prompts, or at least for a segment of a text-to-speech prompt, it is specified in which area of the vehicle the output begins and in which area the audio animation output ends. For example, it can be specified that a speech output initially moves towards the speaker, then travels within the vehicle, giving the speaker the impression that the voice assistant is in the vehicle and moves from a position facing the speaker to a position in the rear of the vehicle in order to comply with a request from the speaker, such as opening the vehicle window.
[0016] In one implementation, the audio signal is selected in response to spoken content based on audio files stored in a database and then output in the vehicle. This process checks how closely the content spoken by a speaker matches the stored audio files in order to react optimally to the spoken content.
[0017] In another implementation, the audio signal is selected in response to spoken content based on audio files stored in a central computer unit that is linked to the vehicle or can be linked to it, and then output in the vehicle. This allows for a significantly larger number of audio files to be used, enabling an optimized response to the speaker's content and a satisfactory acoustic output in the vehicle.
[0018] In one possible implementation, the audio files are decoded and played back using a prompter function unit of the voice assistant system. The audio files are transferred from the central computer unit to a main control unit of the vehicle and / or passed to the main control unit as a link to a data stream address. Specifically, a final result regarding the speaker's content is processed using a dialogue management system within the voice assistant system.Dialogue management decides, based on recognized intents, which represent a task and / or an action that a voice assistant system performs for a speaker (i.e., user), and slots, which denote variables that are passed from a speaker to the voice assistant system in the context of an intent, and based on dialogue resources, which describe possible dialogue sequences, including so-called text prompts that are read aloud to the speaker (i.e., user), as well as on the current dialogue status, which action should be executed next.
[0019] In another possible implementation, a start value is assigned to a vehicle area at the beginning of the output, and an end value is assigned to a vehicle area at the end of the audio animation output. Based on the start and end values and information on whether a position of the vehicle area is to be evaluated relative or absolute to a speaker's position, a start vector and an end vector are derived, which are stored as metadata for the respective text data and / or terms in a database for dialogue resources located on the vehicle side and / or in the central computing unit.
[0020] Thus, in particular, a distance between the vehicle areas can be determined, for example taking into account further intermediate values, so that the output of an audio animation, especially its runtime, can be determined.
[0021] In a further embodiment of the procedure, the duration of the audio file is determined using the prompter functional unit, and an audio animation and / or a data stream of individual vectors is generated by interpolating further vectors between the start vector and the end vector for the audio animation, corresponding to a runtime of the audio file.
[0022] Furthermore, in one version, an audio data stream and / or an audio animation data stream is transmitted via the prompter functional unit to an audio amplifier of an audio system of the vehicle, and the audio data stream and / or the audio animation data stream is converted into an audio data stream by means of the audio amplifier, and the respective audio data stream is output in the vehicle via vehicle-side loudspeakers.
[0023] In a further development of the method, a machine-learning-pre-trained language model is used to generate an audio data stream and / or an audio animation data stream. Such a language model is a computational linguistic probability model that has learned statistical word and sentence sequence relations from a large number of text documents through computationally intensive training.
[0024] Exemplary embodiments of the invention are explained in more detail below with reference to drawings.
[0025] This shows: Fig. 1 schematically a vehicle with a voice assistant system and central computer unit, Fig. 2. Schematically illustrates a process flow based on components of the voice assistant system and Fig. 3 schematically illustrates a process flow based on components of the voice assistant system with audio animations matching generated text-to-speech prompts in relation to a runtime.
[0026] Corresponding parts are marked with the same reference symbols in all figures.
[0027] Fig. Figure 1 shows, in a highly simplified example, a vehicle 1 with a voice assistant system 2 and a central computing unit 3 connected to the vehicle 1, in particular to the voice assistant system 2, via data transmission. Fig. Figure 2 shows a procedure for operating the voice assistant system 2 based on its components.
[0028] The voice assistant system 2 is part of a main control unit 4 of the vehicle 1
[0029] According to the in Fig. In the embodiment shown in Figure 1, four microphones 5 and four loudspeakers 6 are arranged regularly in the vehicle 1, wherein the microphones 5 and the loudspeakers 6 are connected to an audio amplifier 7, a so-called amplifier.
[0030] In general, a voice assistance system 2, in particular a voice dialogue system or a voice assistant, is known for a vehicle 1. In this system, information and / or follow-up questions are typically output to a user, i.e., a speaker, in response to their spoken content, both acoustically, based on a text-to-speech function, and visually in a display area as a graphical user interface of a display unit, in particular the main control unit 4.
[0031] Furthermore, the position of a speaker, i.e., a user of the voice assistant system 2, is determined based on the input signals recorded by a respective microphone 5. In addition, an audio system 8, as part of the main control unit 4, is designed to output audio signals directed within the vehicle 1.
[0032] Prior art includes voice assistant systems 2 in which text-to-speech outputs related to a voice dialogue between the user and the voice assistant system 2 are directed based on the determined position of the user (i.e., the speaker) and a defined loudspeaker-dependent audio output zone in relation to the speaker. This allows the system to identify which occupant of the vehicle 1 is currently using the voice assistant system 2, thus enabling a personalized user experience.
[0033] Furthermore, an audio playback method known as spatial audio or 3D audio is available. This audio playback method makes it possible to create an immersive user experience. For example, using such an audio playback method, a voice assistant, with which the user interacts via voice, can acoustically move through the vehicle 1 and / or the user can be specifically directed to objects within the vehicle 1 using acoustic localization.
[0034] For example, if a passenger, acting as the user, asks the voice assistant system 2, "Hello vehicle, open the rear left window!", the response will be, "Okay, I'll just go back." The voice assistant will then audibly move from a determined position of the speaker to the rear left window to be opened. This creates an immersive experience for the user when using the voice assistant system 2. Similarly, if the user asks the voice assistant, "Hello vehicle, where do I turn on the hazard warning lights?", it can answer, "Here, below your screen." The audible output of the voice assistant system 2 will be audibly delivered to the user in the immediate vicinity of the hazard warning light activation button on the vehicle.The localization of speech output is extended by an additional modality to support the user and reduce distraction. For example, if there are multiple occupants in vehicle 1, the voice assistant can audibly move from one occupant to the next, such as during a quiz game. The voice assistant can whisper questions or information into the ear of each occupant for each round of the game.
[0035] The following describes a function of prior art known voice assistance systems 2. The microphones 5 arranged in the vehicle 1 convert sound into acoustic signals S, which are fed to the audio amplifier 7 of a signal preprocessing unit (SPP) via reference channels. This signal preprocessing unit cleans the input signals from the microphones 5 and determines the position of a speaker in the vehicle 1 based on sound pressure level and time-of-arrival differences. One or more output signals from the signal preprocessing unit (SPP) are fed to a functional unit of an automatic speech recognition (ASR) system, which processes the respective output signal from the signal preprocessing unit (SPP) and converts it into text and / or another machine-readable format, whereby several hypotheses can be output. In particular, the cleaned output signal represents an audio signal.
[0036] Building upon this, intents and slots are classified using a functional unit for computational linguistic processing of natural language (NLU). A slot is a variable that a user passes to a voice assistant system 2 within the context of an intent. Slots provide the voice assistant system 2 with several pieces of information about an intent. An intent is a task or action that the voice assistant system 2 performs or is intended to perform for a user.
[0037] At the same time, the cleaned audio signal is transmitted, if necessary, to the central computing unit 3 and processed there in a comparable manner, so that a result relating to an output of the voice assistant system 2 is also provided by means of the central computing unit 3.
[0038] If the cleaned audio signal is transmitted to the central computing unit 3 and processed there, in a next step an arbitration A takes place between the results in order to determine which of the results is to be used, i.e. output, as a reaction of the voice assistant system 2 to the spoken content of the user.
[0039] A final result is processed using a Dialog Management System (DLM), which, based on recognized intents and slots stored in a database (DB), determines which action is executed next. This DLM uses Dialog Resource Specifications (DRS) to describe possible dialog flows, including text prompts that are read aloud to the user, and the current dialog status. An action can be an audio and / or graphical output to the user, or a system action, such as a user interface call.
[0040] If the Dialog Management (DLM) decides that a prompt should be displayed, a suitable output text (AT) is selected from the Dialog Resources (DRS) as a so-called concept and sent to a Text-to-Speech (TTS) function unit. This TTS unit then uses the selected output text (AT) as its input string to generate an audio file (AD). This audio file (AD) is played back by a Prompter Function Unit (PTR), which communicates with an audio layer, and output acoustically via the loudspeakers (6) present in the vehicle (1). The term "audio layer" refers to audio management, audio handling, and other software components within the main control unit (4) that are responsible for the actual audio output.
[0041] Since vehicles can have different equipment configurations, it is common for some components to need to be configured at system startup, depending on the various equipment options. This can be done using coding and / or a so-called feature toggle.
[0042] Individual speech dialogues are modeled and designed using a utility program D for dialogue designers. This allows, for example, the definition of individual dialogue sequences for specific intent / slot combinations using decision tables. The output texts AT, which are to be displayed when a specific condition occurs, are also defined. Utility D generates dialogue resource files, which the dialogue management system (DLM) later uses at runtime to derive the next dialogue step. Runtime refers to the duration during which the software applications of the main control unit (MCU) 4 are executed. The voice assistant system 2 is switched on, and its software components are activated. In this specific case, a user interacts with the voice assistant system 2, and the dialogue management system (DLM) accesses the dialogue resource data.This dialog resource data was generated outside of main control unit 4 using utility D and stored in main control unit 4, for example, via a software update. Therefore, the dialog resource data is not generated at runtime.
[0043] The following describes various implementation options for generating an immersive user experience when using the voice assistant system 2. For all implementation options, it is important to note that the vehicle 1's configuration is checked upon system startup. Specifically, this check determines whether the vehicle 1 has an audio system 8 with three-dimensional audio output.
[0044] If the vehicle 1 does not have such an audio system 8, the acoustic output of the voice assistant system 2 is based on the text-to-speech function. Since such information may also need to be processed in the central computing unit 3, corresponding metadata is transmitted from the main control unit 4 to the central computing unit 3 when the central computing unit 3 requests it.
[0045] One initial implementation envisions using existing audio files (AD) that have been specifically mixed and encoded for such a three-dimensional audio playback method. These audio files (AD) can be stored locally in the database (DB) of the main control unit 4.
[0046] If the voice assistant system 2 detects an intent / slot combination, the dialog management system DLM decides to transfer the audio file AD to the audio system 8, a so-called onboard player. A voice dialog ends, so the audio system 8 starts playing the audio file AD.
[0047] This first version can be implemented in vehicle 1 with minimal effort, enabling simple three-dimensional playback of audio files AD when using the voice assistant system 2.
[0048] For example, this embodiment can be used to implement a request to the voice assistant system 2, which reads: “Hello vehicle, what does a helicopter sound like?”.
[0049] In a second version, the audio files AD can be transferred from the central computing unit 3 to the main control unit 4 or passed as a reference to a data stream address. This allows new use cases to be implemented in the central computing unit 3 without additional embedding changes to the voice assistant system 2.
[0050] A third option involves using existing audio files (AD) that have been specifically mixed and encoded for the three-dimensional audio system 8. These audio files (AD) can be stored locally in the database (DB) of the main control unit 4. The prompter (PTR) function unit of the voice assistant system 2 then decodes and plays back the audio files (AD).
[0051] If the corresponding intent / slot combination is recognized by the voice assistant system 2, the dialog management DLM decides to decode and play back the audio file AD using the prompter functional unit PTR.
[0052] In this case, the Prompter Function Unit PTR of the voice assistant system 2 takes over the connection to the audio layer.
[0053] In a fourth version, the audio files AD can be transferred from the central computer unit 3 to the main control unit 4, or passed as a reference as a link to a data stream address.
[0054] In a Fig. In the fifth iteration shown in Figure 3, audio animations are generated at runtime to match generated text-to-speech prompts. This fifth iteration is comparatively complex, but offers a way to optimally utilize the potential of the three-dimensional audio system 8 in conjunction with the voice assistant system 2.
[0055] To achieve this, the utility program D for designing dialogues is first enhanced with an additional function that allows dialogue designers to define, for individual text-to-speech prompts or segments thereof, in which area of the vehicle the audio animation should begin and end. For example, a simple XY coordinate system can be used to define a position as a top-down view within vehicle 1. Additionally, a height (z-coordinate) within vehicle 1 can be set using a slider. The coordinates thus provide a bird's-eye view of vehicle 1. If the origin (zero point) is located in the center of vehicle 1 and the axes each have a maximum value from 10 to -10, then the position of a driver in vehicle 1, especially in a left-hand drive vehicle, corresponds to approximately X = -5 and Y = 5.
[0056] Furthermore, it is possible to define whether one of these values, in particular the XY coordinates, is to be set relative to the speaker's position or is to be located absolutely within vehicle 1. Vectors V are derived from start and end values, and possibly further intermediate values, along with information on whether the position is relative or absolute to the speaker's position. These vectors are stored as metadata for the respective output texts AT in the dialog resources DRS.
[0057] If a text-to-speech prompt is to be output, the output text AT is passed to the text-to-speech functional unit TTS, which generates an audio file AD.
[0058] In the next step, the prompter function unit PTR is called, and essentially the audio file AD and the associated vectors V for the audio animation are passed to PTR. PTR then determines the duration of the audio file AD and generates an audio animation, or rather a stream of individual vectors V. This is achieved by interpolating additional vectors V between the provided start, intermediate, and end vectors, corresponding to the runtime of the audio file AD. The rate of the generated vectors V is configurable to adapt to different three-dimensional audio formats. This rate refers to the sample rate, i.e., the frequency with which the vectors V are generated, and specifically sent, per second. For example, the rate might be 60 Hz.
[0059] Both data streams, that is, an audio data stream and an animation data stream, are transferred to the audio layer via the prompter functional unit PTR. From there, the corresponding data streams reach the audio amplifier 7, which converts them into audio streams that are output via the loudspeakers 6 in the vehicle 1. Thus, all use cases specified by utility D can be extended with the three-dimensional audio playback method.
[0060] In addition, there are speech dialogues that take place in the central computer unit 3 and whose output texts cannot be predefined by AT, such as: "Hello vehicle, what can you tell me about "XY"?".
[0061] To make such dynamic content perceptible to a user as a sixth iteration of the three-dimensional audio playback process, dynamically generated output texts AT are provided with metadata in the central processing unit 3, which includes at least a start and end vector. If these output texts AT are to be displayed using the text-to-speech function unit TTS, this data is processed as described previously. However, in this case, the vectors V from a response of the central processing unit 3 are passed to the prompter function unit PTR.
[0062] Additionally, there are use cases where text-to-speech output is generated as a seventh execution in the central computing unit 3 and then transmitted to the main control unit 4.
[0063] This is particularly useful for long, dynamic prompts, as the central processing unit can employ more comprehensive text-to-speech models, often resulting in noticeably higher quality. These data streams are output via the Prompter Function Unit (PTR) of the main control unit 4. To implement such use cases using the three-dimensional audio playback method, another parallel data stream is initiated from the central processing unit 3, containing the audio animation vectors (V) corresponding to the text prompt. A multiplex of both signals is also conceivable.
[0064] Both the audio and animation data streams are received and synchronized by the Prompter functional unit PTR. Output then occurs as described previously.
[0065] Another approach involves using a language model pre-trained using machine learning.
[0066] When using a speech model, it can be assumed that the text prompts read aloud to the user are generated completely dynamically. Therefore, predefining audio animations is either not possible or only possible to a very limited extent.
[0067] Accordingly, the language model must be able to understand the start, intermediate, and end points of an audio animation in order to provide appropriate animation points along with the solution to a given task. This requires providing the language model with a suitable context. This context can, for example, include various static pieces of information about specific parts in vehicle 1 and their positions within the vehicle. This includes, in particular, vehicle windows, seats, and / or controls and operating elements. Furthermore, the context can include dynamic information, such as the number and placement of occupants in vehicle 1.
[0068] If it is an embedded solution, this information is queried at runtime and prepared for use by the language model. If the language model is intended to be used in the central computing unit 3, this dynamic information is sent from the main control unit 4 to the central computing unit 3 upon request and can then be processed there.
[0069] If this dynamic information is available to the language model, it is possible to add a subtask to the task assigned to the language model, which consists of describing how a sound source, i.e., an audio source, should behave during an output in accordance with content, i.e., a generated output prompt, for example: "How could a suitable animation of an audio output in vehicle 1 be designed to match the answer?".
[0070] Since it can be assumed that the language model cannot provide vectors V, but rather a description, the answer for this subtask is subsequently analyzed again. Keywords are compared with a pre-existing data pool, and if a match is found, they are replaced by vectors V from information already present in the data pool.
[0071] A character string intended for the response, a so-called string, is provided to the prompter functional unit PTR along with the vectors V. A comparatively large number of transmission paths are conceivable here, depending on the embedding and / or solution using the central computing unit 3. The audio data AD is then processed and output by the prompter functional unit PTR as in the previously described procedures. Reference symbol list 1 vehicle 2 Voice Assistant System 3 central computing unit 4 Main control unit 5 microphones 6 speakers 7 audio amplifiers 8 Audio system Arbitration AD audio file ASR automatic speech recognition AT output text D Utility DB database DLM Dialogue Management DRS Dialogue Resources NLU functional unit for computational linguistic processing of natural language PTR Prompter Functional Unit S acoustic signal SPP Signal Preprocessing TTS Text-to-Speech Function Unit V vector
Claims
[1] Method for operating a voice assistant system (2) in a vehicle (1), wherein - based on recorded acoustic signals (S) a position of a speaker in the vehicle (1) is determined and the speaker receives an acoustic response in reaction to his spoken content, - if spatial audio output is available, an acoustic response of the voice assistant system (2) is spatially directed towards a vehicle area and / or an occupant, which is related to the content of the speaker, - an audio signal to be output as a response of the voice assistant system (2) is linked with position information of the speaker and / or a vehicle area, - for the output of the audio signal, audio data are generated as audio animations corresponding to generated text-to-speech prompts according to a specified runtime and - at least for individual text-to-speech prompts or at least for a segment of a text-to-speech prompt, it is specified in which vehicle area an output begins and in which vehicle area the output of the audio animation ends. [2] Method according to claim 1, characterized by , that the audio signal is selected in response to spoken content based on audio files (AD) stored in a database (DB) and output in the vehicle (1). [3] Method according to claim 1 or 2, characterized by , that the audio signal is selected in response to spoken content based on audio files (AD) stored in a central computer unit (3) which is linked or linkable to the vehicle (1) and is output in the vehicle (1). [4] Method according to claim 3, characterized by, that the audio files (AD) are decoded and played back by means of a prompter functional unit (PTR) of the voice assistant system (2), wherein the audio files (AD) are transferred from the central computing unit (3) to a main control unit (4) of the vehicle (1) and / or are passed to the main control unit (4) as a link to a data stream address. [5] Method according to any one of the preceding claims, characterized by, that a vehicle area at the beginning of the output is assigned a start value and a vehicle area at the end of the audio animation output is assigned an end value, whereby a start vector and an end vector are derived based on the start value and the end value and information on whether a position of the vehicle area is to be evaluated relative or absolute in relation to a position of a speaker, which are stored as metadata for respective text data and / or terms in a database (DB) for DialogResources (DRS) located on the vehicle side and / or in the central computer unit (3). [6] Method according to claim 5, characterized by, that by means of the Prompter functional unit (PTR) a duration of the audio file (AD) is determined and an audio animation and / or a data stream of individual vectors (V) is generated, or is generated by interpolating further vectors (V) between the start vector and the end vector for the audio animation corresponding to a runtime of the audio file (AD). [7] Method according to claim 6, characterized by , that an audio data stream and / or an audio animation data stream is transmitted by means of the prompter functional unit (PTR) to an audio amplifier (7) of an audio system (8) of the vehicle (1) and that the audio data stream and / or the audio animation data stream is converted into an audio data stream by means of the audio amplifier (7) and that the respective audio data stream is output in the vehicle (1) by means of vehicle-side loudspeakers (6). [8] Method according to any one of the preceding claims, characterized by, that a machine learning-pretrained language model is used to generate an audio data stream and / or an audio animation data stream.
Citation Information
Patent Citations
Voice control system and motor vehicle
DE102020003922A1
Agent device, agent presenting method, and storage medium
US20200111489A1