Voice assistance system and method
Through the variational automatic encoder, the speech recognition that does not rely on vocabulary tokenization is realized in the vehicle, which solves the problem of difficult expression recognition of homophones and natural languages in the prior art, and improves the accuracy and efficiency of speech recognition.
Patent Information
- Application Number
- CN202311813935.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
Existing speech recognition technologies rely on tokenization, making it difficult to deal with traditional expressions in homophones and natural languages, resulting in low recognition accuracy and inefficiency.
The encoder and decoder using a variational autoencoder encoder encodes the audio data into the latent space and generates expressions in combination with context data to achieve natural language recognition that does not rely on vocabulary tokenization.
It improves the accuracy and efficiency of speech recognition, can handle complex expressions in natural language, reduces background noise interference, and enhances the performance of voice assistance applications in vehicles.
Smart Images

Figure CN120220667A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to voice assistance systems and methods. More specifically, the present disclosure relates to voice assistance systems and methods in vehicles. Background Art
[0002] This introduction generally presents the background of the present disclosure. The work of the currently named inventors, to the extent it is described in this introduction, and aspects that may not conform to the description of the prior art at the time of filing, are neither expressly nor impliedly admitted to be prior art against the present disclosure.
[0003] Tokenization-based sentence generation and context understanding for speech recognition have quite a number of constraints and limitations. For example, homophone tokens may introduce ambiguity during the speech recognition process. Therefore, there is a need to develop voice assistance methods and systems that do not rely on tokenization. Summary of the Invention
[0004] The present disclosure describes a voice assistance method. The method includes receiving, by a vehicle controller of a vehicle, audio data. The audio data represents a voice command issued by a vehicle user in natural language. The method further includes encoding, by an encoder of a variational autoencoder, the audio data into a latent space to generate encoded data. The method further includes receiving context data related to the voice command issued by the vehicle user. The method further includes generating, by a decoder of the variational autoencoder, an expression from the encoded data and the context data. The expression represents the audio data. The method further includes commanding, by the vehicle controller, the vehicle to generate a response based on the expression generated by the decoder of the variational autoencoder. The method described in this paragraph improves speech recognition technology and vehicle technology by recognizing natural language uttered by a user that does not rely on lexical tokenization, while lexical tokenization is prone to errors when the user utters conventional expressions specific to a language (e.g., idioms, poems, slang, etc.). Since the method described in this paragraph does not use lexical tokenization, the method improves natural language recognition through voice assistance applications, thereby improving speech recognition technology and voice assistance applications in vehicles.
[0005] In certain aspects of the present disclosure, the method does not include performing lexical tokenization of the audio data. The method may include reducing background noise in the audio data and identifying speech in the audio data. The encoder of the variational autoencoder is a first neural network that maps the audio data to a latent space. The audio data is located in an input space. The decoder is a second neural network that maps the encoded data to the input space. Context data may be used as input. The context data includes user speech data and external factor data. The user speech data includes information about the user's speech intonation when the user issued the speech command. The external factor data includes the traffic conditions around the vehicle when the user issued the speech command, the date when the user issued the speech command, and the time when the user issued the speech command. The context data includes dialogue history data. The dialogue history data includes information about the dialogue history of the user who issued the speech command. The context data is used as input to the second neural network. The response is generated based on multiple constraints. The multiple constraints include response time and sentence length. The method may include controlling an actuator of the vehicle based on the response.
[0006] The present disclosure also describes a voice assistance system. The voice assistance system includes a user interface that includes a microphone. The microphone is configured to capture speech commands issued by a vehicle user. The voice assistance system also includes a plurality of sensors. Each of the plurality of sensors is configured to collect context data. The voice assistance system also includes a vehicle controller that communicates with the user interface and the plurality of sensors. The vehicle controller is programmed to perform the above method.
[0007] The present disclosure also describes a tangible, non-transitory machine-readable medium that includes machine-readable instructions that, when executed by a processor, cause the processor to perform the above method.
[0008] Based on the following detailed description, other application areas of the present disclosure will become apparent. It should be understood that these descriptions and specific examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure.
[0009] When combined with the drawings, the above features and advantages of the currently disclosed systems and methods, as well as other features and advantages, are apparent from the detailed description that includes the claims and exemplary embodiments. Description of the Drawings
[0010] The present disclosure will be more fully understood in conjunction with the detailed description and the drawings, wherein:
[0011] Figure 1 is a schematic diagram of a vehicle including a voice assistance system.
[0012] Figure 2 is a flowchart of a voice assistance method. Detailed Description
[0013] Reference will now be made in detail to several examples of the present disclosure shown in the drawings. Whenever possible, the same or similar reference numerals are used in the drawings and the description to refer to the same or similar components or steps.
[0014] Referring Figure 1 , vehicle 10 includes a voice assistance system 39. Vehicle 10 generally includes a body 12 and a plurality of wheels 14 coupled to the body 12. Vehicle 10 may be an autonomous vehicle. In the illustrated embodiment, vehicle 10 is depicted as a sedan in the illustrated embodiment, but it should be understood that other vehicles, such as trucks, coupes, sport utility vehicles (SUVs), boats, airplanes, recreational vehicles (RVs), etc. may also be used.
[0015] Vehicle 10 also includes one or more sensors 24 connected to the body 12. Sensors 24 sense observable conditions of the external environment and / or the internal environment of vehicle 10. As a non-limiting example, sensors 24 may include one or more cameras, one or more light detection and ranging (LIDAR) sensors, one or more proximity sensors, one or more ultrasonic sensors, one or more thermal imaging sensors, a global positioning system (GPS) transceiver, and / or other sensors. Each sensor 24 is configured to generate a signal indicative of the sensed observable conditions (i.e., sensor data) of the external environment and / or the internal environment of vehicle 10. The signal represents the sensor data collected by sensors 24.
[0016] Vehicle 10 includes a vehicle controller 34 that communicates with sensors 24. Vehicle controller 34 includes at least one vehicle processor 44 and a vehicle non-transitory computer-readable storage device or medium 46. Processor 44 may be a custom or commercially available processor, a central processing unit (CPU), a graphics processing unit (GPU), an auxiliary processor among several processors associated with vehicle controller 34, a semiconductor-based microprocessor (in the form of a microchip or chipset), a macroprocessor, a combination thereof, or a device generally used to execute instructions. Vehicle-readable storage device or medium 46 may include, for example, volatile and non-volatile storage in read-only memory (ROM), random access memory (RAM), and keep-alive memory (KAM). KAM is persistent or non-volatile memory that can be used to store various operating variables when vehicle processor 44 is powered down. Vehicle-readable storage device or medium 46 may be implemented using a variety of storage devices, such as PROM (programmable read-only memory), EPROM (electric PROM), EEPROM (electrically erasable PROM), flash memory, or any other electrical, magnetic, optical, or combination storage device capable of storing data used by vehicle controller 34 in controlling vehicle 10, some of which data represents executable instructions. Vehicle controller 34 is specifically programmed to execute method 100 ( Figure 2 ).
[0017] These instructions can include one or more separate programs, each program including an ordered list of executable instructions for implementing a logical function. When executed by the vehicle processor 44, the instructions receive and process signals from sensors, execute logic, calculations, methods, and / or algorithms for automatically controlling components of the vehicle 10, and generate control signals to automatically control components of the vehicle 10 based on the logic, calculations, methods, and / or algorithms. Although Figure 1 a single vehicle controller 34 is shown in
[0018] Embodiments of the vehicle 10 may include multiple vehicle controllers 34 that communicate via a suitable communication medium or combination of communication media and cooperate to process sensor signals, execute logic, calculations, methods, and / or algorithms, and generate control signals to automatically control features of the vehicle 10. The vehicle controller 34 is part of the voice assistance system 39.
[0019] The vehicle 10 also includes one or more actuators 26 that communicate with the vehicle controller 34. The actuators 26 control one or more vehicle features, such as but not limited to the propulsion system, transmission system, steering system, radio, air conditioning system, and braking system of the vehicle 10. In various embodiments, the vehicle features may also include interior and / or exterior vehicle features, such as but not limited to doors, trunks, and cab features, such as air, music, lighting, etc.
[0020] As discussed below, the voice assistance system 39 uses a generative model (e.g., variational autoencoder). As described above, the encoder of the variational autoencoder encodes the audio data collected by the microphone 50 into a latent space to search for and establish the most relevant information to understand the natural language uttered by the user of the vehicle 10. The variational autoencoder also uses context data and external factor data to maximize the accuracy of the natural language conversion from the user's utterance to a digital audio signal. In this way, the voice assistance system 39 can also understand traditional expressions of a specific language (e.g., poetry, idioms, slang, or other spoken language) by considering external and context factors. Since the voice assistance system 39 considers external and context factors, there is no need to translate the source language into standard English and then back into the source language. Therefore, the voice assistance system 39 accurately converts the user's utterance into a digital audio form. In addition, the voice assistance system 39 does not rely on lexical tokenization. Tokenization-based sentence generation and context understanding for speech recognition have quite a number of constraints and limitations. For example, homophone tokens may introduce ambiguity during the speech recognition process. Therefore, the voice assistance system 39 accurately converts the user's utterance into a digital audio form.
[0021] Figure 2 is a flowchart of a voice assistance method 100. Method 100 does not use lexical tokenization and starts at block 102. At block 102, the vehicle controller 34 receives the audio data collected by the microphone 50. The audio data represents a voice command uttered by the user of the vehicle 10 in natural language. In addition, at block 102, the vehicle controller 34 reduces background noise from the audio data and identifies the user's voice in the audio data. As a non-limiting example, a two-sided companding noise reduction system can be used to reduce background noise from the audio data. A suitable automatic speech recognition (ASR) system can be used to identify the user's voice in the audio data collected by the microphone 50. In method 100, the voice assistance system 39 continuously listens. Therefore, noise reduction is iteratively performed in each cycle of voice assistance (e.g., every 1 millisecond). Then, method 100 continues to block 104.
[0022] At block 104, the vehicle controller 34 encodes the audio data collected by the microphone 50 into the latent space using the encoder of the variational autoencoder. Thus, the audio data serves as the input to the encoder. Encoding the audio data generates encoded data, which is compressed data. The audio data lies in the input space. The latent space is a low-dimensional space relative to the input space to minimize the speech conversion burden from the user's speech to the digital waveform. The encoder of the variational autoencoder is the first neural network that maps the audio to the latent space. In other words, the encoder of the variational autoencoder encodes the speech wave signal into latent variables and vectors. The encoder of the variational autoencoder has been trained beforehand to be applicable to various situations. For example, the encoder is trained to recognize homophones and their meanings. One sound may have multiple meanings. In method 100, the voice assistant system 39 introduces a third language for cross-reference to obtain a better meaning mapping. The encoder of the variational autoencoder is adaptable, so it can use the latent variables and vectors to track the speech of different dialects. In addition, the encoder of the variational autoencoder identifies language-specific traditional expressions (e.g., idioms, poems, slang, etc.) by using the context data. The context data may include the conversation history data about the user. The conversation history data includes information about the conversation history of the user who issued the voice command. Thus, in addition to the audio data, the context data can be the input to the first neural network that forms the encoder of the variational autoencoder. In other words, the audio data and the context data can be the input to the first neural network that forms the encoder of the variational autoencoder. The voice assistant system 39 also recognizes and understands sounds that have meanings even though there are no words. Then, method 100 proceeds to block 106.
[0023] At block 106, the vehicle controller 34 understands the context based on the encoded data. Specifically, the vehicle controller 34 receives context data related to the voice command issued by the vehicle user from, for example, the sensor 24. The context data is in the latent space and does not have an explicit language-based meaning. The context data includes the user speech data. Further, the user speech data includes information about the speech tone and mood of the user when the voice command is issued. The context data may also include external factor data. The external factor data includes the traffic conditions around the vehicle 10 when the user issues the voice command, the date when the user issues the voice command, and the time when the user issues the voice command, etc. The context data may also include the conversation history data. The conversation history data includes information about the conversation history of the user who issued the voice command to help understand the language-based traditional expressions. Then method 100 proceeds to block 108.
[0024] At block 108, the vehicle controller 34 uses a pre-trained decoder of the variational autoencoder to generate an expression (in speech wave format) from the encoded data and the context data. The expression represents the audio data collected by the microphone 50. The decoder is a second neural network that maps the encoded data to the input space. The decoder can be trained using the audio data and the context data from the vehicle 10. The decoder creates a poll of all contexts and generates all possible responses in parallel in real time. By considering the speech intonation, context, and preferences, some of the less likely responses can be filtered out, thereby improving the accuracy of the decoder. Then, method 100 proceeds to block 110.
[0025] At block 110, the vehicle controller 34 commands the vehicle to generate a response based on the expression generated by the decoder of the variational autoencoder. The response has some limitations (e.g., response time, sentence length, speech intonation, etc.) in order to generate a response that is as natural as possible according to the user's preferences and characteristics. Even when no speech information is captured, the voice assistant system 39 can actively generate a response. The vehicle controller 34 can command the speaker 49 to emit the response in speech. Additionally, the vehicle controller 34 can control one or more actuators 26 to automatically control one or more vehicle operations (e.g., air conditioning system, text message, music, shopping, anxiety relief, switching to autonomous driving mode, etc.) after obtaining user confirmation. The response can also explain the decisions made when the vehicle 10 is in autonomous driving mode.
[0026] The accompanying drawings are in simplified form and are not drawn to exact scale. Directional terms, such as top, bottom, left, right, up, above, over, below, beneath, rear, and front, may be used with respect to the drawings for convenience and clarity only. Such similar directional terms should not be construed to limit the scope of the present disclosure in any way.
[0027] Embodiments of the present disclosure are described herein. However, it should be understood that the disclosed embodiments are merely examples and that other embodiments may take various alternative forms. The accompanying drawings are not necessarily to scale; certain features may be enlarged or minimized to show details of particular components. Thus, the specific structural and functional details disclosed herein should not be construed as limiting, but merely as a representative basis for teaching one skilled in the art to employ the systems and methods of the present disclosure in different ways. As will be understood by one of ordinary skill in the art, the various features shown and described with reference to any one of the drawings may be combined with features shown in one or more other drawings to produce embodiments that are not explicitly shown or described. Combinations of the shown features provide representative embodiments for typical applications. However, various combinations and modifications of the features may be required for a particular application or implementation consistent with the teachings of the present disclosure.
[0028] Embodiments of the present disclosure may be described in terms of functional and / or logical block components and various processing steps. It should be understood that such block components may be implemented by a plurality of hardware, software, and / or firmware components configured to perform the specified functions. For example, embodiments of the present disclosure may employ various integrated circuit components, such as memory elements, digital signal processing elements, logic elements, look-up tables, etc., which may perform various functions under the control of one or more microprocessors or other control devices. Additionally, those skilled in the art will appreciate that embodiments of the present disclosure may be practiced in conjunction with a variety of systems, and the systems described herein are merely exemplary embodiments of the present disclosure.
[0029] For simplicity, techniques related to the signal processing, data fusion, signaling, control, and other functional aspects of the system (as well as the various operating components of the system) may not be described in detail herein. Additionally, the connecting lines shown in the various figures included herein are intended to represent example functional relationships and / or physical couplings between the various elements. It should be noted that alternative or additional functional relationships or physical connections may exist in embodiments of the present disclosure.
[0030] The foregoing description is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses. The broad teachings of the present disclosure may be implemented in a variety of forms. Thus, while the present disclosure includes specific examples, the true scope of the present disclosure should not be so limited because other modifications will become apparent upon study of the drawings, the specification, and the appended claims.
Claims
1. A voice assistance method, comprising: Receiving, by a vehicle controller of a vehicle, audio data, wherein the audio data represents a voice command uttered by a vehicle user in natural language; Encoding the audio data into a latent space using an encoder of a variational autoencoder to generate encoded data; Receiving context data related to the voice command uttered by the vehicle user; Generating, using a decoder of the variational autoencoder, an expression from the encoded data and the context data, wherein the expression represents the audio data; and Commanding, using the vehicle controller, the vehicle to generate a response based on the expression generated by the decoder of the variational autoencoder.
2. The method according to claim 1, wherein the method does not include performing lexical tokenization of the audio data.
3. The method according to claim 2, further comprising: Reducing background noise in the audio data; And Identifying speech in the audio data.
4. The method according to claim 3, wherein the encoder of the variational autoencoder is a first neural network that maps the audio data to the latent space, the audio data is located in an input space, and the decoder is a second neural network that maps the encoded data to the input space, the context data includes user speech data, and the user speech data includes information about the speech intonation of the user when the user uttered the voice command.
5. The method according to claim 4, wherein the context data includes external factor data, wherein, The external factor data includes the traffic conditions around the vehicle when the user uttered the voice command, the date when the user uttered the voice command, and the time when the user uttered the voice command.
6. The method according to claim 5, wherein the context data includes conversation history data, wherein, The conversation history data includes information about the conversation history of the user who uttered the voice command.
7. The method according to claim 6, wherein the context data is used as an input to the second neural network.
8. The method according to claim 7, wherein the response is generated based on multiple constraints.
9. The method according to claim 8, wherein the multiple constraints include response time and sentence length.
10. The method according to claim 9, wherein the response includes controlling an actuator of the vehicle based on the response.