Vehicle-mounted devices and methods for processing speech

By combining audio and image data processing, using acoustic models with neural network architectures and speaker feature vector techniques, the accuracy and robustness issues of in-vehicle speech processing are solved, enabling more efficient speech transcription and control command parsing.

CN112420033BActive Publication Date: 2025-10-31SOUNDHOUND INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010841021.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-23
Filing Date
2020-08-20
Publication Date
2025-10-31
Estimated Expiration
2041-01-05

AI Technical Summary

Technical Problem

Providing an efficient and accurate voice processing system within a vehicle presents challenges, including limited processing resources, noise interference, and complex acoustic environments, making it difficult for existing systems to achieve human-level responsiveness and intelligence.

Method used

By combining audio and image data processing, speech processing is improved by obtaining speaker feature vectors. An acoustic model with a neural network architecture is used for parsing, and lip reading and facial recognition technologies are combined to generate speaker feature vectors to improve parsing accuracy.

Benefits of technology

It improves the accuracy and robustness of speech processing within vehicles, overcomes the limitations of noise and limited resources, and enables more efficient speech transcription and control command parsing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112420033B_ABST
    Figure CN112420033B_ABST
Patent Text Reader

Abstract

This application relates to in-vehicle devices and methods for processing speech. Systems and methods for processing speech are described. Specific examples use visual information to improve speech processing. This visual information may be image data obtained from inside the vehicle. In the example, the image data describes the characteristics of a person inside the vehicle. Specific examples use the image data to obtain a speaker feature vector for use by an adapted speech processing module. The speech processing module can be configured to use the speaker feature vector to process audio data describing the characteristics of the speech. The audio data may be audio data derived from an audio capture device inside the vehicle. Specific examples use a neural network architecture to provide an acoustic model for processing the audio data and the speaker feature vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technology relates to the field of speech processing. A specific example involves processing speech captured from inside a vehicle. Background Technology

[0002] Recent advances in computing have increased the possibility of realizing many long-sought-after voice control applications. For example, improvements in statistical models, including practical frameworks for efficient neural network architectures, have significantly improved the accuracy and reliability of previous voice processing systems. This is combined with the rise of wide-area computer networks, which provide a range of modular services that can be easily accessed using application programming interfaces (APIs). Voice is rapidly becoming a viable option for providing user interfaces.

[0003] While voice control devices have become popular in homes, providing voice processing in vehicles presents additional challenges. For example, vehicles typically have limited processing resources for accessibility features (e.g., voice interfaces), suffer from significant noise (e.g., high levels of road and / or engine noise), and are constrained by acoustic environments. Furthermore, any user interface is limited by the safety concerns associated with controlling the vehicle. These factors make in-vehicle voice control difficult to implement in practice.

[0004] Moreover, despite advancements in speech processing, even users of advanced computing devices frequently report a lack of human-level responsiveness and intelligence in current systems. Translating fluctuations in air pressure into parsed commands is incredibly difficult. Speech processing typically involves complex processing pipelines, where errors at any stage can derail successful machine interpretation. Many of these challenges are not readily apparent to humans, who are able to process speech unconsciously using their cortical and subcortical structures. However, engineers in the field have quickly recognized the gap between human capabilities and state-of-the-art speech processing.

[0005] US 8,442,820 B2 (Patent Document 1) describes a combined lip-reading and speech recognition multi-modal interface system. This system can issue navigation operation commands solely through speech and lip movements, allowing the driver to look forward during navigation operations and reducing vehicle accidents related to navigation operations while driving. The combined lip-reading and speech recognition multi-modal interface system described in US 8,442,820 B2 includes: an audio speech input unit; a speech recognition unit; a speech recognition command and estimated probability output unit; a lip video image input unit; a lip-reading unit; a lip-reading recognition command output unit; and a speech recognition and lip-reading recognition result combination unit that outputs the speech recognition command. While US 8,442,820 B2 provides a solution for in-vehicle control, the proposed system is complex, and the numerous interoperable components offer more opportunities for errors and parsing failures.

[0006] There is a desire to provide speech processing systems and methods that can more accurately transcribe and parse human utterance. Furthermore, there is a desire to provide speech processing methods that can be practically implemented in real-world devices, such as embedded computing systems for vehicles. Implementing practical speech processing solutions is very difficult due to the numerous challenges that vehicles present in terms of system integration and connectivity.

[0007] [Patent Document 1] US Patent Specification Serial No. 8,422,820 Summary of the Invention

[0008] The specific examples described in this paper provide methods and systems for processing speech. These specific examples use both audio and image data to process speech. They are applicable to addressing the challenge of processing utterance captured inside a vehicle. The specific examples obtain a speaker feature vector based on image data that describes features of at least facial regions of a person (e.g., a person inside a vehicle). Visually derived information (which depends on the speaker of the utterance) is then used to perform speech processing. This can improve accuracy and robustness.

[0009] In one aspect, an apparatus for a vehicle includes: an audio interface configured to receive audio data from an audio capture device located within the vehicle; an image interface configured to receive image data from an image capture device for capturing image data within the vehicle, the image data describing features of a facial region of a person within the vehicle; and a speech processing module configured to parse human speech based on the audio data and the image data. The speech processing module includes an acoustic model configured to process the audio data and predict phoneme data for parsing the speech, wherein the acoustic model includes a neural network architecture. The apparatus also includes a speaker preprocessing module implemented by a processor, configured to receive image data and obtain a speaker feature vector based on the image data, wherein the acoustic model is configured to receive the speaker feature vector and the audio data as input and is trained to predict the phoneme data using the speaker feature vector and the audio data.

[0010] In the above aspects, the speaker feature vector is obtained using image data describing features of the speaker's facial regions. This speaker feature vector is fed as input to a neural network architecture of an acoustic model, which is configured to use this input along with audio data describing the features of the utterance. In this way, the acoustic model is provided with additional visually derived information, which the neural network architecture can use to improve utterance resolution, such as compensating for adverse acoustic and noise characteristics within a vehicle. For example, configuring the acoustic model based on a specific person and / or that person's mouth regions (as determined from image data) can improve the identification of ambiguous phonemes (e.g., phonemes that might be incorrectly transcribed based on vehicle conditions without additional information).

[0011] In one variant, the speaker preprocessing module is configured to perform facial recognition on image data to identify people inside the vehicle and obtain a speaker feature vector associated with the identified person. For example, the speaker preprocessing module may include a facial recognition module for identifying users speaking inside the vehicle. In cases where the speaker feature vector is determined based on audio data, the identification of a person can allow the retrieval of a predetermined (e.g., pre-computed) speaker feature vector from memory. This can improve the processing latency of constrained embedded vehicle control systems.

[0012] In one variant, the speaker preprocessing module includes a lip-reading module, implemented by a processor, configured to generate one or more speaker feature vectors based on lip movements within a person's facial region. This can be used in conjunction with a face recognition module or independently of it. In this case, one or more speaker feature vectors provide a representation of the speaker's mouth or lip region, which can be used by the neural network architecture of the acoustic model to improve processing.

[0013] In some cases, the speaker preprocessing module may include a neural network architecture configured to receive one or more derived data from audio and image data and predict speaker feature vectors. For example, this approach can combine a vision-based neural lip-reading system with an acoustic “x-vector” system to improve acoustic processing. When using one or more neural network architectures, these architectures can be trained using a training set that includes image data, audio data, and a ground truth set of linguistic features (e.g., a ground truth set of phoneme data and / or text transcription).

[0014] In some cases, the speaker preprocessing module is configured to compute multiple speaker feature vectors for a predefined number of utterances, and to compute static speaker feature vectors based on these multiple speaker feature vectors for the predefined number of utterances. For example, the static speaker feature vectors may include the average of a set of speaker feature vectors linked to a specific user using image data. The static speaker feature vectors can be stored in the vehicle's memory. This again improves speech processing capabilities within resource-constrained vehicle computing systems.

[0015] In one embodiment, the apparatus includes a memory configured to store one or more user profiles. In this case, the speaker preprocessing module can be configured to perform facial recognition on image data to identify user profiles associated with a person inside the vehicle, calculate a speaker feature vector for that person, store the speaker feature vector in the memory, and associate the stored speaker feature vector with the identified user profile. Facial recognition can provide a fast and convenient mechanism for obtaining useful information (e.g., speaker feature vectors) for acoustic processing associated with a specific person. In one embodiment, the speaker preprocessing module can be configured to determine whether the number of stored speaker feature vectors associated with a given user profile is greater than a predefined threshold. If so, the speaker preprocessing module can calculate a static speaker feature vector based on that number of stored speaker feature vectors; store the static speaker feature vector in the memory; associate the stored static speaker feature vector with a given user profile; and signal that the static speaker feature vector will be used for future speech parsing instead of calculating a speaker feature vector for that person.

[0016] In one variant, the apparatus includes an image capture device configured to capture electromagnetic radiation with an infrared wavelength, the image capture device being configured to send image data to an image interface. This can provide an illumination-invariant image for improved image data processing. A speaker preprocessing module can be configured to process the image data to extract one or more portions of the image data, wherein the extracted one or more portions are used to obtain a speaker feature vector. For example, the one or more portions may relate to facial regions and / or mouth regions.

[0017] In one scenario, one or more of the audio interface, image interface, speech processing module, and speaker preprocessing module may be located within the vehicle, for example, as part of a local embedded system. In this case, the processor may be located within the vehicle. In another scenario, the speech processing module may be remote relative to the vehicle. In this case, the device may include a transceiver for sending data derived from audio and image data to the speech processing module and receiving control data derived from the parsing of dialogue. Different distributed configurations are possible. For example, in one scenario, the device may be implemented locally within the vehicle, but another copy of at least one component of the device may be implemented on a remote server device. In this case, certain functions may be performed remotely, for example, in addition to or in lieu of local processing. The remote server device may have enhanced processing resources, which may improve accuracy but may increase processing latency.

[0018] In one case, the acoustic model includes a hybrid acoustic model comprising a neural network architecture and a Gaussian mixture model, wherein the Gaussian mixture model is configured to receive a vector of class probabilities output by the neural network architecture and output phoneme data for utterance parsing. The acoustic model may additionally or alternatively include, for example, a Hidden Markov Model (HMM) and a neural network architecture. In one case, the acoustic model may include a connectionist temporal classification (CTC) model, or another form of neural network model with a recurrent neural network architecture.

[0019] In one variant, the speech processing module includes a language model communicatively coupled to an acoustic model for receiving phoneme data and generating transcriptions representing utterances. In this variant, for example, in addition to the acoustic model, the language model can be configured to generate transcriptions representing utterances using speaker feature vectors. Where the language model includes a neural network architecture (e.g., a recurrent neural network or a transformer architecture), this can be used to improve the accuracy of the language model.

[0020] In one variant, the acoustic model includes: a database of acoustic model configurations; an acoustic model selector for selecting an acoustic model configuration from the database based on a speaker feature vector; and an acoustic model instance for processing audio data, the acoustic model instance being instantiated based on the acoustic model configuration selected by the acoustic model selector, the acoustic model instance being configured to generate phoneme data for parsing utterances.

[0021] In a specific example, the speaker feature vector is one or more of an i-vector and an x-vector. The speaker feature vector may include a synthetic vector, for example, comprising two or more of the following: a first part related to the speaker generated based on audio data; a second part related to the speaker's lip movements generated based on image data; and a third part related to the speaker's face generated based on image data.

[0022] According to another approach, there is a method for processing speech, comprising: receiving audio data from an audio capture device located within a vehicle, the audio data describing features of speech spoken by a person within the vehicle; receiving image data from an image capture device located within the vehicle, the image data describing features of a person's facial regions; obtaining a speaker feature vector based on the image data; and parsing the speech using a speech processing module implemented by a processor. Parsing the speech includes providing the speaker feature vector and audio data as input to an acoustic model of the speech processing module, the acoustic model including a neural network architecture; and at least using the neural network architecture to predict phoneme data based on the speaker feature vector and audio data.

[0023] This method can provide similar improvements to in-vehicle speech processing. In some cases, obtaining speaker feature vectors includes: performing facial recognition on image data to identify a person inside the vehicle; obtaining user profile data for that person based on the facial recognition; and obtaining speaker feature vectors based on the user profile data. The method may further include comparing the number of stored speaker feature vectors associated with the user profile data with a predefined threshold. In response to the number of stored speaker feature vectors being less than the predetermined threshold, the method may include using one or more of audio data and image data to compute speaker feature vectors. In response to the number of stored speaker feature vectors being greater than the predetermined threshold, the method may include obtaining static speaker feature vectors associated with the user profile data, which are generated using the number of stored speaker feature vectors. In one case, speaker feature vectors include: processing image data to generate one or more speaker feature vectors based on lip movements within a human facial region. Parsing utterance may include: providing phoneme data to a language model of the speech processing module; using the language model to predict the transcription of the utterance; and using the transcription to determine control commands for the vehicle.

[0024] According to another aspect, there exists a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause at least one processor to perform the following operations: receiving audio data from an audio capture device; receiving a speaker feature vector obtained based on image data from an image capture device, the image data describing features of a user's facial region; and parsing speech using a speech processing module, including: providing the speaker feature vector and audio data as input to an acoustic model of the speech processing module, the acoustic model including a neural network architecture, at least using the neural network architecture, predicting phoneme data based on the speaker feature vector and audio data, providing the phoneme data to a language model of the speech processing module, and using the language model to generate a transcription of the speech.

[0025] At least one processor may include a computing device, for example, a computing device remote relative to a motor vehicle, wherein audio data and speaker feature vectors are received from the motor vehicle. Instructions may cause the processor to perform automatic speech recognition with a low error rate. In some cases, the speaker feature vector includes one or more of the following: speaker-related vector elements generated based on audio data; vector elements related to the speaker's lip movements generated based on image data; and vector elements related to the speaker's face generated based on image data. Attached Figure Description

[0026] Figure 1A This is a schematic diagram showing the interior of a vehicle based on an example.

[0027] Figure 1B This is a schematic diagram illustrating a device for a vehicle according to an example.

[0028] Figure 2 This is a schematic diagram illustrating a device for a vehicle with a speaker preprocessing module, according to an example.

[0029] Figure 3 This is a schematic diagram illustrating the components of the speaker preprocessing module according to the example.

[0030] Figure 4 This is a schematic diagram illustrating the components of a speech processing module according to an example.

[0031] Figure 5 This is a schematic diagram illustrating the neural speaker preprocessing module and the neural speech processing module according to the example.

[0032] Figure 6 This is a schematic diagram illustrating the components of an acoustic model for configuring a speech processing module, based on an example.

[0033] Figure 7 This is a schematic diagram illustrating an image preprocessor based on an example.

[0034] Figure 8 This is a schematic diagram illustrating image data from different image capture devices, based on an example.

[0035] Figure 9 This is a schematic diagram illustrating the components of a speaker preprocessing module configured to extract lip features, based on an example.

[0036] Figure 10A and Figure 10B This is a schematic diagram illustrating a motor vehicle with a device for voice processing, according to an example.

[0037] Figure 11 This is a schematic diagram illustrating components of a user interface for a motor vehicle, based on an example.

[0038] Figure 12 This is a schematic diagram illustrating an example computing device used in a vehicle.

[0039] Figure 13 This is a flowchart illustrating the method of processing discourse based on the example.

[0040] Figure 14 This is a schematic diagram illustrating a non-transitory computer-readable storage medium according to an example. Detailed Implementation

[0041] The following describes various examples of this technique that illustrate various aspects of interest. Typically, the described aspects can be used in any combination.

[0042] The specific examples described herein use visual information to improve speech processing. This visual information can be obtained from inside the vehicle. In the examples, the visual information describes the characteristics of people inside the vehicle (e.g., a driver or passenger). The specific examples use the visual information to generate speaker feature vectors for use by an adapted speech processing module. The speech processing module can be configured to use the speaker feature vectors to improve the processing of associated audio data (e.g., audio data derived from an audio capture device inside the vehicle). These examples can improve the responsiveness and accuracy of in-vehicle voice interfaces. Computing devices can use the specific examples to improve speech transcription. Thus, the described examples can be viewed as extending a speech processing system to have multimodal capabilities that can improve the accuracy and reliability of audio processing.

[0043] The specific examples described in this paper provide different approaches for generating speaker feature vectors. Some of these approaches are complementary and can be used together to synergistically improve speech processing. In one example, image data obtained from inside the vehicle (e.g., from driver and / or passenger cameras) is processed to identify a person and determine feature vectors that numerically represent certain characteristics of that person. These characteristics may include audio properties, such as a numerical representation of the expected variance within the audio data for an acoustic model. In another example, image data obtained from inside the vehicle (e.g., from driver and / or passenger cameras) is processed to determine feature vectors that numerically represent certain visual characteristics of the person (e.g., characteristics associated with the person's speech). In one case, visual characteristics may be associated with the person's mouth region, such as representing the position and / or movement of the lips. In both examples, the speaker feature vectors may have similar formats and are therefore easily integrated into the input pipeline of the acoustic model used to generate phoneme data. Specific examples can provide improvements over some of the challenges of in-vehicle automatic speech recognition, such as the limited interior space of the vehicle, the possibility that multiple people may be speaking within that limited interior space, and high levels of engine and ambient noise.

[0044] Example vehicle context

[0045] Figure 1A An example scenario of a speech processing device is shown. Figure 1A In this context, the vehicle is a motor vehicle. Figure 1A This is a schematic diagram of the interior 100 of a motor vehicle. The interior 100 shown is directed towards the driver's side of the motor vehicle. A person 102 is shown inside the interior 100. Figure 1AIn this design, the person is the driver of the motor vehicle. The driver faces forward and observes the road through the windshield 104. The person uses the steering wheel 106 to control the vehicle and observes vehicle status indicators through the instrument panel or dashboard 108. Figure 1A In this embodiment, the image capture device 110 is located inside the interior 100 of the motor vehicle near the bottom of the dashboard 108. The image capture device 110 has a field of view 112 that captures the facial region 114 of a person 102. In this example, the image capture device 110 is positioned to capture an image through an opening in the steering wheel 106. Figure 1A An audio capture device 116 located within the interior 100 of a motor vehicle is also shown. The audio capture device 116 is arranged to capture the sound emitted by a person 102. For example, the audio capture device 116 may be arranged to capture speech from the person 102, i.e., sound emitted from the person's facial area 114. The audio capture device 116 is shown mounted to the windshield 104; for example, the audio capture device 116 may be mounted near or on a rearview mirror, or mounted on a door frame to the side of the person 102. Figure 1A A voice processing device 120 is also shown. The voice processing device 120 can be installed in a motor vehicle. The voice processing device 120 may include, or form part of, a control system for a motor vehicle. Figure 1A In the example, image capture device 110 and audio capture device 116 are communicatively coupled to voice processing device 120, for example, via one or more wired and / or wireless interfaces. Image capture device 110 may be located outside the vehicle to capture images inside the vehicle through the vehicle's windows.

[0046] Figure 1A The scenarios and configurations are provided as examples to aid in understanding the following description. It should be noted that these examples are not limited to motor vehicles, but can be similarly implemented with other forms of vehicles, including, but not limited to: maritime transport such as boats and ships; air transport such as helicopters, airplanes, and gliders; rail transport such as trains and trams; spacecraft, engineering vehicles, and heavy equipment. Motor vehicles can include automobiles, trucks, sports utility vehicles, motorcycles, buses, and motorized cars, etc. The use of the term "vehicle" in this document also includes certain heavy equipment that can be motorized while remaining stationary, such as cranes, lifting equipment, and drilling equipment. Vehicles can be manually controlled and / or have autonomous functions. Although Figure 1A The example describes the features of steering wheel 106 and dashboard 108, but other control arrangements may be provided (e.g., an autonomous vehicle may not have the depicted steering wheel 106). While in Figure 1AThe driver's seat scenario is shown, but a similar configuration can be provided for one or more passenger seats (e.g., front and rear seats). Figure 1A Provided for illustrative purposes only, and for clarity, certain features that may also be found within a motor vehicle have been omitted. In some cases, the methods described herein can be used outside of a vehicle context, for example, by computing devices such as desktop or laptop computers, smartphones, or embedded devices.

[0047] Figure 1B yes Figure 1A A schematic diagram of the speech processing device 120 shown. Figure 1B In the speech processing device 120, there is a speech processing module 130, an image interface 140, and an audio interface 150. The image interface 140 is configured to receive image data 145. The image data may include data generated by... Figure 1A Image data captured by image capture device 110. Audio interface 150 is configured to receive audio data 155. Audio data 155 may include data captured by image capture device 110. Figure 1A Audio data is captured by audio capture device 116. Speech processing module 130 is communicatively coupled to both image interface 140 and audio interface 150. Speech processing module 130 is configured to process image data 145 and audio data 155 to generate a set of linguistic features 160 that can be used to parse the speech of person 102. Linguistic features may include phonemes, word parts (e.g., stems or proto-words) and words (including text features such as pauses mapped to punctuation marks), as well as probabilities and other values ​​associated with these linguistic units. In one case, the linguistic features may be used to generate text output representing the speech. In this case, the text output may be used as is, or it may be mapped to a predefined set of commands and / or command data. In another case, the linguistic features may be directly mapped to a predefined set of commands and / or command data (e.g., no explicit text output).

[0048] A person (e.g., person 102) can use Figure 1A and Figure 1BThe configuration allows for issuing voice commands while operating a motor vehicle. For example, person 102 can speak inside the vehicle (e.g., generate utterances) to control the motor vehicle or obtain information. The utterances in this context are associated with spoken sounds generated by a person that represent linguistic information such as speech. For example, utterances can include speech produced from the throat of person 102. The utterances can include voice commands, such as verbal requests from a user. Voice commands can include, for example: requests to perform an action (e.g., “play music,” “turn on the air conditioning,” “activate cruise control”); further information related to the request (e.g., “album XY,” “68 degrees Fahrenheit,” “60 mph for 30 minutes”); speech to be transcribed (e.g., “add to my to-do list…” or “send the following message to user A…”); and / or information requests (e.g., “What’s the traffic like on C?”, “What’s the weather like today?”, or “Where’s the nearest gas station?”).

[0049] Depending on the implementation, the audio data 155 can take various forms. Typically, the audio data 155 can be generated from one or more audio capture devices (e.g., one or more microphones). Figure 1A The audio data 155 is derived from time-series measurements of the audio capture device 116. In some cases, audio data 155 can be captured from a single audio capture device; in others, it can be captured from multiple audio capture devices, for example, multiple microphones may be located at different locations within the interior 100. In the latter case, the audio data may include one or more channels of time-correlated audio data from each audio capture device. The audio data at the capture point may include one or more channels of pulse-code modulation (PCM) data at a predefined sampling rate (e.g., 16 kHz), where each sample is represented by a predefined number of bits (e.g., 8, 16, or 24 bits per sample, where each sample includes an integer or floating-point value).

[0050] In some cases, audio data 155 may be processed after capture but before reception at audio interface 150 (e.g., preprocessed relative to speech processing). Processing may include one or more of the following: filtering in one or more of the time and frequency domains, applying noise reduction, and / or normalization. In one case, audio data may be converted to a time-varying measurement in the frequency domain, for example, by performing a Fast Fourier Transform to create one or more spectrogram data frames. In some cases, filter banks may be used to determine the values ​​of one or more frequency domain features, such as Mel filter banks or Mel-Frequency Cepstral Coefficients. In these cases, audio data 155 may include the outputs of one or more filter banks. In other cases, audio data 155 may include time-domain sampling and may be preprocessed within speech processing module 130. Different combinations of methods are possible. Therefore, audio data received at audio interface 150 may include any measurements performed along the audio processing pipeline.

[0051] The image data described herein can take various forms, similar to audio data, depending on the implementation. In one case, image capture device 110 may include a video capture device, wherein the image data comprises one or more video data frames. In another case, image capture device 110 may include a still image capture device, wherein the image data comprises one or more still image frames. Thus, image data can be derived from both video and still sources. References to image data herein may refer to image data derived, for example, from a two-dimensional array having height and width (e.g., equivalent to rows and columns of an array). In one case, image data may have multiple color channels, for example, including three color channels, red, green, and blue (RGB), wherein each color channel has a two-dimensional array of associated color values ​​(e.g., each array element is 8, 16, or 24 bits). Color channels may also be referred to as different image "planes". In some cases, only a single channel may be used, for example, to represent a "grayscale" or luminance channel. Depending on the application, different color spaces may be used. For example, an image capture device may naturally generate YUV image data frames that describe the characteristics of the luminance channel Y (e.g., illuminance) and the two opposing color channels U and V (e.g., two chromaticity components roughly aligned with blue-green and red-green). Similar to the audio data 155, the image data 145 can be processed after capture; for example, one or more image filtering operations may be applied, and / or the image data 145 may be resized and / or cropped.

[0052] refer to Figure 1A and Figure 1BFor example, one or more of the image interface 140 and audio interface 150 may be local to the hardware within the vehicle. For instance, each of the image interface 140 and audio interface 150 may include a wired coupling of a corresponding image and audio capture device to at least one processor configured to implement the voice processing module 130. In one embodiment, the image and audio interfaces 140, 150 may include a serial interface through which image and audio data 145, 155 can be received. In a distributed vehicle control system, the image and audio capture devices 140, 150 may be communicatively coupled to a central system bus, wherein the image and audio data 145, 155 may be stored in one or more storage devices (e.g., random access memory or solid-state storage). In the latter embodiment, the image and audio interfaces 140, 150 may include a communicative coupling of at least one processor configured to implement the voice processing module to one or more storage devices; for example, at least one processor may be configured to read data from a given memory location to access each of the image and audio data 145, 155. In some cases, the image and audio interfaces 140, 150 may include a wireless interface, wherein the voice processing module 130 may be remote relative to the vehicle. Different methods and combinations are possible.

[0053] Although Figure 1A An example is shown where person 102 is the driver of a motor vehicle; however, in other applications, one or more image and audio capture devices may be arranged to capture image data describing characteristics of a person who does not control the motor vehicle (e.g., a passenger). For example, a motor vehicle may have multiple image capture devices arranged to capture image data relating to a person present in one or more passenger seats within the vehicle (e.g., in different locations within the vehicle, such as the front and rear seats). Audio capture devices may also be similarly arranged to capture speech from different people; for example, microphones may be located in each door or door frame of the vehicle. In one case, multiple audio capture devices may be provided within the vehicle, and audio data may be captured from one or more of these devices to provide the data to audio interface 150. In one case, preprocessing of the audio data may include: selecting audio data from the channel considered closest to the person uttering the speech, and / or combining audio data from multiple channels within the motor vehicle. As described later, some of the examples described herein facilitate speech processing in vehicles with multiple passengers.

[0054] Example Speaker Preprocessing Module

[0055] Figure 2 An example speech processing device 200 is shown. For example, the speech processing device 200 can be used to implement... Figure 1A and Figure 1BThe voice processing device 120 is shown. The voice processing device 200 can be part of an in-vehicle automatic voice recognition system. In other cases, the voice processing device 200 can be used outside the vehicle, such as in a home or office.

[0056] The speech processing device 200 includes a speaker preprocessing module 220 and a speech processing module 230. The speech processing module 230 can be similar to... Figure 1B The speech processing module 230 is configured to receive image data 245 and output speaker feature vector 225. The speech processing module 230 is configured to receive audio data 255 and speaker feature vector 225, and use the audio data 255 and speaker feature vector 225 to generate linguistic features 260. (Note: The image interface 140 and audio interface 150 are omitted in this example for clarity; however, they can form part of the image input of the speaker preprocessing module 220 and the audio input of the speech processing module 230, respectively.)

[0057] The speech processing module 230 is implemented by a processor. The processor can be a processor of a local embedded computing system within the vehicle, and / or a processor of a remote server computing device (a so-called "cloud" processing device). In one case, the processor may include a portion of dedicated speech processing hardware, such as one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and so-called "system-on-a-chip" (SoC) components. In another case, the processor may be configured to process computer program code (e.g., firmware, etc.) stored in an accessible storage device and loaded into memory for processor execution. The speech processing module 230 is configured to parse the speech of a person (e.g., person 102) based on audio data 225 and image data 245. In this case, the speaker preprocessing module 220 preprocesses the image data 245 to generate a speaker feature vector 225. Similar to the speech processing module 230, the speaker preprocessing module 220 can be any combination of hardware and software. In one case, the speaker preprocessing module 220 and the speech processing module 230 may be implemented on a common embedded circuit board for the vehicle.

[0058] In one embodiment, the speech processing module 230 includes an acoustic model configured to process audio data 255 and predict phoneme data for parsing utterances. In this embodiment, linguistic features 260 may include phoneme data. The phoneme data may be associated with one or more phoneme symbols, for example, from a predefined alphabet or dictionary. In one embodiment, the phoneme data may include a predicted sequence of phonemes; in another embodiment, the phoneme data may include a set of phoneme components (e.g., phoneme symbols and / or sub-symbols from a predefined alphabet or dictionary) and probabilities of one or more of a set of state transitions (e.g., for a hidden Markov model). The acoustic model may be configured to receive audio data in the form of an audio feature vector. The audio feature vector may include numerical values ​​representing one or more of the Mel-frequency cepstral coefficients (MFCCs) and filter bank outputs. In some cases, the audio feature vector may be associated with the current time window (often referred to as a “frame”) and include differences related to feature changes between the current window and one or more other time windows (e.g., previous windows). The width of the current window may be in the range w milliseconds; for example, in one embodiment, w may be approximately 25 milliseconds. Other features may include signal energy metrics and logarithmically scaled outputs, etc. After preprocessing, the audio data 255 may include frames (e.g., vectors) with multiple elements (e.g., from 10 to over 1000 elements), each element including a numerical representation associated with a specific audio feature. In some examples, there may be approximately 25-50 Mel filter bank features, a set of internal features of similar size, a set of delta features of similar size (e.g., representing the first derivative), and a set of double delta features of similar size (e.g., representing the second derivative).

[0059] Speaker preprocessing module 220 can be configured to obtain speaker feature vector 225 in a variety of different ways. In one case, speaker preprocessing module 220 can obtain at least a portion of speaker feature vector 225 from memory, for example, via a lookup operation. In another case, a portion of speaker feature vector 225 comprising vectors i and / or x as described below can be obtained from memory. In this case, image data 245 can be used to determine the specific speaker feature vector 225 to be obtained from memory. For example, image data 245 can be classified by speaker preprocessing module 220 to select a specific user from a set of registered users. In this case, speaker feature vector 225 can include a numerical representation of features associated with the selected specific user. In another case, speaker preprocessing module 220 can compute speaker feature vector 225. For example, speaker preprocessing module 220 can compute a compressed or dense numerical representation of salient information within image data 245. This can include a vector with multiple elements that is smaller than the size of image data 245. In this context, the speaker preprocessing module 220 can implement an information bottleneck to compute the speaker feature vector 225. In one case, the computation is determined based on a set of parameters, such as a set of weights, biases, and / or probability coefficients. The values ​​of these parameters can be determined through a training phase using a training dataset. In another case, the speaker feature vector 225 can be buffered or stored as a static value after a set of computations. In this case, the speaker feature vector 225 can be retrieved from memory in subsequent utterances based on image data 245. Further examples illustrating how the speaker feature vector can be computed are given below. Where the speaker feature vector 225 includes components related to lip movements, these components can be provided in real-time or near real-time, and may not need to be retrieved from a data storage device.

[0060] In one case, the speaker feature vector 225 may comprise a fixed-length one-dimensional array of values ​​(e.g., a vector), where each element of the array has a value. In other cases, the speaker feature vector 225 may comprise a multi-dimensional array, for example, having two or more dimensions representing multiple one-dimensional arrays. The values ​​may comprise integer values ​​(e.g., 8 bits, ranging from 0 to 255, within a range set by a specific length) or floating-point values ​​(e.g., defined as 32-bit or 64-bit floating-point values). Floating-point values ​​may be used if normalization is applied to a visual feature tensor, for example, if values ​​are mapped to a range of 0 to 1 or -1 to 1. As an example, the speaker feature vector 225 may comprise an array with 256 elements, where each element is an 8-bit or 16-bit value, although the form may vary depending on the implementation. Typically, the speaker feature vector 225 has less information content than the corresponding image data frame. For example, using the aforementioned example, a speaker feature vector 225 with 8-bit values ​​and a length of 256 is less than a 640 x 480 video frame with three 8-bit channels – a 2048-bit comparison of 7,372,800 bits. The information content can be measured in the form of bit or entropy measurements.

[0061] In one embodiment, the speech processing module 230 includes an acoustic model, and the acoustic model includes a neural network architecture. For example, the acoustic model may include one or more of the following: a deep neural network (DNN) architecture with multiple hidden layers; a mixture model, including a neural network architecture and one or more of Gaussian mixture models (GMMs) and hidden Markov models (HMMs); ​​and a connectionist temporal classification (CTC) model, for example, including one or more recurrent neural networks that operate on the input sequence and generate a sequence of linguistic features as output. The acoustic model may output frame-level predictions (e.g., for phoneme symbols or sub-symbols) and use previous (and in some cases future) predictions to determine the possible or most probable phoneme data sequence for the utterance. Methods such as beam search and the Viterbi algorithm may be used at the output of the acoustic model to further determine the phoneme data sequence output from the acoustic model. Training of the acoustic model may be performed step-by-step over time.

[0062] In cases where the speech processing module 230 includes an acoustic model and the acoustic model includes a neural network architecture (e.g., the acoustic model is a "neural" acoustic model), the speaker feature vector 225 can be provided as input to the neural network architecture along with the audio data 255. The speaker feature vector 225 and the audio data 255 can be combined in various ways. In a simple case, the speaker feature vector 225 and the audio data 255 can be concatenated into a longer combined vector. In another case, different input preprocessing can be performed on each of the speaker feature vector 225 and the audio data 255; for example, one or more attention, feedforward, and / or embedding layers can be applied, and then the results of these layers can be combined. Different sets of layers can be applied to different inputs. In other cases, the speech processing module 230 can include another form of statistical model, such as a probabilistic acoustic model, where the speaker feature vector 225 includes one or more numerical parameters (e.g., probability coefficients) for configuring the speech processing module 230 for a specific speaker.

[0063] Example speech processing device 200 provides improvements for speech processing within a vehicle. High levels of ambient noise, such as road and engine noise, may exist within a vehicle. Sound distortion may also occur due to the enclosed interior space of a motor vehicle. In the comparative example, these factors may make audio data processing difficult; for example, speech processing module 230 may fail to generate linguistic features 260 and / or generate poorly matching sequences of linguistic features 260. However, Figure 2The arrangement allows the speech processing module 230 to be configured or adapted based on speaker features determined from image data 245. This provides additional information to the speech processing module 230, enabling it to select linguistic features consistent with a particular speaker, such as by utilizing correlations between appearance and acoustic characteristics. These correlations can be long-term temporal correlations such as overall facial appearance and / or short-term temporal correlations such as specific lip and mouth positions. This can lead to higher accuracy despite challenging noise and acoustic contexts. This can help reduce speech parsing errors, for example, by improving end-to-end transcription paths and / or improving audio interfaces used to execute voice commands. In some cases, this example is able to utilize existing driver-facing cameras, which are typically configured to monitor the driver for drowsiness and / or inattention. In some cases, there may be speaker-related feature vector components acquired based on the identified speaker, and / or speaker-related feature vector components that include mouth movement features. The latter component can be determined based on features not configured for a single user; for example, a generic feature applicable to all users can be applied, but mouth movements are associated with the speaker. In some other cases, the extraction of mouth movement features can be configured based on the specific user being identified.

[0064] Facial recognition example

[0065] Figure 3 An example speech processing device 300 is shown. The speech processing device 300 illustrates what can be used to implement... Figure 2 An additional component of the speaker preprocessing module 220 in the system. Figure 3 Some of the components shown are with Figure 2 The corresponding components shown are similar and have similar reference numbers. (Referring to the above...) Figure 2 The described features can also be applied to Figure 3 Example 300. Similar to... Figure 2 Example speech processing device 200, Figure 3 The example speech processing device 300 includes a speaker preprocessing module 320 and a speech processing module 330. The speech processing module 330 receives audio data 355 and a speaker feature vector 325, and calculates a set of linguistic features 360. This can be done according to the above reference... Figure 2 The example described is configured in a similar way to the speech processing module 330.

[0066] exist Figure 3 The image shows several sub-components of the speaker preprocessing module 320. These include a face recognition module 370, a vector generator 372, and a data repository 374. While these are shown in... Figure 3The components shown are sub-components of the speaker preprocessing module 320, but in other examples, they can be implemented as separate components. Figure 3 In this example, speaker preprocessing module 320 receives image data 345 describing features of a person's facial regions. This person may include a driver or passenger in a vehicle, as described above. Face recognition module 370 performs face recognition on the image data to identify the person, such as a driver or passenger in a vehicle. Face recognition module 370 may include any combination of hardware and software for performing face recognition. In one case, face recognition module 370 may be implemented using readily available hardware components such as the B5T-007001 provided by Omron Electronics Inc. In this example, face recognition module 370 detects the user based on image data 345 and outputs a user identifier 376. User identifier 376 is passed to vector generator 372. Vector generator 372 uses user identifier 376 to obtain a speaker feature vector 325 associated with the identified person. In some cases, vector generator 372 may obtain speaker feature vector 325 from data repository 374. Speaker feature vector 325 is then passed to speech processing module 330 for use, as referenced. Figure 2 As described.

[0067] exist Figure 3 In the example, vector generator 372 can obtain speaker feature vectors 325 in different ways based on a set of operating parameters. In one case, the operating parameters include a parameter indicating whether a specific number of speaker feature vectors 325 have been computed for a particular identified user (e.g., the user identified by user identifier 376). In another case, a threshold is defined that is associated with the number of previously computed speaker feature vectors. If the threshold is 1, speaker feature vectors 325 can be computed for the first utterance and then stored in data repository 374; for subsequent utterances, speaker feature vectors 325 can be retrieved from data repository 374. If the threshold is greater than 1, for example, n, n speaker feature vectors 325 can be generated, and then the (n+1)th speaker feature vector 325 can be obtained as a composite function of the previous n speaker feature vectors 325 retrieved from data repository 374. The composite function can include an average or interpolation. In one case, once the (n+1)th speaker feature vector 325 has been calculated, it is used as a static speaker feature vector for a configurable number of future utterances.

[0068] In the example above, using a data repository 374 to store the speaker feature vector 325 can reduce the runtime computational requirements of the in-vehicle system. For example, the data repository 374 may include a local data storage device within the vehicle, and therefore, the speaker feature vector 325 can be retrieved from the data repository 374 for a specific user, instead of being computed by the vector generator 372.

[0069] In one scenario, at least one computational function used by the vector generator 372 may involve cloud processing resources (e.g., remote server computing devices). In this case, given the limited connectivity between the vehicle and the cloud processing resources, the speaker feature vector 325 can be obtained as a static vector from local storage without relying on any functionality provided by the cloud processing resources.

[0070] In one scenario, the speaker preprocessing module 320 can be configured to generate a user profile for each newly identified person within the vehicle. For example, before or at the time of detecting utterances, such as those captured by an audio capture device, the face recognition module 370 can attempt to match image data 345 with previously observed faces. If no match is found, the face recognition module 370 can generate (or instruct to generate) a new user identifier 376. In another scenario, components of the speaker preprocessing module 320 (e.g., the face recognition module 370 or vector generator 372) can be configured to generate a new user profile when no match is found, wherein the new user identifier can be used to index the new user profile. The speaker feature vector 325 can then be associated with the new user profile, and the new user profile can be stored in a data repository 374 for retrieval during future matching by the face recognition module 370. In this way, the in-vehicle image capture device can be used for face recognition to select a user-specific speech recognition profile. The user profile can be calibrated through a registration process (e.g., when a driver first uses the vehicle) or learned based on data collected during use.

[0071] In one scenario, the speaker processing module 320 can be configured to perform a reset of the data repository 374. At manufacturing time, the data repository 374 may not have user profile information. During use, as described above, a new user profile can be created and added to the data repository 374. The user can command a reset of the stored user identifier. In some cases, a reset can only be performed during professional service (e.g., when a car is being maintained at a service shop or sold through an authorized dealer). In other cases, a reset can be performed at any time using a password provided by the user.

[0072] In an example where the vehicle includes multiple image capture devices and multiple audio capture devices, the speaker preprocessing module 320 can provide further functionality to determine appropriate facial regions from one or more captured images. In one case, audio data from multiple audio capture devices can be processed to determine the closest audio capture device associated with the utterance. In this case, the closest image capture device associated with the determined closest audio capture device can be selected, and image data 345 from that device can be sent to the face recognition module 370. In another case, the face recognition module 370 can be configured to receive multiple images from multiple image capture devices, wherein each image includes an associated flag indicating whether the image will be used to identify the currently speaking user. In this way, Figure 3 The voice processing device 300 can be used to identify speakers from among multiple people inside the vehicle and configure the voice processing module 330 according to the specific characteristics of the speaker. This can also improve voice processing within the vehicle when multiple people are speaking in the confined interior space of the vehicle.

[0073] vector i

[0074] In some of the examples described herein, the speaker feature vector (e.g., speaker feature vector 225 or 325) may include audio data-based features (e.g., Figure 2 and Figure 3 The data is generated from the audio data (255 or 355). This is in Figure 3The dashed lines represent the speaker's eigenvectors. In one case, at least a portion of the speaker's eigenvectors may include vectors generated based on factor analysis. In this case, the utterance can be represented as a vector M, which is a linear function of one or more factors. These factors can be combined in linear and / or nonlinear models. One of these factors may include a speaker- and session-independent hypervector m. This can be based on a Universal Background Model (UBM). Another factor may include a speaker-dependent vector w. This latter factor may also depend on the channel or session, or may provide another factor that depends on the channel and / or session. In one case, factor analysis is performed using a Gaussian Mixture Model (GMM). In a simple case, the speaker's utterance can be represented by a hypervector M, which is determined as M = m + Tw, where T is a matrix defining at least the speaker's subspace. The speaker-dependent vector w may have multiple elements with floating-point values. In this case, the speaker's eigenvectors may be based on the speaker-dependent vector w. Najim Dehak, Patrick Kenny, Reda Dehak, Pierre Dumouchel, and Pierre Ouellet describe a method for calculating w (sometimes referred to as the "i-vector") in their paper "Front-End Factor Analysis For Speaker Verification," published in IEEE Transactions on Audio, Speech and Language Processing 19, Vol. 4, pp. 788-798, 2010, which is incorporated herein by reference. In some examples, at least a portion of the speaker feature vector includes at least a portion of the i-vector. The i-vector can be viewed as a speaker-related vector determined for the utterance from the audio data.

[0075] exist Figure 3In the example, vector generator 372 can compute i-vectors for one or more utterances. If no speaker feature vectors are stored in data repository 374, vector generator 372 can compute i-vectors based on one or more audio data frames for utterance 355. In this example, vector generator 372 can repeat i-vector computation for each utterance (e.g., each voice query) until a threshold number has been calculated for a specific user (e.g., a user identified using user identifier 376 determined from facial recognition module 370). In this case, after a specific user has been identified based on image data 345, the i-vector for each utterance of that user is stored in data repository 374. The i-vectors are also used to output speaker feature vector 325. Once the threshold number has been calculated, for example, approximately 100 i-vectors have been computed, vector generator 372 can be configured to use the i-vectors stored in data repository 374 to compute a profile for a specific user. The profile can use the user identifier 376 as an index and can include static (e.g., invariant) i-vectors, which are computed as a composite function of the stored i-vectors. The vector generator 372 can be configured to compute the profile upon receiving the (n+1)th query or as part of a background or periodic function. In one case, the static i-vectors can be computed as the average of the stored i-vectors. Once the profile is generated by the vector generator 372 and stored in the data repository 374, for example, by associating the profile with a specific user using the user identifier, the profile can be retrieved from the data repository 374 and used for future speech parsing instead of computed i-vectors for the user. This reduces the computational overhead of generating speaker feature vectors and reduces the variance of the i-vectors.

[0076] vector x

[0077] In some examples, a neural network architecture can be used to compute speaker feature vectors, such as speaker feature vector 225 or 325. For example, in one case, Figure 3The vector generator 372 of the speaker preprocessing module 320 may include a neural network architecture. In this case, the vector generator 372 can compute at least a portion of the speaker feature vector by reducing the dimension of the audio data 355. For example, the vector generator 372 may include one or more deep neural network layers configured to receive one or more frames of the audio data 355 and output fixed-length vector outputs (e.g., one vector per language). One or more pooling, non-linear functions, and SoftMax layers may also be provided. In one case, the speaker feature vector may be generated based on an x-vector, as described in the paper “Spoken Language Recognition using X-vectors” by David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, published in the Odyssey 2018 (pp. 105–111), which is incorporated herein by reference.

[0078] The x vector can be used in a similar manner to the i vector described above, and the above method applies to speaker feature vectors generated using both the x vector and the i vector. In one case, both the i vector and the x vector can be determined, and the speaker feature vector may include a hypervector comprising elements from both the i vector and the x vector. Because both the i vector and the x vector include numerical elements, typically floating-point numbers and / or values ​​normalized within a given range, the i vector and the x vector can be combined by concatenation or weighted summation. In this case, the data repository 374 may include stored values ​​for one or more of the i vector and x vector, whereby a static value is calculated and stored along with a specific user identifier for future retrieval once a threshold is reached. In one case, interpolation can be used to determine the speaker feature vector from one or more i vectors and x vectors. In one case, interpolation can be performed by averaging different speaker feature vectors from the same vector source.

[0079] When the speech processing module includes a neuroacoustic model, a fixed-length format for the speaker feature vector can be defined. The defined speaker feature vector (e.g., as shown by...) can then be used... Figure 2 and Figure 3The neuroacoustic model is trained using the speaker preprocessing module 220 or 320 (as determined in the original text). If the speaker feature vector includes one or more elements derived from vectors i and x, the neuroacoustic model can "learn" to configure acoustic processing based on speaker-specific information embodied or embedded in the speaker feature vector. This can improve the accuracy of acoustic processing, particularly within vehicles such as motor vehicles. In this case, image data provides a mechanism for quickly associating a specific user with computed or stored vector elements.

[0080] Example speech processing module

[0081] Figure 4 An example speech processing module 400 is shown. The speech processing module 400 can be used to implement Figure 1. Figure 2 and Figure 3 The speech processing module is 130, 230, or 330. In other examples, other speech processing modules can be configured.

[0082] As in the previous example, the speech processing module 400 receives audio data 455 and a speaker feature vector 425. The audio data 455 and the speaker feature vector 425 can be configured according to any of the examples described herein. Figure 4 In the example, the speech processing module 400 includes an acoustic model 432, a language model 434, and a speech parser 436. As previously described, the acoustic model 432 generates phoneme data 438. The phoneme data may include one or more predicted sequences of phoneme symbols or sub-symbols, or other forms of raw language units. In some cases, multiple predicted sequences may be generated together with probability data indicating the likelihood of a particular symbol or sub-symbol at each time step.

[0083] Phoneme data 438 is transmitted to language model 434, for example, acoustic model 432 is communicatively coupled to language model 434. Language model 434 is configured to receive phoneme data 438 and generate transcription 440. Transcription 440 may include text data, such as characters, word parts (e.g., stems, endings, etc.), or sequences of words. Characters, word parts, and words can be selected from a predefined dictionary (e.g., a predefined set of possible outputs at each time step). In some cases, phoneme data 438 may be processed before being transmitted to language model 434, or may be preprocessed by language model 434. For example, beamforming may be applied to the probability distribution (e.g., for phonemes) output from acoustic model 432.

[0084] Language model 434 is communicatively coupled to speech parser 436. Speech parser 436 receives transcription 440 and uses transcription 440 to parse speech. In some cases, as a result of parsing speech, speech parser 436 generates speech data 442. Speech parser 436 can be configured to determine commands and / or command data associated with the speech based on the transcription. In one case, language model 434 can, for example, utilize probabilistic information of units within the text to generate multiple possible text sequences, and speech parser 436 can be configured to determine the final text output (e.g., in ASCII or Unicode character encoding), or spoken commands or command data. If transcription 440 is determined to include a spoken command, speech parser 436 can be configured to execute the command or instruct the execution of the command based on the command data. This can produce response data, which is output as speech data 442. Speech data 442 can include a response to be relayed to the person who spoke the speech, for example, a command instruction for providing output on dashboard 108 and / or via the vehicle's audio system. In some cases, language model 434 may include a statistical language model, and discourse parser 436 may include a separate "meta" language model configured to re-evaluate the alternative hypotheses output by the statistical language model. This can be achieved using an ensemble model that uses voting to determine the final output (e.g., the final transcription or command identifier).

[0085] Figure 4 An example is illustrated with solid lines, where acoustic model 432 receives speaker feature vector 425 and audio data 455 as input and uses this input to generate phoneme data 438. For example, acoustic model 432 may include a neural network architecture (including a hybrid model with other non-neural components), and speaker feature vector 425 and audio data 455 may be provided as input to the neural network architecture, where phoneme data 438 is generated based on the output of the neural network architecture.

[0086] Figure 4The dashed lines in the diagram illustrate additional couplings that can be configured in certain implementations. In a first case, the speaker feature vector 425 can be accessed by one or more of the language model 434 and the utterance parser 436. For example, if the language model 434 and the utterance parser 436 also include corresponding neural network architectures, these architectures can be configured to receive the speaker feature vector 425 as additional input, for example, in addition to the phoneme data 438 and the transcription 440, respectively. If the utterance data 442 includes a command identifier and one or more command parameters, then the complete speech processing module 400 can be trained end-to-end given a training set with a base truth output and training samples for the audio data 455 and the speaker feature vector 425.

[0087] In the second implementation, Figure 4 The speech processing module 400 may include one or more recursive connections. In one case, the acoustic model may include a recursive model, such as an LSTM. In other cases, feedback may exist between modules. Figure 4 In the diagram, there are dashed lines indicating a first recursive coupling between the discourse parser 436 and the language model 434, and dashed lines indicating a second recursive coupling between the language model 434 and the acoustic model 432. In the second case, the current state of the discourse parser 436 can be used to configure the future predictions of the language model 434, and the current state of the language model 434 can be used to configure the future predictions of the acoustic model 432. In some cases, the recursive coupling can be omitted to simplify the processing pipeline and achieve simpler training. In one case, the recursive coupling can be used to compute the attention or weighting vector applied at the next time step.

[0088] Neural Speaker Preprocessing Module

[0089] Figure 5 An example speech processing device 500 is shown, which uses a neural speaker preprocessing module 520 and a neural speech processing module 530. Figure 5 In the middle, the speaker preprocessing module 520 (which can implement) Figure 2 and Figure 3 Modules 220 or 320 in the module include a neural network architecture 522. Figure 5 In this configuration, the neural network architecture 522 is configured to receive image data 545. In other cases, the neural network architecture 522 can also receive audio data, such as audio data 355, for example, as... Figure 3 The dashed path is shown in the image. In these other cases, Figure 3 The vector generator 372 may include a neural network architecture 522.

[0090] exist Figure 5In this design, neural network architecture 522 includes at least a convolutional neural network architecture. In some architectures, one or more feedforward neural network layers may exist between the last convolutional neural network layer and the output layer of neural network architecture 522. Neural network architecture 522 may include adapted forms of AlexNet, VGGNet, GoogLeNet, or ResNet architectures. Neural network architecture 522 can be replaced in a modular manner when more precise architectures become available.

[0091] The neural network architecture 522 outputs at least one speaker feature vector 525, wherein the speaker feature vector can be derived and / or used as described in any other example. Figure 5 The illustration shows a scenario where image data 545 comprises, for example, multiple frames from a camera, where these frames describe features of a person's facial regions. In this case, a neural network architecture can be used to compute multiple speaker feature vectors 525, for example, one speaker feature vector 525 for each input frame of the image data. In other cases, a many-to-one relationship can exist between input data frames and speaker feature vectors. It should be noted that when using a recurrent neural network system, the samples of input image data 545 and the output speaker feature vectors 525 do not need to be synchronized in time; for example, the recurrent neural network architecture can act as a time-varying encoder (or integrator). In one case, the neural network architecture 522 can be configured to generate x vectors as described above. In another case, the x vector generator can be configured to receive image data 545, process the image data using a convolutional neural network architecture, and then combine the output of the convolutional neural network architecture with an audio-based x vector. In yet another case, a known x vector configuration can be extended to receive both image data and audio data and generate a single speaker feature vector embodying information from both modal paths.

[0092] exist Figure 5 In this context, the neural speech processing module 530 is a speech processing module that includes a neural network architecture, such as one of modules 230, 330, and 400. For example, the neural speech processing module 530 may include a hybrid DNN-HMM / GMM system and / or a fully neural CTC system. Figure 5In this embodiment, the neural speech processing module 530 receives frames of audio data 555 as input. Each frame may correspond to a time window, for example, a window of w ms through time-series data from an audio capture device. Frames of audio data 555 may be asynchronous with frames of image data 545; for example, frames of audio data 555 may have a higher frame rate. Similarly, hold-up mechanisms and / or recurrent neural network architectures may be applied within the neural speech processing module 530 to provide temporal encoding and / or integration of samples. As in other examples, the neural speech processing module 530 is configured to process frames of audio data 555 and speaker feature vectors 525 to generate a set of linguistic features. As discussed herein, references to neural network architectures include one or more neural network layers (in one case, a “deep” architecture with one or more hidden layers and multiple layers), where each layer may be separated from subsequent layers in a non-linear manner (e.g., tanh units or rectified linear units (RELU)). Other functionality may be embodied in layers including pooling operations.

[0093] The neural speech processing module 530 may include, for example: Figure 4 One or more components are shown. For example, the neural speech processing module 530 may include an acoustic model having at least one neural network. Figure 5 In the example, the neural speaker preprocessing module 520 and the neural speech processing module 530 can be trained together. In this case, the training set may include frames of image data 545, frames of audio data 555, and ground truth linguistic features (e.g., ground truth phoneme sequences, text transcriptions, or voice command classifications and command parameter values). Both the neural speaker preprocessing module 520 and the neural speech processing module 530 can be trained end-to-end using this training set. In this case, the error between the predicted linguistic features and the ground truth linguistic features can be backpropagated through the neural speech processing module 530 and then through the neural speaker preprocessing module 520. The parameters of the two neural network architectures can then be determined using gradient descent. In this way, the neural network architecture 522 of the neural speaker preprocessing module 520 can “learn” the following parameter values ​​(e.g., values ​​for weights and / or biases of one or more neural network layers) that generate one or more speaker feature vectors 525 that at least improve the acoustic processing in the vehicle environment, wherein the neural speaker preprocessing module 520 learns to extract features from the human facial region that help improve the accuracy of the output linguistic features.

[0094] Training of the neural network architectures described in this paper is typically not performed on in-vehicle devices (although it can be performed when needed). In one scenario, training can be performed on a computing device with access to substantial processing resources, such as a server computer device with multiple processing units (whether CPU, GPU, FPGA, or other dedicated processor architecture) and a large memory portion for storing batches of training data. In some cases, training can be performed using coupled accelerator devices (e.g., coupled FPGA or GPU-based devices). In other cases, trained parameters can be transferred from a remote server device to an embedded system within the vehicle, for example, as part of an over-the-air update.

[0095] Example of acoustic model selection

[0096] Figure 6 An example speech processing module 600 is shown, which uses a speaker feature vector 625 to configure an acoustic model. Speech processing module 600 can be used to at least partially implement one of the speech processing modules described in other examples. Figure 6 In this embodiment, the speech processing module 600 includes an acoustic model configuration database 632, an acoustic model selector 634, and acoustic model instances 636. The acoustic model configuration database 632 stores multiple parameters used to configure acoustic models. In this example, the acoustic model instance 636 may include a generic acoustic model that is instantiated (e.g., configured or calibrated) using a specific set of parameter values ​​from the acoustic model configuration database 632. For example, the acoustic model configuration database 636 may store multiple acoustic model configurations. Each acoustic model configuration may be associated with a different user, including one or more default acoustic model configurations used when no user is detected or when a user is detected but not explicitly identified.

[0097] In some cases, the speaker feature vector 625 can be used to represent a specific regional accent rather than (and) a specific user. This can be useful in countries where many different regional accents may exist (e.g., India). In this case, the speaker feature vector 625 can be used to dynamically load an acoustic model based on accent recognition performed using the speaker feature vector 625. This is possible, for example, if the speaker feature vector 625 includes an x-vector as described above. This can be useful in cases where multiple accent models are stored in the vehicle's memory (e.g., multiple acoustic model configurations for each accent). This can then allow the use of multiple separately trained accent models.

[0098] In one scenario, the speaker feature vector 625 could include a classification of people inside the vehicle. For example, it could be derived from... Figure 3The user identifier 376 output by the facial recognition module 370 in the system derives the speaker feature vector 625. In another case, the speaker feature vector 625 may include features derived from, for example... Figure 5 The neural speaker preprocessing module, such as module 520, outputs a set of classifications and / or probabilities. In the latter case, the neural speaker preprocessing module may include a softmax layer that outputs a "probability" (including a "not recognized" classification) for a set of potential users. In this case, one or more frames of the input image data 545 can produce a single speaker feature vector 525.

[0099] exist Figure 6 In this process, the acoustic model selector 634 receives, for example, a speaker feature vector 625 from the speaker preprocessing module and selects an acoustic model configuration from the acoustic model configuration database 632. This can be done in accordance with the above. Figure 3 The example operates in a similar manner. If the speaker feature vector 625 includes a set of user classifications, the acoustic model selector 634 can select an acoustic model configuration based on these classifications, for example, by sampling the probability vector and / or selecting the highest probability value as the identified person. Parameter values ​​associated with the selected configuration can be obtained from the acoustic model configuration database 632 and used to instantiate an acoustic model instance 636. Thus, different acoustic model instances can be used for different identified users within the vehicle.

[0100] exist Figure 6 In this context, an acoustic model instance 636, configured by acoustic model selector 634 using configuration obtained from acoustic model configuration database 632, also receives audio data 655. Acoustic model instance 636 is configured to generate phoneme data 660 for parsing utterances associated with audio data 655 (e.g., features of the utterance described within audio data 655). Phoneme data 660 may include, for example, sequences of phoneme symbols from a predefined alphabet or dictionary. Therefore, in Figure 6 In the example, the acoustic model selector 634 selects an acoustic model configuration from the database 632 based on the speaker feature vector, and the acoustic model configuration is used to instantiate an acoustic model instance 636 to process audio data 655.

[0101] Acoustic model instance 636 may include both neural and non-neural architectures. In one case, acoustic model instance 636 may include a non-neural model. For example, acoustic model instance 636 may include a statistical model. The statistical model may use symbol frequencies and / or probabilities. In one case, the statistical model may include a Bayesian model, such as a Bayesian network or a classifier. In these cases, the acoustic model configuration may include a specific set of symbol frequencies and / or prior probabilities that have been measured in different environments. The acoustic model selector 634 thus allows for the determination of the specific source of the utterance (e.g., a person or user) based on visual (and in some cases, visual and audio) information, which can provide an improvement over generating phoneme sequences 660 using only audio data 655 itself.

[0102] In another scenario, acoustic model instance 636 may include a neural model. In this case, both acoustic model selector 634 and acoustic model instance 636 may include a neural network architecture. In this scenario, acoustic model configuration database 632 can be omitted, and acoustic model selector 634 can provide vector inputs to acoustic model instance 636 to configure the instance. In this scenario, training data can be constructed from image data 625, audio data 655, and a ground truth set for generating speaker feature vector 625 and phoneme output 660. Such a system can be trained jointly.

[0103] Image preprocessing examples

[0104] Figure 7 and Figure 8 An example image preprocessing operation is shown, which can be applied to image data obtained from inside a vehicle, such as a motor vehicle. Figure 7 An example image preprocessing pipeline 700 including an image preprocessor 710 is shown. The image preprocessor 710 may include any combination of hardware and software to achieve the functionality described herein. In one case, the image preprocessor 710 may include hardware components forming part of image capture circuitry coupled to one or more image capture devices; in another case, the image preprocessor 710 may be implemented by computer program code (e.g., firmware) executed by a processor of an onboard control system. In one case, the image preprocessor 710 may be implemented as part of a speaker preprocessing module described in the examples herein; in other cases, the image preprocessor 710 may be communicatively coupled to the speaker preprocessing module.

[0105] exist Figure 7 In the image preprocessor 710, image data 745 is received, for example from... Figure 1AThe image capture device 110 captures an image. The image preprocessor 710 processes the image data to extract one or more portions of the image data. Figure 7 The output 750 of the image preprocessor 710 is shown. Output 750 may include one or more image annotations, such as metadata associated with one or more pixels of image data 745, and / or features defined using pixel coordinates within image data 745. Figure 7 In the example, image preprocessor 710 performs face detection on image data 745 to determine a first image region 752. The first image region 752 can be cropped and extracted as an image portion 762. The first image region 752 can be defined using a bounding box (e.g., at least the top-left and bottom-right (x, y) pixel coordinates of a rectangular region). Face detection can be a preliminary step for face recognition; for example, face detection can determine facial regions within image data 745, and face recognition can classify facial regions as belonging to a given person (e.g., within a group of people). Figure 7 In the example, the image preprocessor 710 also identifies the mouth region within the image data 745 to determine the second image region 754. The second image region 754 can be cropped and extracted as an image portion 764. A bounding box can also be used to define the second image region 754. In one case, the first and second image regions 752, 754 can be determined with respect to a set of detected facial features 756. These facial features 756 may include one or more of the eye, nose, and mouth regions. The detection of one or more of the facial features 756 and / or the first and second regions 752, 754 can be performed using neural network methods or known face detection algorithms, such as the Viola-Jones face detection algorithm described by Paul Viola and Michael J Jones in "Robust Real-Time Face Detection" published in the International Journal of Computer Vision 57 in the Netherlands in 2004, pages 137-154, which is incorporated herein by reference. In some examples, one or more of the first and second image regions 752, 754 are used by the speaker preprocessing module described herein to obtain speaker feature vectors. For example, the first image region 752 can be... Figure 3 The facial recognition module 370 provides input image data (i.e., is used to provide image data 345). See below for reference. Figure 9 Describe an example using the second image region 754.

[0106] Figure 8The effect of using an image capture device configured to capture electromagnetic radiation with infrared wavelengths is shown. In some cases, by means of, Figure 1A Image data 810 acquired by an image capture device such as image capture device 110 may be affected by low-light environments. Figure 8 In the image data 810, there are partially obscured facial areas (e.g., including...). Figure 7 The shadow area 820 of the first and second image regions 752, 754 in the image. In these cases, an image capture device configured to capture electromagnetic radiation with infrared wavelengths can be used. This may include providing... Figure 1A The image capture device 110 is adapted (e.g., a removable filter in hardware and / or software) and / or provides a near-infrared (NIR) camera. The output of such an image capture device is schematically shown as image data 830. In image data 830, the facial region 840 is reliably captured. In this case, image data 830 provides an illumination-invariant representation, for example, unaffected by illumination changes (e.g., illumination changes that may occur during nighttime driving). In these cases, image data 830 may be image data provided to image preprocessor 710 and / or speaker preprocessing module as described herein.

[0107] Lip reading example

[0108] In some examples, the speaker feature vector described herein may include at least one set of elements representing the features of a person's mouth or lips. In these cases, the speaker feature vector may be speaker-dependent because it varies based on image data content describing the features of a person's mouth or lip region. Figure 5 In the example, the neural speaker preprocessing module 520 can encode lip or mouth features used to generate the speaker feature vector 525. These can be used to improve the performance of the speech processing module 530.

[0109] Figure 9 Another example speech processing apparatus 900 is shown, which uses lip features to form at least a portion of a speaker feature vector. Similar to the previous examples, the speech processing apparatus 900 includes a speaker preprocessing module 920 and a speech processing module 930. The speech processing module 930 receives audio data 955 (in this case, audio data frames) and outputs linguistic features 960. The speech processing module 930 can be configured according to other examples described herein.

[0110] exist Figure 9In this example, the speaker preprocessing module 920 is configured to receive two different image data sources. The speaker preprocessing module 920 receives a first set of image data 962 describing features of a person's facial regions. This may include data derived from... Figure 7 The image preprocessor 710 extracts a first image region 762. The speaker preprocessing module 920 also receives a second set of image data 964 describing the features of a person's lips or mouth region. This may include data extracted by the image preprocessor 710. Figure 7 The second image region 764 is extracted by the image preprocessor 710. The second set of image data 964 can be relatively small, for example, using... Figure 1A The image capture device 110 acquires a smaller cropped portion of a larger image. In other examples, the first and second sets of image data 962, 964 may be uncropped and may include copies of a set of images from the image capture device. Different configurations are possible; cropping image data can improve processing speed and training, but neural network architectures can be trained to operate on various image sizes.

[0111] Speaker preprocessing module 920 includes Figure 9 The system consists of two components: a feature acquisition component 922 and a lip feature extractor 924. The lip feature extractor 924 forms part of the lip reading module. The feature acquisition component 922 can be pressed and... Figure 3 The speaker preprocessing module 320 is configured in a similar manner. In this example, the feature acquisition component 922 receives a first set of image data 962 and outputs a vector portion 926 consisting of one or more of the i-vectors and x-vectors (e.g., as described above). In one case, the feature acquisition component 922 receives a single image for each utterance, while the lip feature extractor 924 and the speech processing module 930 receive multiple frames within the time of the utterance. In one case, if the face recognition performed by the feature acquisition component 922 has a confidence value below a threshold, the first set of image data 962 can be updated (e.g., by using another / current video frame) and the face recognition can be reapplied until the confidence value meets the threshold (or exceeds a predefined number of attempts). See reference... Figure 3 As described, the vector portion 926 can be calculated based on the audio data 955 for a first number of utterances, and then retrieved from memory as a static value when the number of utterances exceeds the first number.

[0112] The lip feature extractor 924 receives a second set of image data 964. The second set of image data 964 may include cropped frames of image data focused on the mouth or lip region. The lip feature extractor 924 may receive the second set of image data 964 at the frame rate of the image capture device and / or at a subsampled frame rate (e.g., every 2 frames). The lip feature extractor 924 outputs a set of vector portions 928. These vector portions 928 may include the output of an encoder with a neural network architecture. The lip feature extractor 924 may include a convolutional neural network architecture for providing fixed-length vector outputs (e.g., 256 or 512 elements with integer or floating-point values). The lip feature extractor 924 can output a vector portion for each input frame of the image data 964, and / or can encode features over time using a recurrent neural network architecture (e.g., using Long Short Term Memory (LSTM) or Gated Recurrent Unit (GRU)) or a "transformer" architecture. In the latter case, the output of the lip feature extractor 924 can include one or more of the hidden states and outputs of the recurrent neural network. An example implementation of the lip feature extractor 924 is described in "Lip reading sentences in the wild" by Chung, Joon Son, et al., at the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), which is incorporated herein by reference.

[0113] exist Figure 9In this process, the speech processing module 930 receives vector portion 926 from the feature acquisition component 922 and vector portion 928 from the lip feature extractor 924 as input. In one case, the speaker preprocessing module 920 may combine vector portions 926 and 928 into a single speaker feature vector; in another case, the speech processing module 930 may receive vector portions 926 and 928 separately, but treat these vector portions as different parts of the speaker feature vector. Vector portions 926 and 928 may be combined into a single speaker feature vector by one or more of the speaker preprocessing module 920 and the speech processing module 930 using, for example, cascading or more complex attention-based mechanisms. If the sampling rates of one or more of the frames of vector part 926, vector part 928, and audio data 955 are different, a common sampling rate can be achieved by, for example, a receive and hold architecture (where values ​​that change more and more slowly are held constant at a given value until a new sample value is received), recursive time coding (e.g., using the LSTM or GRU described above), or an interest-based system (where the interest weight vector changes with the time step).

[0114] The speech processing module 930 can be configured to use vector portions 926, 928 as described in other examples illustrated herein, for example, these vector portions 926, 928 can be input into the neuroacoustic model as speaker feature vectors along with audio data 955. In the example where the speech processing module 930 includes a neuroacoustic model, the training set can be generated based on: input video from an image capture device, input audio from an audio capture device, and underlying ground truth linguistic features (e.g., ...). Figure 7 The image preprocessor 710 in the image can be used to obtain first and second sets of image data 962, 964 from the original input video.

[0115] In some examples, the vector portion 926 may also include a set of additional elements whose values ​​are, for example, obtained using a neural network architecture (e.g., Figure 5 The 522 in the vector portion 962 is derived from the encoding of the first set of image data. These additional elements can represent "facial encoding," while the vector portion 928 can represent "lip encoding." Facial encoding can remain static for utterance, while lip encoding can change during utterance or include multiple "frames" for utterance. Although Figure 9 An example using both the lip feature extractor 924 and the feature acquisition component 922 is shown; however, in one example, the feature acquisition component 922 can be omitted. In this latter example, it can be done in a manner similar to... Figure 5 The voice processing device 500 is used in a way that allows the lip-reading system for use in vehicles.

[0116] Example motor vehicle

[0117] Figure 10A and Figure 10B An example of a motor vehicle described in this article is shown. Figure 10A A side view 1000 of a vehicle 1005 is shown. The vehicle 1005 includes a control unit 1010 for controlling components of the vehicle 1005. Figure 1B Components of the voice processing device 120 shown (and in other examples) may be incorporated into the control unit 1010. In other cases, components of the voice processing device 120 may be implemented as separate units and optionally connected to the control unit 1010. The vehicle 1005 also includes at least one image capture device 1015. For example, at least one image capture device 1015 may include… Figure 1A The image capture device 110 is shown in the figure. In this example, at least one image capture device 1015 is communicatively coupled to and controlled by the control unit 1010. In other examples, at least one image capture device 1015 communicates with and is remotely controlled by the control unit 1010. In addition to the functions described herein, at least one image capture device 1015 can be used for video communication, such as Internet Protocol voice calls with video data, environmental monitoring, driver alertness monitoring, etc. Figure 10A At least one audio capture device in the form of a side-mounted microphone 1020 is also shown. These can achieve Figure 1A The audio capture device 116 shown in the figure.

[0118] The image capture device described herein may include one or more cameras or video cameras configured to capture image data frames according to command or at a predefined sampling rate. The image capture device may cover the front and rear seats of a vehicle interior. In one case, the predefined sampling rate may be less than the frame rate of full-resolution video; for example, a video stream may be captured at 30 frames per second, but the image capture device may capture the video stream at that rate or a lower rate (e.g., 1 frame per second). The image capture device may capture one or more image data frames having one or more color channels (e.g., RGB or YUV as described above). In some cases, various aspects of the image capture device (e.g., frame rate, frame size and resolution, number of color channels, and sample format) may be configurable. In some cases, the image data frames may be downsampled; for example, a video capture device capturing video at a "4K" resolution of 3840x2160 may be downsampled to 640x480 or lower. Alternatively, for low-cost embedded devices, a low-resolution image capture device capturing image data frames at a resolution of 320x240 or lower may be used. In some cases, even inexpensive, low-resolution image capture devices can provide sufficient visual information to improve speech processing. As mentioned earlier, image capture devices may also include image preprocessing and / or filtering components (e.g., contrast adjustment, noise removal, color adjustment, cropping, etc.). In some cases, low-latency and / or high-frame-rate image cameras that meet the more stringent Automotive Safety Integrity Level (ASIL) requirements of the ISO 26262 automotive safety standard are available. In addition to safety advantages, these low-latency and / or high-frame-rate image cameras can also improve lip-reading accuracy by providing more precise temporal information. This is useful for recurrent neural networks to perform more accurate feature probability estimation.

[0119] Figure 10B A top view 1030 of a vehicle 1005 is shown. The vehicle 1005 includes front seats 1032 and rear seats 1034 for maintaining passengers in an orientation suitable for voice capture by a front-mounted microphone. The vehicle 1005 includes a driver vision control console 1036 with safety-critical display information. The driver vision control console 1036 may include, for example... Figure 1A This is a portion of the dashboard 108 shown. The vehicle 1005 also includes a universal console 1038 with navigation, entertainment, and temperature control functions. The control unit 1010 can control the universal console 1038 and can implement a local voice processing module (e.g., Figure 1AThe vehicle 1005 also includes a side-mounted microphone 1020, a front-mounted overhead multi-microphone voice capture unit 1042, and a rear-mounted overhead multi-microphone voice capture unit 1044. The front and rear voice capture units 1042 and 1044 provide additional audio capture devices for capturing voice audio, eliminating noise, and identifying the speaker's location. In one case, the front and rear voice capture units 1042 and 1044 may also include additional image capture devices to capture image data describing the characteristics of each passenger in the vehicle.

[0120] exist Figure 10B In the example, any one or more of the microphone and voice capture units 1020, 1042, and 1044 can provide audio data to the audio interface (e.g., Figure 1B (140 in the original text). A microphone or microphone array can be configured to capture or record audio samples at a predefined sampling rate. In some cases, various aspects of each audio capture device (e.g., sampling rate, bit resolution, number of channels, and sampling format) can be configurable. The captured audio data can be pulse-code modulated. Any audio capture device may also include audio preprocessing and / or filtering components (e.g., contrast adjustment, noise removal, etc.). Similarly, any one or more image capture devices can provide image data to an image interface (e.g., ...). Figure 1B (150 in the example), and may also include video preprocessing and / or filtering components (e.g., contrast adjustment, noise removal, etc.).

[0121] Figure 11 An example of the interior of the car 1100 as viewed from the front seat 1032 is shown. For example, Figure 11 It can include orientation Figure 1A View of the windshield 104. Figure 11 Steering wheel 1106 is shown (e.g., Figure 1A Steering wheel 106), side microphone 1120 (e.g., Figure 10A and Figure 10B The system includes one of the side microphones 1020, a rearview mirror 1142 (which may include a front-mounted overhead multi-microphone voice capture unit 1042), and a projection device 1130. The projection device 1130 can be used to project images 1140 onto the windshield, for example, as an additional visual output device (e.g., in addition to the driver's visual console 1036 and the general console 1038). Figure 11In the image 1140, directions are included. These directions could be those projected after a voice command such as "Find me directions to the shopping mall." Other examples could use simpler response systems.

[0122] Local and remote voice processing of vehicles

[0123] In some cases, the functionality of the speech processing module as described herein can be distributed. For example, some functions can be computed locally within the vehicle 1005, and some functions can be computed by a remote (“cloud”) server device. In some cases, functions can be replicated on both the vehicle (“client”) side and the remote server device (“server”) side. In these cases, if a connection to the remote server device is unavailable, processing can be performed by the local speech processing module; if a connection to the remote server device is available, one or more of audio data, image data, and speaker feature vectors can be sent to the remote server device for parsing the captured utterances. The remote server device may have processing resources (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and random access memory), and thus can improve local performance when a connection is available. This can be traded off against latency in the processing pipeline (e.g., local processing is more responsive). In one case, the local speech processing module can provide a first output, which can be supplemented and / or enhanced by the results of the remote speech processing module.

[0124] In one scenario, a vehicle (e.g., car 1005) may be communicatively coupled to a remote server device via at least one network. The network may include one or more local area networks (LANs) and / or wide area networks (WANs) capable of being implemented using various physical technologies, such as wired technologies like Ethernet and / or wireless technologies like Wi-Fi IEEE 802.11 standards and cellular communication technologies. In some cases, the network may include a mixture of one or more private networks and public networks (e.g., the Internet). The vehicle and the remote server device may communicate over the network using different technologies and communication paths.

[0125] refer to Figure 3 In one example of the speech processing apparatus 300, vector generation by vector generator 372 can be performed locally or remotely, but the data repository 374 is located locally within the vehicle 1005. In this case, the static speaker feature vector can be computed locally and / or remotely, but is stored locally in the data repository 374. Subsequently, the speaker feature vector 325 can be obtained from the data repository 374 within the vehicle, rather than being received from a remote server device. This can improve speech processing latency.

[0126] In cases where the voice processing module is remote relative to the vehicle, the local voice processing device may include a transceiver for sending data derived from one or more of audio data, image data, and speaker feature vectors to the voice processing module and receiving control data derived from the parsing of dialogue. In one embodiment, the transceiver may include a wired or wireless physical interface and one or more communication protocols providing methods for sending and / or receiving requests in a predefined format. In another embodiment, the transceiver may include an application layer interface running on top of the Internet Protocol Suite (IPLS). In this case, the application layer interface may be configured to receive communications directed to a specific IPL address identifying a remote server device, where routing based on pathnames or web addresses is performed by one or more proxy and / or communication (e.g., "web") servers.

[0127] In some cases, linguistic features generated by the speech processing module can be mapped to voice commands and a set of data used for those commands (e.g., as referenced). Figure 4 (As described in the speech parser 436). In one case, speech data 442 may be used by the control unit 1010 of the vehicle 1005 and used to implement voice commands. In another case, speech parser 436 may be located within a remote server device, and speech parsing may involve identifying appropriate services from the output of the speech processing module to execute voice commands. For example, speech parser 436 may be configured to issue an application programming interface (API) request to the identified server, the request including commands identified from the output of the language model and any command data. For example, the speech “Where is the shopping mall?” may produce the text output “Where is the shopping mall?”, which may be mapped to a direction service API request for vehicle map data with the desired location parameter “shopping mall” and the vehicle’s current location (e.g., derived from a positioning system such as GPS). A response may be acquired and communicated to the vehicle, which may be located at the vehicle, such as... Figure 11 The area shown is displayed.

[0128] In one scenario, the remote speech parser 436 transmits response data to the control unit 1010 of the vehicle 1005. The response data may include machine-readable data transmitted to the user, for example, via a user interface or audio output. The response data can be processed, and the response to the user may be output on one or more of the driver vision console 1036 and the general console 1038. Providing a response to the user may include displaying text and / or images on the display screen of one or more of the driver vision console 1036 and the general console 1038, or outputting sound via a text-to-speech module. In some cases, the response data may include audio data, which may be processed at the control unit 1005 and used to generate audio output, for example, via one or more speakers. The response may be spoken to the user via speakers installed inside the vehicle 1005.

[0129] Example Embedded Computing System

[0130] Figure 12 An example embedded computing system 1200 is shown that can implement the speech processing apparatus described herein. Systems similar to embedded computing system 1200 can be used to implement the control unit 1010 in FIG. 10. The example embedded computing system 1200 includes one or more computer processor (CPU) cores 1210 and zero or more graphics processor (GPU) cores 1220. The processors are connected via board-level interconnect 1230 to a random access memory (RAM) device 1240 for program code and data storage. Embedded computing system 1200 also includes a network interface 1250 to allow the processors to communicate with remote systems and specific vehicle control circuitry 1260. The CPU 1210 and / or GPU 1220 can perform the functions described herein by executing instructions stored in the RAM device via interface 1230. In some cases, constrained embedded computing devices may have a similar overall component layout; however, in other cases, they may have fewer computing resources and may not have a dedicated graphics processor 1220.

[0131] Example speech processing method

[0132] Figure 13 An example method 1300 for processing speech, which can be executed to improve in-vehicle speech recognition, is shown. Method 1300 begins at box 1305, where audio data is received from an audio capture device. The audio capture device may be located inside the vehicle. The audio data may describe the characteristics of the user's utterance. Box 1305 includes data from one or more microphones (e.g.,...). Figure 10A and Figure 10BDevices 1020, 1042, and 1044 in the diagram capture data. In one case, block 1305 may include receiving audio data via a local audio interface; in another case, block 1305 may include receiving audio data via a network, for example, at an audio interface that is remote relative to the vehicle.

[0133] At box 1310, image data is received from an image capturing device. The image capturing device may be located inside the vehicle; for example, it may include… Figure 10A and Figure 10B The image capture device 1015 is included. In one case, block 1310 may include receiving image data via a local image interface; in another case, block 1310 may include receiving image data via a network, for example at an image interface that is remote relative to the vehicle.

[0134] At box 1315, a speaker feature vector is obtained based on the image data. This could, for example, include implementing any of speaker preprocessing modules 220, 320, 520, and 920. Box 1315 can be executed by a local processor of the vehicle 1005 or by a remote server device. At box 1320, a speech processing module is used to parse the speech. This could, for example, include implementing any of speech processing modules 230, 330, 400, 530, and 930. Box 1320 includes several sub-boxes. These include, at sub-box 1322, providing the speaker feature vector and audio data as input to an acoustic model of the speech processing module. This could include reference... Figure 4 The operations described are similar. In some cases, the acoustic model includes a neural network architecture. At subbox 1324, phoneme data is predicted based on the speaker feature vector and audio data, using at least a neural network architecture. This can include using a neural network architecture trained to receive the speaker feature vector as input in addition to the audio data. Because both the speaker feature vector and the audio data include numerical representations, similar processing can be performed by the neural network architecture. In some cases, existing CTC or hybrid acoustic models can be configured to receive the concatenation of the speaker feature vector and audio data, and then trained using a training set that additionally includes image data (e.g., for deriving the speaker feature vector).

[0135] In some cases, box 1315 includes performing facial recognition on image data to identify people inside the vehicle. For example, this could be as described in the reference... Figure 3 The facial recognition module 370 in the document performs as described. Thereafter, user profile data for that person can be obtained (e.g., in a vehicle) based on facial recognition. For example, user profile data can be retrieved from data repository 374 using user identifier 376, as described in reference [reference missing]. Figure 3The speaker feature vector can then be obtained from the user profile data. In one case, the speaker feature vector can be obtained from the user profile data as a set of static element values. In another case, the user profile data may indicate, for example, that one or more of the audio data and image data received at boxes 1305 and 1310 will be used to compute the speaker feature vector. In some cases, box 1315 includes comparing the number of stored speaker feature vectors associated with the user profile data with a predefined threshold. For example, the user profile data may indicate how many previous voice queries a user identified using facial recognition has performed. In response to the number of stored speaker feature vectors being less than the predefined threshold, one or more of the audio data and image data may be used to compute the speaker feature vector. In response to the number of stored speaker feature vectors being greater than the predefined threshold, static speaker feature vectors, such as those stored within or accessible through the user profile data, can be obtained. In this case, the number of stored speaker feature vectors can be used to generate the static speaker feature vector.

[0136] In some examples, box 1315 may include processing image data to generate one or more speaker feature vectors based on lip movements within a human facial region. For example, a lip reading module, such as lip feature extractor 924 or a suitably configured neural speaker preprocessing module 520, may be used. The output of the lip reading module may be used to provide one or more speaker feature vectors to the speech processing module, and / or may be combined with other values ​​(e.g., i-vectors or x-vectors) to generate larger speaker feature vectors.

[0137] In some examples, box 1320 includes providing phoneme data to a language model of a speech processing module, using the language model to predict the transcription of utterances, and using the transcription to determine control commands for the vehicle. For example, box 1320 may include references to... Figure 4 The operations described are similar to those described.

[0138] Example of discourse parsing

[0139] Figure 14 An example processing system 1400 is shown, including a non-transitory computer-readable storage medium 1410 storing instructions 1420, which, when executed by at least one processor 1430, cause at least one processor to perform a series of operations. The operations of this example use the previously described methods to generate transcriptions of speech. These operations can be performed within a vehicle (e.g., as previously described), or the vehicle-mounted example can be extended to non-vehicle-based situations (e.g., it can be implemented using desktop, laptop, mobile, or server computing devices, etc.).

[0140] via instruction 1432, processor 1430 is configured to receive audio data from an audio capture device. This may include accessing local memory containing the audio data and / or receiving a data stream or a set of array values ​​over a network. The audio data may have a form described with reference to other examples herein. Via instruction 1434, processor 1430 is configured to receive a speaker feature vector. The speaker feature vector is obtained based on image data from an image capture device that describes features of a user's facial regions. For example, the speaker feature vector may use a reference... Figure 2 , Figure 3 , Figure 5 and Figure 9 The speaker feature vector can be obtained using any of the described methods. The speaker feature vector can be computed locally by processor 1430, accessed from local memory, and / or received via a network interface (etc.). Via instruction 1436, processor 1430 is instructed to use the speech processing module to parse the speech. The speech processing module may include references... Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 9 Any module described by any of them.

[0141] Figure 14 It is shown that instruction 1436 can be broken down into many other instructions. Via instruction 1440, processor 1430 is instructed to provide the speaker feature vector and audio data as input to the acoustic model of the speech processing module. This can be used with reference to... Figure 4 The described approach is implemented in a similar manner. In this example, the acoustic model includes a neural network architecture. Via instruction 1442, processor 1430 is instructed to use at least the neural network architecture to predict phoneme data based on speaker feature vectors and audio data. Via instruction 1444, processor 1430 is instructed to provide the phoneme data to the language model of the speech processing module. This can also be used with... Figure 4 The process is executed in a similar manner to that shown. Via instruction 1446, processor 1430 is instructed to use a language model to generate a transcription of the speech. For example, the transcription can be generated as the output of the language model. In some cases, the transcription can be used by a control system to execute voice commands, such as control unit 1010 in car 1005. In other cases, the transcription may include the output of a speech-to-text system. In the latter case, image data can be acquired from a webcam or similar device communicatively coupled to a computing device including processor 1430. For mobile computing devices, image data can be acquired from a forward-facing image capture device.

[0142] In some examples, the speaker feature vector received according to instruction 1434 includes one or more of the following: speaker-related vector elements generated based on audio data (e.g., i-vector or x-vector components); vector elements related to the speaker's lip movements generated based on image data (e.g., as generated by a lip-reading module); and vector elements related to the speaker's face generated based on image data. In one case, processor 1430 may include part of a remote server device, and the audio data and speaker feature vector may be received from a motor vehicle, for example, as part of a distributed processing pipeline.

[0143] Example implementation

[0144] Examples related to speech processing, including automatic speech recognition, are described. Some examples involve processing certain spoken languages. Various examples similarly operate on other languages ​​or combinations of languages. Some examples improve the accuracy and robustness of speech processing by incorporating additional information derived from an image of the person uttering the utterance. This additional information can be used to improve linguistic models. Linguistic models may include one or more of acoustic models, articulation models, and language models.

[0145] Some of the examples described in this paper can be implemented to address the unique challenges of performing automatic speech recognition within vehicles such as automobiles. In some combined examples, image data from a camera can be used to determine lip-reading features and recognize faces, enabling the creation and selection of i-vector and / or x-vector profiles. By implementing the methods described in this paper, automatic speech recognition can be performed in the noisy, multi-channel environment of a motor vehicle.

[0146] Some of the examples described in this paper can improve the efficiency of speech processing by including one or more features (e.g., lip location or movement) derived from image data in a speaker feature vector, which is provided as input to an acoustic model that also receives audio data as input (a single model), for example, instead of having an acoustic model that only receives audio input or separate acoustic models for audio and image data.

[0147] Certain methods and sets of operations can be executed by instructions stored on a non-transitory computer-readable medium. The non-transitory computer-readable medium stores code including instructions that, when executed by one or more computers, cause the computers to perform the steps of the methods described herein. The non-transitory computer-readable medium may include one or more of a spinning disk, a spinning optical disk, a flash memory random access memory (RAM) chip, and other mechanically movable or solid-state storage media. Any type of computer-readable medium is suitable for storing code including instructions according to various examples.

[0148] Some of the examples described herein can be implemented as so-called System-on-Chip (SoC) devices. SoC devices control many embedded automotive systems and can be used to implement the functions described herein. In one case, one or more of the speaker preprocessing module and the speech processing module can be implemented as an SoC device. An SoC device may include one or more processors (e.g., CPU or GPU), random access memory (RAM, e.g., off-chip dynamic RAM or DRAM), and network interfaces for wired or wireless connectivity (e.g., Ethernet, WiFi, 3G, 4G LTE, 5G, and other wireless interface standard radios). SoC devices may also include various I / O interface devices, such as touchscreen sensors, geolocation receivers, microphones, speakers, Bluetooth peripherals, and USB devices (e.g., keyboards and mice), depending on the different peripherals required. By executing instructions stored in the RAM device, the processor of the SoC device can perform the steps of the methods described herein.

[0149] This paper has described some examples, and it will be noted that different combinations of different components from different examples are possible. Significant features have been presented to better explain the examples; however, it is obvious that certain features can be added, modified, and / or omitted without altering the functional aspects of the examples described.

[0150] Various examples are methods that utilize human, machine, or a combination of both actions. These method examples are complete wherever the constituent steps occur in most parts of the world. Some examples are one or more non-transitory computer-readable media arranged to store such instructions for the methods described herein. Examples can be implemented on any machine storing a non-transitory computer-readable medium including any necessary code. Some examples can be implemented as: physical devices such as semiconductor chips; hardware description language representations of the logical or functional behavior of such devices; and one or more non-transitory computer-readable media arranged to store such hardware description language representations. The descriptions of principles, aspects, and embodiments herein cover their structural and functional equivalents. The elements described herein as coupled have effective relationships that can be achieved through direct connection or indirectly using one or more other intermediate elements.

Claims

1. A vehicle-mounted device, comprising: An audio interface is configured to receive audio data from an audio capture device in the vehicle, the audio data describing the characteristics of speech of a person inside the vehicle; An image interface is configured to receive image data from an image capture device for capturing images from the vehicle, the image data describing features of the facial region of the person; The speech processing module is configured to parse human speech based on the audio data and the image data; as well as A speaker preprocessing module is configured to receive the image data and obtain a speaker feature vector based on the image data to predict phoneme data, wherein obtaining the speaker feature vector includes: Perform facial recognition on the image data to identify the person inside the vehicle; Based on the facial recognition, user profile data for the person is obtained; and The speaker feature vector is obtained based on the user profile data; The process of obtaining the speaker feature vector based on the user profile data includes: The number of stored speaker feature vectors associated with the user profile data is compared with a predefined threshold; In response to the number of stored speaker feature vectors falling below a predefined threshold, the speaker feature vectors are computed using one or more of the audio data and the image data; and In response to the number of stored speaker feature vectors being greater than the predefined threshold, a static speaker feature vector associated with the user profile data is obtained, the static speaker feature vector being generated using the number of stored speaker feature vectors.

2. The vehicle-mounted device according to claim 1, wherein, The speech processing module includes an acoustic model configured to process the audio data and predict the phoneme data for parsing the speech.

3. The vehicle-mounted device according to claim 2, wherein, The acoustic model includes a neural network architecture.

4. The vehicle-mounted device according to claim 2 or 3, wherein, The acoustic model is configured to receive the speaker feature vector and the audio data as input, and is trained to use the speaker feature vector and the audio data to predict the phoneme data.

5. The vehicle-mounted device according to any one of claims 1 to 3, wherein, The speaker preprocessing module includes: The lip-reading module, implemented by the processor, is configured to generate one or more speaker feature vectors based on lip movements within the person's facial region.

6. The vehicle-mounted device according to any one of claims 1 to 3, wherein, The speaker preprocessing module includes a neural network architecture configured to receive data derived from one or more of the audio data and the image data and predict the speaker feature vector.

7. The vehicle-mounted device according to claim 1, wherein, The static speaker feature vector is stored in the vehicle's memory.

8. The vehicle-mounted device according to any one of claims 1 to 3, comprising: An image capture device is configured to capture electromagnetic radiation having an infrared wavelength, and the image capture device is configured to send the image data to the image interface.

9. The vehicle-mounted device according to any one of claims 1 to 3, wherein, The speaker preprocessing module is configured to process the image data to extract one or more portions of the image data. One or more of the extracted portions are used to obtain the speaker feature vector.

10. The vehicle-mounted device according to claim 5, wherein, The processor is located inside the vehicle.

11. The vehicle-mounted device according to any one of claims 1 to 3, wherein, The voice processing module is remote relative to the vehicle, and the device includes: A transceiver is used to send data derived from the audio data and the image data to the speech processing module, and to receive control data obtained from the parsing of the speech.

12. The vehicle-mounted device according to claim 3, wherein, The acoustic model includes a hybrid acoustic model, which comprises the neural network architecture and a Gaussian mixture model, wherein the Gaussian mixture model is configured to receive a vector of category probabilities output by the neural network architecture and output phoneme data for parsing the utterance.

13. The vehicle-mounted device according to claim 2 or 3, wherein, The acoustic model includes a connection-time classification (CTC) model.

14. The vehicle-mounted device according to claim 2 or 3, wherein, The voice processing module includes: A language model, communicatively coupled to the acoustic model, is used to receive the phoneme data and generate a transcription representing the utterance.

15. The vehicle-mounted device according to claim 14, wherein, The language model is configured to use the speaker feature vector to generate a transcription representing the utterance.

16. The vehicle-mounted device according to claim 2 or 3, wherein, The acoustic model includes: Database for acoustic model configuration; An acoustic model selector is used to select an acoustic model configuration from the database based on the speaker feature vector; and An acoustic model instance for processing the audio data, the acoustic model instance being instantiated based on an acoustic model configuration selected by the acoustic model selector, the acoustic model instance being configured to generate the phoneme data for parsing the utterance.

17. The vehicle-mounted device according to any one of claims 1 to 3, wherein, The speaker feature vector is one or more of the i-vector and the x-vector.

18. The vehicle-mounted device according to any one of claims 1 to 3, wherein, The speaker feature vector includes: A first portion related to the person, generated based on the audio data; and A second part related to the person's lip movements, generated based on the image data.

19. The vehicle-mounted device according to claim 18, wherein, The speaker feature vector includes: A third part related to the person's face, generated based on the image data.

20. A method for processing discourse, comprising: Audio data is received from an audio capture device located inside the vehicle, the audio data describing the characteristics of speech spoken by a person inside the vehicle; Image data is received from an image capture device used to capture images inside the vehicle, the image data describing features of the person's facial region; The speaker feature vector is obtained based on the image data; as well as The speech is parsed using a speech processing module implemented by a processor, including: The speaker feature vector and the audio data are provided as input to the acoustic model of the speech processing module. The acoustic model includes a neural network architecture. At least using the aforementioned neural network architecture, phoneme data is predicted based on the speaker feature vector and the audio data. The process of obtaining the speaker feature vector includes: Perform facial recognition on the image data to identify the person inside the vehicle; Based on the facial recognition, user profile data for the person is obtained; and The speaker feature vector is obtained based on the user profile data; The process of obtaining the speaker feature vector based on the user profile data includes: The number of stored speaker feature vectors associated with the user profile data is compared with a predefined threshold; In response to the number of stored speaker feature vectors falling below a predefined threshold, the speaker feature vectors are computed using one or more of the audio data and the image data; and In response to the number of stored speaker feature vectors being greater than the predefined threshold, a static speaker feature vector associated with the user profile data is obtained, the static speaker feature vector being generated using the number of stored speaker feature vectors.

21. The method according to claim 20, wherein, The speaker feature vector is obtained by: The image data is processed to generate one or more speaker feature vectors based on lip movements within the person's facial region.

22. The method according to claim 20, wherein, The parsing of the utterances includes: The phoneme data is provided to the language model of the speech processing module; The language model is used to predict the transcription of the discourse; and The transcription is used to determine the control commands for the vehicle.

23. A computer-readable storage medium including instructions that, when executed by at least one processor, cause the at least one processor to perform the following operations: Audio data is received from the vehicle's audio capture device, the audio data describing the characteristics of human speech within the vehicle; Image data is received from an image capture device used to capture images inside the vehicle, the image data describing features of the person's facial region; The speaker feature vector is obtained based on the image data; as well as Use a speech processing module to parse speech, including: The speaker feature vector and the audio data are provided as input to the acoustic model of the speech processing module, and the acoustic model includes a neural network architecture. At least using the aforementioned neural network architecture, phoneme data is predicted based on the speaker feature vector and the audio data. The phoneme data is provided to the language model of the speech processing module, and The language model is used to generate a transcription of the discourse. The process of obtaining the speaker feature vector includes: Perform facial recognition on the image data to identify the person inside the vehicle; Based on the facial recognition, user profile data for the person is obtained; and The speaker feature vector is obtained based on the user profile data; The process of obtaining the speaker feature vector based on the user profile data includes: The number of stored speaker feature vectors associated with the user profile data is compared with a predefined threshold; In response to the number of stored speaker feature vectors falling below a predefined threshold, the speaker feature vectors are computed using one or more of the audio data and the image data; and In response to the number of stored speaker feature vectors being greater than the predefined threshold, a static speaker feature vector associated with the user profile data is obtained, the static speaker feature vector being generated using the number of stored speaker feature vectors.

24. The computer-readable storage medium according to claim 23, wherein, The speaker feature vector includes one or more of the following: Vector elements related to the user generated based on the audio data; Vector elements related to the user's lip movements generated based on the image data; as well as Vector elements related to the user's face generated based on the image data.

Citation Information

Patent Citations

  • Methods using recurrence quantification analysis to analyze and generate images

    US8422820B2

  • Combined lip reading and voice recognition multimodal interface system

    US8442820B2

  • Voice recognition control system

    JP2017090612A

  • System and method for multi-modal focus detection, referential ambiguity resolution and mood classification using multi-modal input

    US6964023B2