Computing devices
A system combining audio and image data using neural networks and facial recognition improves speech processing in vehicles, addressing resource and noise challenges to enhance accuracy and reliability.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SOUNDHOUND INC
- Filing Date
- 2024-07-18
- Publication Date
- 2026-06-22
Smart Images

Figure 0007877398000001 
Figure 0007877398000002 
Figure 0007877398000003
Abstract
Description
Technical Field
[0001] The technical field of the present invention This technology is in the field of audio processing. One example relates to processing audio captured from inside a vehicle.
Background Art
[0002] Background Recent advances in computing have increased the likelihood of realizing many long-desired speech control applications. For example, improvements in statistical models, including practical frameworks for effective neural network architectures, have significantly increased the accuracy and reliability of previous audio processing systems. This combined with the boom in wide area computer networks provides a range of modular services that can simply be accessed using an application programming interface. Speech is rapidly becoming a viable option for providing a user interface.
[0003] Speech control devices are common in the home, but additional difficulties are presented in providing audio processing within a vehicle. For example, vehicles often have limited processing resources for auxiliary functions such as a speech interface, are subject to significant noise (such as high-level road noise and / or engine noise), and present constraints in the acoustic environment. Any user interface is further constrained by the impact on the safety of controlling the vehicle. These factors have made it difficult to actually achieve speech control within a vehicle.
[0004] Furthermore, despite advances in speech processing, even users of advanced computing devices often report that current systems lack human-level responsiveness and intelligence. Converting air pressure fluctuations into analyzed commands is incredibly difficult. Speech processing typically involves complex processing pipelines, and errors at any stage can undermine the success of machine interpretation. Many of these difficulties are not immediately apparent to humans, who can process speech using cortical or subcortical structures without conscious thought. However, engineers working in this field are rapidly becoming aware of the gap between human capabilities and current speech processing technology.
[0005] US8,442,820B2 (Patent Document 1) describes a lip-reading and speech recognition combined multimodal interface system. This system can issue navigation operation commands solely by voice and lip movements, thereby allowing the driver to look ahead during navigation operations and reducing vehicle accidents related to navigation operations while driving. The lip-reading and speech recognition combined multimodal interface system described in US8,442,820B2 includes an audio voice input unit, a speech recognition unit, a speech recognition command and estimated probability output unit, a lip video image input unit, a lip-reading unit, a lip-reading command output unit, and a speech recognition and lip-reading result combined unit that outputs speech recognition commands. While US8,442,820B2 offers one solution for in-vehicle control, the proposed system is complex, and the numerous interacting components increase the chances of errors and analysis failures.
[0006] A speech processing system that more accurately transcribes and analyzes human speech. It is desirable to provide methods for this. Furthermore, it is desirable to provide voice processing methods that can actually be implemented by real-world devices such as embedded computing systems for vehicles. Realizing practical voice processing solutions is difficult because there are many difficulties in vehicle system integration and connectivity. [Prior art documents] [Patent Documents]
[0007] [Patent Document 1] U.S. Patent Specification No. 8,422,820 [Overview of the Initiative] [Means for solving the problem]
[0008] Summary of the Invention Examples described in this specification provide methods and systems for processing speech. Some examples utilize both audio and image data to process speech. Some examples are adapted to address the challenges of processing speech captured within a vehicle. One example obtains speaker feature vectors based on image data that features at least the facial area of a person, such as a person in a vehicle. Speech processing is then performed using visually derived information dependent on the speaker of the speech. This can improve accuracy and robustness.
[0009] In one aspect, the device for the vehicle includes an audio interface configured to receive audio data from an audio capture device located inside the vehicle, an image interface configured to receive image data featuring the facial area of a person inside the vehicle from an image capture device in order to capture image data inside the vehicle, and a speech processing module configured to analyze human speech based on the audio and image data. The speech processing module includes an acoustic model configured to process the audio data and predict phoneme data used to analyze speech, and the acoustic model includes a neural network architecture. The device further includes a speaker preprocessing module implemented by a processor, which is configured to receive image data and obtain speaker feature vectors based on the image data, and the acoustic model is configured to receive speaker feature vectors and audio data as input and is trained to use speaker feature vectors and audio data to predict phoneme data.
[0010] In the above scenario, the speaker feature vector is obtained using image data that features the facial area of the person speaking. This speaker feature vector is provided as input to the neural network architecture of the acoustic model, which is configured to use this input as well as audio data that features utterances. This provides the acoustic model with additional visually derived information that the neural network architecture can use to improve the analysis of utterances, for example, to compensate for undesirable acoustic and noise features in the vehicle. For example, by configuring the acoustic model based on a specific person and / or their mouth area determined from the image data, the determination of ambiguous phonemes that might otherwise be mistranscribed based on the vehicle's context can be improved.
[0011] In one variation, the speaker preprocessing module is configured to perform facial recognition on image data to identify people in the vehicle and extract speaker feature vectors associated with the identified people. For example, the speaker preprocessing module may include a facial recognition module used to identify users speaking in the vehicle. If the speaker feature vectors are determined based on audio data, person identification may allow a predetermined (e.g., pre-calculated) speaker feature vector to be extracted from memory. This limits the Processing latency for a certain embedded vehicle control system can be improved.
[0012] In one variation, the speaker preprocessing module includes a processor-implemented lip-reading module, configured to generate one or more speaker feature vectors based on the movement of the lips within a person's face area. This may be used in conjunction with or independently of a face recognition module. In this case, the one or more speaker feature vectors provide a representation of the speaker's mouth or lip area that can be used by a neural network architecture of the acoustic model, thereby improving processing.
[0013] In some cases, the speaker preprocessing module may include a neural network architecture configured to receive data from one or more audio and image data sources and predict speaker feature vectors. For example, this approach could combine a visual-based neural lip-reading system with an acoustic "x-vector" system to improve acoustic processing. If one or more neural network architectures are used, they may be trained using a training set that includes image data, audio data, and a ground truth set of linguistic features, such as a ground truth set of phoneme data and / or text transcription.
[0014] In some cases, the speaker preprocessing module is configured to compute speaker feature vectors for a predetermined number of utterances and to compute a static speaker feature vector based on multiple speaker feature vectors for a predetermined number of utterances. For example, the static speaker feature vector may include the average of a set of speaker feature vectors linked to a particular user using image data. The static speaker feature vector may be stored in the vehicle's memory. This can also improve speech processing capabilities within a resource-constrained vehicle computing system.
[0015] In one case, the device includes memory configured to store one or more user profiles. In this case, the speaker preprocessing module is configured to perform facial recognition on image data to identify user profiles in memory associated with people in the vehicle, to compute speaker feature vectors for the people, to store the speaker feature vectors in memory, and to associate the stored speaker feature vectors with the identified user profiles. Facial recognition can provide a quick and easy mechanism for extracting useful information (e.g., speaker feature vectors) for acoustic processing that is specific to a particular person. In one case, the speaker preprocessing module may be configured to determine whether the number of stored speaker feature vectors associated with a given user profile is greater than a predetermined threshold. If this is the case, the speaker preprocessing module may compute a static speaker feature vector based on the number of stored speaker feature vectors, store the static speaker feature vector in memory, associate the stored static speaker feature vector with a given user profile, and indicate that the static speaker feature vector should be used for future speech analysis instead of computing a speaker feature vector for the person.
[0016] In one variation, the apparatus includes an image capture device configured to capture electromagnetic radiation having infrared wavelengths, and the image capture device is configured to send image data to an image interface. This may provide illumination-invariant images that enhance image data processing. A speaker preprocessing module may be configured to process the image data to extract one or more portions of the image data, and the extracted one or more portions are used to obtain speaker feature vectors. For example, one or more portions may relate to the face area and / or mouth area.
[0017] In one case, one or more of the audio interface, image interface, speech processing module, and speaker pre-processing module may be located within the vehicle, for example, including part of a local embedded system. In this case, the processor may be located within the vehicle. In another case, the speech processing module may be remote from the vehicle. In this case, the device may include transceivers that transmit data derived from audio and image data to the speech processing module and receive control data from speech analysis. Different distributed configurations are possible. For example, in one case, the device may be implemented locally within the vehicle, but yet another copy of at least one component of the device may run on a remote server device. In this case, certain functions may be performed remotely, for example, together with or instead of local processing. The remote server device may have increased processing resources to improve accuracy, but may increase processing latency.
[0018] In one case, the acoustic model includes a hybrid acoustic model comprising a neural network architecture and a Gaussian mixture model, where the Gaussian mixture model is configured to receive a vector of class probabilities output by the neural network architecture and output phoneme data for analyzing utterances. Additionally or alternatively, the acoustic model may include, for example, a Hidden Markov Model (HMM) along with the neural network architecture. In one case, the acoustic model may include a connectionist temporal classification (CTC) model or another form of a neural network model having a recurrent neural network architecture.
[0019] In one variation, the speech processing module includes a language model that is communicatively coupled to an acoustic model to receive phoneme data and generate transcriptions representing utterances. In this variation, the language model may be configured to use, for example, speaker feature vectors to generate transcriptions representing utterances, in addition to the acoustic model. This can be used to improve the accuracy of the language model if the language model includes a neural network architecture such as a recurrent neural network or a transformer architecture.
[0020] In one variation, the acoustic model includes a database of acoustic model configurations, an acoustic model selector that selects an acoustic model configuration from the database based on speaker feature vectors, and acoustic model instances that process audio data, the acoustic model instances being instantiated based on the acoustic model configuration selected by the acoustic model selector, and the acoustic model instances being configured to generate phoneme data used to analyze speech.
[0021] In one example, the speaker feature vector is one or more of the i and x vectors. The speaker feature vector may include a composite vector, which may include, for example, two or more of the following: a first part that depends on the speaker and is generated based on audio data; a second part that depends on the speaker's lip movements and is generated based on image data; and a third part that depends on the speaker's face and is generated based on image data.
[0022] In another aspect, there exists a method for processing speech, which includes receiving audio data featuring the speech of a person inside the vehicle from an audio capture device located inside the vehicle, receiving image data featuring the face area of the person from an image capture device located inside the vehicle, obtaining a speaker feature vector based on the image data, and analyzing the speech using a speech processing module implemented by a processor. Analyzing the speech involves inputting the speaker feature vector as input to the acoustic model of the speech processing module. The method includes providing voice and audio data, the acoustic model includes a neural network architecture, and the analysis of speech further includes predicting phoneme data using at least a neural network architecture based on speaker feature vectors and audio data.
[0023] The above method may provide similar improvements in voice processing within a vehicle. In some cases, obtaining a speaker feature vector includes performing face recognition on image data to identify a person within the vehicle, obtaining user profile data for the person based on the face recognition, and obtaining a speaker feature vector according to the user profile data. The above method may further include comparing the number of stored speaker feature vectors associated with the user profile data with a predefined threshold. In response to the number of stored speaker feature vectors falling below the predefined threshold, the above method may include calculating a speaker feature vector using one or more of the audio data and the image data. In response to the number of stored speaker feature vectors being greater than the predefined threshold, the above method may include obtaining a static speaker feature vector associated with the user profile data, the static speaker feature vector being generated using the number of stored speaker feature vectors. In one case, the speaker feature vector includes processing image data to generate one or more speaker feature vectors based on the movement of the lips within the face area of the person. Analyzing the utterance may include providing phoneme data to a language model of the voice processing module, predicting a transcript of the utterance using the language model, and determining a control command for the vehicle using the transcript.
[0024] According to another aspect, there is a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to receive audio data from an audio capture device, receive a speaker feature vector obtained based on image data from an image capture device, the image data characterizing a face area of a user, and further cause the at least one processor to analyze an utterance using an audio processing module, which includes providing the speaker feature vector and the audio data as inputs to an acoustic model of the audio processing module, the acoustic model including a neural network architecture, and further predicting phoneme data using at least the neural network architecture based on the speaker feature vector and the audio data, providing the phoneme data to a language model of the audio processing module, and generating a transcript of the utterance using the language model.
[0025] The at least one processor can include, for example, a computing device such as a motor vehicle from which the audio data and the speaker image vector are received remotely. By the instructions, the processor may be enabled to perform automatic speech recognition with a lower error rate. In some cases, the speaker feature vector includes one or more of a speaker-dependent vector element generated based on the audio data, a vector element dependent on the lip movement of the speaker generated based on the image data, and a vector element dependent on the face of the speaker generated based on the image data.
Brief Description of the Drawings
[0026] [Figure 1A] A schematic diagram showing the interior of a vehicle according to an example. [Figure 1B] A schematic diagram showing a device for a vehicle according to an example. [Figure 2] This is a schematic diagram showing a device for a vehicle having a speaker pre-processing module according to a certain example. [Figure 3] This is a schematic diagram showing the components of a speaker preprocessing module according to a certain example. [Figure 4] This is a schematic diagram showing the components of an audio processing module according to a certain example. [Figure 5] This is a schematic diagram showing a neural speaker preprocessing module and a neural speech processing module following a certain example. [Figure 6] This is a schematic diagram showing the components that make up the acoustic model of a speech processing module, following a certain example. [Figure 7] This is a schematic diagram showing an image preprocessor following a certain example. [Figure 8] This is a schematic diagram showing image data from different image capture devices following a certain example. [Figure 9] This is a schematic diagram showing the components of a speaker preprocessing module configured to extract lip features according to a given example. [Figure 10A] This is a schematic diagram showing an automated vehicle equipped with a device for voice processing according to a certain example. [Figure 10B] This is a schematic diagram showing an automated vehicle equipped with a device for voice processing according to a certain example. [Figure 11] This is a schematic diagram showing the components of a user interface for an autonomous vehicle, following a specific example. [Figure 12] This is a schematic diagram showing an exemplary computing device for a vehicle. [Figure 13] This is a flowchart illustrating how to process vocalizations according to a specific example. [Figure 14] This is a schematic diagram illustrating a non-temporary computer-readable storage medium according to a certain example. [Modes for carrying out the invention]
[0027] Detailed explanation The following describes various examples of this technique that demonstrate a variety of interesting situations. In general, the examples can be used in any combination of the situations described.
[0028] One example described in this specification uses visual information to improve speech processing. This visual information may be obtained from inside a vehicle. In the example, the visual information features people inside the vehicle, such as the driver or passengers. Another example uses visual information to generate speaker feature vectors for use by a adapted speech processing module. The speech processing module may be configured to use the speaker feature vectors to improve the processing of associated audio data, such as audio data from an audio capture device inside the vehicle. The example may improve the responsiveness and accuracy of an in-vehicle speech interface. Another example may be used by a computing device to improve speech transcription. Thus, the described examples may appear to extend speech processing systems with a multimodal capability to improve the accuracy and reliability of speech processing.
[0029] Some examples described in this specification provide different approaches to generating speaker feature vectors. These approaches are complementary and can be used together to synergistically improve speech processing. In one example, image data obtained from inside a vehicle, for example from driver and / or passenger cameras, is processed to identify a person and determine feature vectors that numerically represent certain features of that person. These features may include audio features, such as numerical representations of predicted variations in audio data for an acoustic model. In another example, image data obtained from inside a vehicle, for example from driver and / or passenger cameras, is processed to determine feature vectors that numerically represent certain visual features of a person, such as features associated with speech utterances made by that person. In one case, the visual features are associated with the mouth area of the person and represent, for example, the position and / or movement of the lips. In both examples, the speaker feature vectors may have a similar format, so that phoneme data can be processed. It can be easily integrated into the input pipeline of the acoustic model used for generation. One example could provide an improvement in overcoming some of the challenges of automated speech recognition inside vehicles, such as the confined interior of a vehicle, the possibility of multiple people speaking in this limited space, and high levels of engine and environmental noise.
[0030] Context of the example vehicle Figure 1A provides an exemplary context for an audio processing device. In Figure 1A, the context is an automated vehicle. Figure 1A is a schematic diagram of the interior 100 of the automated vehicle. The interior 100 is shown with respect to the driver's side of the front of the automated vehicle. A person 102 is shown to be present inside the interior 100. In Figure 1A, the person is the driver of the automated vehicle. The driver is facing forward in the vehicle and observing the road through the windshield 104. The person controls the vehicle using the steering wheel 106 and observes vehicle status indicators through the dashboard or instrument panel 108. In Figure 1A, an image capture device 110 is located within the interior 100 of the automated vehicle near the bottom of the dashboard 108. The image capture device 110 has a field of view 112 that captures the face area 114 of the person 102. In this example, the image capture device 110 is positioned to capture an image through the aperture of the steering wheel 106. Figure 1A further shows an audio capture device 116 located within the interior 100 of the automated vehicle. The audio capture device 116 is positioned to capture sounds emitted by a person 102. For example, the audio capture device 116 may be positioned to capture speech from person 102, i.e., sounds emitted from the person's face area 114. The audio capture device 116 is shown mounted on the windshield 104. For example, the audio capture device 116 may be mounted near or on the rearview mirror, or on the door frame on the side of person 102. Figure 1A further shows a voice processing device 120. The voice processing device 120 may be mounted on the automated vehicle. The voice processing device 120 may include or form part of a control system for the automated vehicle. In the example in Figure 1A, the image capture device 110 and the audio capture device 116 are communicably coupled to the voice processing device 120, for example, via one or more wired and / or wireless interfaces. The image capture device 110 may be located outside the automated vehicle to capture images inside the automated vehicle through the windows of the automated vehicle.
[0031] The context and configuration of Figure 1A are provided as an example to aid in understanding the following description. This example is not limited to automated vehicles and may be similarly implemented in other forms of vehicles. Other forms of vehicles include, but are not limited to, marine vehicles such as boats and ships, aerial vehicles such as helicopters, airplanes and gliders, rail vehicles such as trains and trams, spacecraft, construction vehicles and heavy equipment. Automated vehicles may include, for example, cars, trucks, sports utility vehicles, motorbikes, buses and motorized carts. The use of the term “vehicle” in this specification further includes certain heavy equipment that can be motor-driven while stationary, such as cranes, lifting equipment and boring equipment. Vehicles may be manually controlled and / or have autonomous functions. The example in Figure 1A features a steering wheel 106 and a dashboard 108, but other control configurations may be provided (for example, an autonomous vehicle may not have a steering wheel 106 as shown). While the context of the driver's seat is shown in Figure 1A, similar configurations may be provided for one or more passenger seats (e.g., both front and rear). Figure 1A is provided for illustrative purposes only and omits certain features that may be present in an automated vehicle for clarity. In some cases, the approach described herein may be used outside the context of a vehicle and may be implemented by computing devices such as a desktop or laptop computer, a smartphone, or an embedded device.
[0032] Figure 1B is a schematic diagram of the speech processing device 120 shown in Figure 1A. In Figure 1B, the speech processing device 120 includes a speech processing module 130, an image interface 140, and an audio interface 150. The image interface 140 is configured to receive image data 145. The image data may include image data captured by the image capture device 110 in Figure 1A. The audio interface 150 is configured to receive audio data 155. The audio data 155 may include audio data captured by the audio capture device 116 in Figure 1A. The speech processing module 130 is communicatively coupled to both the image interface 140 and the audio interface 150. The speech processing module 130 is configured to process the image data 145 and audio data 155 to generate a set of linguistic features 160 that can be used to analyze the speech of a person 102. Linguistic features may include phonemes, word parts (e.g., stems or proto-words), and words (including text features such as pauses that map to punctuation), as well as probabilities and other values relating to these linguistic units. In one case, linguistic features may be used to generate text output representing utterances. In this case, the text output may be used as is, or it may be mapped to a predefined set of commands and / or command data. In another case, linguistic features may be directly mapped to a predefined set of commands and / or command data (e.g., without explicit text output).
[0033] A person (like person 102) may issue voice commands while operating an automated vehicle using the configurations in Figures 1A and 1B. For example, person 102 may speak internally, for example, generate utterances, to control the automated vehicle or obtain information. In this context, utterances are associated with vocal sounds produced by a person that represent linguistic information, such as speech. For example, utterances may include sounds coming from person 102's larynx. Utterances may include voice commands, such as requests spoken by the user. For example, voice commands may include requests to perform an action (e.g., "play music," "turn on the air conditioning," "activate cruise control"), further information related to the request (e.g., "album XY," "68 degrees Fahrenheit," "60 mph for 30 minutes"), speech to be transcribed (e.g., "add ... to my to-do list" or "send the following message to user A"), and / or requests for information (e.g., "what is the traffic volume on C?", "what is the weather like today?", or "where is the nearest gas station?").
[0034] The audio data 155 can take various forms depending on the implementation. Generally, the audio data 155 may originate from time-series measurements from one or more audio capture devices (e.g., one or more microphones) such as the audio capture device 116 in Figure 1A. In some cases, the audio data 155 may be captured from a single audio capture device. In other cases, the audio data 155 may be captured from multiple audio capture devices, for example, multiple microphones may be located at different positions within the internal 100. In the latter case, the audio data may include one or more channels of time-correlated audio data from each audio capture device. At the time of capture, the audio data may include, for example, one or more channels of pulse code modulation (PCM) data at a predetermined sampling rate (e.g., 16 kHz), where each sample is represented by a predetermined number of bits (e.g., 8 bits, 16 bits, or 24 bits per sample, if each sample contains an integer or floating-point value).
[0035] In some cases, the audio data 155 may be processed after capture but before reception at the audio interface 150 (e.g., pre-processed with respect to audio processing). Processing may include one or more filtering in one or more of the time and frequency domains, and noise reduction and / or normalization may be applied. In one case, the audio data may be converted into time-span measurements in the frequency domain by performing a Fast Fourier Transform, for example, to produce one or more frames of spectrogram data. In some cases, a Mel filter bank or Mel frequency may be used to determine values for features in one or more frequency domains. Filter banks, such as the Mel-Frequency Cepstral Coefficient, may be applied. In these cases, the audio data 155 may contain the outputs of one or more filter banks. In other cases, the audio data 155 may contain time-domain samples that may be pre-processed within the speech processing module 130. Different combinations of approaches are possible. Therefore, the audio data received by the audio interface 150 may contain any measurements made along the speech processing pipeline.
[0036] Similar to audio data, the image data described in this specification may take various forms depending on the implementation. In one case, the image capture device 110 may include a video capture device, and the image data may include one or more frames of video data. In another case, the image capture device 110 may include a still image capture device, and the image data may include one or more frames of still images. Thus, the image data may originate from both video and still image sources. References to image data in this specification may, for example, refer to image data derived from a two-dimensional array having height and width (e.g., rows and columns of an array). In one case, the image data may have multiple color channels, for example, three color channels for each of the red, green, and blue (RGB) colors, and each color channel has a two-dimensional array to which color values are associated (e.g., 8 bits, 16 bits, or 24 bits per array element). The color channels may be referred to as different image "planes". In some cases, only a single channel may be used, for example, to represent a "gray" or brightness channel. Different color spaces may be used depending on the application. For example, an image capture device may inherently generate a frame of YUV image data characterized by a brightness channel Y (e.g., luminance) and two opposing color channels U and V (e.g., two chrominance components aligned roughly by blue-green and red-green). Similar to the audio data 155, the image data 145 may be processed after capture. For example, one or more image filtering operations may be applied, and / or the image data 145 may be resized and / or cropped.
[0037] Referring to the examples in Figures 1A and 1B, one or more of the image interface 140 and audio interface 150 may be local to the hardware within the automated vehicle. For example, each of the image interface 140 and audio interface 150 may include a wired connection of the respective image and audio capture devices to at least one processor configured to implement an audio processing module 130. In one case, the image and audio interfaces 140,150 may include a serial interface from which image and audio data 145,155 can be received. In a distributed vehicle control system, the image and audio capture devices 140,150 may be communicatively coupled to a central system bus, and the image and audio data 145,155 may be stored in one or more storage devices (e.g., random access memory or solid-state storage). In the latter case, the image and audio interfaces 140,150 may include a communicative connection of at least one processor configured to implement an audio processing module to one or more storage devices. For example, at least one processor may receive image and audio data 145,155. Each of the 55 can be configured to access data by reading it from a given memory location. In some cases, the image and audio interfaces 140, 150 may include a wireless interface, and the voice processing module 130 may be remote from the automated vehicle. Different approaches and combinations are possible.
[0038] Figure 1A shows an example where person 102 is the driver of the automated vehicle, but in other applications, one or more image and audio capture devices may be provided to capture image data featuring people who are not controlling the automated vehicle, such as passengers. For example, an automated vehicle may have multiple image capture devices provided to capture image data of people present in one or more passenger seats in the vehicle (e.g., different locations within the vehicle, such as the front and rear). Audio capture devices may similarly be provided to capture speech from different people, for example, microphones may be located in each door or door frame of the vehicle. In one case, multiple audio capture devices may be provided within the vehicle, and audio data may be captured from one or more of these for supplying data to the audio interface 150. In one case, preprocessing of the audio data may include selecting audio data from channels considered to be closest to the person speaking, and / or combining audio data from multiple channels within the automated vehicle. As will be discussed later, one example described in this specification facilitates speech processing in a vehicle with multiple passengers.
[0039] Example speaker preprocessing module Figure 2 shows an exemplary voice processing device 200. For example, the voice processing device 200 may be used to implement the voice processing device 120 shown in Figures 1A and 1B. The voice processing device 200 may form part of an in-vehicle automatic voice recognition system. In other cases, the voice processing device 200 may be adapted for use outside a vehicle, such as in a home or office.
[0040] The speech processing device 200 includes a speaker preprocessing module 220 and a speech processing module 230. The speech processing module 230 may be similar to the speech processing module 130 in Figure 1B. In this example, the image interface 140 and the audio interface 150 are omitted for clarity, but they may form the image input portion of the speaker preprocessing module 220 and the audio input portion for the speech processing module 230, respectively. The speaker preprocessing module 220 is configured to receive image data 245 and output a speaker feature vector 225. The speech processing module 230 is configured to receive audio data 255 and the speaker feature vector 225 and use them to generate linguistic features 260.
[0041] The speech processing module 230 is implemented by a processor. The processor may be a processor in a local embedded computing system within the vehicle and / or a processor in a remote server computing device (a so-called "cloud" processing device). In one case, the processor may include dedicated speech processing hardware parts such as one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and so-called "system on chip" (SoC) components. In another case, the processor may be configured to process computer program code, such as firmware, which is stored in an accessible storage device and loaded into memory for execution by the processor. The speech processing module 230 is configured to analyze human speech, such as that of person 102, based on audio data 225 and image data 245. In the present case, the image data 245 is preprocessed by the speaker preprocessing module 220 to generate speaker feature vectors 225. Similar to the speech processing module 230, the speaker preprocessing module 220, Any combination of hardware and software is possible. In one case, the speaker pre-processing module 220 and the speech processing module 230 may be implemented on a common embedded circuit board for the vehicle.
[0042] In one case, the speech processing module 230 includes an acoustic model configured to process audio data 255 and predict phoneme data used to analyze speech. In this case, the linguistic features 260 may include phoneme data. The phoneme data may relate to one or more phoneme symbols from, for example, a predefined alphabet or dictionary. In one case, the phoneme data may include a predicted sequence of phonemes. In another case, the phoneme data may include probabilities for one or more sets of phoneme components, such as phoneme symbols and / or subsymbols from a predefined alphabet or dictionary, and sets of state transitions (for example, for a hidden Markov model). The acoustic model may be configured to receive audio data in the form of an audio feature vector. The audio feature vector may include numerical values representing one or more of the Mel-frequency cepstrum coefficients (MFCCs) and filter bank outputs. In one case, the audio feature vector may relate to the current window in time (often referred to as a "frame") and include differences in feature changes between the current window and one or more other windows in time (for example, previous windows). The current window can have a width of w milliseconds, for example, in one case w may be approximately 25 milliseconds. Other features may include, for example, the signal energy metric and logarithmic scaling output. The audio data after preprocessing may contain frames (e.g., vectors) of multiple elements (e.g., from 10 to over 1000 elements), each element containing a numerical representation associated with a particular audio feature. In one example, there may be approximately 25-50 MEL filter bank features and a similarly sized set of intra features, and (for example) There may be a set of delta features of similar size (representing the first-order derivative) and a set of double delta features of similar size (representing, for example, the second-order derivative).
[0043] The speaker preprocessing module 220 can be configured to obtain speaker feature vectors 225 in many different ways. In one case, the speaker preprocessing module 220 can obtain at least a portion of the speaker feature vectors 225 from memory, for example, via a lookup operation. In one case, a portion of the speaker feature vectors 225 containing the i and / or x vectors described below can be extracted from memory. In this case, image data 245 can be used to determine a particular speaker feature vector 225 to be extracted from memory. For example, image data 245 can be classified by the speaker preprocessing module 220 to select a specific user from a set of registered users. In this case, the speaker feature vectors 225 may contain numerical representations of features correlated to the selected specific user. In another case, the speaker preprocessing module 220 can compute the speaker feature vectors 225. For example, the speaker preprocessing module 220 can compute a compressed or dense numerical representation of significant information within the image data 245. This may contain a vector with many elements that is smaller in size than the image data 245. In this case, the speaker preprocessing module 220 may introduce an information bottleneck in computing the speaker feature vector 225. In one case, the computation is determined based on a set of parameters such as weights, biases, and / or probability coefficients. The values for these parameters may be determined through a training phase using a set of training data. In one case, the speaker feature vector 225 may be buffered or stored as a static value after the set of computations. In this case, the speaker feature vector 225 may be retrieved from memory based on image data 245 during subsequent utterances. Another example illustrating how the speaker feature vector may be computed is described below. If the speaker feature vector 225 includes a component relating to lip movements, this component may be provided on a real-time or near-real-time basis, and the data storage It cannot be extracted from the page.
[0044] In one case, the speaker feature vector 225 may contain a fixed-length one-dimensional array (e.g., a vector) of numbers, such as one value per element of the array. In other cases, the speaker feature vector 225 may contain a multi-dimensional array, such as two or more dimensions representing multiple one-dimensional arrays. The numbers may be integer values (for example, within a range set by a specific bit length (8 bits give a range of 0 to 255)) or floating-point values (for example, defined as 32-bit or 64-bit floating-point values). Floating-point values may be used when normalization is applied to the visual feature tensor, for example, when the values are mapped to a range of 0 to 1 or -1 to 1. As an example, the speaker feature vector 225 may contain an array of 256 elements, each element being an 8 or 16-bit value, but its form may vary depending on the implementation. In general, a speaker feature vector 225 has less information content than the corresponding frame of image data, for example, using the example above. A speaker feature vector 225 of length 256 with 8-bit values is smaller than a 640x480 video frame with three channels of 8-bit values. That is, 2048 bits vs 7,372,800 bits. Information content can be measured in bits or in the form of entropy measurement.
[0045] In one case, the speech processing module 230 includes an acoustic model, and the acoustic model includes a neural network architecture. For example, the acoustic model may include one or more of the following: a deep neural network (DNN) architecture with multiple hidden layers; a hybrid model including a neural network architecture and one or more of a Gaussian mixture model (GMM) and a hidden Markov model (HMM); and a connectionist temporal classification (CTC) model including one or more recurrent neural networks that operate on a sequence of inputs and generate a sequence of linguistic features as an output. Predictions can be output at the frame level (for example, for phoneme symbols or subsymbols), and previous (and in some cases future) predictions can be used to determine possible or most likely sequences of phoneme data for a utterance. Approaches such as beam search and the Viterbi algorithm are used at the output end of the acoustic model. This can be used to further determine the sequence of phoneme data output from the acoustic model. The acoustic model can be trained step by step in time.
[0046] If the speech processing module 230 includes an acoustic model, and the acoustic model includes a neural network architecture (for example, a "neural" acoustic model), then the speaker feature vector 225, along with the audio data 255, may be provided as input to the neural network architecture. The speaker feature vector 225 and the audio data 255 can be combined in many ways. In the simplest case, the speaker feature vector 225 and the audio data 255 may be concatenated to form a longer combined vector. In other cases, different input preprocessing may be performed on each of the speaker feature vector 225 and the audio data 255, for example, one or more attention layers, feedfor A feed-forward layer and / or an embedding layer may be applied, Subsequently, the results of these layers can be combined. Different sets of layers can be applied to different inputs. In other cases, the speech processing module 230 may include another form of statistical model, such as a stochastic acoustic model, and the speaker feature vector 225 may include one or more numerical parameters (e.g., probability coefficients) to configure the speech processing module 230 for a particular speaker.
[0047] An exemplary voice processing device 200 provides an improvement for voice processing within a vehicle. In this context, high levels of environmental noise, such as road and engine noise, may be present. Furthermore, acoustic distortion caused by the enclosed interior space of the automated vehicle may also be present. These factors can make processing audio data difficult in the comparative example. For example, the speech processing module 230 may fail to generate linguistic features 260 and / or generate sequences that do not adequately match the linguistic features 260. However, the configuration in Figure 2 allows the speech processing module 230 to be configured or fitted based on speaker features determined based on image data 245. This provides the speech processing module 230 with additional information, such as being able to select linguistic features consistent with a particular speaker by utilizing correlations between appearance and acoustic features. These correlations can be long-term temporal correlations, such as general facial appearance, and / or short-term temporal correlations, such as specific lip and mouth positions. This leads to higher accuracy despite challenging noise and acoustic contexts. This can help reduce speech analysis errors, for example, by improving the end-to-end transcription path and / or improving the audio interface for voice commands. In some cases, this example may utilize a driver-facing camera that is typically configured to monitor the driver for drowsiness and / or distraction. In some cases, there may be speaker-dependent feature vector components extracted based on the recognized speaker, and / or speaker-dependent feature vector components that include lip-movement features. The latter component may be determined based on features not configured for individual users, for example, common features for all users may be applied, but lip movements are associated with the speaker. In other cases, the extraction of lip-movement features may be configured based on a specific identified user.
[0048] Examples of facial recognition Figure 3 shows an exemplary speech processing device 300. The speech processing device 300 shows additional components that may be used to implement the speaker preprocessing module 220 in Figure 2. Some components shown in Figure 3 are similar to their counterparts shown in Figure 2 and have the same reference numbers. Features described above with reference to Figure 2 may also apply to Example 300 in Figure 3. Like the exemplary speech processing device 200 in Figure 2, the exemplary speech processing device 300 in Figure 3 includes a speaker preprocessing module 320 and a speech processing module 330. The speech processing module 330 receives audio data 355 and speaker feature vectors 325 and computes a set of linguistic features 360. The speech processing module 330 may be configured in a similar manner to the example described above with reference to Figure 2.
[0049] Figure 3 shows several subcomponents of the speaker preprocessing module 320. These include the face recognition module 370, the vector generator 372, and the data store 374. While these are shown as subcomponents of the speaker preprocessing module 320 in Figure 3, they may be implemented as separate components in other examples. In the example in Figure 3, the speaker preprocessing module 320 receives image data 345 featuring a person's face area. The person may include a driver or passenger in a vehicle, as described above. The face recognition module 370 performs face recognition on the image data to identify a person, such as a driver or passenger in a vehicle. The face recognition module 370 may include any combination of hardware and software to perform face recognition. In one case, the face recognition module 370 may be implemented using off-the-shelf hardware components such as the B5T-007001 supplied by Omron Electronics. In this example, the face recognition module 370 detects a user based on the image data 345 and outputs a user identifier 376. The user identifier 376 is passed to the vector generator 372. The vector generator 372 uses the user identifier 376 to obtain a speaker feature vector 325 associated with the identified person. In some cases, the vector generator 372 may extract the speaker feature vector 325 from the data store 374. The speaker feature vector 325 is then shown in Figure 2. As described in reference, it is passed to the speech processing module 330 for use.
[0050] In the example in Figure 3, the vector generator 372 may obtain speaker feature vectors 325 in different ways depending on a set of operating parameters. In one case, the operating parameters include a parameter indicating whether a particular number of speaker feature vectors 325 have been computed for a particular identified user (e.g., identified by user identifier 376). In another case, a threshold is defined associated with a certain number of previously computed speaker feature vectors. If this threshold is 1, speaker feature vectors 325 may be computed for the first utterance and then stored in the data store 374, and for subsequent utterances, speaker feature vectors 325 may be extracted from the data store 374. If the threshold is 1 or greater, such as n, n speaker feature vectors 325 may be generated, and the (n+1)th speaker feature vector 325 may be obtained as a composite function of the previously n speaker feature vectors 325 extracted from the data store 374. The composite function may include averaging or interpolation. In one case, once the (n+1)th speaker feature vector 325 is computed, it is used as a static speaker feature vector for a constitutable number of future utterances.
[0051] In the example above, the use of a data store 374 to store the speaker feature vector 325 can reduce the runtime computation requirements for the in-vehicle system. For example, the data store 374 may include a local data storage device within the vehicle, and therefore the speaker feature vector 325 may be extracted from the data store 374 for a particular user rather than being computed by the vector generator 372.
[0052] In one case, at least one computing function used by the vector generator 372 may involve a cloud processing resource (e.g., a remote server computing device). In this case, where there are limitations on connectivity between the vehicle and the cloud processing resource, the speaker feature vector 325 can be extracted from local storage as a static vector, rather than relying on any functionality provided by the cloud processing resource.
[0053] In one case, the speaker preprocessing module 320 may be configured to generate a user profile for each newly recognized person in the vehicle. For example, prior to or during speech detection, such as being captured by an audio capture device, the face recognition module 370 may attempt to match image data 345 against previously observed faces. If no match is found, the face recognition module 370 may generate (or be instructed to generate) a new user identifier 376. In one case, a component of the speaker preprocessing module 320, such as the face recognition module 370 or a vector generator 372, may be configured to generate a new user profile if no match is found, and the new user profile may be indexed using the new user identifier. The speaker feature vector 325 is then associated with the new user profile, and the new user profile may be stored in the data store 374, ready to be extracted if a future match is made by the face recognition module 370. Thus, an in-vehicle image capture device may be used for face recognition to select a user-specific speech recognition profile. User profiles can be calibrated through an enrollment process, for example, when a driver first uses the vehicle, or they can be learned based on data collected during use.
[0054] In one case, the speaker processing module 320 may be configured to reset the data store 374. At the time of manufacture, the device 374 may not contain user profile information. A new user profile is created during use as described above. It may be added to datastore 374. The user may command a reset of the stored user identifier. In some cases, the reset may only be performed during professional service, such as when a car is serviced at a service shop or sold by an authorized dealer. In some cases, the reset may be provided at any time through a password provided by the user.
[0055] In an example where the vehicle includes multiple image capture devices and multiple audio capture devices, the speaker preprocessing module 320 may provide further functionality to determine the appropriate face area from one or more captured images. In one case, audio data from multiple audio capture devices may be processed to determine the nearest audio capture device associated with the utterance. In this case, the nearest image capture device associated with the determined nearest audio capture device may be selected, and image data 345 from this device may be sent to the face recognition module 370. In another case, the face recognition module 370 may be configured to receive multiple images from multiple image capture devices, each image including an associated flag to indicate whether it should be used to identify the user currently speaking. Thus, the speech processing device 300 in Figure 3 may be used to identify a speaker from multiple people in the vehicle, and the speech processing module 330 may be configured for specific features of that speaker. This can further improve speech processing within the vehicle when multiple people are speaking in the confined space of the vehicle.
[0056] i vector In one example described in this specification, speaker feature vectors such as speaker feature vector 225 or 325 may include data generated based on audio data such as audio data 255 or 355 in Figures 2 and 3, which are shown by dashed lines in Figure 3. In one case, at least a portion of the speaker feature vectors may include vectors generated based on factor analysis. In this case, speech may be represented as a vector M which is a linear function of one or more factors. The factors may be combined in linear and / or nonlinear models. One of these factors may include a speaker and session independent supervector m, which may be based on a Universal Background Model (UBM). Another of these factors may include a speaker-dependent vector w, which may further depend on the channel or session, or provide yet another factor that depends on the channel and / or session. In one case, the factor analysis is performed using a Gaussian mixture (GMM) mixture. In the simplest case, speaker utterance can be represented by a supervector M determined as M = m + Tw, where T is at least a matrix defining the speaker subspace. The speaker-dependent vector w may have multiple elements with floating-point values. In this case, the speaker feature vector may be based on the speaker-dependent vector w. One method for calculating w, sometimes referred to as the "i-vector," is described in the paper "Front-End Factor Analysis For Speaker Verification" by Najim Dehak, Patrick Kenny, Reda Dehak, Pierre Dumouchel, and Pierre Ouellet.The paper in question was published in 2010 in IEEE Transactions On Audio, Speech And Language Processing 19, no.4, pages 788-798, and is incorporated herein by reference. In one example, at least a portion of the speaker feature vector includes at least a portion of the i vector, which can be seen as a speaker-dependent vector determined for utterances from audio data.
[0057] In the example in Figure 3, the vector generator 372 can compute i-vectors for one or more utterances. If no speaker feature vectors are stored in the data store 374, The i-vector can be calculated by the vector generator 372 based on one or more frames of audio data for the utterance 355. In this example, the vector generator 372 may repeatedly calculate the i-vector for each utterance (e.g., each voice query) until a threshold number calculation is performed for a particular user, such as being identified using a user identifier 376 determined by the face recognition module 370. In this case, after a particular user is identified based on image data 345, an i-vector for each user is stored in the data store 374 for each utterance. The i-vector is also used to output the speaker feature vector 325. Once a threshold number calculation is performed, such as calculating around 100 i-vectors, the vector generator 372 may be configured to calculate a profile for a particular user using the i-vectors stored in the data store 374. The profile may use the user identifier 376 as an index and may include static (e.g., unchanging) i-vectors calculated as a composite function of the stored i-vectors. The vector generator 372 may be configured to calculate the profile upon receiving the (n+1)th query, or as part of a background or periodic function. In one case, a static i-vector can be calculated as the average of the stored i-vectors. For example, once a profile is generated by the vector generator 372 and stored in the data store 374 using a user identifier to associate the profile with a specific user, the i-vector for the user can be extracted from the data store 374 instead of being calculated and used for future speech analysis. This can reduce the computational overhead of generating speaker feature vectors and reduce i-vector variability.
[0058] x vector In one example, speaker feature vectors, such as speaker feature vectors 225 or 325, may be computed using a neural network architecture. For example, in one case, the vector generator 372 of the speaker preprocessing module 320 in Figure 3 may include a neural network architecture. In this case, the vector generator 372 may compute at least a portion of the speaker feature vectors by reducing the dimensionality of the audio data 355. For example, the vector generator 372 may include one or more deep neural network layers configured to receive one or more frames of the audio data 355 and output fixed-length vector outputs (e.g., one vector per language). One or more pooling, nonlinear function, and softmax layers may further... In one case, speaker feature vectors can be provided to the paper "Spoken Language Recognition using X-vectors" by David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur. It can be generated based on the x vector as described in ). The paper was published in Odyssey (pp. 105-111) in 2018 and is incorporated by reference herein.
[0059] The x vector may be used in a similar manner to the i vector described above, and the above approach applies to the speaker feature vector generated using the x and i vectors. In one case, both the i and x vectors may be determined, and the speaker feature vector may include a supervector containing elements from both the i and x vectors. Both the i and x vectors may be combined by concatenation or weighted summation, for example, typically floating-point values and / or values normalized within a given range. It includes elements of a possible numerical value. In this case, datastore 374 may contain values stored for one or more of the i and x vectors, so that once a threshold is reached, a static value is calculated and stored along with a specific user identifier for future extraction. In one case, interpolation may be used to determine the speaker feature vector from one or more i and x vectors. In one case, interpolation is performed from different vector sources from the same vector source. This can be done by averaging the resulting speaker feature vectors.
[0060] If the speech processing module includes a neural acoustic model, a fixed-length format for speaker feature vectors may be defined. The neural acoustic model can then be trained using the defined speaker feature vectors, for example, as determined by speaker preprocessing modules 220 or 320 in Figures 2 and 3. If the speaker feature vectors include elements derived from one or more i-vector and x-vector computations, the neural acoustic model can "learn" to construct acoustic processing based on speaker-specific information embodied or embedded within the speaker feature vectors. This can increase acoustic processing accuracy, particularly in vehicles such as autonomous vehicles. In this case, image data provides a mechanism for quickly associating a specific user with the computed or stored vector elements.
[0061] Example of a speech processing module Figure 4 shows an exemplary speech processing module 400. Speech processing module 400 can be used to implement speech processing modules 130, 230, or 330 in Figures 1, 2, and 3. Other speech processing module configurations may be used in other examples.
[0062] As in previous examples, the speech processing module 400 receives audio data 455 and speaker feature vectors 425. The audio data 455 and speaker feature vectors 425 may consist of any of the examples described herein. In the example of Figure 4, the speech processing module 400 includes an acoustic model 432, a language model 434, and a speech parser 436. As described above, the acoustic model 432 generates phoneme data 438. The phoneme data may include one or more predicted sequences of phoneme symbols or subsymbols, or other forms of proto-language units. In this process, multiple predicted sequences may be generated, along with probability data indicating the likelihood of a particular symbol or sub-symbol at each time step.
[0063] Phoneme data 438 is communicated to language model 434, and for example, acoustic model 432 is coupled to language model 434 in a communicative manner. Language model 434 is configured to receive phoneme data 438 and generate transcription 440. Transcription 440 may include text data such as strings, word parts (e.g., stems and suffixes) or words. Characters, word parts and words may be selected from a predefined dictionary, such as a predefined set of possible outputs at each time step. In some cases, phoneme data 438 may be processed before being passed to language model 434, or may be preprocessed by language model 434. For example, beamforming may be applied to the probability distribution (e.g., about phonemes) output from acoustic model 432.
[0064] The language model 434 is communicatively coupled to the speech parser 436. The speech parser 436 receives the transcription 440 and uses it to parse the speech. In some cases, the speech parser 436 generates speech data 442 as a result of the speech analysis. The speech parser 436 may be configured to determine the command and / or command data associated with the speech based on the transcription. In one case, the language model 434 may generate multiple possible text sequences, for example, by probabilistic information about units in the text, and the speech parser 436 may be configured to determine the final text output, for example, in the form of ASCII or Unicode character encoding or a voice command or command data. If it is determined that the transcription 440 contains a voice command, the speech parser 436 may be configured to execute the command according to the command data or to instruct the execution of the command. This may result in response data output as speech data 442. The speech data 442 may include, for example, a command instruction, which is relayed to the person making the speech. The output is then provided on the dashboard 108 and / or via the vehicle's voice system. In some cases, the language model 434 may include a statistical language model, and the speech parser 436 may include a separate “meta” language model configured to rescore alternative assumptions as outputs by the statistical language model. This may be via an ensemble model that uses voting to determine the final output, such as final transcription or command identification.
[0065] Figure 4 shows, with solid lines, an example in which the acoustic model 432 receives speaker feature vectors 425 and audio data 455 as input and uses these inputs to generate phoneme data 438. For example, the acoustic model 432 may include a neural network architecture (including a hybrid model with other non-neural components), the speaker feature vectors 425 and audio data 455 may be provided as input to the neural network architecture, and the phoneme data 438 may be generated based on the output of the neural network architecture.
[0066] The dashed lines in Figure 4 indicate additional connections that may be configured in a given implementation. In the first case, the speaker feature vector 425 may be accessed by one or more of the language model 434 and the speech parser 436. For example, if the language model 434 and the speech parser 436 further include their respective neural network architectures, these architectures may be configured to receive the speaker feature vector 425 as an additional input, in addition to, for example, the phoneme data 438 and the transcription 440, respectively. If the speech data 442 includes a command identifier and one or more command parameters, the complete speech processing module 400 may be trained in an end-to-end manner, given a training set with ground truth output and training samples for the audio data 455 and the speaker feature vector 425.
[0067] In a second implementation, the speech processing module 400 in Figure 4 may include one or more recurrent connections. In one case, the acoustic model may include a recurrent model such as an LSTM. In other cases, there may be feedback between modules. In Figure 4, a dashed line shows a first recurrent connection between the speech parser 436 and the language model 434, and a dashed line shows a second recurrent connection between the language model 434 and the acoustic model 432. In this second case, the current state of the speech parser 436 may be used to construct a future prediction for the language model 434, and the current state of the language model 434 may be used to construct a future prediction for the acoustic model 432. In some cases, recurrent connections may be omitted to simplify the processing pipeline and enable easier learning. In one case, a recurrent connection may be used to compute an attention or weighting vector to be applied in the next time step.
[0068] Neural Speaker Preprocessing Module Figure 5 shows an exemplary speech processing device 500 that uses a neural speaker preprocessing module 520 and a neural speech processing module 530. In Figure 5, the speaker preprocessing module 520, which can implement modules 220 or 320 in Figures 2 and 3, includes a neural network architecture 522. In Figure 5, the neural network architecture 522 is configured to receive image data 545. In other cases, the neural network architecture 522 may further receive audio data such as audio data 355, as shown by the dashed pathway in Figure 3, for example. It can receive optical data. In these other cases, the vector generator 372 in Figure 3 may include a neural network architecture 522.
[0069] In Figure 5, the neural network architecture 522 has at least convolutional This includes neural architectures. In a given architecture, there may be one or more feedforward neural network layers between the last convolutional neural network layer and the output layer of the neural network architecture 522. The neural network architecture 522 may include adapted forms of AlexNet, VGGNet, GoogLeNet, or ResNet architectures. The neural network architecture 522 may be replaced in a modular manner as more precise architectures become available.
[0070] The neural network architecture 522 outputs at least one speaker feature vector 525, which can be derived and / or used as described in one of the other examples. Figure 5 shows, for example, a case where image data 545 contains multiple frames from a video camera, and each frame features a person's face area. In this case, multiple speaker feature vectors 525 can be computed using the neural network architecture, for example, one for each input frame of the image data. In other cases, there may be a many-to-one relationship between the input data frames and the speaker feature vectors. Note that it is not necessary to temporally synchronize the input image data 545 and the samples of the output speaker feature vector 525 using a recurrent neural network system. For example, a recurrent neural network architecture can act as an encoder (or integrator) over time. In one case, the neural network architecture 522 may be configured to generate the x vector as described above. In one case, the x-vector generator may be configured to receive image data 545, process this image data using a convolutional neural network architecture, and then combine the output of the convolutional neural network architecture with an audio-based x-vector. In another case, a known x-vector configuration may be extended to receive image data and audio data and generate a single speaker feature vector that embodies information from both modal pathways.
[0071] In Figure 5, the neural speech processing module 530 is a speech processing module such as one of modules 230, 330, or 400, which include a neural network architecture. For example, the neural speech processing module 530 may include a hybrid DNN-HMM / GMM system and / or a fully neural CTC system. In Figure 5, the neural speech processing module 530 receives frames of audio data 555 as input. Each frame may correspond to a time window, such as a window of w milliseconds elapsed with respect to time-series data from an audio capture device. Frames of audio data 555 may be asynchronous with frames of image data 545, and for example, frames of audio data 555 are likely to have a higher frame rate. Also, retention mechanisms and / or recurrent neural network architectures may be applied within the neural speech processing module 530 to provide temporal encoding and / or sample integration. In another example, the neural speech processing module 530 is configured to process frames of audio data 555 and speaker feature vectors 525 to generate a set of linguistic features. As discussed in this specification, a reference to a neural network architecture includes one or more neural network layers (in one case, one or more hidden layers and a “deep” architecture having multiple layers), each layer which may be separated from the next layer by a nonlinearity such as a hyperbolic tangent unit (tanh unit) or a normalized linear function unit (RELU). Other functions may be embodied within layers that include pooling operations.
[0072] The neural speech processing module 530 may include one or more components as shown in Figure 4. For example, the neural speech processing module 530 may include at least one The acoustic model may include a neural network. In the example in Figure 5, the neural network architectures of the neural speaker preprocessing module 520 and the neural speech processing module 530 may be learned jointly. In this case, the training set may include frames of image data 545, frames of audio data 555, and ground truth linguistic features (e.g., ground truth phoneme sequences, text transcriptions or voice command classifications, and command parameter values). Both the neural speaker preprocessing module 520 and the neural speech processing module 530 may be learned in an end-to-end manner using this training set. In this case, errors between predicted linguistic features and ground truth linguistic features may be propagated after the neural speech processing module 530 and back through the neural speaker preprocessing module 520. Parameters for both neural network architectures may then be determined using a gradient descent approach. The neural network architecture 522 of the neural speaker preprocessing module 520 can "learn" parameter values (such as values for weights and / or biases for one or more neural network layers) to generate one or more speaker feature vectors 525 that improve acoustic processing in the in-vehicle environment, and the neural speaker preprocessing module 520 learns to extract features from human face areas that help improve the accuracy of the output linguistic features.
[0073] Training of neural network architectures as described in this specification is not typically performed on in-vehicle devices (although it may be done if desired). In one case, training may be performed on a computing device with access to considerable processing resources, such as a server computer device having multiple processing units (regardless of whether it is a CPU, GPU, field-programmable gate array (FPGA) or other dedicated processor architecture) and a large memory portion to hold sets of training data. In another case, training may be performed using coupled accelerator devices, such as coupled FPGA or GPU-based devices. In yet another case, the trained parameters may be transferred from a remote server device to the in-vehicle device, for example, as part of an over-the-air update. It can communicate with the embedded system.
[0074] Examples of acoustic model selection Figure 6 shows an exemplary speech processing module 600 that uses speaker feature vectors 625 to construct an acoustic model. Speech processing module 600 may be used to at least partially implement one of the speech processing modules described in other examples. In Figure 6, speech processing module 600 includes a database of acoustic model configurations 632, an acoustic model selector 634, and an acoustic model instance 636. The database of acoustic model configurations 632 stores a number of parameters for constructing an acoustic model. In this example, an acoustic model instance 636 may contain a general acoustic model that is instantiated (e.g., configured or calibrated) using a specific set of parameter values from the database of acoustic model configurations 632. For example, the database of acoustic model configurations 636 may store multiple acoustic model configurations. Each acoustic model configuration may be associated with a different user and include one or more default acoustic model configurations used when the user is not detected or when the user is detected but not particularly recognized.
[0075] In some cases, speaker feature vectors 625 can be used to represent a specific regional accent on behalf of (or with) a particular user. This can be useful in countries like India, where many different regional accents may exist. In this case, speaker feature vectors 625 can be used to dynamically load an acoustic model based on accent recognition performed using speaker feature vectors 625. For example, this could be used to represent a speech This may be possible if the feature vector 625 includes the x vector as described above. This may be useful if there are multiple accent models (e.g., multiple acoustic model configurations for each accent) stored in the vehicle's memory. This may then allow multiple separately trained accent models to be used.
[0076] In one case, the speaker feature vector 625 may include classifications of people in the vehicle. For example, the speaker feature vector 625 may originate from a user identifier 376 output by the face recognition module 370 in Figure 3. In another case, the speaker feature vector 625 may include a set of classifications and / or probabilities output by a neural speaker preprocessing module, such as module 520 in Figure 5. In the latter case, the neural speaker preprocessing module may include a softmax layer that outputs "probabilities" for a set of potential users (including a classification for "unrecognized"). In this case, one or more frames of the input image data 545 may be reduced to a single speaker feature vector 525.
[0077] In Figure 6, the acoustic model selector 634 receives, for example, speaker feature vectors 625 from a speaker preprocessing module and selects an acoustic model configuration from a database of acoustic model configurations 632. This may operate in a similar manner to the example in Figure 3 described above. If the speaker feature vectors 625 include a set of user classifications, the acoustic model selector 634 may select an acoustic model configuration based on these classifications, for example, by sampling probability vectors and / or selecting the one with the highest probability value as the determined person. Parameter values for the selected configuration may be extracted from the database of acoustic model configurations 632 and used to instantiate an acoustic model instance 636. Thus, different acoustic model instances may be used for different identified users in the vehicle.
[0078] In Figure 6, an acoustic model instance 636 also receives audio data 655, such as an acoustic model instance 636 configured by the acoustic model selector 634 using a configuration extracted from a database of acoustic model configurations 632. The acoustic model instance 636 is configured to generate phoneme data 660 for use in analyzing the utterances associated with the audio data 655 (for example, those featured within the audio data 655). The phoneme data 660 may include, for example, a sequence of phoneme symbols from a predefined alphabet or dictionary. Thus, in the example in Figure 6, the acoustic model selector 634 selects an acoustic model configuration from the database 632 based on the speaker feature vector, and the acoustic model configuration is used to instantiate the acoustic model instance 636 and process the audio data 655.
[0079] The acoustic model instance 636 may include both neural and non-neural architectures. In one case, the acoustic model instance 636 may include a non-neural model. For example, the acoustic model instance 636 may include a statistical model. A statistical model may use symbol frequencies and / or probabilities. In one case, the statistical model may include a Bayesian network or a Bayesian model such as a classifier. In these cases, the acoustic model configuration may include a specific set of symbol frequencies and / or previous probabilities measured in different environments. Thus, the acoustic model selector 634 allows a specific source of utterance (e.g., a person or a user) to be determined based on both visual (and in some cases, audio) information, which may provide an improvement over using audio data 655 alone to generate the phoneme sequence 660.
[0080] In another case, acoustic model instance 636 may include a neural model. In this case, acoustic model selector 634 and acoustic model instance 636 may include a neural network architecture. In this case, the database of acoustic model configuration 632 is The acoustic model selector 634, which may be omitted, may supply vector inputs to the acoustic model instance 636 to construct an instance. In this case, the training data may be constructed from image data used to generate a ground truth set of speaker feature vectors 625, audio data 655, and phoneme outputs 660. Such a system may be trained collaboratively.
[0081] Exemplary image preprocessing Figures 7 and 8 illustrate exemplary image preprocessing operations that may be applied to image data obtained from inside a vehicle, such as an automated vehicle. Figure 7 shows an exemplary image preprocessing pipeline 700 including an image preprocessor 710. The image preprocessor 710 may include any combination of hardware and software to implement its functions as described in this specification. In one case, the image preprocessor 710 may include hardware components that form part of an image capture circuit coupled to one or more image capture devices; in another case, the image preprocessor 710 may be implemented by computer program code (such as firmware) executed by a processor in the vehicle's control system. In one case, the image preprocessor 710 may be implemented as part of a speaker preprocessing module as illustrated in this specification; in another case, the image preprocessor 710 may be communicatively coupled to a speaker preprocessing module.
[0082] In Figure 7, the image preprocessor 710 receives image data 745, such as the image from the image capture device 110 in Figure 1A. The image preprocessor 710 processes the image data to extract one or more portions of the image data. Figure 7 shows the output 750 of the image preprocessor 710. The output 750 may include one or more image annotations, such as metadata associated with one or more pixels of the image data 745, and / or features defined using pixel coordinates within the image data 745. In the example in Figure 7, the image preprocessor 710 performs face detection on the image data 745 to determine a first image area 752. The first image area 752 may be cropped and extracted as an image portion 762. The first image area 752 may be defined using a bounding box (for example, at least the top-left and bottom-right (x,y) pixel coordinates for a rectangular area). Face detection may be a preliminary step for face recognition. For example, face detection may determine a face area in image data 745, and face recognition may classify a face area as belonging to a given person (e.g., within a set of people). In the example in Figure 7, the image preprocessor 710 further identifies a mouth area in image data 745 to determine a second image area 754. The second image area 754 may be cropped and extracted as an image portion 764. The second image area 754 may also be defined using a bounding box. In one case, the first and second image areas 752,754 may be determined in relation to a set of detected facial features 756. These facial features 756 may include one or more of the eye, nose, and mouth areas.Detection of one or more of the facial features 756 and / or the first and second areas 752,754 may be performed using a neural network approach, or the Viola-Jones face detection algorithm, as described in "Robust Real-Time Face Detection" by Paul Viola and Michael J Jones, published in 2004 in the International Journal of Computer Vision 57, pp. 137-154, Netherlands, which is incorporated by reference herein. Known face detection algorithms such as M may be used. In one example, one or more of the first and second image areas 752, 754 are used by a speaker preprocessing module as described herein to obtain speaker feature vectors. For example, the first image area 752 may provide input image data for the face recognition module 370 in Figure 3 (i.e., it may be used to supply image data 345). An example of using the second image area 754 is described below with reference to Figure 9.
[0083] Figure 8 illustrates the effect of using an image capture device configured to capture electromagnetic radiation having infrared wavelengths. In some cases, image data 810 obtained by an image capture device such as the image capture device 110 in Figure 1A may be affected by low-light conditions. In Figure 8, the image data 810 includes an area of shadow 820 that partially obscures the face area (including, for example, the first and second image areas 752, 754 in Figure 7). In these cases, an image capture device configured to capture electromagnetic radiation having infrared wavelengths may be used. This provides adaptations (such as filters that can be removed in hardware and / or software) to the image capture device 110 in Figure 1A, and This may include providing a near-infrared (NIR) camera. The output from such an image capture device is schematically shown as image data 830. In image data 830, the face area 840 is reliably captured. In this case, image data 830 provides an illumination-invariant representation that is not affected by changes in illumination, such as changes in illumination that may occur during nighttime driving. In these cases, image data 830 is provided to an image preprocessor 710 and / or a speaker preprocessing module, as described in this specification.
[0084] lip reading example In some cases, the speaker feature vector described in this specification may include at least a set of elements representing the features of a person's mouth or lips. In these cases, the speaker feature vector may be speaker-dependent because it changes based on the content of image data featuring the mouth or lip area of a person. In the example in Figure 5, the neural speaker preprocessing module 520 may encode lip or mouth features used to generate the speaker feature vector 525. These may be used to improve the performance of the speech processing module 530.
[0085] Figure 9 shows another exemplary speech processing device 900 that uses lip features to form at least a portion of the speaker feature vector. Similar to the previous example, the speech processing device 900 includes a speaker preprocessing module 920 and a speech processing module 930. The speech processing module 930 receives audio data 955 (frames of audio data in this case) and outputs linguistic features 960. The speech processing module 930 may be configured as in other examples described herein.
[0086] In Figure 9, the speaker preprocessing module 920 is configured to receive two different sources of image data. In this example, the speaker preprocessing module 920 receives a first set of image data 962 featuring the area of a person's face. This may include a first image area 762, as extracted by the image preprocessor 710 in Figure 7. The speaker preprocessing module 920 also receives a second set of image data 964 featuring the area of a person's lips or mouth. This may include a second image area 764, as extracted by the image preprocessor 710 in Figure 7. The second set of image data 964 may be relatively small, such as a small cropped portion of a larger image obtained using the image capture device 110 in Figure 1A. In other examples, the first and second sets of image data 962,964 do not have to be cropped and may include a copy of the image set from the image capture device. Different configurations are possible. That is, cropping the image data may provide improvements in processing speed and learning, but the neural network architecture can be trained to work with a wide range of different image sizes.
[0087] The speaker preprocessing module 920 includes two components in Figure 9: a feature extraction component 922 and a lip feature extractor 924. The lip feature extractor 924 forms part of the lip-reading module. The feature extraction component 922 is part of the speaker preprocessing module in Figure 3. It may be configured in a manner similar to Joule 320. In this example, the feature extraction component 922 receives a first set of image data 962 and outputs a vector portion 926 consisting of one or more i-vectors and x-vectors (for example, as described above). In one case, the feature extraction component 922 receives a single image per utterance, and the lip feature extractor 924 and the speech processing module 930 receive multiple frames over the duration of the utterance. In one case, if the face recognition performed by the feature extraction component 922 has a confidence value below a threshold, the first set of image data 962 may be updated (for example, by using another / current frame of the video), and face recognition is reapplied until the confidence value reaches the threshold (or exceeds a predetermined number of trials). As described with reference to Figure 3, the vector portion 926 may be calculated based on audio data 955 for a first number of utterances, and once the first number of utterances is exceeded, it is extracted from memory as a static value.
[0088] The lip feature extractor 924 receives a second set of image data 964. The second set of image data 964 may include cropped frames of the image data that focus on the mouth or lip area. The lip feature extractor 924 may receive the second set of image data 964 at the frame rate of the image capture device and / or at the frame rate at which it is subsampled (e.g., every two frames). The lip feature extractor 924 outputs a set of vector parts 928. These vector parts 928 may include the output of an encoder, which may include a neural network architecture. The lip feature extractor 924 may include a convolutional neural network architecture to provide a fixed-length vector output (e.g., 256 or 512 elements with integer or floating-point values). The lip feature extractor 924 may output a vector portion for each input frame of the image data 964, and / or use a recurrent neural network architecture (e.g., using Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), or "Transformer" architecture) to extract features over time steps. This can encode the following. In the latter case, the output of the lip feature extractor 924 may include one or more of the hidden states and outputs of the recurrent neural network. One exemplary implementation for the lip feature extractor 924 is described in "Lip reading sentences in the wild" by Chung, Joon Son et al. at the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), which is incorporated by reference herein.
[0089] In Figure 9, the speech processing module 930 receives a vector portion 926 from the feature extraction component 922 and a vector portion 928 from the lip feature extractor 924 as input. In one case, the speaker preprocessing module 920 may combine the vector portions 926 and 928 to form a single speaker feature vector. In another case, the speech processing module 930 may receive the vector portions 926 and 928 separately but treat them as different parts of the speaker feature vector. The vector portions 926 and 928 may be combined by one or more of the speaker preprocessing module 920 and the speech processing module 920 to form a single speaker feature vector, for example, using concatenation or a more complex attention-based mechanism. If the sample rates of one or more frames of the vector portions 926, 928 and audio data 955 are different, for example, a receive-and-hold architecture (new Recurrent temporal encoding (where a more variable value is kept constant at a given value until a sample value is received), (for example, using LSTM or GRU as described above), or attention weighting per time step. A common sample rate can be achieved through an attention-based system where the vector changes.
[0090] The speech processing module 930 may be configured to use vector portions 926,928 as described in other examples provided in this specification, for example, these may be input to a neural acoustic model as speaker feature vectors along with audio data 955. In an example in which the speech processing module 930 includes a neural acoustic model, a training set may be generated based on input video from an image capture device, input audio from an audio capture device, and ground truth linguistic features (for example, the image preprocessor 710 in Figure 7 may be used to obtain first and second sets of image data 962,964 from raw input video).
[0091] In one example, the vector portion 926 may include an additional set of elements. The values of this additional set of elements are derived from the encoding of a first set of image data 962 using a neural network architecture, such as 522 in Figure 5. These additional elements may represent a “face encoding,” and the vector portion 928 may represent a “lip encoding.” The face encoding may remain static for a speech, while the lip encoding may change during a speech or may include multiple “frames” for a speech. Figure 9 shows an example using both the lip feature extractor 924 and the feature extraction component 922, although in one example the feature extraction component 922 may be omitted. In subsequent examples, a lip-reading system for use in a vehicle may be used in a similar manner to the speech processing device 500 in Figure 5.
[0092] Exemplary Automobile Vehicle Figures 10A and 10B illustrate an example in which the vehicle described in this specification is an automated vehicle. Figure 10A shows a side view 1000 of the automobile 1005. The automobile 1005 includes a control unit 1010 for controlling the components of the automobile 1005. Components of a voice processing device 120, as shown in Figure 1B (and other examples), may be incorporated into this control unit 1010. In other cases, components of the voice processing device 120 may be implemented as separate units, depending on the option of connection with the control unit 1010. The automobile 1005 further includes at least one image capture device 1015. For example, at least one image capture device 1015 may include the image capture device 110 shown in Figure 1A. In this example, at least one image capture device 1015 may be communicatively coupled to the control unit 1010 and controlled by the control unit 1010. In other examples, at least one image capture device 1015 communicates with a control unit 1010 and is remotely controlled. Along with the functions described herein, at least one image capture device 1015 may be used for video communications such as voice over Internet Protocol calls, environmental monitoring, and driver alertness monitoring using video data. Figure 10A further shows at least one audio capture device in the form of a side-mounted microphone 1020. These may realize the audio capture device 116 shown in Figure 1A.
[0093] The image capture device described in this specification may include one or more still or video cameras configured to capture frames of image data in response to commands or at a predetermined sampling rate. The image capture device may provide coverage of both the front and rear of the vehicle interior. In one case, the predetermined sampling rate may be less than the frame rate for full-resolution video; for example, a video stream may be captured at 30 frames / second, but the sampling rate of the image capture device may capture at this rate or at a lower rate such as 1 frame / second. The image capture device may capture one or more frames of image data having one or more color channels (e.g., RGB or YUV as described above). In one case, the frame rate and the frame rate may be... Aspects of an image capture device can be configured, such as size and resolution, the number of color channels, and the sample format. In some cases, frames of image data may be downsampled. For example, a video capture device that captures video at a "4K" resolution of 3840 x 2160 may be downsampled to 640 x 480 or lower. Alternatively, for low-cost embedded devices, low-resolution image capture devices that capture frames of image data at 320 x 240 or lower may be used. In some cases, even inexpensive low-resolution image capture devices may provide enough visual information to improve audio processing. As mentioned above, image capture devices may further include image preprocessing and / or filtering components (e.g., contrast adjustment, noise reduction, color correction, cropping, etc.). In some cases, low-latency and / or high-frame-rate image cameras that meet the more stringent Automotive Safety Integrity Level (ASIL) of the ISO 26262 automotive safety standard are available. Apart from their safety advantages, lip-reading accuracy may be improved by providing more temporal information. This could be useful for recurrent neural networks for more accurate feature probability estimation.
[0094] Figure 10B shows an overhead view 1030 of the vehicle 1005. The vehicle 1005 includes a front seat 1032 and a rear seat 1034 for holding passengers facing a front-mounted microphone for voice capture. The vehicle 1005 includes a driver visual console 1036 having safety critical display information. The driver visual console 1036 may include a portion of the dashboard 108 as shown in Figure 1A. The vehicle 1005 further includes a general console 1038 having navigation, entertainment, and environmental control functions. A control unit 1010 may control the general console 1038 and may implement a local voice processing module and a wireless network communication module as shown in Figure 1A 120. The wireless network communication module may transmit one or more of the image data, audio data, and speaker feature vectors generated by the control unit 1010 to a remote server for processing. The vehicle 1005 further includes a side-mounted microphone 1020, a front overhead multi-microphone speech capture unit 1042, and a rear overhead multi-microphone speech capture unit 1044. The front and rear speech capture units 1042, 1044 provide additional audio capture devices for capturing speech audio, canceling noise, and identifying the speaker's location. In one case, the front and rear speech capture units 1042, 1044 may further include additional image capture devices for capturing image data featuring each of the passengers in the vehicle.
[0095] In the example in Figure 10B, one or more of the microphones and speech capture units 1020, 1042, and 1044 may provide audio data to an audio interface such as 140 in Figure 1B. The microphone or array of microphones may be configured to capture or record audio samples at a predetermined sampling rate. In some cases, aspects of each audio capture device, such as sampling rate, bit resolution, number of channels, and sample format, may be configurable. The captured audio data may be pulse code modulated. Any of the audio capture devices may further include audio preprocessing and / or filtering components (e.g., contrast adjustment, noise reduction). Similarly, one or more of the image capture devices may provide image data to an image interface such as 150 in Figure 1B, and may further include video preprocessing and / or filtering components (e.g., contrast adjustment, noise reduction).
[0096] Figure 11 shows an example of the interior of an automobile 1100 as seen from the front seat 1032. For example, Figure 11 may include a view toward the windshield 104 in Figure 1A. Figure 11 shows a steering wheel 1106 (like the steering wheel 106 in Figure 1), a side microphone 1120 (like one of the side microphones 1020 in Figures 10A and 10B), a rearview mirror 1142 (which may include a front overhead multi-microphone speech capture unit 1042), and a projection device 1130. The projection device 1130 may be used to project an image 1140 onto the windshield for use, for example, as an additional visual output device (in addition to the driver's visual console 1036 and the general console 1038). In Figure 11, the image 1140 includes directions. These may be directions projected after a voice command such as "Tell me the directions to Mallmart." Other examples may use a simpler response system.
[0097] Local and remote voice processing for vehicles In some cases, the functions of the speech processing module may be distributed as described in this specification. For example, some functions may be computed locally within the vehicle 1005, and other functions may be computed by a remote ("cloud") server device. In some cases, functions may be replicated on both the vehicle ("client") side and the remote server device ("server" side). In these cases, if a connection to the remote server device is not available, processing may be performed by the local speech processing module; if a connection to the remote server device is available, one or more of the audio data, image data, and speaker feature vectors may be sent to the remote server device for analysis of the captured speech. The remote server device may have processing resources (e.g., a central processing unit i.e., a CPU, a graphical processing unit i.e., a GPU, and random access memory), and therefore may provide an improvement over local performance if a connection is available. This may be a trade-off with respect to latency in the processing pipeline (e.g., local processing is faster). In one case, the local speech processing module may provide a first output, which may be supplemented and / or enhanced by the results of the remote speech processing module.
[0098] In one case, a vehicle, such as automobile 1005, may be communicably coupled to a remote server device via at least one network. The network may include one or more local and / or wide-area networks that can be implemented using various physical technologies (e.g., wired technologies such as Ethernet®, and / or wireless technologies such as Wi-Fi® (IEEE 802.11) standards and cellular communication technologies). In some cases, the network may include a mixture of one or more private and public networks, such as the Internet. The vehicle and the remote server device may communicate over the network using different technologies and communication pathways.
[0099] Referring to the exemplary speech processing device 300 in Figure 3, in one case, vector generation by the vector generator 372 can be performed locally or remotely, but the data store 374 resides locally within the vehicle 1005. In this case, static speaker feature vectors can be computed locally and / or remotely, but are stored locally within the data store 374. Subsequently, the speaker feature vectors 325 can be extracted from the data store 374 within the vehicle, rather than being received from a remote server device. This can improve speech processing latency.
[0100] If the voice processing module is remote from the vehicle, the local voice processing unit includes a transceiver that transmits data derived from one or more of the following to the voice processing module: audio data, image data, and speaker feature vectors, and receives control data from the speech analysis. In one case, the transceiver may include a wired or wireless physical interface and one or more communication protocols that provide a method for sending and / or receiving requests in a predetermined format. In another case, the transceiver may include an application layer interface that runs on an Internet Protocol suite. In this case, the application layer interface may be configured to receive communications directed to a specific Internet Protocol address that identifies a remote server device, with routing based on pathnames or web addresses performed by one or more proxy and / or communication (e.g., “web”) servers.
[0101] In some cases, linguistic features generated by the speech processing module may be mapped to a set of voice commands and data about those voice commands (as described, for example, with reference to the speech parser 436 in Figure 4). In one case, the speech data 442 may be used by the control unit 1010 of the automobile 1005 to execute voice commands. In another case, the speech parser 436 may reside in a remote server device, and the speech analysis may involve identifying the appropriate service to execute the voice commands from the output of the speech processing module. For example, the speech parser 436 may provide an application programming interface (API) to the identified server. The application programming interface (application programming interface) can be configured to make a request. This request includes a command identified from the output of the language model and any command data. For example, the utterance "Where is Mallmart?" may yield the text output "Where is Mallmart," which can be mapped to a Directional Services API request for vehicle mapping data having the desired location parameter "Mallmart" and the vehicle's current location derived from a positioning system, such as a global positioning system. The response can be extracted, communicated to the vehicle, and displayed as shown in Figure 11.
[0102] In one case, the remote speech parser 436 communicates response data to the control unit 1010 of the vehicle 1005. This may include machine-readable data that is communicated to the user, for example, via a user interface or audio output. The response data may be processed, and the response to the user may be output on one or more of the driver visual console 1036 and the general console 1038. Providing a response to the user may include displaying text and / or images on the display screens of one or more of the driver visual console 1036 and the general console 1038, or including sound output via a text-to-speech module. In some cases, the response data may be processed in the control unit 1005 and may include audio data used to generate audio output, for example, via one or more speakers. The response may be spoken to the user via a speaker mounted inside the vehicle 1005.
[0103] Exemplary embedded computing system Figure 12 shows an exemplary embedded computing system 1200 that can implement a voice processing device as described in this specification. A system similar to the embedded computing system 1200 may be used to implement the control unit 1010 in Figure 10. The exemplary embedded computing system 1200 includes one or more computer processor (CPU) cores 1210 and zero or more graphics processor (GPU) cores 1220. The processors are connected to a random access memory (RAM) device 1240 for program code and data storage via board-level wiring 1230. The embedded computing system 1200 further includes a network interface 1250 to enable the processors to communicate with remote systems and specific vehicle control circuits 1260. By executing instructions stored in the RAM device through interface 1230, the CPU 1210 and / or GPU 1220 can perform functions as described in this specification. In some cases, constrained embedded An integrated computing device may have a similar general arrangement of components, but in some cases may have fewer computing resources and may not have a dedicated graphics processor 1220.
[0104] Example of a speech processing method Figure 13 shows an exemplary method 1300 for processing speech that may be performed to improve in-vehicle speech recognition. Method 1300 begins with block 1305, in which audio data is received from an audio capture device. The audio capture device may be located inside the vehicle. The audio data may feature utterances from a user. Block 1305 includes capturing data from one or more microphones, such as devices 1020, 1042, and 1044 in Figures 10A and 10B. In one case, block 1305 may include receiving audio data via a local audio interface, and in another case, block 1305 may include receiving audio data via a network at an audio interface that is remote from the vehicle, for example.
[0105] In block 1310, image data is received from an image capture device. The image capture device may be located inside the vehicle and may include, for example, the image capture device 1015 shown in Figures 10A and 10B. In one case, block 1310 may receive image data via a local image interface, and in another case, block 1310 may receive image data via a network via an image interface that is remote from the vehicle, for example.
[0106] In block 1315, speaker feature vectors are obtained based on image data. This may include, for example, implementing one of the speaker preprocessing modules 220, 320, 520, and 920. Block 1315 may be performed by the local processor of the vehicle 1005 or by a remote server device. In block 1320, speech is analyzed using a speech processing module. This may include, for example, implementing one of the speech processing modules 230, 330, 400, 530, and 930. Block 1320 includes a number of subblocks, which include, in subblock 1322, providing speaker feature vectors and audio data as input to the acoustic model of the speech processing module. This may include operations similar to those described with reference to Figure 4. In some cases, the acoustic model includes a neural network architecture. In subblock 1324, phoneme data is predicted based on the speaker feature vectors and audio data, using at least a neural network architecture. This may involve using a neural network architecture that is trained to accept speaker feature vectors as input, in addition to audio data. Since both speaker feature vectors and audio data contain numerical representations, they can similarly be processed by a neural network architecture. In some cases, an existing CTC or hybrid acoustic model may be configured to accept a concatenation of speaker feature vectors and audio data, and may be trained using a training set that additionally includes image data (for example, used to derive the speaker feature vectors).
[0107] In some cases, block 1315 includes performing facial recognition on image data to identify people inside the vehicle. For example, this may be done as described with reference to the facial recognition module 370 in Figure 3. After this, user profile data about the people (e.g., inside the vehicle) may be obtained based on facial recognition. For example, user profile data may be extracted from data store 374 using user identifier 376, as described with reference to Figure 3. Subsequently, speaker feature vectors are used in the user profile data The speaker feature vector can be obtained according to the following. In one case, the speaker feature vector can be extracted from the user profile data as a static set of element values. In another case, the user profile data may indicate that the speaker feature vector is to be computed using one or more of the audio and image data received in blocks 1305 and 1310, for example. In one case, block 1315 includes comparing a number of stored speaker feature vectors associated with the user profile data to a predetermined threshold. For example, the user profile data may indicate how many previous voice queries have been made by a user identified using facial recognition. In response to a number of stored speaker feature vectors falling below the predetermined threshold, the speaker feature vector can be computed using one or more of the audio and image data. In response to a number of stored speaker feature vectors being greater than the predetermined threshold, a static speaker feature vector can be obtained, for example, a speaker feature vector stored in or accessible through the user profile data. In this case, the static speaker feature vector can be generated using a number of stored speaker feature vectors.
[0108] In one example, block 1315 may include processing image data to generate one or more speaker feature vectors based on lip movements within a person's face area. For example, a lip-reading module such as a lip feature extractor 924 or a preferably configured neural speaker preprocessing module 520 may be used. The output of the lip-reading module may be used to supply one or more speaker feature vectors to a speech processing module and / or be combined with other values (e.g., i or x vectors) to generate a larger speaker feature vector.
[0109] In one example, block 1320 includes providing phoneme data to a language model of a speech processing module, using the language model to predict a transcript of the utterances, and using the transcript to determine control commands for the vehicle. For example, block 1320 may include operations similar to those described with reference to Figure 4.
[0110] Examples of vocal analysis Figure 14 shows an exemplary processing system 1400, which includes a non-temporary computer-readable storage medium 1410 for storing instruction 1420. When instruction 1420 is executed by at least one processor 1430, it causes at least one processor to perform a series of operations. The operations in this example use the approach previously described for generating a transcription of the utterance. These operations may be performed, for example, in a vehicle as described above, or the in-vehicle example may be extended to non-vehicle-based situations, which may be implemented using, for example, a desktop, laptop, mobile, or server computing device.
[0111] Through instruction 1432, processor 1430 is configured to receive audio data from an audio capture device. This may include accessing local memory containing the audio data and / or receiving a data stream or set of array values over a network. The audio data may be in a form as described with reference to other examples in this specification. Through instruction 1434, processor 1430 is configured to receive a speaker feature vector. The speaker feature vector is obtained based on image data from an image capture device, which features the user's face area. For example, the speaker feature vector may be obtained using an approach described with reference to any of Figures 2, 3, 5, and 9. The speaker feature vector may be computed locally by processor 1430 and stored in local memory or It may be accessed and / or received via (for example) a network interface. Through instruction 1436, the processor 1430 is instructed to analyze the speech using the speech processing module. The speech processing module may include any of the modules described with reference to any one of Figures 2, 3, 4, 5 and 9.
[0112] Figure 14 shows that instruction 1436 can be broken down into a number of further instructions. Through instruction 1440, processor 1430 is instructed to provide speaker feature vectors and audio data as input to the acoustic model of the speech processing module. This can be achieved in a manner similar to that described with reference to Figure 4. In this example, the acoustic model includes a neural network architecture. Through instruction 1442, processor 1430 is instructed to predict phoneme data based on the speaker feature vectors and audio data, using at least the neural network architecture. Through instruction 1444, processor 1430 is instructed to provide phoneme data to the language model of the speech processing module. This can further be done in a manner similar to that shown in Figure 4. Through instruction 1446, processor 1430 is instructed to generate a transcript of the utterance using the language model. For example, the transcript may be generated as the output of the language model. In some cases, the transcript may be used by a control system, such as the control unit 1010 in automobile 1005, to execute voice commands. In other cases, the transcript may include output for a speech-to-text system. In the latter case, image data may be extracted from a webcam or similar device that is communicatively coupled to a computing device including processor 1430. For mobile computing devices, image data may be obtained from a forward-facing image capture device.
[0113] In one example, the speaker feature vector received in accordance with instruction 1434 includes one or more of the following: a speaker-dependent vector element (e.g., an i-vector or x-vector component) generated based on audio data; a speaker lip-movement-dependent vector element (e.g., generated by a lip-reading module) generated based on image data; and a speaker face-dependent vector element generated based on image data. In one case, the processor 1430 may include a portion of a remote server device, and the audio data and speaker image vector may be received from an automated vehicle, for example, as part of a distributed processing pipeline.
[0114] Exemplary implementation examples An example of speech processing, including automatic speech recognition, is described. One example concerns processing a particular spoken language. Various examples operate similarly for other languages or combinations of languages. One example improves the accuracy and robustness of speech processing by incorporating additional information derived from an image of the speaker. This additional information can be used to improve the linguistic model. The linguistic model may include one or more acoustic, phonetic, and linguistic models.
[0115] One example described in this specification may be implemented to address the unique challenges of performing automated speech recognition in vehicles such as automobiles. In one combined example, image data from a camera may be used to recognize a face in order to determine lip-reading features and enable the construction and selection of i-vector and / or x-vector profiles. By implementing the approach described in this specification, it may be possible to perform automated speech recognition in the noisy, multi-channel environment of an automated vehicle.
[0116] Examples described in this specification include, for example, audio input or audio Instead of having an acoustic model that only accepts a separate acoustic model for image data, the efficiency of speech processing can be increased by including one or more features derived from image data, such as lip position or movement, in the speaker feature vector provided as input to an acoustic model that also accepts audio data as input (singular model).
[0117] A certain method and set of operations may be performed by instructions stored on a non-temporary computer-readable medium. The non-temporary computer-readable medium stores code containing instructions that, when executed by one or more computers, cause the computers to perform the steps of the method described in this specification. The non-temporary computer-readable medium may include one or more of the following: a rotating magnetic disk, a rotating optical disk, a flash random-access memory (RAM) chip, and other mechanically moving storage media or solid-state storage media. Any type of computer-readable medium is suitable for storing code containing instructions according to various examples.
[0118] Some examples described in this specification may be implemented as so-called system-on-chip (SoC) devices. SoC devices may be used to control many embedded in-vehicle systems and to implement the functions described in this specification. In one case, one or more of the speaker pre-processing module and speech processing modules may be implemented as an SoC device. An SoC device comprises one or more processors (e.g., CPU or GPU), random access memory (e.g., RAM such as off-chip dynamic RAM or DRAM), and Ethernet®, WiFi®, 3G, 4G Long-Term Evolution (LTE), and 5G. The SoC device may also include a network interface for wired or wireless connectivity, such as other wireless interface standards. The SoC device may further include various I / O interface devices as required for different peripherals, such as touchscreen sensors, geolocation receivers, microphones, speakers, Bluetooth® peripherals, and USB devices such as keyboards and mice. By executing instructions stored in the RAM device, the processor of the SoC device may perform the steps of the method as described in this specification.
[0119] While some examples are described in this specification, it should be noted that different combinations of components may be possible, as different examples may be different. Notable features are shown to better illustrate the examples, but it is clear that certain features may be added, modified, and / or omitted without altering the functional aspects of these examples as described.
[0120] Various examples are methods that utilize the behavior of either humans and / or machines, or a combination thereof. The method examples are complete, as most configuration steps occur anywhere in the world. Some examples are one or more non-temporary computer-readable media configured to store such instructions for the methods described herein. The examples can be realized by any machine holding a non-temporary computer-readable media containing any of the required codes. Some examples can be realized as physical devices such as semiconductor chips, hardware description language representations of the logical or functional behavior of such devices, and one or more non-temporary computer-readable media configured to store such hardware description language representations. In principle, descriptions in this specification describing aspects and embodiments encompass their structural and functional equivalents. Elements described in this specification to be combined either have a valid relationship that can be realized by direct connection or indirectly by one or more other intervening elements.
Claims
1. An audio interface configured to receive audio data from an audio capture device, An image interface configured to receive image data from an image capture device, Includes a speech processing module configured to analyze human speech based on the aforementioned audio data and image data, The speech processing module includes a language model that is communicatively coupled to an acoustic model in order to receive phoneme data and generate a transcription representing the utterance. It further includes a speaker preprocessing module, The speaker preprocessing module is configured to compute speaker feature vectors calculated based on the audio data and the image data, The speech processing module is a computing device that includes an acoustic model selector that selects the configuration of the acoustic model based on the speaker feature vector.
2. An audio interface configured to receive audio data from an audio capture device, An image interface configured to receive image data from an image capture device, Includes a speech processing module configured to analyze human speech based on the aforementioned audio data and image data, The speech processing module includes a language model that is communicatively coupled to an acoustic model in order to receive phoneme data and generate a transcription representing the utterance. It further includes a speaker preprocessing module, The speaker preprocessing module is configured to compute a speaker feature vector calculated based on at least one of the audio data and the image data. The speech processing module includes an acoustic model selector that selects the configuration of the acoustic model based on the speaker feature vector, The aforementioned speaker feature vector is, A first speaker-dependent portion generated based on the aforementioned audio data, A computing device comprising a second part that depends on the movement of the speaker's lips and is generated based on the aforementioned image data.
3. The computing device according to claim 2, wherein the speaker feature vector further comprises a third portion that depends on the speaker's face and is generated based on the image data.
4. The computing device according to any one of claims 1 to 3, further comprising an acoustic model configured to process the audio data and predict the phoneme data used to analyze the utterance.
5. The computing device according to any one of claims 1 to 4, wherein the image data includes a human face area.
6. A computing device according to any one of claims 1 to 5, comprising an image capture device configured to capture electromagnetic radiation having an infrared wavelength as image data, wherein the image capture device is configured to send the image data to the image interface.
7. The aforementioned audio capture device captures sound from inside the vehicle, The voice processing module is remote from the vehicle, and the computing device is A computing device according to any one of claims 1 to 5, comprising a transceiver that transmits data derived from the audio data and the image data to the voice processing module and receives control data based on the analysis of the speech.
8. An audio interface configured to receive audio data from an audio capture device, An image interface configured to receive image data from an image capture device, Includes a speech processing module configured to analyze human speech based on the aforementioned audio data and image data, The aforementioned audio processing module is It includes a language model that is communicatively coupled to an acoustic model in order to receive phoneme data and generate a transcription representing the utterance, It further includes a speaker preprocessing module, The aforementioned speaker preprocessing module is: Receiving the aforementioned image data, The system is configured to calculate a speaker feature vector based on the image data in order to predict the phoneme data, The aforementioned speaker feature vector is, A first speaker-dependent portion generated based on the aforementioned audio data, It includes a second part that is generated based on the aforementioned image data and depends on the movement of the speaker's lips, The aforementioned audio processing module is A database of acoustic model configurations, An acoustic model selector that selects an acoustic model configuration from the database based on the speaker feature vector, Includes an acoustic model instance that processes the aforementioned audio data, A computing device wherein the acoustic model instance is instantiated based on the acoustic model configuration selected by the acoustic model selector, and the acoustic model instance is configured to generate the phoneme data used to analyze the utterance.
9. An audio interface configured to receive audio data from an audio capture device, An image interface configured to receive image data from an image capture device, Includes a speech processing module configured to analyze human speech based on the aforementioned audio data and image data, The aforementioned audio processing module is It includes a language model that is communicatively coupled to an acoustic model in order to receive phoneme data and generate a transcription representing the utterance, It further includes a speaker preprocessing module, The aforementioned speaker preprocessing module is: Receiving the aforementioned image data, The system is configured to calculate a speaker feature vector based on the image data in order to predict the phoneme data, The aforementioned speaker feature vector is, A first speaker-dependent portion generated based on the aforementioned audio data, It includes a second part that is generated based on the aforementioned image data and depends on the movement of the speaker's lips, A computing device in which the first part is one or more of the i-vector and the x-vector.