Adaptive visual speech recognition

The adaptive visual speech recognition model efficiently adapts to new speakers with minimal data and computational resources, addressing the limitations of traditional networks by leveraging speaker-specific embeddings and fine-tuning on consumer devices.

JP2025102754APending Publication Date: 2025-07-08DEEPMIND TECH LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2025026055
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-18
Filing Date
2025-02-20
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing visual speech recognition neural networks require extensive training data and computational resources, making them impractical for adapting to new speakers on consumer devices.

Method used

A sample-efficient, adaptive visual speech recognition model that uses a distributed computing system for initial training and allows adaptation on less powerful hardware with minimal data, such as a few minutes of video recording, by learning speaker-specific embedding vectors and optionally fine-tuning neural network parameters.

Benefits of technology

Enables rapid adaptation to new speakers using significantly less data and computational resources than traditional methods, allowing deployment on consumer devices like mobile phones or laptops.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025102754000001_ABST
    Figure 2025102754000001_ABST
Patent Text Reader

Abstract

To provide adaptive visual speech recognition.SOLUTION: Methods, systems, and apparatus include computer programs encoded on computer storage media, for processing video data using an adaptive visual speech recognition model. One of the methods includes the steps of: receiving a video that includes a plurality of video frames that depict a first speaker; obtaining a first embedding characterizing the first speaker; and processing a first input including (i) the video and (ii) the first embedding using a visual speech recognition neural network having a plurality of parameters. The visual speech recognition neural network is configured to process the video and the first embedding in accordance with trained values of the parameters to generate a speech recognition output that defines a sequence of one or more words being spoken by the first speaker in the video.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to a visual speech recognition neural network.

Background Art

[0002] A neural network is a machine learning model that uses one or more layers of non-linear units to predict the output of received inputs. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, such as the next hidden layer or the output layer. Each layer of the network generates an output from the received inputs according to the current values of its respective set of parameters.

[0003] An example of a neural network is a visual speech recognition neural network. A visual speech recognition neural network decodes speech from the movements of a speaker's mouth. In other words, a visual speech recognition neural network receives a video of a speaker's face as input and generates as output a text representing the words spoken by the speaker depicted in the video.

[0004] An example of a visual speech recognition neural network is LipNet. LipNet was first described in LipNet: End-to-End Sentence-Level Lipreading by Assael et al. in arXiv preprint arXiv:1611.01599 (2016), which is available at arxiv.org. LipNet is a deep neural network that uses spatio-temporal convolution and recurrent neural networks to map a variable-length sequence of video frames to text.

[0005] Another example of a visual speech recognition neural network is described in "Large-Scale Visual Speech Recognition" by Shillingford et al. in arXiv preprint arXiv:1807.05612 (2018), available at arxiv.org. Large-scale visual speech recognition describes a deep visual speech recognition neural network that maps a video of lips to a sequence of phoneme distributions, and an audio decoder that outputs a sequence of words from the sequence of phoneme distributions generated by the deep neural network.

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Non-Patent Document 4

Summary of the Invention

Means for Solving the Problems

[0007] This specification describes a system implemented as a computer program on one or more computers in one or more locations that can generate a sample-efficient, adaptive visual speech recognition model. In this context, being sample-efficient and adaptive means that the model can be customized to recognize the speech of a new speaker using far less training data than was used to train the adaptation model. For example, while training the adaptation model may require several hours of video recording per individual speaker, adapting the model to a new speaker may require only a few minutes of video recording of the new speaker.

[0008] The training system can train a visual speech recognition model using multiple embedding vectors for individual speakers and a visual speech recognition neural network. Due to the computationally intensive nature of the training process, the training can be performed by a distributed computing system with hundreds or thousands of computers, such as a data center.

[0009] The output of the training process is an adaptive visual speech recognition model that can efficiently adapt to a new speaker. Model adaptation generally involves learning new embedding vectors for the new speaker and optionally may include fine-tuning the parameters of the neural network for the new speaker. The adaptation data can be a few seconds or minutes of video of the new speaker and the corresponding text transcript. For example, the video can be the video of the speaker while the speaker is speaking the text on a text prompt presented to the user on a user device.

[0010] Accordingly, the adaptation process is significantly less computationally intensive than the original training process. Accordingly, the adaptation process can be executed on much less powerful hardware such as, to name a few, a mobile phone or another wearable device, a desktop or laptop computer, or another internet-enabled device installed in the user's home.

[0011] In one aspect, the method includes receiving a video including a plurality of video frames depicting a first speaker, obtaining a first embedding characterizing the first speaker, and using a visual speech recognition neural network having a plurality of parameters to process a first input comprising (i) the video and (ii) the first embedding, the visual speech recognition neural network being configured to process the video and the first embedding according to the trained values of the parameters to generate an audio recognition output defining a sequence of one or more words spoken by the first speaker in the video.

[0012] In some implementations, the visual speech recognition neural network is configured to generate additional input channels from the first embedding and combine the additional channels with one or more of the frames in the video before processing the frames in the video to generate the audio recognition output.

[0013] In some implementations, the visual-audio recognition neural network comprises a plurality of hidden layers, and for at least one of the hidden layers, the neural network is configured to generate an additional hidden channel from a first embedding and to combine the hidden channel with the output of the hidden layer before providing an output for processing by another hidden layer of the visual-audio recognition neural network.

[0014] In some implementations, the method further comprises obtaining adaptation data for a first speaker, the adaptation data comprising one or more videos of the first speaker and a ground truth transcription for each of the videos, and determining a first embedding for the first speaker using the adaptation data.

[0015] In some implementations, the method further comprises obtaining pre-trained values of model parameters determined by training a visual-audio recognition neural network on training data comprising training examples corresponding to a plurality of speakers different from the first speaker, and the step of determining the first embedding comprises determining the first embedding using the pre-trained values and the adaptation data.

[0016] In some implementations, the step of determining the first embedding comprises initializing the first embedding, and the step of updating the first embedding by repeatedly performing operations comprises processing each of one or more video segments and the first embedding using the visual-audio recognition neural network according to current values of the parameters to generate respective audio recognition outputs for each of the one or more video segments, and updating the first embedding to minimize a loss function that measures a respective error between the ground truth transcription of the video segment and the respective audio recognition output of the video segment for each of the one or more video segments.

[0017] In some implementations, for each of one or more video segments, to minimize a loss function that measures a respective error between the ground truth transcription of the video segment and the respective speech recognition output of the video segment, a step of updating a first embedding includes a step of backpropagating a gradient of the loss function through a visual speech recognition neural network to determine a gradient of the loss function with respect to the first embedding, and a step of updating the first embedding using the gradient of the loss function with respect to the first embedding.

[0018] In some implementations, the current value is equal to the pre-trained value and the trained value, and the model parameters are fixed when determining the first embedding.

[0019] In some implementations, the operation further includes a step of updating a current value of the parameters of the visual speech recognition neural network based on a gradient of the loss function with respect to the parameters of the visual speech recognition neural network, and the trained value is equal to the current value after determining the first embedding vector.

[0020] In some implementations, the method further includes a step of applying a decoder to the speech recognition output of the video to generate a sequence of one or more words spoken by a first speaker in the video.

[0021] In some implementations, the speech recognition output comprises a respective probability distribution over the vocabulary of text elements for each of the video frames.

[0022] Particular implementations of the subject matter described herein can be implemented to realize one or more of the following advantages.

[0023] To quickly adapt to a new speaker, the adaptive visual-audio recognition models described herein can be used with data orders of magnitude less than the data used to train the model. This enables the adaptation process to be performed by the end-user consumer hardware rather than in a data center.

[0024] Furthermore, when a multi-speaker visual-audio recognition model is trained on a large dataset representing videos of multiple speakers, a large number of data samples from the training data tend to be underfit. This can be due to small imbalances in the collected video data, or it can be due to the finite capacity of the model to capture all the scenarios represented in a large video dataset. The techniques described address these issues by first training a speaker-conditioned visual-audio recognition model conditioned on (i) a speaker's video and (ii) a speaker embedding, and then adapting the speaker-conditioned visual-audio recognition model by learning the embedding of the new speaker (optionally, by fine-tuning the weights of the model).

[0025] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0026]

Figure 1

Figure 2

Figure 3

Best Mode for Carrying Out the Invention

[0027] Like reference numerals and names in the various drawings indicate like elements.

[0028] FIG. 1 is a diagram showing an exemplary architecture 100 for training an adaptive visual-audio recognition model.

[0029] The architecture 100 includes a visual-audio recognition neural network 110a that is trained using an embedding table 120 that stores embedding vectors for a plurality of different individual speakers.

[0030] The visual-audio recognition neural network 110a can be any suitable visual-audio recognition neural network that receives, as input, a video of a speaker and a speaker's embedding vector, as will be described later, and processes the video and the embedding vector to generate, as output, an audio recognition output representing a predicted transcription of the audio spoken by the speaker within the video.

[0031] As used herein, "video" includes only a sequence of video frames and does not include audio corresponding to the video frame sequence. Thus, the visual-audio recognition neural network 110a generates an audio recognition output without accessing the audio data of the audio actually spoken by the speaker.

[0032] An example of a visual speech recognition neural network is LipNet. LipNet was first described in LipNet: End-to-End Sentence-Level Lipreading by Assael et al. in the arXiv preprint arXiv:1611.01599 (2016), which is available on arxiv.org. LipNet is a deep neural network that uses spatio-temporal convolution and recurrent neural networks to map a variable-length sequence of video frames to text.

[0033] Another example of a visual speech recognition neural network is described in Large-Scale Visual Speech Recognition by Shillingford et al. in the arXiv preprint arXiv:1807.05612 (2018), which is available on arxiv.org. Large-Scale Visual Speech Recognition also uses spatio-temporal convolution and recurrent neural networks to describe a deep visual speech recognition neural network that maps lip videos to a sequence of phoneme distributions.

[0034] In general, either of the two above-described visual speech recognition neural network architectures, or any other visual speech recognition neural network architecture, can be modified to accept as input an embedding vector along with a video depicting a speaker, e.g., the speaker's face, or the speaker's mouth or lips, at each of a plurality of time steps. The architecture can be modified to process the embedding vector in any of a variety of ways.

[0035] As an example, the system can generate additional input channels using an embedding vector having the same spatial dimensions as the video frames, and then combine the additional channels with the video frames, for example, by concatenating the additional channels along the channel dimension to each video frame (e.g., as if the input channels were color intensity values added to the color (e.g., RGB) of the video frames). For example, the system can generate additional channels using the embedding vector by applying a broadcast operation to the values in the embedding vector to generate a two-dimensional spatial map having the same dimensions as the video frames.

[0036] As another example, the system can generate additional hidden channels using an embedding vector having the same spatial dimensions as a particular one of the outputs of the hidden layer of a neural network, e.g., one of the spatio-temporal convolutional layers in the neural network, and then combine the additional hidden channels with the output of the hidden layer, for example, by adding the additional channels to the output of the hidden layer, by concatenating the additional channels along the channel dimension to the output of the hidden layer, by multiplying the additional channels and the output of the hidden layer element-wise, or by applying a gating mechanism between the output of the hidden layer and the additional channels.

[0037] The components shown in FIG. 1 can be implemented by a distributed computing system comprising a plurality of computers configured to train a visual-audio recognition neural network 110a.

[0038] The computing system can train a visual-audio recognition neural network 110a on training data 130 that includes a plurality of training examples 132. Each training example 132 corresponds to a respective speaker and includes (i) a video 140 of the corresponding speaker, and (ii) a respective ground truth transcription 150 of the audio spoken by the corresponding speaker within the video.

[0039] Each speaker corresponding to one or more training examples 132 has a respective embedding vector stored in the embedding table 120. The embedding vector is a vector of numerical values, for example, a floating-point value or a quantized floating-point value having a fixed dimension (number of components).

[0040] During training, the embedding vector of a given speaker can be generated in various ways.

[0041] As an example, the embedding vector can be generated based on one or more characteristics of the speaker.

[0042] As a specific example, the computer system can use an embedding neural network, for example, a trained convolutional neural network, to generate a face embedding vector of the speaker's face that can be used to distinguish people or to generate a face embedding that reflects other properties of people's faces, for example, to process one or more images of the speaker's face cropped from the speaker's video in a corresponding training example. For example, the computer system can generate a respective image embedding vector for each image and then use an embedding neural network to process multiple images of the speaker's face to combine, for example, average, the image embedding vectors to generate a face embedding vector of the speaker's face.

[0043] As another specific example, the computer system can measure specific properties of the appearance of the speaker while speaking and map each measured property to a respective property embedding vector using, for example, a predefined mapping. An example of such a property is the frequency of opening the mouth while speaking. Another example of such a property is the maximum opening of the mouth while speaking. Yet another example of such a property is the average opening of the mouth while speaking.

[0044] When the system generates respective embedding vectors for each of a plurality of speaker characteristics, such as a face embedding vector and one or more respective property embedding vectors, the system can combine, e.g., average, sum, or concatenate, the respective embedding vectors of the plurality of characteristics to generate a speaker embedding vector.

[0045] As another example, the system can randomly initialize each speaker embedding (embedding vector) in the embedding table 120 and then update the speaker embeddings in conjunction with the training of the neural network 110a.

[0046] In each iteration of training, the system samples a mini-batch of one or more training examples 132 and, for each training example 132, uses the neural network 110a to process the respective speaker's video 140 in the training example and the corresponding speaker's embedding vector from the embedding table 120 to generate a predicted speech recognition output, e.g., a probability distribution over a set of text elements such as characters, phonemes, or word pieces of the training example 132.

[0047] Next, the system uses gradient-based techniques, such as stochastic gradient descent, Adam, or rmsProp, to train the neural network 110a to minimize a loss function that measures the respective error between the ground truth transcription 150 of the audio of the training example 132 and the predicted speech recognition output of the training example 132 for each training example 132 within the mini-batch. For example, the loss function can be a connectionist temporal classification (CTC) loss function. The CTC loss is described in more detail by Alex Graves, Santiago Fernandez, Faustino Gomez, and Jurgen Schmidhuber in Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks, International Conference on Machine Learning pp. 369 - 376, 2006.

[0048] In some implementations, the system also updates the embedding vectors of the speaker videos within the mini-batch, for example, by backpropagating the gradients through the neural network 110a to appropriate embedding vectors. More specifically, when the embedding vectors within the embedding table 120 are randomly initialized, the system also updates the embedding vectors. When the embedding vectors within the embedding table 120 are generated based on the characteristics or properties of the corresponding speaker, in some implementations, the system keeps the embedding vectors fixed during training, while in other implementations, the system fine-tunes the embedding vectors by updating the embedding vectors within the embedding table 120 in conjunction with the training of the neural network 110a.

[0049] After training, to generate a transcription of a new speaker video, the system (or another system) can process the new speaker video and the new embedding of the speaker as inputs to generate a predicted speech recognition output for the new speaker video. Optionally, the system can then apply a decoder, such as a beam search decoder or a finite state transducer (FST)-based decoder, to the predicted speech recognition output to map the speech recognition output to a sequence of words.

[0050] However, during training, embeddings are generated using the characteristics of the speaker, but these characteristics are usually not available for a new speaker. Thus, the system can adapt the trained neural network 110a before using the trained neural network 110a to generate a transcription for a new speaker.

[0051] FIG. 2 is a diagram showing an exemplary architecture 200 for adapting an adaptive visual speech recognition model to a new individual speaker. During the adaptation process, the embedding of the new speaker is adjusted so that the neural network 110a adapts to the characteristics of a particular individual. In other words, the purpose of the training process shown in FIG. 1 is to learn in advance. During adaptation, this prior information is combined with new data to quickly adapt to the characteristics of the new speaker.

[0052] Typically, the training process shown in FIG. 1 is executed on a distributed computing system having a plurality of computers. And, as described above, the adaptation process can be executed on hardware with a much lower computational cost, such as a desktop computer, a laptop computer, or a mobile computing device. For convenience, the adaptation process will be described as being executed by a system of one or more computers.

[0053] The architecture 200 includes, for example, a trained version of the visual-audio recognition neural network 110b corresponding to the trained version of the visual-audio recognition neural network 110a trained using the process described above with reference to FIG. 1.

[0054] To adapt the model to a new individual speaker, the system uses a set of videos 230 (the "video segments") of the new individual speaker speaking and corresponding transcriptions 240 of the text spoken in each video, represented as adaptation data 220. For example, each video 230 can be a video of the speaker speaking text written on a text prompt or text presented to the user on a user device.

[0055] Generally, the adaptation data 220 used in the adaptation process may be orders of magnitude smaller than the training data 130 used in the training process. In some implementations, the training data 130 includes video recordings of multiple hours for each of a plurality of different individual speakers, while the adaptation data 220 can use video recordings of less than 10 minutes of the new individual speaker.

[0056] Furthermore, the adaptation process generally requires significantly less computational effort than the training process. Thus, as shown above, in some implementations, the training process is executed in a data center with dozens, hundreds, or thousands of computers, while the adaptation process is executed on a mobile device or a single Internet-enabled device.

[0057] To initiate the adaptation phase, the system can initialize a new embedding vector 210 for the new speaker. Generally, the new embedding vector 210 may be different from any of the embedding vectors used during the training process.

[0058] For example, the system can initialize a new embedding vector 210 randomly or using available data that characterizes the new speaker. In particular, the system can initialize the new embedding vector 210 randomly or from the adaptation data 220 using one of the above-described techniques for generating speaker embeddings within the table 120. Even if any of the above techniques for generating embeddings using speaker characteristics are used, the adaptation data 220 generally has less data than the data available in the training data for any given speaker, so the newly generated embedding vector 210 generally has less information about the speaker's voice than the speaker embedding vectors used during training.

[0059] The adaptation process can be performed in multiple ways. In particular, the system can use non-parametric or parametric techniques.

[0060] Non-parametric techniques include using the adaptation data 220 to adapt the new speaker embedding 210 and optionally the model parameters of the neural network 110b, or both.

[0061] In particular, when performing non-parametric techniques, for each iteration of the adaptation phase, the system processes one or more video segments 230 within the adaptation data 220 and the current embedding vector 210 of the new speaker using the neural network 110b to generate a predicted speech recognition output for each video segment 230.

[0062] Next, the system updates the embedding vector 210 by backpropagating the gradient of the loss, e.g., the CTC loss, between the ground truth transcription 240 of the video segment 230 and the predicted speech recognition output of the video segment 230 through the neural network 110b to calculate the gradient with respect to the embedding vector 210, and then updates the embedding vector 210 using an update rule such as the stochastic gradient descent update rule, the Adam update rule, or the rmsProp update rule.

[0063] In some of these cases, the system fixes and holds the values of the model parameters of the neural network 110b during the adaptation phase. In other cases, the system also updates the model parameters at each iteration of the adaptation phase by using gradient-based techniques to update the model parameters using, for example, the same loss used to update the embedding vector 210.

[0064] Alternatively, the system can use parametric techniques involving training an auxiliary network to predict the embedding vector of a new speaker using a set of demonstration data different from that in the training data used to train the neural network 110b, e.g., a set of videos. Then, when the speaker adaptation data 220 is provided, the trained auxiliary neural network can be used to predict the embedding vector 210 of the new speaker.

[0065] Figure 3 is a flowchart of an exemplary process 300 for generating and using an adaptive visual speech recognition model. As described above, the process includes three stages: training, adaptation, and inference.

[0066] Typically, the training stage is executed on a distributed computing system having multiple computers.

[0067] And, as described above, the other two stages can be executed on hardware with a much lower computational cost, such as a desktop computer, a laptop computer, or a mobile computing device.

[0068] For convenience, exemplary process 300 is described as being executed by a system of one or more computers, but it will be understood that different steps of process 300 can be executed by different computing devices having different hardware functions.

[0069] The system generates an adaptive visual-audio recognition model (310) using training data representing videos of audio by a plurality of different individual speakers. As described above with reference to FIG. 1, the system can generate different embedding vectors for a plurality of individual speakers. Next, the system can train the parameter values of the neural visual-audio recognition model using training data including text and video data representing a plurality of different individual speakers who are speaking a part of the text. Each of the embedding vectors generally represents the respective characteristics of one of the plurality of different individual speakers.

[0070] The system adapts the adaptive visual-audio recognition model to a new individual speaker using adaptation data representing a video of audio spoken by the new individual speaker (320). As described above with reference to FIG. 2, the adaptation process uses video data representing a new individual speaker who is speaking a part of the text.

[0071] During the adaptation phase, the system can generate a new embedding vector for the new speaker using the adaptation data and optionally fine-tune the model parameters of the trained visual-audio recognition neural network.

[0072] After adaptation, the system executes an inference process (330) to convert the video of the new speaker and the new speaker's embedding into a transcription of the text being spoken within the video. Generally, the system uses a visual speech recognition model adapted to the new individual speaker, which includes using the new embedding vector of the individual speaker and the new video determined during the adaptation phase as inputs. As described above, the system can generate a transcription as a sequence of words and, in some cases, as punctuation by applying a decoder to the speech recognition output generated by the adapted visual speech recognition model. Executing the inference can also include one or more of displaying the transcription on a user interface, translating the transcription into another language, or providing audio data representative of the verbalization of the transcription for playback on one or more audio devices.

[0073] Although the above description is of adaptive visual speech recognition, the techniques described can also be applied to generate an adaptive audio-visual speech recognition model, where the inputs to the model are a video sequence of the speaker speaking, the corresponding audio sequence of the speaker speaking (although these two may not be temporally aligned), and the speaker's embedding, and the output is a transcription of the text being spoken in the audio-video pair. Examples of audio-visual speech recognition models that can be modified to accept the embedding and adapted as described are those described in Recurrent Neural Network Transducer for Audio-Visual Speech Recognition by Makino et al. in arXiv preprint arXiv:1911.04890 (2019), available at arxiv.org. When the input also includes audio data, the speaker embedding can also be generated by processing the audio using, for example, an audio embedding neural network trained to generate an embedding that uniquely identifies the speaker, either using the audio data or instead of it.

[0074] The proposed adaptive visual-audio recognition model, or audio-visual-audio recognition model, for example, can be useful for a user with hearing impairment to generate text representing the speech of a new speaker so that the user can read it. In another example, the user can be a new speaker, and the visual-audio recognition model, or audio-visual-audio recognition model, may be used in a dictation system to generate text, or may be used in a control system to generate text commands implemented by another system such as an electronic system or an electromechanical system. The adaptive visual-audio recognition model can be implemented to execute steps 320 and / or 330 by a computer system comprising at least one video camera for capturing a video of the new speaker processed in steps 320 and / or 330. In the case of an audio-visual-audio recognition model, the video camera may comprise a microphone for capturing an audio track associated with the captured video.

[0075] In situations where the systems described herein use data that may contain personal information, the data may be processed in one or more ways, such as aggregation and anonymization, before being stored or used so that such personal information cannot be determined from the data being stored or used. Additionally, such information may be used such that no information that can identify an individual is determined from the output of the system using such information.

[0076] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more combinations of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded in an artificially generated propagated signal, i.e., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiver device for execution by a data processing apparatus.

[0077] The term “data processing apparatus” refers to data processing hardware and includes any kind of apparatus, device, and machine for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include code that creates an execution environment for computer programs, e.g., processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0078] A computer program, which may also be referred to as or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language, including compiler-type or interpreter-type languages, declarative or procedural languages, and can be deployed in any form, such as as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not correspond to a file in the file system. The program can be stored in other programs or data, such as part of a file that holds one or more scripts stored in a markup language document, a single file dedicated to the program in question, or multiple coordinated files, such as files that hold one or more modules, subprograms, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.

[0079] This specification uses the term "configured" in relation to system and computer program components. One or more computer systems configured to perform a particular operation or action means that the system has software, firmware, hardware, or a combination thereof installed that causes the system to perform the operation or action during operation. One or more computer programs configured to perform a particular operation or action means that the one or more programs include instructions that cause an apparatus to perform the operation or action when executed by a data processing apparatus.

[0080] As used herein, the term "database" is widely used to refer to any collection of data, where the data need not be structured in any particular way, or structured at all, and can be stored on storage devices in one or more locations. Thus, for example, an indexed database can include multiple collections of data, each of which can be organized and accessed in a different way.

[0081] Similarly, as used herein, the term "engine" is widely used to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Typically, an engine is implemented as one or more software modules or components and installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines can be installed and run on the same computer.

[0082] The processes and logical flows described herein can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data to generate output. The processes and logical flows can also be performed by, for example, a special-purpose logic circuit such as an FPGA or ASIC, or by a combination of special-purpose logic circuits and one or more programmed computers.

[0083] A computer suitable for the execution of a computer program can be based on a general-purpose microprocessor or a special-purpose microprocessor or both, or other types of central processing units. Generally, the central processing unit receives instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a central processing unit for executing or performing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented or incorporated by dedicated logic circuits. Generally, a computer also includes or is operatively coupled to one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, to receive data therefrom, to transfer data thereto, or to do both. However, a computer need not be equipped with such devices. Further, a computer can be incorporated into another device, such as, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive.

[0084] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD ROM and DVD-ROM disks.

[0085] To provide interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, including acoustic, voice, or tactile input. Further, the computer can interact with the user by sending and receiving documents between the devices used by the user, for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser, or by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message from the user.

[0086] A data processing apparatus for implementing a machine learning model can also include, for example, a dedicated hardware accelerator unit for processing general and computationally intensive parts of machine learning training or production, such as inference, workloads.

[0087] The machine learning model can be implemented and deployed using, for example, a machine learning framework such as the TensorFlow framework.

[0088] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes, for example, backend components as a data server, or includes middleware components, such as an application server, or includes frontend components, such as a graphical user interface, a web browser, or a client computer equipped with an app with which a user can interact with an implementation of the subject matter described in this specification, or includes any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0089] A computing system can include clients and servers. Typically, clients and servers are remotely located and usually interact through a communication network. The relationship between a client and a server arises by computer programs that run on respective computers and have a client - server relationship to each other. In some embodiments, a server sends data, such as an HTML page, to a user device for purposes such as displaying data to a user interacting with a device that functions as a client and receiving user input from the user. Data generated at the user device, such as the result of a user interaction, can be received at the server from the device.

[0090] This specification includes details of many specific implementations, which should not be construed as limitations on the scope of the invention or the claims, but rather as descriptions of features specific to particular embodiments of a particular invention. The specific features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features are described above as acting in certain combinations and even claimed as such initially, one or more features from the claimed combination may in some cases be excluded from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0091] Similarly, operations are shown in the drawings in a particular order and recited in the claims, but this should not be understood as requiring that the operations be performed in the particular order or sequence shown, or that all of the operations shown be performed, to achieve a desirable result. In certain circumstances, multitasking and parallel processing may be advantageous. Further, the separation of the various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.

[0092] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes shown in the accompanying figures do not necessarily require the particular order or sequential order shown to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

Description of Reference Numerals

[0093] 100 Architecture 110a Visual-Audio Recognition Neural Network 110b Visual-Audio Recognition Neural Network 120 Embedding Table 130 Training Data 132 Training Example 140 Video 150 Ground Truth Transcription 200 Architecture 210 New Embedding Vector 210 Current Embedding Vector 220 Adaptive Data 230 Video 240 Transcription 300 Process

Claims

**Claim 1** A method executed by one or more computers, comprising: receiving a video including a plurality of video frames depicting a first speaker; obtaining a first embedding characterizing the first speaker; processing a first input comprising (i) the video and (ii) the first embedding using a visual-audio recognition neural network having a plurality of parameters; wherein the visual-audio recognition neural network is configured to process the video and the first embedding according to trained values of the parameters to generate an audio recognition output defining a sequence of one or more words spoken by the first speaker in the video. **Claim 2** The method of claim 1, wherein the visual-audio recognition neural network is configured to: generate an additional input channel from the first embedding; and combine the additional channel with one or more of the frames in the video before processing the frames in the video to generate the audio recognition output. **Claim 3** The method of claim 1 or 2, wherein the visual-audio recognition neural network comprises a plurality of hidden layers, and for at least one of the hidden layers, the visual-audio recognition neural network is configured to: generate an additional hidden channel from the first embedding; and combine the hidden channel with an output of the hidden layer before providing an output for processing by another hidden layer of the visual-audio recognition neural network. **Claim 4** The method according to any one of claims 1 to 3, further comprising: obtaining adaptation data for the first speaker, the adaptation data comprising one or more videos of the first speaker and a ground truth transcription for each of the videos; and determining the first embedding for the first speaker using the adaptation data. **Claim 5** ​ ​ Further comprising the step of obtaining a pre-trained value of the model parameters determined by training the visual-audio recognition neural network on training data comprising training examples corresponding to a plurality of speakers different from the first speaker, wherein the step of determining the first embedding comprises the step of determining the first embedding using the pre-trained value and the adaptation data, the method according to claim 4.

6. The step of determining the first embedding comprises comprising the step of initializing the first embedding, the step of updating the first embedding by repeatedly performing an operation comprises processing each of the one or more video segments in the adaptation data and the first embedding using the visual-audio recognition neural network according to the current value of the parameters to generate respective audio recognition outputs for each of the one or more video segments; updating the first embedding to minimize a loss function that measures a respective error between the ground truth transcription of the video segment and the respective audio recognition output of the video segment for each of the one or more video segments The method according to claim 5, comprising.

7. The step of updating the first embedding to minimize a loss function that measures a respective error between the ground truth transcription of the video segment and the respective audio recognition output of the video segment for each of the one or more video segments comprises backpropagating the gradient of the loss function through the visual-audio recognition neural network to determine the gradient of the loss function with respect to the first embedding; updating the first embedding using the gradient of the loss function with respect to the first embedding The method according to claim 6, comprising.

8. The method according to claim 6 or claim 7, wherein the current value is equal to the pre-trained value and the trained value, and the model parameters are fixed when determining the first embedding.

9. The operation is Further comprising the step of updating the current value of the parameters of the visual-audio recognition neural network based on the gradient of the loss function with respect to the parameters of the visual-audio recognition neural network, wherein the trained value is equal to the current value after determining the first embedding vector, the method according to claim 6 or claim 7.

10. The method according to any one of claims 1 to 9, further comprising the step of applying a decoder to the speech recognition output of the video to generate the sequence of one or more words spoken by the first speaker in the video.

11. The method according to any one of claims 1 to 10, wherein the speech recognition output comprises a respective probability distribution over the vocabulary of text elements for each of the video frames.

12. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the respective methods according to any one of claims 1 to 11.

13. One or more computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the respective methods according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Imaging apparatus, imaging method and program

    JP2011014985A

  • Learning apparatus, learning method, program, learnt model and lip reading apparatus

    JP2019204147A

  • Face recognition device, learning device and program

    JP2021009571A

  • Lip reading device and lip reading method

    JP2021086274A

  • Joint neural network for speaker recognition

    US20190341058A1