Contrastive Learning and Mask Modeling for End-to-End Self-Supervised Pre-Training

The end-to-end self-supervised pre-training framework using contrastive and masked modeling loss terms enhances model performance by optimizing unlabeled data utilization, achieving state-of-the-art results in tasks such as automatic speech recognition and reducing fine-tuning requirements.

JP7711305B2Active Publication Date: 2025-07-22GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024505150
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-28
Filing Date
2022-07-28
Publication Date
2025-07-22
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

Existing machine learning models struggle to effectively utilize unlabeled data for self-supervised pre-training, leading to suboptimal performance in various tasks.

Method used

An end-to-end self-supervised pre-training framework combining contrastive loss and masked modeling loss terms to optimize models by discretizing input data and learning contextual representations through masked prediction tasks.

Benefits of technology

The framework enables state-of-the-art performance on tasks like automatic speech recognition and improves real-world recognition tasks, reducing the need for extensive fine-tuning and computational resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711305000003
    Figure 0007711305000003
  • Figure 0007711305000004
    Figure 0007711305000004
  • Figure 0007711305000005
    Figure 0007711305000005
Patent Text Reader

Abstract

An improved end-to-end self-supervised pre-training framework is provided that leverages a combination of contrastive loss and mask modeling loss terms. In particular, the present disclosure provides a framework that combines contrastive learning and mask modeling, where the former trains a model to discretize input data (e.g., a continuous signal such as a continuous speech signal) into a finite set of differentiable tokens, and the latter trains a model to learn contextualized representations by solving a masked prediction task that consumes the discretized tokens. In contrast to some existing mask modeling-based pre-training frameworks that rely on an iterative reclustering and retraining process or other existing frameworks that concatenate two separate trained modules, the proposed framework can enable a model to be optimized in an end-to-end manner by simultaneously solving two self-supervised tasks (contrast task and mask modeling).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to machine learning. More specifically, the present disclosure relates to an improved end-to-end self-supervised pre-training framework that utilizes a combination of a contrastive loss term and a mask modeling loss term.

Background Art

[0002] To improve the performance of machine learning models in various tasks, the development of techniques for leveraging large-scale unannotated data has been an ongoing research topic for many years. So far, there have been two main approaches for using unlabeled data to tackle such semi-supervised tasks.

[0003] The first line of work is self-training, also known as pseudo-labeling, where the system first starts training a teacher model using the available labeled data. The teacher model is then used to label the unlabeled data. The combination of the labeled data and the pseudo-labeled data is then used to train a student model. The pseudo-labeling process may be repeated multiple times to improve the quality of the teacher model. Self-training has been a well-studied technique that is actually useful for several different tasks and domains.

[0004] A second direction that utilizes unlabeled data is unsupervised pre-training or self-supervised pre-training. In unsupervised pre-training, the model is first trained (hence the name "unsupervised") to complete proxy tasks that are designed to consume only unlabeled data. Such proxy tasks are generally thought to be able to initialize the model's parameters at an effective starting point before being trained with labeled data. Important recent research efforts have been made to develop proxy tasks that perform well for the model when it is fine-tuned on some downstream tasks. Studies have also been conducted to show that the gains brought about by self-training and unsupervised pre-training are additional for some downstream tasks.

Prior Art Documents

Non-Patent Documents

[0005]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Means for Solving the Problems

[0006] Aspects and advantages of embodiments of the present disclosure will be described in part in the following description, or will be apparent from the description, or will be learned through the practice of the embodiments.

[0007] One example described in this disclosure is directed to a computer-implemented method for performing end-to-end self-supervised pre-training. The method includes obtaining, by a computing system comprising one or more computing devices, a series of input data. The method includes processing, by the computing system, the series of input data using a first encoder portion of a machine learning model to generate a plurality of encoded features. The method includes quantizing, by the computing system, the plurality of encoded features to generate a plurality of target quantization vectors and a plurality of discretization identifiers associated with the plurality of target quantization vectors. The method includes masking, by the computing system, one or more of the plurality of encoded features. After the masking step, the method includes processing, by the computing system, the plurality of encoded features using a second encoder portion of the machine learning model to generate a first set of context vectors. The method includes processing, by the computing system, the first set of context vectors using a third encoder portion of the machine learning model to generate a second set of context vectors. The method includes evaluating, by the computing system, a loss function that includes a contrastive loss term and a mask modeling loss term, the contrastive loss term evaluating a contrastive pre-training output generated based on the first set of context vectors and the plurality of target quantization vectors, and the mask modeling loss term evaluating a mask modeling pre-training output generated based on the second set of context vectors and the plurality of discretization identifiers. The method includes training, by the computing system, the machine learning model end-to-end based on the loss function.

[0008] For each of the one or more masked positions, the control pre-training output may include a predicted selection from a set of candidate vectors, the predicted selection being generated based on one of a first set of context vectors corresponding to the masked position. The set of candidate vectors may include the true target quantization vector and one or more scrambled vectors. The control loss term may evaluate whether the predicted selection corresponds to the true target quantization vector.

[0009] For each of the one or more masked positions, the mask modeling pre-training output may include a predicted identifier generated based on one of a second set of context vectors corresponding to the masked position, and the mask modeling loss term may evaluate whether the predicted identifier corresponds to the true discretization identifier among a plurality of discretization identifiers corresponding to the masked position.

[0010] The second encoder portion of the machine learning model may include one or more conformer blocks. Similarly, the encoder portion may include one or more conformer blocks.

[0011] The step of training the machine learning model based on the loss function by the computing system may include the step of modifying, by the computing system, one or more values of one or more parameters of the third encoder portion, the second encoder portion, and the first encoder portion of the machine learning model based on the mask modeling loss term. The step of training the machine learning model may further include the step of modifying, by the computing system, one or more values of one or more parameters of the second encoder portion and the first encoder portion of the machine learning model based on a combination of the mask modeling loss term and the control loss term.

[0012] The method may include the step of modifying a codebook used to perform quantization based on the loss function.

[0013] A series of input data may include audio data or a spectral representation of audio data. For example, the audio data may include voice data. For example, the machine learning model may be a model for performing voice-related tasks such as speech recognition and / or voice conversion. The series of input data may include, additionally or alternatively, text data, sensor data, and / or image data.

[0014] Another example described in this disclosure is directed to one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations. For example, the operations may include operations for performing any of the methods described herein. For example, the operations may include obtaining task-specific training inputs. The operations include processing the task-specific training inputs using a machine learning model to generate task-specific training outputs, wherein at least an encoder portion of the machine learning model is end-to-end trained using a loss function that includes a contrastive loss term and a masking modeling loss term, the contrastive loss term evaluating a contrastive pre-training output generated based on a first set of context vectors and a plurality of target quantization vectors, the first set of context vectors being generated by an encoder portion of the machine learning model after masking of an input or intermediate output of the encoder portion of the machine learning model, the plurality of target quantization vectors being generated by quantization of an input or intermediate output of the encoder portion of the machine learning model, the masking modeling loss term evaluating a masking modeling pre-training output generated based on a second set of context vectors and a plurality of discretization identifiers, the second set of context vectors being generated by the encoder portion of the machine learning model from the first set of context vectors, the plurality of discretization identifiers being generated by quantization of an input or intermediate output of the encoder portion of the machine learning model. The operations include evaluating a task-specific loss function based on the task-specific training outputs. The operations include training the machine learning model based on the task-specific loss function.

[0015] Another exemplary aspect of the present disclosure is directed to a computing system. The computing system includes one or more processors and one or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more processors of the computing system, cause the computing system to perform operations. The operations may include any of the methods described herein. For example, the operations may include obtaining task-specific inference inputs. The operations include processing the task-specific inference inputs using a machine learning model to generate task-specific inference outputs, wherein at least an encoder portion of the machine learning model is end-to-end trained using a loss function that includes a contrastive loss term and a masking modeling loss term, the contrastive loss term evaluating a contrastive pre-training output generated based on a first set of context vectors and a plurality of target quantization vectors, the first set of context vectors being generated by an encoder portion of the machine learning model after masking of an input or intermediate output of the encoder portion of the machine learning model, the plurality of target quantization vectors being generated by quantization of an input or intermediate output of the encoder portion of the machine learning model, the masking modeling loss term evaluating a masking modeling pre-training output generated based on a second set of context vectors and a plurality of discretization identifiers, the second set of context vectors being generated from the first set of context vectors by an encoder portion of the machine learning model, the plurality of discretization identifiers being generated by quantization of an input or intermediate output of the encoder portion of the machine learning model. The operations include providing the task-specific inference outputs as an output.

[0016] Other examples described in the present disclosure are directed to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.

[0017] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the relevant principles.

[0018] A detailed description of embodiments directed to those skilled in the art is set forth herein with reference to the accompanying drawings.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Modes for Carrying Out the Invention

[0020] Reference numbers repeated across multiple figures identify the same features in various implementations.

[0021] Overview Generally, the present disclosure is directed to an improved end-to-end self-supervised pre-training framework that leverages a combination of a contrastive loss term and a masked modeling loss term. In particular, the present disclosure provides a framework that combines contrastive learning and masked modeling, where the former trains a model to discretize input data (e.g., a continuous signal such as a continuous audio signal) into a finite set of differentiable tokens, and the latter trains the model to learn contextualized representations by solving a masked prediction task that consumes the discretized tokens. In contrast to some existing masked modeling-based pre-training frameworks that rely on iterative reclustering and retraining processes or other existing frameworks that concatenate two separately trained modules, the proposed framework enables the model to be optimized in an end-to-end manner by simultaneously solving two self-supervised tasks (the contrastive task and the masked modeling).

[0022] More specifically, an exemplary aspect of the present disclosure focuses on improving the ability to perform unsupervised pre-training by proposing a novel pre-training framework. An exemplary implementation of the proposed framework uses a contrastive pre-training task to obtain a catalog of a finite set of differentiable, discretized speech units, and then uses those speech units as targets in a masked prediction task. The masked prediction task requires the model to consume tokens that will be learned by first solving the contrastive task, but the present disclosure demonstrates that the two objectives can actually be optimized simultaneously.

[0023] The pre-training framework described in this specification can be applied to many different tasks, domains, and / or data modalities. One specific exemplary task is automatic speech recognition. Another exemplary task is voice conversion. Thus, the input data can include audio data such as speech data (e.g., represented in raw form or using spectrograms). In other examples, the input data can include other forms of data such as text data (e.g., natural language data), sensor data, image data, biological or chemical data, and / or other forms of data. Once pre-trained, the model can be fine-tuned to perform any number of different tasks.

[0024] The systems and methods of the present disclosure provide several technical effects and benefits. As one exemplary technical advancement, the present disclosure provides a pre-training framework that simultaneously and directly optimizes contrastive loss and masked prediction loss for end-to-end self-supervised representation learning. The pre-training framework can elicit state-of-the-art performance on various tasks. For example, the framework has been shown to elicit state-of-the-art performance on the well-benchmarked LibriSpeech task and to improve the performance of real-world recognition tasks (voice search) significantly over existing state-of-the-art methods. Thus, the pre-training framework described in this specification can enable improved model performance on several different tasks, which corresponds to an improvement in the computing system itself.

[0025] As another exemplary technical effect, the improved pre-training framework described herein can lead to better pre-trained models that can be more quickly or easily fine-tuned for various downstream tasks. That is, by providing an improved pre-trained model as a potential checkpoint from which to start for a given task, fewer fine-tuning operations are required to be performed to achieve comparable performance. As a result, less fine-tuning training, corresponding to savings in computational resources such as processor usage, memory usage, network bandwidth, etc., may need to be performed. Similarly, the improved pre-trained model may enable fine-tuning of models for tasks where only a limited amount of fine-tuning training data is available. Thus, the proposed framework can enable the application of machine learning techniques to various domains or tasks that were previously not available.

[0026] Reference is now made to the drawings, in which exemplary embodiments of the disclosure are discussed in more detail.

[0027] Exemplary Pre-Training Framework FIG. 1 shows a block diagram of an exemplary pre-training framework that can be used to pre-train a machine learning model 14. The machine learning model 14 can include a first encoder portion 16, a second encoder portion 18, and a third encoder portion 20. The exemplary architecture shown in FIG. 1 is provided by way of example only. The model 14 and its portions 16, 18, and 20 can have various architectures that are the same as or different from those shown in FIG. 1.

[0028] The pre-training process can be performed on a series of input data 22. As an example, the input data 22 may be samples from a continuous signal. By way of example, the input data 22 may be audio data (e.g., raw audio data or audio data represented using a spectrogram) (e.g., the audio data may be voice data), text data, image data, sensor data, biological or chemical data, and / or combinations thereof. The input data 22 can be formatted as a sequence having several positions (e.g., positions 1, 2, 3, ..., j).

[0029] The first encoder portion 16 of the machine learning model 14 can process the input data 22 to generate a plurality of encoded features 24. In one example, as shown in FIG. 1, the first encoder portion 16 may be, for example, a convolutional subsampling block including two 2D convolutional layers both having a stride of (2, 2), such that as a result, the input sequence length is reduced by a factor of 4. For example, when a log-mel spectrogram is provided as input, the first encoder portion 16 can extract a latent representation that is taken as input by the second encoder portion 18.

[0030] The quantization technique 26 can be performed on the plurality of encoded features 24 to generate a plurality of target quantization vectors 28 and a plurality of discretization identifiers 30 associated with the plurality of target quantization vectors 28.

[0031] As an example, in some implementations, quantization technique 26 may include performing direct product quantization. Direct product quantization may include selecting quantization representations from multiple codebooks and concatenating them. Given some codebooks, or groups, each having some entries, quantization technique 26 may include selecting one entry from each codebook, concatenating the resulting vectors, and then applying a linear transformation to obtain quantization vector 28. In some implementations, the use of Gumbel softmax may enable selecting individual codebook entries in a fully differentiable manner. In some implementations, a straight-through estimator can be used and G Gumbel softmax operations can be set up. The feature encoder output can be mapped to some logits corresponding to different codebook entries. In the backward pass, the true gradient of the Gumbel softmax output can be used.

[0032] Referring back to FIG. 1, one or more of the plurality of encoded features 24 can be masked at 32. For example, masking 32 can be performed at one or more of some positions associated with the input data 22, and the positions where masking 32 is performed can then be referred to as "masked positions". In one example, masking 32 may include setting the feature values equal to zero. In another example, masking 32 may include changing the feature values to some other value (e.g., a random noise value).

[0033] After the masking 32, the second encoder portion 18 of the machine learning model 14 can process the plurality of encoded features 24 to generate a first set of context vectors 34. In one example, the second encoder portion 18 can include a linear projection layer followed by a stack of conformer blocks. See Gulati et al., Conformer: Convolution-augmented transformer for speech recognition, Interspeech, 2020. Each of the conformer blocks can include a series of multi-head self-attention (Vaswani et al., Attention is all you need, NIPS, 2017), depthwise convolution, and feed-forward layers.

[0034] In the framework shown in the figure, one goal of the second encoder portion 18 is to discretize the (masked) encoded features 24 into a finite set of representative units. For this purpose, the second encoder portion 18 can interact with the quantization mechanism 26. Specifically, the encoded features 24 output by the first encoder portion 16 can, on the one hand, after masking, be fed to a linear projection layer followed by a stack of conformer blocks to yield a first set of context vectors 34. On the other hand, the encoded features 24 can be passed to the quantizer 26 without masking to yield quantization vectors 28 and the token IDs 30 assigned to them. The quantization vectors 28 can be used together with the first set of context vectors 34 corresponding to the masked positions to resolve the contrast task. The assigned token IDs 30 can be used later as prediction targets by subsequent masked prediction scenarios.

[0035] Specifically, still referring to FIG. 1, the third encoder portion 20 of the machine learning model 14 can process the first set of context vectors 34 to generate a second set of context vectors 36. As an example, as shown in FIG. 1, the third encoder portion 20 can include a stack of conformable blocks, each block having the same configuration as that from the second encoder portion 18. The third encoder portion 20 can directly take in the first set of context vectors 34 and extract highly contextualized representations.

[0036] The pre-training process can include evaluating a loss function that includes both a contrastive loss term 38 and a mask modeling loss term 40. Specifically, the contrastive loss term 38 can evaluate a contrastive pre-training output generated based on the first set of context vectors 34 and a plurality of target quantization vectors 28.

[0037] Specifically, in one example, for the context vector c corresponding to the masked time step (position) t t the model 14 (including the auxiliary pre-training prediction component) is relied upon to identify its true quantization vector q t from a set of K confounding elements

[0038]

Number

[0039] For example, the confounding elements can be quantization vectors uniformly sampled from other masked positions in the general input set (e.g., the utterance for voice input). This part of the loss can be denoted as L w and can be denoted as such.

[0040] In one example, L w is

[0041]

Number

[0042] becomes as follows, and in the above formula, sim(a,b)=a T b / ||a|| ||b|| is the cosine similarity between the context vector and the quantization vector.

[0043] The above-mentioned loss L w can be further augmented by the codebook diversity loss L d so as to encourage the uniform use of the code. Therefore, one exemplary final contrast loss can be defined as follows. L c =L w +α·L d

[0044] In one example, α = 0.1. However, other values may be used.

[0045] Therefore, in some implementations, for each of one or more masked positions, the contrast pre-training output can include a predicted selection from a set of candidate vectors, the predicted selection being generated based on one of a first set of context vectors 34 corresponding to the masked position. Further, the set of candidate vectors may include the true target quantization vector 28 and one or more perturbation vectors, and the contrast loss term 38 can evaluate whether the predicted selection corresponds to the true target quantization vector 28.

[0046] The contrast loss 38 can be used to train the first and second encoder portions 16 and 18 together with the quantizer 26, whereby the first and second encoder portions 16, 18 produce a valid context vector 34 that is taken as input by the third encoder portion 20, and the quantizer 26 produces a differentiable discretization token that is used as a target by the third encoder portion 20.

[0047] In particular, the mask modeling loss term 40 can evaluate the mask modeling pre-training output generated based on the second set of context vectors 36 and the plurality of discretized identifiers 30.

[0048] In particular, in one example, a softmax layer is added on top of the third encoder portion 20. When the context vector 36 in the final layer corresponds to the masked position, the softmax layer takes the context vector 36 as input and attempts to predict the corresponding token ID 30, which has been previously assigned by the quantizer 26. An exemplary cross-entropy loss for this masked prediction task is L m which can be denoted as.

[0049] Thus, in some implementations, for each of one or more masked positions, the mask modeling pre-training output can include a predicted identifier generated based on one of the second set of context vectors 36 corresponding to the masked position. The mask modeling loss term 40 can evaluate whether the predicted identifier corresponds to the true discretized identifier of the plurality of discretized identifiers 30 corresponding to the masked position.

[0050] The machine learning model 14 can be trained end-to-end based on a loss function that includes both the contrastive loss term 38 and the mask modeling loss term 40. Thus, the model 14 may be trained to resolve two self-supervised tasks at the same time. One exemplary final training loss to be minimized can be as follows. L p = β · L c + γ · L m

[0051] In some examples, both β and γ may be set equal to 1. However, other values may be used.

[0052] Exemplary fine-tuning techniques Figure 2 shows a block diagram of an exemplary fine-tuning training method that can be used to train the machine learning model 200. The model 200 may include a pre-trained encoder model 14 (e.g., pre-trained as shown in FIG. 1). For example, the model 14 may include encoder portions 16, 18, and 20 that are pre-trained with a loss function that includes both a contrastive loss 38 and a mask modeling loss 40, as shown in FIG. 1.

[0053] Referring now to FIG. 2, the model may also include a decoder portion 202. The training scheme shown in FIG. 2 can operate over a number of training examples. One training example 204 shown in FIG. 2 includes a task-specific training input 206 and a ground truth 210 (e.g., a label).

[0054] The training example 204 can be a specific example for any different task, domain, and / or data modality. Exemplary data modalities include text, audio, images, sensor data, and / or other forms of data. Exemplary tasks can include recognition tasks, translation tasks, detection tasks, synthesis tasks, prosody classification, sentiment or mood classification, and / or various other tasks that can be performed in any of the data modalities listed above. Two specific exemplary tasks include automatic speech recognition and voice conversion.

[0055] The pre-trained encoder 14 can process the input 206 to generate a context representation. The decoder can process the context representation to generate a task-specific training output 208.

[0056] The objective function 212 can simply compare the task-specific training output 208 with the ground truth 210. The model 200 can be trained based on the objective function 212 (e.g., by backpropagation of the objective function through the decoder 202 and / or the pre-trained encoder 14).

[0057] Exemplary inference method FIG. 3 shows a block diagram of an exemplary inference method that can be used after training the machine learning model 200. In particular, the model 200 can include a pre-trained encoder model 14 (e.g., pre-trained as shown in FIG. 1 and fine-tuned as shown in FIG. 2) and a decoder portion 202 (e.g., fine-tuned as shown in FIG. 2).

[0058] Referring now to FIG. 3, the inference method shown in FIG. 3 can operate over a number of inference inputs. One task-specific inference input 302 is shown in FIG. 3. The inference input 302 can be a specific input for any different task, domain, and / or data modality. Exemplary data modalities include text, audio, images, sensor data, and / or other forms of data. Exemplary tasks can include recognition tasks, translation tasks, detection tasks, synthesis tasks, prosody classification, sentiment or mood classification, and / or various other tasks that can be performed in any of the data modalities listed above. Two specific exemplary tasks include automatic speech recognition and voice conversion.

[0059] The pre-trained encoder 14 can process the inference input 302 to generate a context representation. The decoder can process the context representation to generate a task-specific inference output 304.

[0060] Exemplary devices and systems FIG. 4A shows a block diagram of an exemplary computing system 100. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled via a network 180.

[0061] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0062] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0063] In some implementations, the user computing device 102 can store or include one or more machine learning models 120. For example, the machine learning model 120 can be various machine learning models, such as a neural network (e.g., a deep neural network) or other types of machine learning models including non-linear models and / or linear models, or alternatively, may include those machine learning models. The neural network can include a feed-forward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Some exemplary machine learning models can utilize attention mechanisms such as self-attention. For example, some exemplary machine learning models can include a multi-head self-attention model (e.g., a transformer model). The exemplary machine learning model 120 will be discussed with reference to FIGS. 1-3.

[0064] In some implementations, one or more machine learning models 120 are received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then can be used or otherwise implemented by one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single machine learning model 120 (e.g., to perform parallel predictions across multiple instances of the input).

[0065] Additionally or alternatively, one or more machine learning models 140 may be included in a server computing system 130 that communicates with the user computing device 102 according to a client - server relationship, or otherwise may be stored and implemented by the server computing system 130. For example, the machine learning model 140 may be implemented by the server computing system 140 as part of a web service. Thus, one or more models 120 may be stored and implemented on the user computing device 102, and / or one or more models 140 may be stored and implemented on the server computing system 130.

[0066] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch - sensitive component (such as a touch - sensitive display screen or a touchpad) that is sensitive to the touch of a user input object (such as a finger or a stylus). The touch - sensitive component can serve to implement a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means by which a user can provide user input.

[0067] Server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0068] In some implementations, the server computing system 130 includes or is implemented by one or more server computing devices. In cases where the server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0069] As described above, the server computing system 130 can store one or more machine learning models 140, or alternatively, can include the model 140. For example, the model 140 can be, or alternatively can include, various machine learning models. Exemplary machine learning models include neural networks or other multi-layer non-linear models. Exemplary neural networks include feed-forward neural networks, deep neural networks, regression neural networks, and convolutional neural networks. Some exemplary machine learning models can utilize attention mechanisms such as self-attention. For example, some exemplary machine learning models can include multi-head self-attention models (e.g., transformer models). The exemplary model 140 will be discussed with reference to FIGS. 1-3.

[0070] The user computing device 102 and / or the server computing system 130 can train the model 120 and / or 140 through interaction with a training computing system 150 communicatively coupled via the network 180. The training computing system 150 can be separate from the server computing system 130 or can be a part of the server computing system 130.

[0071] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 154 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices or, alternatively, is implemented by one or more server computing devices.

[0072] The training computing system 150 can include a model trainer 160 that trains a machine learning model 120 and / or 140 stored in the user computing device 102 and / or the server computing system 130 using various training or learning techniques such as, for example, backpropagation of error. For example, a loss function can be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters for a number of training iterations.

[0073] In some implementations, performing backpropagation of errors may include performing reduced backpropagation over time. The model trainer 160 can perform some generalization techniques (such as weight decay, dropout, etc.) to improve the generalization ability of the model being trained.

[0074] In particular, the model trainer 160 can train the machine learning models 120 and / or 140 based on a set of training data 162. In some implementations, when the user gives consent, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 with respect to user-specific data received from the user computing device 102. In some cases, this process may be referred to as model personalization.

[0075] The model trainer 160 includes computer logic used to provide the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.

[0076] Network 180 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication via Network 180 can be carried over any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0077] The machine learning models described herein may be used in a variety of tasks, applications, and / or use cases.

[0078] In some implementations, the input to the machine learning models of the present disclosure may be image data. The machine learning models may process the image data to generate an output. By way of example, the machine learning models may process the image data to generate an image recognition output (e.g., recognition of the image data, embedding of the image data's latent features, encoded representation of the image data, hash of the image data, etc.). As another example, the machine learning models may process the image data to generate an image segmentation output. As another example, the machine learning models may process the image data to generate an image classification output. As another example, the machine learning models may process the image data to generate an image data modification output (e.g., modification of the image data, etc.). As another example, the machine learning models may process the image data to generate an encoded image data output (e.g., encoded and / or compressed representation of the image data, etc.). As another example, the machine learning models may process the image data to generate an upscaled image data output. As another example, the machine learning models may process the image data to generate a prediction output.

[0079] In some implementations, the input to the machine learning model of the present disclosure may be text or natural language data. The machine learning model may process the text or natural language data to generate an output. By way of example, the machine learning model may process natural language data to generate a language encoding output. As another example, the machine learning model may process text or natural language data to generate a latent text embedding output. As another example, the machine learning model may process text or natural language data to generate a conversion output. As another example, the machine learning model may process text or natural language data to generate a classification output. As another example, the machine learning model may process text or natural language data to generate a text segmentation output. As another example, the machine learning model may process text or natural language data to generate a semantic intent output. As another example, the machine learning model may process text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is of higher quality than the input text or natural language, etc.). As another example, the machine learning model may process text or natural language data to generate a prediction output.

[0080] In some implementations, the input to the machine learning model of the present disclosure may be audio data. The machine learning model may process the audio data to generate an output. As an example, the machine learning model may process the audio data to generate an audio recognition output. As another example, the machine learning model may process the audio data to generate an audio conversion output. As another example, the machine learning model may process the audio data to generate a latent embedding output. As another example, the machine learning model may process the audio data to generate an encoded audio output (e.g., an encoded and / or compressed representation of the audio data, etc.). As another example, the machine learning model may process the audio data to generate an upscaled audio output (e.g., audio data of higher quality than the input audio data, etc.). As another example, the machine learning model may process the audio data to generate a text representation output (e.g., a text representation of the input audio data, etc.). As another example, the machine learning model may process the audio data to generate a prediction output.

[0081] In some implementations, the input to the machine learning model of the present disclosure may be latent encoded data (e.g., a latent space representation of the input, etc.). The machine learning model may process the latent encoded data to generate an output. As an example, the machine learning model may process the latent encoded data to generate a recognition output. As another example, the machine learning model may process the latent encoded data to generate a reconstruction output. As another example, the machine learning model may process the latent encoded data to generate a search output. As another example, the machine learning model may process the latent encoded data to generate a reclustering output. As another example, the machine learning model may process the latent encoded data to generate a prediction output.

[0082] In some implementations, the input to the machine learning model of the present disclosure may be statistical data. Statistical data can be data that is calculated and / or derived from some other data source, represents this, or otherwise includes this. The machine learning model can process the statistical data to generate an output. By way of example, the machine learning model can process the statistical data to generate a recognition output. As another example, the machine learning model can process the statistical data to generate a prediction output. As another example, the machine learning model can process the statistical data to generate a classification output. As another example, the machine learned model can process the statistical data to generate a segmentation output. As another example, the machine learning model can process the statistical data to generate a visualization output. As another example, the machine learning model can process the statistical data to generate a diagnostic output.

[0083] In some implementations, the input to the machine learning model of the present disclosure may be sensor data. The machine learning model can process the sensor data to generate an output. By way of example, the machine learning model can process the sensor data to generate a recognition output. As another example, the machine learning model can process the sensor data to generate a prediction output. As another example, the machine learning model can process the sensor data to generate a classification output. As another example, the machine learning model can process the sensor data to generate a segmentation output. As another example, the machine learning model can process the sensor data to generate a visualization output. As another example, the machine learning model can process the sensor data to generate a diagnostic output. As another example, the machine learning model can process the sensor data to generate a detection output.

[0084] In some cases, the machine learning model can be configured to perform tasks that include encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task may be an audio compression task. The input may include audio data and the output may include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task may include generating an embedding for the input data (e.g., input audio or visual data).

[0085] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data for one or more images and the task is an image processing task. For example, the image processing task may be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images show an object belonging to the object class. The image processing task may be object detection, where the image processing output identifies one or more regions in one or more images and, for each region, the likelihood that the region shows a target object. As another example, the image processing task may be image segmentation, where the image processing output defines, for each pixel in one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories may be foreground and background. As another example, the set of categories may be object classes. As another example, the image processing task may be depth estimation, where the image processing output defines a respective depth value for each pixel in one or more images. As another example, the image processing task may be motion estimation, where the network input includes multiple images and the image processing output defines, for each pixel in one of the input images, the motion of the scene shown at the pixel between the images in the network input.

[0086] In some cases, the input includes audio data representing speech, and the task is an automatic speech recognition task. The output may include a text output mapped to the speech. In some cases, the task includes encrypting or decrypting the input data. In some cases, the task includes microprocessor-implemented tasks such as branch prediction or memory address translation.

[0087] FIG. 4A shows one exemplary computing system that can be used to implement the present disclosure. Other computing systems may be used. For example, in some implementations, the user computing device 102 may include a model trainer 160 and a training dataset 162. In such implementations, the model 120 can be both trained locally and used on the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to customize the model 120 based on user-specific data.

[0088] FIG. 4B shows a block diagram of an exemplary computing device 10 that performs the operations described in the present disclosure. The computing device 10 may be a user computing device or a server computing device.

[0089] The computing device 10 includes several applications (e.g., applications 1 to N). Each application includes its own machine learning library and machine learning-based model. For example, each application may include a machine learning-based model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.

[0090] As shown in FIG. 4B, each application can communicate with some other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0091] FIG. 4C shows a block diagram of an exemplary computing device 50 that performs the operations described in this disclosure. The computing device 50 can be a user computing device or a server computing device.

[0092] The computing device 50 includes some applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).

[0093] The central intelligence layer includes several machine learning models. For example, as shown in FIG. 4C, each machine learning model can be provided to each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model to all applications. In some implementations, the central intelligence layer is included in or otherwise implemented by the operating system of the computing device 50.

[0094] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As shown in FIG. 4C, the central device data layer can communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0095] Additional disclosure The technology described in this specification refers to servers, databases, software applications, and other computer-based systems, as well as actions taken and information transmitted between such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functions among components. For example, the processes described in this specification can be implemented using a single device or component, or multiple devices or components operating in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0096] Although the subject matter has been described in detail with respect to various specific examples, each example is provided by way of illustration and not limitation of the disclosure. Those skilled in the art will, upon understanding the foregoing, readily be able to make modifications, variations, and equivalents of such examples. Accordingly, the disclosure is not intended to exclude such modifications, variations, and / or additional incorporations of the subject matter that will readily become apparent to those skilled in the art. For example, features illustrated or described as part of one example can be used with another example to yield a further example. Accordingly, the disclosure is intended to cover such modifications, variations, and equivalents.

Description of the Reference Numerals

[0097] 14 Machine learning model, model, pre-trained encoder model, pre-trained encoder 16 First encoder portion, encoder portion 18 Second encoder portion, encoder portion 20 Third encoder portion, encoder portion 22 Input data 24 Encoded features 26 Quantization technique, quantization mechanism, quantizer 28 Target quantization vector, quantization vector 30 Discretization identifier, token ID 32 Masking 34 Context Vector 36 Context Vector 38 Contrastive Loss Term 40 Mask Modeling Loss Term 100 Computing System 102 User Computing Device 112 Processor 114 Memory, Computing Device Memory 120 Machine Learning Model 122 User Input Component 130 Server Computing System 132 Processor 134 Memory 140 Machine Learning Model 150 Training Computing System 152 Processor 154 Memory 160 Model Trainer 162 Training Data 180 Network 200 Machine Learning Model, Model 202 Decoder Portion, Decoder

Claims

1. A computer-implemented method for performing self-supervised pre-training, comprising: obtaining, by a computing system comprising one or more computing devices, a series of input data; processing, by the computing system, the series of input data using a first encoder portion of a machine learning model to generate a plurality of encoded features; quantizing, by the computing system, the plurality of encoded features to generate a plurality of target quantization vectors and a plurality of discretization identifiers associated with the plurality of target quantization vectors; masking, by the computing system, one or more of the plurality of encoded features; after the masking step, processing, by the computing system, the plurality of encoded features using a second encoder portion of the machine learning model to generate a first set of context vectors; processing, by the computing system, the first set of context vectors using a third encoder portion of the machine learning model to generate a second set of context vectors; evaluating, by the computing system, a loss function comprising a contrastive loss term and a masking modeling loss term, wherein the contrastive loss term evaluates a contrastive pre-training output generated based on the first set of context vectors and the plurality of target quantization vectors, and the masking modeling loss term evaluates a masking modeling pre-training output generated based on the second set of context vectors and the plurality of discretization identifiers; training, by the computing system, the machine learning model based on the loss function.

2. For each of one or more masked positions, the contrastive pre-training output comprises a predicted selection from a set of candidate vectors, the predicted selection being generated based on one of the first set of context vectors corresponding to the masked position. The set of candidate vectors includes a true target quantization vector and one or more scrambling vectors, The computer-implemented method according to claim 1, wherein the contrast loss term evaluates whether the prediction selection corresponds to the true target quantization vector.

3. For each of the one or more masked positions, The mask modeling pre-training output includes a prediction identifier generated based on one of the second set of context vectors corresponding to the masked position, The computer-implemented method according to claim 1, wherein the mask modeling loss term evaluates whether the prediction identifier corresponds to the true discretization identifier among the plurality of discretization identifiers corresponding to the masked position.

4. The computer-implemented method according to claim 1, wherein the second encoder portion and / or the third encoder portion of the machine learning model includes one or more conformable blocks.

5. The computer-implemented method according to claim 1, wherein the third encoder portion of the machine learning model includes one or more conformable blocks.

6. The step of training the machine learning model by the computing system based on the loss function includes: The step of modifying, by the computing system, one or more values of one or more parameters of the third encoder portion, the second encoder portion, and the first encoder portion of the machine learning model based on the mask modeling loss term; The computer-implemented method according to claim 1, further comprising the step of modifying, by the computing system, one or more values of one or more parameters of the second encoder portion and the first encoder portion of the machine learning model based on a combination of the mask modeling loss term and the contrast loss term.

7. The computer-implemented method according to claim 1, further comprising the step of modifying, by the computing system, a codebook used to perform the quantization based on the loss function.

8. The computer-implemented method according to claim 1, wherein the series of input data includes audio data or a spectrogram representation of the audio data.

9. The computer-implemented method according to claim 8, wherein the audio data includes voice data.

10. The computer-implemented method according to claim 1, wherein the series of input data includes text data.

11. The computer-implemented method according to any one of claims 1 to 10, wherein the series of input data includes sensor data or image data.

12. One or more computer-readable storage media for collectively storing instructions, which, when executed by one or more processors of a computing system, cause the computing system to perform operations, the operations including obtaining task-specific training inputs; processing the task-specific training inputs using a machine learning model to generate task-specific training outputs, wherein at least an encoder portion of the machine learning model is trained using a loss function including a contrastive loss term and a masking modeling loss term, the contrastive loss term evaluating a contrastive pre-training output generated based on a first set of context vectors and a plurality of target quantization vectors, the first set of context vectors being generated by the encoder portion of the machine learning model after masking an input or intermediate output of the encoder portion of the machine learning model, the plurality of target quantization vectors being generated by quantization of the input or intermediate output of the encoder portion of the machine learning model, the masking modeling loss term evaluating a masking modeling pre-training output generated based on a second set of context vectors and a plurality of discretization identifiers, the second set of context vectors being generated by the encoder portion of the machine learning model from the first set of context vectors, the plurality of discretization identifiers being generated by quantization of the input or intermediate output of the encoder portion of the machine learning model, generating; evaluating a task-specific loss function based on the task-specific training output; training the machine learning model based on the task-specific loss function. One or more computer-readable storage media including.

13. For each of one or more masked positions The control pre-training output includes a predicted selection from a set of candidate vectors, the predicted selection being generated based on one of the first set of context vectors corresponding to the masked position, the set of candidate vectors including a true target quantization vector and one or more scrambled vectors, The control loss term is one or more computer-readable storage media according to claim 12, which evaluates whether the predicted selection corresponds to the true target quantization vector.

14. For each of the one or more masked positions, the mask modeling pre-training output includes a predicted identifier generated based on one of the second set of context vectors corresponding to the masked position, The mask modeling loss term is one or more computer-readable storage media according to claim 12, which evaluates whether the predicted identifier corresponds to the true discretization identifier among the plurality of discretization identifiers corresponding to the masked position.

15. The encoder portion of the machine learning model includes one or more conformable blocks, one or more computer-readable storage media according to claim 12.

16. The machine learning model includes a decoder portion configured to process the output of the encoder portion to generate the task-specific training output, one or more computer-readable storage media according to any one of claims 12 to 15.

17. The task-specific training input includes audio data, The task-specific training output includes the audio data or a translation of speech recognition for the audio data, one or more computer-readable storage media according to claim 12.

18. A computing system comprising one or more processors and one or more computer-readable storage media collectively storing instructions, the instructions when executed by the one or more processors of the computing system cause the computing system to perform operations, the operations including obtaining a task-specific inference input, ​ Processing the task-specific inference input using a machine learning model to generate a task-specific inference output, wherein at least an encoder part of the machine learning model is trained using a loss function including a contrast loss term and a mask modeling loss term, the contrast loss term evaluating a contrast pre-training output generated based on a first set of context vectors and a plurality of target quantization vectors, the first set of context vectors being generated by the encoder part of the machine learning model after masking an input or intermediate output of the encoder part of the machine learning model, the plurality of target quantization vectors being generated by quantization of the input or intermediate output of the encoder part of the machine learning model, the mask modeling loss term evaluating a mask modeling pre-training output generated based on a second set of context vectors and a plurality of discretization identifiers, the second set of context vectors being generated from the first set of context vectors by the encoder part of the machine learning model, the plurality of discretization identifiers being generated by quantization of the input or intermediate output of the encoder part of the machine learning model, and generating; A computing system including providing the task-specific inference output as an output. Claim 19 For each of one or more masked positions, The contrast pre-training output includes a prediction selection from a set of candidate vectors, the prediction selection being generated based on one of the first set of context vectors corresponding to the masked position. The set of candidate vectors includes a true target quantization vector and one or more scrambled vectors. The computing system according to claim 18, wherein the contrast loss term evaluates whether the prediction selection corresponds to the true target quantization vector. Claim 20 For each of one or more masked positions, The mask modeling pre-training output includes a prediction identifier generated based on one of the second set of context vectors corresponding to the masked position. The computing system according to claim 18, wherein the mask modeling loss term evaluates whether the predicted identifier corresponds to a true discretization identifier among the plurality of discretization identifiers corresponding to the masked position. **Claim 21** The computing system according to claim 18, wherein the encoder portion of the machine learning model includes one or more conformable blocks. **Claim 22** The computing system according to claim 18, wherein the machine learning model includes a decoder portion configured to process an output of the encoder portion to generate the task-specific inference output. **Claim 23** The task-specific inference input includes audio data, The computing system according to any one of claims 18 to 22, wherein the task-specific inference output includes the audio data or a translation of speech recognition for the audio data.