Training neural networks by intermediate block stacking
Patent Information
- Application Number
- PCT/US2025/030579
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-05-22
- Publication Date
- 2026-10-01
Smart Images

Figure US2025030579_01102026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: 45288-0633WO1TRAINING NEURAL NETWORKS BY INTERMEDIATE BLOCK STACKINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 778,304, filed on March 26, 2025. The disclosure of the prior application is considered part of and is incorporated by reference in its entirety in the disclosure of this application.BACKGROUND
[0002] This specification relates to training neural networks.
[0003] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., another hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.SUMMARY
[0004] This specification describes a training system implemented as computer programs on one or more computers in one or more locations that implements a multi-stage training process that includes a sequence of training stages to train a neural network to perform one or more machine learning tasks on a network input.
[0005] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0006] The described techniques address challenges in training large neural networks, particularly deep architectures like Transformers, which are often computationally intensive. Conventional approaches to progressively growing network depth, for instance by stacking new layers at network extremities or using random initialization for new deeper layers, can suffer from suboptimal training convergence or fail to establish beneficial inductive biases for complex tasks. The embodiments described herein aim to provide a more technically efficient method for training progressively deeper neural networks, focusing on accelerating convergence, reducing computational cost, and fostering a network structure conducive to improved performance onAttorney Docket No.: 45288-0633WO1challenging technical applications, such as those demanding complex reasoning or the generation of structured technical outputs like computer code or control sequences for a technical system.
[0007] A training system implementing the intermediate stacking technique described in this specification can train a neural network by progressively expanding the model size (depth) of the neural network across a plurality of training stages and using the values of parameters of an existing, intermediate block from a smaller model in an earlier training stage to initialize the values of parameters of a new block in a next training stage.
[0008] By initializing the values of parameters of the new block using the trained values of parameters of the existing, intermediate block that have been learned in the earlier training stage rather than random initialization, the training system can improve computational resource efficiency during the training process.
[0009] More efficient usage of computing resources lowers energy consumption, which in turn reduces carbon emissions and environmental impact. More efficient usage of computing resources means faster training times, i.e., increased speed of the training process of the neural network, without causing performance drop. This accelerates development cycles and allows for quicker iteration and deployment of neural networks. More efficient usage of computing resources also makes it easier to scale up experiments, handle larger datasets, or train more complex neural networks without requiring exponential increases in hardware.
[0010] As a particular example, implementations of the intermediate stacking technique can speed up the training process of a language model neural network by up to 40% compared to existing stacking techniques that involve stacking the first or last block of the neural network while expanding the model size. This increase in the speed of the training process would substantially reduce the computational resources required by the training process, and simultaneously increase the availability of computational resources for other tasks.
[0011] Furthermore, expanding the model size by stacking the intermediate block rather than the first or last block provides the neural network with a desirable inductive bias toward performance improvement on many downstream tasks, especially including complex reasoning tasks, generative tasks, and arithmetic tasks.
[0012] This advantageous inductive bias may be attributed to the technical strategy of selecting an intermediate block for duplication, as such blocks are hypothesized to capture a balance of abstract yet adaptable representations. When these representations are used as a basisAttorney Docket No.: 45288-0633WO1for deeper layers by duplicating the intermediate block, they can be particularly effective for building the complex hierarchical features needed for these tasks. This technical choice can differentiate the approach from methods that merely grow the network at its extremities (first or last blocks), which may not yield the same improvement in inductive bias for complex reasoning or generative capabilities relevant to technical applications.
[0013] The neural network trained using the described intermediate stacking technique can thus be more easily adapted to any of a range of downstream tasks. Once adapted, the neural network can exceed the performance of other neural networks trained using existing stacking techniques on many downstream tasks, despite an adaptation process that consumes fewer computing resources, is faster in terms of wall-clock time, or both.
[0014] The details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] FIG. 1 shows an example training system.
[0016] FIG. 2 is an example illustration of training a neural network across a sequence of training stages.
[0017] FIG. 3 is an example illustration of expanding a neural network using an intermediate stacking technique during a training stage in comparison to expanding a neural network using a gradual stacking technique during the training stage.
[0018] FIG. 4 is a flow diagram of an example process for training a neural network at a given training stage in a sequence of training stages.
[0019] FIG. 5 is a flow diagram of an example process for training a neural network at another given training stage in a sequence of training stages.
[0020] FIG. 6 shows four downstream evaluations vs validation log perplexity isoplots.
[0021] FIG. 7 shows test accuracy improvements of a neural network trained using the intermediate stacking technique compared to a baseline neural network on the same tasks.
[0022] Like reference numbers and designations in the various drawings indicate like elements.Attorney Docket No.: 45288-0633WO1DETAILED DESCRIPTION
[0023] FIG. 1 shows an example training system 100. The training system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations that implements a multi-stage training process that includes a sequence of training stages to train a neural network 110 to perform one or more machine learning tasks on a network input 102 to generate a network output 152.
[0024] The machine learning tasks can include any machine learning task that (i) operates on a network input 102 that is an input sequence, (ii) generates a network output 152 that is an output sequence, or (iii) both.
[0025] Some examples of machine learning tasks that the neural network 110 can be configured through training to perform are discussed below.
[0026] As one example, the task may be a neural machine translation task. For example, if the input to the neural network is a sequence of text, e.g., a sequence of words, phrases, characters, or word pieces, in one language, the output generated by the neural network may be a translation of the sequence of text into another language, i.e., a sequence of text in the other language that is a translation of the input sequence of text. As a particular example, the task may be a multi-lingual machine translation task, where a single neural network is configured to translate between multiple different source language - target language pairs. In this example, the source language text may be augmented with an identifier that indicates the target language into which the neural network should translate the source language text.
[0027] As another example, the task may be an audio processing task. For example, if the input to the neural network is a sequence representing a spoken utterance, the output generated by the neural network may be a score for each of a set of pieces of text, each score representing an estimated likelihood that the piece of text is the correct transcript for the utterance. As another example, if the input to the neural network is a sequence representing a spoken utterance, the output generated by the neural network can indicate whether a particular word or phrase (“hotword”) was spoken in the utterance. As another example, if the input to the neural network is a sequence representing a spoken utterance, the output generated by the neural network can identify the natural language in which the utterance was spoken.
[0028] As another example, the task can be a natural language processing or understanding task, e.g., an entailment task, a paraphrase task, a textual similarity task, a sentiment task, aAttorney Docket No.: 45288-0633WO1sentence completion task, a grammaticality task, and so on, that operates on a sequence of text in some natural language.
[0029] As another example, the task can be a text to speech task, where the input is text in a natural language or features of text in a natural language and the network output is a spectrogram, a waveform, or other data defining audio of the text being spoken in the natural language.
[0030] As another example, the task can be a health prediction task, where the input is a sequence derived from electronic health record data for a patient and the output is a prediction that is relevant to the future health of the patient, e.g., a predicted treatment that should be prescribed to the patient, the likelihood that an adverse health event will occur to the patient, or a predicted diagnosis for the patient.
[0031] As another example, the task can be a text generation task, where the input is a sequence of text, and the output is another sequence of text, e.g., a completion of the input sequence of text, a response to a question posed in the input sequence, or a sequence of text that is about a topic specified by the first sequence of text. In this example, both the input sequence of text and the output sequence of text can include tokens from a vocabulary of text tokens that includes, e.g., one or more of characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in a natural language or a computer language.
[0032] As a similar example, the task can be an automatic code generation task, where the input is a sequence of words, wordpieces or characters in a first natural language and the output is a sequence of tokens that represent instructions in a computer programming or markup language, or instructions for controlling an application program to perform a task e.g. build a data item such as an image or web page.
[0033] As a particular example of this, the input can represent a context input that includes a text description of a desired piece of code or a snippet of computer code in a programming language and the output can be an output sequence that includes computer code, e.g., a snippet of code that is described by the context input or a snippet of code that follows the context input in a computer program.
[0034] As another example, the input to the text generation task can be an input other than text, e.g., an image or audio, and the output sequence can be text that describes the input. In this example, both the input sequence of text and the output sequence of text can include tokens fromAttorney Docket No.: 45288-0633WO1a vocabulary of tokens that includes tokens that can represent data other than text, in addition to the text tokens mentioned above.
[0035] For example, the vocabulary of tokens can additionally include image tokens that represent a discrete set of image patch embeddings of an image that can be generated by an image encoder neural network based on processing the image patches of the image. As another example, the vocabulary of tokens can additionally include audio tokens that represent code vectors in a codebook of a quantizer, e.g., a residual vector quantizer.
[0036] As another example, the task can be an image generation task, where the input is a conditioning input, e.g., text, a lower-resolution image, or a partial image, and the output is a sequence of intensity value inputs for the pixels of an image.
[0037] As another example, the task can be an image processing task. For example, the input can be the intensity values of the pixels of the image or an encoded representation of the intensity values of the pixels generated by an encoder neural network, and the network output can be (i) an image classification output that classifies the input image into one of a plurality of object categories (ii) an object detection output, i.e., a sequence that specifies the coordinates of one or more bounding boxes in the image that are predicted to encompass objects or (iii) a segmentation output that classifies each pixel in the input image into one of a plurality of categories. As another example, the input can include the intensity values of the pixels of the image or an encoded representation of the intensity values of the pixels generated by an encoder neural network and optionally text, and the network output can be text that characterizes the image, e.g., captions the image or answers a question posed by the text in the input about the image.
[0038] As another example, the task can be an audio generation task, where the input is a conditioning input, e.g., text, an image, or context audio, and the output is a sequence of tokens that represents audio.
[0039] As another example, the task can be an audio processing task. For example, the input can include audio or an encoded representation of the audio generated by an encoder neural network, and the network output can be text or an image that characterizes the audio.
[0040] As another example, the task can be an agent control task, where the input is a sequence of observations or other data characterizing states of an environment and the output defines an action to be performed by the agent in response to the most recent data in theAttorney Docket No.: 45288-0633WO1sequence. The agent can be, e.g., a real-world or simulated robot, a control system for an industrial facility, or a control system that controls a different kind of agent.
[0041] As another example, the task can be a genomics task, where the input is a sequence representing a fragment of a DNA sequence or other molecule sequence and the output is either an embedding of the fragment for use in a downstream task, e.g., by making use of an unsupervised learning technique on a data set of DNA sequence fragments, or an output for the downstream task. Examples of downstream tasks include promoter site prediction, methylation analysis, predicting functional effects of non-coding variants, and so on.
[0042] In some cases, the machine learning task is a combination of multiple individual machine learning tasks, i.e., the neural network is configured to perform multiple different individual machine learning tasks, e.g., two or more of the machine learning tasks mentioned above. For example, the neural network can be configured to perform multiple individual natural language understanding tasks, with the network input including an identifier for the individual natural language understanding task to be performed on the network input.
[0043] In some cases, the machine learning task is a multi-modal processing task that requires processing multi-modal data. In general, multi-modal data is a combination of two or more different types of data, e.g., two or more of audio data, image data, text data, or graph data. As one example the multi-modal data may comprise audio-visual data, comprising a combination of pixels of an image or of video and audio data representing values of a digitized audio waveform. As another example the multi-modal data may comprise a combination of i) text data representing text in a natural language and ii) pixels of an image or of video or audio data representing values of an audio waveform. Optionally, but not necessarily, the different types of data may represent the same or overlapping objects using the different modalities (types), and when processing multi-modal data the data may be mapped into a common embedding space.
[0044] As a particular example, the task is a multi-modal processing task that requires processing both text and image inputs, so that the neural network includes both a computer vision neural network and a text processing neural network. That is, the target output to be generated by the computer vision neural network for a given image depends on one or more outputs generated by the text processing neural network for one or more corresponding text inputs (and vice versa). Examples of such tasks include open-vocabulary image classification,Attorney Docket No.: 45288-0633WO1open -vocabulary object detection, image captioning, text-based image search, image-based retrieval, and so on.
[0045] More generally, the multi-modal processing task may correspond to any of the tasks previously described for any of the types of data making up the multi-modal combination. For example, an accuracy of the previously described tasks may be increased when the task is applied to multi-modal data combining the data for which the task has been previously described and another type of data. For example detection or classification of an object or event may be improved when data of multiple different types (modalities) is processed. As another example, the quality (e.g., accuracy, fidelity, or intelligibility) of a generated image, video, or audio may be improved when data of multiple different types (modalities) is processed.
[0046] In some cases, once trained, the neural network can perform tasks that it was not explicitly trained to perform. For example the neural network can perform translation tasks (provided that the training corpus included words in different languages), generative tasks, and many other tasks.
[0047] In these cases, the neural network can be made to perform a particular task by providing a natural language description of the desired response as a part of the input or “prompt”. The prompt may be a few-shot prompt where a few, e.g., 1 to 10, examples of a query and an example output are provided in the text prior to the actual query.
[0048] Additional description of generative tasks that the neural network (when configured as a “generative” neural network) can perform are discussed below.
[0049] Generally, the generative neural network is configured to process a conditioning input (“prompt”) to generate a data item. The data item can include data in any of a variety of modalities, e.g., text data, image data, video data, or audio data. Generally, the data item represents a response to the conditioning input which may be, e.g. a “prompt” for the generative neural network. For example, the conditioning input can characterize one or more desired properties for the generated data item.
[0050] In some implementations the generative neural network generates an output token sequence from an input token sequence including the conditioning input. The generative neural network may then be configured to process the input token sequence to generate for each position in the output token sequence, a respective score for each token in a vocabulary of output tokens, that is used to select an output token for the output token sequence.Attorney Docket No.: 45288-0633WO1
[0051] In some implementations the tokens can represent text, e.g., words, wordpieces or characters, in a natural or computer language. For example, text may be received, e.g., as a series of encoded characters, e.g. UTF-8 encoded characters; such “characters” can include Chinese and other similar characters, as well as logograms, syllabograms and the like. A text encoder, i.e. a tokenizer, can process a sequence of text to represent the text as a series of text tokens from a vocabulary of text tokens, e.g. that each represent words, wordpieces or characters in a natural or computer language. The computer language may be any formal language used to communicate with a computer, e.g. a markup language, or a command or configuration language, or a data exchange language such as JSON, or a programming language. The tokenizer can, e.g., implement BPE (Byte Pair Encoding) or Wordpiece tokenization. Optionally the text can be obtained from audio data representing speech; the output tokens may be converted into audio data that represent speech corresponding to the text.
[0052] Also, or instead the tokens may represent an image. For example, a set (sequence) of input or output tokens can represent an image. Each image token may comprise a block encoding of values of the pixels in a different region of an image that maps a set of values of the pixels to a respective image token. The block encoder may comprise a neural network, e.g. having one or more (self-)attention layers, such as a Transformer neural network.
[0053] Also, or instead the tokens may represent an audio waveform. For example, a set (sequence) of input or output tokens can represent audio data representing a waveform e.g. instantaneous audio amplitude values or time-frequency audio data. Each image token may comprise a block encoding of the audio waveform in a different time segment of the audio that maps a set of values representing the audio waveform to a respective image token. The block encoder may comprise a neural network, e.g. having one or more (self-)attention layers, such as a Transformer neural network.
[0054] In a multimodal system audio data or an image may be flagged by a start-of-audio token or start-of-image token.
[0055] In some implementations the generative neural network can be a multimodal network that is configured to process a conditioning input comprising one or more of text data, audio data defining an audio signal (e.g. as amplitude values of the audio signal or as a time-frequency representation of the audio signal), or a still or moving image (e.g. as image pixel values), to generate a data item that can similarly comprise text data, audio data, or a still or moving image.Attorney Docket No.: 45288-0633WO1
[0056] For example, the conditioning input may comprise text and the data item may comprise an image or an audio signal that represents speech an image generated in response to the text, e.g. described by the text. Also, or instead the conditioning input may comprise an audio signal that represents speech, or an image, and the data item may comprise text, e.g. that describes the conditioning input.
[0057] As another example the conditioning input may comprise an observation, e.g. of a real world environment, e.g. from sensor such as a camera or other image sensor; and optionally additional information such as information defining a particular task to be deformed. The output data item may comprise agent control data that defines one or more actions to be performed by an agent, e.g. by a mechanical agent such as a robot or autonomous vehicle, to perform a task. The reward model(s) may, e.g., define a preferred trajectory of motion of the mechanical agent in the (real-world) environment.
[0058] In general, the neural network 110 can have any appropriate architecture to enable it to perform these machine learning tasks, e.g., to perform a generative task by processing the conditioning input to generate the data item.
[0059] The neural network 110 includes a sequence (or stack) of blocks. A block refers to a group of one or more contiguous neural network layers in a neural network.
[0060] For example, the neural network 110 can include a plurality of neural network layers, and each block can include a different subset of the plurality of neural network layers.
[0061] Each block includes a respective set of parameters. The respective set of parameters includes parameters of each of the one or more contiguous neural network layers included in the block.
[0062] The plurality of neural network layers can include any kind of neural network layers, for example, attention layers, e.g., self-attention layers, multi-head self-attention layers, or crossattention layers, convolutional layers, fully-connected layers, embedding layers, activation layers, or recurrent layers, e.g., Long Short-Term Memory (LSTM) layers or gated recurrent unit (GRU) layers.
[0063] An attention layer is a neural network layer that applies an attention mechanism. Generally, to apply the attention mechanism, the attention layer uses one or more attention heads. Each attention head generates a set of queries, a set of keys, and a set of values, and then applies any of a variety of variants of query-key-value (QKV) attention using the queries, keys,Attorney Docket No.: 45288-0633WO1and values to generate an output. When there are multiple attention heads, the attention layer then combines the outputs of the multiple attention heads, e.g., by concatenating the outputs and, optionally, processing the concatenated outputs through a linear layer.
[0064] Examples of QKV attention variants are described in Vaswani, et al., attention Is All You Need, arXiv: 1706.03762, Raffel, et al, Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, arXiv: 1910.10683, Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, arXiv: 1810.04805, Kitaev, et al., Reformer: The efficient transformer, arXiv preprint arXiv: 2001.04451, and Chowdhery, et al., Palm: Scaling language modeling with pathways. Journal of Machine Learning Research 24.240 (2023): 1-113, the entire contents of which are hereby incorporated by reference herein in their entirety.
[0065] The selection of an intermediate block for duplication and re-insertion, as opposed to an initial or final block, provides significant technical advantages. Initial blocks in a deep neural network typically learn to extract low-level, general features from the input data (e.g., edges or basic textures in an image, or fundamental n-gram statistics in text processed by the neural network for natural language understanding or machine translation tasks). Duplicating these early blocks may not efficiently contribute to learning higher-level abstractions needed for complex tasks and could lead to redundancy without proportional performance gain in terms of, for example, prediction accuracy or processing speed. Conversely, final blocks tend to learn features that are highly specialized to the specific training objective and output format being optimized at that stage of training (e.g., specific class logits for a classification task). Duplicating these late-stage, specialized blocks might prematurely narrow the network's learning trajectory or hinder its ability to adapt its deeper layers to more nuanced or complex patterns relevant for improved technical performance.
[0066] An intermediate block, situated between these extremes, is hypothesized to have learned a balance of sufficiently abstract yet still adaptable feature representations. These representations are more sophisticated than those in early blocks but less over-specialized than those in late blocks. Replicating an intermediate block and its learned parameters provides a strong, relevant initialization for new layers, allowing the network to more effectively explore deeper functional compositions and hierarchies of features. This strategy aims to create a “scaffold” of well-initialized parameters in the middle of the network, fostering the developmentAttorney Docket No.: 45288-0633WO1of complex feature interactions fortasks requiring deep reasoning (e.g., in medical diagnosis support or complex system control) or intricate generative capabilities (e.g., generating precise technical documentation or synthesizing control parameters for a physical process). This may improve the network’s learning capacity and the efficiency of the training process as it grows, leading to reduced computational load and faster convergence to a high-performing model.
[0067] In the example of FIG. 1, the neural network 110 has a Transformer based architecture. The neural network 110 includes a sequence of blocks that include an attention block 120. The attention block 120 includes one or more attention layers, e g., one or more selfattention layers, one or more multi-head self-attention layers, or one or more cross-attention layers. The attention block 120 can also include additional components, e.g., one or more feedforward layers, one or more residual connection layers, one or more normalization layers, and so on.
[0068] For clarity, the “sequence of blocks” referred to herein typically constitutes the main processing pathway or “stack” of the neural network, often comprising repeating structural units like multiple attention layers. The neural network can include additional neural network layers before and / or after the sequence of blocks.
[0069] This sequence generally excludes one or more initial input processing layers (e.g., an embedding layer such as a text token embedding layer or an image patch embedding layer, a positional encoding layer, etc.) that convert a raw network input 102 into an initial representation to be received by the first block in the sequence, and one or more final output processing layers (e.g., a de-embedding layer, a projection layer, layers in a classification head, a softmax layer or another activation layer, etc.) that generate the final network output 152 from the block output produced by the last block in the sequence.
[0070] This distinction is relevant because the intermediate stacking process focuses on deepening the core computational pathway of the neural network, where complex feature transformations and hierarchical representations for solving technical problems may be learned. The input embedding and output projection layers often perform more fixed roles related to data format conversion (e.g., converting input sensor values into an initial vector representation, or mapping final hidden states to output control signals) and their parameters may not benefit from, or be suitable for, the same duplication and progressive deepening strategy applied to the intermediate processing blocks. A technical effect can be a more targeted and efficient increaseAttorney Docket No.: 45288-0633WO1in model depth, focused on the parts of the network most critical for learning complex functions, potentially optimizing the use of computational resources for improving the model’s problemsolving capacity.
[0071] Thus, in some cases, the first block (or more generally, any block in the sequence of blocks) excludes, i.e., does not include, an embedding layer. In some cases, the last block (or more generally, any block in the sequence of blocks) does not include a de-embedding layer. In some cases, the last block (or more generally, any block in the sequence of blocks) does not include the softmax layer.
[0072] In some cases, the first block (or more generally, any block in the sequence of blocks) does not include any neural network layer from a separately trained neural network. That is, the first block does not include any layer that is initialized based on one of the layers from another separately trained neural network.
[0073] In some cases, the neural network 110 does not include a separately pre-trained neural network integrated into the sequence of blocks. That is, the sequence of blocks does not include any block that is initialized based on a block from another separately trained neural network.
[0074] The neural network 110 is configured to perform a machine learning task by performing one or more forward passes through the sequence of blocks in the neural network.
[0075] In some cases, the neural network can receive a network input and perform a single forward pass through the sequence of blocks using the network input to generate a network output for the machine learning task.
[0076] In these cases, the network input is processed by the initial input processing layers that are arranged preceding the first block in the sequence of blocks to generate an initial representation of the network input — and the initial representation of the network input is then received as a block input by the first block in the sequence of blocks to generate block output ( to be received as a block input by the second block in the sequence of blocks). The block output generated by the last block in the sequence of blocks is processed by the final output processing layers there are arranged subsequent to the last block in the sequence of blocks to generate the network output.
[0077] In some cases where the task is a generative task, the neural network can execute an auto-regressive generation process across a plurality of generation steps to generate a data item that is an output sequence made up of tokens from a vocabulary, conditioned on a conditioningAttorney Docket No.: 45288-0633WO1input (“prompt”) that provides context for the output sequence. In these cases, each generation step corresponds to a respective forward pass through the sequence of blocks.
[0078] More specifically, the auto-regressively generated output sequence is created by, at each generation step, generating a particular token in the output sequence conditioned on a current input sequence that includes any tokens that precede the particular token in the output sequence, i.e., the tokens that have already been generated for any previous positions in the output sequence that precede the particular position of the particular token, and the prompt.
[0079] For example, the current input sequence when generating a token at any given position in the output sequence can include the input sequence and the tokens at any preceding positions that precede the given position in the output sequence. As a particular example, the current input sequence can include the prompt followed by the tokens at any preceding positions that precede the given position in the output sequence.
[0080] More specifically, at each generation step, to generate a particular token at a particular position within an output sequence, the neural network can process the current input sequence to generate a score distribution, e.g., a probability distribution, that assigns a respective score, e.g., a respective probability, to each token in the vocabulary of tokens. The neural network can then select, as the particular token, a token from the vocabulary using the score distribution. For example, the neural network can greedily select the highest-scoring token or can sample, e.g., using nucleus sampling or another sampling technique, a token from the distribution.
[0081] In these cases, at each generation step, an initial representation of the current input sequence that is generated by one or more initial input processing layers is received by the first block in the sequence of blocks, and the score distribution is generated by one or more final output processing layers from the block output of the last block in the sequence of blocks after the neural network has performed a forward pass through the sequence of blocks using the current input sequence.
[0082] In some cases where the task is a generative task, the neural network can execute a reverse diffusion process across a plurality of updating iterations to generate a data item. In these cases, each updating iteration corresponds to a respective forward pass through the sequence of blocks.Attorney Docket No.: 45288-0633WO1
[0083] At each updating iteration, the neural network is configured to map a diffusion input that includes a current (noisy) representation of the data item and a conditioning input and to generate an updated representation of the data item. The diffusion input at each updating iteration can also include data defining a noise level for the updating iteration.
[0084] In some implementations, the neural network performs a reverse diffusion process in output space, e.g., pixel space in the case of images. In this example, the representations operated on and generated by the neural network have values for each pixel that specify color values, e.g., RGB values or another color encoding scheme.
[0085] In some other implementations, the diffusion neural network performs a reverse diffusion process in latent space, e.g., in a latent space that is lower-dimensional than the output space. That is, the representations operated on by the neural network are latent representations and the values in the representations are learned, latent values, e.g., rather than color values.
[0086] The updated representation of the data item can either be generated directly, e.g., where a denoising output of the neural network for the updating iteration includes the updated representation, or indirectly, e.g., where the denoising output includes a noise term computed by the neural network for the updating iteration, and the updated representation can then be generated by removing at least some of the noise from the current representation in accordance with the noise term. The noise term is an estimate of the noise, as computed by the neural network, that has been added to the data item to arrive at the current representation.
[0087] In these cases, at each updating iteration, an initial representation of the current representation that is generated by one or more initial input processing layers is received by the first block in the sequence of blocks, and the denoising output (from which the updated representation can be derived) is generated by one or more final output processing layers from the block output of the last block in the sequence of blocks after the neural network has performed a forward pass through the sequence of blocks using the current representation.
[0088] To implement the multi-stage training process that includes a sequence of training stages to train the neural network 110, the training system 100 uses a stacking engine 130 and a training engine 140.
[0089] Each training stage includes a respective plurality of training iterations. In particular, the training process includes a plurality of training iterations, and each training stage corresponds to a different subset of all of the plurality of training iterations included in the training process.Attorney Docket No.: 45288-0633WO1
[0090] In some implementations, each training stage in the multi-stage training process includes the same number of training iterations whereas, in other implementations, the numbers of training iterations in different training stages differ.
[0091] At each training stage, the training system 100 first uses the stacking engine 130 to expand the architecture of the neural network 110, and then uses the training engine 140 to train the neural network 110 having the expanded architecture by performing the respective plurality of training iterations that correspond to the training stage.
[0092] At a given training stage in the sequence of training stages, the stacking engine 130 implements an intermediate stacking technique to expand the architecture of the neural network 110 by identifying an intermediate block from the sequence of blocks as of the given training stage and then generating a new block based on the intermediate block for inclusion in the expanded architecture.
[0093] The intermediate block is arranged between the first block in the sequence of blocks and the last block in the sequence of blocks. The intermediate block is not the first or last block. In various implementations, the intermediate block does not immediately follow the first block in the sequence of blocks, and the last block does not immediately follow the intermediate block in the sequence of blocks.
[0094] The stacking engine 130 updates the sequence of blocks to include the intermediate block and the new block, with both the intermediate block and the new block arranged between the first block and the last block in the sequence of blocks. There are many ways in which the stacking engine 130 can do this.
[0095] In some implementations, the stacking engine 130 updates the sequence of blocks by stacking the new block atop the intermediate block, such that the new block immediately follows the intermediate block. Thus, during each forward pass through the sequence of blocks, the intermediate block receives a block input that includes a hidden state from a preceding block in the sequence and provides a block output that includes an updated hidden state to a subsequent block (the new block) in the sequence.
[0096] A hidden state is a vector of numeric values having a fixed dimensionality. The hidden state is generally the output of the last neural network layer of the one or more neural network layers included in a block, or, when the layers include skip connections or residualAttorney Docket No.: 45288-0633WO1connections, a combination, e.g., a sum, concatenation, or average, of the outputs of two or more of the neural network layers included in the block.
[0097] In other implementations, the system updates the sequence of blocks by stacking the intermediate block atop the new block, such that the intermediate block immediately follows the new block. Thus, during each forward pass through the sequence of blocks, the new block receives a block input that includes a hidden state from a preceding block in the sequence and provides a block output that includes an updated hidden state to a subsequent block (the intermediate block) in the sequence.
[0098] In yet other implementations, the system updates the sequence of blocks by inserting the new block into any intermediate position within the sequence of blocks.
[0099] Thus, in various implementations, the intermediate block is not a block that receives the initial representation of the network input (or, analogously, the initial representation of the current input sequence or the diffusion input). Nor is the intermediate block a block that generates the block output that is received by one or more final output processing layers to generate the network output (or, analogously, the score distribution or the denoising output).
[0100] Likewise, in various implementations, after being included in the sequence of blocks, the new block is also not a block that receives the initial representation of the network input (or, analogously, the initial representation of the current input sequence or the diffusion input). Nor is the new block a block that generates the block output that is received by one or more final output processing layers to generate the network output (or, analogously, the score distribution or the denoising output).
[0101] The stacking engine 130 generates the new block such that the parameters of the new block have values that are derived from trained values of parameters of the intermediate block as of the given training stage, i.e., from the values of the parameters of the intermediate block that have been learned during the training process up to and not including the given training stage.
[0102] Thus, in some implementations, the new block can be viewed as a duplication of the intermediate block: the intermediate block and the new block include the same number and type of layers, and the new block includes a set of parameters having values that are identical to the trained values of the parameters of the intermediate block.
[0103] At the given training stage in the sequence of training stages, after the stacking engine 130 has updated the sequence of blocks to include the new block, the training engine 140 trainsAttorney Docket No.: 45288-0633WO1the neural network, which has the updated sequence of blocks that has been updated as of the given training stage, on training data based on optimizing one or more objective functions.
[0104] In general, the training engine 140 trains the blocks in the updated sequence, including the new block and the intermediate block, independently from each other, so that the values of the parameters of the new block and the intermediate block will be different after the given training stage (although they may begin with being identical to each other).
[0105] By initializing the values of parameters of the new block using the trained values of parameters of the existing, intermediate block that have been learned in the earlier training stage rather than random initialization, the training system 100 can improve computational resource efficiency during the training process.
[0106] The use of trained parameters from an intermediate block means the newly added deeper portion of the network can start from a more informed state than random initialization, potentially avoiding lengthy and resource-intensive initial learning phases where basic feature extractors for that depth would have to be learned from scratch. This, in turn, can reduce the number of training epochs and associated processor cycles, e.g., CPU / GPU / TPU cycles. It can also help avoid the potential instability or slow convergence that might occur if a randomly initialized block is inserted deep within an already partially trained network, which could degrade the overall learning process. Furthermore, because the intermediate block's parameters are already adapted to the data distribution (e.g., statistical properties of sensor readings or image features) and the task (to some extent), the fine-tuning required for the expanded network can be more focused and may converge more rapidly. This can lead to the technical effect of reduced training iterations, typically measurable by a decrease in the number of epochs or total floatingpoint operations (FLOPs) required to reach a target performance level (e.g., a specific validation loss or task-specific accuracy), and potentially, lower energy consumption and faster availability of the trained model for deployment in a technical system.
[0107] Furthermore, expanding the model size by stacking the intermediate block rather than the first or last block provides the neural network with a desirable inductive bias toward performance improvement on many downstream tasks, as will be discussed in greater detail below with reference to FIG. 6.
[0108] The “desirable inductive bias” may stem from the specific choice of duplicating an intermediate block that has already learned meaningful, yet not overly specialized,Attorney Docket No.: 45288-0633WO1transformations of the input data (e.g., transformations of sensor data from a monitored industrial process, or intermediate representations of code syntax for an automatic code generation task). By re-inserting this block (or its copy) deeper into the network, the training process encourages the network to build upon these existing learned functions to discover more complex, hierarchical relationships in the data. This can be contrasted with stacking at the end, which might simply refine existing high-level features without adding significant new representational depth for intermediate processing stages, or stacking at the beginning, which might require substantial retraining of subsequent layers to adapt to new low-level features. The intermediate placement and initialization can facilitate the creation of deeper processing pathways for these mid-level representations, which may be particularly advantageous for technical tasks where the solution involves multiple stages of abstraction or reasoning (e.g., multi-step control decisions in an autonomous system, or generation of complex structured data like a chemical formula). This targeted deepening is a technical aspect that can contribute to faster convergence on such tasks and potentially lead to a more robust internal representation that is less susceptible to noise in input data (e.g., noisy sensor readings). This can manifest as improved technical performance metrics, e.g., higher accuracy in detecting fault conditions in an industrial system, improved precision / recall in identifying objects in medical images, or reduced perplexity in generative language tasks that produce technical specifications, all potentially achievable for a given computational budget for training.
[0109] Such a multi-stage training process implemented by using the stacking engine 130 and the training engine 140 is applicable to all of the above use cases (where the neural network 110 performs a machine learning task by performing one or more forward passes through the sequence of blocks in the neural network 110).
[0110] FIG. 2 is an example illustration of training the neural network 110 of FIG. 1 by the training system 100 which implements a multi-stage training process that includes a sequence of training stages.
[0111] As mentioned previously, at a given training stage in the sequence of training stages, the stacking engine 130 identifies an intermediate block of the neural network 110 and uses the identified intermediate block to generate a new block to be included as part of the neural network 110.Attorney Docket No.: 45288-0633WO1
[0112] In this way, the architecture of the neural network 110 and, correspondingly, the model size (depth) of the neural network 110 is gradually expanded using the intermediate stacking technique across the sequence of training stages. In particular, by implementing the intermediate stacking technique, the depth of the neural network 110 grows linearly with the training stages.
[0113] FIG. 2 illustrates that the sequence of training stages includes three training stages. At the first training stage (“stage 1”), the stacking engine 130 is configured to expand the architecture of the neural network 110 by identifying an intermediate block from a sequence of blocks included in the neural network 110 as of the first training stage and using the identified intermediate block to generate a new block for inclusion in an updated sequence of blocks included in the neural network 110 as of the first training stage.
[0114] The first training stage follows a preceding training stage in the sequence of training stages. For example, when the preceding training stage is the initial training stage, the sequence of blocks as of the first training stage can include the blocks included as part of an initial architecture of the neural network 110, and include parameters having values resulting from the initial training stage.
[0115] Suppose that the initial architecture includes a total of three blocks (block 1, block 2, block 3) stacked one after another, i.e., arranged in a sequence with the block output of any block except the last being a block input to another of the blocks, implementing the intermediate stacking technique by the stacking engine 130 will result in the generation of a new block (block 4) for inclusion in the updated sequence of blocks as of the first training stage based on the middle block (block 2) in the stack of three blocks.
[0116] The updated sequence of blocks as of the first training stage will thus include four blocks. The four blocks include all three blocks (block 1, block 2, block 3) included in the initial architecture and the new block (block 4), where the new block (block 4) is arranged immediately preceding or immediately subsequent to the middle block (block 2) within the updated sequence of blocks.
[0117] For example, the updated sequence of blocks as of the first training stage can either include block 1, followed by block 2, followed by block 4, followed by block 3 (when the new block is arranged immediately preceding the middle block), or can alternatively include block 1,Attorney Docket No.: 45288-0633WO1followed by block 4, followed by block 2, followed by block 3 (when the new block is arranged immediately subsequent to the middle block).
[0118] The updated sequence of blocks as of the first training stage will be used as a sequence of blocks included in the neural network 110 as of the second training stage that follows the first training stage in the sequence of training stages.
[0119] At the second training stage (“stage 2”), the stacking engine 130 is configured to further expand the architecture of the neural network 110 by identifying an intermediate block from the sequence of blocks included in the neural network 110 as of the second training stage and using the identified intermediate block to generate a new block for inclusion in an updated sequence of blocks included in the neural network 110 as of the second training stage.
[0120] Continuing with the example above, implementing the intermediate stacking technique by the stacking engine 130 will result in the generation of a new block (block 5) for inclusion in the updated sequence of blocks as of the second training stage based on either block 2 or block 4.
[0121] The updated sequence of blocks as of the second training stage will thus include five blocks. The five blocks include all three blocks (block 1, block 2, block 3) included in the initial architecture, the new block added in the first training stage (block 4), and the new block added in the second training stage (block 5), where the new block added in the second training stage is inserted before or after any intermediate block within the updated sequence of blocks.
[0122] For example, the updated sequence of blocks as of the second training stage can include block 1, followed by block 2, followed by block 4, followed by block 5, followed by block 3. As another example, the updated sequence of blocks as of the second training stage can include block 1, followed by block 2, followed by block 5, followed by block 4, followed by block 3. As another example, the updated sequence of blocks as of the second training stage can include block 1, followed by block 5, followed by block 2, followed by block 4, followed by block 3.
[0123] The updated sequence of blocks as of the second training stage will be used as a sequence of blocks included in the neural network 110 as of the third training stage that follows the second training stage in the sequence of training stages.
[0124] At the third training stage (“stage 3”), the stacking engine 130 is configured to further expand the architecture of the neural network 100 by identifying an intermediate block from theAttorney Docket No.: 45288-0633WO1sequence of blocks included in the neural network 110 as of the third training stage and using the identified intermediate block to generate a new block for inclusion in an updated sequence of blocks included in the neural network 110 as of the third training stage.
[0125] Continuing with the example above, implementing the intermediate stacking technique by the stacking engine 130 will result in the generation of a new block (block 6) for inclusion in the updated sequence of blocks as of the third training stage based on one of: block 2, block 4, or block 5.
[0126] The updated sequence of blocks as of the third training stage will thus include six blocks. The six blocks include all three blocks (block 1, block 2, block 3) included in the initial architecture, the new block added in the first and second training stages (block 4, block 5), and the new block added in the third training stage (block 6), where the new block added in the third training stage is inserted before or after any intermediate block within the updated sequence of blocks.
[0127] For example, the updated sequence of blocks as of the third training stage can include block 1, followed by block 2, followed by block 4, followed by block 5, followed by block 6, followed by block 3. As another example, the updated sequence of blocks as of the second training stage can include block 1, followed by block 2, followed by block 5, followed by block 6, followed by block 4, followed by block 3. As another example, the updated sequence of blocks as of the second training stage can include block 1, followed by block 6, followed by block 5, followed by block 2, followed by block 4, followed by block 3.
[0128] The updated sequence of blocks as of the third training stage will be used as a sequence of blocks included in the neural network 110 as of the next training stage that follows the third training stage in the sequence of training stages (if the third training stage in not the last training stage in the sequence of training stages).
[0129] At each training stage, after the stacking engine 130 has expanded the architecture by updating the sequence of blocks to include a new block, the training engine 140 trains the neural network on training data to update the values of the parameters of the blocks included in the updated sequence of blocks.
[0130] The system 100 can continue in this order until the last training stage in the multistage training process, or until another termination criteria for the multi-stage training process ofAttorney Docket No.: 45288-0633WO1the neural network have been satisfied, e.g., until the parameters have converged, or until a threshold amount of wall clock time has elapsed.
[0131] FIG. 3 is an example illustration 300 of expanding a neural network using the intermediate stacking technique in comparison to expanding a neural network using a gradual stacking technique at a given training stage in the sequence of training stages.
[0132] In the example of FIG. 3, the neural network includes a sequence of blocks 302 that includes a total of three blocks (block 1, block 2, block 3) stacked one after another. Using the gradual stacking technique will result in the generation of a new block for inclusion in an updated sequence of blocks based on the last block (block 3) in the stack of three blocks.
[0133] For example, the new block can include the same number and type of layers, and a set of parameters having values that are identical to the trained values of the parameters of the last block (block 3). FIG. 3 thus illustrates that the updated sequence of blocks 304 generated by using the gradual stacking technique can include a duplication of the last block (block 3) that is arranged subsequent to the last block (block 3) within the updated sequence of blocks.
[0134] In contrast, using the intermediate stacking technique will result in the generation of a new block for inclusion in an updated sequence of blocks based on the middle block (block 2) in the stack of three blocks. Further, using the intermediate stacking technique will result in the placement of the new block at a different position within the updated sequence of blocks compared to using the gradual stacking technique.
[0135] For example, the new block can include the same number and type of layers, and a set of parameters having values that are identical to the trained values of the parameters of the middle block (block 2). FIG. 3 thus illustrates that, the updated sequence of blocks 306 generated by using the intermediate stacking technique can include a duplication of the middle block (block 2) that is arranged subsequent to the middle block (block 2) within the updated sequence of blocks.
[0136] Thus, expanding a neural network using the intermediate stacking technique differs from expanding a neural network using a gradual stacking technique at a given training stage in not only an original block from a neural network based on which a new block is generated, e.g., block 3 vs block 2, but also the position within the neural network at which the new block will be placed at, e.g., after block 3 vs before block 3.Attorney Docket No.: 45288-0633WO1
[0137] FIG. 4 is a flow diagram of an example process 400 for training a neural network at a given training stage in a sequence of training stages. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a training system, e.g., the training system 100 depicted in FIG. 1, appropriately programmed in accordance with this specification, can perform the process 400.
[0138] The system obtains training data for the given training stage (step 402). The training data includes a plurality of training network inputs. For example, the system can receive training data as an upload from a remote user of the system over a data communication network, e.g., using an application programming interface (API) made available by the system. The system can then use at least a subset of the uploaded training data as the training data for the given training stage. As another example, the system can receive an input from a user specifying which data that is already maintained by the system or another system that is accessibly by the system should be used as the training data.
[0139] For example, in implementations where the neural network is configured as an autoregressive neural network, the training network inputs can be training input sequences. Each training input sequence has a plurality of positions. Each position has a token selected from a vocabulary of tokens.
[0140] As mentioned above, the vocabulary of tokens can include one or more of characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in a corpus of natural language text and / or computer code. Additionally, or alternatively, the vocabulary of tokens can include tokens that can represent data other than text. For example, the vocabulary of tokens can include image tokens that represent a discrete set of image embeddings of an image that can be generated by an image encoder neural network based on processing the image. As another example, the vocabulary of tokens can include audio tokens that represent code vectors in a codebook of a quantizer, e.g., a residual vector quantizer.
[0141] For example, the training input sequences included in the training data can be generated from a large dataset of text in one or more natural languages, e g., text that is publicly available from the Internet or another text corpus, a large dataset of computer code in one or more programming languages, e.g., Python, C++, C#, Java, Ruby, PHP, and so on, e.g., computer code that is publicly available from the Internet or another code repository, a large dataset of audio samples, e.g., audio recordings or waveforms that represent the audioAttorney Docket No.: 45288-0633WO1recordings, a large dataset of images where each image includes an array of pixels, a large dataset of videos where each video includes a temporal sequence of frames, or a large multimodal dataset that includes a combination of two or more of these datasets.
[0142] As another example, in implementations where the neural network is configured as a diffusion neural network, the training network inputs can be training data items. The training data items can include, e.g., image data items, video data items, audio data items, and so on.
[0143] In some implementations, the training data includes unlabeled training network inputs, e.g., unlabeled training input sequences, whereas, in other implementations, the training data is labeled, e.g., such that the system has access to the labels or annotations associated with the training input sequences included in the training data, or the prompts that provide context for the training data items included in the training data.
[0144] The system identifies an intermediate block from the neural network (step 404). The neural network includes a sequence of blocks as of the given training stage. The intermediate block is arranged between a first block in the sequence of blocks and a last block in the sequence of blocks. The system refrains from identifying either the first block or the last block in the sequence of blocks as the intermediate block.
[0145] The rationale for selecting an intermediate block, rather than one at the extremities of the sequence, can be based on the technical consideration that intermediate blocks are likely to have learned feature representations (e g., of time-series data from sensors, or of spatial features in images) that are sufficiently abstract to be built upon for greater depth, yet not so specialized to the final output layer’s objective as to hinder further flexible learning. This can offer a more effective “seed” for network growth aimed at improving performance on complex technical tasks (such as anomaly detection in manufacturing or route optimization in logistics) or enhancing training efficiency by potentially reducing the total computational cost to reach a desired level of accuracy.
[0146] The given training stage follows a preceding training stage in the sequence of training stages. For example, when the preceding training stage is the initial training stage, the sequence of blocks as of the given training stage can include the blocks included as part of an initial architecture of the neural network.
[0147] The system generates a new block based on the intermediate block (step 406). The new block includes parameters having values that are derived from trained values of parametersAttorney Docket No.: 45288-0633WO1of the intermediate block as of the given training stage, i.e., from the values of the parameters of the intermediate block that have been learned during the training process up to and not including the given training stage.
[0148] In some implementations, the new block is generated as a duplication of the intermediate block. That is, the system initializes the new block to include the same number and type of layers as the intermediate block, and the system initializes the parameters of the new block to have values that are identical to the trained values of the parameters of the intermediate block as of the given training stage.
[0149] For example, the new block can be generated based on:Replicationwhere f represents the neural network, n represents the total number of layers, ft represents the ithlayer in the total of n layers included in the neural network, b represents the number of layers included in a block.
[0150] In this example, the new block is generated as a duplication of the intermediate block. The intermediate block is identified based on applying a ceiling function to the total number of layers divided by two, such that the sequence of blocks includes about equal numbers of blocks between the first block and intermediate block and between the intermediate block and last block.
[0151] In some implementations, when generating the new block based on the intermediate block, the system adds random noise to the trained values of the parameters of the intermediate block as of the given training stage, such that the new block includes the same number and type of layers as the intermediate block, but the parameters of the new block will have different values than those of the intermediate block.
[0152] The system expands the architecture of the neural network by updating the sequence of blocks to include both the intermediate block and the new block (step 408). That is, the system generates an updated sequence of blocks as of the given training stage. The updated sequence of blocks includes the new block, in addition to all of the blocks from the sequence of blocks as of the given training stage.Attorney Docket No.: 45288-0633WO1
[0153] Within the updated sequence of blocks, the intermediate block and the new block are arranged after the first block and before the last block in the updated sequence of blocks. When updating the sequence of blocks to include the new block, the system refrains from prepending the new block to the beginning of the sequence of blocks, and also refrains from appending the new block to the end of the sequence of blocks.
[0154] After updating the sequence of blocks, the system trains the neural network that now has an expanded architecture that includes the updated sequence of blocks as of the given training stage on the training data to adjust (update) the values of the parameters of the blocks included in the updated sequence of blocks, i.e., to determine the trained values of the parameters of the blocks (step 410). In doing so, the system obtains the trained values of the parameters of the blocks by concurrently training the updated sequence of blocks together.
[0155] The system trains the neural network across a plurality of training iterations that correspond to the given training stage based on optimizing one or more objective functions. The system will generally obtain different training network inputs at different training iterations, e.g., by sampling a fixed number of training network inputs from a larger number of training network inputs included in the obtained training data at each iteration.
[0156] The one or more objective functions can include any of a variety of objective functions, including unsupervised or self-supervised objective functions and supervised objective functions, that are appropriate for training the neural network.
[0157] For example, in implementations where the neural network is configured as an autoregressive neural network, the objective functions can include a next token prediction objective function, e.g., a maximum-likelihood objective function. Examples of the next token prediction objective functions include those described in Tay, et al. U12: Unifying language learning paradigms. In The Eleventh International Conference on Learning Representations, 2022.
[0158] As another example, in implementations where the neural network is configured as a diffusion neural network, the objective functions can include a score matching objective function. Examples of the score matching objective functions include those described in Song, et al., Generative modeling by estimating gradients of the data distribution. Advances in neural information processing systems, 32, 2019, and Song, et al., Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2020.Attorney Docket No.: 45288-0633WO1
[0159] In implementations where the new block is generated as a duplication of the intermediate block, when training the neural network on the training data, the system adjusts the values of the parameters of the new block beginning from the trained values of the parameters of the intermediate block as of the given training stage, and then adjusts, over the course of the training, the values of parameters of the new block to be different from the trained values of the parameters of the intermediate block as of the given training stage.
[0160] In general, the system trains the blocks in the updated sequence, including the new block and the intermediate block, independently from each other, so that the values of the parameters of new block and intermediate block will be different after the given training stage (although they may begin with being identical to each other). In other words, the system trains the neural network without constraining the values of parameters of one block based on the values of parameters of another block.
[0161] The system can repeatedly perform an iteration of the process 400 for each training stage in the sequence of training stages in a multi-stage training process.
[0162] FIG. 5 is a flow diagram of an example process 500 for training a neural network at another given training stage in the sequence of training stages. For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a training system, e.g., the training system 100 depicted in FIG. 1, appropriately programmed in accordance with this specification, can perform the process 500.
[0163] The system obtains training data for the other given training stage (step 502). The training data includes a plurality of training network inputs. How the system can obtain the training data as well as what can be included in the training data is described above with reference to step 402 of FIG. 4.
[0164] The system identifies an other intermediate block from the neural network (step 504). The neural network includes a sequence of blocks as of the other given training stage. The other intermediate block is arranged between a first block in the sequence of blocks and a last block in the sequence of blocks. The system refrains from identifying either the first block or the last block in the sequence of blocks as the other intermediate block.
[0165] For example, the other given training stage can be a next training stage that immediately follows the given training stage discussed above with reference to FIG. 4 in the sequence of training stages, and the sequence of blocks as of the other given training stage canAttorney Docket No.: 45288-0633WO1include the blocks included as part of the expanded architecture of the neural network that was generated in that given training stage. In this example, the system can identify either the intermediate block or the new block mentioned above with reference to FIG. 4 as the other intermediate block.
[0166] The system generates an other new block based on the other intermediate block (step 506). The other new block includes parameters having values that are derived from trained values of parameters of the other intermediate block as of that given training stage, i.e., from the values of the parameters of the other intermediate block that have been learned during the training process up to and not including the other given training stage. In some implementations, the other new block is generated as a duplication of the other intermediate block.
[0167] The system expands the architecture of the neural network by updating the sequence of blocks to include the other new block (step 508). That is, the system generates an updated sequence of blocks as of the other given training stage. The updated sequence of blocks includes the other new block, in addition to all of the blocks from the sequence of blocks as of the other given training stage. For example, the updated sequence of blocks can also include the other intermediate block mentioned above with reference to FIG. 5 and the new block and intermediate block mentioned further above with reference to FIG. 4.
[0168] Within the updated sequence of blocks, the other intermediate block and the other new block are arranged after the first block and before the last block in the updated sequence of blocks. When updating the sequence of blocks to include the other new block, the system refrains from prepending the other new block to the beginning of the sequence of blocks, and also refrains from appending the other new block to the end of the sequence of blocks.
[0169] For example, the other new block can be inserted immediately preceding the other intermediate block in the updated sequence of blocks. As another example, the other new block can be inserted immediately subsequent to the other intermediate block in the updated sequence of blocks.
[0170] After updating the sequence of blocks, the system trains the neural network that now has an expanded architecture that includes the updated sequence of blocks as of the other given training stage on the training data to adjust (update) the values of the parameters of the blocks included in the updated sequence of blocks, i.e., to determine the trained values of the parametersAttorney Docket No.: 45288-0633WO1of the blocks (step 510). How the system can perform the training is described above with reference to step 410 of FIG. 4.
[0171] In some implementations, after having trained the neural network across the sequence of training stages in the multi-stage training process, the training system or another system, e.g., a fine-tuning system, fine-tunes, i.e., further trains, the neural network to perform one or more downstream tasks. The one or more downstream tasks can include any of the tasks mentioned above, and possibly other tasks.
[0172] That is, a neural network that can be deployed to compute inference for the one or more downstream tasks, e.g., performed generative tasks in response to prompts received from users, can be generated from a neural network that has been trained by a training system implementing the multi-stage training process. The neural network includes the updated sequence of blocks that has been updated as of the last training stage in the sequence of training stages and that includes parameters having values resulting from the sequence of training stages.
[0173] In these implementations, after the last training stage in the sequence of training stages has completed, the system or the other system proceeds to adapt the neural network, e.g., through supervised fine-tuning or reinforcement learning from human feedback (RLHF), on labeled or unlabeled training data that is specific to the downstream task to perform a downstream task.
[0174] Notably, the intermediate stacking technique provides the neural network with a desirable inductive bias toward performance improvement on many downstream tasks, especially including complex reasoning tasks, generative tasks, and arithmetic tasks. In particular a neural network trained using the intermediate stacking technique will achieve improved downstream evaluations when trained on the same number of tokens as baseline training.
[0175] FIG. 6 shows four downstream evaluations vs validation log perplexity isoplots for a baseline neural network trained by a training system implementing a conventional training technique and a neural network trained by a training system implementing the intermediate stacking technique described in this specification on the same data.
[0176] The baseline neural network corresponds to a neural network trained by a training system implementing a standard training technique, e.g., which keeps the architecture of the neural network fixed throughout the training, and the MIDAS neural network corresponds to aAttorney Docket No.: 45288-0633WO1neural network trained by a training system implementing the intermediate stacking technique. The MIDAS neural network can be trained 24% faster than the baseline neural network.
[0177] On the y-axis, each isoplot shows the performance of the baseline and MIDAS neural networks on various task groups: closed book question answering (QA) tasks, open book QAs, math word problems, and reasoning primitives. On the x-axis, each isoplot shows the log perplexity in the reverse order, thus the downstream performance for both neural networks improves as log perplexity decreases.
[0178] For closed book QA tasks, the MIDAS neural network has largely similar trends to the baseline neural network. For open book QA tasks, math word problems, and reasoning primitives, the MIDAS neural network has much better downstream performance at an equivalent log perplexity. This showcases the inductive bias of the MIDAS neural network towards better overall quality and better reasoning abilities.
[0179] The observed improvements shown in FIG. 6, particularly for open book QA (which can be analogous to retrieving and reasoning over technical manuals), math word problems (reflecting structured reasoning applicable to engineering calculations or logistical planning), and reasoning primitives (foundational logical operations), may be attributed to a technical effect of the intermediate stacking method. By duplicating an intermediate block, the neural network is encouraged to develop deeper and more refined processing pathways for mid-level semantic and structural representations of the input data. These mid-level representations can be important for tasks that require chaining multiple reasoning steps or integrating information from broader contexts. Unlike stacking at the very end, which might just refine a final output representation, or at the beginning, which adds to low-level feature extraction without necessarily improving higher-order reasoning, intermediate stacking focuses on enhancing the network’s capacity for complex transformations within its “core reasoning” layers. This can lead to a trained model that may be computationally cheaper to obtain (due to potentially faster training convergence, e.g., requiring fewer training epochs or less wall-clock time on a given hardware platform) and may also possess an internal architecture more adept at these specific types of technical information processing tasks.
[0180] The reasoning primitives are synthetic tasks that represent building blocks of a reasoning task. They include: induction copying (a task which presents a sequence of words, followed by a subsequence selected randomly from within this original sequence, and requiresAttorney Docket No.: 45288-0633WO1the neural network to output the next word in the sequence), variable assignment (a task which requires the neural network to associate a value with a variable name), and pre-school math (a task which requires the neural network to solve a math problem by correctly associating multiple values and variables simultaneously and applying this association to a particular task).
[0181] An example of the induction copying task is: “pum nyj gdq ocu rzk jbw mlz eny kyx uni rzk jbw mlz eny kyx”, and the expected output is “uni”.
[0182] The variable assignment tasks have different “depths”. An example of the depth-0 variable assignment task is “u=l; t=0; v=13; y=4; f=22; y- ’, and the expected output is 4. An example of the depth-2 variable assignment task is “y=7; f=0; z=3; b=9; x=8; q=y; l=f; m=z; h=x; a=b; n=h; j=m; t=a; i=l; g=q; n=”, and the expected output is 8.
[0183] An example of the pre-school math task is “z=6; b=5; i=-z+b; i=”, and the expected answer (with chain-of-thought) is “-6+5=-l”.
[0184] FIG. 7 shows test accuracy improvements of a neural network trained by a training system implementing the intermediate stacking technique described in this specification compared to a baseline neural network trained by a training system implementing a conventional training technique and a neural network on reasoning primitives. The MIDAS neural network shows clear improvements over the baseline neural network on most of the reasoning primitives.
[0185] In the context of the intermediate stacking technique, where a new block may be initialized based on an existing intermediate block, a further technical consideration for achieving or enhancing a desirable inductive bias can relate to the subsequent evolution of functional diversity between different blocks within the neural network. A greater similarity measure between two constituent blocks of the neural network generally indicates a greater functional similarity between the two blocks. In practice, training two blocks to become functionally dissimilar to each other is one way of further providing the neural network with the desirable inductive bias toward performance improvement on many downstream tasks. For instance, while initializing a new block as a copy of an intermediate block provides a strong starting point by leveraging learned parameters, the subsequent independent training may facilitate these blocks in differentiating and specializing their functions. If blocks were to remain too functionally similar after such initialization and further training, it might suggest that the added depth is not contributing as efficiently as possible to new representational capacity, potentially leading to underutilized computational resources during inference. Therefore,Attorney Docket No.: 45288-0633WO1allowing or encouraging constituent blocks, including a newly added block and its “parent” intermediate block, to become functionally dissimilar during the training process subsequent to the stacking operation can be one way of helping the network effectively leverage the increased depth. This may contribute to making the model more computationally effective and further enhance its performance on various downstream tasks.
[0186] For example, the similarity between the two blocks can be computed as a cosine similarity, a dot product, or some other similar similarity measure, between the trained values of the respective parameters of the two blocks. As another example, the similarity between the two blocks can be computed as a cosine similarity, a dot product, or some other similar similarity measure, between the respective hidden states generated by the two blocks.
[0187] The intermediate stacking approach, by initializing a new block with learned parameters from an already functional intermediate block, might initially create high functional similarity between the source block and the newly added block. However, the subsequent independent training of these blocks facilitates their divergence and specialization, effectively tailoring the network's deeper architecture to the specifics of the technical problem being solved. Monitoring this divergence (e.g., by periodically calculating similarity metrics between the parameters or outputs of these blocks) can be a useful diagnostic tool. For example, if the newly added block and its 'parent' intermediate block remain highly similar in function even after substantial further training, it might indicate that the current network depth or architecture is not effectively utilizing the additional parameters for the given task or data, potentially suggesting adjustments to the training regime or the choice of which intermediate block to duplicate. This monitoring can provide a technical insight into the training dynamics and may guide further optimization of the training process or the model architecture itself.
[0188] Thus, in some implementations, the system monitors the similarity measure between constituent blocks of the neural network throughout the training process and, in the event of an increased similarity measure, takes remedial actions to lower the similarity measure.
[0189] For example, the system can compute a similarity measure between the intermediate block and the last block based on the trained values of the parameters of the intermediate block as of the given training stage mentioned above with reference to FIG. 4 and trained values of the parameters of the last block as of the other given training stage mentioned above with reference to FIG. 5. In response to determining that the similarity measure between is greater than aAttorney Docket No.: 45288-0633WO1threshold similarity measure, the system can pause the training process to perform a remedial action to lower the similarity measure.
[0190] For example, the remedial action can be a restoration action that restores the parameters of the last block to have trained values learned as of the given training stage. As another example, the remedial action can be a reinitialization action that reinitializes the parameters of the intermediate block to have different values.
[0191] In this specification, the term “configured” is used in relation to computing systems and environments, as well as computer program components. A computing system or environment is considered "configured" to perform specific operations or actions when it possesses the necessary software, firmware, hardware, or a combination thereof, enabling it to carry out those operations or actions during operation. For instance, configuring a system might involve installing a software library with specific algorithms, updating firmware with new instructions for handling data, or adding a hardware component for enhanced processing capabilities. Similarly, one or more computer programs are “configured” to perform particular operations or actions when they contain instructions that, upon execution by a computing device or hardware, cause the device to perform those intended operations or actions.
[0192] The embodiments and functional operations described in this specification can be implemented in various forms, including digital electronic circuitry, software, firmware, computer hardware (encompassing the disclosed structures and their structural equivalents), or any combination thereof. The subject matter can be realized as one or more computer programs, essentially modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by or to control the operation of a computing device or hardware. The storage medium can be a storage device such as a hard drive or solid-state drive (SSD), a storage medium, a random or serial access memory device, or a combination of these.Additionally or alternatively, the program instructions can be encoded on a transmitted signal, such as a machine-generated electrical, optical, or electromagnetic signal, designed to carry information for transmission to a receiving device or system for execution by a computing device or hardware. Furthermore, implementations may leverage emerging technologies like quantum computing or neuromorphic computing for specific applications, and may be deployed in distributed or cloud-based environments where components reside on different machines or within a cloud infrastructure.Attorney Docket No.: 45288-0633WO1
[0193] The term “computing device or hardware” refers to the physical components involved in data processing and encompasses all types of devices and machines used for this purpose. Examples include processors or processing units, computers, multiple processors or computers working together, graphics processing units (GPUs), tensor processing units (TPUs), and specialized processing hardware such as field-programmable gate arrays (FPGAs) or applicationspecific integrated circuits (ASICs). In addition to hardware, a computing device or hardware may also include code that creates an execution environment for computer programs. This code can take the form of processor firmware, a protocol stack, a database management system, an operating system, or a combination of these elements. Embodiments may particularly benefit from utilizing the parallel processing capabilities of GPUs, in a General-Purpose computing on Graphics Processing Units (GPGPU) context, where code specifically designed for GPU execution, often called kernels or shaders, is employed. Similarly, TPUs excel at running optimized tensor operations crucial for many machine learning algorithms. By leveraging these accelerators and their specialized programming models, the system can achieve significant speedups and efficiency gains for tasks involving artificial intelligence and machine learning, particularly in areas such as computer vision, natural language processing, and robotics.
[0194] A computer program, also referred to as software, an application, a module, a script, code, or simply a program, can be written in any programming language, including compiled or interpreted languages, and declarative or procedural languages. It can be deployed in various forms, such as a standalone program, a module, a component, a subroutine, or any other unit suitable for use within a computing environment. A program may or may not correspond to a single file in a file system and can be stored in various ways. This includes being embedded within a file containing other programs or data (e.g., scripts within a markup language document), residing in a dedicated file, or distributed across multiple coordinated files (e.g., files storing modules, subprograms, or code segments). A computer program can be executed on a single computer or across multiple computers, whether located at a single site or distributed across multiple sites and interconnected through a data communication network. The specific implementation of the computer programs may involve a combination of traditional programming languages and specialized languages or libraries designed for GPGPU programming or TPU utilization, depending on the chosen hardware platform and desired performance characteristics.Attorney Docket No.: 45288-0633WO1
[0195] In this specification, the term “engine” broadly refers to a software-based system, subsystem, or process designed to perform one or more specific functions. An engine is typically implemented as one or more software modules or components installed on one or more computers, which can be located at a single site or distributed across multiple locations. In some instances, one or more dedicated computers may be used for a particular engine, while in other cases, multiple engines may operate concurrently on the same one or more computers. Examples of engine functions within the context of Al and machine learning could include data preprocessing and cleaning, feature engineering and extraction, model training and optimization, inference and prediction generation, and post-processing of results. The specific design and implementation of engines will depend on the overall architecture and the distribution of computational tasks across various hardware components, including CPUs, GPUs, TPUs, and other specialized processors.
[0196] The processes and logic flows described in this specification can be executed by one or more programmable computers running one or more computer programs to perform functions by operating on input data and generating output. Additionally, graphics processing units (GPUs) and tensor processing units (TPUs) can be utilized to enable concurrent execution of aspects of these processes and logic flows, significantly accelerating performance. This approach offers significant advantages for computationally intensive tasks often found in Al and machine learning applications, such as matrix multiplications, convolutions, and other operations that exhibit a high degree of parallelism. By leveraging the parallel processing capabilities of GPUs and TPUs, significant speedups and efficiency gains compared to relying solely on CPUs can be achieved. Alternatively or in combination with programmable computers and specialized processors, these processes and logic flows can also be implemented using specialized processing hardware, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), for even greater performance or energy efficiency in specific use cases.
[0197] Computers capable of executing a computer program can be based on general-purpose microprocessors, special -purpose microprocessors, or a combination of both. They can also utilize any other type of central processing unit (CPU). Additionally, graphics processing units (GPUs), tensor processing units (TPUs), and other machine learning accelerators can be employed to enhance performance, particularly for tasks involving artificial intelligence andAttorney Docket No.: 45288-0633WO1machine learning. These accelerators often work in conjunction with CPUs, handling specialized computations while the CPU manages overall system operations and other tasks. Typically, a CPU receives instructions and data from read-only memory (ROM), random access memory (RAM), or both. The elements of a computer include a CPU for executing instructions and one or more memory devices for storing instructions and data. The specific configuration of processing units and memory will depend on factors like the complexity of the Al model, the volume of data being processed, and the desired performance and latency requirements. Embodiments can be implemented on a wide range of computing platforms, from small embedded devices with limited resources to large-scale data center systems with high-performance computing capabilities. The system may include storage devices like hard drives, SSDs, or flash memory for persistent data storage.
[0198] Computer-readable media suitable for storing computer program instructions and data encompass all forms of non-volatile memory, media, and memory devices. Examples include semiconductor memory devices such as read-only memory (ROM), solid-state drives (SSDs), and flash memory devices; hard disk drives (HDDs); optical media; and optical discs such as CDs, DVDs, and Blu-ray discs. The specific type of computer-readable media used will depend on factors such as the size of the data, access speed requirements, cost considerations, and the desired level of portability or permanence.
[0199] To facilitate user interaction, embodiments of the subject matter described in this specification can be implemented on a computing device equipped with a display device, such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display, for presenting information to the user. Input can be provided by the user through various means, including a keyboard), touchscreens, voice commands, gesture recognition, or other input modalities depending on the specific device and application. Additional input methods can include acoustic, speech, or tactile input, while feedback to the user can take the form of visual, auditory, or tactile feedback. Furthermore, computers can interact with users by exchanging documents with a user's device or application. This can involve sending web content or data in response to requests or sending and receiving text messages or other forms of messages through mobile devices or messaging platforms. The selection of input and output modalities will depend on the specific application and the desired form of user interaction.Attorney Docket No.: 45288-0633WO1
[0200] Machine learning models can be implemented and deployed using machine learning frameworks, such as TensorFlow or JAX. These frameworks offer comprehensive tools and libraries that facilitate the development, training, and deployment of machine learning models.
[0201] Embodiments of the subject matter described in this specification can be implemented within a computing system comprising one or more components, depending on the specific application and requirements. These may include a back-end component, such as a back-end server or cloud-based infrastructure; an optional middleware component, such as a middleware server or application programming interface (API), to facilitate communication and data exchange; and a front-end component, such as a client device with a user interface, a web browser, or an app, through which a user can interact with the implemented subject matter. For instance, the described functionality could be implemented solely on a client device (e.g., for on-device machine learning) or deployed as a combination of front-end and back-end components for more complex applications. These components, when present, can be interconnected using any form or medium of digital data communication, such as a communication network like a local area network (LAN) or a wide area network (WAN) including the Internet. The specific system architecture and choice of components will depend on factors such as the scale of the application, the need for real-time processing, data security requirements, and the desired user experience.
[0202] The computing system can include clients and servers that may be geographically separated and interact through a communication network. The specific type of network, such as a local area network (LAN), a wide area network (WAN), or the Internet, will depend on the reach and scale of the application. The client-server relationship is established through computer programs running on the respective computers and designed to communicate with each other using appropriate protocols. These protocols may include HTTP, TCP / IP, or other specialized protocols depending on the nature of the data being exchanged and the security requirements of the system. In certain embodiments, a server transmits data or instructions to a user's device, such as a computer, smartphone, or tablet, acting as a client. The client device can then process the received information, display results to the user, and potentially send data or feedback back to the server for further processing or storage. This allows for dynamic interactions between the user and the system, enabling a wide range of applications and functionalities.Attorney Docket No.: 45288-0633WO1
[0203] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment.Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0204] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0205] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0206] What is claimed is:
Claims
Attorney Docket No.: 45288-0633WO1CLAIMS1. A method performed by one or more computing devices for training a neural network across a sequence of training stages, the neural network comprising a sequence of blocks, and the method comprising, at a first training stage in the sequence of training stages:obtaining training data for the first training stage;identifying an intermediate block from the sequence of blocks as of the first training stage, wherein the intermediate block is arranged between a first block in the sequence of blocks and a last block in the sequence of blocks;generating a new block based on the intermediate block, wherein the new block comprises parameters having values that are derived from trained values of parameters of the intermediate block as of the first training stage;updating the sequence of blocks to include both the intermediate block and the new block arranged between the first block and the last block in the sequence of blocks; andafter updating the sequence of blocks, training the neural network on the training data.
2. The method of claim 1, wherein the intermediate block and the new block comprise a same number and type of layers.
3. The method of any one of claims 1-2, wherein generating the new block comprises initializing parameters of the new block to have values that are identical to the trained values of the parameters of the intermediate block as of the first training stage.
4. The method of claim 3, wherein training the neural network on the training data comprises adjusting the values of the parameters of the new block beginning from the trained values of the parameters of the intermediate block as of the first training stage.
5. The method of any one of claims 3-4, wherein training the neural network on the training data comprises adjusting the values of parameters of the new block to be different from the trained values of the parameters of the intermediate block as of the first training stage.Attorney Docket No.: 45288-0633WO16. The method of any one of claims 1-5, further comprising, at a second training stage that follows the first training stage in the sequence of training stages:obtaining training data for the second training stage;identifying an other intermediate block from the sequence of blocks as of the second training stage, wherein the other intermediate block is arranged between the first block in the sequence of blocks and the last block in the sequence of blocks;generating an other new block based on the other intermediate block, wherein the other new block comprises parameters having values that are derived from trained values of parameters of the other intermediate block as of the second training stage;updating the sequence of blocks to include the intermediate block, the new block, the other intermediate block, and the other new block arranged between the first block and the last block in the sequence of blocks; andafter updating the sequence of blocks, training the neural network on the training data.
7. The method of claim 6, wherein the other intermediate block is either the intermediate block or the new block.
8. The method of one of claims 6-7, wherein the other new block is immediately after the other intermediate block in the sequence of blocks.
9. The method of one of claims 6-7, wherein the other new block is immediately before the other intermediate block in the sequence of blocks.
10. The method of any one of claims 6-9, wherein the other intermediate block and the other new block comprise a same number and type of layers.
11. The method of any one of claims 1-10, wherein the first block does not comprise an embedding layer.
12. The method of any one of claims 1-11, wherein the first block does not comprise one or more neural network layers from a separately trained neural network.Attorney Docket No.: 45288-0633WO113. The method of any one of claims 1-12, wherein the last block does not comprise a deembedding layer.
14. The method of any one of claims 1-13, wherein the last block does not comprise a softmax layer.
15. The method of any one of claims 1-14, wherein the sequence of training stages follows an initial training stage, and wherein prior to the first training stage in the sequence of training stages, the sequence of blocks comprises parameters having values resulting from the initial training stage.
16. The method of any one of claims 1-15, wherein the intermediate block receives a hidden state from a preceding block in the sequence of blocks and generates an updated hidden state for a subsequent block in the sequence of blocks.
17. The method of any one of claims 1-16, wherein the intermediate block and the first block comprise a same number and type of layers.
18. The method of any one of claims 1-18, wherein the first block is not a separately pretrained neural network model integrated into the sequence of blocks.
19. The method of any one of claims 1-18, wherein the trained values of the parameters were obtained by concurrently training the sequence of blocks together.
20. The method of any one of claims 1-19, further comprising, after training the neural network across the sequence of training stages:training the neural network on task-specific training data to perform a downstream task, wherein the neural network comprises the sequence of blocks that has been updated after a last training stage in the sequence of training stages.Attorney Docket No.: 45288-0633WO121. The method of claim 20, wherein the training data used across the sequence of training stages comprises unlabeled data, and wherein training the neural network across the sequence of training stages comprises training the neural network based on optimizing one or more unsupervised or self-supervised objective functions.
22. The method of any one of claims 20-21, further comprising, after training the neural network to perform the downstream task:receiving a network input; andprocessing the network input using the neural network to generate a network output for the downstream task.
23. A method comprising:receiving a network input for a downstream task; andprocessing the network input using a neural network to generate a network output for the downstream task, wherein the neural network comprises a sequence of blocks, wherein the neural network has been previously trained according to a process comprising a sequence of training stages, wherein at each of the plurality of training stages, a duplication of an intermediate block in the sequence of blocks is added to the neural network, and wherein the intermediate block is arranged between a first block in the sequence of blocks and a last block in the sequence of blocks.
24. The method of any preceding claim, wherein the downstream task is a generative task, wherein the network input comprises a conditioning input, and wherein the network output comprises a data item.
25. The method of claim 24, wherein the data item comprises one or more of: text data, image data, video data, or audio data.
26. The method of any preceding claim, wherein the neural network has a Transformer based architecture, and wherein the intermediate block comprises one or more attention layers.Attorney Docket No.: 45288-0633WO127. The method of any preceding claim, further comprising, at the second training stage in the sequence of training stages:computing a similarity measure between the intermediate block and the last block based on the trained values of the parameters of the intermediate block as of the second training stage and trained values of the parameters of the last block as of the second training stage; andin response to determining that the similarity measure between is greater than a threshold similarity measure, pausing the training to perform a remedial action to lower the similarity measure.
28. The method of claim 27, wherein the remedial action comprises one of:restoring the parameters of the last block to have trained values learned as of the first training stage, orreinitializing the parameters of the intermediate block to have different values.
29. A system comprising:one or more computers; andone or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the respective operations of any one of claims 1-28.
30. One or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the respective operations of the method of any one of claims 1-28.