Low-rank side adaption for machine learning models

A parallel side network with a low-rank adaptor function efficiently adapts backbone networks to downstream tasks, addressing the computational and memory challenges of finetuning, enhancing accuracy and privacy on devices with limited resources.

WO2025151511A1PCT designated stage expired Publication Date: 2025-07-17GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/010729
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-11
Filing Date
2025-01-08
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing machine learning models, particularly foundation models, require computationally expensive finetuning for downstream tasks, which is time-consuming and memory-intensive, especially on devices with constrained resources.

Method used

Employ a parallel side network that processes the outputs of a backbone network to adapt to downstream tasks without backpropagating gradients through the backbone, using a low-rank adaptor function to refine the backbone network outputs, allowing efficient training and inference on multiple tasks with reduced memory and time requirements.

Benefits of technology

This approach reduces memory and training time while maintaining high accuracy, enabling efficient adaptation and personalization of backbone networks on devices with limited resources, while ensuring privacy and security by preventing access to the backbone network parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025010729_17072025_PF_FP_ABST
    Figure US2025010729_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates generally to the use of a parallel side network to adapt a backbone network to a task. According to a first aspect of this specification, there is provided a method implemented by one or more data processing apparatus, the method comprising generating, by a backbone network comprising a plurality of layers, a plurality of backbone network outputs from a network input. The plurality of backbone network outputs comprises: a final backbone output comprising the output of a final layer of the backbone network; and one or more intermediate backbone network outputs, wherein an intermediate backbone network output comprises an output of a layer of the backbone network prior to the final layer of the backbone network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] LOW-RANK SIDE ADAPTION FOR MACHINE LEARNING MODELS

[0002] Related Applications

[0003] This application claims priority to and the benefit of United States Provisional Patent Application Number 63 / 620,054, filed January 11, 2024. United States Provisional Patent Application Number 63 / 620,054 is hereby incorporated by reference in its entirety.

[0004] Field

[0005] The present disclosure relates generally to the use of a parallel side network to adapt a backbone network to a task.

[0006] Background

[0007] This specification relates to performing a machine learning task on a network input using neural networks.

[0008] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.

[0009] Summary

[0010] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description and claims, or can be learned from the description, or can be learned through practice of the embodiments.

[0011] According to a first aspect of this specification, there is provided a method implemented by one or more data processing apparatus, the method comprising generating, by a backbone network comprising a plurality of layers, a plurality of backbone network outputs from a network input. The plurality of backbone network outputs comprises: a final backbone output comprising the output of a final layer of the backbone network; and one or more intermediate backbone network outputs, wherein an intermediate backbone network output comprises an output of a layer of the backbone network prior to the final layer of the backbone network. The method further comprises generating, by a parallel side network comprising a plurality of layers, a parallel side network output from the final backbone network output, comprising, for each of a plurality of layers of the parallel side network: receiving a respective backbone network output; and processing a respective layer input and the respective backbone network output using a respective layer of the parallel side network to generate a respective parallel side network layer output.

[0012] The method further comprises generating a final network output based on the parallel side network output.

[0013] According to a further aspect of this specification, there is provided a method implemented by one or more data processing apparatus, the method comprising generating a candidate network output from a training input using a backbone network and a parallel side network, comprising : generating, using the backbone network, a plurality of backbone network outputs from the training input, the plurality of backbone network outputs comprising a final backbone network output and one or more intermediate backbone network outputs; and generating, using the parallel side network, the candidate network output from the plurality of backbone network outputs.

[0014] The method further comprises updating parameters of the parallel side network based on an objective function depending on the candidate network output. The updating comprises: determining gradients of the objective function with respect to the parameters of the parallel side network, comprising backpropagating gradients of the objective function through the parallel side network without backpropagating gradients through the backbone network; and determining updates to the parameters of the parallel side network based on the gradients of the objective function with respect to the parameters of the parallel side network.

[0015] Generating, using the parallel side network, the candidate network output from the plurality of backbone network outputs may comprise, for each of a plurality of layers of the parallel side network: receiving a respective backbone network output; combining a respective layer input and the respective backbone network output to generate a combined input; and processing the combined input using a respective layer of the parallel side network to generate a respective parallel side network layer output. The objective function may further depend on a ground truth output corresponding to the training input.

[0016] These, and other, aspects of this specification may include one or more of the following features, either alone or in combination.

[0017] Processing the respective layer input and the respective backbone network output using a respective layer of the parallel side network to generate a respective parallel side network layer output may comprise applying an adaptor function comprising a low rank factorisation of a projection of a combined input comprising a combination of the respective layer input and the respective backbone network output. Applying the adaptor function may comprise: generating a low-dimensional representation of the combined input from the combined input for the respective layer using a downprojection operation; applying a nonlinear function to low-dimensional representation of the combined input; and generating an adaptor function output from the output of the nonlinear function using an up-projection operation.

[0018] The down-projection operation may comprise applying a down-projection matrix of dimension d x r to the combined input, wherein d is a dimension of the combined input, r is a dimension of the low-dimensional representation of the combined input and d > r. The up-projection operation comprises applying an up-projection matrix of dimension r x d to the adaptor function output.

[0019] Processing the combined input using a respective layer of the parallel side network to generate a respective parallel side network layer output may further comprises combining the combined input for the respective layer with the adaptor function output.

[0020] Generating the parallel side network output from the final backbone network output may comprise alternating between applying adaptor functions in a channel dimension of respective combined inputs and applying adaptor functions in a token dimension of respective combined inputs.

[0021] The respective layer input for an initial layer of the parallel side network may comprise the final backbone network output. The respective layer inputs for subsequent layers of the parallel side network may comprise a respective parallel side network layer output from a preceding layer of the parallel side network. The respective backbone network output for a final layer of the parallel side network may comprise the final backbone network output.

[0022] Generating the final network output based on the parallel side network output may comprise generating the final neural network output from the parallel side network output using a classifier network.

[0023] The backbone network and the parallel side network may each comprise neural networks. The backbone network may comprise a transformer network.

[0024] The neural network input may comprise one or more images. The final neural network output may comprise: one or more image classification outputs; one or more object detection outputs; one or more image segmentations; one or more denoised images; one or more infilled images; one or more enhanced images; one or more depth estimations; and / or one or more image captions.

[0025] Other embodiments of these aspects include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.

[0026] The subject matter described in this specification can be implemented so as to realize one or more of the following advantages. The use of a model (e.g., a neural network) comprising a parallel side network in addition to a backbone network allows the backbone network of the model to be adapted to downstream tasks without backpropagating gradients through the backbone network. Consequently, the size of a memory space required for adapting the backbone network to a task, as well as the time for adapting the backbone network to the task, can be reduced, since gradients for the backbone network are not calculated and stored. This can, for example, allow the backbone network to be adapted to a task on a device / system with a constrained memory space, such as on a mobile device. When the model is executed in a distributed manner, with the backbone network being executed remotely and the parallel side network executed locally on a user device, this can further allow the model to be personalised / adapted to a task at the user device, increasing security and privacy. In such examples, privacy and security of the backbone model can also be increased by preventing third parties having access to the backbone network parameters, while allowing some finetuning / adaption of the backbone model output by third parties through use of the parallel side network(s).

[0027] Furthermore, the use of the backbone network output as input to the parallel side network can increase the accuracy of the model on the downstream task. The use of skip connections in the parallel side network can ensure that knowledge of the backbone network output is propagated to later layers of the parallel side network, which can further increase the accuracy of the model on the downstream task.

[0028] Additionally, the use of a low rank adaptor function can allow the backbone network to be adapted to a downstream task with a low number of additional parameters. This provides a lightweight parallel side network that can be trained efficiently while maintaining high accuracy, requiring less memory during training and faster training than comparable approaches.

[0029] Inference and training efficiency for multiple tasks can also be improved by allowing a single pass of the backbone network to be used for inference by / training of multiple parallel side models, each performing a different task on the same input data. In such examples, inference by multiple parallel side models and / or the training of multiple parallel side models can be performed in parallel with only a single forward pass of the backbone network being used. Alternatively or additionally, the results of a forward pass through the backbone network (i.e., final and intermediate outputs of the backbone network) can be stored and later retrieved for use to train parallel side networks for particular tasks on demand.

[0030] The use of a single backbone network for multiple tasks can reduce the memory requirements for storing a set of models for a corresponding set of tasks, as the models shares a common backbone network that only needs to be stored once on a device.

[0031] Brief Description of the Drawings

[0032] FIG. 1 shows an overview of an example system / method for generating a network output from a network input using a backbone network and a parallel side network; FIG. 2 shows an overview of an example adaptor function; FIG. 3A shows an example of a computational graph for a self-attention layer of a backbone network during full fine-tuning;

[0033] FIG. 3B shows an example of a computational graph for a self-attention layer of a backbone network using a low rank adaption approach;

[0034] FIG. 3C shows an example of a computational graph for a self-attention layer of a backbone network using a parallel side network approach;

[0035] FIG. 4 shows a flow diagram of an example method for generating a network output from a network input using a backbone network and a parallel side network;

[0036] FIG. 5 shows a flow diagram of an example method for training a parallel side network to adapt the output of a backbone network for a task;

[0037] FIG. 6 shows a schematic example of a computing system / apparatus for performing any of the methods, operations or processes described herein; and

[0038] FIG. 7 shows an example non-transitory computer readable medium according to some implementations.

[0039] Detailed Description

[0040] Foundation models / neural networks are large machine-learned models, trained on massive datasets, that typically have diverse abilities across a range of applications. The use of foundation models in machine learning is becoming increasingly common. While such models can perform well on a range of tasks, performance can be improved by finetuning the models on individual tasks. However, such finetuning can be computationally expensive, for example, in terms of time taken to finetune the foundation model and / or the memory footprint used when finetuning the model.

[0041] This specification describes methods, systems and apparatus that can enable a machine-learning model (such as a foundation neural network) to be used as a backbone network and adapted for a downstream task efficiently and, once adapted, can allow output of the backbone network to be refined for the downstream task, increasing accuracy of the machine-learning model on the task. A parallel side network is used that processes the output of the backbone network and intermediate outputs from the backbone network to generate a refined output relating to a downstream task. The backbone network may be frozen (i.e. its parameters may be fixed) while the parallel side network is trained. Thus, the parallel side network can be trained to adapt the output of the backbone network to downstream task without backpropagating gradients through the backbone network, resulting in efficient training / finetuning on the downstream task. FIG. 1 shows an overview of an example system / method 100 for generating a network output from a network input using a backbone network 104 (e.g., a foundation model) and a parallel side network 110. Data flow during a forward pass through the networks is shown with solid arrows; data flow during a backward pass (i.e., during backpropagation of gradients) is shown with dashed arrows.

[0042] At inference time, the backbone network 104 processes a neural network input 102, 1, through a plurality of layers 104A-L to generate an output of the backbone network 106, bi_ (also referred to herein as a "final backbone network output"). The output 106 of the backbone network 104 is, in some examples, output from a final layer 104L of the backbone network 104. The backbone network 104 also outputs one or more intermediate outputs 108A, 108B, where an intermediate output corresponds to output of an intermediate layer 104A-104L-1 of the backbone network 104 (i.e., a layer before the final layer 104L). In some examples, every intermediate layer of the backbone network outputs a respective intermediate output, e.g., bi - bi_-i. In other examples, a proper subset of the intermediate layers of the backbone network 104 outputs a respective intermediate output. The final output 106 of the backbone layer may also be referred to herein as an intermediate output, since it is an intermediate output of the overall method 100.

[0043] For example, given a backbone network 204, B, having L layers, the backbone network 104 produces L-l intermediate outputs, bi, b?, . . . bi_-i, and a final output 106, bi, each comprising n tokens with a hidden dimensionality of d, i.e., b, e IRnxd. As an example, the number of tokens, n, may be between 128 and 1024 (e.g., 256), and the number of hidden dimensions may be between 1000 and 2000 (e.g., 1408).

[0044] The backbone network 104 may be a pre-trained neural network, e.g., a foundation model. One or more of the layers of the backbone network 104 may be a transformer layer. For example, the backbone network 104 can be a vision transformer. Examples of such vision transformers include, but are not limited to, the ViT-Base (86 million parameters), ViT-H (632 million parameters), ViT-g (1 billion parameters), ViT-G (1.8 billion parameters) and ViT-e (4 billion parameters) models. A vision transformer may have been pretrained on an image dataset, such as the ImageNet-21K dataset, iNaturalist2018 and / or 2012 datasets, and / or Places365 dataset. For video transformers, the ViViT-e model (636 million parameters) may be used as a backbone. Alternatively or additionally, one or more of the layers of the backbone network 104 may comprise: a convolutional layer; a recurrent layer, such as an LSTM layer; and / or a fully connected layer.

[0045] The final backbone network output 106 and the one or more intermediate outputs 108 are processed by a further network 110 (referred to herein as a "parallel side network") to generate a refined network output (also referred to herein as a "parallel side network output"), from which a final network output 112 for a task is generated. The parallel side network 110 comprises a plurality of layers 110A-L.

[0046] One or more of the layers 110A-L (e.g., each layer, or a subset of the layers) of the parallel side network 110 receives input comprising an intermediate output 108 of the backbone layer 104. The input to a layer 110A-L may comprise a combined input comprising a combination of the received intermediate output 108 of the backbone layer and a further layer input. The further layer input for the initial layer 110A, gi, of the parallel side network 110 may comprise the final output 106 of the backbone network 104. The further layer input for the subsequent layers of the parallel side network 110 may comprise the output of the previous layer of the parallel side network. The combined input is, in some examples, generated by adding the intermediate output 108 of the backbone layer and a further layer input. Alternatively, the combined input may be generated by concatenating the intermediate output 108 of the backbone layer and a further layer input. A final layer 110L, gi_, of the parallel side network 110 receives input comprising the final output 106 of the backbone network 104 and the output of the penultimate layer (not shown) of the parallel side network 110.

[0047] The final output 106 of the backbone network 104, bi_, is used as the initial input to the parallel side network 110 rather than using the input 102 to the backbone network 104 itself (e.g., tokenised image patches). The final output 106 of the backbone network 104 is, in some implementations, combined with an intermediate output 108A of the backbone network 104, e.g., the output 108A of a first layer 104A of the backbone network 104, before input into the parallel side network 110. The use of bi_ as input to the parallel side network 110 allows the parallel side network 110 to refine the outputs of the backbone network 104, given further intermediate activations 108 from the backbone network 104. This approach can improve the accuracy of the output of the parallel side network. In some implementations, every layer of the backbone network 104 has a corresponding layer in the parallel side network 110. Alternatively, the number of layers in the parallel side network 110 may be less than the number of layers in the backbone network 104. For example, instead of operating on activations from each layer of the backbone network 104, the parallel side network 110 may have k < L evenly spaced layers instead.

[0048] One or more of the layers of the parallel side network 110 (e.g., each layer, or a subset of the layers) may implement a low-rank factorisation of a projection of their respective inputs. An example of such a layer is shown in FIG. 2. These one or more of the layers 210 may comprise a down-sampling operation 202 that reduces the size of the layer input, e.g., reduces the size along one dimension of the layer input. The down sampling operation 202 is, in some examples, followed by a non-linear function 204 (e.g., GeLU, ReLU, sigmoid, tanh or the like). The nonlinear function 204 is, in some examples, applied to each element of the downsampled input in a pointwise manner. The one or more layers 210 may further comprise an up-sampling operation 206 that increases the size of the output of the nonlinear function 204. In some implementations, the output of the up-sampling operation 206 is the same dimension as the layer input. In some implementations, the one or more layers 210 may comprise a scaling operation 208 that scales the layer output by a factor, a. In some implementations, the one or more layers 210 may comprise a skip connection 212 that combines 214 the layer input of the output of the previous layer with the (scaled) layer output.

[0049] As an example, the layer 210 may generate an output, y / , from the output of a previous layer y / -7 and an intermediate output of the backbone network, bi, using:

[0050] Yi = S'i(bi+ yi-i) + yi-i (1) where g, is an adaptor function. The adaptor function applies down-sampling operation 202, Wd: IRd-> HF, to its input, x, where d is the dimension of the input to the adapter function 210, r is the dimension of the output of the down-sampling operation 202 and r < d (typically r<<d). In some examples, r may be between 10 and 100, e.g., 48. The down-sampling operation 202 is followed by a GeLU nonlinearity 204. The nonlinearity is followed by an upsampling operation 206, Wu: UF -> IRd. A scaling term 208, a, is applied to the output of the upsampling operation 206. The of the adaptor function, ^(x), can be written symbolically as: gt(x) = aWuGeLU(Wd(x)), (2) for input x and layer index / '. The down-sampling operation 202, Wd, and the upsampling operation 206, Wu, may be referred to as "weight matrices".

[0051] Each of such parallel adaptor layers 210 contains 2rd + 1 parameters due to the low- rank decomposition of the weight matrices. In some examples, the weight matrices each further comprise a set of biases, which adds a further r + d parameters.

[0052] The use of such layers (also referred to as "low-rank mixers") can allow the parallel side network to operate efficiently at the same hidden dimension, d, as the backbone network, and does not require down-projecting backbone activations to the latent space of the transformer, saving compute.

[0053] The adaptor function of equation (2) consists of projections and a point-wise nonlinearity, and operates on the hidden dimension of the backbone features. It does not model interactions along the spatial and, in the case of videos, temporal axes of the input data. To address this, in some implementations, application of the adaptor function is alternated between applying the adaptor function along the channel- and token-dimensions, respectively. These operations can be denoted as: if i mod 2 ^ 0 (31 otherwise where / is the backbone layer index. Thus, even layers correspond to "token mixing" and odd layers to "channel mixing" (or vice versa in some examples). Alternating between "channel mixing" and "token mixing" with the adaptor function of equation (2) can result in an improvement in the accuracy. Moreover, this can also improve efficiency metrics, because for e.g., images, the number of tokens (e.g., 256) is less than the hidden dimension size (e.g., d = 1408) for the backbone network. Therefore, every second mixer block has fewer parameters and GFLOPs.

[0054] In implementations where the input is a video, the backbone features can also have spatiotemporal dimensions, that is bj e IRnt'nsXdwhere nsand nt denote the spatial and temporal-dimensions, respectively. In such examples, the token dimension may be factorised further into separate spatial- and temporal axes, that is bj e [R"txnsxd_ Theadaptor function can be applied separately along each of the spatial and temporal dimensions, and an overall value for the adaptor function for a token layer obtained by combining the result of the spatial and temporal dimensions, e.g., by summing them.

[0055] For example, the adaptor function for each token layer may be given by: where gsPatial(x^ denotes application of the adaptor function along the spatial token dimension and gtempoorai (xj denotes application of the adaptor function along the temporal token dimension. Such processing may also be applied to audio data.

[0056] In some implementations, one or more of the layers of the parallel side network 110 may alternatively or additionally comprise: a convolutional layer; a recurrent layer, such as an LSTM layer; and / or a fully connected layer.

[0057] Referring back to FIG. 1, in some implementations, one or more of the layers of the parallel side network 110 further comprises a transpose operation 114A-L that transposes the combined input of the layer (e.g., b, + prior to input into the adaptor function of the layer.

[0058] In some examples, the final output 112, O, for a task is generated from the output of the parallel side network 110 using a classifier 116, e.g., a neural network. The classifier 116 may, in some examples, be a fully connected neural network. In some alternative examples, the output of the parallel side network 110 is used as the final output 112 for a task directly, i.e., there is no additional classifier 116.

[0059] During training of the parallel side network for a task, the backbone network 104 generates a final backbone network output 106 and one or more intermediate backbone network outputs 108 from a training network input, i.e., an input 102 from a training dataset. The parallel side network 110 generates a candidate output from the final backbone network output 106 and one or more intermediate backbone network outputs 108. The value of an objective function for the task is determined from the candidate output and, in some examples, a ground truth output corresponding to the training network input. In some examples, ground truth outputs are not used in the objective function, e.g., for self-supervised learning, reinforcement learning and / or generative models. Gradients of the objective function are determined with respect to parameters of the parallel side network using backpropagation of gradients through the parallel side network, without backpropagating gradients through the backbone network. These gradients are used to update parameters of the parallel side network using a gradient-based optimisation routine, such as stochastic gradient descent or the like. In other words, the backbone network 104 is frozen and does not have gradients backpropagated through it.

[0060] By keeping the original backbone network 104, B, frozen, and training a parallel side network 110, the storage requirements of the adapted model are small, as the parameters of our side network can be stored for each task. Moreover, this makes deploying numerous parallel side networks 110 adapted from the same backbone 104, B, simple and efficient in terms of storage. Furthermore, the method 100 does not require changing the internal model architecture of the backbone network 104 at all, which makes its practical implementation straightforward.

[0061] Moreover, the initial input to the parallel side network 110, yo, comprises the output 106 of the backbone network 104, bi_. Together with residual connections to the output of the previous layer in the parallel side network 110, this ensures that when the parallel side network is trained, the parallel side network 110 can act in part as an identity function, with the knowledges of the original backbone output 106 being maintained through the parallel side network 110. The rationale for this is that the parallel side network 110 can be viewed as a function that refines the features representations from the backbone network 104, using the original backbone features 108 to do so.

[0062] In some inventions, parameters of the classifier 116 are also updated during the fine- tuning process.

[0063] In some implementations, the model is executed in a distributed manner. For example, the parallel side network is executed on a user / client device, such as mobile user device, and the backbone network is executed at a remote server / in a network. During training, the user device sends input data to the system executing the backbone network. This system inputs the input data to the backbone network, which generates a final backbone network output and one or more intermediate backbone network outputs, and transmits them to the user device. The user device uses the received final backbone network output and one or more intermediate backbone to train the parallel side network locally, for example as described in the preceding paragraph. In some examples, ground truth data for the training is stored locally at the user device. This can allow efficient personalisation of a backbone network for a particular user, while maintaining data privacy and security for the user. In some implementations, a plurality of parallel side networks is trained, each parallel side network specialising in a particular task or sub-task, i.e., an ensemble of expert parallel side networks is trained. At inference time, one or more of the experts may be selected to perform a task based on properties of the input data and / or a user task selection. For example, each parallel side network can be associated with one or more labels indicating what task or sub-task it has been trained to performed, and the one or more models selected based on a match between one or more of the labels and the properties of the input data and / or the user task selection.

[0064] FIGs. 3A-C show a comparison of computation graphs during training of a selfattention layer of a backbone network using a full fine-tuning approach, illustrated in FIG. 3A, a low-rank adaption model (for example, as described in "Low-rank adaptation of large language models" in ICLR, 2022, the contents of which are incorporated herein by reference), illustrated in FIG. 3B, and an example of the methods described herein, illustrated in FIG. 3C.

[0065] FIG. 3A shows an example of a computational graph for a self-attention layer 302A of a backbone network during full fine-tuning, i.e., when all parameters of the layer are adapted during fine tuning. The self-attention layer 302A process a layer input 304A, X / , to generate a layer output 306A, Xi+i. Double headed arrows represent data flows that occur during both a forward pass through the layer and during a backward pass. Solid circles represent intermediate activations whose gradient is required for a backward pass, i.e., that are either cached or computed during a backward pass.

[0066] In such examples, parameters of the weight matrices Q, Wq, Wk and IV^ are updated during finetuning. Such full finetuning requires caching or recomputing gradients with respect to large activation tensors (the solid circles), which is both memory- and compute intensive.

[0067] FIG. 3B shows an example of a computational graph for a self-attention layer 302B of a backbone network using a low rank adaption approach, e.g., as described in "Low- rank adaptation of large language models" in ICLR, 2022. The self-attention layer 302B process a layer input 304B, X / , to generate a layer output 306B, Xi+i. Double headed arrows represent data flows that occur during both a forward pass through the layer and during a backward pass. Single headed arrows represent data flows that occur during a forward pass through the network, but not during a backward pass. Solid circles represent intermediate activations whose gradient is required for a backward pass, i.e., that are either cached or computed during a backward pass. In such examples, only a small number of parameters are updated during finetuning per self-attention block - namely the weight matrices Aq308A and Bq 308B. The weight matrices WQ, Wq, Wk and Wvremain fixed. While this improves parameter efficiency when compared to full finetuning, it still requires backpropagating gradients through the entire backbone network. Thus the computational graph is quite similar to full finetuning.

[0068] FIG. 3C shows an example of a computational graph for a self-attention layer 302C of a backbone network using a parallel side network approach. The self-attention layer 302C process a layer input 304C, X / , to generate a layer output 306C, Xi+i. The parallel side network 310 processes the layer input 304C, X / , and a parallel layer input 312, y / , to generate a parallel layer output 314, y / +i.

[0069] Double headed arrows represent data flows that occur during both a forward pass through the layer and during a backward pass. Single headed arrows represent data flows that occur during a forward pass through the network, but not during a backward pass. Solid circles represent intermediate activations whose gradient is required for a backward pass, i.e., that are either cached or computed during a backward pass. Dashed circles represent intermediate activations that can be ignored during a backward pass.

[0070] In such examples, the self-attention layer 302C of the backbone network is completely frozen during fine tuning, and backpropagation and parameter updates only occur in the parallel side network 310. This results in significant reductions in training time and memory.

[0071] FIG. 4 shows a flow diagram of an example method 400 for adapting the output of a backbone network to a task. For convenience, the operations of the method 400 are described with reference to a system that performs the operations. This system of the method 400 includes one or more processors, memory, and / or other component(s) of computing device(s).

[0072] At operation 402, the system generates, using a backbone network comprising a plurality of layers, a plurality of backbone network outputs from a network input. The plurality of backbone network outputs comprises a final backbone output comprising the output of a final layer of the backbone network. The plurality of backbone network outputs further comprises one or more intermediate backbone network outputs. An intermediate backbone network output comprises an output of a layer of the backbone network prior to the final layer of the backbone network. For example, for an L-layer backbone network, an intermediate backbone network output comprises the output of one of layers 1 to L-l of the backbone network.

[0073] At operation 404, the system generates, using a parallel side network comprising a plurality of layers, a parallel side network output from the final backbone network output. Generating the parallel side network output from the final backbone network output comprises, for each of a plurality of layers of the parallel side network: receiving 404A a respective backbone network output (e.g., a final backbone output and / or an intermediate backbone output); and processing 404B a respective layer input and the respective backbone network output using a respective layer of the parallel side network to generate a respective parallel side network layer output.

[0074] In some implementations, the respective layer input and the respective backbone network output are combined to generate a combined input prior to processing by their respective layer of the parallel side network. The combined input comprises a combination of the respective layer input, e.g., yi-i, and the respective backbone network output, bi. For example, the combined input may be a sum of the respective layer input and the respective backbone network output for the layer, e.g., y, + b,. Alternatively, in some implementations the combine input is a concatenation of the respective layer input and the respective backbone network output for the layer. A transpose operation may be applied to the combined input before processing by its respective backbone network output for the layer.

[0075] In some implementations, the input to the first layer of the parallel side network comprises an intermediate network output corresponding to a first layer of the backbone network and the final output of the backbone network. The input to the final layer of the parallel side network comprises the output of a penultimate layer of the parallel side network and the final output of the backbone network. Input to intermediate layers of the parallel side network comprise a respective intermediate network output of a corresponding intermediate layer of the backbone network and the output of a previous layer of the parallel side network, e.g., the immediately prior layer.

[0076] For one or more of the layers of the parallel side network, processing the respective layer input and the respective backbone network output comprises applying an adaptor function comprising a low rank factorisation of a projection of a combined input to the layer. Applying an adaptor function comprises, for example, generating a low-dimensional representation of the combined input from the combined input for the respective layer using a down-projection operation, Wd. Following the down-projection operation, a nonlinear function is applied to low-dimensional representation of the combined input, such as a GeLU function, a tanh function, a sigmoid function, a ReLU function, or the like. The non-linear function may be applied in a pointwise manner. An adaptor function output is generated from the output of the nonlinear function using an up-projection operation, Wu. In some implementations, a scaling function, a, is applied to the output of the up-projection operation.

[0077] In some examples, the parallel side network alternates between applying adaptor functions in a channel dimension of respective combined inputs and applying adaptor functions in a token dimension of respective combined inputs. For example, odd layers of the parallel side network apply the adaptor function in the channel dimension, and even layers of the parallel side network apply the adaptor function in the token dimension (or vice versa). For input data with both a spatial and temporal dimension, such as video or audio data, applying the adaptor function in the token dimension may comprise applying the adaptor function in the temporal dimension and applying the adaptor function in the spatial dimension, and then combining the results, e.g., summing the result of applying the adaptor function in the temporal dimension and applying the adaptor function in the spatial dimension.

[0078] In some implementations, each layer of the parallel side network comprises a residual / skip connection between at least one of its inputs (i.e., the respective layer input, the respective backbone network output, or the combined input) and the output of the layer. For example, the residual connection acts to combine the combined input for a layer with the output of that layer, e.g., the adaptor function output.

[0079] At operation 406, the system generates a final network output based on the parallel side network output. In some implementation, the final network is the parallel side network output. In other implementations, the parallel side network output undergoes further processing to generate the final network output, for example, using a classifier network, e.g., a fully connected neural network.

[0080] FIG. 5 shows a flow diagram of an example method 500 training a parallel side network for adapting the output of a backbone network to a task. For convenience, the operations of the method 400 are described with reference to a system that performs the operations. This system of the method 400 includes one or more processors, memory, and / or other component(s) of computing device(s).

[0081] At operation 502, the system generates a candidate network output from a training input using a backbone network and a parallel side network.

[0082] The system generates 502A a plurality of backbone network outputs from the training input using the backbone network, for example as described in relation to operation 402 of FIG. 4. The plurality of backbone network outputs comprising a final backbone network output (e.g., the output of a final layer of the backbone network) and one or more intermediate backbone network outputs (e.g., respective outputs of one or more intermediate layers of the backbone network).

[0083] The system generates 502B the candidate network output from the plurality of backbone network outputs using the parallel side network, for example as described in relation to operations 404 and 406 of FIG. 4.

[0084] At operation 504, the system updates parameters of the parallel side network based on an objective function depending on the candidate network output. In examples with labelled training data, the objective function may further depend on a ground truth output corresponding to the training input.

[0085] The system determines 504A gradients of the objective function with respect to the parameters of the parallel side network. Determining gradients of the objective function comprises backpropagating gradients of the objective function through the parallel side network without backpropagating gradients through the backbone network. The objective function used depends on the task the network is being trained for. For example, for a classification task, the objective function may comprise a classification loss, such as a cross entropy loss or a Kullback-Leibler divergence loss. For an image / video / audio reconstruction / enhancement task, the objective function may comprise an LI or L2 loss. For a generative task, the objective function may be a GAN loss, which may be based on the output of a discriminator network applied to the candidate network output. Many other examples are possible.

[0086] The system determines 504B updates to the parameters of the parallel side network based on the gradients of the objective function with respect to the parameters of the parallel side network. The parameter updates may be determined from the gradients using a gradient based optimisation procedure, e.g., by applying stochastic gradient descent to the objective function.

[0087] Example tasks that a neural network may be trained to perform will now be described. It will be appreciated that the neural networks described above may be trained to perform any appropriate task and the below examples are illustrative and not intended to be limiting. The neural network may be configured to receive and process any type of digital input and to provide any type of digital output including a score, classification, or regression output. The input to and / or the output of the neural network may be sequential in nature as appropriate for the task.

[0088] Implementations of the neural network find particular use in the context of consumer applications, such as video sharing platforms or for social media applications.

[0089] In some examples, a neural network may be configured to perform an image processing task. The neural network may be configured to receive and process input data comprising image data, e.g., an input image. The image data may comprise one or more values for each of a plurality of pixels, such as intensity values. Alternatively, the image data may comprise a low-level feature encoding of an image. The image data may comprise colour or monochrome data. The image data may be captured by an image sensor of a digital camera, LIDAR, infra-red camera, or any other camera type.

[0090] The image processing task may be an image classification task or an object detection task. The neural network may be configured to process input image data to provide an output indicating the presence of one or more object categories in the input image data. The indication may, for example, be a probability, a score, or a binary indicator for a particular object category.

[0091] The image processing task may be an object detection task. The neural network may be configured to process input image data to provide an output indicating a location of one or more objects that have been detected in the input image data. The indication may be a bounding box, set of co-ordinates or other location indicator and the output may further comprise a label indicating the corresponding detected object.

[0092] The image processing task may be image segmentation. The neural network may be configured to process input image data to provide an output indicating an object category that a pixel (or each pixel) in the input image data belongs to. The image processing task may be an image encoding task. The neural network may be configured to process input image data to provide a representation of the input image data. The representation may be a compressed encoding of the image data.

[0093] The image processing task may be a depth estimation task. The neural network may be configured to process input image data to provide an output indicating an estimated depth of objects depicted in the image data. The output may be a depth map comprising an estimated depth value for each pixel of the input image data.

[0094] The image processing task may be a captioning task. The neural network may be configured to process input image data to provide an output comprising a sequence of text in natural language describing the objects depicted within the image.

[0095] The image processing task may be a de-noising or in-filling task. The neural network may be configured to process input image data to provide as output a version of the input image data that has reduced noise or artefacts or has missing data imputed.

[0096] In another example, the neural network may be configured to receive and process audio data as an input. The audio data may comprise any type of digital audio signal and may comprise raw digital samples of a waveform (e.g. amplitude values) or may be an encoding derived from an audio signal such as a spectrogram, mel-frequency cepstral coefficients or other acoustic features / time-frequency domain representation. The audio data may be obtained from an audio transducer such as a microphone.

[0097] The audio data may comprise a speech signal. The audio processing task may be a speech processing task such as speech recognition. The neural network may be configured to process an input speech signal to provide output data comprising one or more probabilities or scores indicating that one or more words or sub-word units comprise a correct transcription of the speech contained within. Alternatively, the output data may comprise a transcription itself.

[0098] The audio processing task may be a keyword ("hotword") spotting task. The neural network may be configured to process input audio data to provide an indication of whether a particular word or phrase is spoken in the input audio data.

[0099] The audio processing task may be a speaker recognition or verification task. The neural network may be configured to process input audio data to provide an indication of the identity of a speaker or to provide an indication of whether a particular speaker is present in the audio data.

[0100] The audio processing task may be a language recognition task. The neural network may be configured to process input audio data to provide an indication or delineation of one or more languages present in the input audio data.

[0101] The audio processing task may be a control task. The neural network may be configured to process input audio data comprising a spoken command for controlling a device to generate output data that causes the device to carry out actions corresponding to the spoken command.

[0102] In a further example, the neural network may be configured to receive and process text data. The neural network may be configured to perform a text-to-speech task. The neural network may be configured to process input text data to generate audio data comprising a spoken utterance corresponding to the input text data. This output audio data may be in a similar format to the input audio data for audio processing tasks described above.

[0103] The neural network may be configured to perform an image generation task. The neural network may be configured to process input text data to generate image data comprising objects corresponding to the input text data. The output image data may be in a similar format to the input image data for image processing tasks described above.

[0104] The neural network may be configured to perform a neural machine translation task. The neural network may be configured to process input text data comprising text in one language to provide output data comprising one or more probabilities or scores indicating that one or more words or sub-word units in a second language is comprised in a proper translation of the input text data into the second language. Alternatively, the output data may comprise the translation itself.

[0105] The neural network may be configured to process input text data comprising instructions for controlling a device to generate output data that causes the device to carry out actions corresponding to the instructions.

[0106] The neural network may be part of a dialogue system. The neural network may be configured to perform a conditional text generation task. The neural network may be configured to process input text data comprising a user prompt to provide output text data that is a response to the user prompt. The user prompt may be a question and the response may be an answer to the question. The user prompt may be a text description of the function of computer code and the response may be computer code in a programming language that is configured to perform the function. The user prompt may be an initial sequence of computer code in a programming language and the response may be a further sequence of computer code that completes the initial sequence to perform a particular function.

[0107] The neural network may be configured to perform a natural language processing or understanding task. For example, the task may be an entailment task, a paraphrase task, a textual similarity task, a sentiment task, a sentence completion task, a grammar task, or other similar task.

[0108] In another example, the neural network may be configured to receive and process video data. The video data may comprise a plurality of frames, e.g. image data, and audio data. The above-described image and audio processing tasks may therefore also be applicable to video data.

[0109] In addition, the neural network may be configured to perform an action recognition or detection task. The neural network may be configured to process input video data to provide an output indicating the detection of one or more actions being performed in the video data and / or the spatial and temporal locations of the detected actions within the video data. The indications may be a probability or score indicating a particular action has been detected, the spatial location may be a bounding box or co-ordinates in the frames of the video, the temporal location may be a set of timestamps or frame numbers of the video.

[0110] In another example, the neural network may be configured to receive and process digital documents. The digital document may be an Internet resource, or one or more portions or features extracted from an Internet resource. The neural network may be configured to classify the document. The neural network may be configured to process the digital document to provide an output indicating a classification of the digital document. The indication may be a probability, a score, or a binary indicator for a particular category. The classification may be a topic of the document.

[0111] In a further example, the neural network may be configured to receive and process an impression context for a particular advertisement. The neural network may generate an output comprising an estimated probability or score that the particular advertisement will be clicked on.

[0112] The neural network may be configured to receive and process features of a personalized recommendation for a user. For example, features characterizing the context for the recommendation such as features characterizing previous actions taken by the user. The neural network may generate an output comprising a probability or score for a particular content item representing the estimated likelihood that the user will respond favourably to being recommended the particular content item.

[0113] In another example, the neural network may be configured to perform a health prediction task. The neural network may be configured to receive and process electronic health record data. The neural network may be configured to generate an output comprising an indication of a diagnosis of a particular disease or health condition, or a prediction of the occurrence of an adverse health event, or a predicted treatment for the patient. The indication may be a probability or score.

[0114] In a further example, the neural network may be part of a data compression system. The neural network may be configured to process input data to provide a compressed version of the input data as output. The compressed data may be stored at an appropriate storage or memory device or transmitted to another device.

[0115] In another example, the neural network may be part of a reinforcement learning system. In general, in a reinforcement learning system an agent interacts with an environment in order to carry out a particular task. Observations characterizing the state of an environment may be received. The reinforcement learning system may select an action for an agent to perform based upon the received observations. The action performed by the agent may cause the environment to transition to a new state and the environment may provide a reward signal in response to agent's action. Further observations of the environment may be received, and a new action may be selected based upon the further observations and the received reward if applicable. Actions for the agent to carry out may be selected based upon a history of observations of the environment and actions performed by the agent.

[0116] The reinforcement learning system may be configured to select an action for an agent to perform based upon the output of the neural network. For example, the neural network may be a policy network and may be configured to process the observation characterizing the state of the environment to provide a probability distribution (or set of scores) over a set of possible actions. An action may be selected by sampling an action from the probability distribution or the action with the highest probability / score may be selected. The reinforcement learning system may cause the agent to perform the selected action. Neural networks in reinforcement learning systems may perform other functions such as provide an estimate of the value of a being in a given state of the environment, provide an estimate of the value of carrying out a particular action, provide a model of the environment to predict state transitions and observations, or provide an estimate of the reward function of an environment amongst others. It will be appreciated that in all these cases, whilst a neural network may not directly select an agent action, the output of the neural network feeds into the process for the selection of an action.

[0117] The environment may be a real-world environment. The agent may be a mechanical or electronic agent interacting with the real-world environment to carry out a particular task.

[0118] FIG. 6 illustrates a schematic example of a computing system / apparatus 600 for performing any of the methods, operations or processes described and / or for implementing any of the systems, units and / or apparatus as described. The computing system / apparatus 600 shown is an example of a computing device or platform. It will be appreciated by the skilled person that other types of computing devices / systems / platforms can alternatively be used to implement the methods described, such as a distributed computing system. The computing system / apparatus 600 is, in some examples, a UE, a subsystem of a UE, a network node / entity and / or a subsystem of a network node / entity.

[0119] The apparatus (or system) 600 includes one or more data processing apparatus, for example one or more processors 602. The one or more processors 602 control operation of other components of the system / apparatus 600. The system / apparatus 600 can be part of a computing device, computing system, distributed computing system, cloud computing platform and the like for implementing the functionality of the systems / apparatus and / or one or more methods / operations / processes as described. The one or more processors 602 , for example, include a general-purpose processor. The one or more processors 602 can be a single core device or a multiple core device. The one or more processors 602 can include a Central Processing Unit (CPU) or a graphical processing unit (GPU). Alternatively, or in addition, the one or more processors 602 may include one or more neural network accelerators, or other specialized processing hardware, for instance a RISC processor or programmable hardware with embedded firmware. In some examples, multiple processors are included. In some embodiments, the one or more processors 602 are part of a distributed computing system such as a cloud computing system and / or cloud computing platform.

[0120] The system / apparatus includes memory system or memory 604 including a working or volatile memory 606. The one or more processors access the volatile memory 606 in order to process data and control the storage of data 607 in memory. The volatile memory 606 can include RAM of any type, for example, Static RAM (SRAM), Dynamic RAM (DRAM), or include Flash memory, such as an SD-Card. In some embodiments, the memory 604 and / or one or more volatile memories 606 include a plurality of memories 604 forming part of the distributed computing system such as the cloud computing system and / or cloud computing platform and the like.

[0121] The system / apparatus includes a non-volatile memory 608. The non-volatile memory 608 stores a set of operation or operating system instructions 609a for controlling the operation of the processors 602 in the form of computer readable instructions and / or software instructions 609b in the form of computer readable instructions, which when executed on the one or more processors 602 cause the processors to implement the methods, processes, operations and / or functionality described. The non-volatile memory 608 can be a memory of any kind such as a Read Only Memory (ROM), a Flash memory, SD drive, a magnetic drive memory or magnetic disc drive memory and the like as the application demands. In some embodiments, the non-volatile memory 608 includes a plurality of non-volatile memories 608 forming part of the distributed computing system such as the cloud computing system and / or cloud computing platform and the like.

[0122] The one or more processors 602 are configured to execute operating instructions 609a and / or software instructions 609b to cause the system / apparatus to perform any of the methods or processes described. The operating instructions 609a include, for example, code (i.e., drivers) relating to the hardware components of the system / apparatus 600, as well as code relating to the basic operation of the system / apparatus 600. Generally speaking, the one or more processors 602 execute one or more instructions of the operating instructions 609a and / or software instructions 609b, which are stored permanently or semi-permanently in the nonvolatile memory 608, using the volatile memory 606 to store temporarily data generated during execution of said operating instructions 609a and / or software instructions 609b.

[0123] In some implementations, the one or more processors 602 are connected to a network interface 608 including a transmitter (TX) and a receiver (RX) for communicating over a network with other apparatus and systems. The one or more processors 602 are, in some examples, connected with a user interface (UI) 610 for user or operator input for instructing or using the computing system and / or for outputting data therefrom. The one or more processors 602 are, in some examples, connected with a display 612 for displaying output to a user or operator. The at least one processor 602, with the at least one memory 604 and the computer program code 609a, 609b are arranged to cause the computing system 600 to at least perform at least the operations, methods, and / or processes, for example as disclosed in relation to the schematic diagrams, flow diagrams or operations as described with any of figures 1 to 5 and related features thereof.

[0124] FIG. 7 shows an example non-transitory computer readable medium 700 according to some implementations. The non-transitory medium 700 includes a computer readable storage medium 702 and / or input / output mechanism 704 for enabling a computing system 600 to access said computer-readable medium 702. Although in this example the non-transitory medium is USB stick, this is by way of example only and the invention is not so limited, the skilled person would appreciate the non-transitory media 700 could be any other type of computer readable media or medium such as, for example, a CD, a DVD, a USB stick, a blue ray disk, flash drive, computer memory such as a hard drive or flash memory etc. and / or any other computer readable medium as the application demands. The non-transitory medium 700 stores computer program code, causing an apparatus to perform one or more of the methods, operations, processors of any preceding process for example as disclosed in relation to the flow diagrams and schematic diagrams of figures 1 to 5 and related features thereof.

[0125] Any system feature as described herein may also be provided as a method feature, and vice versa. As used herein, means plus function features may be expressed alternatively in terms of their corresponding structure. In particular, method aspects may be applied to system aspects, and vice versa.

[0126] Furthermore, any, some and / or all features in one aspect can be applied to any, some and / or all features in any other aspect, in any appropriate combination. It should also be appreciated that particular combinations of the various features described and defined in any aspects of the invention can be implemented and / or supplied and / or used independently. Although several embodiments have been shown and described, it would be appreciated by those skilled in the art that changes may be made in these embodiments without departing from the principles of this disclosure, the scope of which is defined in the claims.

Claims

ClaimsWhat is claimed is:

1. A method implemented by one or more data processing apparatus, the method comprising: generating, by a backbone network comprising a plurality of layers, a plurality of backbone network outputs from a network input, the plurality of backbone network outputs comprising: a final backbone output comprising the output of a final layer of the backbone network; and one or more intermediate backbone network outputs, wherein an intermediate backbone network output comprises an output of a layer of the backbone network prior to the final layer of the backbone network; generating, by a parallel side network comprising a plurality of layers, a parallel side network output from the final backbone network output, comprising, for each of a plurality of layers of the parallel side network: receiving a respective backbone network output; and processing a respective layer input and the respective backbone network output using a respective layer of the parallel side network to generate a respective parallel side network layer output; and generating a final network output based on the parallel side network output.

2. The method of claim 1, wherein processing the respective layer input and the respective backbone network output using a respective layer of the parallel side network to generate a respective parallel side network layer output comprises applying an adaptor function comprising a low rank factorisation of a projection of a combined input comprising a combination of the respective layer input and the respective backbone network output.

3. The method of claim 2, wherein applying the adaptor function comprises: generating a low-dimensional representation of the combined input from the combined input for the respective layer using a down-projection operation; applying a nonlinear function to low-dimensional representation of the combined input; and generating an adaptor function output from the output of the nonlinear function using an up-projection operation.

4. The method of any of claims 2 or 3, wherein the down-projection operation comprises applying a down-projection matrix of dimension d x r to the combined input, wherein d is a dimension of the combined input, r is a dimension of the lowdimensional representation of the combined input and d > r.

5. The method of claim 4, wherein the up-projection operation comprises applying an up-projection matrix of dimension r x d to the adaptor function output.

6. The method of any of claims 2 to 5, wherein processing the combined input using a respective layer of the parallel side network to generate a respective parallel side network layer output further comprises combining the combined input or the respective layer input for the respective layer with the adaptor function output.

7. The method of any of claims 2 to 6, wherein generating the parallel side network output from the final backbone network output comprises alternating between applying adaptor functions in a channel dimension of respective combined inputs and applying adaptor functions in a token dimension of respective combined inputs.

8. The method any preceding claim, wherein : the respective layer input for an initial layer of the parallel side network comprises the final backbone network output; and the respective layer inputs for subsequent layers of the parallel side network comprise a respective parallel side network layer output from a preceding layer of the parallel side network.

9. The method of any preceding claim, wherein the respective backbone network output for a final layer of the parallel side network is the final backbone network output.

10. The method of any preceding claim, wherein generating the final network output based on the parallel side network output comprises generating the final network output from the parallel side network output using a classifier network.

11. The method of any preceding claim, wherein the backbone network comprises a transformer network.

12. The method of any preceding claim, wherein : the neural network input comprises one or more images; andthe final network output comprises: one or more image classification outputs; one or more object detection outputs; one or more image segmentations; one or more denoised images; one or more infilled images; one or more enhanced images; one or more depth estimations; and / or one or more image captions.

13. The method of any preceding claim, further comprising: updating parameters of the parallel side network based on an objective function depending on the final neural network output, the updating comprising: determining gradients of the objective function with respect to the parameters of the parallel side network, comprising backpropagating gradients of the objective function through the parallel side network without backpropagating gradients through the backbone network; and determining updates to the parameters of the parallel side network based on the gradients of the objective function with respect to the parameters of the parallel side network.

14. A method implemented by one or more data processing apparatus, the method comprising: generating a candidate network output from a training input using a backbone network and a parallel side network, comprising : generating, using the backbone network, a plurality of backbone network outputs from the training input, the plurality of backbone network outputs comprising a final backbone network output and one or more intermediate backbone network outputs; and generating, using the parallel side network, the candidate network output from the plurality of backbone network outputs; and updating parameters of the parallel side network based on an objective function depending on the candidate network output , the updating comprising: determining gradients of the objective function with respect to the parameters of the parallel side network, comprising backpropagating gradients of the objective function through the parallel side network without backpropagating gradients through the backbone network; and determining updates to the parameters of the parallel side network based on the gradients of the objective function with respect to the parameters of the parallel side network.

15. The method of claim 15, wherein generating, using the parallel side network, the candidate network output from the plurality of backbone network outputs comprises, for each of a plurality of layers of the parallel side network: receiving a respective backbone network output; combining a respective layer input and the respective backbone network output to generate a combined input; and processing the combined input using a respective layer of the parallel side network to generate a respective parallel side network layer output.

16. The method of claim 15, wherein processing the combined input using a respective layer of the parallel side network to generate a respective parallel side network layer output comprises applying an adaptor function comprising a low rank factorisation of a projection of the combined input.

17. The method of claim 16, wherein applying the adaptor function comprises: generating a low-dimensional representation of the combined input from the combined input for the respective layer using a down-projection operation; applying a nonlinear function to low-dimensional representation of the combined input; and generating an adaptor function output from the output of the nonlinear function using an up-projection operation.

18. The method of any of claims 16 or 17, wherein the down-projection operation comprises applying a down-projection matrix of dimension d x r to the combined input, wherein d is a dimension of the combined input, r is a dimension of the lowdimensional representation of the combined input and d < r.

19. The method of claim 18, wherein the up-projection operation comprises applying an up-projection matrix of dimension r x d to the adaptor function output.

20. The method of any of claims 16 to 19, wherein processing the combined input using a respective layer of the parallel side network to generate a respective parallel side network layer output further comprises combining the combined input for the respective layer with the adaptor function output.

21. The method of any of claims 16 to 20, wherein generating the parallel side network output from the final backbone network output comprises alternating betweenapplying adaptor functions in a channel dimension of respective combined inputs and applying adaptor functions in a token dimension of respective combined inputs.

22. The method any of claims 15 to 21, wherein : the respective layer input for an initial layer of the parallel side network comprises the final backbone network output; and the respective layer inputs for subsequent layers of the parallel side network comprise a respective parallel side network layer output from a preceding layer of the parallel side network.

23. The method of any of claims 15 to 22, wherein the respective backbone network output for a final layer of the parallel side network is the final backbone network output.

24. The method of any of claims 14 to 23, wherein generating the candidate network output further comprises generating the candidate network output from output of the parallel side network using a classifier network.

25. The method of any of claims 14 to 24, wherein the backbone network comprises a transformer network.

26. The method of any of claims 14 to 25, wherein: the network input comprises one or more images; and the final network output comprises: one or more image classification outputs; one or more object detection outputs; one or more image segmentations; one or more denoised images; one or more infilled images; one or more enhanced images; one or more depth estimations; and / or one or more image captions.

27. The method of any of claims 14 to 26, wherein the objective function further depends on a ground truth network output corresponding to the training input.

28. A system comprising one or more processors and a memory, the memory storing computer readable instructions that, when executed by the one or more processors, causes the system to perform a method according to any preceding claim.

29. A computer program product comprising computer readable instructions that, when executed by data processing apparatus, causes the data processing apparatus to perform a method according to any of claims 1 to 27.