Continuous learning neural network system training for classification type tasks

CN116997908BActive Publication Date: 2026-09-29GDM HOLDINGS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280020905.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-05-27
Filing Date
2022-05-27
Publication Date
2026-09-29
Estimated Expiration
2042-05-27

AI Technical Summary

Benefits of technology

[0029]本说明书描述了一种用于训练基于神经网络的系统的方法,该方法在连续学习设置中特别有利。在一些现有技术的连续学习方法中,连续学习新任务或数据分布不稳定的情况要求任务结构和任务边界的知识。在本训练方法中,不要求任务和任务边界的知识。使用本方法训练的系统可以在连续学习基准上胜过当前现有技术系统很大的裕量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116997908B_ABST
    Figure CN116997908B_ABST
Patent Text Reader

Abstract

A computer-implemented method for training a neural network-based system is disclosed. The method includes receiving a training data item and target data associated with the training data item. The training data item is processed using an encoder to generate an encoding of the training data item. A subset of neural networks is selected from a plurality of neural networks stored in a memory based on the encoding; wherein the plurality of neural networks are configured to process the encoding to generate output data indicative of a classification of an aspect of the training data item. The encoding is processed using the selected subset of neural networks to generate output data. An update to parameters of the selected subset of neural networks is determined based on a loss function, the loss function comprising a relationship between the generated output data and the target data associated with the training data item. The parameters of the selected subset of neural networks are updated based on the determined update.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This manual relates to training a neural network system in a continuous learning setup. The neural network system can be trained to perform classification tasks. Background Technology

[0002] A neural network is a machine learning model that uses one or more non-linear units to predict the output of a received input. Some neural networks also include one or more hidden layers in addition to the output layer. The output of each hidden layer is used as the input to the next layer in the network (i.e., the next hidden layer or output layer). Each layer of the network generates an output from the received input based on the current values ​​of its corresponding parameter set. Summary of the Invention

[0003] This specification describes a system implemented as a computer program on one or more computers at one or more locations for training a neural network-based system. Typically, neural network-based systems are trained to perform classification-type tasks. The system can be a continuously learning system, as it can continuously learn to perform new tasks or adapt to changes in data distribution as new tasks arise. However, training does not require detecting or knowing the task boundaries as some prior art continuous learning methods do.

[0004] According to one aspect, a method for training a computer-based system based on neural networks is provided. The method includes receiving training data items and target data associated with the training data items. The training data items are processed using an encoder to generate an encoding of the training data items. A subset of neural networks is selected from a plurality of neural networks stored in memory based on the encoding. The plurality of neural networks are configured to process the encoding to generate output data indicating a classification of aspects of the training data items. The selected subset of neural networks is used to process the encoding to generate output data. An update to the parameters of the selected subset of neural networks is determined based on a loss function, said loss function including a relationship between the generated output data and the target data associated with the training data items. The parameters of the selected subset of neural networks are updated based on the determined update.

[0005] The training method may also include repeating the above steps for multiple training data items. The multiple training data items may include a first training data item extracted from a first data distribution and a second training data item extracted from a second data distribution, wherein the first and second data distributions are different. For example, the training method may use training data items associated with a first task or sub-task and training data items associated with a second task or sub-task. Therefore, a neural network-based system can be trained to perform either a task or sub-task or to adapt to changes in the data distribution.

[0006] Multiple training data items can include training data items extracted from a second data distribution and training data items extracted from a first data distribution. That is, there may not be a clear boundary between changes in data distribution or between changes in task or subtask. Changes in data distribution or task / subtask can be gradual. This change can be achieved by extracting training data items associated with different data distributions / tasks / subtasks based on a specific probability distribution conditioned on training time. For example, training data items can be extracted from a first data distribution with peak probability at time t1, or from a second data distribution with peak probability at time t2, etc., where the temporal overlap in the probability distributions is used to select training data items from either the first or second data distribution. The probability distribution can be a Gaussian distribution. Training methods can also be used in incremental learning settings, where training data items are presented sequentially, one complete task / subtask / class at a time.

[0007] As described above, neural network-based systems can be trained to perform classification-type tasks. For example, when the training data item is image data, a neural network-based system can be trained for object classification, i.e., predicting or determining the presence of objects in image data. In another example, the task could be object detection, i.e., determining whether an aspect of image data (such as a pixel or region) is part of an object. Another image-based task could be object pose estimation. Training data items can be video data items. Possible video tasks include action recognition, i.e., determining what action is being performed in a video or a segment of video (aspect), and action detection to determine whether an action is being performed in a video segment. Training data items can be audio data items. Possible audio tasks include speech recognition and speaker recognition, etc.

[0008] The relationship between the generated output data and the target data of the training data items can be based on the dot product between the generated output data and the target data. The loss function can be the negative of the dot product between the generated output data and the target data (depending on whether it is minimized or maximized for training). Compared with conventional methods, the loss function may not require a normalization term and can use the raw values ​​from the dot product.

[0009] The target data can be in the form of one-hot vectors. That is, a vector where the element corresponding to the correct class / selection is set to 1 and all other elements are set to 0. In the case where the loss function is based on the dot product between the one-hot vectors and the output data, this means that during training, only the parameters of the neural network on the paths to neurons representing the correct class are modified. The parameters of the paths associated with incorrect classes are not modified. These paths used for incorrect classes may be important for previously learned data / tasks / subtasks. This helps mitigate "catastrophic forgetting," thereby degrading performance on earlier learned data / tasks / subtasks to benefit newer data / tasks / subtasks. In the context of catastrophic forgetting, it has been observed that for a single neural network, helpful updates generally outweigh unhelpful updates. By providing multiple neural networks, this general statistical trend can be amplified and further mitigated.

[0010] The encoder can be pre-trained using a dataset different from the dataset to which the training data items belong. In one example, the encoder is pre-trained on the Omniglot dataset, while the training data items could belong to the MNIST dataset. In another example, the encoder can be pre-trained on the ImageNet dataset, and the training data items could belong to one of the CIFAR family datasets.

[0011] Self-supervised learning techniques can be used to pre-train the encoder. These techniques can include training based on transformed views of the training data items. For example, training could be based on instance discrimination and discrimination between positive and negative pairs of transformed versions of data items. Further details can be found in Mitrovic et al., “Representation learning via invariant causal mechanisms”, arXiv:2010.07922 (https: / / arxiv.org / abs / 2010.07922), the entire contents of which are incorporated herein by reference, and in Chen et al., “A simple framework for contrastive learning of visual representations”, arXiv:2002.05709 (https: / / arxiv.org / abs / 2002.05709), the entire contents of which are incorporated herein by reference. In Grill et al.'s "Bootstrap your own latent: A new approach to self-supervised learning," arXiv:2006.07733, only positive examples are used, which is available at https: / / arxiv.org / abs / 2006.07733, the full content of which is incorporated herein by reference.

[0012] The encoder can be based on a variational autoencoder; for example, it can include the encoder portion of a variational autoencoder. In another example, the encoder can be based on a ResNet architecture. In some other instances, the encoder can be based on BYOL (arXiv:2006.07733), SimCLR (arXiv:2002.05709), or ReLIC (arXiv:2010.07922).

[0013] The encoder parameters can remain fixed. That is, the encoder parameters do not need to be changed during training, and only the parameters of multiple neural networks are updated. This provides a stable representation from which a subset of neural networks can be selected based on the encoding of the training data items, and allows for the specialization of neural networks according to the encoding of the training data items, which can provide useful knowledge for the task associated with the training data items.

[0014] Each of the multiple neural networks can be associated with a corresponding key. Selecting a subset of neural networks can also be based on the corresponding key. The method may further include determining the similarity between the encoding and each corresponding key; and selecting a subset of neural networks can be based on the determined similarity. The similarity can be based on the cosine distance between the encoding and the corresponding key. Therefore, the k nearest neighbors of the encoded key can be selected.

[0015] The corresponding keys can be generated by sampling the probability distribution based on the embedding space represented by the encoder. That is, the key can be a vector sampled from the embedding space represented by the encoder. In one example, the embedding space has a dimension of 512. In another example, the embedding space has a dimension of 2048. It should be understood that the embedding space can have dimensions that are considered appropriate by those skilled in the art.

[0016] The probability distribution can be determined based on samples of the encoded data. That is, samples of encoded data can be generated by processing sample data items with an encoder to produce corresponding codes for those data items. In one example, the number of sample data items is 256, although it should be understood that other numbers can be used. Sample data items can be extracted from a dataset different from the dataset to which the training data items belong. The dataset for sample data items can be the same dataset used to pre-train the encoder, or a dataset different from both the pre-training and training data item datasets. The probability distribution can be generated based on statistics of the data encoded by the samples. For example, the sample mean and variance can be determined against a Gaussian distribution. This can provide a more uniform key distribution across the encoder embedding space.

[0017] Processing the encoding of training data items using a selected subset of neural networks may include: processing the encoding through each corresponding neural network in the subset to generate intermediate data for each corresponding neural network, and aggregating the intermediate data for each corresponding neural network to generate output data indicating the classification of aspects of the training data items. The intermediate data may be in the same form as the output data, and the intermediate data may be equivalent to the output data of a corresponding neural network.

[0018] Aggregating intermediate data can include weighting the intermediate data for each corresponding neural network in a subset of neural networks. When the neural network is a classifier, aggregating intermediate data can include weighting the intermediate data for each corresponding classifier in a subset of classifiers. For example, aggregation can be a weighted sum. It should be understood that other forms of aggregation are also possible. Weighted sums can be normalized by summing the weights.

[0019] Weighting can be based on the similarity between the key associated with the neural network / classifier and the encoding of the training data items. In other words, weighting can be based on the similarity between the key and the encoding of the training data items used to select a subset of the neural network.

[0020] Multiple neural networks can serve as neural network classifiers. A neural network classifier can be a single-layer classifier. That is, a neural network classifier can have only an input layer and an output layer, without any hidden layers. A neural network classifier can include neurons with a hyperbolic tangent activation function, which can also include a scaling factor (a scaled hyperbolic tangent function). Multiple neural networks can be considered an ensemble. In one example, the number of neural networks could be 1024, although as few as 16 can be used. In another example, the selected subset of neural networks could be 32, although other numbers can be used.

[0021] According to another aspect, a system is provided, comprising one or more computers and one or more storage devices storing instructions, which, when executed by the one or more computers, cause the one or more computers to perform the operations of the corresponding training method described above.

[0022] According to another aspect, one or more computer storage media are provided that store instructions, which, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding training method described above.

[0023] According to another aspect, a neural-based system is provided, comprising: a memory configured to store a plurality of neural networks and keys associated with each respective neural network; wherein each of the plurality of neural networks is configured to process an encoding of a data item to generate output data indicating a classification of aspects of the data item. The system further includes one or more computers and one or more storage devices storing instructions, which, when executed by the one or more computers, cause the one or more computers to perform operations including: receiving a data item; processing the data item using an encoder to generate an encoding of the data item; selecting a subset of neural networks from the memory based on similarity between the encoding and the keys associated with each respective neural network; processing the encoding through each respective neural network in the selected subset of neural networks to generate intermediate data for each respective neural network; aggregating the intermediate data for each respective neural network to generate output data, wherein the aggregation includes weighting the intermediate data for each respective neural network based on similarity between the encoding and the keys associated with the respective neural network; and outputting output data; wherein the output data indicates a classification of aspects of the data item.

[0024] It should be understood that features described in the context of one aspect can be combined with features described in the context of another aspect.

[0025] In some of the described examples, the data items include images, but generally any type of data item can be processed. Examples of different types of data items are described later. This method can be used to train neural network-based systems to perform any type of task involving processing the same type of data items (e.g., images) used in training.

[0026] Where image data items, as used herein, include video data items, the task can include any kind of image processing or vision task, such as image classification or scene recognition tasks, image segmentation tasks (e.g., semantic segmentation tasks), object localization or detection tasks, and depth estimation tasks. When performing such tasks, the input can include pixels of an image or pixels derived from an image. For image classification or scene recognition tasks, the output can include a classification output that provides a score for each of a plurality of image or scene categories, such as an estimated probability that an input data item or an object or element within an input data item or an action within a video data item belongs to a category. For image segmentation tasks, for each pixel, the output can include an assigned segmentation category or the probability that the pixel belongs to a segmentation category (e.g., an object or action represented in an image or video). For object localization or detection tasks, the output can include data that defines the coordinates of bounding boxes or regions representing one or more objects in an image. For depth estimation tasks, for each pixel, the output can include an estimated depth value such that the output pixel defines a (3D) depth map of the image. Such tasks can also facilitate higher-level tasks, such as object tracking across video frames; or pose recognition, i.e., recognizing poses performed by entities depicted in a video.

[0027] Another example of an image processing task could be an image keypoint detection task, where the output includes the coordinates of one or more image keypoints (such as landmarks representing objects in an image), for example, a human pose estimation task where keypoints define the positions of body joints. Another example is an image similarity determination task, where the output could include values ​​representing the similarity between two images, for example, as part of an image search task.

[0028] The subject matter described in this specification can be implemented in specific embodiments to achieve one or more of the following advantages.

[0029] This specification describes a method for training neural network-based systems, which is particularly advantageous in continuous learning settings. In some prior art continuous learning methods, continuous learning of new tasks or situations with unstable data distributions requires knowledge of the task structure and task boundaries. In this training method, knowledge of the task and task boundaries is not required. Systems trained using this method can outperform current prior art systems by a significant margin on continuous learning benchmarks.

[0030] Specific training methods can also help mitigate "catastrophic forgetting," where performance degradation occurs on tasks learned earlier to benefit newer tasks. The availability of multiple neural networks and the selection of those networks based on the encoding of data items allow specific neural networks to be specialized for handling certain tasks / data.

[0031] This training method is also particularly well-suited for online learning settings, where each training data item is seen only once or a limited number of times. Therefore, compared to conventional techniques, this training method is computationally efficient and requires significantly less training time. Attached Figure Description

[0032] Figure 1 An example neural network training system is shown;

[0033] Figure 2 A schematic diagram of an example network training system is shown;

[0034] Figure 3 This is a flowchart illustrating the process used to train a neural network;

[0035] Figure 4 This is a flowchart illustrating the process used to generate output data;

[0036] Figure 5 Six graphs illustrating the performance of an exemplary neural network system using a continuous learning protocol over time are shown.

[0037] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation

[0038] Figure 1 An example neural network training system 100 for training multiple neural networks is shown. System 100 can be implemented as one or more computer programs on one or more computers in one or more locations.

[0039] System 100 is configured to receive training data item 105 and target data 110 associated with training data item 105. Target data 110 can provide a corresponding output expected by system 100 in response to receiving and processing training data item 105. Target data 110 can include data indicating the classification of aspects of training data item 105. For example, training data item 105 can include image data, and the associated target data 110 can indicate the type of object present in the image data, and / or target data 110 can indicate the location of the object in the image data via, for example, a bounding box or pixel-by-pixel marker, and / or target data 110 can indicate the pose of the object. Training data item 105 can be video data. Target data 110 can indicate an action performed in the video data. The temporal location of the action in the video can also be indicated in the target data 110. Training data item 105 can be audio data, such as an audio signal or waveform. Target data 110 can provide an indication of the words spoken in the audio data and / or provide an indication of the speaker's identity in the audio signal. The following provides further examples of the types of training data items.

[0040] System 100 can retrieve training data item 105 and target data 110 from a local data storage system, or it can receive training data item 105 and target data 110 from a remote system.

[0041] System 100 is configured to process training data item 105 using encoder 115 to generate an encoding 120 for training data item 105. Encoder 115 can be a pre-trained neural network with any suitable neural network architecture, such as a variational autoencoder or a ResNet architecture. Encoder 115 can be pre-trained to generate latent representations of the input data items. Encoder 115 can be pre-trained using a training dataset different from the dataset to which training data item 105 belongs. Encoder 115 can be pre-trained using self-supervised learning techniques, for example, training can be based on transformations or "augmented" views of the training data items. The pre-training of encoder 115 is described in further detail below. When used in system 100, the parameters of encoder 115 can remain fixed; that is, system 100 does not update the parameters of encoder 115.

[0042] System 100 also includes a memory 125 configured to store multiple neural networks. Each neural network stored in memory 125 is configured to process encoding 120 to generate output data indicating the classification of aspects of training data item 105. System 100 is configured to select a subset of neural networks 130 from the multiple neural networks stored in memory 125 based on encoding 120. For example, each of the multiple neural networks may be associated with a corresponding key, and the selection of a subset of neural networks 130 may be based on the corresponding key. The similarity (such as cosine distance) between encoding 120 and each corresponding key may be determined, and the selection of a subset of neural networks 130 may be based on the determined similarity. For example, neural networks associated with the k most similar keys of encoding 120 may be selected as a subset of neural networks 130. Thus, the corresponding keys and encoding 120 may reside in the same latent embedding space. The corresponding keys may be generated by sampling a probability distribution based on the embedding space represented by encoder 115, and the probability distribution may be determined based on samples of the encoded data. The keys may be fixed and not updated by system 100 after generation. The following describes further details regarding the selection of a subset of neural networks.

[0043] System 100 is configured to process code 120 using a selected subset of neural networks 130 to generate output data 135 corresponding to training data item 105. For example, code 120 can be processed by each corresponding neural network in the selected subset of neural networks 130 to generate intermediate data. The intermediate data can be in the same form as the output data, and the intermediate data can be equivalent to the output data of a corresponding neural network. For example, the intermediate data can be the initial classification provided by each corresponding neural network. The intermediate data can be aggregated to generate output data 135. Aggregation can include weighting the intermediate data of each corresponding neural network in the subset of neural networks 130. For example, aggregation can be a weighted sum, although it will be understood that other forms of aggregation can be used. Where the similarity between code 120 and the corresponding key associated with the neural network has been determined, the weighting can be based on the determined similarity. Further details regarding the generation of output data 135 are described below.

[0044] System 100 is configured to determine updates 145 to the parameters of a selected subset 130 of a neural network based on a loss function that includes the relationship between the generated output data 135 and the target data 110 associated with the training data item 105. As described in more detail below, the relationship between the generated output data 135 and the target data 110 in the loss function can be based on the dot product between the generated output data 135 and the target data 110. The target data 110 can be in the form of a one-hot vector, i.e., each element in the vector can represent a specific class, and the element corresponding to the target class of the associated training data item can be set to one, while all other elements are set to zero. The dot product with the one-hot target vector will only produce values ​​for the correct target class. In this way, it can be ensured that only connections and associated parameters in the neural network that contribute to the correct target class are updated. As described below, this helps mitigate catastrophic forgetting in online or continuous learning settings.

[0045] Update 145 can be determined by parameter update computation subsystem 140. Update 145 can be determined based on the gradient of each parameter to be updated relative to the loss function. Stochastic gradient descent or other suitable optimization techniques can be used to compute update 145 and the gradient. However, in one example, update 145 is determined based on the sign of the gradient, and the magnitude is discarded. The sign of the determined gradient can be applied with a fixed step size to provide the update value 145.

[0046] System 100 is configured to update the parameters of the selected subset of neural networks 130 based on the determined update 145. Once updated, the neural network can be written back to memory 125 if a copy of the selected subset of neural networks 130 is made instead of using the neural network directly from memory 125.

[0047] System 100 can be configured to process additional training data items from the training dataset. However, the training data items can be extracted from different data distributions. That is, a first training data item can be extracted from a first data distribution, and a second training data item can be extracted from a second data distribution. Different data distributions can represent different tasks to be performed. System 100 can be provided with training data items from one task at a time, where there are hard boundaries between tasks, or the tasks can change gradually over time and mix training data items from the first and second tasks in the transition. That is, the training dataset and multiple training data items can include training data items extracted from the first data distribution, which are scattered with training data items extracted from the second data distribution. This change can be achieved by extracting training data items associated with different data distributions / tasks / subtasks according to a specific probability distribution conditioned on training time. For example, training data items can be extracted from a first data distribution with a peak probability at time t1, and training data items can be extracted from a second data distribution with a peak probability at time t2, etc., where the temporal overlap in the probability distribution is used to select training data items from either the first or second data distribution. The probability distribution can be a Gaussian distribution.

[0048] System 100 can provide effective learning without needing to detect task boundaries or knowledge of task boundaries or task structure. System 100 can work effectively even when the data distribution is not fixed. In such continuous learning settings, the system may suffer from catastrophic forgetting, i.e., the performance of previously learned tasks may degrade when learning new tasks. The described technique can mitigate catastrophic forgetting. System 100 is also effective in online learning settings, where each training data item in the training dataset is presented to the training system 100 only once, or only a single passthrough of the training dataset is performed.

[0049] Now for reference Figure 2 This illustrates a potential embedding space 200, in which encoders (such as...) Figure 1 The encoder 115 maps the input data items to a latent embedding space 200. For visualization purposes, the latent embedding space 200 is depicted as two-dimensional; however, it should be understood that the latent embedding space 200 can have a larger number of dimensions. For example, in one embodiment, the latent embedding space has 512 dimensions. In another embodiment, the dimensions are 2048. The dimensions can be appropriately chosen based on the dimensions of the input data items and the complexity of the task to be learned.

[0050] As described above, each neural network in memory 125 can be associated with a corresponding key, where each key has a value in the latent embedding space 200. Figure 2In this diagram, each key is represented as a circle in the latent embedding space 200. Keys can be generated based on sampling a probability distribution over the latent embedding space 200. The probability distribution can be generated based on the encoding of the sample training data; for example, the sample encoding can be used to determine the parameters of a Gaussian distribution. In another example, a uniform probability distribution can be generated based on the range of possible values ​​in the latent embedding space. By generating keys through random sampling, the keys can be distributed across the latent embedding space 200. In this way, each of the multiple neural networks can cover a specific region of the latent embedding space 200 and utilize any class-specific clustering generated by the encoder 115. Each neural network can be encouraged to specialize based on the latent embedding space 200, and therefore each neural network can have lower complexity and fewer parameters than a network that covers the entire latent embedding space. In one example, each neural network is a single-layer classifier. More formally, each classifier has a trainable weight matrix. Where m is the number of output classes, and d is the 200-bit dimension of the latent embedding space. The classifier can also have a set of trainable biases. An encoder can be represented as Here, x is the input data item provided to the encoder, which implements the function f to provide the encoding z. Each classifier can take the encoding z as input, and the output of each classifier can be as follows:

[0051] v(W,b,z)=[φ(ψ1(z)),...,φ(ψ m (z))]

[0052] Where v is the w of W with row i. i Classifier ψ i (z)=w i ·z T +b i The output is φ(x) = τtanh(x / τ), and the scaling factor τ is a hyperparameter. In one example, τ is set to 250; however, it should be understood that other values ​​can be used appropriately. The scaling tanh function allows the neuron's output to grow to near τ without reaching τ.

[0053] Memory 125 can be considered to include n pairs of bonds and a neural network. For example, memory 125 can be represented as having (That is, n keys, each key having dimension d) and (That is, n pairs of corresponding classifier weights and biases) M = (M key M cfier ).

[0054] As described above, a subset of neural networks 130 can be selected based on the similarity between the encoding 120 of the input data items and the corresponding keys associated with the multiple neural networks stored in memory 125. In one example, the similarity is based on cosine distance; however, any other suitable distance metric, such as Euclidean distance, can be used. Figure 2 As shown, the potential embedding space 200 is determined according to the distance metric and the encoding 120 (in Figure 2 The k nearest neighbors (shown as crosses) are selected, and the neural networks associated with these keys are chosen as a subset of neural network 130. Figure 2 In this example, for illustrative purposes, k = 3. Other values ​​of k may be used as appropriate. In one example, with 1024 neural networks in memory 125, k is set to 32.

[0055] As described above, the code 120 can be processed by each neural network in the selected subset 130 of neural networks. The overall output data 135 can be generated based on the combination of the outputs of the individual neural networks in the selected subset 130. The output of each individual neural network is referred to herein as intermediate data, and as described above, it can take the same form as the overall output data 135. The intermediate data can be aggregated based on a weighted sum with weights having similarities based on the similarity between the code 120 and its corresponding key. For example, the output data 135 can be generated as follows:

[0056]

[0057] Among them, V M (z) represents the output data 135, and γ(x,y) is the distance metric, such as cosine distance. It is M key The index of the key based on the determined similarity ranking, and It corresponds to The classifier. Therefore, the neural network closest to encoding 120 will have the highest weight in the above formula. Neural networks can be thought of as working together as an integrated system and can be specialized and cooperative.

[0058] As described above, the update 145 of the parameters of the selected subset 130 of the neural network can be calculated based on a loss function. The loss function can be based on the dot product between the generated output data 135 and the target data 110 associated with the training data item 105. For example, the loss function can take the following form:

[0059]

[0060] Where y∈{0,1} mIt is the one-hot encoding of the class to which the input training data item x belongs, and The output data is the predicted label of x generated by the system. In this formula, there are no other normalization or softmax terms in the loss function, and the loss function depends only on the original dot product value.

[0061] As mentioned above, the update 145 of the parameters of a subset of the neural network can be determined based on minimizing the loss function, and the gradient of the parameters can be determined accordingly. However, unlike the conventional method, in one example, the update 145 is determined by applying the sign of the gradient to a fixed step size (e.g., the learning rate), and the magnitude of the gradient is not used. In one example, the learning rate is set to 0.0001.

[0062] As described above, encoder 115 can take any suitable form. In one example, encoder 115 is based on a variational autoencoder. For example, when the input is an image, the encoder part of the variational autoencoder may include two linear layers following two convolutional layers, and a corresponding decoder part with two transposed convolutional layers following the two linear layers. After training the variational autoencoder, the weights of the encoder part are frozen and can be used in system 100. The decoder part is not used and can be discarded. Further details about variational autoencoders can be found in Kingma and Welling's "Auto-encoder variational bayes", arXiv:1312.6114, which is available at https: / / arxiv.org / abs / 1312.6114, the entire contents of which are incorporated herein by reference.

[0063] In another example, the encoder is based on the ResNet architecture. Self-supervised learning techniques can be used to train the encoder; in particular, contrastive learning-based techniques such as ReLIC (“Representation learning via invariant causal mechanisms”, arXiv: 2010.07922) and BYOL (“Bootstrap you ownlatent”, arXiv: 2006.07733) can be used. In short, ReLIC is based on instance discrimination. For a batch of training images, each image can be transformed or “enhanced” to produce two views of the training image. Instance enhancements can include one or more of the following: cropping, rotation, scaling, shearing, flipping, color distortion, and modifications to brightness, contrast, saturation, and hue. Each view can be processed by the encoder to generate an encoding. For a specific training image, the encoding of the first view of that specific training image can be compared with the encoding of the second views of all images in the batch. The training task is instance discrimination, i.e., given the first view of a training image, determining which of the second views matches and which does not match the specific training image. This process is repeated for a second view of a specific training image compared to the first view of the images in the batch. Two probability distributions can be constructed based on comparisons between encodings, one for comparing the first and second views, and one for comparing the second and first views. The encoder parameters can then be tuned to minimize the error in an instance discrimination task that follows two probability distributions that maintain similarity. For example, a constraint can be used to keep the Kullback-Leibler divergence between the two probability distributions within a threshold. In this way, the encoder can be (pre)trained to generate useful representations (encodings) for other tasks. It should be understood that the above-described method for training the encoder can be applied to data items of other modalities, such as audio with appropriate transformations. Further details about the ResNet architecture can be found in He et al., “Deep residual learning for image recognition,” Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, the entire contents of which are incorporated herein by reference.

[0064] Encoder 115 can be pre-trained using a different dataset than that used in System 100. For example, encoder 115 can be pre-trained as a variational autoencoder on the Omniglot dataset, while System 100 can use the MNIST dataset. In another example, encoder 115 can be pre-trained on the ImageNet dataset using ReLIC with a ResNet-50 encoder, where System 100 uses the CIFAR-10 / 100 dataset.

[0065] Now for reference Figure 3 The processing for training neural network-based systems will now be described. This processing can use... Figure 1 The system 100 is used to achieve this.

[0066] At box 305, training data item 105 and target data 110 associated with training data item 105 are received. Target data 110 may be the corresponding output expected to be produced by the neural network system in response to receiving and processing training data item 105. As described above, training data item 105 may include image data, video data, audio data, data representing environmental states, etc. Target data 110 may indicate the classification of aspects of the training data item.

[0067] At box 310, encoder 115 is used to process training data item 105 to generate encoding 120 of training data item 105. Encoding 120 may be a latent representation of training data item 105, and encoder 115 may be pre-trained and kept fixed, as described above.

[0068] At box 315, a subset of neural networks 130 is selected from a plurality of neural networks stored in memory 125 based on encoding 120. As described above, the plurality of neural networks are configured to process encoding 120 to generate output data 135 indicating the classification of aspects of training data item 105. Each neural network may be associated with a key, and the selection of the subset of neural networks 130 may be based on the similarity between encoding 120 and each corresponding key, as discussed above.

[0069] At box 320, a subset of the selected neural networks 130 is used to process the encoding 120 to generate output data 135. Each neural network in the selected subset can generate intermediate data such as initial classification, and the intermediate data can be aggregated to generate output data 135, as discussed in further detail above.

[0070] At box 325, an update 145 for the parameters of the selected subset 130 of the neural network is determined based on a loss function that includes the relationship between the generated output data 135 and the target data 110 associated with the training data item 105. For example, the gradient of each parameter to be updated can be computed based on optimization of the loss function using stochastic gradient descent or other optimization methods.

[0071] At box 330, the parameters of the selected subset of neural networks 130 are updated based on the determined update 145. The update can be performed according to the selected optimization method. However, in one example, the sign of the computed gradient is applied to a fixed step size, and the magnitude of the computed gradient is not used.

[0072] This can be repeated for further training data items. Figure 3 This processing can be used for online learning, where each training data item is provided to the system only once, or a single pass is performed using only the training dataset. This processing can also be used in continuous learning settings, where the data distribution of the training data items can change over time, or where the system learns new tasks over time.

[0073] Now for reference Figure 4 This illustrates the process used to generate output data from input data items. Figure 4 The processing can be used Figure 1 The system 100 is used to achieve this.

[0074] At box 405, a data item is received. The data item can have the same form as the training data item 105.

[0075] At box 410, encoder 115 is used to process the data item to generate the encoding of the data item.

[0076] At box 415, a subset of neural networks from memory 125 is selected based on the similarity between the encoding and the key associated with each corresponding neural network. For example, the similarity could be based on cosine similarity. Further details regarding similarity and key generation have been described above in the context of the training data items, but are equally applicable here.

[0077] At box 420, the encoding is processed by each corresponding neural network in the selected subset of neural networks to generate intermediate data for each corresponding neural network. As mentioned above, the intermediate data can be the initial classification for each corresponding neural network in the selected subset.

[0078] At box 425, the intermediate data for each corresponding neural network is aggregated to generate the output data. The aggregation involves weighting the intermediate data for each corresponding neural network, and the weighting is based on the similarity between the encoding and the key associated with the corresponding neural network.

[0079] At box 430, output data is provided as output. As mentioned above, the output data indicates the classification of aspects of the data items.

[0080] Now for reference Figure 5 Six graphs are provided illustrating the classification accuracy of an exemplary neural network system trained over time using the techniques described above. The system is trained using the CIFAR-10 dataset according to a continuous learning protocol. More specifically, the ten classes to be identified are divided into five subsets, each containing two classes, to create five tasks. Each task is presented to the system one at a time. The goal of this type of continuous learning protocol is to maintain classification performance on earlier tasks while learning new tasks. Only one pass through the dataset is performed. Therefore, the system does not see data from previous tasks again.

[0081] Each plot shows the classification accuracy for each task over time as training progresses. Task 1 is presented to the system first, while Task 5 is presented last. Therefore, the accuracy for Task 5 is very low until the Task 5 data is presented to the system in the final stage of training. In Task 1, the accuracy is high at the beginning because it is the first task presented, but it decreases as new tasks are learned and Task 1 data is no longer seen. The final plot, labeled "All Tasks," shows the combined classification accuracy for all ten classes.

[0082] In each graph, line 501 corresponds to the performance of the exemplary system trained using the techniques described above. Line 502 corresponds to the performance of a system including a single classifier using only the hyperbolic tangent activation function, combined with the same encoder used in the exemplary system of line 501. Line 503 corresponds to the performance of a system including a single softmax classifier combined only with the same encoder. The encoders were pre-trained using the ImageNet dataset.

[0083] As can be seen from the figure, the exemplary system corresponding to line 501 is able to maintain the classification accuracy of earlier tasks while learning new tasks. For the hyperbolic tangent classifier 502 and the softmax classifier 503, performance degrades significantly, exhibiting signs of catastrophic forgetting. Therefore, the system trained according to the technique described herein can mitigate catastrophic forgetting.

[0084] Further examples will now be provided describing the inputs (data items or training data items) of System 100 and the types of tasks that System 100 can perform.

[0085] Neural network-based systems can be configured to receive any kind of digital data input (as data items) and generate any kind of score, classification, or regression output based on the input.

[0086] For example, if the input to a neural network-based system is an image or features already extracted from an image, the output generated by the neural network-based system for a given image can be a score for each of a set of object categories, where each score represents an estimated probability that the image contains an object belonging to that category.

[0087] As another example, if the input to a neural network-based system is a sequence representing spoken utterances, the output generated by the neural network could be a score for each of a set of text segments, each score representing an estimated probability that the text segment is a correct transcription of the utterance. As another example, if the input to a neural network-based system is a sequence representing spoken utterances, the output generated by the neural network-based system could indicate whether a particular word or phrase (“hot word”) was spoken in the utterance. As another example, if the input to a neural network-based system is a sequence representing spoken utterances, the output generated by the neural network-based system could identify the natural language of the spoken utterances. Therefore, typically, network input can include audio data used to perform an audio processing task, and network output can provide the result of the audio processing task, such as recognizing words or phrases or converting audio to text.

[0088] As another example, the task could be a health prediction task, where the input is a sequence derived from a patient’s electronic health record data, and the output is a prediction related to the patient’s future health, such as a predicted treatment that should be prescribed for the patient, the likelihood of the patient experiencing an adverse health event, or a predicted diagnosis for the patient.

[0089] As another example, if the input to a neural network-based system is a sequence of texts in one language, the output generated by the neural network can be a score for each of a set of text fragments in another language, where each score represents an estimated probability that the text fragment in the other language is a correct translation of the input text into the other language.

[0090] As another example, if the input to a neural network-based system is an internet resource (e.g., a webpage), a document, or a portion of a document, or features extracted from an internet resource, document, or portion of a document, then the output generated by the neural network for a given internet resource, document, or portion of a document can be a score for each of a set of topics, where each score represents an estimated probability of the internet resource, document, or portion of a document with respect to that topic.

[0091] As another example, if the input to a neural network-based system is features of the impression context of a particular ad, the output generated by the neural network could be a score representing the estimated probability that the particular ad will be clicked.

[0092] As another example, if the input to a neural network-based system is features for personalized recommendations to a user, such as features characterizing the context of the recommendation, or features characterizing the user's previous actions, then the output generated by the neural network could be a score for each of the set of content items, where each score represents an estimated probability that the user will respond favorably to the recommended content item.

[0093] For a system of one or more computers configured to perform a specific operation or action, this means that the system has software, firmware, hardware, or a combination thereof installed thereon, which, in operation, cause the system to perform the operation or action. For one or more computer programs configured to perform a specific operation or action, this means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.

[0094] Embodiments of the subject matter and functional operation described in this specification may be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware (including the structures disclosed in this specification and their equivalents), or combinations thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier, for performing or controlling operations of data processing by the device's data processing. Alternatively or additionally, the program instructions may be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof. However, the computer storage medium is not a propagating signal.

[0095] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines used for processing data, including, for example, programmable processors, computers, or multiple processors or computers. Apparatus may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.

[0096] A computer program (which may also be referred to or described as a program, software, software application, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program may be stored as a portion of a file that holds other programs or data, for example, as one or more scripts stored in a markup language document, as a single file dedicated to the program in question, or as multiple harmonized files, for example, as a file storing one or more modules, subroutines, or code portions. A computer program can be deployed to execute on a single computer or on multiple computers located at a site or distributed across multiple sites and interconnected by a communication network.

[0097] As used herein, "engine" or "software engine" refers to a software-implemented input / output system that provides outputs distinct from its inputs. An engine can be a coded functional block, such as a library, platform, software development kit ("SDK"), or object. Each engine can be implemented on any suitable type of computing device, including one or more processors and computer-readable media, such as a server, mobile phone, tablet computer, laptop computer, music player, e-book reader, laptop or desktop computer, PDA, smartphone, or other fixed or portable device. Furthermore, two or more engines can be implemented on the same computing device or on different computing devices.

[0098] The processes and logic flows described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuits, and the device can also be implemented as dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). For example, the processes and logic flows can be executed by a graphics processing unit (GPU), and the device can also be implemented as a GPU.

[0099] Computers suitable for executing computer programs include, for example, those based on general-purpose or special-purpose microprocessors or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory or random access memory or both. The basic components of a computer are the central processing unit for executing or running instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from or transfer data to or both. However, a computer does not need to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, personal digital assistant (PDA), mobile audio or video player, game console, GPS receiver, or portable storage device such as a Universal Serial Bus (USB) flash drive, to name a few.

[0100] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.

[0101] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser on the user's client device in response to a request received from a web browser.

[0102] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer with a graphical user interface or web browser through which a user can interact with embodiments of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), such as the Internet.

[0103] A computing system may include clients and servers. Clients and servers are typically geographically separated and usually interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other.

[0104] While this specification contains numerous specific details of implementation, these should not be construed as limiting any invention or the scope of any possible claims, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, one or more features from a claimed combination may be removed from the combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.

[0105] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or to perform all the shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0106] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired result. As an example, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing may be advantageous.

Claims

1. A computer-implemented method for training a neural network-based system, the method comprising the following training steps: Receive training data items and target data associated with the training data items; The training data items are processed using an encoder to generate the encoding of the training data items in the latent embedding space of the encoder; Based on the encoding, a subset of neural networks is selected from a plurality of neural networks stored in memory; wherein the plurality of neural networks are configured to process the encoding to generate output data indicating the classification of aspects of the training data items; Use a selected subset of neural networks to process the encoding to generate output data; The update of parameters for a selected subset of the neural network is determined based on a loss function that depends on the relationship between the generated output data and the target data associated with the training data items; and Update the parameters of the selected subset of neural networks based on a deterministic update; Each of the plurality of neural networks is associated with a corresponding key, and each key has a value in the latent embedding space of the encoder; The method further includes: determining the similarity between the encoding and each corresponding key; wherein, a subset of the neural network is selected based on the determined similarity; and wherein... The training data item includes image data of the pixels of the image, and the target data and the output data indicate at least one of the following: the type of object present in the image data; the location of the object in the image data; the pose of the object in the image data; the segmentation category assigned to each pixel in the image data or the probability that each pixel in the image data belongs to a segmentation category; the estimated depth value of each pixel in the image data; the coordinates of one or more image keypoints in the image data; a similarity value representing the relationship between two images; a score for each of a set of object categories, where each score represents the estimated probability that the image data contains an image of an object belonging to that category; or The training data item includes video data of pixels in a video image, and the target data and the output data indicate at least one of the following: an action being performed in the video data; a detection result indicating whether an action is being performed in a segment of the video data; tracking information of objects in the video data; a pose performed by an entity depicted in the video data; or The training data item includes audio data of an audio signal, and the target data and the output data provide at least one of the following: an indication of the words spoken in the audio signal; an indication of the speaker's identity in the audio signal; a score for each of a set of text segments of the audio data, each score representing an estimated probability that the text segment is a correct transcription of a word spoken in the audio signal; the natural language type of the words spoken in the audio signal; text converted from the audio data; or The training data items comprise sequences derived from the patient's electronic health record data, and the target data and the output data indicate predictions related to the patient's future health; or The training data items comprise text sequences in one language, and the target data and the output data comprise scores for each of a set of text segments in another language, where each score represents an estimated probability that the text segment in the other language is a correct translation of the input text into the other language; or The training data items include features from internet resources, documents, or portions of documents, or features extracted from internet resources, documents, or portions of documents, and the target data and the output data include scores for each of a set of topics, where each score represents an estimated probability of an internet resource, document, or portion of a document with respect to that topic; or The training data items include features of the impression context of a specific advertisement, and the target data and the output data include scores representing the estimated probability that a specific advertisement will be clicked; or The training data items include features for personalized recommendations to users, and the target data and the output data include scores for each of the content items in the set, where each score represents an estimated probability that a user will respond favorably to the recommended content item.

2. The method according to claim 1, further comprising: The training steps are repeated for multiple training data items; wherein the multiple training data items include a first training data item extracted from a first data distribution and a second training data item extracted from a second data distribution, and wherein the first data distribution and the second data distribution are different.

3. The method according to claim 2, wherein, The plurality of training data items include training data items extracted from a first data distribution, and the training data items are distributed among training data items extracted from a second data distribution.

4. The method according to claim 1, wherein, The loss function, which depends on the generated output data and the target data for the training data items, includes the dot product between the generated output data and the target data.

5. The method according to claim 1, wherein, The target data is in the form of a one-hot vector.

6. The method according to claim 1, wherein, The encoder is pre-trained using a dataset different from the dataset to which the training data items belong.

7. The method according to claim 6, wherein, The encoder is pre-trained using self-supervised learning techniques.

8. The method according to claim 7, wherein, The self-supervised learning technique includes training based on transformed views of training data items.

9. The method according to claim 1, wherein, The parameters of the encoder remain constant.

10. The method according to claim 1, wherein, The encoder is based on a variational autoencoder.

11. The method according to claim 1, wherein, The encoder is based on the ResNet architecture.

12. The method according to claim 1, wherein, The similarity is based on the cosine distance between the encoding and the corresponding key.

13. The method according to any one of claims 1 to 12, wherein, The corresponding keys are generated by sampling the probability distribution based on the latent embedding space represented by the encoder.

14. The method according to claim 13, wherein, The probability distribution is determined based on samples of the encoded data.

15. The method according to claim 14, wherein, The samples of the encoded data include encoded data generated by processing data items using the encoder, wherein the data items are extracted from a dataset different from the dataset to which the training data items belong.

16. The method according to any one of claims 1 to 12, wherein, Processing the encoding of training data items using a selected subset of neural networks includes: processing the encoding through each corresponding neural network in the subset of neural networks to generate intermediate data for each corresponding neural network; and aggregating the intermediate data for each corresponding neural network to generate output data indicating the classification of aspects of the training data items.

17. The method according to claim 16, wherein, Aggregating the intermediate data includes weighting the intermediate data of each corresponding neural network in the neural network subset.

18. The method according to claim 17, wherein, The weighting is based on the similarity between the keys associated with the corresponding neural network and the encoding of the training data items.

19. The method according to any one of claims 1 to 12, wherein, The multiple neural networks are neural network classifiers.

20. The method according to claim 19, wherein, The neural network classifier is a single-layer classifier.

21. The method according to claim 19 or 20, wherein, The neural network classifier includes neurons with a hyperbolic tangent activation function.

22. A system comprising one or more computers and one or more storage devices storing instructions, the instructions, when executed by the one or more computers, causing the one or more computers to perform operations according to any one of claims 1-21.

23. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operation of the corresponding method according to any one of claims 1-21.

24. A neural network-based system, comprising: The memory is configured to store a plurality of neural networks and keys associated with each respective neural network; wherein each of the plurality of neural networks is configured to process the encoding of data items to generate output data indicating the classification of aspects of the data items; One or more computers and one or more storage devices storing instructions, the instructions causing the one or more computers to perform operations when executed by the one or more computers, the operations including: Receive data items; The data item is processed using an encoder to generate an encoding of the data item in the potential embedding space of the encoder; A subset of neural networks is selected from memory based on the similarity between the encoding and the keys associated with each corresponding neural network. The encoding is processed by each corresponding neural network in the selected subset of neural networks to generate intermediate data for each corresponding neural network; The intermediate data of each corresponding neural network is aggregated to generate output data, wherein the aggregation includes weighting the intermediate data of each corresponding neural network based on the similarity between the encoding and the key associated with the corresponding neural network; and Output the output data; wherein the output data indicates the classification of aspects of the data items; Each of the plurality of neural networks is associated with a corresponding key, and each key has a value in the latent embedding space of the encoder; The operation further includes: determining the similarity between the encoding and each corresponding key; wherein, a subset of the neural network is selected based on the determined similarity; and wherein... The training data item includes image data of the pixels of the image, and the target data and the output data indicate at least one of the following: the type of object present in the image data; the location of the object in the image data; the pose of the object in the image data; the segmentation category assigned to each pixel in the image data or the probability that each pixel in the image data belongs to a segmentation category; the estimated depth value of each pixel in the image data; the coordinates of one or more image keypoints in the image data; a similarity value representing the relationship between two images; a score for each of a set of object categories, where each score represents the estimated probability that the image data contains an image of an object belonging to that category; or The training data item includes video data of pixels in a video image, and the target data and the output data indicate at least one of the following: an action being performed in the video data; a detection result indicating whether an action is being performed in a segment of the video data; tracking information of objects in the video data; a pose performed by an entity depicted in the video data; or The training data item includes audio data of an audio signal, and the target data and the output data provide at least one of the following: an indication of the words spoken in the audio signal; an indication of the speaker's identity in the audio signal; a score for each of a set of text segments of the audio data, each score representing an estimated probability that the text segment is a correct transcription of a word spoken in the audio signal; the natural language type of the words spoken in the audio signal; text converted from the audio data; or The training data items comprise sequences derived from the patient's electronic health record data, and the target data and the output data indicate predictions related to the patient's future health; or The training data items comprise text sequences in one language, and the target data and the output data comprise scores for each of a set of text segments in another language, where each score represents an estimated probability that the text segment in the other language is a correct translation of the input text into the other language; or The training data items include features from internet resources, documents, or portions of documents, or features extracted from internet resources, documents, or portions of documents, and the target data and the output data include scores for each of a set of topics, where each score represents an estimated probability of an internet resource, document, or portion of a document with respect to that topic; or The training data items include features of the impression context of a specific advertisement, and the target data and the output data include scores representing the estimated probability that a specific advertisement will be clicked; or The training data items include features for personalized recommendations to users, and the target data and the output data include scores for each of the content items in the set, where each score represents an estimated probability that a user will respond favorably to the recommended content item.

Citation Information

Patent Citations

  • Mixture of experts neural networks

    CN109923558A