Training machine-trained models by directly specifying gradient elements

By using gradient objective information to create machine-trained parameter values outside of traditional loss function derivation, this technique addresses the inaccuracies of traditional loss functions, resulting in efficient and accurate model creation with reduced errors.

JP2025515554APending Publication Date: 2025-05-20MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024555336
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-10
Filing Date
2023-02-19
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

Traditional loss functions in machine learning do not accurately model training objectives, leading to poorly trained models that can cause artifacts in synthetic data, misrecognition in control systems, and other errors.

Method used

A computer-implemented technique that receives instances of gradient objective information to create sets of machine-trained parameter values without deriving them from individual loss functions, allowing for the measurement of performance and identification of optimal parameter sets.

Benefits of technology

This technique rapidly creates accurate machine-trained models, reducing the time and effort required for developers and minimizing computational resources, while also reducing errors in application systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025515554000001_ABST
    Figure 2025515554000001_ABST
Patent Text Reader

Abstract

A computer-implemented technique performs machine learning that avoids traditional designs of loss functions. The technique includes receiving a plurality of instances of gradient objective information, each of the plurality of instances including a particular combination of a plurality of gradient elements. The technique creates a plurality of sets of machine-trained parameter values ​​using the plurality of respective instances of gradient objective information. The technique does this based on the plurality of instances of gradient objective information provided without using a loss function to calculate the plurality of instances of gradient objective information. The technique then measures the performance of the plurality of sets of machine-trained parameter values ​​on an application system. Based on the measured performance, the technique provides output information that identifies a particular set of machine-trained parameter values ​​that meets a specified test.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] background Traditional approaches to building machine-trained models start with designing a differentiable loss function that defines a test to determine the accuracy of the machine-trained model's output. The training system trains the machine-trained model by iteratively (1) using the machine-trained model to map training examples to output results, (2) measuring the error of the output results using the loss function, and (3) using backpropagation and stochastic gradient descent to adjust the weights of the machine-trained model based on the determined error. Many training systems use a hinge-based triplet loss as the loss function. Triplet loss attempts to place mismatched data item pairs close together while placing mismatched data item pairs far apart.

[0002] However, traditional loss functions do not always accurately model the training objectives that developers are trying to achieve. This problem may be due in part to the difficulty developers have in understanding all the aspects of complex training objectives and / or in expressing the training objectives in a mathematical form. A poorly selected loss function may degrade the performance of any machine-trained model that a training system creates based on that loss function. For example, in some cases, a machine-trained model is applied to the task of generating synthetic image or audio items. A poorly trained machine-trained model may cause artifacts in the image or audio items. In other cases, a machine-trained model is integrated into a control system. A poorly trained machine-trained model may cause the control system to take inappropriate actions. For example, a flawed face detection model may misrecognize a person's face, allowing that person to inappropriately enter a secured facility. A flawed object detection model installed in a moving vehicle may fail to detect an obstacle in the vehicle's path, causing a collision. These are just exemplary problems that may result from selecting a flawed loss function in various application environments. Summary of the Invention

[0003] overview Described herein is a computer-implemented technique for performing machine learning that avoids the traditional design of loss functions. The technique includes receiving a plurality of instances of gradient objective information, each of which includes a particular combination of a plurality of gradient elements. The technique uses the plurality of respective instances of gradient objective information to create a plurality of sets of machine-trained parameter values. That is, the technique does this based on the plurality of instances of gradient objective information given, without deriving the plurality of instances of gradient objective information from individual loss functions by differentiation. The technique then measures the performance of the plurality of sets of machine-trained parameter values ​​in an application system. Based on the measured performance, the technique identifies a particular set of machine-trained parameter values ​​that meets a prescribed test.

[0004] In some implementations, each particular instance of the gradient objective information includes a first gradient element that is part of a gradient of a first loss function and a second gradient element that is part of a gradient of a second loss function, the second loss function being different from the first loss function.

[0005] In some implementations, the application system executes the application using the selected set of machine-trained parameter values. For example, the application may correspond to a search application that identifies text items that match an input image. In other cases, the search application identifies images that match the input text item.

[0006] Among its technical merits, the technique rapidly creates accurate machine-trained models. For example, the technique eliminates the need for developers to devise loss functions appropriate for the problem domain, which can be a very challenging task in some complex problem domains that are not easy to conceptualize using loss functions. The technique also provides a systematic method of experimentation with various training objectives, instead of the ad-hoc trial-and-error approach traditionally used to improve model accuracy. These advantages benefit developers by reducing the time and effort required to create accurate machine-trained models. These advantages also lead to a training process that uses fewer computational resources compared to ad-hoc trial-and-error approaches.

[0007] The accuracy of the machine-trained model produced by the present technique also contributes to reducing errors occurring in application systems that use the machine-trained model. In some cases, the reduced errors are manifested in the presentation of inaccurate search results, irrelevant digital advertisements, or inaccurate bot responses, etc. In other cases, the reduced errors appear as artifacts in the corrected or synthetic images. In other cases, the reduced errors take the form of noise in audio output information. In other cases, the reduced errors correspond to the misrecognition of faces or other objects. When such errors occur, they can lead to inappropriate control actions, such as the failure to detect objects in the path of a moving vehicle or the erroneous entry of unauthorized persons into restricted areas.

[0008] The techniques summarized above may be manifested in various types of systems, apparatus, components, methods, computer-readable storage media, data structures, graphical user interface presentations, articles of manufacture, and the like.

[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. [Brief description of the drawings]

[0010] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] 1 illustrates an example training system for training a machine-trained model. [Diagram 2] 2 illustrates exemplary similarity information used in the training system of FIG. 1; [Diagram 3] 2 illustrates an example gradient processing component that is an element of the training system of FIG. 1. [Figure 4] 2 illustrates additional functionality used by the training system of FIG. 1; [Diagram 5] 2 illustrates an example application system that uses a machine-trained model created by the training system of FIG. 1 . [Figure 6] 2 illustrates a first model architecture that can be used in the training system of FIG. 1. [Figure 7] 2 illustrates a second model architecture that can be used in the training system of FIG. 1. [Figure 8] Three different gradient components are characterized, each of which depends on similarity information based on the three input items of a triplet. [Figure 9] Five gradient elements are characterized, with each gradient element depending on similarity information based on two input items of a triplet. [Figure 10] 1 is a table showing the relationship between a loss function and a gradient element. [Figure 11] 2 is a table showing the performance of different machine-trained models produced by the training system of FIG. 1 . [Figure 12] 2 is a flow chart outlining one aspect of the operation of the training system of FIG. 1; [Figure 13] 2 is a flowchart showing further details regarding the operation of the training system of FIG. 1. [Figure 14] 2 illustrates a computing device that can be used to implement the training system and any application system of FIG. 1. [Figure 15]1 illustrates an exemplary type of computing system that can be used to implement any aspect of the features illustrated in the previous figures. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] The same numbers are used throughout this disclosure and the drawings to reference like components and features, with the 100 series numbers referring to features first seen in Figure 1, the 200 series numbers referring to features first seen in Figure 2, the 300 series numbers referring to features first seen in Figure 3, and so on.

[0012] Detailed Description This disclosure is organized as follows: Section A provides an overview of a training system for training a machine-trained model. Section A also describes an application system that uses the machine-trained model. Section B describes an example method that illustrates the operation of the system of Section A. Section C describes example computational functionality that can be used to implement any aspect of the features described in Sections A and B.

[0013] A. Exemplary Systems A.1. Overview FIG. 1 illustrates a training system 102 for creating a machine-trained model. The machine-trained model performs a prescribed operation configured by machine-trained parameter values ​​created by the training system 102. In some implementations, the training system 102 creates a machine-trained model for use in a visual semantic embedding (VSE) system. The VSE system identifies relationships between images and text items. The VSE system can leverage these relationships to retrieve images that match a specified text item or retrieve text items that match a specified image. As used herein, a "text item" refers to a natural language expression having one or more tokens (e.g., words).

[0014] The training system 102 may also be applied to create other types of models. For example, some implementations of the training system 102 may create machine-trained models that determine relationships between pairs of images. Other implementations of the training system 102 may create machine-trained models that determine relationships between pairs of text items. In still other cases, the training system 102 may create machine-trained models that perform object detection tasks, face recognition tasks, image or audio correction tasks, speech recognition tasks, image synthesis tasks, etc. However, for ease of explanation, the following description refers primarily to exemplary cases in which the machine-trained models created by the training system 102 determine and exploit relationships between images and text items.

[0015] In overview, the training system 102 includes a training component 104, a neural network 106, and a gradient processing component 108. The training component 104 includes functionality for managing the training of the machine-trained model. The neural network 106 represents the logic that makes the operation of the machine-trained model subject to the machine-trained parameter values ​​being trained. Here, the neural network 106 includes a first neural network 110 that maps input images to image vectors x in a shared embedding space. A second neural network 112 maps text items to text vectors y in the same shared embedding space. In general, any vector in this embedding space represents the semantics of the input item using a distributed set of values. The vectors in the embedding space can also be said to represent the features of the corresponding input item (e.g., image or text item). Once the training system 102 has completed training, vectors in the embedding space that are close to each other correspond to semantically related items, while vectors in the embedding space that are far from each other correspond to semantically unrelated items.

[0016] The gradient processing component 108 receives an instance of gradient objective information from the training component 104. The gradient objective information refers to the counterpart of the gradient of a loss function traditionally produced by the loss function component 114. The loss function thus provides the logic for determining the degree to which an image vector x matches a text vector y. The gradient of a loss function can generally be expressed as (∂loss / ∂x, ∂loss / ∂y) and represents the effect that changes in vectors x and y have on the loss function (loss).

[0017] However, in this case, the training system 102 does not use the loss function component 114, and therefore the training system 102 does not involve computing the gradient of the loss function. Figure 1 illustrates this point by eliminating the loss function component 114. Instead of using the loss function component 114, the gradient processing component 108 receives the gradient objective information directly from the training component 104. The gradient processing component 108 generates instantiated gradient information based on the received gradient objective information. This operation includes substituting the similarity information calculated based on the particular vectors x and y into the placeholder variables of the gradient objective information.

[0018] The gradient selection component 116 provides an instance of the gradient objective information that is sent to the gradient processing component 108. The gradient selection component 116 selects the gradient elements (G e1 ,G e2 ,G e3 ,...,G n ), we construct an instance of the gradient objective based on the gradient elements. The gradient elements thus correspond to portions of the gradient created by differentiating the loss function. For example, a first loss function loss with a gradient that includes multiple terms a , for example G a1 , G a2 , G a3 and a second loss function loss with a gradient that includes multiple terms b , for example G b1 , G b2 , G b3Each of these terms in the gradient constitutes a gradient element. In this simplified case, the gradient element G a1 and G b2 The gradient selection component 116 can construct an instance of gradient objective information that combines one or more gradient elements drawn from two loss functions, such as a combination of G and G. Note, however, that the data store 118 may also include developer-created gradient elements that do not arise from the gradient of an existing loss function. Additionally, the data store 118 may also include gradient elements that are modifications of, and not direct copies of, terms found in the gradient of a loss function. For example, the data store 118 may include the actual gradient term G created by differentiating the loss function. a The gradient element G that represents the correction of a’ may include.

[0019] In some implementations, the gradient selection component 116 can combine multiple gradient elements by forming a product of the multiple gradient elements. In other implementations, the gradient selection component 116 can combine the gradient elements by forming a sum of the gradient elements. Other implementations can combine the gradient elements in other ways, such as by forming a weighted sum of the gradient elements.

[0020] From a more general perspective, the training component 104 creates multiple sets of machine-trained parameter values ​​based on multiple respective instances of gradient objective information. For example, in the simplified example above, G a1 and G b1 A first set of machine-trained parameter values ​​for the combinations of a1 and G b2 A second set of machine-trained parameter values ​​for the combinations of a1 and G b3The training component 104 may create nine sets of machine-trained values ​​based on different permutations of the gradient factors, such as a third set of machine-trained parameter values ​​for the combination of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 108, 109, 101, 109, 108, 109,

[0021] The backpropagation management component 120 and the optimization component 122 cooperate to create each set of machine trained parameter values ​​based on various instances of gradient objective information. The training component 104 can perform this task in a parallel or serial manner. The training component 104 calculates the product of two gradient elements (G e1 G e2 Suppose we generate a first set of machine-trained values ​​for a first instance of gradient objective information that includes

[0022] The backpropagation management component 120 performs training based on training examples drawn from a data store 124. The data store 124 includes, among other things, a set of images 126 and a set of text items 128. The data store 124 also stores information indicating established relationships between the images 126 and the text items 128. For example, for a given image, the data store 124 may store information indicating that a particular text item matches the image and another text item does not match the image. This relationship information may be created in any manner, such as based on labels applied by a human annotator, based on the juxtaposition of images and descriptive labels in a document, etc.

[0023] Each training example includes at least a triplet having a particular anchor item, a positive item, and a negative item. In some cases, the anchor item corresponds to an image item, and the positive item corresponds to a text item that is predetermined to match the image. The negative item corresponds to a text item that is predetermined to not match the image. In other cases, the anchor item corresponds to a text item, and the positive item corresponds to an image that is predetermined to match the text item. The negative item corresponds to an image that is predetermined to not match the text item. For ease of explanation, the following description focuses on training performed on a first type of triplet, for example, where an image serves as the anchor item. However, all principles described below apply equally to cases where a text item serves as the anchor item.

[0024] In practice, the training system 102 may train on a batch X of images and a batch Y of text items, which the set of neural networks maps to a set of image vectors and text vectors, respectively. The gradient processing component 108 then mines one or more triplets from these vectors in an online manner, i.e., without pre-establishing these triplets. In some implementations, online mining may specifically include selecting "hard" negative items (e.g., text items) with respect to a given anchor item (e.g., an image). Each hard negative item is an item that is known not to match the corresponding anchor item, but is closer to the anchor in the embedding space than the corresponding positive item. Hard negative items are particularly "hard" insofar as they present the training system 102 with non-trivial training examples, allowing the training system 102 to learn parameter values ​​more effectively. However, for simplicity, the following description takes an agnostic view of when and where triplets are established. The triplets may be pre-established by some offline process, or may be defined during training in an online manner. This further means that references to "training examples" encompass cases where corresponding triplets are established in the course of a training operation as well as cases where triplets are established in advance.

[0025] The training component 104 trains itself in successive forward and backward passes. Figure 1 shows a particular forward pass 130 and a particular backward pass 132. In the forward pass 130, the neural network 106 maps image items to image vectors x in the embedding space, maps matching text items to vectors y in the embedding space, and maps non-matching text items to vectors y' in the embedding space. The gradient processing component 108 then calculates a score S that represents the semantic distance (e.g., the degree of semantic match) between the image items and the matching text items. x,y and a score S representing the semantic distance between the image item and the mismatched text item. x,y’Again, in other cases the roles of x and y are reversed. The gradient processing component 108 then computes at least the score S x,y and S x,y’ Based on,the gradient objective information is used to calculate the instantiated gradient,information.

[0026] 2 illustrates the relationship between an image vector 202, a positive text vector 204, and a negative text vector 206 in the embedding space. Assume that the image vector 202 represents features of an image that serves as an anchor item, the positive text vector 204 represents features of text items that match the image, and the negative text vector 206 represents features of text items that do not match the image. In other cases, the gradient processing component 108 may construct instantiated gradient information based on additional vectors, such as other positive vectors 208 and / or other negative vectors 210, that have a relationship with the image vector 202.

[0027] Returning to FIG. 1, in the backward path 132, the backpropagation management component 120 uses the chain rule to backpropagate the instantiated gradient information backwards through the levels of the neural network 106, level by level. That is, for a given level of the neural network 106, the backpropagation management component 120 maps the input gradient information to output gradient information. That output gradient information serves as the input for the next level of the neural network. When this backpropagation operation is complete, it creates final gradient information that describes how changes in the machine-trained parameter values ​​of the neural network 106 affect changes in the image and text vectors. The optimization component 122 may use the final gradient information to adjust the parameter values ​​using some type of optimization algorithm, such as stochastic gradient descent. Note that the backward path 132 follows the same path as the forward path 130 through the neural network 106. FIG. 1 shows the backward path 132 next to the components of the neural network 106 simply for ease of illustration.

[0028] Expressed mathematically, let θ denote the parameter values ​​of the first neural network 110 used to create the image vectors, and φ denote the parameter values ​​of the second neural network 112 that creates the text vectors. Let f denote the mapping function applied by the first neural network 110. θ (·) and the mapping function applied by the second neural network 112 is g θ Let (·) denote the image vector created for a batch of images X. batch Let y be the text vector created for a batch of text items Y. batch Let x = y ≠ 0, which we simplify as x and y in the following two equations: Conventionally, training systems create loss information using a loss function component 114 in the forward pass based on a loss function L(·). loss = L(x,y) (where x = f θ (X) and y=g φ (1) (Y)

[0029] The weights of the neural network 106 are then updated using the following formula:

number

[0030] The first set of derivative terms (∂loss / ∂x and ∂loss / ∂y) represent how varying the embedding features (x,y) of the image and text items affects the loss. The second set of derivative terms (∂x / ∂θ and ∂y / ∂φ) represent how varying the parameters of the model affects the embedding features. The symbol η represents a learning rate that governs how quickly the training run converges to an optimal set of parameter values. As mentioned above, in this case, the training system 102 omits the loss function component 114 that would traditionally compute the first set of derivative terms. Instead, the gradient processing component 108 receives gradient objective information directly from the gradient selection component 116.

[0031] The training system 102 addresses challenges that arise in developing accurate machine-trained models in many problem domains, including the visual semantic embedding domain. These challenges stem from the difficulty for developers to conceptualize the loss function up front as a concise formula that expresses the training objective in many problem domains, and the additional constraint that the loss function be differentiable. A flawed loss function can result in the creation of a machine-trained model that provides substandard performance results, as measured, for example, by accuracy and / or other performance metrics. Traditionally, developers could address this situation by modifying the loss function and retraining the machine-trained model based on the new loss function. However, this is an unstructured, ad-hoc approach that can consume significant amounts of time and effort, and can lead to the consumption of a correspondingly large amount of computational resources. Computational resources include processor-related resources, memory-related resources, power, etc.

[0032] In part, the training system 102 addresses the above problems based on the insight that training goals in some problem domains can often be more accurately expressed by directly specifying gradient objective information rather than a loss function. The training system 102 also provides a structured way to try different combinations of gradient elements. The training system 102 ultimately provides output information that reveals the combination of gradient elements that yields the most favorable results (e.g., the most accurate results). This reduces the time and effort required for a developer to create a final machine-trained model, and also reduces the computational resources used to create the final machine-trained model.

[0033] Referring to the first neural network 110, the input encoding component 134 converts image items into one or more encoding vectors. For example, the input encoding component 134 can perform this task using a preliminary neural network that converts images into feature information. This preliminary neural network can include one or more layers of convolutional layers that perform convolution operations. Alternatively, the input encoding component 134 can divide the image into sub-images (patches) and then map the sub-images to individual encoding vectors. This mapping function can be implemented as a trainable linear projection.

[0034] The image processing component 136 maps the encoding vector to output feature information. Different implementations can implement the image processing component 136 using different neural network architectures. For example, the image processing component 136 can be implemented as a ResNet model, an example of which is described in He, et al., “Deep Residual Learning for Image Recognition,” arXiv:1512.03385v1 [cs.CV], December 10, 2015, page 12. Or, the first neural network 110 can be implemented as a transformer-based vision model, an example of which is described in Dosovitskiy, et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” arXiv:2010.11929v2 [cs.CV], June 3, 2021, page 22. Further details regarding one implementation of the image processing component 136 are provided below with reference to FIG.

[0035] The mapping component 138 maps the feature information produced by the image processing component 136 to an image vector x. The mapping component 138 may use a fully connected neural network with any number of layers to perform this task. The goal of the mapping component 138 is to convert the output feature information into a form that is directly comparable to the text vector y produced by the second neural network 112, for example, by creating an image vector x that has the same dimensions as the text vector y.

[0036] The second neural network 112 includes an input encoding component 140 that maps text items to one or more encoding vectors. The input encoding component 140 accomplishes this task by splitting the text items into one or more tokens. As used herein, a "token" or "text token" refers to a unit of text having any granularity, such as an individual word, a word fragment created by byte pair encoding (BPE), a character n-gram, a word fragment identified by the WordPiece algorithm, etc. The input encoding component 140 can use a machine-trained linear transformation to map the tokens to individual encoding vectors.

[0037] The text processing component 142 maps the encoding vector to output feature information. Various implementations can implement the text processing component 142 using various neural network architectures. For example, the text processing component 142 can be implemented as a transformer-based model, an example of which is described in Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv:1810.04805v2 [cs.CL] May 24, 2019, p. 16. Alternatively, the second neural network 112 can be implemented using a convolutional neural network, such as the network described in commonly assigned U.S. Published Patent Application No. 20150278200 to He et al., published October 1, 2015, and entitled “Convolutional Latent Semantic Models and their Applications.” Further details regarding one implementation of the text processing component 142 are provided below with reference to FIG. 7.

[0038] The mapping component 144 maps the feature information produced by the text processing component 142 to a text vector y. The mapping component 144 may use a fully connected neural network with any number of layers to perform this task. The goal of the mapping component 144 is to convert the output feature information into a form that is directly comparable to the image vector x produced by the first neural network 110, among others. Finally, it is noted that the neural network 106 may normalize the vectors x and y using, for example, the L2 (Euclidean) norm.

[0039] In some environments, the training system 102 of Figure 1 creates a machine-trained model that subsequently functions as a general-purpose pre-trained model. Another training system (not shown) performs further training on the pre-trained model to adapt the model to perform a specific application task. Prior to this supplemental training, a developer can add one or more neural network layers "on top" of the machine-trained model that are configured to perform further processing of the output results produced by the machine-trained model.

[0040] Proceeding to Figure 3, this figure illustrates one implementation of the gradient processing component 108. The similarity assessment component 302 estimates the similarity S between vector x and vector y. x,y and the similarity S between vector x and vector y' x,y’ where vector x represents an anchor item (e.g., an image). In the case where vector y represents an anchor item, the similarity assessment component 302 computes a corresponding similarity score S y,x and S y,x’ The similarity assessment component 302 may calculate each similarity score using any distance metric, such as cosine similarity, Manhattan distance, etc. The output of the similarity assessment component 302 may be more generally referred to as similarity information.

[0041] More specifically, in some cases, the triplets (x, y, y') are defined in advance, e.g., in an offline manner. In other cases, an optional online triplet mining component 304 performs the additional task of creating triplets based on batches of image and text vectors and based on the similarity information generated by the similarity generation component 302, e.g., by selecting negative text items that have a defined "hard" relationship with the image. Figure 3 should be broadly construed to encompass at least these two cases.

[0042] The gradient instantiation component 306 creates an instance of the instantiated gradient information based on the similarity information provided by the similarity assessment component 302 and the gradient objective information provided by the gradient selection component 116. The gradient instantiation component 306 performs this operation by substituting the similarity information into the placeholder variables of the gradient objective information.

[0043] FIG. 4 illustrates another aspect of the training system 102 of FIG. 1 that is not captured in FIG. 1. In this figure, an application system 402 uses a collection of machine trained models 404 created by the training component 104 to perform an application task, such as a search task, which is described in more detail below with respect to FIG. 5. A performance measurement component 406 measures the performance of each machine trained model using any performance metric or combination of performance metrics. For example, the performance measurement component 406 can use a Recall@1 metric to measure the retrieval accuracy of a search application. The Recall@1 metric measures whether the top-ranked target items identified by the application system 402 accurately match the input query. Additionally or alternatively, the performance measurement component 406 can measure the amount of computational resources consumed by the application system 402 in processing a user's query.

[0044] The performance measurement component 406 can provide output information that allows a user to determine the relative merits of various machine-trained models. The performance measurement component 406 can also identify one or more machine-trained models that provide the most favorable results, for example, by annotating the models that provide the most accurate results. Based on this guidance, a developer can install the most favorable models into a production-stage version of the application system 402 that is available to end users.

[0045] 5 illustrates an example of an application system 502 performing a search application. A query processing component 504 uses a trained machine trained model 506 created by the training system 102 to map queries to query vectors in the shared embedding space. For example, if the query is an image, the query processing component 504 uses a trained version of the first neural network 110 to map the image to an image vector in the embedding space. If the query is a text item, the query processing component 504 can use a trained version of the second neural network 112 to map the text item to a text vector in the embedding space.

[0046] The data store 508 stores feature information for multiple items of interest. For example, if the query is an image, the data store 408 includes entries with text vectors for each candidate text item. An offline feature generation system (not shown) can use a trained version of the second neural network 112 to provide these text vectors. If the query is a text item, the data store 508 includes entries with image vectors for each candidate image. The offline feature generation system can use a trained version of the first neural network 110 to provide these image vectors. In other cases, the query processing component 504 can compute vectors of items of interest in a dynamic manner in response to the submission of a user query.

[0047] The query processing component 504 can retrieve one or more target items having a target item vector that is closest to the query vector in the shared embedding space. The query processing component 504 can evaluate the similarity in any manner, for example using a cosine similarity metric. The query processing component 504 can also use any algorithm to speed up the search for matching target items, for example any type of approximate nearest neighbor (ANN) search algorithm. Although not shown, the query processing component 504 can perform additional mapping analysis using other machine trained models. In other words, the above-described matching can be part of a more comprehensive matching process.

[0048] Another application system may use the trained model created by the training system 102 to detect objects in an image. In some cases, the application system may incorporate a control system that takes appropriate action based on the detected objects. For example, the application system may use the trained model to detect objects in the path of the vehicle based on video captured by an on-board video camera. The application system may take a control action based on the detection, such as by applying the brakes of the vehicle.

[0049] Another application system may recognize human faces using the trained model provided by the training system 102. In some cases, the application system may incorporate a control system that takes action based on facial recognition or failure to recognize a face. For example, the application system may control a lock or gate to allow or deny access to a restricted premises based on the output results provided by the trained model.

[0050] Another application system may use the trained model to convert an input image into an output image, for example, by enhancing details in the input image, removing red eye or glare, etc. Another application system may use the trained model to synthesize an image based on input information. Another application system may use the trained model to recognize speech, etc. The above application systems are provided by way of example and not by way of limitation.

[0051] Whatever form the application system takes, the training system 102 contributes to reducing errors in output results based on the application system. In some cases, the errors are manifested in inaccurate search results, or irrelevant digital ads, or delivery of inaccurate bot responses, etc. In other cases, the errors appear as artifacts in corrected or synthetic images. In other cases, the errors take the form of noise in audio output information. In other cases, the errors correspond to misrecognition of faces or other objects. In addition to erroneous output results, these types of errors can lead to inappropriate control actions, such as failure to detect objects in the path of a moving vehicle or erroneous admission of unauthorized persons into restricted areas.

[0052] More specifically, as explained above, it is difficult for developers to formulate loss functions that represent what constitutes good and bad model outputs in many application domains. The training system 102 addresses this challenge by saving the task of pre-developing a comprehensive loss function and providing a structured way to try different combinations of gradient elements. Because the set of selected gradient elements is tailored to the application domain, the application system can perform the assigned task more accurately, thereby reducing errors in its outputs and the control actions taken based on those outputs (compared to models created in a traditional way that use hand-crafted loss functions).

[0053] Furthermore, application systems that use models trained by training system 102 utilize computational resources efficiently. For example, in the context of a search application, the application system may enable an end user to efficiently retrieve matching text items given an image, or matching images given a text item. These advantages result in a correspondingly more efficient use of computational resources by the application system. For example, an application system that quickly provides relevant answers to queries consumes less computational resources on average per task than an application system that requires users to engage in lengthy trial-and-error techniques to obtain information.

[0054] 6 illustrates a convolutional neural network (CNN) 602 that can be used to implement the first neural network 110 and / or the second neural network 112. The CNN 602 provides a pipeline that includes multiple convolutional blocks (e.g., encoder blocks 604, 606) optionally interspersed with pooling components (e.g., representative pooling component 608). FIG. 6 simply illustrates an exemplary case in which the encoding block 604 includes a pair of convolutional components (610, 612). FIG. 6 also illustrates a residual connection 614 that adds input information provided to the first convolutional component 610 to output information produced by the second convolutional component 612.

[0055] Each convolution component performs a convolution operation that involves moving an n×m kernel (e.g., a 3×3 kernel) over the feature information provided to the convolution component. For input images, the feature information represents image information. For input text items, the feature information represents text information. At each location of the kernel, the encoding subcomponent generates a dot product of the kernel value and the underlying value of the feature information. Each pooling component downsamples the result of the previous convolution operation using some sampling function, such as a max operation that selects the maximum value in a subset of values.

[0056] 7 illustrates a transformer-based neural network 702 that may be used to implement the first neural network 110 and / or the second neural network 112. The transformer-based neural network 702 provides a pipeline that includes multiple encoder blocks (e.g., encoder blocks 704, 706). FIG. 7 illustrates a representative architecture of the first encoder block 704. Although not shown, the other encoder blocks share the same architecture as the first encoder block 704.

[0057] The encoder block 704 includes, in order, an attention component 708, a summation normalization component 710, a feedforward neural network (FFN) 712, and a second summation normalization component 714. The attention component 708 performs a self-attention analysis using the following equation:

number

[0058] That is, the attention component 708 divides the input vector provided to the attention component 708 into three individual machine-trained matrices W Q , W K and W V, which creates the query information Q, the key information K, and the value information V. More specifically, the attention component 708 takes the dot product of the transpose of K and Q, and then divides the dot product by a scaling factor √d to create a scaled result. The symbol d represents the dimensionality of the transformer-based neural network 702. The attention component 718 takes the Softmax (a normalized exponential function) of the scaled result, and then multiplies the result of the Softmax operation by V to create the attention output information. More generally, the attention component 708 determines the importance of each input vector under consideration with respect to all other input vectors. Background information on the general concept of attention can be found in Vaswani, et al., “Attention Is All You Need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017, p. 11.

[0059] The additive normalization component 710 includes residual connections that combine (e.g., sum) the input information provided to the attention component 708 with the output information generated by the attention component 708. The additive normalization component 710 then performs a layer normalization operation on the output information generated by the residual connections, for example, by normalizing the values ​​in the output information based on the mean and standard deviation of those values. The other additive normalization component 714 performs the same function as the first-mentioned additive normalization component 710. The FFN 712 converts the input information into output information using a feedforward neural network with any number of layers.

[0060] A.2 Example Gradient Elements This subsection describes example gradient elements that the gradient selection component 116 can use to construct an instance of gradient objective information. There are many possible gradient elements that may be included in the data store 124. By way of example and not limitation, this subsection specifically describes examples of gradient components that result from gradients of known loss functions used in distance metric learning (DML) applications.

[0061] In the following description, S x,y =x T y and S x,y’ =x T y' represents the cosine similarity score calculated for the positive pair (x, y) and the negative pair (x, y') of the normalized encoding vectors, respectively. S y,x =y T x and S y,x’ =y T Let x' denote the cosine similarity scores calculated for the positive pair (y,x) and the negative pair (y,x') of the normalized encoding vectors, respectively. The conventional symmetric triplet loss function based on the similarity scores described above is loss = L(S x,y ,S x,y’ )+L(S y,x ,S y,x’ ) The gradient of this loss function is:

number

[0062] The above formula includes the first group, ∂L(S x,y ,S x,y’ ) / ∂S x,y , ∂L(S x,y ,S x,y’ ) / ∂S x,y’ , ∂L(S y,x ,S y,x’ ) / ∂S y,x and ∂L(S y,x ,S y,x’ ) / ∂S y,x’There are two types of gradient terms: the first group (x, y, x', and y') that represent the orientation in the embedding space, and the second group (x, y, x', and y') that represent the orientation in the embedding space. Different specific triplet loss functions vary primarily by the use of different logic used to compute the scalar gradient weights. Therefore, the following discussion focuses on the scalar component of the gradient elements.

[0063] More specifically, the following description first describes three exemplary gradient elements of a first type, each of which depends on all three terms of the triplet under consideration (e.g., {x, y, and y'} or {y, x, and x'}). This type of gradient element is called triplet-based gradient weights and is represented by the symbol T. The description then describes another five exemplary gradient elements of a second type, each of which specifies a relationship between two of the three terms of the triplet under consideration. This type of gradient element is called pair-based triplet weights and is represented by the symbol P. The following description more specifically describes the gradient weights when x is the anchor term. The corresponding gradient weights when y is the anchor term are derived therefrom, for example, by replacing y with x and x' with y'.

[0064] Constant Triplet Gradient Weights The gradient of the standard triplet loss function can be derived as follows:

number

[0065] In these equations, m is the margin parameter, δ(·) is the Heaviside step function, and H(x) is 1 if x>0 and 0 if x≦0.

[0066] In the gradient of the triplet loss function, each of the scalars is a triplet-based gradient weight because it depends on the similarity scores of both the positive and negative pairs of the triplet. The triplet weight T con is as follows: T con = δ(m+S x,y’ -S x,y ) (6)

[0067] When the Heaviside step function is activated, T con reduces to a constant 1. When the Heaviside step function is not activated, T con is 0, indicating that the triplet under consideration does not affect the gradient. con For example, S x,y S y,x Replace with S x,y’ S y,x’ By substituting , the corresponding expression in equation (6) is given. Background information on the general topic of triplet loss functions can be found in Scroff, et al., “FaceNet: A Unified Embedding for Face Recognition and Clustering,” arXiv:1503.03832v3 [cs.CV], June 17, 2015, p. 10.

[0068] NT-Xent gradient weights The second common loss function is loss nca The NT-Xent loss derived from neighborhood component analysis (NCA), denoted as: This loss function can be expressed as

number

[0069] In this equation, τ is a scaling parameter. nca Each scalar in is given by the following formula:

number

[0070] T nca is S x,y and S. x,y’If the difference is greater than zero (for example, S x,y -S x,y’ >0), T nca is relatively small. Otherwise, T nca is relatively large. Background information on the general topic of NT-Xent loss functions can be found in Sohn, Kihyuk, “Improved Deep Metric Learning with Multi-class N-pair Loss Objective,” in Proceedings of the 30th Conference on Neural Information Processing Systems (NIPS 2016), 2016, p. 9. Background information on the general topic of neighborhood component analysis can be found in Goldberger, et al., “Neighbourhood Components Analysis,” in Proceeding of Advances in Neural Information Processing Systems 17 (NIPS 2004), 2004, p. 8.

[0071] Circular Gradient Weights The gradient weight T mentioned above nca is the difference in similarity scores (S x,y -S x,y’ ) This relationship depends on S x,y and S x,y’ are relatively large, or S x,y and S x,y’ The loss function does not fully represent some special cases where both are relatively small. cir Use to address these cases.

number

[0072] T cir is, among other things, the S in the exponential term, which gives more weight to the special cases mentioned above. x,y and S x,y’We introduce a nonlinear mapping for . Background information on the general topic of circular loss functions can be found in Sun, et al., “Circle Loss: A Unified Perspective of Pair Similarity Optimization,” arXiv:2002.10857v2 [cs.CV], June 15, 2020, p. 10.

[0073] FIG. 8 shows three diagrams illustrating the properties of triplet-based weights described above. The horizontal axis of each diagram specifically represents the similarity scores of positive image-text pairs, and the vertical axis of each diagram represents the similarity scores of the corresponding negative image-text pairs. The darkness of the dots in the diagrams represents the magnitude of the gradient weights (0 to 1), so that, for example, relatively large gradient weights map to relatively dark dots and relatively small gradient weights map to relatively light dots. Each of these diagrams generally places triplets in which the anchor item features, the positive item features, and the negative item features are all very similar in the upper right corner of the diagram, and triplets in which the positive pairs are similar while the corresponding negative pairs are dissimilar in the lower right corner of the diagram.

[0074] More specifically, diagram (a) of FIG. 8 shows a triplet gradient weight T con , and Fig. 8(b) shows the behavior of the NT-Xent weight T nca , and Fig. 8(c) shows the behavior of the circular gradient weight T cir Figure 8(a) shows the behavior of the uniformly high T con value and uniformly low T con It shows a sharp equally weighted line that distinguishes the values. nca Figure 8(b) shows the high T nca Low T from value nca Although the transition from high to low values ​​includes a more gradual transition to low values, the transition from high to low values ​​generally runs linearly across the plot. cir How T nca Indicates whether to modify T ncaThe equal weighted boundary in x,y -S x,y’ =const. In contrast, T cir The equal weighted boundary at

number

[0075] Pair-based constant gradient weights Next, we move on to the description of the pair-based triplet weight P. We define the gradient weight for the positive pair as P + and the gradient weight for negative pairs is P - In pair-based constant weighting, the gradient weights of both positive and negative pairs are set to a constant of 1. In other words,

number

[0076] The training system 102 introduces pair-based constant weights primarily as a way to instantiate gradient objective information in which pair-based weights play no role: multiplying triplet-based weights by pair-based constant weights of 1 gives the original triplet-based weights.

[0077] Pair-based linear gradient weights If the vectors of the pair are close to each other in the embedding space, good training results are achieved by assigning a relatively large gradient weight to the negative pair. Otherwise, the training system 102 may rapidly converge to a false minimum. The circular loss function is lin We address this issue by applying the following: More specifically, for negative pairs, P lin is relatively large when the similarity between the component vectors of the negative pair is relatively large, and is relatively small when the similarity is relatively small. For the positive pair, P linis relatively large when the similarity between the component vectors is relatively small, and is relatively small when the similarity is relatively large. In other words, the following equation holds:

number

[0078] Binomial deviation gradient weights Binomial deviation gradient weight P sig represents the same kind of relationship as the pair-based linear weights, but incorporates the influence of the nonlinear sigmoid.

number

[0079] In this formula, α, β, and λ are the three hyperparameters. Background information on the general topic of binomial variance loss functions can be found in Yi, et al., “Deep Metric Learning for Practical Person Re-Identification,” in arXiv:1407.4979v1 [cs.CV], July 18, 2014, p. 11.

[0080] Multi-Similar Gradient Weights Multi-similarity (MS) gradient weight P sig_ms is the similarity score S of a positive pair and the corresponding negative pair. x,y and S x,y’ In the following, these scores are called self-similarity scores. The multi-similarity gradient weight P sig_ms also depends on other positive and negative pairs in the batch for the same anchor. We define each such other positive pairing as

number

number

number

number

number

number

number

number

number

number

number

[0081] The symbols α, β, λ, and ε are hyperparameters. The two terms in equation (14) are sig_ms We dynamically change the value of p. The first term (self-similarity term) depends on the self-similarity score, and the second term (relative similarity term) depends on the relative similarity score. The self-similarity term is a function of the sigmoid pair weights p sigIn particular, it has the same effect as having the same effect on the relative similarity terms as having the effect of increasing or decreasing the maximum magnitude of the pair weights.

[0082] More specifically, given a negative pair, the relative similarity term

number

number

number

number

[0083] Given a positive pair, the relative similarity term

number

number

number

number

[0084] lastly,

number

[0085] In practice, training using the MS loss involves tuning the four hyperparameters α, β, λ, and ε to fit different datasets. This operation makes training a machine-trained model complicated and inefficient. To address this issue, we use linear pair weights P lin_ms A linear version of the gradient weights mentioned above that shares the same general behavior can be used, called gradient weights = ∑ i ...

number

number

[0086] Figure 9 shows the behavior of the pair-based gradient weights discussed above. More specifically, Figure 9(a) shows the behavior of the pair-based constant weight gradient P conFigure 9(c) shows the behavior of the linear-based weight gradient P lin Figure 9(c) shows the behavior of the binomial deviation weight gradient P sig Figure 9(d) shows the behavior of

number

number

[0087] Again, the gradient elements mentioned above are provided by way of example and not limitation. In other implementations, the gradient elements can be derived from the following loss functions known in the technical literature: angle loss, arc face loss, base metric loss function, centroid triplet loss, contrast loss, Cos Face loss, cross batch memory, fast AP loss, generalized pair loss, lift structure loss and its variants, intra pair variance loss, large margin Softmax loss, margin loss, regularized Softmax loss, N pair loss, proxy anchor loss, proxy NCA loss, signal to noise ratio contrast loss, soft triple loss, sphere loss, subcenter arc face loss, supervised contrast learning, triplet margin loss, tuple margin loss, weight regularized mixture loss, VICReg loss, etc.

[0088] FIG. 10 shows nine different gradient element combinations that can be obtained from the triplet-based and pair-based gradient weights described above. In some cases, a particular combination corresponds to a set of gradient elements that are created by differentiating the loss function. For example, differentiating the circular loss function gives T cir Gradient weight and P linIn other cases, a particular combination may not have a direct counterpart in the derivation of any loss function. Such combinations have not yet been explored in the technical literature and are labeled "None" in FIG. 10. Although not shown, the training system 102 can create combinations of gradient objective information that combine three or more gradient elements. Additionally or alternatively, the training system 102 can combine gradient elements derived from the same class, for example, by combining two or more pair-based gradient elements.

[0089] FIG. 11 shows the accuracy of the nine different combinations shown in FIG. 10 when mapping images to text and when mapping text to images. nca (or T cir ) and P sig_ms Note that the combination of has superior performance compared to other combinations. This is an insight gained by the training system 102 that a developer would not be able to reach independently based on conceptual analysis of the problem domain or ad-hoc unstructured experimentation. Furthermore, in some cases, certain combinations may not be synthesizable, which would prevent a developer from discovering such combinations by differentiating the loss function. Thus, FIG. 10 demonstrates the usefulness of the training system 102 in creating satisfying machine-trained models in complex problem domains.

[0090] B. Exemplary Process 12 and 13 show a process illustrating the operation of the system of section A in flowchart form according to some implementations. Because the principles underlying the operation of the system have already been described in section A, this section highlights certain operations in a summarized manner. Each flowchart is depicted as a series of operations performed in a particular order. However, the order of these operations is merely representative and may be varied in other implementations. Furthermore, any two or more of the operations described below may be performed in parallel. In some implementations, the blocks shown in the flowcharts relating to processing-related functions are implemented by hardware logic circuitry described in section C, which may in turn be implemented by one or more hardware processors and / or other logic units that include a collection of task-specific logic gates.

[0091] FIG. 12 illustrates a process 1202 for performing machine learning. At block 1204, the training system 102 receives a plurality of instances of gradient objective information, each of the plurality of instances of gradient objective information including a particular combination of a plurality of gradient elements, the plurality of instances of gradient objective information including different respective combinations of gradient elements. At block 1206, the training system 102 creates a plurality of sets of machine-trained parameter values ​​using the plurality of respective instances of gradient objective information, the creating operation avoiding computing the plurality of instances of gradient objective information using a loss function. At block 1208, the training system 102 measures performance of the plurality of sets of machine-trained parameter values ​​on the application system. At block 1210, the training system 102 creates output information identifying a selected set of machine-trained parameter values ​​from the plurality of sets of machine-trained parameter values ​​based on the test results generated by the measuring operation. The selected set provides more effective performance than others of the plurality of sets of machine-trained parameter values ​​with respect to any criterion of performance (e.g., accuracy, latency, memory utilization, or any combination thereof). A selected parameter value set is created using a corresponding selected instance of gradient information having a selected combination of gradient elements. At block 1212, an application system, such as application system 502 of FIG. 5, performs an application task using the machine-trained parameter value set identified at block 1210. More specifically, the application system creates output results using the selected machine-trained parameter value set that have less error compared to output results created by other considered instances of gradient information having other corresponding combinations of gradient elements.

[0092] 13 illustrates a process 1302 for creating a first set of machine-trained parameter values ​​for a first instance of gradient objective information. In block 1304, the training system 102 uses the neural network 106 to map training examples including at least anchor terms, positive terms, and negative terms into an embedding space to create at least an anchor term vector, a positive term vector, and a negative term vector, respectively, where the positive terms match the anchor terms and the negative terms do not match the anchor terms. In block 1306, the training system 102 generates similarity information based on at least the anchor term vector, the positive term vector, and the negative term vector. The operations of blocks 1304 and 1306 should be interpreted to encompass cases where triplet relationships between the anchor terms, positive terms, and negative terms are established prior to the mapping and generation of the similarity information, as well as cases where the relationships are established after the mapping and generation of the similarity information in an online manner, for example, by mining a batch of vectors to find suitable negative vectors that meet specified criteria. At block 1308, the training system 102 creates instantiated gradient information based on the similarity information and the first instance of the gradient input information, where the act of creating the instantiated gradient information uses the first instance of the gradient objective information as received and avoids computing the first instance of the gradient objective information from a loss function. At block 1310, the training system 102 backpropagates the instantiated gradient information through the neural network 106 and performs optimization to create a model update and uses the model update to update the first set of machine trained values. Loop 1312 indicates that the training system 102 repeats the acts of mapping, generating, creating instantiated gradient information, and backpropagating multiple times for other training examples to create trained versions of the first set of machine trained parameter values.

[0093] C. Representative calculation functions 14 illustrates an example of a computing device that can be used to implement any of the systems summarized above. The computing device includes a set of user computing devices 1402 coupled to a set of servers 1404 via a computer network 1406. Each user computing device may correspond to any device that performs computing functions, including a desktop computing device, a laptop computing device, any type of handheld computing device (e.g., smartphone, tablet computing device, etc.), a mixed reality device, a wearable computing device, an Internet of Things (IoT) device, a gaming system, etc. The computer network 1406 may be implemented as a local area network, a wide area network (e.g., the Internet), one or more point-to-point links, or any combination thereof.

[0094] FIG. 14 also illustrates that the training system 102 and any application systems (e.g., application system 502 of FIG. 5) may be spread across user computing devices 1102 and / or servers 1404 in any manner. For example, in some cases, the application system 502 is implemented entirely by one or more servers 1104. Each user may interact with the servers 1404 by a user computing device. In other cases, the application system 502 is implemented entirely by a user computing device in a local manner, in which case no interaction with the servers 1404 is necessary. In other cases, functionality related to the application system 502 is distributed among the servers 1404 and each user computing device in any manner.

[0095] Figure 15 illustrates a computing system 1502 that can be used to implement any aspect of the mechanisms described in the previous figures. For example, the type of computing system 1502 illustrated in Figure 15 could be used to implement any of the user computing devices or any of the servers illustrated in Figure 14. In all cases, computing system 1502 represents a physical and tangible processing mechanism.

[0096] The computing system 1502 may include one or more hardware processors 1504. The hardware processors 1504 may include, without limitation, one or more central processing units (CPUs), and / or one or more graphic processing units (GPUs), and / or one or more application specific integrated circuits (ASICs), and / or one or more neural processing units (NPUs), etc. More generally, any hardware processor may correspond to a general purpose processing unit or an application specific processing unit.

[0097] The computing system 1502 may also include a computer readable storage medium 1506 corresponding to one or more computer readable media hardware units. The computer readable storage medium 1506 holds any type of information 1508, such as machine readable instructions, settings, data, etc. The computer readable storage medium 1506 may include, without limitation, one or more solid state devices, one or more magnetic hard disks, one or more optical disks, magnetic tape, etc. Any instance of the computer readable storage medium 1506 may use any technology for storing and retrieving information. Furthermore, any instance of the computer readable storage medium 1506 may represent a fixed or removable unit of the computing system 1502. Furthermore, any instance of the computer readable storage medium 1506 may provide volatile or non-volatile retention of information.

[0098] More generally, any of the storage resources described herein or any combination of storage resources may be considered a computer-readable medium. In many cases, a computer-readable medium represents some form of physical and tangible entity. The term computer-readable medium also encompasses, for example, propagating signals transmitted and received over physical conduits and / or the air or other wireless media. However, the specific term "computer-readable storage medium" explicitly excludes propagating signals in transmission, per se, while including all other forms of computer-readable media.

[0099] Computing system 1502 may utilize any instance of computer-readable storage medium 1506 in a variety of ways. For example, any instance of computer-readable storage medium 1506 may represent hardware storage (such as random access memory (RAM)) for storing information during execution of a program by computing system 1502 and / or hardware storage (such as a hard disk) for more permanently retaining / archiving information. In the latter case, computing system 1502 also includes one or more drive mechanisms 1510 (such as a hard drive mechanism) for storing and retrieving information from the instance of computer-readable storage medium 1506.

[0100] Computing system 1502 can perform any of the above functions when hardware processor 1504 executes computer readable instructions stored in any instance of computer readable storage medium 1506. For example, computing system 1502 can execute computer readable instructions for performing each block of the process described in Section B.

[0101] Alternatively or in addition, the computing system 1502 may rely on one or more other hardware logic units 1512 to perform operations using a collection of task-specific logic gates. For example, the hardware logic unit 1512 may include a fixed configuration of hardware logic gates that are created and configured, e.g., at the time of manufacture and cannot be changed thereafter. Alternatively or in addition, the other hardware logic unit 1512 may include a collection of programmable hardware logic gates that can be configured to perform various application-specific tasks. The latter class of devices includes, but is not limited to, programmable array logic devices (PALs), generic array logic devices (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), and the like.

[0102] FIG. 15 generally illustrates that the hardware logic circuitry 1514 includes any combination of a hardware processor 1504, a computer-readable storage medium 1506, and / or other hardware logic units 1512. That is, the computing system 1502 may employ any combination of a hardware processor 1504 executing machine-readable instructions provided in the computer-readable storage medium 1506 and / or one or more other hardware logic units 1512 that perform operations using a fixed and / or programmable collection of hardware logic gates. More generally, the hardware logic circuitry 1514 corresponds to one or more hardware logic units of any type that perform operations based on logic stored and / or embodied in the hardware logic units. Furthermore, in some circumstances, the terms "component," "module," "engine," "system," and "tool" each refer to a portion of the hardware logic circuitry 1514 that performs a particular function or combination of functions.

[0103] In some cases (e.g., cases in which the computing system 1502 represents a user computing device), the computing system 1502 also includes an input / output interface 1516 for receiving various inputs (via input devices 1518) and providing various outputs (via output devices 1520). Exemplary input devices include a keyboard device, a mouse input device, a touch screen input device, a digitizing pad, one or more still image cameras, one or more video cameras, one or more depth camera systems, one or more microphones, a voice recognition mechanism, any position determining device (e.g., a GPS device), any motion detection mechanism (e.g., an accelerometer, a gyroscope, etc.), etc. A particular output mechanism may include a display device 1522 and associated graphical user interface presentation (GUI) 1524. The display device 1522 may correspond to a liquid crystal display device, a light emitting diode display (LED) device, a cathode ray tube device, a projection mechanism, etc. Other output devices include a printer, one or more speakers, a haptic output mechanism, an archive mechanism (for storing output information), etc. The computing system 1502 may also include one or more network interfaces 1526 for exchanging data with other devices via one or more communication conduits 1528. One or more communication buses 1530 communicatively couple the above-mentioned units.

[0104] The communications conduit 1528 may be implemented in any manner, for example, by a local area computer network, a wide area computer network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communications conduit 1528 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., managed by any protocol or combination of protocols.

[0105] FIG. 15 illustrates a computing system 1502 that is composed of a discrete collection of separate units. In some cases, the collection of units corresponds to discrete hardware units provided in a chassis of a computing device having any form factor. FIG. 15 illustrates an example form factor in its bottom portion. In other cases, the computing system 1502 may include a hardware logic unit that integrates the functions of two or more of the units illustrated in FIG. 1. For example, the computing system 1502 may include a system on a chip (SoC or SOC) that corresponds to an integrated circuit that integrates the functions of two or more of the units illustrated in FIG. 15.

[0106] The following summary provides examples for a non-exhaustive set of descriptions of the technology described herein. (A1) According to a first aspect, some implementations of the technology described herein include a computer-implemented method (e.g., process 1202) for performing machine learning. The method includes receiving (e.g., at block 1204) a plurality of instances of gradient objective information, each of the plurality of instances of gradient objective information including a particular combination of a plurality of gradient elements, and the plurality of instances of gradient objective information including different respective combinations of gradient elements. The method further includes: creating (e.g., at block 1206) a plurality of sets of machine-trained parameter values ​​using a plurality of respective instances of gradient objective information, avoiding computing the plurality of instances of gradient objective information using a loss function; measuring (e.g., at block 1208) the performance of the plurality of sets of machine-trained parameter values ​​in an application system (e.g., application system 502); and creating (e.g., at block 1210) output information that identifies a selected machine-trained parameter value set from the plurality of sets of machine-trained parameter values ​​that most effectively meets a specified criterion of performance based on test results generated by the measurement, the selected parameter value set being created using a corresponding selected instance of gradient information having a selected combination of gradient elements. The application system creates output results using the selected machine-trained parameter value set, the output results having less error compared to output results created by other considered instances of gradient information having other corresponding combinations of gradient elements. The method is advantageous for providing a time- and resource-efficient method of developing machine-trained models. The methods may also enable an application system to produce output results that have a reduced number of errors compared to application systems having models that are trained using other techniques. (A2) According to some implementations of the method of A1, a particular instance of gradient objective information among the multiple instances of gradient objective information includes a first gradient element that is a portion of a gradient of a first loss function and a second gradient element that is a portion of a gradient of a second loss function, the second loss function being different from the first loss function. (A3) According to some implementations of any of the methods of A1 or A2, a particular instance of gradient objective information among the multiple instances of gradient objective information includes a first pair-based gradient element based on similarity information dependent on a comparison of two input items of a triplet, and a second triplet-based gradient element based on similarity information dependent on a comparison of three input items of the triplet. (A4) According to some implementation forms of any of the methods of A1 to A3, for a first instance of gradient objective information among the multiple instances of gradient objective information, the first set of machine-trained parameter values ​​includes: using a neural network to map training examples including at least anchor terms, positive terms, and negative terms into an embedding space to create at least an anchor term vector, a positive term vector, and a negative term vector, respectively; and generating similarity information based on at least the anchor term vector, the positive term vector, and the negative term vector, wherein the triplet relationships between the anchor terms, the positive terms, and the negative terms are determined prior to or after the operation of generating the similarity information. is established after the operation of generating similarity information; generating, creating instantiated gradient information based on the similarity information and the first instance of gradient input information, where the first instance of gradient objective information is used as received and avoids computing the first instance of gradient objective information from a loss function; back-propagating the instantiated gradient information through the neural network and optimizing to create a model update and use the model update to update the first set of machine-trained values; and repeating the operations of mapping, generating, creating instantiated gradient information, and back-propagating multiple times for other training examples. (A5) According to some implementations of the method of A4, the training examples include at least one image. (A6) According to some implementations of the method of A5, the training examples also include at least one text item. (A7) According to some implementation forms of any of the methods of A4 to A6, the neural network includes a first neural network for mapping anchor items to anchor item vectors, and a second neural network for mapping positive items and negative items to positive item vectors and negative item vectors, respectively. (A8) According to some implementations of the method of A7, the first neural network processes images and the second neural network processes text items. (A9) According to some implementations of any of the methods of A7 or A8, the first neural network and / or the second neural network have a transformer-based architecture. (A10) According to some implementations of any of the methods of A7 or A8, the first neural network and / or the second neural network have a convolutional neural network architecture. (A11) According to some implementations of any of the methods of A1 to A10, the application system is a search application that identifies items of interest that match an input query. (A12) According to some implementation forms of any of the methods of A1 to A11, the application system performs a control action based on an output result produced by a selected set of parameter values.

[0107] (B1) According to a second aspect, another implementation of the technology described herein includes a computing system (e.g., computing system 1502) having a computer-implemented application system (e.g., application system 502) that performs an application task based on a machine-trained model (e.g., trained model 506), where the machine-trained model uses a set of selected machine-trained parameter values ​​created by a computer-implemented training system (e.g., training system 102). The selected parameter value set is created by the training system using hardware logic provided by the training system (e.g., hardware logic 1514) by receiving (e.g., at block 1204) a plurality of instances of gradient objective information, each of the plurality of instances of gradient objective information including a particular combination of a plurality of gradient elements, the plurality of instances of gradient objective information including different respective combinations of gradient elements; creating (e.g., at block 1206) a plurality of sets of machine-trained parameter values ​​using the plurality of respective instances of gradient objective information, avoiding calculating the plurality of instances of gradient objective information using a loss function; measuring (e.g., at block 1208) performance of the plurality of sets of machine-trained parameter values ​​on the application system; and creating (e.g., at block 1210) output information that identifies a selected machine-trained parameter value set from the plurality of sets of machine-trained parameter values ​​that most effectively meets a specified criteria of performance based on test results generated by the measuring operation, the selected parameter value set being created using a corresponding selected instance of gradient information having a selected combination of gradient elements. The application system uses the selected set of machine-trained parameter values ​​to produce output results that have less error compared to output results produced by other considered instances of gradient information having other corresponding combinations of gradient elements.

[0108] In yet another aspect, some implementations of the techniques described herein include another computing system (e.g., computing system 1502) that includes hardware logic (e.g., hardware logic 1514) configured to perform any of the methods described herein (e.g., any of the methods A1-A12).

[0109] In yet another aspect, some implementations of the techniques described herein include a computer-readable storage medium (e.g., computer-readable storage medium 1506) for storing computer-readable instructions (e.g., instructions 1508). One or more hardware processors (e.g., hardware processor 1504) execute the computer-readable instructions to perform any of the methods described herein (e.g., any of methods A1-A12).

[0110] More generally, any of the individual elements and steps described herein may be combined without limitation in any logically consistent permutation or subset. Furthermore, any such combination may be manifested without limitation as a method, apparatus, system, computer-readable storage medium, data structure, article of manufacture, graphical user interface presentation, or the like. The technology may also be expressed in the claims as a series of means-plus-format elements, although this format should not be considered as invoked in the claims unless the "means for" step is explicitly used.

[0111] With respect to terms used herein, the phrase "configured to" encompasses various physical and tangible mechanisms for performing an identified operation. This mechanism may be configured to perform an operation using hardware logic circuitry 1514 of section C. The term "logic" similarly encompasses various physical and tangible mechanisms for performing a task. For example, each processing-related operation shown in the flowcharts of section B corresponds to a logical component for performing that operation.

[0112] In this description, one or more features may be identified as "optional" or other language may be used to indicate that one or more features may be used in some implementations but not in others. This type of description should not be construed as an exhaustive list of features that may be considered optional; other features may be considered optional even if not explicitly identified in the text. Furthermore, any description of a single entity is not intended to preclude the use of a plurality of such entities, and similarly, a description of a plurality of entities is not intended to preclude the use of a single entity. Furthermore, although this specification may describe certain features as alternative ways of performing an identified function or implementing an identified mechanism, those features may also be combined in any combination. Furthermore, the term "plurality" refers to two or more items and does not necessarily imply "all" items of a particular type unless expressly specified otherwise. Furthermore, descriptive terms such as "first," "second," "third," and the like are used to distinguish between different items and do not imply any ordering between the items unless otherwise specified. The phrase "A and / or B" means A or B, or A and B. Additionally, the terms "comprise," "include," and "have" are open-ended terms used to identify at least a portion of a larger whole, but not necessarily every part of a whole. Finally, the terms "exemplary" or "illustrative" refer to one implementation of potentially many implementations.

[0113] Finally, this specification may describe various concepts in the context of example problems or problems. This method of description is not intended to imply that others have understood and / or articulated those problems or problems in the manner set forth herein. Moreover, this method of description is not intended to suggest that the subject matter recited in the claims is limited to solving the problems or problems identified, i.e., the subject matter recited in the claims may be applied in connection with problems or problems other than those described herein.

[0114] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter set forth in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms for implementing the claims.

Claims

1. 1. A computer-implemented method for performing machine learning, comprising: receiving a plurality of instances of gradient objective information, each of the plurality of instances of gradient objective information including a particular combination of a plurality of gradient elements, the plurality of instances of gradient objective information including different respective combinations of gradient elements; generating a plurality of sets of machine-trained parameter values ​​using a plurality of respective instances of the gradient objective information, avoiding using a loss function to calculate the plurality of instances of the gradient objective information; measuring performance of the sets of machine-trained parameter values ​​on an application system; generating output information that identifies a selected machine-trained parameter value set from the plurality of machine-trained parameter value sets that most effectively meets a specified criterion of performance based on test results generated by the measurements, the selected parameter value set being generated using a corresponding selected instance of gradient information having a selected combination of gradient elements; wherein the application system generates output results using the selected set of machine-trained parameter values, the output results having less error as compared to output results generated by other considered instances of gradient information having other corresponding combinations of gradient elements.

2. A particular instance of gradient objective information among the plurality of instances of gradient objective information comprises: a first gradient element that is a portion of the gradient of the first loss function; a second gradient element that is a portion of the gradient of the second loss function; and wherein the second loss function is different from the first loss function.

3. A particular instance of gradient objective information among the plurality of instances of gradient objective information comprises: a first pair-based gradient element based on similarity information dependent on a comparison of two input items of a triplet; a second triplet-based gradient element based on similarity information dependent on a comparison of the three entries of the triplet; The computer-implemented method of claim 1 , comprising:

4. For a first instance of gradient objective information among the plurality of instances of gradient objective information, a first set of machine-trained parameter values ​​is Mapping training examples including at least anchor terms, positive terms, and negative terms into an embedding space using a neural network to generate at least an anchor term vector, a positive term vector, and a negative term vector, respectively; generating similarity information based on at least the anchor term vector, the positive term vector, and the negative term vector, where a triplet relationship between the anchor term, the positive term, and the negative term is established before or after generating the similarity information; creating instantiated gradient information based on the similarity information and a first instance of the gradient input information, wherein the first instance of the gradient objective information is used as received to create the instantiated gradient information and avoids computing the first instance of the gradient objective information from a loss function; backpropagating the instantiated gradient information through the neural network and performing optimization to generate a model update and to update the first set of machine-trained values ​​using the model update; repeating the mapping, generating, creating instantiated gradient information, and backpropagating multiple times for other training examples; The computer-implemented method of claim 1 , produced by:

5. The computer-implemented method of claim 4 , wherein the training examples include at least one image.

6. The computer-implemented method of claim 5 , wherein the training examples also include at least one text item.

7. 5. The computer-implemented method of claim 4, wherein the neural networks include a first neural network for mapping the anchor terms to the anchor term vectors, and a second neural network for mapping the positive terms and the negative terms to the positive term vectors and negative term vectors, respectively.

8. 8. The computer-implemented method of claim 7, wherein the first neural network processes images and the second neural network processes text items.

9. 8. The computer-implemented method of claim 7, wherein the first neural network and / or the second neural network have a transformer-based architecture.

10. 8. The computer-implemented method of claim 7, wherein the first neural network and / or the second neural network have a convolutional neural network architecture.

11. The computer-implemented method of claim 1 , wherein the application system is a search application that identifies items of interest that match an input query.

12. The computer-implemented method of claim 1 , wherein the application system performs a control action based on an output result produced by the selected set of parameter values.

13. A computing system configured to perform the method according to any one of claims 1 to 12.

14. A computer readable storage medium for storing computer readable instructions which, when executed by one or more hardware processors, perform the method of any one of claims 1 to 12.

15. 1. A computing system, comprising: a computer implemented application system for performing an application task based on a machine trained model, the machine trained model using a selected set of machine trained parameter values ​​created by a computer implemented training system; The selected set of parameter values ​​is generated by the training system using hardware logic provided by the training system: receiving a plurality of instances of gradient objective information, each of the plurality of instances of gradient objective information including a particular combination of a plurality of gradient elements, the plurality of instances of gradient objective information including different respective combinations of gradient elements; generating a plurality of sets of machine-trained parameter values ​​using a plurality of respective instances of the gradient objective information, avoiding using a loss function to calculate the plurality of instances of the gradient objective information; measuring performance of a plurality of sets of the machine-trained parameter values ​​on the application system; generating output information identifying the selected machine-trained parameter value set from the plurality of machine-trained parameter value sets that most effectively meets a specified criterion of performance based on test results generated by the measurements, the selected parameter value set being generated using a corresponding selected instance of gradient information having a selected combination of gradient elements; Created by A computing system, wherein the application system creates an output result using the selected set of machine-trained parameter values, the output result having less error compared to output results created by other considered instances of gradient information having other corresponding combinations of gradient elements.